oxedyne/daimond/hand/REVIEW.md
102 KiB, 1 run
created by r2519314175:895, which is this file's identity for as long as the history lasts, whatever it is later renamed to
download · who wrote it · its history
| 1 | # Adversarial review, 2026-08-02 |
| 2 | |
| 3 | The hand was written in a single session by five parallel agents plus the lead. |
| 4 | It was then reviewed by six independent adversarial passes, one per area, each |
| 5 | told to prove findings by making them happen rather than by reading. What |
| 6 | follows is what they found. |
| 7 | |
| 8 | **The short version: the hand is not close to shippable, and the two guarantees |
| 9 | the product would make about it — the compartment and the journal — currently do |
| 10 | not hold.** Three independent escapes from the fence were demonstrated against |
| 11 | this kernel, and a tampered journal was forged three ways. Nothing is exposed to |
| 12 | anyone, because `main.rs` has no message loop and the host cannot serve a |
| 13 | browser at all; but every claim in `README.md`'s release gates is further off |
| 14 | than it looked when they were written. |
| 15 | |
| 16 | Findings are CONFIRMED (the reviewer made it happen) or PLAUSIBLE (reasoned, not |
| 17 | reproduced). Line numbers are as of commit `ad19a62`. |
| 18 | |
| 19 | --- |
| 20 | |
| 21 | ## Where this stands |
| 22 | |
| 23 | **Everything above this line is the review as it was written on the morning of |
| 24 | 2026-08-02, and it is left exactly as written.** The paragraph beginning "The |
| 25 | short version" was true of the code that morning and is no longer true of the |
| 26 | code today — `main.rs` has a message loop, the host serves a browser, and most of |
| 27 | what was found has been repaired. It stays because a finding without its original |
| 28 | verdict is a finding with its teeth pulled, and because the reasoning is why the |
| 29 | code now looks as it does. |
| 30 | |
| 31 | **What was added afterwards is a state line on every finding**, marked **CLOSED** |
| 32 | (with what answers it and how that was proved) or **OPEN** (with what the |
| 33 | exposure is meanwhile). Nothing has been softened. Reproductions are left intact: |
| 34 | a closed finding with its reproduction still in it is the strongest thing this |
| 35 | document can carry, and an escape that no longer works is not a secret worth |
| 36 | keeping. |
| 37 | |
| 38 | The state lines were written by reading the working tree, driving the release |
| 39 | binary over a pipe where behaviour rather than code was the question, and running |
| 40 | the crate's tests. They are accurate as of the commit this file ships in and |
| 41 | nowhere else; a reader taking any of them on trust should check the code cited, |
| 42 | which is why each one cites code. |
| 43 | |
| 44 | | | Finding | State | |
| 45 | |---|---|---| |
| 46 | | 1.1 | Symlink in a carved directory grants its target | **closed** | |
| 47 | | 1.2 | Metadata syscalls ungoverned | **closed** | |
| 48 | | 1.3 | Unix socket is a way out of the compartment | **closed** | |
| 49 | | 1.4 | A fence root of `""` grants everything | **closed** | |
| 50 | | 1.5 | The received fence is not clamped to the grant | **closed** | |
| 51 | | 1.6 | Localhost dev origins in the manifest | **closed** | |
| 52 | | 1.7 | One global grant, not per-origin | **closed** | |
| 53 | | 1.8 | `apply_fence` a no-op returning success | **closed** | |
| 54 | | 2.1 | A truncated history verifies as intact | **closed** | |
| 55 | | 2.2 | One planted filename destroys the record | **closed** | |
| 56 | | 2.3 | Credentials through eight unredacted fields | **closed** | |
| 57 | | 2.4 | `redact_argv` caught 0 of 12 | **closed** | |
| 58 | | 2.5 | No locking; two hands break the chain | **closed** | |
| 59 | | 2.6 | A write that reaches nothing returns `Ok` | **closed** | |
| 60 | | 2.7 | The two verifiers disagree on CRLF | **closed** | |
| 61 | | 2.8 | Nothing calls the journal | **closed** | |
| 62 | | 3.1 | The shipping path can emit an unsendable frame | **closed** | |
| 63 | | 3.2 | Trailing content silently discarded | **closed** | |
| 64 | | 3.3 | An oversized frame poisons the stream | **closed** | |
| 65 | | 3.4 | `argv[0]` never vetted against the fence | **closed** | |
| 66 | | 3.5 | `LD_PRELOAD` settable by the caller | **closed** | |
| 67 | | 3.6 | A duplicate `id` makes a run unkillable | **closed** | |
| 68 | | 3.7 | One noisy run blocks every later one | **closed** | |
| 69 | | 3.8 | No total output cap | **closed** | |
| 70 | | 3.9 | Registry entries leak | **closed** | |
| 71 | | 3.10 | Group signalling degrades on BusyBox | **closed** | |
| 72 | | 3.11 | `u64` fields do not survive JavaScript | **closed** | |
| 73 | | 3.12 | The decoder accepts JDAT, not JSON | **closed** | |
| 74 | | 4.1 | `www/js/hand.js` is dead code | **closed** | |
| 75 | | 4.2 | A host can hold the page open for ever | **closed** | |
| 76 | | 4.3 | A quiet command killed at 30 seconds | **closed** | |
| 77 | | 4.4 | Chunks accumulate with no cap | **closed** | |
| 78 | | 1.9 | `no_write` never populated in the browser build | **open** — being closed elsewhere | |
| 79 | | 1.10 | `fence_spec` not a faithful restatement of `Bound` | **closed** | |
| 80 | | 1.11 | A killed command reported as `exit code: 0` | **closed** | |
| 81 | | 1.12 | `handler.rs` assigns rather than composes | **closed** | |
| 82 | | 1.13 | The first command of a turn costs the turn its network | **open** — a decision for the user | |
| 83 | | 1.14 | No check that the hand's root is the workspace | **closed** at both ends | |
| 84 | | 1.15 | Release gate 1 is nobody's job | **closed** | |
| 85 | | 1.16 | Every disconnect reported as "not installed" | **closed** | |
| 86 | | 1.17 | A lost output tail is detectable and not detected | **closed** | |
| 87 | | 1.18 | Scoped workers cannot run with a default `cwd` | **closed** | |
| 88 | | 1.19 | `id` neither unique nor bounded | **closed** | |
| 89 | | 1.20 | A Diamond's crystal agent could reach another Diamond | **closed** | |
| 90 | | 1.21 | A verifier's trustworthiness is reported, not enforced | **closed** — with one residual named for the owner | |
| 91 | | 3.13 | `screen_env` is past through `argv`, not through `env` | **closed by correction** — the claim, not the code | |
| 92 | |
| 93 | **One open, and it is not an escape.** 1.9 and 1.12 were one thing and both are |
| 94 | now closed: a worker is scoped by its own Diamond, and a second bound composes |
| 95 | with the first rather than replacing it. 1.13 is a design decision, not a |
| 96 | defect, and it is written up as a recommendation. |
| 97 | |
| 98 | **And one more, found afterwards, which is also not an escape.** 3.13 was found |
| 99 | while wiring the git toolkit and is written up in full at the end of Severity 3. |
| 100 | It does not widen the compartment: everything it reaches is inside the same |
| 101 | ruleset and the same filter. What it breaks is the neighbouring claim — that |
| 102 | what a command runs *with* is the user's decision — because `/usr/bin/env` and |
| 103 | `/bin/sh` are inside every fence and both take an environment out of their own |
| 104 | arguments, which no screen of the `env` field can see. The paragraph above about |
| 105 | the release gates was written before it and is left as written. |
| 106 | |
| 107 | **3.13 is now closed the other way round: the claim was corrected, not the code** |
| 108 | (2026-08-07). It is the only finding in this document resolved by changing what |
| 109 | `README.md` promises rather than what the hand does, and the reasoning is at the |
| 110 | end of its entry. Behaviour is unchanged and every measurement in the entry still |
| 111 | reproduces. |
| 112 | |
| 113 | **Both compartment escapes are closed, and each closed by a different layer.** |
| 114 | 1.1 was Landlock's own carve; 1.2 and 1.3 needed a mechanism Landlock does not |
| 115 | have, and the filter that provides it had been written, tested and never |
| 116 | installed — the most valuable thing this status pass found. 1.5 was the page |
| 117 | being believed about which folders it was allowed to name. |
| 118 | |
| 119 | **The release gates in `README.md` are close but not all met**, because gate 1 |
| 120 | now holds at both ends while 1.9 keeps the fence wider than the guide describes. |
| 121 | The consent window's wording can be revisited once 1.9 lands; on this machine the |
| 122 | sentence "only inside the folders the workspace already allows" now has the |
| 123 | exceptions `--report` prints and no others. |
| 124 | |
| 125 | --- |
| 126 | |
| 127 | ## Severity 1 — the compartment does not hold |
| 128 | |
| 129 | **1.1 A symlink in a carved directory grants its target.** CONFIRMED. |
| 130 | `fence.rs:1151-1157`. `carve()` enumerates a carved parent with `read_dir` and |
| 131 | grants each child without canonicalising it; `apply_linux` (`fence.rs:884`) then |
| 132 | opens it with `PathFd::new`, which follows symlinks and binds the rule to the |
| 133 | target inode. Spec-named paths go through `canonical()`; carved children never |
| 134 | do. |
| 135 | |
| 136 | This is not a corner case. `fence_spec` always adds `.daimond` to `deny` inside |
| 137 | the workspace `rw` root, so **every real fence carves the workspace**. A daimon |
| 138 | with ordinary write access to its own Diamond drops `ln -s /home/u ws/escape` |
| 139 | on one turn; on the next, the fence grants read-write on the whole home |
| 140 | directory. Deterministic, persistent, attacker-controlled. Proved twice: a file |
| 141 | was created in an outside directory and a private file there was read. |
| 142 | |
| 143 | **CLOSED.** `fence.rs` canonicalises every carved child and refuses one that is a |
| 144 | symbolic link, so the link is dropped from the plan rather than followed; the |
| 145 | drop is reported in `plan.dropped` and named in `plan.caveats()`, so a command |
| 146 | that stops working says why. Proved at both ends by |
| 147 | `a_symlink_in_a_carved_directory_is_not_granted` (the plan carries neither the |
| 148 | link nor its target) and `a_symlink_escape_is_refused_by_the_kernel` (the kernel |
| 149 | refuses the target under a real ruleset), because a plan that looks right and a |
| 150 | fence that is wrong is the failure this file is written against. |
| 151 | |
| 152 | **1.2 Metadata syscalls are ungoverned, including inside the denied subtree.** |
| 153 | CONFIRMED. Landlock's `AccessFs` has no right covering `chmod`, `chown`, |
| 154 | `utimensat` or `setxattr`, so none are mediated. Under a full ABI-8 fence, all |
| 155 | four succeeded on a file outside every root, and `chmod 777` succeeded on a file |
| 156 | *inside* the denied `.daimond` subtree — taking it from 600 to 777 — even though |
| 157 | reading it is refused. A fenced `cargo test` can world-write the home directory |
| 158 | or strip protection from the exact secrets the deny exists to protect. |
| 159 | `Fence::holes()` mentions none of this. |
| 160 | |
| 161 | **CLOSED — and it was open for a reason worth recording.** `seccomp.rs` |
| 162 | implemented the answer, had tests for both measured halves, and **was called from |
| 163 | nowhere**: `fence.rs`, `exec.rs`, `main.rs` and the launcher never referenced it, |
| 164 | and the crate's only other mention of `seccompiler` was in `Cargo.toml`. The |
| 165 | module's unit tests passed the entire time the launcher ran unfiltered, so a |
| 166 | passing unit test on the filter was precisely the evidence that failed. A defence |
| 167 | that exists in the source and not in the process is worse than none, because it |
| 168 | has already been counted as done. |
| 169 | |
| 170 | `launch_inner` now installs it, and the escape is the test. Against the release |
| 171 | binary, over a pipe, with the whole fence in force and the file **inside the |
| 172 | denied subtree**: |
| 173 | |
| 174 | ``` |
| 175 | before argv: /bin/chmod 777 <ws>/.daimond/secret.txt |
| 176 | {"t":"ended","id":"e","exit":0,"timed_out":false,"killed":false,…} |
| 177 | secret.txt mode: 0600 -> 0777 CHANGED |
| 178 | |
| 179 | after argv: /bin/chmod 777 <ws>/.daimond/secret.txt |
| 180 | {"t":"chunk","stream":"err","data":"chmod: Operation not permitted (os error 1)"} |
| 181 | {"t":"ended","id":"e","exit":1,…} |
| 182 | secret.txt mode: 0600 -> 0600 unchanged |
| 183 | ``` |
| 184 | |
| 185 | The trade `Meta::NoLoosening` makes is unchanged and is still in `holes()`: |
| 186 | `chmod 777` is refused, `chmod 644` is not, because `cargo` sets 644 on every |
| 187 | file it unpacks and the utime family has to stay for the same reason. |
| 188 | |
| 189 | Also a test now, at `exec::tests::the_filter_is_installed_and_not_merely_written` |
| 190 | — the `chmod` half, because it needs no session bus and can therefore run |
| 191 | anywhere. It drives the real spawn path and the real launcher, and it fails with |
| 192 | "chmod 777 succeeded behind the filter" the moment the install is removed. |
| 193 | |
| 194 | **1.3 The unix-socket hole is arbitrary unfenced execution, not a leak.** |
| 195 | CONFIRMED. With `net:false` fully applied, `connect()` to |
| 196 | `/run/user/1000/bus` succeeds, and `systemd-run --user … /bin/cat <denied-file>` |
| 197 | runs **outside the fence** and returns the contents. `SysBase::Minimal` supplies |
| 198 | the tools by granting `/usr` read-only. `holes()` and `README.md` describe this |
| 199 | as "can still talk to the session bus"; the true consequence is that the |
| 200 | compartment can be stepped out of entirely. The under-statement is itself the |
| 201 | defect. |
| 202 | |
| 203 | **CLOSED, with 1.2 — and closing it needed a second decision.** The filter |
| 204 | refuses `socket(AF_UNIX, …)`, which every `connect()` to the bus needs first. |
| 205 | Against the release binary, with the fence in force and the bus's address handed |
| 206 | to the command explicitly (the hand clears the environment, and the escape is |
| 207 | `connect()` to a path a compiled attacker carries in its own text, so supplying |
| 208 | it is what makes this a test of the fence rather than of `env_clear`): |
| 209 | |
| 210 | ``` |
| 211 | before argv: /usr/bin/systemd-run --user --quiet --pipe --wait /bin/cat <denied file> |
| 212 | {"t":"chunk","stream":"out","data":"the private thing\n"} |
| 213 | {"t":"ended","id":"e","exit":0,…} |
| 214 | |
| 215 | after {"t":"chunk","stream":"err","data":"Failed to connect to user scope bus |
| 216 | via local transport: Operation not permitted …"} |
| 217 | {"t":"ended","id":"e","exit":1,…} |
| 218 | ``` |
| 219 | |
| 220 | **The second decision.** `Spec::for_fence(net)` refused `AF_UNIX` only when the |
| 221 | fence refused the network, reasoning that refusing the bus buys nothing from a |
| 222 | command that may reach outward anyway. That reasoning is wrong and it was |
| 223 | measured wrong: wired in as written, with `net:true`, the filter installed and |
| 224 | the fence in force, `systemd-run` **still returned the file's contents**. And |
| 225 | `fence_spec` sets `net: !tainted`, so `net:true` is the *first command of every |
| 226 | turn* — the compartment would have been open on the first command of every |
| 227 | ordinary session. |
| 228 | |
| 229 | The escape is not a network escape. It is a filesystem escape wearing a socket: |
| 230 | the bus starts a process Landlock never bound, and that process reads a path this |
| 231 | fence denies. Whether the command was allowed to fetch a crate has nothing to do |
| 232 | with it, and the same socket reaches `ssh-agent`, which can sign with the user's |
| 233 | keys without the key ever being read. `fence.rs` already scopes *abstract* unix |
| 234 | sockets unconditionally from ABI 6, so refusing the pathname ones unconditionally |
| 235 | is what makes the two layers agree. |
| 236 | |
| 237 | So `Spec::for_command()` refuses `AF_UNIX` for every command, and the third line |
| 238 | of the table above is `net:true`, refused. The cost was measured rather than |
| 239 | assumed: a from-scratch `cargo build --offline` succeeds behind it, because |
| 240 | `cargo`, `rustc` and `ld` use `socketpair`, which is a different call and is left |
| 241 | alone. What a command cannot have is a local socket it names — a database, a |
| 242 | container daemon, X11, an `ssh-agent`-authenticated fetch. That is stated in |
| 243 | `Unix::Refuse` and printed by `--report`, and `Unix::Allow` survives as an arm |
| 244 | nothing chooses, so an operator setting has a shape to take. |
| 245 | |
| 246 | **1.4 A fence root of `""` grants the whole filesystem.** CONFIRMED. |
| 247 | `exec.rs:709-711,739-742`. The empty-fence guard tests `rw.is_empty()`, not |
| 248 | whether the roots mean anything, and `under()` is `Path::starts_with`, for which |
| 249 | `Path::new("/etc/ssh").starts_with("")` is true. `FenceSpec{rw:[""]}` ran a |
| 250 | command in `/etc/ssh` and returned success. |
| 251 | |
| 252 | **This one reaches the app.** `Tool::run` (`tools.rs`) reads the hand's granted |
| 253 | root with `extract_json_string(&st, "root")` and refuses only when the key is |
| 254 | *absent*; a root of `""` returns `Some("")`, passes the check, and `fence_spec` |
| 255 | pushes it into `rw`. The lead wrote both the guard and the hole. |
| 256 | |
| 257 | **CLOSED, at both ends.** In the hand, `under()` refuses an empty or relative |
| 258 | prefix outright rather than asking `starts_with` (`exec.rs`, proved by |
| 259 | `test_containment_is_by_component_not_by_prefix`, which asserts |
| 260 | `!under("/etc/ssh", "")`). In the app, `Machine` refuses an absent root and an |
| 261 | empty one alike, and `test_an_unusable_root_fences_nothing_rather_than_everything` |
| 262 | asserts that `""`, `relative/path`, `./ws` and `C:\ws` each yield *no* roots and |
| 263 | no network — failing closed rather than open. |
| 264 | |
| 265 | **1.5 The fence is computed inside the page, and the page is not trusted.** |
| 266 | CONFIRMED. `ext/hand.js:506-517,493` validates only `id`, `argv` and `cwd`; |
| 267 | `env`, `fence`, `timeout_ms` and `capture` are forwarded to the executing host |
| 268 | verbatim, and an exec with no `fence` key at all is forwarded too. A page sent |
| 269 | `fence:{rw:["/"],net:true}` with its own `LD_PRELOAD` and the host received it |
| 270 | byte-for-byte. |
| 271 | |
| 272 | This is sound against the *model* — the LLM cannot choose its own fence, because |
| 273 | `fence_spec` derives it from `ToolContext` — but not against a hostile origin. |
| 274 | Combined with 1.6 it is a full compromise. **The durable fix is that the host |
| 275 | must clamp any received fence to its own grant and refuse anything wider**; it |
| 276 | cannot know a Diamond's bounds, but it can refuse `rw:["/"]`. |
| 277 | |
| 278 | **CLOSED.** First re-measured, because a finding worth fixing is worth |
| 279 | reproducing: driven over a pipe, `fence:{rw:["/"]}` *was* refused, but only |
| 280 | incidentally — the journal lives somewhere under `/`, and `Journal::check_fence` |
| 281 | refuses any fence that reaches the record. Narrow the fence to a root that misses |
| 282 | the journal and it was honoured in full: `fence:{rw:["/etc"]}`, `cwd:"/etc"`, |
| 283 | `argv:["/bin/ls","/etc/ssh"]` started, ran and returned the directory listing, |
| 284 | with the granted root nowhere in it. A coincidence is not a boundary, and that |
| 285 | one stops holding the moment somebody moves the journal. |
| 286 | |
| 287 | `exec::vet_roots` now answers the question, and the question is the whole of the |
| 288 | fix. "Every root must be under the granted root" is the obvious rule and it is |
| 289 | wrong: a toolchain does not live in the workspace — `cargo` is under `~/.cargo`, |
| 290 | `node` under `~/.nvm` — so that rule refuses every real build, and a security |
| 291 | check that breaks `cargo` is one somebody switches off. So the clamp asks |
| 292 | whether a root is one **the grant could imply**, against a set that is closed and |
| 293 | knowable: the granted workspace, the hand's own scratch directory, and the |
| 294 | toolchain folders `Toolkit::grants` names in `src/tools.rs`. `deny` is not |
| 295 | clamped and must not be — a deny only ever takes access away. |
| 296 | |
| 297 | The two copies of that toolchain list can drift, and the drift fails safe and |
| 298 | loud: a path the hand does not know is refused in a sentence naming both the path |
| 299 | and the constant to add it to, so it is one line to fix rather than a hole to |
| 300 | find. |
| 301 | |
| 302 | Measured against the release binary, over a pipe, after the change: |
| 303 | |
| 304 | ``` |
| 305 | refused rw:[/etc] cwd:/etc ls /etc/ssh |
| 306 | refused rw:[/] cwd:/etc |
| 307 | refused ro:[~/.ssh] |
| 308 | refused rw:[~] the whole home directory |
| 309 | refused rw:[/tmp] |
| 310 | ran, 0 rw:[workspace] output: in the workspace |
| 311 | ran, 0 rw:[workspace/sub] a subtree of it |
| 312 | ran, 0 rust toolkit ro:[~/.cargo/bin ~/.rustup] output: cargo 1.90.0 |
| 313 | ran, 0 node toolkit ~/.nvm ~/.npm |
| 314 | ran, 0 python toolkit ~/.pyenv ~/.local ~/.cache/pip |
| 315 | ran, 0 go toolkit ~/sdk ~/go ~/.cache/go-build |
| 316 | ``` |
| 317 | |
| 318 | Only Rust is installed on the machine this was measured on; the other three |
| 319 | toolchain directories were created empty for the run and removed afterwards, |
| 320 | because without them the *planner* refuses the fence ("cannot be resolved") long |
| 321 | before the clamp is asked, and a pass for that reason would prove nothing. |
| 322 | |
| 323 | Also by `exec::tests::a_fence_may_only_name_roots_the_grant_implies`, which walks |
| 324 | every entry in `TOOLKIT_ROOTS` and needs no directory to exist, and which fails |
| 325 | three ways against broken code: with no clamp at all, with the naive |
| 326 | under-the-root clamp (which refuses `~/.cargo` and takes `cargo` with it), and |
| 327 | with `deny` wrongly clamped too. |
| 328 | |
| 329 | **1.6 The two localhost dev origins ship in the extension manifest.** CONFIRMED. |
| 330 | `ext/manifest.json:30-31`. A bare hostile HTML file served from |
| 331 | `http://127.0.0.1:8777` opened the port and completed an exec. Any process that |
| 332 | binds that port — a stray dev server, a static server rooted in `~/Downloads`, |
| 333 | another account on a shared machine, user-level malware — obtains content-script |
| 334 | injection and unfenced execution. DNS rebinding does not apply (Chrome matches |
| 335 | the origin string), so the requirement is genuinely "own that port". |
| 336 | |
| 337 | **CLOSED.** `ext/manifest.json` now lists exactly one origin in both |
| 338 | `externally_connectable` and `content_scripts`: `https://daimond.oxedyne.com/*`. |
| 339 | The two localhost entries are gone. |
| 340 | |
| 341 | **1.7 The machine-hand grant is one global boolean, not per-origin.** CONFIRMED. |
| 342 | `ext/hand.js:82,188-195`. Granted from `127.0.0.1:8777`, `localhost:8777` then |
| 343 | reached the host with no window shown at all. Compare `background.js:159-208`, |
| 344 | where site grants are real per-origin `chrome.permissions` patterns that Chrome |
| 345 | itself enforces: the new path is markedly laxer than the old one it was supposed |
| 346 | to match. |
| 347 | |
| 348 | **CLOSED.** The grant is a per-origin map (`ext/hand.js`, `{ '<origin>': { at, |
| 349 | caps } }`) rather than one boolean, and the window is shown per origin, so a |
| 350 | grant made from one origin is not a grant to another. |
| 351 | |
| 352 | **1.8 `apply_fence` is a no-op that returns success.** `exec.rs:695-700`, called |
| 353 | at `:255`. Gate 1 says a command that cannot be fenced must be *refused*; this |
| 354 | one runs. `fence.net` is never consulted anywhere in `exec.rs`, so the |
| 355 | tainted-turn network rule — the whole answer to prompt injection — is not |
| 356 | enforced at all. A no-op returning `Ok` is the fail-open shape the README warns |
| 357 | against; it should be a refusal or a compile-time gate. |
| 358 | |
| 359 | --- |
| 360 | |
| 361 | **CLOSED.** The fence is really applied. `exec.rs` re-executes the hand as a |
| 362 | launcher, applies the plan there, and becomes the command through a safe |
| 363 | `CommandExt::exec`; `Runner::spawn` plans with `Unfenced::Refuse` and refuses |
| 364 | where the plan is not fenced, and `Desk::exec` makes the same check before |
| 365 | anything is written down or run. `fence.net` reaches the plan. This is release |
| 366 | gate 4, and it landed. |
| 367 | |
| 368 | ## Severity 2 — the journal does not hold |
| 369 | |
| 370 | **2.1 An arbitrarily truncated history verifies as intact.** CONFIRMED. |
| 371 | `journal.rs:1831-1851`. `verify_dir` walks whatever files exist and never checks |
| 372 | that the newest is the newest. From a 29-file journal, deleting the last three |
| 373 | files gave `Intact { entries: 79 }`, and the hand then resumed appending and |
| 374 | *stayed* intact. Blanking the final file to zero bytes does the same |
| 375 | (`verify_file:1740` treats an empty file as intact). Nothing persists a |
| 376 | high-water seq or head, so nothing could detect it. The module doc claims only |
| 377 | the live file's tail is unprotected; in fact **any suffix of history erases |
| 378 | silently** — which is precisely the property a tamper-evident log exists to deny. |
| 379 | |
| 380 | **CLOSED.** The journal keeps a high-water mark of its own (`head.json`, written |
| 381 | alongside the chain), so the newest file and the newest entry are recorded rather |
| 382 | than inferred from whatever files happen to exist. Proved by three tests: deleting |
| 383 | whole files off the end no longer verifies as intact, a hand reopened on that |
| 384 | history refuses to carry on as though nothing were missing, and blanking the final |
| 385 | file to zero bytes is caught as the erasure it is. |
| 386 | |
| 387 | **2.2 One planted filename destroys the record, permanently and silently.** |
| 388 | CONFIRMED. `journal_files:1365-1385` accepts only 8-digit names, but |
| 389 | `maybe_rotate:1195` and `open:1074` increment past them. Planting |
| 390 | `hand-99999999.jsonl` — which any process running as the user can do, and which |
| 391 | until gate 4 lands means *every command the hand runs* — makes the hand write to |
| 392 | `hand-100000000.jsonl`, which `journal_files` never returns. Nothing is verified |
| 393 | and every later launch rotates into it again from a fresh chain. If the planted |
| 394 | file is empty instead, the hand appends a second chain starting at seq 0 and |
| 395 | `verify_dir` reports Broken forever, making real tampering indistinguishable |
| 396 | from the plant. A *directory* with that name makes both `open` and `verify_dir` |
| 397 | return `Err`, so under "journal before acting" the hand can never act again. |
| 398 | |
| 399 | **CLOSED.** The directory is read by one function (`survey`) that accounts for |
| 400 | every file it finds rather than returning only the ones it liked, so a planted |
| 401 | name is visible instead of invisible. All three shapes the finding names have |
| 402 | tests: the highest name the format allows must not push the hand into a file |
| 403 | nothing verifies, an empty plant must not fork a second chain from zero, and a |
| 404 | *directory* wearing a journal file's name must not stop the hand from ever acting |
| 405 | again. |
| 406 | |
| 407 | **2.3 Credential values reach the journal through eight unredacted fields.** |
| 408 | CONFIRMED. `Event::from_resp:574-598` applies no redaction whatever: |
| 409 | `Refused.reason` and `Error.message` are recorded verbatim, and a refusal that |
| 410 | quotes the offending command — the natural wording, and what the app's own |
| 411 | refusals do — writes the secret straight in. Also verbatim: `cwd`, `id`, env |
| 412 | *keys*, `fence.rw/ro/deny` paths, `mechs`, `Hello.client`. One secret reached |
| 413 | the file eight times in a single probe. |
| 414 | |
| 415 | **CLOSED.** Every free-text field now goes through the same scrubber before it is |
| 416 | written: `Refused.reason`, `Error.message`, `cwd`, `id`, environment keys, `mechs` |
| 417 | and `Hello.client`, and the fence's `rw`/`ro`/`deny` paths as well — a directory |
| 418 | can be named after a token. The count of redactions is recorded in the entry. |
| 419 | Tested through the constructors the message loop actually calls, not only through |
| 420 | the helpers. |
| 421 | |
| 422 | **2.4 `redact_argv` caught 0 of 12 real credential shapes.** CONFIRMED. |
| 423 | `journal.rs:754-802`. All twelve returned `cut=0`, including |
| 424 | `https://oauth2:ghp_…@github.com` (the prefix list uses `starts_with`, so a |
| 425 | token *inside* a URL is missed — the commonest way a token lands in argv), |
| 426 | `postgres://u:pw@db`, `--header=Authorization: Bearer …`, `--PASSWORD=x` and |
| 427 | `--Token x` (`SECRET_FLAGS` compares case-sensitively while the header check |
| 428 | lowercases), `--secret-access-key`, `--private-key`, `-phunter2`. Latent bug: |
| 429 | `&a[..h.len()]` at `:784,795` indexes the original string by the lowercased |
| 430 | length; no panic is reachable with today's constants, but one added constant |
| 431 | containing `k` or `s` makes it a slice panic. |
| 432 | |
| 433 | **CLOSED.** `redact_argv` was rewritten around a scrubber with separate passes for |
| 434 | credentials inside a URL, `--flag=value` pairs and known prefixes, case-folded |
| 435 | where the review found it case-sensitive. The twelve real credential shapes that |
| 436 | all came back `cut=0` are a test, and the latent slice panic — indexing the |
| 437 | original string by the lowercased length — is gone with the code that did it. |
| 438 | |
| 439 | **2.5 No locking; two hands permanently break the chain.** CONFIRMED. Nothing |
| 440 | takes a lock. Two `Journal::open` on one directory both resume at the same head, |
| 441 | and interleaved appends produced `Broken`. Chrome can launch more than one host, |
| 442 | and the design explicitly expects one per tab. |
| 443 | |
| 444 | **CLOSED.** The journal takes an exclusive `flock` through `File::try_lock` (safe |
| 445 | Rust, no dependency; it is why `Cargo.toml` names a minimum toolchain). A second |
| 446 | hand on the same directory is refused rather than allowed to interleave, and there |
| 447 | is a test for two hands both thinking they own the chain. |
| 448 | |
| 449 | **2.6 A journal write that reaches nothing returns `Ok`.** CONFIRMED. |
| 450 | `flush:1128` writes to a held descriptor; after the journal directory was removed |
| 451 | mid-run, `append` returned `Ok` and the refusal vanished. |
| 452 | |
| 453 | **CLOSED.** `flush` calls `sync_data` and returns the failure, so a write that |
| 454 | reaches nothing is an error and "journal before acting" refuses rather than |
| 455 | proceeds. |
| 456 | |
| 457 | **2.7 The Rust verifier and the documented shell verifier disagree.** CONFIRMED. |
| 458 | `str::lines()` strips a trailing `\r`, so a CRLF-converted journal verifies as |
| 459 | intact in Rust while the documented `sed | sha256sum` mismatches every line. Two |
| 460 | "independent" checks that disagree is the one failure this product cannot |
| 461 | afford. |
| 462 | |
| 463 | **CLOSED.** The verifier splits on `split_inclusive('\n')` rather than `lines()`, |
| 464 | so a `\r` is part of the line it is part of, and the Rust and shell verifiers |
| 465 | agree on a CRLF-converted journal. There is a test that converts one and checks |
| 466 | both. |
| 467 | |
| 468 | **2.8 Nothing calls the journal.** `main.rs:57` has no message loop, and neither |
| 469 | `exec.rs` nor `fence.rs` references `Journal`. "Journal before acting" is |
| 470 | asserted nowhere in code and by no test. |
| 471 | |
| 472 | --- |
| 473 | |
| 474 | **CLOSED.** `main.rs` has the message loop, and it opens the journal before it |
| 475 | serves anything — no journal, no service. The handshake, every command, every |
| 476 | signal and the closing line are written, and a command whose record cannot be |
| 477 | written is refused rather than run. |
| 478 | |
| 479 | ## Severity 3 — protocol and process |
| 480 | |
| 481 | **3.1 `chunk_fit` has no callers, and the shipping path can emit an unsendable |
| 482 | frame.** CONFIRMED. `codec.rs:1045`; `exec.rs:606` splits at a fixed `CHUNK_MAX` |
| 483 | without measuring. The `id` is caller-supplied and echoed on every chunk, and |
| 484 | `tools.rs` builds it as `run-{argv[0]}` from **model-chosen text**. A 300 KB |
| 485 | `argv[0]` plus one `CHUNK_MAX` run of control bytes measures 1,086,504 bytes; |
| 486 | `write_resp` returns `FrameTooBig` and the run's output is dropped. The |
| 487 | "measured, not calculated" property is true of a function nothing calls. |
| 488 | |
| 489 | **CLOSED.** The writer measures. `main.rs` calls `resp_fits` before writing and |
| 490 | `chunk_fit` to decide where to cut, so the caller-supplied `id` is paid for rather |
| 491 | than assumed; an oversized chunk is cut and the loss is reported, with tests for |
| 492 | both. |
| 493 | |
| 494 | **3.2 Trailing content after the first JSON value is silently discarded.** |
| 495 | CONFIRMED. `codec.rs:565,668`. A frame containing two `exec` objects runs the |
| 496 | first and never sees the second, while any reviewer, journal or policy layer |
| 497 | reading the same bytes with a real JSON parser rejects the frame outright or |
| 498 | sees something different. |
| 499 | |
| 500 | **CLOSED.** `want_strict_json` scans exactly one value and refuses any trailing |
| 501 | bytes, naming the offset where the message ended. A frame carrying two objects is |
| 502 | refused rather than half-obeyed. |
| 503 | |
| 504 | **3.3 An oversized frame poisons the stream permanently.** CONFIRMED. |
| 505 | `codec.rs:891-896`. `INBOUND_MAX` is 1 MB but `ext/hand.js:104` forwards |
| 506 | anything under 60 MB. On `LengthTooBig` the body is never consumed, so the next |
| 507 | four body bytes are read as a length prefix. Every queued request behind it is |
| 508 | lost, with no resynchronisation and nothing telling the page the real ceiling. |
| 509 | |
| 510 | **CLOSED.** The reader consumes an oversized body before refusing it, so the next |
| 511 | read starts at a frame boundary, and the page is told the real ceiling. Beyond |
| 512 | `RESYNC_MAX` the connection is ended instead, because a prefix that large is not |
| 513 | a message that went wrong. |
| 514 | |
| 515 | **3.4 `argv[0]` is never vetted against the fence.** CONFIRMED. `exec.rs:213`. |
| 516 | An absolute path outside the fence ran; `../outside/evil` ran; and a bare name |
| 517 | resolved through a caller-supplied `PATH`, because `env_clear()` then `execvp` |
| 518 | resolves against the *child's* environment, which the caller writes. Even with an |
| 519 | empty environment, glibc's `confstr(_CS_PATH)` fallback finds `/bin:/usr/bin`. |
| 520 | |
| 521 | **CLOSED.** `vet_program` resolves `argv[0]` once, to an absolute path, checks it |
| 522 | against the plan, and hands the launcher something already resolved so nothing |
| 523 | resolves it a second time. An absolute path outside the fence, a `..` spelling |
| 524 | and a bare name through a caller-supplied `PATH` are all refused. |
| 525 | |
| 526 | **3.5 `LD_PRELOAD` is settable by the caller.** CONFIRMED. `exec.rs:218-221`. |
| 527 | `README.md:39-41` states that a model able to name environment variables could |
| 528 | set `LD_PRELOAD`, and gives that as the reason the environment is not the |
| 529 | model's. The loop applies caller pairs verbatim with no screen. The guarantee is |
| 530 | stated and not implemented. (The app half sends `env:[]`, so this is reachable |
| 531 | only through 1.5/1.6 today.) |
| 532 | |
| 533 | **CLOSED.** `screen_env` refuses the loader variables outright before anything is |
| 534 | spawned, so the guarantee `README.md` states is now implemented rather than |
| 535 | merely stated. |
| 536 | |
| 537 | **3.6 A duplicate caller-chosen `id` makes a run unkillable and invisible.** |
| 538 | CONFIRMED. `exec.rs:282,472-474`. Two execs sharing an id leave one registry |
| 539 | slot: `live_count` reports 1 with two children alive, `stop_all` reaps one, and |
| 540 | the survivor outlives `Bye`. Conversely, when the shorter run ends it removes the |
| 541 | *live* run's entry, after which `signal` answers `Finished` for a process that is |
| 542 | still running. |
| 543 | |
| 544 | **CLOSED.** An identifier already in use is refused before there is a second |
| 545 | child, in a sentence that tells the caller to pick another or signal the run it |
| 546 | already has. There is no repair after the fact, so it is prevented instead. |
| 547 | |
| 548 | **3.7 `spawn` awaits the shared response channel, so one noisy run blocks every |
| 549 | later one.** CONFIRMED. With a flooding producer, a second `spawn` had not |
| 550 | returned after four seconds while its child was already running and unannounced. |
| 551 | A message loop that awaits `spawn` stops reading, so the `Signal` that would stop |
| 552 | the flood never arrives — head-of-line blocking on the one channel carrying both |
| 553 | control acknowledgements and bulk output. |
| 554 | |
| 555 | **CLOSED.** The loop is three parts that cannot block each other: a reader thread |
| 556 | on stdin, a dispatcher that never awaits `spawn`, and a writer holding stdout, |
| 557 | with the hand's own responses on a separate channel from bulk output. Proved by a |
| 558 | test that floods one run and shows the next request still being served. |
| 559 | |
| 560 | **3.8 No total output cap.** CONFIRMED. `yes` with a three-second timeout |
| 561 | delivered 3,406,442,688 bytes in 52,067 chunks. Memory is bounded, which is |
| 562 | good; nothing bounds the total, so the journal and the extension pipe absorb |
| 563 | gigabytes from one command. |
| 564 | |
| 565 | **CLOSED.** `OUTPUT_TOTAL_MAX` bounds a run at 20 MB across both streams |
| 566 | together, with one marker saying what was dropped. The *true* totals still travel |
| 567 | in `Ended`, so nothing is hidden — which is what 1.17 then uses. |
| 568 | |
| 569 | **3.9 Registry entries leak when `Started` cannot be sent.** CONFIRMED. |
| 570 | `exec.rs:280-289`. The insert precedes the send; on send failure `spawn` returns |
| 571 | `Err` without removing. Five attempts left five permanent entries that no signal |
| 572 | can clear. |
| 573 | |
| 574 | **CLOSED.** The registry entry is removed when `Started` cannot be sent, so a |
| 575 | failed announcement leaves nothing behind. |
| 576 | |
| 577 | **3.10 Group signalling degrades silently on BusyBox.** CONFIRMED for the |
| 578 | behaviour, PLAUSIBLE for the consequence. `busybox kill` rejects the `--` form, |
| 579 | and both call sites (`exec.rs:442-445,450-453`) discard the result with `let _ =`. |
| 580 | On Alpine — a realistic Cloud-tier host — `Kill` degrades to killing only the |
| 581 | direct child while the page is told `Ended{killed:true}`. `exec.rs:670-675` also |
| 582 | returns on the first binary that *spawns*, so a working `/usr/bin/kill` is never |
| 583 | reached. |
| 584 | |
| 585 | **CLOSED.** A real BusyBox 1.37 was put in front of the code path, linked as |
| 586 | `kill` so its own applet answers, and the finding was measured rather than |
| 587 | reasoned about. The behaviour is confirmed and the *consequence* was wrong in a |
| 588 | way worth recording: |
| 589 | |
| 590 | ``` |
| 591 | busybox kill -s TERM -- -<pgid> rc=1 stderr "kill: invalid number '--'" group killed |
| 592 | busybox kill -s TERM -<pgid> rc=0 group killed |
| 593 | procps kill -s TERM -- -<pgid> rc=0 group killed |
| 594 | procps kill -s TERM -<pgid> rc=1 (silent) group killed |
| 595 | ``` |
| 596 | |
| 597 | BusyBox counts the unreadable operand as an error and carries on to the next |
| 598 | one, so the group *does* die — what is lost is the exit status, not the signal. |
| 599 | That made the consequence the opposite of the one recorded: with the discarded |
| 600 | result now kept, BusyBox produced `Degraded`, and `supervise` would have sent the |
| 601 | page “anything it had started may still be running” about a group that was |
| 602 | already gone, and escalated a `Term` to a hard kill of the child on the strength |
| 603 | of it. No system takes both spellings, so one hard-coded form cannot serve both. |
| 604 | |
| 605 | `signal_group` now asks each `kill` which spelling it takes, by sending signal 0 |
| 606 | — the null signal, which validates arguments and delivers nothing — to the hand's |
| 607 | own process, and then signals once in the form that binary accepts. The obvious |
| 608 | alternative, sending with `--` and retrying without it, was tried and rejected: |
| 609 | it sends a second signal to a group the first has already emptied, and whether |
| 610 | that reports success turns on whether the leader has been reaped, so the sentence |
| 611 | the user reads would depend on the caller's bookkeeping rather than on what |
| 612 | happened to their command. |
| 613 | |
| 614 | Proved by `exec::tests::a_busybox_kill_reaches_the_group_and_says_so`, which |
| 615 | asserts the fixture really does reject `--` (so it stands for the finding rather |
| 616 | than standing in for it), that the probe answers `Bare` for BusyBox and |
| 617 | `Separated` for the system's `kill`, and that a real group with a real grandchild |
| 618 | dies and is reported `Sent`. Against the shipped code that test fails with |
| 619 | `Degraded("… exited 1 (kill: invalid number '--')")`. The test skips silently |
| 620 | where no BusyBox is installed rather than claiming a proof it did not perform. |
| 621 | |
| 622 | **3.11 `u64` fields do not survive the JavaScript half.** CONFIRMED. `seq`, |
| 623 | `out_bytes`, `err_bytes` and `timeout_ms` exceed `Number.MAX_SAFE_INTEGER`; |
| 624 | `seq: u64::MAX` reads back as `18446744073709552000` under `JSON.parse`. |
| 625 | `www/js/hand.js:98` compares `msg.seq !== want`, so the gap detection depends on |
| 626 | a value the wire cannot faithfully carry. |
| 627 | |
| 628 | **CLOSED, by refusing rather than clamping.** `wire.rs` now states the ceiling |
| 629 | as part of the contract — `SAFE_INT_MAX`, 2^53 − 1 — and it binds all five `u64` |
| 630 | fields the wire has: `Chunk.seq`, `Output.seq`, `Ended.out_bytes`, |
| 631 | `Ended.err_bytes` and `Exec.timeout_ms`. Every one of them is read through one |
| 632 | function in `codec.rs`, so a `u64` added later inherits the rule by being read the |
| 633 | same way, and every one is checked again on the way out, before a byte is |
| 634 | written. The rest of the wire's numbers are `u32`, `i32` or `u16` and cross |
| 635 | unharmed; `exit` is `i32` at both ends, so `-1` survives. |
| 636 | |
| 637 | Clamping was the obvious fix and is the wrong one, because it is the shape of |
| 638 | §1.11: a value quietly replaced by a plausible one, which is how a killed |
| 639 | `cargo test` was read as a green build. A clamped `out_bytes` would say nothing |
| 640 | went missing — which is precisely what §1.17 uses that field to detect — and a |
| 641 | clamped `seq` would make two frames compare equal, breaking the gap detection it |
| 642 | exists for. String encoding was the other candidate and moves the same ceiling |
| 643 | into every reader's `BigInt` handling, changing the contract for five fields to |
| 644 | fix a case none of them can reach. So the rule is the one with no silent arm: a |
| 645 | named `Fault` on decode, a refusal to write on encode, and no invented value at |
| 646 | either end. Nothing legitimate is refused — the ceiling is 800 exabytes of |
| 647 | terminal output, nine petabytes down one pipe, or a wall-clock limit of 285,000 |
| 648 | years. |
| 649 | |
| 650 | Proved by `codec::tests::every_wire_number_is_refused_past_what_javascript_can_hold`, |
| 651 | which fails against each of the four guards removed in turn, and end to end |
| 652 | against the release binary over a pipe: an exec carrying `timeout_ms: 2^60` is |
| 653 | answered with the named fault, the command does not run, and the next exec on the |
| 654 | same connection runs normally. |
| 655 | |
| 656 | Two residuals, neither of them this finding. A frame that fails to decode is |
| 657 | answered with `Error{id:null}`, so a page waiting on that run learns nothing |
| 658 | until its own grace timer fires — true of every malformed field, not just this |
| 659 | one, and unreachable through the extension, which screens `timeout_ms` before |
| 660 | forwarding. And `exec.rs`'s `clamp_timeout` still narrows a limit above 24 hours |
| 661 | silently; that is a policy ceiling on an input rather than a misreported result — |
| 662 | the run is killed and `timed_out` says so — but it is the same family and worth a |
| 663 | sentence to the model one day. |
| 664 | |
| 665 | **3.12 The decoder accepts JDAT, not JSON.** CONFIRMED. `codec.rs:260`. |
| 666 | `{'t':'bye'}`, `{t:"bye"}`, a trailing comma, a `#` comment and typed values are |
| 667 | all accepted and all rejected by `JSON.parse`. Reachable in Cloud mode, where |
| 668 | the bytes are not Chrome-serialised. |
| 669 | |
| 670 | --- |
| 671 | |
| 672 | **CLOSED, with 3.2, by the same function.** `want_strict_json` is a JSON scanner |
| 673 | run before the JDAT decoder, so single quotes, unquoted keys, trailing commas, |
| 674 | `#` comments and typed values are refused rather than accepted by a superset. Each |
| 675 | of the accepted spellings the finding lists is a case in the test. |
| 676 | |
| 677 | **3.13 `screen_env` is a speed bump, not a boundary: the caller sets |
| 678 | `LD_PRELOAD` through `argv` instead.** CONFIRMED. `exec.rs:1681-1703` (the |
| 679 | screen), applied at `exec.rs:447` in `Runner::spawn`, at `exec.rs:2464` in |
| 680 | `launch_inner`, and at `pty.rs:447` for a terminal session. Line numbers here are |
| 681 | as of the working tree this file ships in, not of `ad19a62`. |
| 682 | |
| 683 | `README.md:39-41` states the guarantee as *"The environment is not the model's to |
| 684 | set … a model that could name environment variables could set `LD_PRELOAD`, or |
| 685 | carry a stolen value out through one."* 3.5 made that true of the `env` **field** |
| 686 | of the request. It is not true of the request. `/usr` is in the hand's read-only |
| 687 | system base, so `/usr/bin/env` and `/bin/sh` are inside every fence, and both of |
| 688 | them set an environment out of their own arguments — which `screen_env` never |
| 689 | looks at, because `argv` is not an environment. |
| 690 | |
| 691 | Measured through the real `Runner::spawn`, the real launcher, and a real |
| 692 | Landlock ruleset with the syscall filter installed: |
| 693 | |
| 694 | ``` |
| 695 | 0 env field env:[("LD_PRELOAD","/x/nope.so")] |
| 696 | Refused: this command asked to run with LD_PRELOAD set. … |
| 697 | |
| 698 | 1 argv ["/usr/bin/env","FOO=bar","/usr/bin/printenv","FOO"] |
| 699 | out: "bar\n" exit 0 |
| 700 | |
| 701 | 2 argv ["/usr/bin/env","LD_PRELOAD=/x/nope.so","/bin/echo","hi"] |
| 702 | out: "hi\n" |
| 703 | err: "ERROR: ld.so: object '/x/nope.so' from LD_PRELOAD |
| 704 | cannot be preloaded (cannot open shared object file): |
| 705 | ignored." exit 0 |
| 706 | |
| 707 | 3 argv ["/bin/sh","-c","export LD_PRELOAD=/x/nope.so; exec /bin/echo hi"] |
| 708 | err: the same ld.so line exit 0 |
| 709 | |
| 710 | 4 argv ["/usr/bin/env","LD_PRELOAD=<ws>/copy.so","/bin/echo","hi"] |
| 711 | (copy.so is a copy of libm.so.6, written INSIDE the fenced |
| 712 | workspace and named from there) |
| 713 | out: "hi\n" err: "" exit 0 |
| 714 | ``` |
| 715 | |
| 716 | Case 0 is the screen doing its job. Case 1 is the model setting an environment |
| 717 | variable anyway. Cases 2 and 3 are the *dynamic loader* acting on a variable the |
| 718 | model chose — the diagnostic is proof the value was read, not merely carried. |
| 719 | Case 4 is the one that matters: a real shared object, in a directory the command |
| 720 | can write, loaded with **no diagnostic and exit 0**. The contrast between 4 and |
| 721 | 2 is what makes 4 a success rather than a silent ignore. |
| 722 | |
| 723 | **The exposure**, which the resolution below narrows the claim to rather than |
| 724 | removes: |
| 725 | |
| 726 | * **Not an escape from the fence.** Everything in cases 1–4 runs inside the same |
| 727 | Landlock ruleset and the same seccomp filter as any other command; the preloaded |
| 728 | object has to be readable and executable through the same grants, so it is the |
| 729 | workspace or the read-only system base and nowhere else. The compartment holds. |
| 730 | * **It is a hole in a different claim** — that what a command runs *with* is the |
| 731 | user's decision. A `.so` in the workspace, loaded into every later command, |
| 732 | makes `cargo test` do something the user never granted and the journal records |
| 733 | the argv it was asked for. A model that can write a file and compile it (the |
| 734 | Rust toolkit grants `cargo`) can do this on one turn and use it on the next. |
| 735 | * **`TMPDIR` and `TERM` are screened the same way and are past the same way** |
| 736 | (`screen_scratch`, `screen_term`), so the private scratch directory is likewise |
| 737 | a default rather than a decision. |
| 738 | * **The pty half is worse**, not better: `Pty::open` calls the identical |
| 739 | `screen_env` and then runs an interactive shell, where `export LD_PRELOAD=…` is |
| 740 | the ordinary way to set a variable. |
| 741 | |
| 742 | What a fix would have to be, so that the next person does not reach for the |
| 743 | cheap one: screening `argv` for `env` and `sh` is a fence made of string matching, |
| 744 | which `README.md:31` already refuses on its own terms — `/usr/bin/env` has as many |
| 745 | spellings as `argv[0]` had before 3.4, and every interpreter in the base can |
| 746 | export a variable. The two candidates that are not string matching are (a) an |
| 747 | allow-list of programs the base may execute, which is a much larger decision than |
| 748 | this finding, and (b) refusing `LD_*` at the *loader* rather than at the request, |
| 749 | which on Linux means the filter, not `exec.rs`. |
| 750 | |
| 751 | **CLOSED BY CORRECTING THE CLAIM, 2026-08-07 — the behaviour is unchanged.** |
| 752 | |
| 753 | Of the two candidates above, (b) is not available. `execve` carries its |
| 754 | environment as a pointer to an array of pointers, and `seccomp.rs:27` states the |
| 755 | reason it cannot be read there: *"a seccomp filter cannot dereference a |
| 756 | pointer."* The kernel hands the filter `seccomp_data` — the syscall number, the |
| 757 | instruction pointer and six argument registers — and nothing they point at. So |
| 758 | "refuse `LD_*` at the loader" is not a smaller version of this fix; on a |
| 759 | seccomp-bpf filter it is not a fix at all, and reaching it would mean a `ptrace` |
| 760 | supervisor or an LSM of our own. Candidate (a), an allow-list of executables, |
| 761 | decides what the Machine tier is *for* — a general shell for the user's own |
| 762 | toolchain, or a menu — and that is a product decision, not a defect repair. It |
| 763 | is not made here. |
| 764 | |
| 765 | Which leaves the third option, and on the evidence it is the right one: the |
| 766 | guarantee is narrower than `README.md` states, so `README.md` should state the |
| 767 | narrower guarantee. A promise that is 90% true is worse than a smaller one that |
| 768 | is wholly true, because it is the 10% nobody checks. The sentence at |
| 769 | `README.md:39-41` — *"The environment is not the model's to set, for the same |
| 770 | reason: a model that could name environment variables could set `LD_PRELOAD`, or |
| 771 | carry a stolen value out through one."* — should read: |
| 772 | |
| 773 | > The `env` field is not the model's to set: a model that could name environment |
| 774 | > variables through it could set `LD_PRELOAD`, or carry a stolen value out |
| 775 | > through one. That screen covers the field and not the request. `/usr/bin/env` |
| 776 | > and `/bin/sh` are in the read-only system base, and both take an environment |
| 777 | > out of their own arguments, so a command can still choose what it runs with — |
| 778 | > from inside the fence, using only what the fence already grants. What bounds |
| 779 | > that is the compartment, not the screen: see `REVIEW.md` §3.13. |
| 780 | |
| 781 | Three reasons this is the resolution and not a retreat: |
| 782 | |
| 783 | * **The compartment is what the document is about, and it holds.** Cases 1–4 run |
| 784 | inside the same Landlock ruleset and the same filter as any other command. The |
| 785 | finding never was an escape, and closing it does not make one. |
| 786 | * **The false half of the sentence was the dangerous half.** A reader deciding |
| 787 | whether to grant the Machine tier weighs what it promises. "The environment is |
| 788 | not the model's to set" invites them to believe a command cannot choose its own |
| 789 | loader, which is exactly what a `.so` written into the workspace does. Naming |
| 790 | that is worth more than a screen that string-matching would defeat. |
| 791 | * **Nothing is silently dropped.** `TMPDIR` and `TERM` are past the same way and |
| 792 | the corrected sentence covers them, because it stops claiming the request is |
| 793 | screened at all. The pty half is likewise honest: `export LD_PRELOAD=…` in an |
| 794 | interactive shell is the ordinary way to set a variable, and the corrected text |
| 795 | no longer implies otherwise. |
| 796 | |
| 797 | The `README.md` edit itself is the lead's — this document does not own that file |
| 798 | — and it is the one line this finding leaves outstanding. |
| 799 | |
| 800 | **`screen_git_push` is NOT closed by this**, and the paragraph below still |
| 801 | stands. The two are the same shape but not the same finding: one is a claim about |
| 802 | the environment, and the other is a credential guard whose bypass has a |
| 803 | consequence outside the fence. |
| 804 | |
| 805 | **The same shape appears in the new `screen_git_push`** (`exec.rs`, the section |
| 806 | comment says so at the point of definition): it matches `argv[0]` by basename, so |
| 807 | `sh -c 'git push --force'` is past it. There the page's answer — a command whose |
| 808 | `argv[0]` is not `git` gets no credential, so it cannot authenticate — does not |
| 809 | transfer, because the harm that guard exists for is a repository holding |
| 810 | credentials of its own. Same defect, same reason, and it is written down at both |
| 811 | ends rather than counted as covered. |
| 812 | |
| 813 | ## Severity 4 — the page relay |
| 814 | |
| 815 | **4.1 `www/js/hand.js` is dead code.** Nothing calls `init`, `setExtId` or |
| 816 | `adopt`; `run()` never sends `hello` and never listens for one, so `state.root` |
| 817 | can never be populated and `Tool::run` always refuses. `dev/verify_hand.mjs` |
| 818 | drives the raw port and never loads this file, so none of it is covered by any |
| 819 | test. Written by the lead and not wired in. |
| 820 | |
| 821 | **CLOSED.** `www/index.html` loads `js/hand.js`, `js/daimond.js` calls |
| 822 | `DaimondHand.init(...)`, and `src/wasm/hand.rs` binds the object the `run` tool |
| 823 | reaches through. It is wired at every joint the finding said it was not. |
| 824 | |
| 825 | **4.2 A hostile or buggy host holds the page open for ever.** CONFIRMED. |
| 826 | `www/js/hand.js:138-146`. Any message carrying a `t` field resets the grace timer |
| 827 | *before* the type is dispatched, so `{"t":"noop"}` every 700 ms leaves the promise |
| 828 | pending indefinitely — the exact failure `REPLY_GRACE` exists to prevent. |
| 829 | |
| 830 | **CLOSED.** An unknown message type touches no timer and no run, so a host |
| 831 | repeating `{"t":"noop"}` cannot hold the promise open. |
| 832 | |
| 833 | **4.3 A genuinely quiet command is killed at 30 seconds.** CONFIRMED. |
| 834 | `www/js/hand.js:66,144`. The grace period is refreshed only by output, so a host |
| 835 | that sends `started` and then nothing is rejected with "stopped part-way through |
| 836 | the command" while the process is still alive. This is precisely the `cargo test` |
| 837 | case the file header says the design is for. |
| 838 | |
| 839 | **CLOSED.** The grace period is no longer refreshed by output alone, so a `cargo |
| 840 | test` that says nothing for a minute is not reported as a host that died |
| 841 | part-way. |
| 842 | |
| 843 | **4.4 Chunks accumulate with no cap.** PLAUSIBLE. `www/js/hand.js:103,167`. |
| 844 | `CHUNK_MAX` bounds a frame; nothing bounds the total, so a command printing many |
| 845 | gigabytes exhausts the tab long before `truncate_output` in `tools.rs` runs. |
| 846 | |
| 847 | --- |
| 848 | |
| 849 | --- |
| 850 | |
| 851 | **CLOSED.** The page keeps 256 kB of each end of a stream and states in the middle |
| 852 | what was dropped, so a runaway command cannot grow an array until the tab dies and |
| 853 | nobody reads a hole as continuity. Both ends rather than the first, because a |
| 854 | build says what it is doing at the start and why it failed at the end. |
| 855 | |
| 856 | ## Severity 1 (continued) — the app half, and the seams |
| 857 | |
| 858 | **1.9 `ToolContext::no_write` is never populated in the browser build.** |
| 859 | CONFIRMED. `src/wasm/app.rs:221` `set_diamond_scope` is the only thing that ever |
| 860 | builds `diamond_bounds`, and it has **no caller anywhere in the repo**; the |
| 861 | daimon's own steering context sets `no_write: Vec::new()` explicitly |
| 862 | (`src/wasm/app.rs:689`). So `fence_spec` always takes the `rw.push(root)` |
| 863 | fallback and every command is fenced to the entire granted folder. |
| 864 | |
| 865 | The claim at `src/tools.rs:1959-1960` — "a daimon's command reaches exactly the |
| 866 | files its `file_read` would have reached" — is therefore **false**: `file_read` |
| 867 | is pinned to OPFS `diamonds/<id>`, and the command gets the whole grant. Every |
| 868 | other divergence below concerns a translation that nothing currently feeds. |
| 869 | |
| 870 | **OPEN at the time of this pass, and being closed elsewhere.** `set_diamond_scope` |
| 871 | now documents this finding by name and is written to scope the whole turn — both |
| 872 | doors, the file tools and the fence — but a search of `www/js` finds no caller, |
| 873 | so the browser build still takes the `rw.push(root)` fallback. A separate agent |
| 874 | was landing the caller as this was written; whoever reads this next should check |
| 875 | `www/js` for the call before trusting either state. |
| 876 | |
| 877 | **1.10 `fence_spec` is not a faithful restatement of `Bound`.** CONFIRMED, three |
| 878 | ways, each reproduced by a test. |
| 879 | |
| 880 | - *The allow-list-before-carve-out ordering is not preserved.* `tools.rs:253` |
| 881 | pushes every `Bound::MayRead` into `ro` unconditionally, while |
| 882 | `ToolContext::may_read` tests `within_allow_list` first and refuses. With a |
| 883 | Diamond's bounds plus `MayRead("elsewhere/secrets")`, the app refuses the read |
| 884 | and the fence grants it. The `.daimond` deny does not rescue it, because |
| 885 | `fence.rs:1050-1057` expresses a deny as the *absence* of a rule, so a narrower |
| 886 | `ro` grant beneath it survives. |
| 887 | - *Any `MayRead` defeats the absent-allow-list guard.* `tools.rs:260` tests |
| 888 | `rw.is_empty() && ro.is_empty()`; a skill's carve-out fills `ro`, so the guard |
| 889 | never fires and `rw` stays empty. A skill-bounded turn that may write the whole |
| 890 | workspace in the app gets a fence with **no writable root at all**, and |
| 891 | `vet_cwd` then refuses every command. The condition is the bug. |
| 892 | - *A `NoWrite` nested inside an `OnlyUnder` is dropped.* `tools.rs:250` demotes |
| 893 | only when the `OnlyUnder` sits inside a `NoWrite`, never the reverse, so |
| 894 | `attached=["proj"], read_only=["proj/docs"]` yields an `rw` covering |
| 895 | `proj/docs`. Not exploitable end to end — `fence.rs` carves `ro` out of an |
| 896 | enclosing `rw` — but the spec read alone says "writable", so any second |
| 897 | consumer (a log line, the grant window, a non-Landlock backend) reads it wrong. |
| 898 | |
| 899 | **CLOSED, all three ways.** The allow-list is tested before the carve-out, so a |
| 900 | `MayRead` outside a Diamond's bounds is dropped from `ro` exactly as `may_read` |
| 901 | would refuse it. The absent-allow-list guard now asks whether an allow-list was |
| 902 | *declared* (`if !scoped { rw.push(root) }`) rather than whether the lists came out |
| 903 | empty, so a skill's carve-out no longer defeats it. And a `NoWrite` nested inside |
| 904 | an `OnlyUnder` is re-stated as a read-only root for the hand to carve, in both |
| 905 | directions, so the spec read alone now says what the fence does. |
| 906 | |
| 907 | **1.11 A killed or crashed command is reported to the model as `exit code: 0`.** |
| 908 | CONFIRMED. `tools.rs:2040` uses `extract_json_number`, which returns |
| 909 | `Option<u64>` and fails to parse `-1`, so `.unwrap_or(0)` makes it zero. |
| 910 | `ext/hand.js:321` sends exactly `{exit:-1, killed:true}` when the native host |
| 911 | dies mid-run, and `wire.rs:206` reserves `-1` for "no exit status". The daimon |
| 912 | reads partial output plus `[exit code: 0]` and reports a broken build as green. |
| 913 | `killed` is never read by `run_result` at all, so nothing else catches it. |
| 914 | |
| 915 | **CLOSED.** `run_result` reads the status with `extract_json_i64`, so `-1` is |
| 916 | `-1`, and the three cases are told apart in the sentence the model reads: timed |
| 917 | out, stopped before it finished, or did not finish and reported no exit code. The |
| 918 | comment at that line records why, because this was the worst defect of the |
| 919 | session. |
| 920 | |
| 921 | **1.12 `src/handler.rs:530` assigns rather than composes.** |
| 922 | `narrowed.ctx.no_write = skill_bounds(&dirs)` discards any allow-list already on |
| 923 | the context. The composition documented at `tools.rs:126-129` — "the two compose, |
| 924 | and the allow-list wins" — describes something no code does. If a Diamond-scoped |
| 925 | registry ever reaches that line, the Diamond's fence is deleted outright. |
| 926 | |
| 927 | **CLOSED, and the documentation was the one telling the truth.** `src/tools.rs` |
| 928 | now has `compose`, and `src/handler.rs:537` calls it: a skill's bound is merged |
| 929 | into whatever the turn already carried rather than written over it. The rule is |
| 930 | the one the doc comment always claimed — a path survives only where BOTH lists |
| 931 | permitted it — and it is now stated as an invariant rather than a description. |
| 932 | **Composing can never widen either input**, so that line can only narrow, which |
| 933 | is what makes it safe to call from anywhere a bound is set. |
| 934 | |
| 935 | Rule by rule. Allow-lists INTERSECT, and the intersection of two prefix sets is |
| 936 | exactly expressible rather than approximated: where one prefix contains the |
| 937 | other the answer is the deeper of the two, and where they are disjoint it is |
| 938 | nothing. Nothing becomes `Bound::Nowhere` and NOT an empty list -- an empty |
| 939 | allow-list reads as no allow-list at all, which is the widest answer there is to |
| 940 | the narrowest question, and that is the empty-prefix trap arriving through the |
| 941 | merge. Denials union. A read carve-out survives only where the other list would |
| 942 | have permitted the whole of its subtree anyway, because `may_read` answers a |
| 943 | carve-out before it looks at any deny, so one carried across would punch through |
| 944 | the other list's denials. A `Toolkit` survives only where both sides granted it: |
| 945 | a toolchain is machine paths a command may reach, and carrying one into a turn |
| 946 | whose other bound granted none would widen that turn's fence. |
| 947 | |
| 948 | One consequence is worth saying plainly rather than discovering. **A skill |
| 949 | running inside a Diamond cannot read its own shipped references.** A carve-out |
| 950 | is a hole punched in its own deny fence, and `Bound::MayRead` has always said it |
| 951 | does not escape an `OnlyUnder`; the Diamond denies the whole of `.daimond/` and |
| 952 | allow-lists none of it, so the hole closes. The skill is refused in those words |
| 953 | rather than the Diamond being quietly widened to fit it. Nothing behaves |
| 954 | differently today -- the native handler's context carries no bounds, so a skill |
| 955 | turn composes to exactly `skill_bounds`, as it did when the line assigned -- and |
| 956 | the day skills are wired into the browser this fails closed and loudly instead |
| 957 | of silently and open. |
| 958 | |
| 959 | Ten tests, and not one of them passes on the code it replaced: eight fail when |
| 960 | the second bound is assigned over the first, seven fail on the other order, and |
| 961 | the union is all ten. The four-way case is |
| 962 | `test_a_turn_bounded_twice_reaches_what_both_permit_00` -- own Diamond readable |
| 963 | and writable, other Diamond refused, Daimond's own directory refused, the |
| 964 | skill's carve-out subordinate to the allow-list and alive again the moment the |
| 965 | turn is not scoped. The walkers are checked too, in both directions: an |
| 966 | allow-list the merge could have dropped, and a deny the merge could have dropped |
| 967 | inside one. And the section checks itself: `composition_checks` holds eighteen |
| 968 | named effects of the merge, and |
| 969 | `test_every_check_here_fails_on_the_assignment_it_replaced_00` runs every one |
| 970 | against both assignment orders and asserts that twelve of them are wrong under |
| 971 | an assignment. The other six are declared as liveness checks, so the count means |
| 972 | what it says. |
| 973 | |
| 974 | `DaimondApp::set_diamond_scope` now composes as well, so that a second caller |
| 975 | can never widen the first: on a freshly built app it is the identity, re-scoping |
| 976 | the same Diamond is idempotent, and re-scoping a *different* one intersects to |
| 977 | `Bound::Nowhere`, which `diamond_scope` reports and the caller already refuses to |
| 978 | start a turn on. |
| 979 | |
| 980 | **1.13 The first command of a turn costs the turn its network.** CONFIRMED. |
| 981 | `run_result` wraps output through `ctx.wrap_untrusted`, which sets `tainted`, and |
| 982 | `Tool::run` reads `is_tainted()` when building the fence. So `cargo fetch` |
| 983 | succeeds and then every later command in the same turn gets `net:false`, with no |
| 984 | explanation the model can act on. The gate itself is sound — one-way flag, read |
| 985 | at fence-build time, no ordering hole — but treating a command's own output as |
| 986 | tainting input makes the network rule fire on the second command of every |
| 987 | ordinary build session. A design decision, not a slip, and it needs revisiting. |
| 988 | |
| 989 | **CLOSED in the engine, and it needs one patch to `www/js/daimond.js` before a |
| 990 | user sees any of it.** Remedies 1 and 2 are built; remedy 3 was explicitly NOT |
| 991 | authorised and has not been touched. What was OPEN here is now decided: the user |
| 992 | authorised remedy 2 as well as remedy 1, and the paragraph below that said |
| 993 | remedy 1 "is the only one being done" was true when it was written and is not |
| 994 | true now. |
| 995 | |
| 996 | *What is done.* The network is no longer taken away in silence. A turn that has |
| 997 | read outside content is ASKED, once, naming the command; the answer holds for |
| 998 | the rest of that turn, whichever way it went; and a command that ends up without |
| 999 | the network is TOLD so, outside the untrusted envelope, in words that say the |
| 1000 | project is not at fault. Nothing was loosened: no answer, no network. A |
| 1001 | dispatched worker is not asked at all -- there is nobody to ask -- and keeps |
| 1002 | exactly the fence it had. |
| 1003 | |
| 1004 | *What is NOT done, and must be said plainly.* The question is put through |
| 1005 | `window.__daimondEgressAllowed`, which lives in `www/js/daimond.js`, and that |
| 1006 | file belongs to another lane. Its gate switches on `req.tool`, and it does not |
| 1007 | yet know the name `run_net`. It therefore falls through to the URL work below, |
| 1008 | where the answer it gives is `allow` -- a command line resolves against the page |
| 1009 | as a same-origin URL, and same-origin is waved through -- so the engine does not |
| 1010 | take `allow` for an answer to this question at all. It takes `allow-net`, which |
| 1011 | only the branch written for this question can say, and everything else means no. |
| 1012 | Today that means the network is withheld exactly as it always was, plus the |
| 1013 | sentence explaining it, and **no user is asked anything: remedy 2 is built and |
| 1014 | unreachable until the branch lands.** The patch is about fifteen lines and the |
| 1015 | three strings it needs (`permmode.net_title`, `permmode.net_body`, |
| 1016 | `permmode.net_ok`) are already in all eight locale files. |
| 1017 | |
| 1018 | *Where it lives.* `NetStep`, `net_step`, `net_needs_consent`, `net_verdict` and |
| 1019 | `RUN_NET_TOOL` in `src/tools.rs`; the answer on `TurnState::net_consent`, read |
| 1020 | and written through `ToolContext::net_consent` / `set_net_consent`; |
| 1021 | `Tool::run` composing the four lines that ask and record; `NO_NET_NOTE` and |
| 1022 | `push_no_net_refusal` saying what happened; and `prompts::machine_note` telling |
| 1023 | the model which of the four states its turn is in. |
| 1024 | |
| 1025 | *What "for the turn" means, exactly.* The answer is kept on `TurnState`, which |
| 1026 | is the same object the taint is kept on, so it lasts precisely as long as the |
| 1027 | condition that provoked the question. In the browser a chat's `ToolContext` is |
| 1028 | built once per chat and never reset, so for the user's own chat "the turn" is |
| 1029 | the life of that chat's app instance -- the same scope `tainted` already has. A |
| 1030 | dispatched worker gets a fresh context, so a worker's turn really is one turn. A |
| 1031 | yes therefore covers more than one message in a chat, and that is the same |
| 1032 | breadth as the taint it answers; it is written down here because "for the turn" |
| 1033 | would otherwise be read as "for one message". |
| 1034 | |
| 1035 | *Where the question sits in the order of checks.* Everything that refuses a |
| 1036 | command for what it is comes first -- no hand, no root, a hand that cannot fence, |
| 1037 | a `cwd` out of the workspace or in browser storage, and bounds that describe |
| 1038 | nowhere to run, which is settled off a fence built with the network withheld |
| 1039 | because the roots do not depend on it. The question comes after all of those and |
| 1040 | BEFORE `git_step`, which is the one place the "ask last" rule is given up on |
| 1041 | purpose: a push refused for having no network is this defect in its most acute |
| 1042 | form, so the question has to precede it. The cost is that a `--force` push, which |
| 1043 | `git_guard` was going to refuse anyway, can cost one dialog -- and not a wasted |
| 1044 | one, since the answer covers every later command in the turn. |
| 1045 | |
| 1046 | *What this does NOT cover.* `wasm::pty` builds its own fence for an interactive |
| 1047 | terminal, out of the same taint and the same rung, and it still withdraws the |
| 1048 | network without asking. That is deliberate and not an oversight: the person |
| 1049 | typing into a terminal is present by definition, sees the failure in front of |
| 1050 | them, and can open a fresh session -- which is the same argument `pty.rs` |
| 1051 | already makes for not putting `Mode::Ask`'s question to a terminal. If it ever |
| 1052 | should ask, it is the same three parts and one more call. |
| 1053 | |
| 1054 | *A wart, recorded rather than hidden.* On the Ask rung a tainted turn's first |
| 1055 | command puts two dialogs in a row: the network question, then "Run this |
| 1056 | command?". They are different questions and both are needed, and the first is |
| 1057 | asked once per turn while the second is asked every time, so it settles after |
| 1058 | one command. Folding them into one dialog needs the page half and was not done. |
| 1059 | |
| 1060 | *Exactly what used to happen* (line numbers as they were when this was found). |
| 1061 | `Tool::run` built the fence with |
| 1062 | `fence_spec(&ctx.no_write, &machine, ctx.is_tainted())` (`tools.rs:2766`) and |
| 1063 | `fence_spec` sets `net: !tainted` (`:904`). Every `run` ends by returning |
| 1064 | `ctx.wrap_untrusted(&origin, &s)` (`:2915`), and `ToolContext::wrap_untrusted` |
| 1065 | sets `tainted = true` (`:1507`). The flag is one-way within a turn. So the first |
| 1066 | command that *ran* cost the turn its network, whatever it was and whether or not |
| 1067 | it printed anything; a command that is *refused* returns before the wrap and |
| 1068 | does not. The same flag also arms `egress_check`, so `web_fetch` and `web_open` |
| 1069 | start asking for consent from that point on -- which is untouched by any of this |
| 1070 | and is still true. |
| 1071 | |
| 1072 | *What did change there, on 2026-08-27, and what did not.* Those two now ask ONCE |
| 1073 | per conversation and the answer covers every site, rather than asking about each |
| 1074 | host and remembering the host for as long as the page lives (`web_step` and |
| 1075 | `TurnState::web_consent` in `src/tools.rs`). The owner ruled the old scope a |
| 1076 | defect after a tester reported being asked about every new address. It does not |
| 1077 | touch this section: the question here is whether a COMMAND keeps its network |
| 1078 | after the turn is tainted, it is answered by `net_step` from `net_consent`, and |
| 1079 | neither function reads the other's field. An address carrying more than an |
| 1080 | address needs is still put to the user on its own terms, whatever the |
| 1081 | conversation has granted, which is what keeps the widening from reopening the |
| 1082 | channel §7 closes. |
| 1083 | |
| 1084 | *What it costs.* `cargo fetch` then `cargo build`; `npm install` then `npm test`; |
| 1085 | `git fetch` then `git pull` — in each pair the second command runs with no |
| 1086 | network and fails with the toolchain's own offline error. The model reads a |
| 1087 | network failure it has no way to attribute, and the most likely thing it does |
| 1088 | next is report the project as broken, which is §1.11's failure mode arriving by a |
| 1089 | different road. The turn that most needs the network is a build, and a build is |
| 1090 | the thing this makes offline. |
| 1091 | |
| 1092 | *Why the rule is nonetheless right.* Command output is not the user's words. A |
| 1093 | build log carries a dependency's name, a test fixture, a fetched page; `curl` is |
| 1094 | an argv like any other. Marking it untrusted is the same judgement `shell` |
| 1095 | already makes and should not be undone. The defect is not the mark — it is what |
| 1096 | the mark is wired to, and that it is silent. |
| 1097 | |
| 1098 | *The three remedies, in the order they were to be done, and what became of each.* |
| 1099 | |
| 1100 | 1. **Say it.** DONE. It loosens nothing, and it turns an unattributable build |
| 1101 | failure into a fact the model can report. |
| 1102 | |
| 1103 | *Where.* The flag is captured in `Tool::run` beside the fence -- `let no_net = |
| 1104 | !fence.net;` -- and passed to `run_result` rather than read again there. |
| 1105 | Asking the context a second time gives the same answer today only because |
| 1106 | nothing between the two calls taints the turn, which is an accident of |
| 1107 | ordering that the next edit breaks silently. It is now read off the fence |
| 1108 | itself, so what the model is told is what the command actually ran with. |
| 1109 | |
| 1110 | *What.* One more arm on the `tail` that already carries `[exit code: 0]` and |
| 1111 | `[timed out; the command was killed]`, appended after it, outside the |
| 1112 | untrusted envelope because this is the app speaking and not the command: |
| 1113 | `NO_NET_NOTE` in `src/tools.rs`. It now names the reason a turn has no network |
| 1114 | -- the user declined, or there was nobody to ask -- because after remedy 2 |
| 1115 | "every command in it runs with the network refused" is no longer the whole |
| 1116 | truth. `push_no_net_refusal` says the same thing for the one command whose |
| 1117 | whole purpose is the network. |
| 1118 | |
| 1119 | *Why unconditionally rather than only on failure.* Deciding which failures are |
| 1120 | network failures means matching prose from `cargo`, `npm`, `git` and every |
| 1121 | other toolchain, which is the guessing this is meant to end. The cost is one |
| 1122 | line of noise on a command that did not need the network; the benefit is that |
| 1123 | the one that did needs no guessing at all. |
| 1124 | |
| 1125 | *Proved by* `test_the_no_network_note_is_true_when_said_and_outside_the_ |
| 1126 | envelope` (the note is outside the envelope, after the exit code, and absent |
| 1127 | from a networked run) and `test_a_refused_command_is_not_also_told_about_the_ |
| 1128 | network`. |
| 1129 | 2. **Ask rather than withdraw.** DONE in the engine, and waiting on the page |
| 1130 | half (see above). The `run` case goes through the door `egress_check` already |
| 1131 | uses -- the same `Verdict`, the same global, the same three parts -- and not a |
| 1132 | second consent mechanism beside it. On a tainted turn a command that would |
| 1133 | have had the network prompts once, naming the command and the folder, and the |
| 1134 | answer holds for the turn. The review's own text says the gate is sound; it is |
| 1135 | the silence that is wrong, and a user watching `cargo build` will say yes |
| 1136 | while a user watching an unexplained `curl` will not. |
| 1137 | |
| 1138 | *What is proved, and how.* The call site is `Tool::run`'s wasm arm, which no |
| 1139 | test can reach, so it was reduced to four lines over pure functions and those |
| 1140 | are held by native tests in `tools.rs`: `net_step` decides, `net_verdict` |
| 1141 | turns silence into a no, and `ToolContext` remembers. Each test was seen |
| 1142 | failing first, against a `net_step` that behaved as the code it replaced. |
| 1143 | Three of them are the ones that matter. **Holds for the turn**: commands two |
| 1144 | and three of the same turn are handed a closure that would say NO, and still |
| 1145 | come back with the network the user restored -- so it cannot pass by nobody |
| 1146 | being consulted. **A new turn asks again**: a fresh context has no answer and |
| 1147 | is asked. **A no is a no**: a declined turn runs its command inside the same |
| 1148 | fence with `net:false`, is not asked again, and gets the note; and a worker |
| 1149 | with a consent forged into its state is still refused, because the reason is |
| 1150 | the actor and not the answer. `test_a_yes_moves_nothing_but_the_network` |
| 1151 | holds the user's answer to the same clause a rung is held to: the whole wire |
| 1152 | fence differs in one field and no other. |
| 1153 | |
| 1154 | *The patch the page half needs*, for whoever owns `www/js/daimond.js`. It goes |
| 1155 | in `egressAllowed`, beside the `req.tool === 'run'` branch and for the same |
| 1156 | reason that one is there: the "url" is a command line, which the URL work |
| 1157 | further down cannot read as an address. It remembers nothing -- the engine |
| 1158 | remembers, on the turn's own state, which is what makes one answer cover the |
| 1159 | turn and a new turn ask again. |
| 1160 | |
| 1161 | ```js |
| 1162 | if (req.tool === 'run_net') { |
| 1163 | var ncmd = String(req.url || ''); |
| 1164 | if (!ncmd.trim()) return 'deny'; |
| 1165 | var shownNet = ncmd.length > 300 ? (ncmd.slice(0, 300) + '…') : ncmd; |
| 1166 | var okNet = await confirmDialog( |
| 1167 | t('permmode.net_body', { cmd: shownNet, cwd: String(req.detail || '').slice(0, 300) }), |
| 1168 | t('permmode.net_ok'), |
| 1169 | { title: t('permmode.net_title'), danger: true }); |
| 1170 | // 'allow-net', NOT 'allow' — see below. Never remembered here. |
| 1171 | return okNet ? 'allow-net' : 'deny'; |
| 1172 | } |
| 1173 | ``` |
| 1174 | |
| 1175 | **A yes to this question is the word `allow-net` and nothing else**, and that |
| 1176 | is not fussiness. `egressHost` resolves its argument against `location.href`, |
| 1177 | so a command line -- which has no scheme -- resolves as a RELATIVE URL whose |
| 1178 | host is our own host, and the same-origin shortcut a few lines down answers |
| 1179 | `allow` outright. A `run_net` question reaching that shortcut, because the |
| 1180 | branch is missing or sits below it, would therefore hand every tainted turn |
| 1181 | the network with nobody asked. That is exactly how §7's egress guarantee was |
| 1182 | voided once already, by the same shortcut, for `web_search`. So the engine's |
| 1183 | edge (`wasm::web::egress_allowed_net`) takes a yes only in a word no other |
| 1184 | path through that function can say: an old bundle, a missing branch and the |
| 1185 | shortcut all say `allow` or `deny`, and all three mean no here. |
| 1186 | |
| 1187 | 3. **Grade the taint** — NOT DONE, and deliberately: the user did not |
| 1188 | authorise it, and it is theirs to authorise. Split |
| 1189 | `TurnState::tainted` into *outside content* (a fetched page, a message, an |
| 1190 | attachment) and *machine output* (a command's own stdout), and let |
| 1191 | `fence_spec` read the first while the envelope and `egress_check` continue to |
| 1192 | read both. A command's output is the machine's words, from programs the user |
| 1193 | installed, in the user's own workspace — a step removed from a stranger's |
| 1194 | words in a page, but not zero, since a dependency's build script prints |
| 1195 | whatever it likes. That is a genuine narrowing of a boundary and belongs to |
| 1196 | the user, not to an agent tidying a review. |
| 1197 | |
| 1198 | **1.14 Nothing checks that the hand's `root` is the app's workspace.** PLAUSIBLE. |
| 1199 | `Tool::run` joins workspace-relative paths onto whatever folder the hand reports. |
| 1200 | With an OPFS-only workspace, or an FSA folder different from the grant, the fence |
| 1201 | names paths on the machine that have nothing to do with the files the model just |
| 1202 | read. No folder-identity token exists on the wire. |
| 1203 | |
| 1204 | **CLOSED at both ends.** There now is a folder-identity token, the page compares |
| 1205 | it against the folder it has open, and a command is REFUSED where the two cannot |
| 1206 | be shown to be the same folder. |
| 1207 | |
| 1208 | *The hand's half.* It writes a random 32-hex token to |
| 1209 | `<root>/.daimond/workspace.id` and publishes it in `caps` as `ws:<token>`, beside |
| 1210 | the `root:` entry that was already there. The two answer different questions: |
| 1211 | `root:` says *where* the hand will work, and `ws:` is what lets the page find out |
| 1212 | whether that is the folder it is looking at — which a path alone cannot settle, |
| 1213 | because the File System Access API gives the page a handle and never a path. |
| 1214 | |
| 1215 | It lives inside `.daimond` deliberately. A fence always denies that directory, so |
| 1216 | a command cannot read the token, and a command that has been talked into helping |
| 1217 | cannot answer a challenge about a folder it is not in. The token is written once |
| 1218 | and kept, so it identifies the folder rather than the run, and a page that |
| 1219 | remembers it notices its workspace being swapped underneath it. Where no token |
| 1220 | can be established — an unwritable grant — the hand publishes `ws:unproven` |
| 1221 | rather than silence, because a page cannot tell silence from an older hand. A |
| 1222 | token can never be the word `unproven`, so one string settles both questions. |
| 1223 | `--report` prints the folder and its identity too, since both are configuration a |
| 1224 | person can get wrong and neither was visible anywhere before. |
| 1225 | |
| 1226 | Proved by `main::tests::the_page_is_told_which_folder_this_is`, five properties |
| 1227 | each of which fails against a deliberately broken version: the token on the wire |
| 1228 | is the token in the file, it survives a restart, two folders never share one, a |
| 1229 | planted line is replaced rather than published, and an unprovable folder is said |
| 1230 | to be unproven rather than given an identity anyway. Confirmed against the |
| 1231 | release binary over a pipe, whose `hello` carries |
| 1232 | `ws:f3540427d40b90dcffc6bb7a7e4feb90` matching the file on disk and the line |
| 1233 | `--report` prints. |
| 1234 | |
| 1235 | *The page's half, in `www/js/hand.js`.* The hand holds one of the two names and |
| 1236 | can only supply evidence; the comparison belongs where both names meet, and that |
| 1237 | is the page. Once per grant — and again whenever the folder changes — it opens |
| 1238 | `.daimond` and then `workspace.id` through the directory handle it holds, passing |
| 1239 | no `{create: true}` to either call, because a page that creates the file is a page |
| 1240 | that has proved nothing. It takes the first line that is neither blank nor a `#` |
| 1241 | comment and compares it with the `ws:` value, exactly, as strings. The file's four |
| 1242 | comment lines are why the token is not simply the first line: |
| 1243 | |
| 1244 | ``` |
| 1245 | # Daimond wrote this so that the browser and the machine hand can tell whether |
| 1246 | # they are talking about the same folder. It is not a secret and not a key. |
| 1247 | # Deleting it costs nothing: the next hand to start writes a new one, and the |
| 1248 | # page will ask you to confirm the folder again. |
| 1249 | 75111c6348d13219899a27405d5a769f |
| 1250 | ``` |
| 1251 | |
| 1252 | The verdict is reached in `status()`, which is the one door every route to a |
| 1253 | command already goes through: `Tool::run` reads it before composing a fence, |
| 1254 | `pty_request` reads it before opening a terminal, and the Terminal panel shows its |
| 1255 | `reason` where it will not open one. So a single refusal closes all of them, and |
| 1256 | what the model is handed is a sentence rather than the output of a command that |
| 1257 | ran somewhere else. |
| 1258 | |
| 1259 | **The four outcomes, and what the user reads.** |
| 1260 | |
| 1261 | - **Equal.** Nothing is said and nothing is shown. This is the ordinary case and |
| 1262 | it stays silent, or the check becomes a dialog people learn to dismiss. |
| 1263 | - **Different, or the file or the `.daimond` directory is missing through the |
| 1264 | handle.** Refused, with: *"The folder you opened in Daimond is not the folder |
| 1265 | the machine hand was told to work in, so a command would run against different |
| 1266 | files from the ones Daimond has been reading. Daimond has «the folder's name as |
| 1267 | the page knows it»; the hand has «the `root:` path». Fix the path in the hand's |
| 1268 | root.txt, or open the other folder here."* Both ends are named, because the two |
| 1269 | fixes are different and the user cannot otherwise tell which end is wrong. |
| 1270 | - **No folder at all — an OPFS-only workspace.** Its own case, and not one to |
| 1271 | skip for want of a handle: there is nothing to read the token through, so the |
| 1272 | check cannot pass, and a check that is skipped when it cannot pass is not a |
| 1273 | check. *"This workspace lives in the browser and not in a folder on this |
| 1274 | machine, so there is nothing for the hand's commands to run against. Open a |
| 1275 | folder for this workspace before using the machine hand."* |
| 1276 | - **`ws:unproven`.** Refused: *"The machine hand could not write its identity file |
| 1277 | into the folder it was granted, so the two ends cannot confirm they mean the |
| 1278 | same folder. The hand's own error output names the path that failed."* |
| 1279 | |
| 1280 | A fifth thing can happen which is not an outcome of the comparison at all: the |
| 1281 | page can fail to say what folder it has. That is refused in the same place and |
| 1282 | for the same reason — a check that could not be made has not passed — and the |
| 1283 | sentence says so, rather than blaming a folder. |
| 1284 | |
| 1285 | **Two decisions inside that, which a later reader should not undo.** |
| 1286 | |
| 1287 | *The folder's name is never compared.* Not as a fallback, not as a tie-break. Two |
| 1288 | projects called `site` on one machine is the ordinary case rather than the exotic |
| 1289 | one, and a check that passes for the wrong folder is worse than no check at all. |
| 1290 | Both directions are tested: a `site` whose identity is wrong is refused, and a |
| 1291 | folder called `moved` whose identity is right is allowed. |
| 1292 | |
| 1293 | *Silence passes.* A hand that publishes no `ws:` at all is an OLDER hand, not a |
| 1294 | mismatch. A page cannot tell an old hand from any other silent thing on that wire, |
| 1295 | and refusing silence would break every mock host permanently while telling the |
| 1296 | user about a folder they can do nothing about. It is a compatibility seam and it |
| 1297 | is recorded as one in the source: a page that meets such a hand is back where it |
| 1298 | was before this entry, and the proper place to close it is the protocol version, |
| 1299 | where "this hand is too old to serve" can be said once and plainly. |
| 1300 | |
| 1301 | The verdict is cached per grant AND per directory handle. Per grant alone was not |
| 1302 | enough: the user opens a different folder in the Workspace panel while the hand |
| 1303 | says nothing at all, so a remembered verdict would answer for a folder it had |
| 1304 | never read — and it would do so in the direction that runs the command. |
| 1305 | |
| 1306 | **How this is tested, given that no automated run can satisfy it.** A page holds a |
| 1307 | real folder only through `showDirectoryPicker()`, a native dialog no harness can |
| 1308 | answer, so every headless run has an OPFS workspace, which is the third outcome |
| 1309 | and a refusal. There is no configuration in which this check passes by accident. |
| 1310 | That is a fact about the browser and not a gap in the tests, and it is met in |
| 1311 | three places: |
| 1312 | |
| 1313 | - `dev/verify_wsident.mjs` is new and tests the refusal itself: 33 checks, of |
| 1314 | which 10 are proved against a deliberately broken `www/js/hand.js` served through |
| 1315 | a patch. The two folders are real `FileSystemDirectoryHandle`s taken from OPFS, |
| 1316 | each with a real `.daimond/workspace.id` in it, so `getDirectoryHandle`, the |
| 1317 | read, the comment-skipping and the compare all run for real; the only thing |
| 1318 | stood in for is the one thing a headless browser cannot have. The headline case |
| 1319 | is a hand granted folder A while the page holds folder B — both real folders, |
| 1320 | both perfectly good workspaces — refused, with the sentence naming `beta` and |
| 1321 | `/home/u/projects/alpha` in the same breath, and `pty_request` on the real wasm |
| 1322 | then refusing to open a terminal and passing that sentence on whole. |
| 1323 | - `dev/verify_ptyedge.mjs` asserts the refusal once against the REAL hand — the |
| 1324 | relay's own `status`, and the engine refusing on the strength of it — and then |
| 1325 | wraps `status` so that the verdict, and only the verdict, is stood in for. Its |
| 1326 | subject is the composition of a terminal request and a real pty on this machine, |
| 1327 | and the hand's own account of itself passes through untouched, so every fence it |
| 1328 | composes is still the real one. |
| 1329 | - `dev/verify_handreal.mjs` does the same at the other end of the chain: its first |
| 1330 | turn is now a REFUSAL, asserted on what the model was actually handed — a |
| 1331 | sentence, no nonce, no exit code — after which the same wrapper is installed and |
| 1332 | the file gets on with proving that a real process runs, that its real output |
| 1333 | reaches the daimon and that the kernel refuses what the fence denies. |
| 1334 | |
| 1335 | Both wrappers substitute the folder verdict and nothing else, and both say so at |
| 1336 | the point of use. `dev/verify_scope.mjs` had already taken the same route for the |
| 1337 | whole of `status`, for the same reason. |
| 1338 | |
| 1339 | **1.15 Release gate 1 is nobody's job.** CONFIRMED. `Tool::run` reads only `root` |
| 1340 | from `status()` and ignores `caps` and `paired`; `apply_fence` is an empty stub; |
| 1341 | `main.rs` has no message loop. Gate 4 is accurate, but gate 1 — refuse where the |
| 1342 | fence is not in force — is unmet on *both* ends, so nothing in the pipeline would |
| 1343 | refuse a command an unfenceable hand offered to run. |
| 1344 | |
| 1345 | **CLOSED, at both ends.** In the hand, `Desk::exec` refuses where the plan is not |
| 1346 | fenced, before the command is journalled or run. In the app, `Tool::run` reads the |
| 1347 | hand's `caps` as the array it is — it was read with a string extractor, which |
| 1348 | found nothing in `["fence:none"]`, so the refusal could not fire at all — and |
| 1349 | `fence_enforced` answers affirmatively or not at all: an absent or empty list is |
| 1350 | refused alongside `fence:none`, because a hand that will not say what it can |
| 1351 | enforce has not said that it can enforce anything. |
| 1352 | |
| 1353 | **1.16 Every disconnect is reported as "the hand is not installed".** CONFIRMED. |
| 1354 | `www/js/hand.js:175-180` rejects with `NO_HAND` on any `onDisconnect`, including |
| 1355 | after `started` and after chunks. A host that crashed, was killed, or blew |
| 1356 | Chrome's 1 MB cap makes the daimon tell the user to install software they already |
| 1357 | have. |
| 1358 | |
| 1359 | **CLOSED.** `HAND_GONE` is kept apart from `NO_HAND`, and which one the user |
| 1360 | reads depends on whether the hand ever greeted the page. Someone whose host |
| 1361 | crashed is no longer told to install software they already have. |
| 1362 | |
| 1363 | **1.17 A lost output tail is already detectable, and is not detected.** |
| 1364 | CONFIRMED. `Resp::Ended` carries `out_bytes`/`err_bytes`, `www/js/hand.js` |
| 1365 | forwards them, and `run_result` never reads them — three bytes were presented to |
| 1366 | the model as a 900 kB stream. This retires the earlier concern about `Ended` |
| 1367 | lacking a final `seq`: **the byte count already closes that hole**, it is simply |
| 1368 | unused. `run.gap` is likewise set and never surfaced. |
| 1369 | |
| 1370 | **CLOSED.** `run_result` compares the bytes that arrived against `out_bytes` and |
| 1371 | `err_bytes` and says so in the text the model reads — "some output did not |
| 1372 | arrive: N of M bytes" — and the relay's own account of a hole or a disconnect is |
| 1373 | shown separately from what the command printed. |
| 1374 | |
| 1375 | **1.18 Scoped workers cannot run a command with a default `cwd`.** CONFIRMED, |
| 1376 | latent. `set_diamond_scope` sets bounds but leaves `path_prefix` empty; |
| 1377 | `Tool::run` defaults `cwd` to `path_prefix`, and `may_read("")` is false under any |
| 1378 | allow-list, so the command is refused as "not in this Diamond's workspace". |
| 1379 | |
| 1380 | **CLOSED.** `path_prefix` is deliberately left empty for a scoped worker, whose |
| 1381 | model writes whole workspace-relative paths, and `default_cwd` takes the first |
| 1382 | `OnlyUnder` from the allow-list instead. A scoped worker with no `cwd` now lands |
| 1383 | inside its own Diamond rather than being refused. |
| 1384 | |
| 1385 | **1.19 `id` is `run-<argv[0]>` and is neither unique nor bounded.** CONFIRMED. |
| 1386 | Two concurrent `cargo` runs share an id; `Req::Signal` and the journal both key |
| 1387 | on it, so a cancel can reach the wrong run. Unbounded, it also drives 3.1. |
| 1388 | |
| 1389 | --- |
| 1390 | |
| 1391 | **CLOSED.** The identifier is `run-<n>-<name>`, where `n` is a per-turn counter |
| 1392 | and `name` is the program's basename filtered to a safe alphabet and cut at |
| 1393 | `RUN_ID_MAX`. Two concurrent `cargo` runs no longer share an id, and an unbounded |
| 1394 | one can no longer drive 3.1. |
| 1395 | |
| 1396 | **1.20 A Diamond's crystal agent could read and write another Diamond, using a |
| 1397 | path the model wrote.** FOUND while closing 1.12, and closed with it. |
| 1398 | |
| 1399 | *The reproduction.* `steer_inner` (`src/wasm/app.rs`) gives the crystal agent |
| 1400 | `no_write: Vec::new()` and relies entirely on `path_prefix` for its compartment. |
| 1401 | `Tool::scoped` joined that prefix to whatever the model wrote -- |
| 1402 | `fmt!("{}/{}", prefix, rel.trim_start_matches("./"))` -- and never normalised the |
| 1403 | result. `wasm::opfs::split_components` then resolved `..` lexically and refused |
| 1404 | only a climb above the OPFS ROOT. A Diamond is not the root. So a daimon |
| 1405 | steering its crystal and asking for `../beta/crystal.json` was handed |
| 1406 | `diamonds/alpha/../beta/crystal.json`, which landed at `diamonds/beta/crystal.json` |
| 1407 | -- another Diamond's private notes, inside OPFS, permitted, read and writable. |
| 1408 | (The file was `crystal.md` when this was found; the leaf is incidental to the |
| 1409 | escape, and the name is kept current so the shape can still be reproduced.) |
| 1410 | `guard` could not catch it: it tests the path as the model wrote it against the |
| 1411 | turn's bounds, and this turn's bounds are empty, which permits everything. |
| 1412 | |
| 1413 | Six more shapes did the same: a bare `..` and `./..` reached the `diamonds` |
| 1414 | directory itself, `notes/../../beta/x.md` and `../../beta/x.md` reached a |
| 1415 | sibling from the middle and from the start, `notes/../..` climbed two, and |
| 1416 | `../alpha2/x.md` reached a Diamond whose name merely *begins* with this one's -- |
| 1417 | which is the shape a string-prefix containment test waves through. Only |
| 1418 | `../../../../etc/passwd` was ever stopped, and by the wrong fence, with a |
| 1419 | message about the workspace rather than about this Diamond. |
| 1420 | |
| 1421 | **This is a model-controlled string leaving its compartment**, the same class as |
| 1422 | the empty-prefix escape and reachable by any daimon steering a Diamond. The |
| 1423 | instruction that produces the path may itself have come from a stranger's words. |
| 1424 | |
| 1425 | **CLOSED.** `Tool::scoped` normalises the join and refuses what is not under the |
| 1426 | prefix, in plain English that names the path, says it is outside this Diamond, |
| 1427 | and says nothing was read, written or run. The containment test is `under`, |
| 1428 | which compares whole segments, so `../alpha2` is the different Diamond it |
| 1429 | actually is. Separators are unified first, so a backslash is not a containment |
| 1430 | test that means one thing on one platform and another elsewhere. An absolute |
| 1431 | path stays relative to the Diamond and cannot escape it, which is what |
| 1432 | `Workspace::resolve` does natively. A turn with no prefix -- the user's own |
| 1433 | workspace agent -- is untouched, and bounded by the OPFS root as it always was. |
| 1434 | |
| 1435 | The doc comment records why this cannot be done with `may_read` instead, because |
| 1436 | that is the repair the next person will reach for: that door tests the path as |
| 1437 | the MODEL wrote it, so a crystal agent asking for `crystal.json` would be measured |
| 1438 | against an allow-list of `diamonds/<id>` and refused for its ordinary work. It |
| 1439 | is 1.18's collision from the other side -- a prefix and an allow-list are two |
| 1440 | ways of saying where a turn lives, and a path can be checked against one, not |
| 1441 | both. |
| 1442 | |
| 1443 | Three tests, and the section counts itself: eighteen path shapes, of which nine |
| 1444 | must now be refused, and three counters assert that seven of those nine landed |
| 1445 | in another compartment before, one was stopped only by the OPFS root jail, and |
| 1446 | one was harmless. The ordinary paths are pinned in the same table, so a fix that |
| 1447 | simply refused everything would fail here. They run NATIVELY against the same |
| 1448 | function the browser calls -- it is compiled for `test` as well as for `wasm32`, |
| 1449 | so there is one implementation rather than two that drift -- but nothing in them |
| 1450 | reaches the OPFS edge, which no native test can; where the old path landed is |
| 1451 | shown through a model of `split_components` and not through the edge itself. |
| 1452 | |
| 1453 | **Still open, and latent: `path_prefix` means two different things.** In the |
| 1454 | browser it confines every file path; natively the file tools ignore it entirely |
| 1455 | and jail on the workspace root instead, so it reaches only `default_cwd`. Every |
| 1456 | native context sets it empty today, so nothing is wrong now -- and the day a |
| 1457 | native turn carries a prefix, it will confine nothing. Wiring `scoped` into the |
| 1458 | native transport is a no-op on today's values and would make the two agree. |
| 1459 | |
| 1460 | **1.21 A verifier's trustworthiness is reported and never enforced, so a daimon can |
| 1461 | run its own script outside the fence.** CONFIRMED 2026-08-25, found while measuring |
| 1462 | B6 rather than while looking for this. |
| 1463 | |
| 1464 | `verify` runs a script OUTSIDE the command fence deliberately, and `verify.rs`'s |
| 1465 | own first section says why: *the fence exists to contain a command a MODEL wrote, |
| 1466 | and a verifier is tracked repository code.* The report says it to the model |
| 1467 | verbatim, on every run — "Run OUTSIDE the command fence, because a verifier is |
| 1468 | tracked repository code and not a command anybody's model wrote." |
| 1469 | |
| 1470 | Trackedness is asked of git and not assumed — `tracked_by_git`, |
| 1471 | `hand/src/verify.rs:434`, `git ls-files --error-unmatch` — and the answer is used |
| 1472 | for one thing: a phrase in the report's first line (`Tracked::phrase`, its only |
| 1473 | call site, `:1298`). Nothing anywhere refuses `Tracked::No`. And the file it is |
| 1474 | asked about is `<granted root>/dev/verify_<name>.mjs`, inside the writable root |
| 1475 | that every fenced command in the turn already has. |
| 1476 | |
| 1477 | *The reproduction.* A root with no git in it, an untracked `dev/verify_probe.mjs` |
| 1478 | whose whole body writes a file OUTSIDE that root, then `resolve` and `conduct` |
| 1479 | from this crate's own public API: |
| 1480 | |
| 1481 | ``` |
| 1482 | trackedness reported: NOT TRACKED -- a clean checkout of this commit would not have it |
| 1483 | marker outside the granted root exists: true |
| 1484 | ``` |
| 1485 | |
| 1486 | Accepted, run with no fence, and it wrote where no fence would have let it. |
| 1487 | |
| 1488 | **Reachable from an ordinary daimon's turn, and every step is one the turn already |
| 1489 | takes.** `Tool::Verify` is in `Tool::daimon()` — `src/tools.rs:16463` asserts it in |
| 1490 | those words. The granted root is writable, so `file_write` or any fenced command |
| 1491 | puts the file in `dev/`. The `verify:dev` capability is computed by looking at |
| 1492 | whether the root holds any `dev/verify_*.mjs` (`verify::available`), which this |
| 1493 | repository satisfies hundreds of times over, so the daimon's own new file is not |
| 1494 | even needed to unlock the door. The one policy the door does carry — a dispatched |
| 1495 | worker is refused, because a worker is the turn with nobody watching — does not |
| 1496 | reach this, since the daimon is not a worker. |
| 1497 | |
| 1498 | **And enforcing what is already measured would narrow it without closing it.** |
| 1499 | `git ls-files --error-unmatch` answers about the INDEX, not about the file: with |
| 1500 | one comment appended to `dev/verify_graph.mjs` it still exits 0, while `git diff |
| 1501 | --quiet` on the same file exits 1. So a daimon that edits a tracked verifier gets |
| 1502 | the same unfenced run, and the report calls it "tracked". The claim the door |
| 1503 | rests on is *this is the repository's code*, and the question that answers it is |
| 1504 | whether the file differs from the commit — not whether git has heard of it. |
| 1505 | |
| 1506 | **OPEN.** The cheapest honest close is a refusal on anything but "unmodified since |
| 1507 | the commit", written as a sentence the model can act on: name the file, say a |
| 1508 | verifier it wrote or changed is not repository code, and say the person can run it |
| 1509 | themselves. Two prices come with it and both look right: a lane writing a NEW |
| 1510 | verifier has to `git add` it before a daimon may run it, which is the same |
| 1511 | discipline `dev/gate.sh` already imposes by building its tree from a commit; and |
| 1512 | a daimon can no longer test a verifier it has just improved, which is a real loss |
| 1513 | and is the thing to weigh. Whoever takes it should keep the probe above and assert |
| 1514 | it goes red first: a check on this that has never been seen to fail proves nothing. |
| 1515 | |
| 1516 | **CLOSED 2026-08-25, and not the way the paragraph above proposed.** The gate is |
| 1517 | `verify::unfenced_refusal`, `hand/src/verify.rs`, and it asks |
| 1518 | `verify::provenance` — three questions of git in order: is there a repository |
| 1519 | here at all, is the path a symlink, and does `git diff --quiet HEAD` say the |
| 1520 | working file is the commit's. `conduct` passes it before the journal and before |
| 1521 | the first process, and `once` asks again immediately before every spawn, because |
| 1522 | the sequence runs for minutes and a command an earlier turn left in the |
| 1523 | background can rewrite the file inside it. The gate is in `verify.rs` and not in |
| 1524 | `main.rs` deliberately: the reproduction above went in through this crate's |
| 1525 | public API, and a gate in the dispatcher is a gate it would have walked past |
| 1526 | again. |
| 1527 | |
| 1528 | **Three things the entry above did not say, each of which changed the design.** |
| 1529 | |
| 1530 | *The strict refusal it recommends does not close it either.* `.git` is inside the |
| 1531 | granted root and no fence denies it — `src/tools.rs` says so in its own words |
| 1532 | where it explains why a Daimond push injects `core.hooksPath`, `protocol.allow` |
| 1533 | and an empty `credential.helper`: *the model can write `.git/config`*. So a |
| 1534 | daimon can commit its own verifier with one `run` call and then be told, truly, |
| 1535 | that the bytes are the commit's. Every check available on this side of the line |
| 1536 | is one `run` call from being satisfied; that is not an argument for leaving the |
| 1537 | door open, but it is the reason this is a narrowing and not a close. What |
| 1538 | changes is that the act is a commit in the repository's history rather than a |
| 1539 | `file_write` nobody sees, and that the report's first line stops being false. |
| 1540 | |
| 1541 | *A committed symlink is the same hole wearing 1.1's clothes.* `resolve` asked |
| 1542 | `std::fs::metadata`, which follows a link, so `dev/verify_x.mjs -> ../notes/x.mjs` |
| 1543 | committed once would run whatever the model last wrote into `notes/x.mjs`, with |
| 1544 | `git diff` reporting no change because the link's target string had not moved. |
| 1545 | `provenance` asks `symlink_metadata` and refuses a link outright. |
| 1546 | |
| 1547 | *A folder with no git in it was the permissive default.* That is the state the |
| 1548 | reproduction was run in, and `Tracked::Unknown` — reported, never refused — is |
| 1549 | what let it through. It refuses now: a granted folder that cannot say whose code |
| 1550 | a file is cannot vouch for it. |
| 1551 | |
| 1552 | **The two prices the entry above weighed are paid differently, and this is the |
| 1553 | part that bears on the objective.** A refusal a model cannot converge on costs |
| 1554 | the run anyway, so the refusal hands over the fenced route in the same breath: |
| 1555 | |
| 1556 | > Refused: dev/verify_about.mjs is not this repository's committed code — it is |
| 1557 | > tracked, and the working copy is not what the commit holds — staging a change |
| 1558 | > is not committing it. Nothing was run. […] RUN IT YOURSELF INSTEAD: 'run' with |
| 1559 | > `["node","dev/verify_about.mjs"]` runs this same file inside the fence every |
| 1560 | > command gets, which is all a verifier that reads the tree needs. It is not |
| 1561 | > enough for one that drives a browser or opens a port — the fence refuses both, |
| 1562 | > so a failure there is the fence and not your script. To get the unfenced run, |
| 1563 | > commit the file and ask again, or ask the person to run it themselves. |
| 1564 | |
| 1565 | So a daimon that has just improved a verifier is not stopped; it is moved into |
| 1566 | the compartment, which is where a script a model wrote belongs. **That covers |
| 1567 | 253 of this tree's 282 verifiers**, which is every one that does not name |
| 1568 | playwright, chromium, webkit or firefox — measured, not estimated. The 29 that |
| 1569 | drive a browser cannot run fenced at all, for the reason this module exists, and |
| 1570 | for those the daimon has two moves it can make and a person it can ask. |
| 1571 | |
| 1572 | **Proved by making it happen, in both directions.** The reproduction is now nine |
| 1573 | tests in `verify.rs`, run through `resolve` and `conduct` exactly as the original |
| 1574 | probe was, over a fixture verifier whose entire body writes a file OUTSIDE the |
| 1575 | granted root — so the marker's existence is the measurement and not a proxy for |
| 1576 | it. Untracked, tracked-and-edited, tracked-edited-and-staged, no-repository and |
| 1577 | committed-symlink all wrote that marker before the gate and all are refused after |
| 1578 | it. One test keeps the rest honest: a verifier that IS the commit must still |
| 1579 | write its marker, so a gate that refused everything would go red rather than |
| 1580 | reading as a pass. One asserts the refusal's four working parts by meaning — the |
| 1581 | argv, the tool that takes it, what commits it, and what fenced running cannot do |
| 1582 | — since a refusal a model cannot converge on costs the run anyway. And one aims |
| 1583 | at the mid-sequence re-ask, with a committed verifier that appends to its own |
| 1584 | source: the clean run is the commit's, the break run is not, and the report says |
| 1585 | `NEVER RAN` with the reason rather than measuring a file that changed underneath |
| 1586 | it. Every one was run against a deliberately broken gate first — six mutations, |
| 1587 | each reddening a different subset, none reddening all — because a check that has |
| 1588 | never been seen to fail proves nothing, which is this document's own rule. The |
| 1589 | crate's suite went 268 → 277, every compiled test executed, by `dev/testcount.mjs`. |
| 1590 | |
| 1591 | **What is left, and it is the owner's.** Two ways to close the commit path, one |
| 1592 | clause each. Take `.git` out of every fence — cheap to state, and it stops a |
| 1593 | daimon using git at all, which is most of what self-development is. Or bind |
| 1594 | consent to the bytes, the way `social_send` already binds publishing consent to |
| 1595 | exactly what would go out and never remembers it — right in shape, and it means |
| 1596 | a daimon working while nobody is awake cannot verify its own new verifier, which |
| 1597 | is the objective this app is for. **Recommendation: neither, yet.** The gate |
| 1598 | above turns an invisible `file_write` into a visible commit, which is most of the |
| 1599 | value, and both closes cost more than the residual is worth until the hand ships |
| 1600 | to somebody who is not the author. |
| 1601 | |
| 1602 | ## What was verified as genuinely sound |
| 1603 | |
| 1604 | Worth recording, so a later reader does not re-litigate what has already been |
| 1605 | checked: |
| 1606 | |
| 1607 | - **`argv`, never a shell string, holds.** Exactly two `Command::new` calls in |
| 1608 | `exec.rs`, no shell anywhere, metacharacters pass through literally. The |
| 1609 | `/bin/kill` argv is fixed strings plus a pgid from `child.id()`; no caller |
| 1610 | value reaches it. |
| 1611 | - **The wire seam agrees end to end.** The exec JSON `Tool::run` composes decodes |
| 1612 | byte-for-byte in the hand, including `"stdin":null` → `None`, `"env":[]`, |
| 1613 | `timeout_ms` as `u64` and `capture:"both"`. Field names and enum spellings |
| 1614 | agree across `tools.rs` → `www/js/hand.js` → `ext/hand.js` → `codec.rs`, and |
| 1615 | the result keys `refused`/`stdout`/`stderr`/`exit`/`timed_out` all exist on |
| 1616 | both sides. |
| 1617 | - **No JSON breakout.** `argv`, `cwd`, `stdin`, `id` and every fence path go |
| 1618 | through `json_escape`, which escapes `"`, `\` and all C0; `timeout_ms` is a |
| 1619 | parsed `u64`; and `extract_json_string`'s prefix test makes a `"refused":` |
| 1620 | planted in a command's own stdout unfindable. Attempted and failed. |
| 1621 | - **The journal's cryptographic core is correct.** The hash was reproduced with |
| 1622 | coreutils `sed | sha256sum` over 200 entries containing emoji, DEL, U+2029, |
| 1623 | quotes, backslashes, tabs, CJK and 300-element argv, with zero mismatches. The |
| 1624 | `,"entry":"` boundary is *not* attacker-controlled: every field is located |
| 1625 | positionally from the end, never by search. Canonicalisation is RFC |
| 1626 | 8785-correct on the values used. Tamper detection *within a present file* is |
| 1627 | genuinely good — edited, rehashed, deleted, reordered, duplicated and torn |
| 1628 | lines were all caught and correctly located. |
| 1629 | - **`check_fence_at` survived every attack brought against it**, including a |
| 1630 | symlink onto the journal, `..` spellings through a symlink, relative roots and |
| 1631 | an `rw` parent with a `deny` naming the journal. |
| 1632 | - **Native-messaging framing is correct** — 4-byte native-endian, `FRAME_MAX` |
| 1633 | below 1 MiB with slack, prefix counted into the total. `MAX_DEPTH` is really |
| 1634 | enforced: 100,000 open brackets return a named fault in milliseconds. Every |
| 1635 | hostile input tried returned a named `Fault`; nothing panicked, hung or |
| 1636 | allocated unboundedly. |
| 1637 | - **Deny-inside-rw carving of real, non-symlink children works** — the denied |
| 1638 | file is unreadable and unwritable and siblings are intact. `..` resolution is |
| 1639 | correct component-wise. The refusal default on Linux never runs unfenced for a |
| 1640 | malformed spec. ABI is never silently rounded up, and partial enforcement is a |
| 1641 | hard error. |
| 1642 | - **UTF-8 chunk handling is subtle and right** — split characters are rejoined, |
| 1643 | and text is re-bounded after `from_utf8_lossy` trebles invalid bytes. |
| 1644 | - **Revocation reaches a run in flight**, with the right sentence. `dismissed` is |
| 1645 | kept distinct from `declined` and neither is treated as allowed. `bye` then |
| 1646 | disconnect leaves no orphan. |
| 1647 | - Style is compliant throughout: no `unwrap()`, no `unsafe`, no `?`. |
| 1648 | |
| 1649 | ## What this means |
| 1650 | |
| 1651 | The two claims the product would make about the hand are the two that failed. |
| 1652 | That is not a coincidence: they are the claims that require an adversary to be |
| 1653 | wrong about something, and the rest of the code only requires the author to be |
| 1654 | right. |
| 1655 | |
| 1656 | Three consequences follow. |
| 1657 | |
| 1658 | 1. **`Fence::holes()` was the most valuable thing in the fence, and it was |
| 1659 | incomplete.** It should list 1.1, 1.2 and the true severity of 1.3, and the |
| 1660 | capability list should stop implying a compartment that metadata syscalls walk |
| 1661 | straight through. |
| 1662 | 2. **The consent wording cannot ship as written**, per gate 3: the sentence is |
| 1663 | not what needs editing. On this kernel, "only inside the folders the workspace |
| 1664 | already allows" is false in at least three ways. |
| 1665 | 3. **A compartment on a kernel below ABI 9, with a reachable user session bus, is |
| 1666 | not a compartment.** That is a statement about Landlock, not about this code, |
| 1667 | and no amount of care in `fence.rs` changes it. Either the bus is made |
| 1668 | unreachable, or the claim is narrowed to what the fence genuinely delivers. |
| 1669 | |
| 1670 | --- |
| 1671 | |
| 1672 | ## And where those three stand |
| 1673 | |
| 1674 | Written with the state lines above, against the same three consequences. |
| 1675 | |
| 1676 | 1. **`Fence::holes()` was made to say all of it, and then to stop saying two of |
| 1677 | them.** It named the symlink race, the metadata syscalls, the true severity of |
| 1678 | the session bus, the breadth of the system base, the leak of existence through |
| 1679 | `stat`, ungoverned UDP, descriptors opened before the fence, hard links, and |
| 1680 | the one place every command may write that the workspace did not name. Two of |
| 1681 | those are now closed by the syscall filter rather than by Landlock, so |
| 1682 | `holes()` is told what the filter refuses and drops exactly those two — a list |
| 1683 | that goes on describing a shut hole is as dishonest as one that leaves an open |
| 1684 | one out. `--report` prints the composed list from both layers, and the test |
| 1685 | asserts entries by keyword in *both* directions: the two must disappear when |
| 1686 | the filter is in force, and nothing else may. |
| 1687 | 2. **The consent wording can be revisited, and has not been rewritten here.** |
| 1688 | "Only inside the folders the workspace already allows" is now true of this |
| 1689 | kernel with the exceptions `--report` prints: the private temporary directory, |
| 1690 | the read-only system base, metadata reads, and the residue of the filter's own |
| 1691 | trade — `chmod 644` and the utime family, which stay so that `cargo` can unpack |
| 1692 | a crate. Gate 3 says the sentence is not what needs editing, and the sentence |
| 1693 | still has not been edited; what changed is that the claim behind it is now |
| 1694 | nearly the one the code keeps. 1.9 is what stands between "nearly" and "is". |
| 1695 | 3. **The bus was made unreachable.** The filter refuses `AF_UNIX` for every |
| 1696 | command, not only for one that was denied the network — the change that |
| 1697 | decision needed is at 1.3. So the honest description of the Machine tier is a |
| 1698 | fence that holds against a model and against a command that has been talked |
| 1699 | into something, on a kernel two ABI levels below the one that would make |
| 1700 | Landlock sufficient on its own. What it is not is a syscall sandbox: the filter |
| 1701 | is a deny-list, and it says so in `holes()`. |