oxedyne/daimond/examples/devcycle_probe.rs
77.3 KiB, 1 run
created by r2519314175:847, which is this file's identity for as long as the history lasts, whatever it is later renamed to
download · who wrote it · its history
| 1 | //! Can a daimon do a day's work in this app? The instrument for the one question |
| 2 | //! nothing else here asks. |
| 3 | //! |
| 4 | //! ## Why it exists |
| 5 | //! |
| 6 | //! On 2026-08-23 the owner gave a daimon one small, exactly specified change to make in |
| 7 | //! this repository. It read a 19 KB memory index nobody asked for, was told by this app |
| 8 | //! that this app's own 1.6 MB UI source was a binary file, fell back to reading it through |
| 9 | //! `run sed`, spent two turns on malformed `grep` calls, and had consumed about 78 KB of |
| 10 | //! context before its first edit. He stopped the turn. |
| 11 | //! |
| 12 | //! **690 library tests and roughly 270 gate checks were green throughout, and every one of |
| 13 | //! them was right.** That is not an oversight in any of them. It is a consequence of the |
| 14 | //! shape they share: each asserts that a named thing does a named thing, and every named |
| 15 | //! thing was working. What failed was the LOOP -- a real model, the real tools, the real |
| 16 | //! tree, and the marks the owner actually uses, all in the room together. Nothing in this |
| 17 | //! repository puts them in a room. |
| 18 | //! |
| 19 | //! So this is not another verifier. It is a probe, for the same reason |
| 20 | //! `dev/probe_details.sh` is one: a scripted mock cannot answer "can a model do this with |
| 21 | //! these tools", because the mock is the part that would have to be intelligent. |
| 22 | //! |
| 23 | //! ## What it measures, and the third one is the point |
| 24 | //! |
| 25 | //! For each task, three things, and a task passes only on all three: |
| 26 | //! |
| 27 | //! 1. **It got there.** The check below the task says what "there" is, in terms of the |
| 28 | //! tree or the answer, never in terms of what the model said about itself. |
| 29 | //! 2. **It did not flail.** No refused call, no failed call. A turn that reaches the |
| 30 | //! right answer through three refusals found a fault and worked around it, which is |
| 31 | //! what the owner has been doing by hand and what this exists to stop. |
| 32 | //! 3. **It stayed inside a budget** -- tool calls, and bytes of tool output taken into |
| 33 | //! context. This is the one that catches the 23rd. Every fault that day was |
| 34 | //! survivable on its own; what made the turn unusable was the total. A harness that |
| 35 | //! asserted only (1) would have passed the run the owner killed. |
| 36 | //! |
| 37 | //! The budget is in BYTES OF TOOL OUTPUT rather than in tokens billed, because that is the |
| 38 | //! quantity the app controls and the one that compounds: a tool result enters the |
| 39 | //! conversation once and is re-sent on every later round of the turn. Reading 39 KB to |
| 40 | //! learn how to invoke a script is not one mistake, it is one mistake times the number of |
| 41 | //! rounds still to come. |
| 42 | //! |
| 43 | //! ## Run it at the owner's marks, never at a fixture |
| 44 | //! |
| 45 | //! `MARK` below is a real directory in the real tree, and two of the tasks are marked at |
| 46 | //! `~/usr/code` -- 590,000 files -- because that is what he marked. A tidy fixture cannot |
| 47 | //! see either of the faults that matter: the binary refusal needs the actual 1.6 MB file |
| 48 | //! with its actual NUL at byte 1,113,118, and the walk cap needs a reach large enough for |
| 49 | //! `WALK_ENTRIES_MAX` to bite. Both would be invisible in a 200-line fixture, which is |
| 50 | //! exactly how they survived this long. |
| 51 | //! |
| 52 | //! ## It spends, and it says so first |
| 53 | //! |
| 54 | //! One turn per task against a real provider. The worst case is printed before the first |
| 55 | //! call, and `PROBE_YES=1` is required to skip the pause. |
| 56 | //! |
| 57 | //! ```bash |
| 58 | //! PROBE_SELFTEST=1 cargo run --example devcycle_probe -p oxedyne_daimond # free |
| 59 | //! DAIMOND_PROBE_KEY=sk-or-v1-... cargo run --example devcycle_probe -p oxedyne_daimond |
| 60 | //! ``` |
| 61 | //! |
| 62 | //! **`PROBE_SELFTEST=1` proves the harness before any money is spent** and is not optional |
| 63 | //! courtesy: this file's checks, its git reset and its budget arithmetic are as capable of |
| 64 | //! being wrong as anything they measure, and `probe_details.sh` has a paragraph in its own |
| 65 | //! header about the day its classifier put a reply in the wrong bucket and printed a tally |
| 66 | //! that was quietly false. The self-test runs every check twice -- against a tree where |
| 67 | //! the task is done, and against one where it is not -- and fails unless each check answers |
| 68 | //! differently. A check that cannot go red is a finding, not a fixture problem. |
| 69 | //! |
| 70 | //! `DAIMOND_PROBE_TASKS=bigfile,bigmark` runs a subset. |
| 71 | //! |
| 72 | //! ## The first live run, 2026-08-23, `anthropic/claude-haiku-4.5` |
| 73 | //! |
| 74 | //! ```text |
| 75 | //! task verdict calls bytes ref fail secs worst read |
| 76 | //! bigfile pass 3 87296 0 0 10.9 file_read 80016 |
| 77 | //! bigmark pass 3 3341 0 0 89.9 file_search 1752 |
| 78 | //! locales pass 10 2919 0 0 22.9 file_search 2311 |
| 79 | //! ranit pass 1 1617 0 0 4.0 shell 1617 |
| 80 | //! TOTAL 17 call(s), 95173 byte(s) of tool output, 128s. |
| 81 | //! ``` |
| 82 | //! |
| 83 | //! **All four passed on correctness, and two of them are the reason this file exists.** |
| 84 | //! `bigfile` took 80,016 bytes in a single `file_read` to learn a line number, and |
| 85 | //! `bigmark` spent 89.9 seconds walking 590,000 files for an answer it got right. The |
| 86 | //! budgets above were set from these figures afterwards -- at what each task is WORTH, not |
| 87 | //! at what it cost -- so both now fail, and will go green when the app is fixed rather than |
| 88 | //! when the numbers are edited. **A budget set above an observed figure measures nothing.** |
| 89 | //! |
| 90 | //! The run before this one failed three of four with every file tool REFUSED, and that was |
| 91 | //! this harness: it rooted the workspace at the mark, so `diamond_bounds` saw `"."`, |
| 92 | //! normalised it away and answered `Bound::Nowhere`. The model said "I have no workspace |
| 93 | //! attached", which was true. Caught by check (2) -- a refusal is never a pass here -- on |
| 94 | //! the instrument's first outing, which is the argument for check (2). |
| 95 | //! |
| 96 | //! ## After the `file_read` peek, same evening |
| 97 | //! |
| 98 | //! ```text |
| 99 | //! bigfile pass 3 19478 0 0 8.8 file_read 12198 |
| 100 | //! bigmark pass 3 2864 0 0 21.6 file_search 1419 |
| 101 | //! locales FAIL 20 14477 0 1 20.7 file_read 6745 |
| 102 | //! ranit pass 1 1617 0 0 3.5 shell 1617 |
| 103 | //! ``` |
| 104 | //! |
| 105 | //! `bigfile` fell from 87,296 bytes to 19,478, its worst single read from 80,016 to 12,198, |
| 106 | //! and it now passes a budget set at what the task is worth. That is what the peek bought. |
| 107 | //! |
| 108 | //! **Two things in that table are NOT findings, and saying so is the point of keeping it.** |
| 109 | //! `bigmark` at 21.6 s against the first run's 89.9 s owes most of the difference to a warm |
| 110 | //! page cache, not to any change: nothing was done to the walk, so the earlier figure was |
| 111 | //! partly an artefact and the claim built on it was worth less than it looked. And |
| 112 | //! `locales`'s one failed call did not reproduce -- a re-run passed at 18 calls and 12,747 |
| 113 | //! bytes -- so it is model variance rather than a regression. |
| 114 | //! |
| 115 | //! ## After the marks became the default starting point, same evening |
| 116 | //! |
| 117 | //! ```text |
| 118 | //! bigfile pass 3 19478 0 0 8.0 file_read 12198 |
| 119 | //! bigmark pass 1 1419 0 0 16.8 file_search 1419 |
| 120 | //! locales pass 10 2919 0 0 23.3 file_search 2311 |
| 121 | //! ranit pass 1 1617 0 0 3.5 shell 1617 |
| 122 | //! TOTAL 15 call(s), 25433 byte(s) of tool output, 52s. |
| 123 | //! ``` |
| 124 | //! |
| 125 | //! Against the first run: **95,173 bytes to 25,433, and 17 tool calls to 15.** `bigmark` is |
| 126 | //! the structural one -- three calls to ONE, because a bare search now begins at the mark |
| 127 | //! instead of at the workspace root above it. Between those two runs it also spent a run |
| 128 | //! reporting that `updateSpend` "does not exist in" a file holding ten of them, which is what |
| 129 | //! searching the wrong tree looks like from the inside. |
| 130 | //! |
| 131 | //! Two faults were found by the probe rather than by a person, and both are fixed: a walk |
| 132 | //! starting above the marks, and `file_read` answering a directory with the operating system's |
| 133 | //! "Is a directory" wrapped in two error frames. The second was invisible until this file |
| 134 | //! learned to NAME the call that failed rather than count it. |
| 135 | //! |
| 136 | //! **Which exposes the instrument's own weakness: n = 1 per task.** A single run can neither |
| 137 | //! confirm a fix nor convict a regression on the noisy tasks, and `locales` has now returned |
| 138 | //! 10, 20 and 18 calls for the same brief. Read a single column as a signal only where the |
| 139 | //! change is large, as `bigfile`'s was. The fix is repeats, and it is not built. |
| 140 | //! |
| 141 | //! ## 2026-08-24: four more tasks, and what this transport cannot see |
| 142 | //! |
| 143 | //! Five faults were carried over from the night of the 23rd and the day after, every one of |
| 144 | //! them found by the owner or by a daimon failing in front of him and none of them by an |
| 145 | //! instrument. Four became tasks -- `commit`, `parses`, `alias`, `world`. The fifth did not, |
| 146 | //! and the reason it did not is the largest thing this file has to say about itself. |
| 147 | //! |
| 148 | //! **This probe is a NATIVE binary, and three of those five faults live in the browser build.** |
| 149 | //! `Tool::Run` and `Tool::Verify` are `#[cfg(target_arch = "wasm32")]`; the native `execute` |
| 150 | //! answers both with `Unimplemented`, and `Tool::defaults()` does not offer them at all. So the |
| 151 | //! hand's fence, the environment a granted [`Toolkit`] hands a command, and the Diamond store |
| 152 | //! are all invisible from here, and a task written against one of them would be GREEN while the |
| 153 | //! product stayed broken -- which is worse than no task, because it is a false all-clear. |
| 154 | //! |
| 155 | //! What each of the five could honestly become: |
| 156 | //! |
| 157 | //! * **`commit`** -- a daimon could not see `.git` at all: `ls -la` showed none, `git status` |
| 158 | //! walked every parent to `/`. That is the hand's fence and it is not reachable here. What |
| 159 | //! IS reachable is the capability itself, which nothing measured: can a daimon change a |
| 160 | //! line and record the change. It runs in a repository of its own under `target/`, for the |
| 161 | //! reason [`check_commit`] gives at length. |
| 162 | //! * **`parses`** -- the fault is reaching past the repository's own checking machinery and |
| 163 | //! rebuilding it. A daimon spent FORTY-ONE tool calls trying to run `verify_vocabulary` |
| 164 | //! through `run`, and could not have succeeded: that verifier needs a dev server and a real |
| 165 | //! browser, which is why `verify` exists. It cannot be run from here either. The same |
| 166 | //! shape at a hundredth of the cost is `node --check`, which EXITS 0 ON A FILE WITH A SYNTAX |
| 167 | //! ERROR IN IT, against `dev/jscheck.sh`, which does not. The wrong route answers wrongly, |
| 168 | //! so this one is caught on the verdict and not only on the budget. |
| 169 | //! * **`alias`** -- fully visible here, and the only one of the five that is. `file_search` |
| 170 | //! is the same code in both builds. |
| 171 | //! * **`world`** -- `run` clears the environment and passes no `HOME` unless the Git toolkit |
| 172 | //! was granted, so scripts under `dev/` die before they print. Native `sh -c` inherits the |
| 173 | //! lot, so the fault cannot reproduce. The task is kept because its BUDGET still bites: |
| 174 | //! `dev/world.sh` is 11,419 bytes and its answer is 330, so reading it instead of running |
| 175 | //! it fails here today, hand or no hand. |
| 176 | //! * **The store boundary** -- `diamonds/<id>`, `chats/<id>/work` and `mail/<address>` are |
| 177 | //! browser storage; the file tools reach them and a command never can, and a daimon tried |
| 178 | //! three times in one turn to `cp` into its own Diamond folder. **No task was written.** |
| 179 | //! `is_store_path` and the OPFS root are wasm-only, and the native `shell` is unfenced, so |
| 180 | //! the wrong route -- the `cp` -- SUCCEEDS here. Any check this file could write would be |
| 181 | //! satisfied by it, the self-test would pass, and the instrument would report a capability |
| 182 | //! the product does not have. It is left out deliberately, and the way to get it is to give |
| 183 | //! this probe a way to speak to the hand, which is the same change that would let `commit` |
| 184 | //! and `world` measure their real faults. |
| 185 | //! |
| 186 | //! **The self-test found a rotted task on its first run with these in.** `check_locales` asked |
| 187 | //! for `spend.period_day`, which landed in `1e294c1`; from that commit the check was green |
| 188 | //! against a tree where nothing had been asked or done. The key is now chosen from the tree at |
| 189 | //! run time -- see [`PERIOD_KEYS`] -- and the first attempt at that was wrong in an instructive |
| 190 | //! way, which [`period_key`] records. |
| 191 | //! |
| 192 | //! ## The first run with all eight, 2026-08-24, `anthropic/claude-haiku-4.5` |
| 193 | //! |
| 194 | //! ```text |
| 195 | //! task verdict calls bytes ref fail secs worst read |
| 196 | //! bigfile FAIL 3 12387 0 0 9.0 file_read 12198 |
| 197 | //! bigmark FAIL 1 1752 0 0 46.0 file_search 1752 |
| 198 | //! locales FAIL 19 7674 0 1 17.0 file_read 817 |
| 199 | //! ranit pass 1 1617 0 0 3.2 shell 1617 |
| 200 | //! commit pass 3 556 0 0 6.3 shell 417 |
| 201 | //! parses FAIL 15 100744 0 0 66.7 file_search 34526 |
| 202 | //! alias FAIL 4 32250 0 0 11.8 file_read 24901 |
| 203 | //! world pass 1 570 0 0 4.1 shell 570 |
| 204 | //! TOTAL 47 call(s), 157550 byte(s) of tool output, 164s. |
| 205 | //! ``` |
| 206 | //! |
| 207 | //! **`parses` and `alias` both got the right answer and both failed, and that pair is the whole |
| 208 | //! argument for check (3).** `parses` found the verdict after FIFTEEN calls and 100,744 bytes |
| 209 | //! against a budget of five and 8,000 -- one `file_search` alone returned 34,526 -- which is the |
| 210 | //! forty-one-call shape of the 23rd, reproduced by an instrument for the first time rather than |
| 211 | //! watched over somebody's shoulder. `alias` named both callers, the aliased one included, and |
| 212 | //! spent 24,901 bytes of its 32,250 on ONE bare `file_read` of `spend.js`: 22,629 bytes taken |
| 213 | //! into context to look at four lines. A harness asserting only correctness would have printed |
| 214 | //! two passes. |
| 215 | //! |
| 216 | //! **`commit` and `world` passed, and neither pass means what it looks like.** Both are written |
| 217 | //! against browser-build faults this transport cannot reach, and both went green in three calls |
| 218 | //! and one -- which is exactly the false all-clear the section above says a task like this risks. |
| 219 | //! They are kept as the budgets they are: `world` at 4,000 bytes still refuses a run that reads |
| 220 | //! `dev/world.sh` rather than running it. Read the two green cells as "the capability exists on |
| 221 | //! a machine with no fence around it", and nothing more. |
| 222 | //! |
| 223 | //! `bigfile` failed on correctness with "The function `updateSpend` does not exist in that file", |
| 224 | //! which the section above records happening once before: it is what searching the wrong tree |
| 225 | //! looks like from the inside, and the file plainly holds ten of them. `bigmark`'s 46.0 s |
| 226 | //! against 21.6 s the evening before is a cold page cache, not a regression -- the same caution |
| 227 | //! the earlier table carries. |
| 228 | //! |
| 229 | //! And one thing this run found that no task was aimed at. `locales`'s failed call was |
| 230 | //! `file_read` on a directory, whose refusal `src/tools.rs` words carefully and at length -- and |
| 231 | //! it arrived at the model as `Error: LocalErr{[Invalid Input] "src/tools.rs:8733: file_read: |
| 232 | //! ...` with the ANSI colour codes still in it. The sentence somebody wrote for a model to read |
| 233 | //! is being delivered inside an error frame addressed to a developer at a terminal. |
| 234 | |
| 235 | use oxedyne_fe2o3_core::prelude::*; |
| 236 | use oxedyne_daimond::agent::{Agent, build_tls_client_config}; |
| 237 | use oxedyne_daimond::executor::Executor; |
| 238 | use oxedyne_daimond::llm::LlmClient; |
| 239 | use oxedyne_daimond::protocol::{AgentEvent, Session}; |
| 240 | use oxedyne_daimond::tools::{CallOutcome, diamond_bounds, FileRoot, Tool, ToolContext, ToolRegistry}; |
| 241 | use oxedyne_daimond::workspace::Workspace; |
| 242 | |
| 243 | use std::path::{Path, PathBuf}; |
| 244 | use std::sync::OnceLock; |
| 245 | use std::time::Instant; |
| 246 | |
| 247 | /// The repository under test, as an absolute path. |
| 248 | /// |
| 249 | /// Read from the environment so the probe can be pointed at a worktree, and defaulting to |
| 250 | /// this crate's own root -- which is the tree the tasks below are written about. |
| 251 | fn repo() -> PathBuf { |
| 252 | match std::env::var("DAIMOND_PROBE_REPO") { |
| 253 | Ok(p) if !p.trim().is_empty() => PathBuf::from(p), |
| 254 | _ => PathBuf::from(env!("CARGO_MANIFEST_DIR")), |
| 255 | } |
| 256 | } |
| 257 | |
| 258 | /// The workspace ROOT, which is not a mark and is above every mark. |
| 259 | /// |
| 260 | /// **This distinction is the one the first run of this probe got wrong**, and it is worth the |
| 261 | /// paragraph because it is the same distinction the product confuses. A `Workspace` is the |
| 262 | /// folder the file tools address paths against; a MARK is a folder inside it that |
| 263 | /// `diamond_bounds` names in an `OnlyWriteUnder`. Rooting the workspace AT the mark and then |
| 264 | /// marking `"."` normalises to the empty string, which `diamond_bounds` counts as no place at |
| 265 | /// all and answers with `Bound::Nowhere` -- so every file tool was refused, and the model |
| 266 | /// reported "I have no workspace attached", which was true and was the harness's fault. |
| 267 | /// |
| 268 | /// Five levels up from `code/web/apps/oxedyne/daimond` is `~/usr`, which holds both `code` and |
| 269 | /// `complement`: the arrangement the owner's own session was in, where a path in a brief reads |
| 270 | /// `code/web/apps/oxedyne/daimond/www/js/ledger.js`. Reproduced rather than tidied, because |
| 271 | /// the length of that path is part of what is being measured. |
| 272 | fn ws_root() -> PathBuf { |
| 273 | let r = repo(); |
| 274 | let mut p = r.as_path(); |
| 275 | for _ in 0..5 { |
| 276 | p = match p.parent() { |
| 277 | Some(up) => up, |
| 278 | None => return r.clone(), |
| 279 | }; |
| 280 | } |
| 281 | p.to_path_buf() |
| 282 | } |
| 283 | |
| 284 | /// The repository as the file tools see it: relative to [`ws_root`], e.g. |
| 285 | /// `code/web/apps/oxedyne/daimond`. Every path in a brief is written against this. |
| 286 | fn app_rel() -> String { |
| 287 | let (root, r) = (ws_root(), repo()); |
| 288 | match r.strip_prefix(&root) { |
| 289 | Ok(rel) => rel.to_string_lossy().replace('\\', "/"), |
| 290 | Err(_) => String::new(), |
| 291 | } |
| 292 | } |
| 293 | |
| 294 | /// The wide mark, relative to [`ws_root`]: `code`, ~590,000 files. The first segment of |
| 295 | /// [`app_rel`], so a worktree marks its own estate and not somebody else's. |
| 296 | fn wide_rel() -> String { |
| 297 | let rel = app_rel(); |
| 298 | match rel.split('/').next() { |
| 299 | Some(first) if !first.is_empty() => first.to_string(), |
| 300 | _ => rel, |
| 301 | } |
| 302 | } |
| 303 | |
| 304 | /// Where a task's workspace root sits, which is what makes a path in the brief long or short. |
| 305 | #[derive(Clone, Copy, PartialEq)] |
| 306 | enum Mark { |
| 307 | /// Marked at the app itself, which is what a careful user does. |
| 308 | App, |
| 309 | /// Marked at `~/usr/code`, which is what the owner did. ~590,000 files. |
| 310 | Wide, |
| 311 | } |
| 312 | |
| 313 | /// What a task's answer is checked against. |
| 314 | /// |
| 315 | /// A function over the tree and the reply, and never over what the model said about its own |
| 316 | /// work. `CONTRACT_CLAIMS.md` gives the reason at length: a turn's account of itself is |
| 317 | /// prose, and this codebase removed thirty-four prose sniffs in one night. |
| 318 | type Check = fn(&Path, &str) -> Result<(), String>; |
| 319 | |
| 320 | struct Task { |
| 321 | /// Short name, for `DAIMOND_PROBE_TASKS` and the report. |
| 322 | name: &'static str, |
| 323 | /// Which fault it exists to catch, printed beside a failure so a red line says why it |
| 324 | /// was worth running. |
| 325 | catches: &'static str, |
| 326 | mark: Mark, |
| 327 | brief: &'static str, |
| 328 | check: Check, |
| 329 | /// Tool calls this task may make. |
| 330 | max_calls: usize, |
| 331 | /// Bytes of tool output it may take into context. |
| 332 | max_bytes: usize, |
| 333 | /// Seconds the turn may take. Here because the first live run answered `bigmark` |
| 334 | /// correctly in 89.9 s: a right answer that costs a minute and a half of walking is a |
| 335 | /// finding, and correctness alone could not see it. |
| 336 | max_secs: f64, |
| 337 | /// Paths, repository-relative, `git checkout --` reverts after the run. Empty for a |
| 338 | /// read-only task, and a read-only task is CHECKED to have written nothing. |
| 339 | touches: &'static [&'static str], |
| 340 | // Fixture, made and unmade |
| 341 | // |
| 342 | // A fixture that is a git repository, or a file that must not parse, cannot be expressed in |
| 343 | // `touches`: both live under `target/`, which is ignored, so `git checkout --` has nothing to |
| 344 | // restore and would fail rather than clean. These make and unmake them, and neither ever |
| 345 | // runs a git verb in the repository under test -- see `setup_commit` for the argument. |
| 346 | setup: Option<fn(&Path) -> Result<(), String>>, // before the turn |
| 347 | teardown: Option<fn(&Path)>, // after the check |
| 348 | } |
| 349 | |
| 350 | // ── The tasks ──────────────────────────────────────────────────────────────────────── |
| 351 | // |
| 352 | // Eight -- four from the 23rd, four from the day after -- each aimed at one thing that went |
| 353 | // wrong, and each small enough that a competent person would call it a minute's work. That is the standard: the |
| 354 | // owner's objective says "without hitting continual snags", so the tasks are deliberately |
| 355 | // dull. A probe made of hard tasks measures the model; this one measures the app. |
| 356 | |
| 357 | /// 1. A large file must be readable, and read in part. |
| 358 | /// |
| 359 | /// `www/js/daimond.js` is 1.6 MB and holds three NUL characters as a composite-key |
| 360 | /// separator. Until 2026-08-23 `file_read` called it binary and refused it outright. The |
| 361 | /// budget is the second half: the file is 1.6 MB, so a run that reads it whole passes the |
| 362 | /// answer and fails the probe, which is correct -- taking 1.6 MB into context to find one |
| 363 | /// line is not a working development loop. |
| 364 | fn check_bigfile(repo: &Path, reply: &str) -> Result<(), String> { |
| 365 | let (line, _) = match first_cell_line(repo) { |
| 366 | Some(f) => f, |
| 367 | None => return Err(fmt!("www/js/daimond.js no longer holds an appendChild(cell( inside \ |
| 368 | updateSpend, so this task has rotted -- re-point it")), |
| 369 | }; |
| 370 | let plain = fmt!("{}", line); |
| 371 | let comma = fmt!("{},{:03}", line / 1000, line % 1000); |
| 372 | if reply.contains(&plain) || reply.contains(&comma) { |
| 373 | return Ok(()); |
| 374 | } |
| 375 | Err(fmt!("the reply does not name line {}: {}", line, |
| 376 | reply.chars().take(200).collect::<String>())) |
| 377 | } |
| 378 | |
| 379 | /// The line the task above asks for: the first `appendChild(cell(` inside `updateSpend`. |
| 380 | /// |
| 381 | /// **Found in the file, because it was written down and it rotted.** The number was 15996, and |
| 382 | /// the comment beside it said the answer was checked against the file rather than against a |
| 383 | /// number written here, which was not true of the code under it. By 2026-08-24 `updateSpend` |
| 384 | /// had moved to line 16026 and the call to 16050, so the check was red against every correct |
| 385 | /// answer -- and the self-test could not see it, because the good case it handed the check was |
| 386 | /// the same 15996 the check was looking for. A pair built out of the thing under test proves |
| 387 | /// that a constant equals itself. |
| 388 | fn first_cell_line(repo: &Path) -> Option<(usize, String)> { |
| 389 | let text = match std::fs::read_to_string(repo.join("www/js/daimond.js")) { |
| 390 | Ok(t) => t, |
| 391 | Err(_) => return None, |
| 392 | }; |
| 393 | let mut inside = false; |
| 394 | for (i, line) in text.lines().enumerate() { |
| 395 | if line.contains("function updateSpend(") { |
| 396 | inside = true; |
| 397 | continue; |
| 398 | } |
| 399 | if inside && line.contains("appendChild(cell(") { |
| 400 | return Some((i + 1, line.trim().to_string())); |
| 401 | } |
| 402 | } |
| 403 | None |
| 404 | } |
| 405 | |
| 406 | /// Where `updateSpend` begins and where the next function does, by line. |
| 407 | /// |
| 408 | /// The independent half of the pair above. [`first_cell_line`] walks forward and never looks |
| 409 | /// for the end of the function, so a `updateSpend` that stopped containing such a call would |
| 410 | /// hand back a line out of whatever came next and say nothing. This finds the closing bound a |
| 411 | /// different way, and the self-test holds the one against the other. |
| 412 | fn update_spend_span(repo: &Path) -> Option<(usize, usize)> { |
| 413 | let text = match std::fs::read_to_string(repo.join("www/js/daimond.js")) { |
| 414 | Ok(t) => t, |
| 415 | Err(_) => return None, |
| 416 | }; |
| 417 | let mut start = None; |
| 418 | for (i, line) in text.lines().enumerate() { |
| 419 | match start { |
| 420 | None => if line.contains("function updateSpend(") { start = Some(i + 1); }, |
| 421 | Some(s) => if line.starts_with("\tfunction ") { |
| 422 | return Some((s, i + 1)); |
| 423 | }, |
| 424 | } |
| 425 | } |
| 426 | start.map(|s| (s, usize::MAX)) |
| 427 | } |
| 428 | |
| 429 | /// 2. A search under a wide mark must not answer from a fraction of it. |
| 430 | /// |
| 431 | /// `WALK_ENTRIES_MAX` is 20,000 and its own comment justifies the figure on the premise |
| 432 | /// that a mark is "an ordinary project". Marked at `code`, that is 3.4% of the reach, and |
| 433 | /// `file_search` reports "0 matches" with the shortfall in a footnote. The task is |
| 434 | /// answerable -- the constant is right there in `src/tools.rs` -- so a failure here is the |
| 435 | /// app answering a smaller question than the one it was asked. |
| 436 | fn check_bigmark(_repo: &Path, reply: &str) -> Result<(), String> { |
| 437 | let says_file = reply.contains("tools.rs"); |
| 438 | let says_num = reply.contains("20000") || reply.contains("20_000") || reply.contains("20,000"); |
| 439 | if says_file && says_num { |
| 440 | return Ok(()); |
| 441 | } |
| 442 | Err(fmt!("wanted the file and the value; got: {}", |
| 443 | reply.chars().take(200).collect::<String>())) |
| 444 | } |
| 445 | |
| 446 | // The eight files the fan-out task has to reach, and the oracle the check reads. |
| 447 | const LOCALES: &[&str] = &["de.js", "en.js", "es.js", "fr.js", "ja.js", "ko.js", "pt-BR.js", |
| 448 | "zh-Hans.js"]; |
| 449 | |
| 450 | // A key the tree does not already carry, and its English value. |
| 451 | // |
| 452 | // A LIST, and not the one constant this used to be. `spend.period_day` landed in 1e294c1, and |
| 453 | // from that commit on `check_locales` was green against a tree where nothing had been asked or |
| 454 | // done -- a task measuring nothing while reporting a pass, which is the failure this whole file |
| 455 | // is about. The self-test found it on its first run afterwards. The task now asks for the |
| 456 | // first of these no locale carries, so the same landing cannot silence it twice. |
| 457 | const PERIOD_KEYS: &[(&str, &str)] = &[ |
| 458 | ("spend.period_day", "Day"), |
| 459 | ("spend.period_hour", "Hour"), |
| 460 | ("spend.period_year", "Year"), |
| 461 | ("spend.period_quarter", "Quarter"), |
| 462 | ]; |
| 463 | |
| 464 | /// The key the fan-out task asks for on this tree, and the English word for it. |
| 465 | /// |
| 466 | /// **Chosen once and then fixed for the life of the process**, which is not an optimisation. |
| 467 | /// The check runs AFTER the turn has added the key, so a second derivation from the tree would |
| 468 | /// find that candidate taken, step to the next one, and report a task that had just been done |
| 469 | /// as not done. The self-test caught precisely that on the first run after the list went in -- |
| 470 | /// green, red, then red again with the answer planted -- which is what a check derived from the |
| 471 | /// evidence it is checking looks like from the outside. |
| 472 | fn period_key(repo: &Path) -> Option<(&'static str, &'static str)> { |
| 473 | static CHOSEN: OnceLock<Option<(&'static str, &'static str)>> = OnceLock::new(); |
| 474 | *CHOSEN.get_or_init(|| choose_period_key(repo)) |
| 475 | } |
| 476 | |
| 477 | /// The first key in [`PERIOD_KEYS`] no locale file carries. |
| 478 | fn choose_period_key(repo: &Path) -> Option<(&'static str, &'static str)> { |
| 479 | let dir = repo.join("www/i18n"); |
| 480 | for (key, eng) in PERIOD_KEYS { |
| 481 | let mut free = true; |
| 482 | for n in LOCALES { |
| 483 | match std::fs::read_to_string(dir.join(n)) { |
| 484 | Ok(t) => if t.contains(key) { free = false; }, |
| 485 | Err(_) => return None, |
| 486 | } |
| 487 | } |
| 488 | if free { |
| 489 | return Some((key, eng)); |
| 490 | } |
| 491 | } |
| 492 | None |
| 493 | } |
| 494 | |
| 495 | /// 3. A change across eight files is one task, not eight. |
| 496 | /// |
| 497 | /// The locale fan-out is the shape most likely to be done four-eighths of the way. The eight |
| 498 | /// files are read here rather than put through `dev/i18ncheck.mjs`, because that script answers |
| 499 | /// a coverage question over the whole catalogue and this task is about ONE key: a run that added |
| 500 | /// it to four files would leave the script red for reasons that were red before. |
| 501 | fn check_locales(repo: &Path, _reply: &str) -> Result<(), String> { |
| 502 | let (key, _) = match period_key(repo) { |
| 503 | Some(k) => k, |
| 504 | None => return Err(fmt!("every key in PERIOD_KEYS is already in the tree, so there is \ |
| 505 | nothing left for this task to ask for -- add one, or re-point it")), |
| 506 | }; |
| 507 | let dir = repo.join("www/i18n"); |
| 508 | let mut absent = Vec::new(); |
| 509 | for n in LOCALES { |
| 510 | let text = match std::fs::read_to_string(dir.join(n)) { |
| 511 | Ok(t) => t, |
| 512 | Err(e) => return Err(fmt!("{}: {}", n, e)), |
| 513 | }; |
| 514 | if !text.contains(key) { |
| 515 | absent.push(*n); |
| 516 | } |
| 517 | } |
| 518 | if absent.is_empty() { |
| 519 | return Ok(()); |
| 520 | } |
| 521 | Err(fmt!("{} is absent from {}", key, absent.join(" "))) |
| 522 | } |
| 523 | |
| 524 | /// 4. A tool is run, not read. |
| 525 | /// |
| 526 | /// `dev/i18ncheck.mjs` is 39 KB, almost all of it commentary, and a daimon read the whole |
| 527 | /// of it to learn how to invoke it. The answer is one line of that script's output, so a |
| 528 | /// run that reads the file at all is spending 39 KB to avoid one `run` call -- which the |
| 529 | /// byte budget, set below the file's size, is what catches. |
| 530 | fn check_ran_it(_repo: &Path, reply: &str) -> Result<(), String> { |
| 531 | let l = reply.to_lowercase(); |
| 532 | // The script says "all 7 locales carry every one of en.js's N keys" when it is happy. |
| 533 | if l.contains("locales carry") || l.contains("out of step") { |
| 534 | return Ok(()); |
| 535 | } |
| 536 | Err(fmt!("the reply does not quote the checker's own verdict: {}", |
| 537 | reply.chars().take(200).collect::<String>())) |
| 538 | } |
| 539 | |
| 540 | // ── Four more, from the night of the 23rd and the day after ────────────────────────── |
| 541 | // |
| 542 | // Each of the four below was found by the owner, or by a daimon failing in front of him, |
| 543 | // and by no instrument in this repository. That is the thing they exist to end. |
| 544 | // |
| 545 | // **Two of them are red for a reason that is not the app's.** `commit` and `world` are |
| 546 | // written against faults that live in the browser build -- the hand's fence, and the |
| 547 | // environment a granted toolkit hands a command -- and this probe is a native binary, where |
| 548 | // `Tool::Run` and `Tool::Verify` are refused at compile time and `sh -c` inherits the whole |
| 549 | // environment. See the header section "What this transport cannot see". They are kept |
| 550 | // because the budget still bites on the native side: reading an 11 KB script instead of |
| 551 | // running it fails `world` today, whatever the hand does. |
| 552 | |
| 553 | /// A directory of the probe's own, under the repository's ignored build output. |
| 554 | /// |
| 555 | /// Everything a task has to BUILD goes here rather than into the tracked tree, and it is |
| 556 | /// removed by the task's own teardown. `touches` cannot help: it is a `git checkout --`, and |
| 557 | /// there is nothing for that to restore under an ignored path. |
| 558 | fn probe_dir(repo: &Path) -> PathBuf { |
| 559 | repo.join("target/probe") |
| 560 | } |
| 561 | |
| 562 | /// The throwaway repository the `commit` task works in. |
| 563 | fn commit_dir(repo: &Path) -> PathBuf { |
| 564 | probe_dir(repo).join("commit") |
| 565 | } |
| 566 | |
| 567 | // What the commit task must produce, checked against the throwaway repository's own log. |
| 568 | const COMMIT_SUBJECT: &str = "probe: bump the count"; |
| 569 | |
| 570 | // The seed file, and the one line in it the task is asked to change. |
| 571 | const SCRATCH_SEED: &str = "// Scratch, made by the probe and thrown away after it.\ncount = 1\n"; |
| 572 | |
| 573 | /// Run `git` in `dir` and answer its standard output, trimmed. |
| 574 | /// |
| 575 | /// Every call names the directory explicitly, and no call here is ever made in the repository |
| 576 | /// under test -- see [`setup_commit`], where that separation is the whole safety argument. |
| 577 | fn git(dir: &Path, args: &[&str]) -> Result<String, String> { |
| 578 | let out = match std::process::Command::new("git").arg("-C").arg(dir).args(args).output() { |
| 579 | Ok(o) => o, |
| 580 | Err(e) => return Err(fmt!("git {}: {}", args.join(" "), e)), |
| 581 | }; |
| 582 | if !out.status.success() { |
| 583 | return Err(fmt!("git {} in {:?}: {}", args.join(" "), dir, |
| 584 | String::from_utf8_lossy(&out.stderr).trim())); |
| 585 | } |
| 586 | Ok(String::from_utf8_lossy(&out.stdout).trim().to_string()) |
| 587 | } |
| 588 | |
| 589 | /// The repository's HEAD and its staged paths, as one line. |
| 590 | /// |
| 591 | /// Read after every task, because four other people's uncommitted work is in this tree while |
| 592 | /// the probe runs. A task that moved HEAD or added to the index did it to THEM, and the only |
| 593 | /// safe thing an instrument can do about that is say so loudly: undoing a commit or an index |
| 594 | /// is exactly the reset that turns one bad turn into somebody's lost afternoon. |
| 595 | fn repo_state(repo: &Path) -> String { |
| 596 | let head = match git(repo, &["rev-parse", "HEAD"]) { |
| 597 | Ok(h) => h, |
| 598 | Err(e) => e, |
| 599 | }; |
| 600 | let staged = match git(repo, &["diff", "--cached", "--name-only"]) { |
| 601 | Ok(s) => s.split('\n').filter(|l| !l.trim().is_empty()).collect::<Vec<&str>>().join(","), |
| 602 | Err(e) => e, |
| 603 | }; |
| 604 | fmt!("HEAD {} staged [{}]", head, staged) |
| 605 | } |
| 606 | |
| 607 | /// 5. A daimon must be able to commit its own work. |
| 608 | /// |
| 609 | /// On 2026-08-23 one could not see `.git` at all: `ls -la` showed none and `git status` walked |
| 610 | /// every parent to `/` before giving up. A machine that can write the change and cannot record |
| 611 | /// it is a machine that cannot do a day's work, whatever else it can do, so this is the hard |
| 612 | /// stop under "full self-development capability" and nothing measured it. |
| 613 | /// |
| 614 | /// **It works in a repository of its own, and that is not squeamishness.** Four other lanes |
| 615 | /// have uncommitted work in the tree under test. A task that committed HERE would need undoing |
| 616 | /// afterwards, and the undo -- `reset`, `checkout`, `stash` -- is the operation that destroys |
| 617 | /// somebody else's morning when it is aimed one word wrong. An inner repository is a WALL |
| 618 | /// rather than a promise: git resolves to the nearest `.git`, so a `git add -A` inside it cannot |
| 619 | /// reach the outer index however the turn is worded, and the teardown is `remove_dir_all` on a |
| 620 | /// directory the probe made, which cannot touch a tracked file at all. What it gives up is the |
| 621 | /// hand's fence, which this transport could not see in any case; what it keeps is the whole of |
| 622 | /// the capability. |
| 623 | fn check_commit(repo: &Path, _reply: &str) -> Result<(), String> { |
| 624 | let dir = commit_dir(repo); |
| 625 | let subject = match git(&dir, &["log", "-1", "--pretty=%s"]) { |
| 626 | Ok(s) => s, |
| 627 | Err(e) => return Err(fmt!("no commit to read: {}", e)), |
| 628 | }; |
| 629 | if subject != COMMIT_SUBJECT { |
| 630 | return Err(fmt!("the last commit is '{}', not '{}'", subject, COMMIT_SUBJECT)); |
| 631 | } |
| 632 | let text = match git(&dir, &["show", "HEAD:scratch.txt"]) { |
| 633 | Ok(t) => t, |
| 634 | Err(e) => return Err(fmt!("scratch.txt is not in that commit: {}", e)), |
| 635 | }; |
| 636 | if !text.contains("count = 2") { |
| 637 | return Err(fmt!("the commit does not carry the change: {}", |
| 638 | text.chars().take(120).collect::<String>())); |
| 639 | } |
| 640 | // Committed, and not merely written. A turn that edits the file and commits nothing leaves a |
| 641 | // dirty tree, which is the near miss worth telling apart from the hit. |
| 642 | match git(&dir, &["status", "--porcelain"]) { |
| 643 | Ok(s) if s.trim().is_empty() => Ok(()), |
| 644 | Ok(s) => Err(fmt!("committed, but left the tree dirty: {}", s.replace('\n', "; "))), |
| 645 | Err(e) => Err(e), |
| 646 | } |
| 647 | } |
| 648 | |
| 649 | /// Build the throwaway repository, seeded and committed once. |
| 650 | fn setup_commit(repo: &Path) -> Result<(), String> { |
| 651 | let dir = commit_dir(repo); |
| 652 | let _ = std::fs::remove_dir_all(&dir); |
| 653 | if let Err(e) = std::fs::create_dir_all(&dir) { |
| 654 | return Err(fmt!("{:?}: {}", dir, e)); |
| 655 | } |
| 656 | let steps: [&[&str]; 4] = [ |
| 657 | &["init", "-q"], |
| 658 | &["config", "user.name", "Daimond probe"], |
| 659 | &["config", "user.email", "probe@localhost"], |
| 660 | // No hook runs in here. The machine's global `core.hooksPath` points at a credential |
| 661 | // scanner with nothing to say about this task, and a hook that refused the commit would |
| 662 | // reach the report as the daimon having failed. |
| 663 | &["config", "core.hooksPath", ".git/hooks-none"], |
| 664 | ]; |
| 665 | for args in steps { |
| 666 | if let Err(e) = git(&dir, args) { |
| 667 | return Err(e); |
| 668 | } |
| 669 | } |
| 670 | if let Err(e) = std::fs::write(dir.join("scratch.txt"), SCRATCH_SEED) { |
| 671 | return Err(fmt!("scratch.txt: {}", e)); |
| 672 | } |
| 673 | if let Err(e) = git(&dir, &["add", "scratch.txt"]) { |
| 674 | return Err(e); |
| 675 | } |
| 676 | if let Err(e) = git(&dir, &["commit", "-q", "-m", "seed"]) { |
| 677 | return Err(e); |
| 678 | } |
| 679 | Ok(()) |
| 680 | } |
| 681 | |
| 682 | fn teardown_commit(repo: &Path) { |
| 683 | let _ = std::fs::remove_dir_all(commit_dir(repo)); |
| 684 | } |
| 685 | |
| 686 | /// 6. Work is checked with the repository's own machinery, not with a check reinvented. |
| 687 | /// |
| 688 | /// A daimon asked to verify a page spent FORTY-ONE tool calls trying to reconstruct the check -- |
| 689 | /// starting a dev server, hunting for `playwright-core`, testing whether it had a network -- for |
| 690 | /// a verifier this repository runs with one command. That is the shape, and the cheapest true |
| 691 | /// instance of it is here: `node --check` EXITS 0 ON A FILE WITH A SYNTAX ERROR IN IT when the |
| 692 | /// file is named `.js` and opens with an `import`, because node then parses it as CommonJS. |
| 693 | /// `dev/jscheck.sh` exists for exactly that, and says so in its own header. |
| 694 | /// |
| 695 | /// **So the wrong route does not merely cost more, it answers wrongly**, which is why this task |
| 696 | /// is checked on the verdict and not only on the budget: a turn that reaches for its own |
| 697 | /// `node --check` is told the file is fine, and reports that the file is fine. |
| 698 | fn check_parses(_repo: &Path, reply: &str) -> Result<(), String> { |
| 699 | let l = reply.to_lowercase(); |
| 700 | let named = l.contains("jscheck"); |
| 701 | let failed = l.contains("do not parse") || l.contains("does not parse") |
| 702 | || l.contains("fail") || l.contains("syntax error"); |
| 703 | if named && failed { |
| 704 | return Ok(()); |
| 705 | } |
| 706 | Err(fmt!("wanted the repository's own checker named and its refusal reported; got: {}", |
| 707 | reply.chars().take(200).collect::<String>())) |
| 708 | } |
| 709 | |
| 710 | /// The file that must not parse, and the `.js` name that makes `node --check` lie about it. |
| 711 | fn setup_parse(repo: &Path) -> Result<(), String> { |
| 712 | let dir = probe_dir(repo).join("parse"); |
| 713 | let _ = std::fs::remove_dir_all(&dir); |
| 714 | if let Err(e) = std::fs::create_dir_all(&dir) { |
| 715 | return Err(fmt!("{:?}: {}", dir, e)); |
| 716 | } |
| 717 | // The `import` is load-bearing: without it node parses the file as a script, finds the |
| 718 | // error and the task stops measuring anything. |
| 719 | let src = "import * as X from './x.js';\nvar broken = (1;\n"; |
| 720 | match std::fs::write(dir.join("broken.js"), src) { |
| 721 | Ok(()) => Ok(()), |
| 722 | Err(e) => Err(fmt!("broken.js: {}", e)), |
| 723 | } |
| 724 | } |
| 725 | |
| 726 | fn teardown_parse(repo: &Path) { |
| 727 | let _ = std::fs::remove_dir_all(probe_dir(repo).join("parse")); |
| 728 | } |
| 729 | |
| 730 | /// The local names a file gives to `what`, e.g. the `L` in `var L = window.DaimondLedger;`. |
| 731 | /// |
| 732 | /// A reference followed by another name character, or by a dot, is a USE and not a binding, so |
| 733 | /// `window.DaimondLedger.totals(` is passed over here and caught by the direct match instead. |
| 734 | fn aliases_of(text: &str, what: &str) -> Vec<String> { |
| 735 | let mut out: Vec<String> = Vec::new(); |
| 736 | for line in text.lines() { |
| 737 | let mut from = 0usize; |
| 738 | while let Some(found) = line[from..].find(what) { |
| 739 | let at = from + found; |
| 740 | from = at + what.len(); |
| 741 | match line[from..].chars().next() { |
| 742 | Some(c) if c.is_alphanumeric() || c == '_' || c == '$' || c == '.' => continue, |
| 743 | _ => {} |
| 744 | } |
| 745 | // Back over the `=`, refusing `==`, `!=`, `<=` and `>=`, which bind nothing. |
| 746 | let head = line[..at].trim_end(); |
| 747 | let head = match head.strip_suffix('=') { |
| 748 | Some(h) => h, |
| 749 | None => continue, |
| 750 | }; |
| 751 | match head.chars().last() { |
| 752 | Some('=') | Some('!') | Some('<') | Some('>') => continue, |
| 753 | _ => {} |
| 754 | } |
| 755 | let mut name: Vec<char> = head.trim_end().chars().rev() |
| 756 | .take_while(|c| c.is_alphanumeric() || *c == '_' || *c == '$') |
| 757 | .collect(); |
| 758 | name.reverse(); |
| 759 | let name: String = name.into_iter().collect(); |
| 760 | if !name.is_empty() && !out.contains(&name) { |
| 761 | out.push(name); |
| 762 | } |
| 763 | } |
| 764 | } |
| 765 | out |
| 766 | } |
| 767 | |
| 768 | /// Every place in `www/js` that calls `DaimondLedger.totals`, alias and all. |
| 769 | /// |
| 770 | /// Computed from the tree on every run rather than written down, so a caller added, moved or |
| 771 | /// renamed cannot turn the task green by having been recorded here once. |
| 772 | fn totals_callers(repo: &Path) -> Result<Vec<(String, usize)>, String> { |
| 773 | let dir = repo.join("www/js"); |
| 774 | let listing = match std::fs::read_dir(&dir) { |
| 775 | Ok(l) => l, |
| 776 | Err(e) => return Err(fmt!("{:?}: {}", dir, e)), |
| 777 | }; |
| 778 | let mut names: Vec<String> = Vec::new(); |
| 779 | for ent in listing.flatten() { |
| 780 | let p = ent.path(); |
| 781 | if p.extension().and_then(|x| x.to_str()) != Some("js") { |
| 782 | continue; |
| 783 | } |
| 784 | if let Some(n) = p.file_name().and_then(|x| x.to_str()) { |
| 785 | names.push(n.to_string()); |
| 786 | } |
| 787 | } |
| 788 | names.sort(); |
| 789 | let mut out: Vec<(String, usize)> = Vec::new(); |
| 790 | for n in &names { |
| 791 | let text = match std::fs::read_to_string(dir.join(n)) { |
| 792 | Ok(t) => t, |
| 793 | Err(e) => return Err(fmt!("{}: {}", n, e)), |
| 794 | }; |
| 795 | let aliases = aliases_of(&text, "window.DaimondLedger"); |
| 796 | for (i, line) in text.lines().enumerate() { |
| 797 | let direct = line.contains("DaimondLedger.totals("); |
| 798 | let through = aliases.iter().any(|a| line.contains(&fmt!("{}.totals(", a))); |
| 799 | if direct || through { |
| 800 | out.push((n.clone(), i + 1)); |
| 801 | } |
| 802 | } |
| 803 | } |
| 804 | Ok(out) |
| 805 | } |
| 806 | |
| 807 | /// The same answer by a different road, for the self-test to hold `check_alias` against. |
| 808 | /// |
| 809 | /// A bare scan for `.totals(` needs no alias resolution and finds both callers, so it is an |
| 810 | /// oracle the check does not share a line of code with -- which is the point, since a self-test |
| 811 | /// that builds its right answer out of the thing under test proves only that the thing agrees |
| 812 | /// with itself. If the two ever disagree, the self-test goes red and that is a finding: either |
| 813 | /// the alias walk has broken, or `.totals(` has grown a second meaning in this tree. |
| 814 | fn bare_totals_scan(repo: &Path) -> String { |
| 815 | let dir = repo.join("www/js"); |
| 816 | let listing = match std::fs::read_dir(&dir) { |
| 817 | Ok(l) => l, |
| 818 | Err(_) => return String::new(), |
| 819 | }; |
| 820 | let mut names: Vec<String> = Vec::new(); |
| 821 | for ent in listing.flatten() { |
| 822 | if let Some(n) = ent.path().file_name().and_then(|x| x.to_str()) { |
| 823 | if n.ends_with(".js") { |
| 824 | names.push(n.to_string()); |
| 825 | } |
| 826 | } |
| 827 | } |
| 828 | names.sort(); |
| 829 | let mut out: Vec<String> = Vec::new(); |
| 830 | for n in &names { |
| 831 | let text = match std::fs::read_to_string(dir.join(n)) { |
| 832 | Ok(t) => t, |
| 833 | Err(_) => continue, |
| 834 | }; |
| 835 | for (i, line) in text.lines().enumerate() { |
| 836 | if line.contains(".totals(") { |
| 837 | out.push(fmt!("www/js/{}:{}", n, i + 1)); |
| 838 | } |
| 839 | } |
| 840 | } |
| 841 | out.join(", ") |
| 842 | } |
| 843 | |
| 844 | /// 7. "Who calls this?" is the question before every non-trivial change. |
| 845 | /// |
| 846 | /// A daimon grepped `DaimondLedger.totals` and reported one caller. There are two: the second |
| 847 | /// goes through `var L = window.DaimondLedger` and reads `L.totals()`. It gave the right answer |
| 848 | /// for a reason that did not hold -- had the brief been RENAME rather than add, the Spending |
| 849 | /// panel would have lost a stat and nothing would have said so. |
| 850 | /// |
| 851 | /// This app answers that question with string matching, so the fault is not the model's: it is |
| 852 | /// what `file_search` can see. The check computes the callers from the tree at run time, so it |
| 853 | /// follows a file that moves and goes red rather than stale when the alias does. |
| 854 | fn check_alias(repo: &Path, reply: &str) -> Result<(), String> { |
| 855 | let callers = match totals_callers(repo) { |
| 856 | Ok(c) => c, |
| 857 | Err(e) => return Err(e), |
| 858 | }; |
| 859 | let mut files: Vec<&str> = callers.iter().map(|(f, _)| f.as_str()).collect(); |
| 860 | files.sort(); |
| 861 | files.dedup(); |
| 862 | // A task pointed at a question that no longer has the shape it was written for measures |
| 863 | // nothing, and must say that rather than pass. |
| 864 | if files.len() < 2 { |
| 865 | return Err(fmt!("the tree holds {} caller(s) of DaimondLedger.totals in {} file(s), so \ |
| 866 | there is no aliased one left and this task has rotted -- re-point it", |
| 867 | callers.len(), files.len())); |
| 868 | } |
| 869 | let missed: Vec<String> = callers.iter() |
| 870 | .filter(|(f, l)| !reply.contains(f.as_str()) || !reply.contains(&fmt!("{}", l))) |
| 871 | .map(|(f, l)| fmt!("{}:{}", f, l)) |
| 872 | .collect(); |
| 873 | if missed.is_empty() { |
| 874 | return Ok(()); |
| 875 | } |
| 876 | Err(fmt!("the reply does not name {}: {}", missed.join(" "), |
| 877 | reply.chars().take(200).collect::<String>())) |
| 878 | } |
| 879 | |
| 880 | /// 8. A script under `dev/` needs `$HOME`, and must be run rather than read. |
| 881 | /// |
| 882 | /// `run` clears the environment and passes no `HOME` unless the user granted the Git toolkit |
| 883 | /// (see `Kit::for_home` in `src/tools.rs`), so nearly every script under `dev/` dies before it |
| 884 | /// prints anything. A daimon met it on `dev/world.sh`, which is the one used here: with `HOME` |
| 885 | /// unset it stops at line 84 with `HOME: unbound variable` and produces no output at all. |
| 886 | /// |
| 887 | /// The scratch path is the whole check, because it is the one line of that output `$HOME` is |
| 888 | /// needed to produce -- the ports beside it are arithmetic and would be printed by a shell that |
| 889 | /// had nothing. |
| 890 | fn check_world(_repo: &Path, reply: &str) -> Result<(), String> { |
| 891 | let home = match std::env::var("HOME") { |
| 892 | Ok(h) if !h.trim().is_empty() => h, |
| 893 | _ => return Err(fmt!("this probe's own HOME is unset, so there is no value to check the \ |
| 894 | reply against; the check cannot run")), |
| 895 | }; |
| 896 | let want = fmt!("{}/.cache/daimond", home.trim_end_matches('/')); |
| 897 | if reply.contains(&want) { |
| 898 | return Ok(()); |
| 899 | } |
| 900 | Err(fmt!("the reply does not carry the scratch path the script prints ({}): {}", want, |
| 901 | reply.chars().take(200).collect::<String>())) |
| 902 | } |
| 903 | |
| 904 | const TASKS: &[Task] = &[ |
| 905 | Task { |
| 906 | name: "bigfile", |
| 907 | catches: "file_read refusing the app's own 1.6 MB UI source, and reading it whole", |
| 908 | mark: Mark::App, |
| 909 | brief: "In {app}/www/js/daimond.js, find the function updateSpend. Tell me the line \ |
| 910 | number of the first el.appendChild(cell(...)) call inside it. Answer \ |
| 911 | with the number and nothing else. Change no files.", |
| 912 | check: check_bigfile, |
| 913 | max_calls: 6, |
| 914 | // 20 KB. The first live run passed this task at 87,296 bytes, of which ONE |
| 915 | // `file_read` was 80,016 -- 80 KB taken into context to learn a line number, which |
| 916 | // is the waste the owner stopped a turn over. A budget set above the observed |
| 917 | // figure measures nothing, so this is set at what the task is worth: a targeted |
| 918 | // search and a small read around the hit. |
| 919 | max_bytes: 20_000, |
| 920 | max_secs: 30.0, |
| 921 | touches: &[], |
| 922 | setup: None, |
| 923 | teardown: None, |
| 924 | }, |
| 925 | Task { |
| 926 | name: "bigmark", |
| 927 | catches: "a search under a 590,000-file mark answering from 20,000 of them", |
| 928 | mark: Mark::Wide, |
| 929 | brief: "Somewhere under this workspace is a Rust constant named \ |
| 930 | WALK_ENTRIES_MAX. Tell me which file defines it and what value it is \ |
| 931 | given. Change no files.", |
| 932 | check: check_bigmark, |
| 933 | max_calls: 8, |
| 934 | max_bytes: 60_000, |
| 935 | // 89.9 s on the first live run, nearly all of it the walk over 590,000 files. |
| 936 | max_secs: 30.0, |
| 937 | touches: &[], |
| 938 | setup: None, |
| 939 | teardown: None, |
| 940 | }, |
| 941 | Task { |
| 942 | name: "locales", |
| 943 | catches: "a fan-out across eight files done part of the way", |
| 944 | mark: Mark::App, |
| 945 | brief: "Add the key {key} to all eight locale files in {app}/www/i18n/, \ |
| 946 | beside the existing spend.period_week. The English value is {eng}; \ |
| 947 | translate it for the other seven. Touch nothing else.", |
| 948 | check: check_locales, |
| 949 | max_calls: 24, |
| 950 | max_bytes: 60_000, |
| 951 | max_secs: 90.0, |
| 952 | touches: &["www/i18n"], |
| 953 | setup: None, |
| 954 | teardown: None, |
| 955 | }, |
| 956 | Task { |
| 957 | name: "ranit", |
| 958 | catches: "reading a 39 KB script instead of running it", |
| 959 | mark: Mark::App, |
| 960 | brief: "In {app}, run node dev/i18ncheck.mjs --frozen and tell me the verdict line \ |
| 961 | it prints. Do not pass any other flag. Change no files.", |
| 962 | check: check_ran_it, |
| 963 | max_calls: 4, |
| 964 | // Below the 39 KB of `dev/i18ncheck.mjs`, deliberately: reading the script is the |
| 965 | // failure, so the budget has to be a figure reading it cannot fit inside. |
| 966 | max_bytes: 30_000, |
| 967 | max_secs: 30.0, |
| 968 | touches: &[], |
| 969 | setup: None, |
| 970 | teardown: None, |
| 971 | }, |
| 972 | Task { |
| 973 | name: "commit", |
| 974 | catches: "a daimon that can write the change and cannot record it", |
| 975 | mark: Mark::App, |
| 976 | brief: "The folder {app}/target/probe/commit is a git repository of its own, and the \ |
| 977 | only one this task concerns. In its scratch.txt, change the line reading \ |
| 978 | 'count = 1' to 'count = 2', and commit that one change with the message \ |
| 979 | 'probe: bump the count'. Run no git command anywhere else.", |
| 980 | check: check_commit, |
| 981 | // A read of a one-line file, an edit, an add, a commit, and a look at the log to see it |
| 982 | // landed. Five is the shape; the sixth is the allowance for landing in the wrong |
| 983 | // directory once, which is what the 23rd looked like. |
| 984 | max_calls: 6, |
| 985 | // The file is 70 bytes and git's own output for these verbs is a few hundred more. Set |
| 986 | // where reading anything large to find out where the repository is will not fit. |
| 987 | max_bytes: 5_000, |
| 988 | max_secs: 45.0, |
| 989 | touches: &[], |
| 990 | setup: Some(setup_commit), |
| 991 | teardown: Some(teardown_commit), |
| 992 | }, |
| 993 | Task { |
| 994 | name: "parses", |
| 995 | catches: "reinventing a check this repository already runs, and being told the wrong thing by it", |
| 996 | mark: Mark::App, |
| 997 | brief: "Does {app}/target/probe/parse/broken.js parse? Answer it with the machinery \ |
| 998 | already in this repository for that question rather than a command of your \ |
| 999 | own, and report the verdict it prints. Change no files.", |
| 1000 | check: check_parses, |
| 1001 | // A targeted glob to find the checker, the checker's own 1,952 bytes if the turn wants |
| 1002 | // to read it, and the run. |
| 1003 | max_calls: 5, |
| 1004 | // Below the ~11 KB a capped listing of `dev/`'s 526 entries costs. Knowing which script |
| 1005 | // answers "does this parse" should not cost a directory. |
| 1006 | max_bytes: 8_000, |
| 1007 | max_secs: 40.0, |
| 1008 | touches: &[], |
| 1009 | setup: Some(setup_parse), |
| 1010 | teardown: Some(teardown_parse), |
| 1011 | }, |
| 1012 | Task { |
| 1013 | name: "alias", |
| 1014 | catches: "\"who calls this?\" answered by string matching, one caller of two", |
| 1015 | mark: Mark::App, |
| 1016 | brief: "Name every place in {app}/www/js that calls DaimondLedger's totals method, \ |
| 1017 | each as a file and a line number. Include the ones that reach it through a \ |
| 1018 | local name. Change no files.", |
| 1019 | check: check_alias, |
| 1020 | // Three searches -- the symbol, the object, the local name it is bound to -- and one |
| 1021 | // targeted read around a hit. |
| 1022 | max_calls: 5, |
| 1023 | // Twenty lines carrying `DaimondLedger` across `www/js` is about 2 KB. This allows the |
| 1024 | // searches and a read with an offset, and refuses a bare read of `spend.js`, which is |
| 1025 | // 22,629 bytes and under the peek threshold, so it would come back whole. |
| 1026 | max_bytes: 8_000, |
| 1027 | max_secs: 40.0, |
| 1028 | touches: &[], |
| 1029 | setup: None, |
| 1030 | teardown: None, |
| 1031 | }, |
| 1032 | Task { |
| 1033 | name: "world", |
| 1034 | catches: "a dev script that dies before it prints, because the command was given no HOME", |
| 1035 | mark: Mark::App, |
| 1036 | brief: "In {app}, run 'bash dev/world.sh 0 --env' and tell me the scratch directory \ |
| 1037 | it names. Pass no other argument, and start nothing. Change no files.", |
| 1038 | check: check_world, |
| 1039 | // One command answers it. Three is two false starts' grace. |
| 1040 | max_calls: 3, |
| 1041 | // The script is 11,419 bytes and its answer is 330. Set between them deliberately: this |
| 1042 | // is `ranit`'s fault in a second place, and the budget has to be a figure the read |
| 1043 | // cannot fit inside. |
| 1044 | max_bytes: 4_000, |
| 1045 | max_secs: 20.0, |
| 1046 | touches: &[], |
| 1047 | setup: None, |
| 1048 | teardown: None, |
| 1049 | }, |
| 1050 | ]; |
| 1051 | |
| 1052 | /// What one task's turn actually did. |
| 1053 | #[derive(Default)] |
| 1054 | struct Spend { |
| 1055 | calls: usize, |
| 1056 | bytes: usize, |
| 1057 | refused: usize, |
| 1058 | failed: usize, |
| 1059 | /// The tool that returned the most, and how much, so a report names the read that hurt |
| 1060 | /// rather than only the total. |
| 1061 | worst: (String, usize), |
| 1062 | /// The FIRST call that was refused or failed, and what it said. |
| 1063 | /// |
| 1064 | /// Added on the instrument's third run, when `locales` reported one failed call and this |
| 1065 | /// file could say nothing about which one. A harness that counts a fault without naming it |
| 1066 | /// sends its reader back to the provider logs, which is the position the probe exists to |
| 1067 | /// get out of -- and "1 failed" with no name is exactly the shape of report this codebase |
| 1068 | /// keeps writing up as answering from the wrong evidence. |
| 1069 | firstbad: Option<(String, String)>, |
| 1070 | prompt: u64, |
| 1071 | completion: u64, |
| 1072 | secs: f64, |
| 1073 | } |
| 1074 | |
| 1075 | fn main() { |
| 1076 | // The closing line names WHICH run it is closing. A self-test that signs off with |
| 1077 | // "every task passed" is a sentence somebody will quote as evidence the app works, |
| 1078 | // when nothing was asked of the app at all. |
| 1079 | let what = if std::env::var("PROBE_SELFTEST").is_ok() { "self-test" } else { "devcycle" }; |
| 1080 | match run() { |
| 1081 | Ok(bad) if bad == 0 => println!("\n{}: nothing failed.", what), |
| 1082 | Ok(bad) => { println!("\n{}: {} failure(s).", what, bad); std::process::exit(1); } |
| 1083 | Err(e) => { eprintln!("{}: {}", what, e); std::process::exit(2); } |
| 1084 | } |
| 1085 | } |
| 1086 | |
| 1087 | fn run() -> Outcome<usize> { |
| 1088 | let repo = repo(); |
| 1089 | if !repo.join("www/js/daimond.js").exists() { |
| 1090 | return Err(err!("{:?} does not look like the Daimond repository.", repo; Init, Invalid)); |
| 1091 | } |
| 1092 | let chosen: Vec<&Task> = match std::env::var("DAIMOND_PROBE_TASKS") { |
| 1093 | Ok(list) if !list.trim().is_empty() => { |
| 1094 | let want: Vec<String> = list.split(',').map(|s| s.trim().to_string()).collect(); |
| 1095 | TASKS.iter().filter(|t| want.iter().any(|w| w == t.name)).collect() |
| 1096 | } |
| 1097 | _ => TASKS.iter().collect(), |
| 1098 | }; |
| 1099 | if chosen.is_empty() { |
| 1100 | return Err(err!("DAIMOND_PROBE_TASKS named no task this file knows."; Init, Invalid)); |
| 1101 | } |
| 1102 | |
| 1103 | if std::env::var("PROBE_SELFTEST").is_ok() { |
| 1104 | return selftest(&repo, &chosen); |
| 1105 | } |
| 1106 | live(&repo, &chosen) |
| 1107 | } |
| 1108 | |
| 1109 | // ── The self-test, which costs nothing and is the first thing to run ────────────────── |
| 1110 | |
| 1111 | /// Prove every check BOTH ways without a provider. |
| 1112 | /// |
| 1113 | /// A check is run against a tree where the task is done and against one where it is not, |
| 1114 | /// and it must answer differently. `probe_details.sh` records what it costs to skip this: |
| 1115 | /// its classifier put a one-character reply in the wrong bucket, and the printed tally said |
| 1116 | /// 9 of 10 when the answer was 9 of 9 -- a true-looking number produced by an instrument |
| 1117 | /// nobody had pointed at a known case. |
| 1118 | /// |
| 1119 | /// The budget arithmetic is proved here too. A budget that cannot be exceeded is not a |
| 1120 | /// budget, and it is the check this whole file rests on. |
| 1121 | fn selftest(repo: &Path, chosen: &[&Task]) -> Outcome<usize> { |
| 1122 | println!("== self-test: no provider, no spend ==\n"); |
| 1123 | let mut bad = 0usize; |
| 1124 | |
| 1125 | for t in chosen { |
| 1126 | // The answer a task's check should accept, and one it must not. Written here |
| 1127 | // rather than derived, because the point is to hand each check a case whose right |
| 1128 | // answer is known independently of the code under test. |
| 1129 | let (good, poor): (String, String) = match t.name { |
| 1130 | "bigfile" => (match first_cell_line(repo) { |
| 1131 | Some((n, _)) => fmt!("{}", n), |
| 1132 | None => fmt!("the file no longer holds that call"), |
| 1133 | }, |
| 1134 | fmt!("I could not read the file; it is binary.")), |
| 1135 | "bigmark" => (fmt!("src/tools.rs, 20_000"), fmt!("No matches found.")), |
| 1136 | "ranit" => (fmt!("i18ncheck: all 7 locales carry every one of en.js's 3512 keys"), |
| 1137 | fmt!("I read the script. It checks locale coverage.")), |
| 1138 | // The two poor answers below are not invented. `parses` is told what |
| 1139 | // `node --check` actually says about that file -- exit 0 and nothing printed -- |
| 1140 | // and `alias` is handed the reply a daimon really gave: the literal caller, alone. |
| 1141 | "parses" => (fmt!("jscheck: 1 file(s) do not parse."), |
| 1142 | fmt!("node --check reported nothing, so the file parses.")), |
| 1143 | "alias" => (bare_totals_scan(repo), fmt!("www/js/daimond.js:16036 is the caller.")), |
| 1144 | "world" => (fmt!("export DAIMOND_SCRATCH={}/.cache/daimond", |
| 1145 | std::env::var("HOME").unwrap_or_default().trim_end_matches('/')), |
| 1146 | fmt!("dev/world.sh: line 84: HOME: unbound variable")), |
| 1147 | // These two read the tree, so their cases are made in it below. |
| 1148 | "locales" | "commit" => (String::new(), String::new()), |
| 1149 | other => return Err(err!("no self-test case for task '{}'", other; Init, Missing)), |
| 1150 | }; |
| 1151 | |
| 1152 | // The commit task's evidence is a repository, so its two cases are a repository with |
| 1153 | // the commit in it and one without. Both are built and thrown away here; nothing |
| 1154 | // tracked is touched, and no git verb below names the repository under test. |
| 1155 | if t.name == "commit" { |
| 1156 | let dir = commit_dir(repo); |
| 1157 | let built = setup_commit(repo); |
| 1158 | match built { |
| 1159 | Err(ref e) => { |
| 1160 | println!(" FAIL {:<9} could not build the fixture: {}", t.name, e); |
| 1161 | bad += 1; |
| 1162 | } |
| 1163 | Ok(()) => { |
| 1164 | match (t.check)(repo, "") { |
| 1165 | Ok(()) => { |
| 1166 | println!(" FAIL {:<9} passes on a repository with no such commit", t.name); |
| 1167 | bad += 1; |
| 1168 | } |
| 1169 | Err(_) => println!(" ok {:<9} red before the commit is made", t.name), |
| 1170 | } |
| 1171 | // And green, with the commit made by hand. |
| 1172 | let done = std::fs::write(dir.join("scratch.txt"), |
| 1173 | SCRATCH_SEED.replace("count = 1", "count = 2")) |
| 1174 | .map_err(|e| fmt!("scratch.txt: {}", e)) |
| 1175 | .and_then(|()| git(&dir, &["add", "scratch.txt"])) |
| 1176 | .and_then(|_| git(&dir, &["commit", "-q", "-m", COMMIT_SUBJECT])); |
| 1177 | let verdict = match done { |
| 1178 | Ok(_) => (t.check)(repo, ""), |
| 1179 | Err(e) => Err(e), |
| 1180 | }; |
| 1181 | // Taken down BEFORE the verdict is printed, as the locale branch does, so a |
| 1182 | // panic in the report cannot leave the fixture behind. |
| 1183 | teardown_commit(repo); |
| 1184 | match verdict { |
| 1185 | Ok(()) => println!(" ok {:<9} green once the commit is there", t.name), |
| 1186 | Err(e) => { |
| 1187 | println!(" FAIL {:<9} still red with the commit made: {}", t.name, e); |
| 1188 | bad += 1; |
| 1189 | } |
| 1190 | } |
| 1191 | } |
| 1192 | } |
| 1193 | teardown_commit(repo); |
| 1194 | continue; |
| 1195 | } |
| 1196 | |
| 1197 | if t.name == "locales" { |
| 1198 | // Red first, against the tree as it stands. |
| 1199 | match (t.check)(repo, "") { |
| 1200 | Ok(()) => { |
| 1201 | println!(" FAIL {:<9} the check passes on a tree where the key is absent", t.name); |
| 1202 | bad += 1; |
| 1203 | } |
| 1204 | Err(_) => println!(" ok {:<9} red on a tree without the key", t.name), |
| 1205 | } |
| 1206 | // Then green, with the key put in every locale by hand and taken out again. |
| 1207 | let dir = repo.join("www/i18n"); |
| 1208 | // The same key the brief will ask for, so the planted case is the task done and not |
| 1209 | // a different task done. |
| 1210 | let (key, _) = match period_key(repo) { |
| 1211 | Some(k) => k, |
| 1212 | None => { |
| 1213 | println!(" FAIL {:<9} every candidate key is already in the tree", t.name); |
| 1214 | bad += 1; |
| 1215 | continue; |
| 1216 | } |
| 1217 | }; |
| 1218 | let mark = fmt!("\n// {} (self-test, removed below)\n", key); |
| 1219 | let mut planted: Vec<PathBuf> = Vec::new(); |
| 1220 | let mut plant_failed = None; |
| 1221 | for n in LOCALES { |
| 1222 | let f = dir.join(n); |
| 1223 | let text = match std::fs::read_to_string(&f) { |
| 1224 | Ok(t) => t, |
| 1225 | Err(e) => { plant_failed = Some(fmt!("{}: {}", n, e)); break; } |
| 1226 | }; |
| 1227 | let with = fmt!("{}{}", text, mark); |
| 1228 | if let Err(e) = std::fs::write(&f, with) { |
| 1229 | plant_failed = Some(fmt!("{}: {}", n, e)); |
| 1230 | break; |
| 1231 | } |
| 1232 | planted.push(f); |
| 1233 | } |
| 1234 | let verdict = match plant_failed { |
| 1235 | Some(e) => Err(e), |
| 1236 | None => (t.check)(repo, ""), |
| 1237 | }; |
| 1238 | // Put the tree back BEFORE reporting, so a panic in the report cannot leave it |
| 1239 | // dirty. `git checkout` is the belt; this is the braces. |
| 1240 | for f in &planted { |
| 1241 | if let Ok(text) = std::fs::read_to_string(f) { |
| 1242 | let back = text.replace(&mark, ""); |
| 1243 | let _ = std::fs::write(f, back); |
| 1244 | } |
| 1245 | } |
| 1246 | match verdict { |
| 1247 | Ok(()) => println!(" ok {:<9} green once every locale carries it", t.name), |
| 1248 | Err(e) => { println!(" FAIL {:<9} still red with the key planted: {}", t.name, e); bad += 1; } |
| 1249 | } |
| 1250 | continue; |
| 1251 | } |
| 1252 | |
| 1253 | match (t.check)(repo, &good) { |
| 1254 | Ok(()) => println!(" ok {:<9} green on a right answer", t.name), |
| 1255 | Err(e) => { println!(" FAIL {:<9} red on a right answer: {}", t.name, e); bad += 1; } |
| 1256 | } |
| 1257 | match (t.check)(repo, &poor) { |
| 1258 | Ok(()) => { println!(" FAIL {:<9} GREEN on a wrong answer, so it checks nothing", t.name); bad += 1; } |
| 1259 | Err(_) => println!(" ok {:<9} red on a wrong answer", t.name), |
| 1260 | } |
| 1261 | } |
| 1262 | |
| 1263 | // Two checks are pointed at a line found in the tree, and the pair above cannot say whether |
| 1264 | // the locator found the right line: both halves would carry the same wrong number. So the |
| 1265 | // line is held against a bound found a different way. |
| 1266 | if chosen.iter().any(|t| t.name == "bigfile") { |
| 1267 | match (first_cell_line(repo), update_spend_span(repo)) { |
| 1268 | (Some((n, text)), Some((from, to))) if n > from && n < to => |
| 1269 | println!(" ok bigfile line {} is inside updateSpend ({}..{}): {}", |
| 1270 | n, from, to, text.chars().take(46).collect::<String>()), |
| 1271 | (Some((n, _)), Some((from, to))) => { |
| 1272 | println!(" FAIL bigfile line {} is outside updateSpend ({}..{})", n, from, to); |
| 1273 | bad += 1; |
| 1274 | } |
| 1275 | _ => { |
| 1276 | println!(" FAIL bigfile updateSpend or the call inside it cannot be found"); |
| 1277 | bad += 1; |
| 1278 | } |
| 1279 | } |
| 1280 | } |
| 1281 | |
| 1282 | // And the budget, which is the assertion the other three rest on. |
| 1283 | println!(); |
| 1284 | let over = Spend { calls: 99, bytes: 1, ..Default::default() }; |
| 1285 | let fat = Spend { calls: 1, bytes: 9_999_999, ..Default::default() }; |
| 1286 | let slow = Spend { calls: 1, bytes: 1, secs: 9_999.0, ..Default::default() }; |
| 1287 | let fine = Spend { calls: 1, bytes: 1, secs: 0.1, ..Default::default() }; |
| 1288 | let t = &TASKS[0]; |
| 1289 | for (label, s, want_over) in [ |
| 1290 | ("too many calls", &over, true), |
| 1291 | ("too many bytes", &fat, true), |
| 1292 | ("too slow", &slow, true), |
| 1293 | ("inside all", &fine, false), |
| 1294 | ] { |
| 1295 | let is_over = s.calls > t.max_calls || s.bytes > t.max_bytes || s.secs > t.max_secs; |
| 1296 | if is_over == want_over { |
| 1297 | println!(" ok budget {} reads as {}", label, |
| 1298 | if want_over { "over" } else { "inside" }); |
| 1299 | } else { |
| 1300 | println!(" FAIL budget {} does not", label); |
| 1301 | bad += 1; |
| 1302 | } |
| 1303 | } |
| 1304 | |
| 1305 | println!(); |
| 1306 | if bad == 0 { |
| 1307 | println!("self-test: every check answers both ways. The instrument is worth spending on."); |
| 1308 | } else { |
| 1309 | println!("self-test: {} check(s) cannot tell a right answer from a wrong one.", bad); |
| 1310 | println!("A check that will not go red is a finding, not a fixture problem. Fix it before spending."); |
| 1311 | } |
| 1312 | Ok(bad) |
| 1313 | } |
| 1314 | |
| 1315 | // ── The live run ───────────────────────────────────────────────────────────────────── |
| 1316 | |
| 1317 | fn live(repo: &Path, chosen: &[&Task]) -> Outcome<usize> { |
| 1318 | let key = res!(std::env::var("DAIMOND_PROBE_KEY").map_err(|_| err!( |
| 1319 | "Set DAIMOND_PROBE_KEY to a provider key. Run PROBE_SELFTEST=1 first -- it is free \ |
| 1320 | and it proves these checks can go red."; |
| 1321 | Init, Missing))); |
| 1322 | let host = std::env::var("DAIMOND_PROBE_HOST") |
| 1323 | .unwrap_or_else(|_| fmt!("openrouter.ai")); |
| 1324 | let path = std::env::var("DAIMOND_PROBE_PATH") |
| 1325 | .unwrap_or_else(|_| fmt!("/api/v1/chat/completions")); |
| 1326 | let model = std::env::var("DAIMOND_PROBE_MODEL") |
| 1327 | .unwrap_or_else(|_| fmt!("anthropic/claude-haiku-4.5")); |
| 1328 | |
| 1329 | println!("== a day's work, {} task(s) ==", chosen.len()); |
| 1330 | println!(" repository {:?}", repo); |
| 1331 | println!(" workspace {:?}", ws_root()); |
| 1332 | println!(" marks {} (app), {} (wide)", app_rel(), wide_rel()); |
| 1333 | println!(" model {}", model); |
| 1334 | println!(" THIS SPENDS: one turn per task, each up to {} tool call(s).", |
| 1335 | chosen.iter().map(|t| t.max_calls).max().unwrap_or(0)); |
| 1336 | if std::env::var("PROBE_YES").is_err() { |
| 1337 | println!("\n Set PROBE_YES=1 to run. Nothing has been called yet."); |
| 1338 | return Ok(0); |
| 1339 | } |
| 1340 | |
| 1341 | let _ = rustls::crypto::ring::default_provider().install_default(); |
| 1342 | let tls = res!(build_tls_client_config()); |
| 1343 | |
| 1344 | let mut bad = 0usize; |
| 1345 | let mut rows: Vec<(String, bool, Spend, String)> = Vec::new(); |
| 1346 | let mut state = repo_state(repo); |
| 1347 | |
| 1348 | for t in chosen { |
| 1349 | // Taken before the task and put back after it, so what the tree held going in is what |
| 1350 | // it holds coming out -- residue from an earlier run included, since with four other |
| 1351 | // people working here there is no way to tell residue from somebody's afternoon. The |
| 1352 | // rot that would otherwise cause is handled where it belongs: `period_key` chooses a |
| 1353 | // key the tree does not already carry. |
| 1354 | let snap = snapshot(repo, t.touches); |
| 1355 | if let Some(build) = t.setup { |
| 1356 | if let Err(e) = build(repo) { |
| 1357 | println!(" {:<9} fixture not built: {}", t.name, e); |
| 1358 | bad += 1; |
| 1359 | rows.push((t.name.to_string(), false, Spend::default(), |
| 1360 | fmt!("the fixture could not be built: {}", e))); |
| 1361 | continue; |
| 1362 | } |
| 1363 | } |
| 1364 | let (ok, spend, why) = one_task(t, repo, &key, &host, &path, &model, tls.clone()); |
| 1365 | revert(repo, t.touches, &snap); |
| 1366 | // After the check, which reads the fixture, and before the next task. |
| 1367 | if let Some(unbuild) = t.teardown { |
| 1368 | unbuild(repo); |
| 1369 | } |
| 1370 | // Four other people have uncommitted work in this tree. A task that moved HEAD or added |
| 1371 | // to the index did it to them, so it is a failure whatever its own verdict was -- and |
| 1372 | // nothing here undoes it, because the undo is what loses the afternoon. |
| 1373 | let now = repo_state(repo); |
| 1374 | let (ok, why) = if now == state { |
| 1375 | (ok, why) |
| 1376 | } else { |
| 1377 | let moved = fmt!("THIS TASK MOVED THE REPOSITORY UNDER TEST — before [{}], after \ |
| 1378 | [{}]. Nothing was undone: other lanes' uncommitted work is in this tree, and a \ |
| 1379 | reset is how that gets lost. Look before fixing it.{}", |
| 1380 | state, now, if why.is_empty() { String::new() } else { fmt!(" Also: {}", why) }); |
| 1381 | state = now; |
| 1382 | (false, moved) |
| 1383 | }; |
| 1384 | if !ok { bad += 1; } |
| 1385 | rows.push((t.name.to_string(), ok, spend, why)); |
| 1386 | } |
| 1387 | |
| 1388 | println!("\n {:<9} {:<7} {:>6} {:>9} {:>5} {:>5} {:>7} {}", |
| 1389 | "task", "verdict", "calls", "bytes", "ref", "fail", "secs", "worst read"); |
| 1390 | println!(" {}", "-".repeat(96)); |
| 1391 | for (name, ok, s, _) in &rows { |
| 1392 | println!(" {:<9} {:<7} {:>6} {:>9} {:>5} {:>5} {:>7.1} {} {}", |
| 1393 | name, if *ok { "pass" } else { "FAIL" }, s.calls, s.bytes, |
| 1394 | s.refused, s.failed, s.secs, s.worst.0, s.worst.1); |
| 1395 | } |
| 1396 | |
| 1397 | let failures: Vec<&(String, bool, Spend, String)> = rows.iter().filter(|r| !r.1).collect(); |
| 1398 | if !failures.is_empty() { |
| 1399 | println!("\n Why each failed, and what that task is for"); |
| 1400 | println!(" {}", "-".repeat(96)); |
| 1401 | for (name, _, _, why) in failures { |
| 1402 | let t = match TASKS.iter().find(|t| &t.name == name) { |
| 1403 | Some(t) => t, |
| 1404 | None => continue, |
| 1405 | }; |
| 1406 | println!(" {}", name); |
| 1407 | println!(" {}", why); |
| 1408 | println!(" catches: {}", t.catches); |
| 1409 | } |
| 1410 | } |
| 1411 | |
| 1412 | // Said on every run, pass or fail. A total that moved is how a regression announces |
| 1413 | // itself, and a run that only prints "pass" cannot show one. |
| 1414 | let calls: usize = rows.iter().map(|r| r.2.calls).sum(); |
| 1415 | let bytes: usize = rows.iter().map(|r| r.2.bytes).sum(); |
| 1416 | let secs: f64 = rows.iter().map(|r| r.2.secs).sum(); |
| 1417 | println!("\n TOTAL {} call(s), {} byte(s) of tool output, {:.0}s.", |
| 1418 | calls, bytes, secs); |
| 1419 | println!(" Compare with the last run rather than with nothing: the numbers are the finding."); |
| 1420 | Ok(bad) |
| 1421 | } |
| 1422 | |
| 1423 | /// Every file under the paths a task will touch, with its bytes, as they stand right now. |
| 1424 | /// |
| 1425 | /// **This replaced a `git checkout --`, and the difference is somebody's afternoon.** Reverting |
| 1426 | /// to HEAD restores the paths to the last commit, which is correct only if nobody else has |
| 1427 | /// uncommitted work in them. On 2026-08-24 four other people were working in this tree and all |
| 1428 | /// eight files under `www/i18n` -- the one path in any task's `touches` -- carried a lane's |
| 1429 | /// unfinished translations. The `locales` task would have run, been reverted, and taken those |
| 1430 | /// with it, and the probe would have printed a clean table over the top. A snapshot restores |
| 1431 | /// what was THERE, which is the only thing an instrument has any business restoring. |
| 1432 | fn snapshot(repo: &Path, touches: &[&str]) -> Vec<(PathBuf, Vec<u8>)> { |
| 1433 | let mut out: Vec<(PathBuf, Vec<u8>)> = Vec::new(); |
| 1434 | for t in touches { |
| 1435 | collect(&repo.join(t), &mut out); |
| 1436 | } |
| 1437 | out |
| 1438 | } |
| 1439 | |
| 1440 | /// Every file at or under `p`, read into `out`. |
| 1441 | fn collect(p: &Path, out: &mut Vec<(PathBuf, Vec<u8>)>) { |
| 1442 | if p.is_file() { |
| 1443 | if let Ok(b) = std::fs::read(p) { |
| 1444 | out.push((p.to_path_buf(), b)); |
| 1445 | } |
| 1446 | return; |
| 1447 | } |
| 1448 | let listing = match std::fs::read_dir(p) { |
| 1449 | Ok(l) => l, |
| 1450 | Err(_) => return, |
| 1451 | }; |
| 1452 | for ent in listing.flatten() { |
| 1453 | collect(&ent.path(), out); |
| 1454 | } |
| 1455 | } |
| 1456 | |
| 1457 | /// Put the snapshot back, and take away anything that was not in it. |
| 1458 | /// |
| 1459 | /// A file is rewritten only where its bytes actually differ, so a task that changed nothing |
| 1460 | /// leaves no modification times behind for the next tool to puzzle over. |
| 1461 | fn revert(repo: &Path, touches: &[&str], snap: &[(PathBuf, Vec<u8>)]) { |
| 1462 | if touches.is_empty() { |
| 1463 | return; |
| 1464 | } |
| 1465 | let mut now: Vec<(PathBuf, Vec<u8>)> = Vec::new(); |
| 1466 | for t in touches { |
| 1467 | collect(&repo.join(t), &mut now); |
| 1468 | } |
| 1469 | for (p, _) in &now { |
| 1470 | if !snap.iter().any(|(q, _)| q == p) { |
| 1471 | let _ = std::fs::remove_file(p); |
| 1472 | } |
| 1473 | } |
| 1474 | for (p, bytes) in snap { |
| 1475 | let same = match std::fs::read(p) { |
| 1476 | Ok(b) => &b == bytes, |
| 1477 | Err(_) => false, |
| 1478 | }; |
| 1479 | if !same { |
| 1480 | if let Some(parent) = p.parent() { |
| 1481 | let _ = std::fs::create_dir_all(parent); |
| 1482 | } |
| 1483 | let _ = std::fs::write(p, bytes); |
| 1484 | } |
| 1485 | } |
| 1486 | } |
| 1487 | |
| 1488 | /// Run one task's turn and measure it. |
| 1489 | fn one_task( |
| 1490 | t: &Task, |
| 1491 | repo: &Path, |
| 1492 | key: &str, |
| 1493 | host: &str, |
| 1494 | path: &str, |
| 1495 | model: &str, |
| 1496 | tls: std::sync::Arc<rustls::ClientConfig>, |
| 1497 | ) |
| 1498 | -> (bool, Spend, String) |
| 1499 | { |
| 1500 | let root = ws_root(); |
| 1501 | let ws = match Workspace::new(root.clone()) { |
| 1502 | Ok(w) => w, |
| 1503 | Err(e) => return (false, Spend::default(), fmt!("workspace at {:?}: {}", root, e)), |
| 1504 | }; |
| 1505 | // The mark, INSIDE the workspace and never equal to it -- see `ws_root`. |
| 1506 | let mark = match t.mark { |
| 1507 | Mark::App => app_rel(), |
| 1508 | Mark::Wide => wide_rel(), |
| 1509 | }; |
| 1510 | // The bounds a Diamond's daimon carries. Built through `diamond_bounds` rather than |
| 1511 | // assembled here, so the probe runs under the rules the product runs under and cannot |
| 1512 | // drift from them -- including the `Nowhere` guard, which is what caught the harness's |
| 1513 | // own first mistake by refusing every call rather than quietly widening. |
| 1514 | let bounds = diamond_bounds("", &[mark], &[]); |
| 1515 | let ctx = ToolContext { |
| 1516 | workspace: ws, |
| 1517 | executor: Executor::local_default(), |
| 1518 | cwd: String::new(), |
| 1519 | path_prefix: String::new(), |
| 1520 | root: FileRoot::Workspace, |
| 1521 | read_seen: oxedyne_daimond::tools::new_read_cache(), |
| 1522 | no_write: bounds, |
| 1523 | daimon_of: String::new(), |
| 1524 | }; |
| 1525 | let registry = ToolRegistry::new(Tool::defaults(), ctx); |
| 1526 | |
| 1527 | let llm = LlmClient::new(host, 443, path, key, model, 4096, tls); |
| 1528 | // The daimon's own prompt, composed the way the app composes it, so the probe measures |
| 1529 | // the instruction the product ships and not a paraphrase of it. |
| 1530 | let system = oxedyne_daimond::prompts::Role::Daimon.compose(""); |
| 1531 | let agent = Agent::new(llm, &system); |
| 1532 | let mut session = Session::new( |
| 1533 | fmt!("devcycle-{}", t.name), t.name.to_string(), model.to_string()); |
| 1534 | |
| 1535 | // The app's own path, spelled as the file tools will see it. Written into the brief at |
| 1536 | // run time rather than hardcoded, so the brief cannot go stale against `ws_root`. |
| 1537 | let brief = t.brief.replace("{app}", &app_rel()); |
| 1538 | // The locale key is chosen from the tree, so the brief cannot ask for something already |
| 1539 | // there. See `period_key`, and the constant that rotted before it. |
| 1540 | let (key, eng) = period_key(repo).unwrap_or(("spend.period_day", "Day")); |
| 1541 | let brief = brief.replace("{key}", key).replace("{eng}", eng); |
| 1542 | let mut s = Spend::default(); |
| 1543 | let mut reply = String::new(); |
| 1544 | let started = Instant::now(); |
| 1545 | { |
| 1546 | let mut on_event = |ev: AgentEvent| { |
| 1547 | match ev { |
| 1548 | AgentEvent::Text(text) => reply.push_str(&text), |
| 1549 | AgentEvent::ToolCall { .. } => s.calls += 1, |
| 1550 | // The OUTCOME, never the prose. `CONTRACT_OUTCOME.md` §3. |
| 1551 | AgentEvent::ToolResult { name, result, outcome } => { |
| 1552 | let (name2, result2) = (name.clone(), result.clone()); |
| 1553 | s.bytes += result.len(); |
| 1554 | if result.len() > s.worst.1 { |
| 1555 | s.worst = (name, result.len()); |
| 1556 | } |
| 1557 | match outcome { |
| 1558 | CallOutcome::Refused => s.refused += 1, |
| 1559 | CallOutcome::Failed => s.failed += 1, |
| 1560 | CallOutcome::Done => {} |
| 1561 | } |
| 1562 | if outcome != CallOutcome::Done && s.firstbad.is_none() { |
| 1563 | s.firstbad = Some((name2, |
| 1564 | result2.chars().take(240).collect::<String>())); |
| 1565 | } |
| 1566 | } |
| 1567 | _ => {} |
| 1568 | } |
| 1569 | }; |
| 1570 | let rt = match tokio::runtime::Runtime::new() { |
| 1571 | Ok(rt) => rt, |
| 1572 | Err(e) => return (false, s, fmt!("runtime: {}", e)), |
| 1573 | }; |
| 1574 | if let Err(e) = rt.block_on( |
| 1575 | agent.run_turn(&mut session, brief, ®istry, &mut on_event)) |
| 1576 | { |
| 1577 | s.secs = started.elapsed().as_secs_f64(); |
| 1578 | return (false, s, fmt!("the turn ended in an error: {}", e)); |
| 1579 | } |
| 1580 | } |
| 1581 | s.secs = started.elapsed().as_secs_f64(); |
| 1582 | s.prompt = session.prompt_tokens as u64; |
| 1583 | s.completion = session.completion_tokens as u64; |
| 1584 | |
| 1585 | // (1) It got there. |
| 1586 | if let Err(e) = (t.check)(repo, &reply) { |
| 1587 | return (false, s, fmt!("did not get there — {}", e)); |
| 1588 | } |
| 1589 | // (2) It did not flail. A right answer reached through refusals found a fault and |
| 1590 | // worked around it, which is the thing this probe exists to stop happening silently. |
| 1591 | if s.refused > 0 || s.failed > 0 { |
| 1592 | let named = match &s.firstbad { |
| 1593 | Some((n, r)) => fmt!(" — first was {}: {}", n, r.replace('\n', " ")), |
| 1594 | None => String::new(), |
| 1595 | }; |
| 1596 | let why = fmt!("got there through {} refused and {} failed call(s){}", |
| 1597 | s.refused, s.failed, named); |
| 1598 | return (false, s, why); |
| 1599 | } |
| 1600 | // (3) It stayed inside the budget. |
| 1601 | if s.calls > t.max_calls { |
| 1602 | let why = fmt!("{} tool calls against a budget of {}", s.calls, t.max_calls); |
| 1603 | return (false, s, why); |
| 1604 | } |
| 1605 | if s.bytes > t.max_bytes { |
| 1606 | let why = fmt!("{} bytes of tool output against a budget of {} -- worst was {} at {}", |
| 1607 | s.bytes, t.max_bytes, s.worst.0, s.worst.1); |
| 1608 | return (false, s, why); |
| 1609 | } |
| 1610 | if s.secs > t.max_secs { |
| 1611 | let why = fmt!("{:.1}s against a budget of {:.0}s", s.secs, t.max_secs); |
| 1612 | return (false, s, why); |
| 1613 | } |
| 1614 | (true, s, String::new()) |
| 1615 | } |