oxedyne/daimond/dev/verify_vision.mjs
46.2 KiB, 1 run
created by r2519314175:791, which is this file's identity for as long as the history lasts, whatever it is later renamed to
download · who wrote it · its history
| 1 | // verify_vision.mjs — a worker that is shown a picture ends up on a model that can |
| 2 | // see it, is billed for both halves, and says so on screen. |
| 3 | // |
| 4 | // WRITTEN BEFORE THE FIX, AND EXPECTED TO BE RED. Defect AA in |
| 5 | // dev/DEFECTS_20260821.md and the design in dev/DESIGN_VISION.md; nothing in §7 of |
| 6 | // that design is implemented at the time of writing. A verifier written after a fix |
| 7 | // asserts what the fix happens to do. This one asserts what was wanted, so the checks |
| 8 | // below are the specification and their failures are the defect, quoted. |
| 9 | // |
| 10 | // THE DEFECT. `Workers.routeFor` decides which model a worker runs on with |
| 11 | // `sees = !supplied && taskWantsVision(task)`, and `taskWantsVision` is nothing but |
| 12 | // `IMAGE_EXT.test(task)` — four extensions, matched against the TASK TEXT. So a worker |
| 13 | // that will certainly be shown a picture runs on the text model unless the daimon |
| 14 | // happened to spell a filename into the task; it then reads the image, the provider |
| 15 | // refuses it, and `LlmClient` quietly takes the pictures out and carries on answering |
| 16 | // about something it cannot see. That is what cost A$17 on 2026-08-20, and the only |
| 17 | // disclosure of it was a `title` attribute. |
| 18 | // |
| 19 | // FOUND WHILE BUILDING THIS, AND WORSE THAN THE DEFECT AS WRITTEN: the spelling rule |
| 20 | // does not run at all. `routeFor`'s first line is |
| 21 | // |
| 22 | // var sees = !supplied && taskWantsVision(task); |
| 23 | // |
| 24 | // and `supplied` is `!!(pick && pick.model)` where `pick` is `dispatch`'s fifth |
| 25 | // argument. `dispatch` has exactly two callers — the daimon's, which passes |
| 26 | // `diamondWorkerModel(diamondId)`, and the chat's, which passes |
| 27 | // `chatWorkerModel(chat)` — and neither can return a pair without a model. So |
| 28 | // `supplied` is TRUE on every dispatch that exists, `sees` is always false, and |
| 29 | // `taskWantsVision`, `IMAGE_EXT` and `diamondVisionModel` decide nothing: `vm` is |
| 30 | // resolved, passed in, and never read. A task that DOES spell `shot.png` goes to the |
| 31 | // text model exactly like one that does not. |
| 32 | // |
| 33 | // It was not noticed because `dev/verify_diamondmodels.mjs:500-540` tests the rule |
| 34 | // through `Workers.routeForDiamond`, which hardcodes `supplied` to FALSE — a path |
| 35 | // production never takes. `routeForDiamond` is called by nothing under `www/` at all; |
| 36 | // its only callers are those two lines of that verifier, and its own doc comment says |
| 37 | // so: "What a verifier drives, and what a future settings preview would." The rule is |
| 38 | // correct, tested twice over, and reachable only from the test harness. BUILT BUT |
| 39 | // UNREACHABLE, and "true of the code, wrong about the reader", in one object. |
| 40 | // |
| 41 | // The measurement, rather than the argument: `--break always` was first written as |
| 42 | // `var sees = !supplied;` and changed nothing at all — no check moved, the worker still |
| 43 | // ran on the text model, and `run.sees` was still false. A patch that cannot change the |
| 44 | // answer proves the term it patched was already dead. |
| 45 | // |
| 46 | // CHECK 7 BELOW IS THE ONE THAT WOULD HAVE CAUGHT THIS, and it is why there are seven |
| 47 | // rather than the six DESIGN_VISION.md asks for. It puts a task naming a picture |
| 48 | // through the door `dispatch` actually opens, rather than through the rule's own |
| 49 | // method. The two green tests that cover this today both drive doors production never |
| 50 | // opens, which is exactly how a rule stays correct and dead for as long as it likes. |
| 51 | // |
| 52 | // SEVEN PROPERTIES — six from DESIGN_VISION.md §7 "Proof", and one the building of it |
| 53 | // turned up: |
| 54 | // |
| 55 | // 1. THE WORK MOVED TO THE IMAGE MODEL. The second leg names the Diamond's image |
| 56 | // model and carries a picture; the first leg never reached that model, and the |
| 57 | // picture it did carry was refused. |
| 58 | // 2. AND IT MOVED ONCE. A Diamond whose IMAGE model is itself blind does not |
| 59 | // ping-pong: two legs, and no third. |
| 60 | // 3. AND EACH LEG IS BILLED TO THE MODEL THAT SPENT IT, both against one Diamond. |
| 61 | // 4. AND IT IS DISCLOSED ON SCREEN, WITHOUT HOVERING. The check that makes 1-3 mean |
| 62 | // something: a re-route nobody can see is the silent fallback again with more |
| 63 | // machinery behind it. |
| 64 | // 5. AND THE SECOND LEG IS A RESUME, not a restart from the top. |
| 65 | // 6. BUT A TASK CARRYING NO PICTURE IS UNTOUCHED. The check that stops "always |
| 66 | // switch to the image model" from passing 1-5. |
| 67 | // 7. AND A TASK THAT NAMES A PICTURE REACHES THE IMAGE MODEL THROUGH THE DOOR |
| 68 | // `dispatch` ACTUALLY USES. Not one of the design's six, and not fixed by the |
| 69 | // design either: §7 adds a capability re-route and never touches `supplied`, so |
| 70 | // this one stays red after that fix lands. That is the point of it — a check that |
| 71 | // survives the approved design is the check that says the design is incomplete. |
| 72 | // |
| 73 | // ── THE BREAKS, AND WHICH OF THEM CAN BE PROVED TODAY ───────────────────────── |
| 74 | // |
| 75 | // House discipline is that a check is proved by reddening it on purpose first. Four |
| 76 | // of these six breaks patch lines the fix has not written yet, so they cannot be |
| 77 | // applied and REFUSE, naming the line they wanted. That is deliberate: the break table |
| 78 | // is the other half of the specification, and a break that silently patched nothing |
| 79 | // would be worse than one that says it could not. |
| 80 | // |
| 81 | // TWO GATES, not one, and the second was earned. `bill-after`'s anchor — `recordSpend` |
| 82 | // in the worker's `finally` — is in the file TODAY, so the anchor test passed it and it |
| 83 | // would have patched in an `if (run._reroute)` that is false for ever: a run identical |
| 84 | // to a plain one, reported under a break name. So a break may also declare what its |
| 85 | // PATCH depends on, and is refused when that is missing. |
| 86 | // |
| 87 | // node dev/verify_vision.mjs --break spelling 1 (and 2-5) — routing by IMAGE_EXT |
| 88 | // alone, which is the tree as it |
| 89 | // stands: patches NOTHING and is |
| 90 | // simply a plain run under a name. |
| 91 | // node dev/verify_vision.mjs --break always 6 — every worker is routed to the |
| 92 | // image model. LIVE TODAY: it patches |
| 93 | // `routeFor`, which exists. |
| 94 | // node dev/verify_vision.mjs --break loop 2 — REFUSES until the fix lands. |
| 95 | // node dev/verify_vision.mjs --break bill-after 3 — REFUSES until the fix lands. |
| 96 | // node dev/verify_vision.mjs --break silent 4 — REFUSES until the fix lands. |
| 97 | // node dev/verify_vision.mjs --break restart 5 — REFUSES until the fix lands. |
| 98 | // |
| 99 | // Check 7 has no break of its own, because it is red on the tree as shipped: a break |
| 100 | // for it would be a patch that made it GREEN, which is a fix and not a break. |
| 101 | // |
| 102 | // One break does not redden one check. `always` reddens 6 AND 1, because routing is |
| 103 | // upstream of everything; `spelling` reddens 1 through 5, and 7 with them. |
| 104 | // DESIGN_VISION.md's table says otherwise and is wrong about it — see the report |
| 105 | // accompanying this file. |
| 106 | // |
| 107 | // ── WHAT THIS FILE CANNOT YET DISTINGUISH ──────────────────────────────────── |
| 108 | // |
| 109 | // Check 4 looks for a DISCLOSURE on the worker tile and the fix has not written one, |
| 110 | // so it cannot name a class. It therefore asserts a property no class name can dodge — |
| 111 | // the tile's own visible text names both the model the worker left and the model it |
| 112 | // moved to — and it prints the tile's whole visible text and every `title` attribute |
| 113 | // in it when it fails, so the failure reads as "the tile says nothing; the only |
| 114 | // mention of the image model is a tooltip" rather than "selector not found". If an |
| 115 | // implementer discloses the move by naming ONLY the destination model, this check will |
| 116 | // redden on a fix that is arguably correct: naming the destination alone does not tell |
| 117 | // a reader that anything moved, which is why it is asked for, but it is a judgement |
| 118 | // and it is written down here rather than hidden in a selector. |
| 119 | // |
| 120 | // Check 5 asserts the SHAPE of the second leg's first request — that its last user |
| 121 | // message is not the bare task, and that it carries the worker's own earlier words. |
| 122 | // The design's stronger sentence, "a worker that already wrote a file must not write |
| 123 | // it twice", is not asserted, and cannot honestly be: whether a model repeats work it |
| 124 | // can see it already did is a property of the model, and this one is a script. What IS |
| 125 | // app-level is restart-versus-resume, and that is what is measured. |
| 126 | // |
| 127 | // AND UNTIL SOMETHING MOVES A WORKER, CHECK 5 CANNOT SPEAK FOR ITSELF. There is no |
| 128 | // second leg to look at, so it reports that, which is honest but is a consequence of |
| 129 | // check 1 rather than a measurement of its own. It becomes a real check the moment the |
| 130 | // re-route exists, and not before. |
| 131 | // |
| 132 | // eval "$(bash dev/world.sh 5 --up)" |
| 133 | // node dev/verify_vision.mjs |
| 134 | // |
| 135 | // Needs dev/serve.mjs and dev/mockllm.mjs — the mock with model-name blindness |
| 136 | // (`mock/blind` refuses a picture, `mock/eyes` takes it, at ONE endpoint). The |
| 137 | // preflight refuses if the mock answering this world is an older one, because every |
| 138 | // check below would then be measuring a fixture that is not there. No gateway, no wasm |
| 139 | // rebuild: the branch under test is JavaScript. |
| 140 | |
| 141 | import fs from 'node:fs'; |
| 142 | import path from 'node:path'; |
| 143 | import { fileURLToPath } from 'node:url'; |
| 144 | import { open, steerDiamond, scratch, shot, mockLog, clearMockLog, MOCK } from './harness.mjs'; |
| 145 | |
| 146 | const HERE = path.dirname(fileURLToPath(import.meta.url)); |
| 147 | const WWW = path.join(HERE, '..', 'www'); |
| 148 | const SRC = fs.readFileSync(path.join(WWW, 'js/daimond.js'), 'utf8'); |
| 149 | const MOCK_LOG_PATH = process.env.DAIMOND_MOCK_LOG || path.join(HERE, 'mockllm.log'); |
| 150 | |
| 151 | // The two blind models and the sighted one. `mock/blind-too` is never listed by the |
| 152 | // mock's catalogue and does not need to be: blindness is a property of the NAME, and |
| 153 | // `DaimondModels.resolve` will resolve any id under a provider that holds a key. |
| 154 | const TEXT_MODEL = 'mock/blind'; |
| 155 | const VISION_MODEL = 'mock/eyes'; |
| 156 | const BLIND_VISION = 'mock/blind-too'; |
| 157 | |
| 158 | let bad = 0; |
| 159 | const check = (pass, name, detail) => { |
| 160 | if (!pass) bad++; |
| 161 | console.log((pass ? ' ok ' : ' FAIL ') + name + (detail != null ? ' — ' + detail : '')); |
| 162 | }; |
| 163 | |
| 164 | // ── The breaks ─────────────────────────────────────────────────────── |
| 165 | // |
| 166 | // Each is an anchor in js/daimond.js and what to put in its place, served over the |
| 167 | // real file through `page.route` so the page under test is the shipped page with one |
| 168 | // line changed. An anchor that is not in the file EXACTLY ONCE stops the run: a break |
| 169 | // that patches nothing would produce a green page under a red name, which is the |
| 170 | // failure this whole file exists to avoid. |
| 171 | const BREAKS = { |
| 172 | // The tree as it stands. There is nothing to patch, because the defect is the |
| 173 | // current behaviour: routing by spelling and no re-route at all. |
| 174 | spelling: { |
| 175 | anchor: 'var sees = !supplied && taskWantsVision(task);', |
| 176 | patched: 'var sees = !supplied && taskWantsVision(task);', |
| 177 | why: 'the routing rule as shipped — this break patches NOTHING and is the ' |
| 178 | + 'defect itself, run under a name so the table in DESIGN_VISION.md has a ' |
| 179 | + 'row that can be executed', |
| 180 | }, |
| 181 | // The plausible over-correction: send every worker to the image model. Checks 1-5 |
| 182 | // would pass and the bill would double for the nineteen workers in twenty that |
| 183 | // never look at anything. |
| 184 | always: { |
| 185 | anchor: 'var sees = !supplied && taskWantsVision(task);', |
| 186 | patched: 'var sees = true;', |
| 187 | why: 'every worker is routed to the image model, whatever it was asked to do', |
| 188 | // `!supplied` was the first spelling of this break and it changed NOTHING, which |
| 189 | // is how the finding below was made: `supplied` is true on every dispatch there |
| 190 | // is, so the whole right-hand side of that line is already dead. The break has to |
| 191 | // overwrite the lot. |
| 192 | }, |
| 193 | // ── Below here the anchors belong to code that does not exist yet ── |
| 194 | // |
| 195 | // Each names the line DESIGN_VISION.md §7 item 5 asks for. They will refuse until |
| 196 | // that line is written, and the refusal names it, so the implementer is told what |
| 197 | // this file will patch rather than being left to guess. |
| 198 | loop: { |
| 199 | anchor: 'if (run._reroute) return;', |
| 200 | patched: 'if (false) return;', |
| 201 | why: 'the once-per-run guard on the re-route is dropped, so a blind image ' |
| 202 | + 'model is moved to again and again', |
| 203 | }, |
| 204 | 'bill-after': { |
| 205 | // ITS ANCHOR IS ALREADY IN THE FILE — `recordSpend`'s call in the worker |
| 206 | // `finally` is live today — so the anchor test alone lets this break through and |
| 207 | // it then patches in an `if (run._reroute)` that is false for ever. It would run |
| 208 | // clean, look exactly like a plain run, and report check 3 red under a break name |
| 209 | // that had done nothing. `needs` is the second gate: the break is refused until |
| 210 | // the field its patch depends on exists. |
| 211 | needs: 'run._reroute', |
| 212 | anchor: 'recordSpend(run.model, _pt, _ct, _ca, _cost, run.provider, run.diamondId || \'\');', |
| 213 | patched: 'if (run._reroute) { run.model = run._reroute.to.model; run.provider = run._reroute.to.provider; }\n\t\t\t\t' |
| 214 | + 'recordSpend(run.model, _pt, _ct, _ca, _cost, run.provider, run.diamondId || \'\');', |
| 215 | why: 'the model is swapped BEFORE the spend is recorded, so the wasted first ' |
| 216 | + 'leg is billed to the image model — the same class of lie as defect A', |
| 217 | }, |
| 218 | silent: { |
| 219 | anchor: 'run.reroutedFrom', |
| 220 | patched: 'run.__no_such_field', |
| 221 | why: 'the tile stops knowing it was re-routed, so the disclosure falls back ' |
| 222 | + 'to the tooltip it was before', |
| 223 | }, |
| 224 | restart: { |
| 225 | anchor: 'Workers.reroute', |
| 226 | patched: 'Workers.__restartFromTheTop', |
| 227 | why: 'the second leg is built fresh and run from the task again, repeating ' |
| 228 | + 'every side effect the first leg had already had', |
| 229 | }, |
| 230 | }; |
| 231 | |
| 232 | const BREAK = (() => { |
| 233 | const eq = process.argv.find(a => a.startsWith('--break=')); |
| 234 | if (eq) return eq.split('=')[1]; |
| 235 | const i = process.argv.indexOf('--break'); |
| 236 | return i > 0 ? String(process.argv[i + 1] || '') : ''; |
| 237 | })(); |
| 238 | if (BREAK && !BREAKS[BREAK]) { |
| 239 | console.error(`unknown break '${BREAK}'; one of: ${Object.keys(BREAKS).join(', ')}`); |
| 240 | process.exit(2); |
| 241 | } |
| 242 | if (BREAK) { |
| 243 | const b = BREAKS[BREAK]; |
| 244 | // An anchor that is present is not the same as a break that can bite: a patch may |
| 245 | // depend on a field the fix has not written. Both gates, and the same refusal. |
| 246 | if (b.needs && !SRC.includes(b.needs)) { |
| 247 | console.error(`\n REFUSED --break ${BREAK} would patch nothing.`); |
| 248 | console.error(` Its anchor is in the file, but the patch depends on ${JSON.stringify(b.needs)},`); |
| 249 | console.error(' which is not. Applying it would produce a run indistinguishable from a'); |
| 250 | console.error(' plain one, under a break name — the exact thing this file exists to stop.'); |
| 251 | console.error(` What it is for: ${b.why}.\n`); |
| 252 | process.exit(2); |
| 253 | } |
| 254 | const n = SRC.split(b.anchor).length - 1; |
| 255 | if (n !== 1) { |
| 256 | console.error(`\n REFUSED --break ${BREAK} cannot be applied.`); |
| 257 | console.error(` It patches: ${JSON.stringify(b.anchor)}`); |
| 258 | console.error(` which is in www/js/daimond.js ${n} time(s), not once.`); |
| 259 | console.error(` What it is for: ${b.why}.`); |
| 260 | console.error(' If this is one of the four breaks that belong to the vision fix,'); |
| 261 | console.error(' the fix has not landed yet and there is nothing to break. Running'); |
| 262 | console.error(' the file with no --break measures the defect, which is the point'); |
| 263 | console.error(' of it today.\n'); |
| 264 | process.exit(2); |
| 265 | } |
| 266 | } |
| 267 | |
| 268 | // ── The routing rule, read out of the app rather than copied ───────── |
| 269 | // |
| 270 | // Check 1's whole premise is that its task names NO image, so today's rule sends the |
| 271 | // worker to the text model and only a capability signal could move it. A copy of the |
| 272 | // regex here would go stale the moment somebody widened the real one — which §5 of the |
| 273 | // design argues against and somebody will eventually try anyway — and this file would |
| 274 | // then be proving something about a rule the app no longer uses. |
| 275 | const IMAGE_EXT = (() => { |
| 276 | const m = /var IMAGE_EXT = (\/.*\/[a-z]*);/.exec(SRC); |
| 277 | if (!m) { |
| 278 | console.error('IMAGE_EXT is not in www/js/daimond.js where this file reads it. ' |
| 279 | + 'Check 1 cannot state its own premise, so nothing below would mean anything.'); |
| 280 | process.exit(2); |
| 281 | } |
| 282 | return new Function('return ' + m[1])(); |
| 283 | })(); |
| 284 | |
| 285 | // ── Preflight: is the mock this run reads the mock this run drives, and is it |
| 286 | // the one that knows how to be blind? ─────────────────────────────── |
| 287 | // |
| 288 | // Both halves, because either alone is the failure this suite has been burned by. |
| 289 | // A mock writing to a log nobody reads makes every check below report "the model was |
| 290 | // never called"; a mock from before this fixture existed answers 200 to a picture on |
| 291 | // `mock/blind`, so the worker never learns anything, no event is ever emitted, and six |
| 292 | // checks fail with a story about the app that is entirely about the fixture. |
| 293 | { |
| 294 | const before = mockLog().length; |
| 295 | let why = ''; |
| 296 | try { |
| 297 | const r = await fetch(MOCK, { |
| 298 | method: 'POST', |
| 299 | headers: { 'content-type': 'application/json' }, |
| 300 | body: JSON.stringify({ model: 'mock/fast', stream: false, |
| 301 | messages: [{ role: 'user', content: 'vision-preflight-probe' }] }), |
| 302 | }); |
| 303 | if (!r.ok) why = `the mock answered ${r.status}`; |
| 304 | else { |
| 305 | await r.text(); |
| 306 | if (mockLog().length <= before) why = 'the probe was not written to the log'; |
| 307 | } |
| 308 | } catch (e) { |
| 309 | why = 'the mock could not be reached: ' + e.message; |
| 310 | } |
| 311 | if (!why) { |
| 312 | // The fixture itself: a picture on a blind model must come back 400, in the |
| 313 | // provider's own words about `image_url`, because that is the shape |
| 314 | // `LlmClient::stream_turn` learns blindness from. |
| 315 | try { |
| 316 | const r = await fetch(MOCK, { |
| 317 | method: 'POST', |
| 318 | headers: { 'content-type': 'application/json' }, |
| 319 | body: JSON.stringify({ model: TEXT_MODEL, stream: false, messages: [{ |
| 320 | role: 'user', content: [ |
| 321 | { type: 'text', text: 'look' }, |
| 322 | { type: 'image_url', image_url: { url: 'data:image/png;base64,AAAA' } }, |
| 323 | ] }] }), |
| 324 | }); |
| 325 | const body = await r.text(); |
| 326 | if (r.status !== 400 || !/image_url/.test(body)) { |
| 327 | why = `${TEXT_MODEL} answered ${r.status} to a picture instead of refusing it — ` |
| 328 | + 'this mock predates the vision fixture in dev/mockllm.mjs'; |
| 329 | } |
| 330 | } catch (e) { |
| 331 | why = 'the blindness probe could not be sent: ' + e.message; |
| 332 | } |
| 333 | } |
| 334 | if (why) { |
| 335 | console.log(' REFUSED ' + why); |
| 336 | console.log(` driving: ${MOCK}`); |
| 337 | console.log(` reading: ${MOCK_LOG_PATH}`); |
| 338 | console.log(' Nothing below could measure the app: it would measure the fixture.'); |
| 339 | console.log(' `bash dev/world.sh N --down` then --up, from THIS worktree.'); |
| 340 | process.exit(2); |
| 341 | } |
| 342 | } |
| 343 | |
| 344 | // ── Reading the wire ───────────────────────────────────────────────── |
| 345 | |
| 346 | /// Whether a logged request is a WORKER's turn on `task` rather than the daimon's own |
| 347 | /// turn about it. |
| 348 | /// |
| 349 | /// The daimon's transcript quotes the task inside its `spawn_agent` call, so every |
| 350 | /// round of the dispatching turn mentions it. A worker's turn is the one where the task |
| 351 | /// IS a user message — and a RESUMED worker's session is seeded with that same user |
| 352 | /// message, which is exactly why it is the right test: both legs of one worker match, |
| 353 | /// and the daimon never does. |
| 354 | const isWorkerTurn = (e, task) => (e.messages || []).some( |
| 355 | (m) => m.role === 'user' && typeof m.content === 'string' && m.content.trim() === task); |
| 356 | |
| 357 | /// Every request one worker's task produced, oldest first. |
| 358 | const turnsFor = (task) => mockLog().filter((e) => isWorkerTurn(e, task)); |
| 359 | |
| 360 | /// The legs of one worker: consecutive requests grouped by the model they named. |
| 361 | /// |
| 362 | /// A leg is a session on one model. Two legs is a re-route; one is the defect; three |
| 363 | /// or more is a ping-pong. Grouping by RUNS of the same model rather than by the set of |
| 364 | /// models is what tells the last two apart. |
| 365 | const legsOf = (task) => { |
| 366 | const out = []; |
| 367 | for (const e of turnsFor(task)) { |
| 368 | const last = out[out.length - 1]; |
| 369 | if (last && last.model === e.model) last.reqs.push(e); |
| 370 | else out.push({ model: e.model, reqs: [e] }); |
| 371 | } |
| 372 | return out.map((l) => ({ |
| 373 | model: l.model, |
| 374 | reqs: l.reqs, |
| 375 | images: l.reqs.reduce((n, e) => n + (e.images || 0), 0), |
| 376 | refused: l.reqs.filter((e) => e.refusedImages).length, |
| 377 | })); |
| 378 | }; |
| 379 | |
| 380 | const legLine = (task) => legsOf(task) |
| 381 | .map((l, i) => `leg ${i + 1}: ${l.model}, ${l.reqs.length} req, ${l.images} image(s)` |
| 382 | + (l.refused ? `, ${l.refused} refused` : '')) |
| 383 | .join(' | ') || 'no worker turn reached the mock at all'; |
| 384 | |
| 385 | /// Wait for a condition, polling. |
| 386 | const until = async (fn, ms = 60000) => { |
| 387 | const t0 = Date.now(); |
| 388 | while (Date.now() - t0 < ms) { |
| 389 | if (await fn()) return true; |
| 390 | await new Promise((r) => setTimeout(r, 300)); |
| 391 | } |
| 392 | return false; |
| 393 | }; |
| 394 | |
| 395 | // ── The session ────────────────────────────────────────────────────── |
| 396 | |
| 397 | const s = await open({ |
| 398 | name: 'vision', |
| 399 | profile: scratch('pw', 'vision' + (BREAK ? '-' + BREAK : '')), |
| 400 | // The damaged file in place of the real one, registered before `goto`. |
| 401 | route: (BREAK && BREAKS[BREAK].patched !== BREAKS[BREAK].anchor) ? (async (page) => { |
| 402 | const body = SRC.replace(BREAKS[BREAK].anchor, BREAKS[BREAK].patched); |
| 403 | await page.route('**/js/daimond.js', (r) => r.fulfill({ |
| 404 | status: 200, contentType: 'application/javascript', body, |
| 405 | })); |
| 406 | }) : null, |
| 407 | }); |
| 408 | const p = s.page; |
| 409 | if (BREAK) { |
| 410 | console.log(`\n*** RUNNING UNDER --break ${BREAK}: ${BREAKS[BREAK].why}.`); |
| 411 | if (BREAKS[BREAK].patched === BREAKS[BREAK].anchor) { |
| 412 | console.log('*** This break patches nothing: it IS the shipped behaviour.'); |
| 413 | } |
| 414 | console.log('*** Failures below are the point ***\n'); |
| 415 | } |
| 416 | |
| 417 | /// A 2x3 PNG, byte for byte the one dev/verify_fileview.mjs uses, so what the app |
| 418 | /// sniffs is a real picture and not a signature with filler behind it. |
| 419 | const PNG = [ |
| 420 | 0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a, 0x00, 0x00, 0x00, 0x0d, |
| 421 | 0x49, 0x48, 0x44, 0x52, 0x00, 0x00, 0x00, 0x02, 0x00, 0x00, 0x00, 0x03, |
| 422 | 0x08, 0x02, 0x00, 0x00, 0x00, 0x36, 0x88, 0x49, 0xd6, 0x00, 0x00, 0x00, |
| 423 | 0x10, 0x49, 0x44, 0x41, 0x54, 0x78, 0xda, 0x63, 0xf8, 0xcf, 0x00, 0x04, |
| 424 | 0xff, 0x19, 0x50, 0x28, 0x00, 0x3e, 0xd6, 0x05, 0xfb, 0xb6, 0xd6, 0xf9, |
| 425 | 0xda, 0x00, 0x00, 0x00, 0x00, 0x49, 0x45, 0x4e, 0x44, 0xae, 0x42, 0x60, |
| 426 | 0x82, |
| 427 | ]; |
| 428 | |
| 429 | /// Put bytes in the workspace through the engine's own door, as dev/verify_fileview |
| 430 | /// does: it applies the path jail and the per-account namespace, and a hand-rolled |
| 431 | /// OPFS walk would write somewhere the worker's fence does not reach. |
| 432 | const put = (file, bytes) => p.evaluate(async ({ file, bytes }) => { |
| 433 | const m = await import('/pkg/oxedyne_daimond.js'); |
| 434 | const app = new m.DaimondApp('http://127.0.0.1/v1/chat/completions', '', 'none', 4096, '', true); |
| 435 | await app.write_bytes(file, new Uint8Array(bytes)); |
| 436 | }, { file, bytes }); |
| 437 | |
| 438 | /// Shut the Admin drawer if it is open. |
| 439 | /// |
| 440 | /// It opens over the rail, so the `+` buttons sit UNDER it and a click on one is |
| 441 | /// swallowed by `.admin-drawer-head` — which Playwright reports as a timeout on a |
| 442 | /// button it can plainly see. It matters on the SECOND run of a fixed profile, because |
| 443 | /// the drawer's state is remembered: run one is fresh and green, run two dies on a |
| 444 | /// button, and the difference is not in the app. |
| 445 | /// |
| 446 | /// THE HARNESS ALREADY KNEW. `harness.newChat` has closed this drawer, with the reason |
| 447 | /// written beside it, since long before this file existed; this file reached past the |
| 448 | /// harness for the rail and re-met the trap on its own. Reach for the harness first — |
| 449 | /// it is the same failure as writing a helper without searching for the machinery that |
| 450 | /// already does the job, and that one was walked past twice more the same night. |
| 451 | const closeAdmin = async () => { |
| 452 | const btn = p.locator('#admin-close'); |
| 453 | if (await btn.isVisible().catch(() => false)) { |
| 454 | await btn.click({ force: true }); |
| 455 | await p.waitForTimeout(300); |
| 456 | } |
| 457 | }; |
| 458 | |
| 459 | /// Make a Diamond through the real dialog and answer its id. |
| 460 | const newDiamond = async (name) => { |
| 461 | await closeAdmin(); |
| 462 | await p.click('#new-diamond-btn'); |
| 463 | await p.waitForSelector('.dlg-input', { timeout: 8000 }); |
| 464 | await p.fill('.dlg-input', name); |
| 465 | await p.click('.dlg-ok'); |
| 466 | await p.waitForTimeout(1200); |
| 467 | return p.evaluate(() => { |
| 468 | const f = window.DaimondAttach && window.DaimondAttach.focus(); |
| 469 | return (f && f.kind === 'diamond') ? String(f.id) : ''; |
| 470 | }); |
| 471 | }; |
| 472 | |
| 473 | /// Put a Diamond back in focus, the way a person does: by clicking its rail box. |
| 474 | const selectDiamond = async (id) => { |
| 475 | await closeAdmin(); |
| 476 | await p.$$eval('.diamond-box', (els, want) => { |
| 477 | for (const e of els) if (e.dataset.id === want) { e.click(); return; } |
| 478 | }, id); |
| 479 | await p.waitForTimeout(900); |
| 480 | }; |
| 481 | |
| 482 | /// Give a Diamond its three models. |
| 483 | /// |
| 484 | /// Written straight into the record `setDiamondModel` owns, rather than driven through |
| 485 | /// the tile dialog's "Workers, images" pulldown. That pulldown is not what this file is |
| 486 | /// about, and dev/verify_diamondmodels.mjs already drives it; what matters here is that |
| 487 | /// the record downstream reads is the one the user's choice produces, which is why the |
| 488 | /// shape is written out in full and read back. |
| 489 | const setModels = (id, rec) => p.evaluate(({ id, rec }) => { |
| 490 | const all = JSON.parse(localStorage.getItem('daimond-diamond-models') || '{}'); |
| 491 | all[id] = rec; |
| 492 | localStorage.setItem('daimond-diamond-models', JSON.stringify(all)); |
| 493 | return all[id]; |
| 494 | }, { id, rec }); |
| 495 | |
| 496 | /// The worker run the app persisted for `task`, as it sees it. |
| 497 | const runFor = (task) => p.evaluate((task) => { |
| 498 | let box = {}; |
| 499 | try { box = JSON.parse(localStorage.getItem('daimond-workers') || '{}'); } catch (e) { return null; } |
| 500 | const runs = (box && box.runs) || []; |
| 501 | return runs.find((r) => (r.task || '').trim() === task) || null; |
| 502 | }, task); |
| 503 | |
| 504 | /// Install the visibility predicate on the page, once, as `window.__visionShown`. |
| 505 | /// |
| 506 | /// VISIBILITY IS ASSERTED PROPERLY, and this is the part of the file that exists |
| 507 | /// because `verify_view.mjs:271` went vacuous (defect K): a chip scrolled out of a |
| 508 | /// horizontal scroller still returns a rect with area, and content under |
| 509 | /// `content-visibility: hidden` keeps its last layout. So an element counts as shown |
| 510 | /// only when its CENTRE lies inside the intersection of every clipping ancestor's box |
| 511 | /// and the viewport, AND a hit test at that centre lands on it or inside it. Both, not |
| 512 | /// either: containment says it is not scrolled away, the hit test says nothing is drawn |
| 513 | /// over it, and neither says anything about the other. |
| 514 | /// |
| 515 | /// One installation shared by the checks and by the instrument's own self-test, so what |
| 516 | /// the self-test proves is the predicate the checks then use and not a copy of it. |
| 517 | const installShown = () => p.evaluate(() => { |
| 518 | // The intersection of the viewport with every ancestor that clips. |
| 519 | const clipRect = (el) => { |
| 520 | let r = { l: 0, t: 0, r: window.innerWidth, b: window.innerHeight }; |
| 521 | for (let a = el.parentElement; a; a = a.parentElement) { |
| 522 | const cs = getComputedStyle(a); |
| 523 | if (cs.contentVisibility === 'hidden') return { l: 0, t: 0, r: -1, b: -1 }; |
| 524 | if (!/auto|scroll|hidden|clip/.test(cs.overflowX + ' ' + cs.overflowY)) continue; |
| 525 | const q = a.getBoundingClientRect(); |
| 526 | r = { l: Math.max(r.l, q.left), t: Math.max(r.t, q.top), |
| 527 | r: Math.min(r.r, q.right), b: Math.min(r.b, q.bottom) }; |
| 528 | } |
| 529 | return r; |
| 530 | }; |
| 531 | window.__visionShown = (el) => { |
| 532 | if (!el) return { ok: false, why: 'no element' }; |
| 533 | const cs = getComputedStyle(el); |
| 534 | if (cs.display === 'none') return { ok: false, why: 'display:none' }; |
| 535 | if (cs.visibility === 'hidden') return { ok: false, why: 'visibility:hidden' }; |
| 536 | if (Number(cs.opacity) === 0) return { ok: false, why: 'opacity:0' }; |
| 537 | // Chrome's own answer, where it has one: it knows about content-visibility and |
| 538 | // about ancestors this walk would have to guess at. |
| 539 | if (typeof el.checkVisibility === 'function' |
| 540 | && !el.checkVisibility({ checkOpacity: true, checkVisibilityCSS: true, |
| 541 | contentVisibilityAuto: true })) { |
| 542 | return { ok: false, why: 'checkVisibility() says no' }; |
| 543 | } |
| 544 | const b = el.getBoundingClientRect(); |
| 545 | if (b.width < 1 || b.height < 1) return { ok: false, why: 'no area' }; |
| 546 | const cx = b.left + b.width / 2, cy = b.top + b.height / 2; |
| 547 | const c = clipRect(el); |
| 548 | if (cx < c.l || cx > c.r || cy < c.t || cy > c.b) { |
| 549 | return { ok: false, why: `centre (${Math.round(cx)},${Math.round(cy)}) is outside its ` |
| 550 | + `clip rect (${Math.round(c.l)},${Math.round(c.t)})-(${Math.round(c.r)},${Math.round(c.b)})` |
| 551 | + ' — it has a rect with area but nobody can see it' }; |
| 552 | } |
| 553 | const hit = document.elementFromPoint(cx, cy); |
| 554 | if (!hit || !(hit === el || el.contains(hit) || hit.contains(el))) { |
| 555 | return { ok: false, why: 'the hit test at its centre lands on ' |
| 556 | + (hit ? '<' + hit.tagName.toLowerCase() + ' class="' + hit.className + '">' : 'nothing') }; |
| 557 | } |
| 558 | return { ok: true, why: '' }; |
| 559 | }; |
| 560 | }); |
| 561 | |
| 562 | /// The worker tile for `task`, and everything the checks read off it. |
| 563 | const tileFor = (task) => p.evaluate((task) => { |
| 564 | const shown = window.__visionShown; |
| 565 | const card = [...document.querySelectorAll('#agents-list .acard')] |
| 566 | .find((c) => ((c.querySelector('.atask') || {}).textContent || '').trim() === task); |
| 567 | if (!card) return { found: false, cards: document.querySelectorAll('#agents-list .acard').length }; |
| 568 | // Every leaf in the tile that carries words of its own, so "the tile says X" can be |
| 569 | // traced to the one element that says it rather than to an innerText soup. |
| 570 | const said = [...card.querySelectorAll('*')] |
| 571 | .filter((e) => e.children.length === 0 && (e.textContent || '').trim()) |
| 572 | .map((e) => ({ cls: e.className || '', text: (e.textContent || '').trim(), vis: shown(e) })); |
| 573 | return { |
| 574 | found: true, |
| 575 | cls: card.className, |
| 576 | visible: shown(card).ok ? card.innerText : '', |
| 577 | why: shown(card).why, |
| 578 | said: said.filter((x) => x.vis.ok).map((x) => x.text), |
| 579 | hidden: said.filter((x) => !x.vis.ok).map((x) => x.text + ' [' + x.vis.why + ']'), |
| 580 | // Everything only a mouse would ever find, which is what the app offers today. |
| 581 | titles: [...card.querySelectorAll('[title]')] |
| 582 | .map((e) => (e.className || e.tagName) + ': ' + e.getAttribute('title')), |
| 583 | }; |
| 584 | }, task); |
| 585 | |
| 586 | /// Every ledger entry and Diamond-turn count since `mark`. |
| 587 | const spendSince = (mark) => p.evaluate((mark) => { |
| 588 | let entries = []; |
| 589 | try { entries = JSON.parse(localStorage.getItem('daimond-ledger') || '[]'); } catch (e) { entries = []; } |
| 590 | let sig = { diamonds: {} }; |
| 591 | try { sig = window.DaimondSignals ? window.DaimondSignals.snapshot() : sig; } catch (e) { /* absent */ } |
| 592 | return { |
| 593 | entries: entries.filter((e) => e && e.t >= mark).map((e) => ({ m: e.m, u: e.u, p: e.p, c: e.c })), |
| 594 | diamonds: Object.keys(sig.diamonds || {}).reduce((o, k) => { |
| 595 | o[k] = (sig.diamonds[k] || {}).turns || 0; return o; |
| 596 | }, {}), |
| 597 | }; |
| 598 | }, mark); |
| 599 | |
| 600 | try { |
| 601 | // ── The instrument's own self-test ─────────────────────────── |
| 602 | // |
| 603 | // Before anything is claimed about the app: the visibility predicate check 4 rests |
| 604 | // on must reject a thing that has a rect with area and is nonetheless invisible. |
| 605 | // This is defect K made deliberately and measured, so that a green check 4 later |
| 606 | // cannot be the vacuity that check was written to escape. |
| 607 | await installShown(); |
| 608 | const control = await p.evaluate(() => { |
| 609 | const wrap = document.createElement('div'); |
| 610 | wrap.style.cssText = 'position:fixed;left:10px;top:10px;width:80px;height:20px;' |
| 611 | + 'overflow:hidden;z-index:99998'; |
| 612 | const away = document.createElement('span'); |
| 613 | away.textContent = 'scrolled out of a clipped scroller'; |
| 614 | away.style.cssText = 'display:inline-block;margin-left:600px;white-space:nowrap'; |
| 615 | const here = document.createElement('span'); |
| 616 | here.textContent = 'plainly there'; |
| 617 | here.style.cssText = 'display:inline-block;position:fixed;left:10px;top:60px;' |
| 618 | + 'background:#000;color:#fff;z-index:99999'; |
| 619 | wrap.appendChild(away); |
| 620 | document.body.appendChild(wrap); |
| 621 | document.body.appendChild(here); |
| 622 | const shown = window.__visionShown; |
| 623 | const b = away.getBoundingClientRect(); |
| 624 | const out = { |
| 625 | awayRects: away.getClientRects().length, |
| 626 | awayArea: b.width * b.height, |
| 627 | away: shown(away), |
| 628 | here: shown(here), |
| 629 | }; |
| 630 | wrap.remove(); here.remove(); |
| 631 | return out; |
| 632 | }); |
| 633 | check(control.awayRects > 0 && control.awayArea > 0 |
| 634 | && !control.away.ok && control.here.ok, |
| 635 | 'INSTRUMENT: the visibility predicate rejects a chip scrolled out of a clipped ' |
| 636 | + 'scroller (a rect with area that nobody can see) and accepts one that is drawn', |
| 637 | `${control.awayRects} rect(s), area ${Math.round(control.awayArea)}; ` |
| 638 | + `hidden → ${control.away.ok ? 'ACCEPTED, which is the vacuity of defect K' : control.away.why}; ` |
| 639 | + `visible → ${control.here.ok ? 'accepted' : 'REJECTED: ' + control.here.why}`); |
| 640 | |
| 641 | // ── The fixture ────────────────────────────────────────────── |
| 642 | const prov = await p.evaluate(() => { |
| 643 | const d = window.DaimondModels.getDefault(); |
| 644 | return d.provider; |
| 645 | }); |
| 646 | const dId = await newDiamond('Vision Routing'); |
| 647 | if (!dId) throw new Error('no Diamond came into focus after the New Diamond dialog'); |
| 648 | const rec = await setModels(dId, { |
| 649 | provider: prov, model: 'mock/fast', |
| 650 | workerProvider: prov, workerModel: TEXT_MODEL, |
| 651 | visionProvider: prov, visionModel: VISION_MODEL, |
| 652 | }); |
| 653 | check(rec.workerModel === TEXT_MODEL && rec.visionModel === VISION_MODEL, |
| 654 | 'the Diamond has a text worker model that cannot see and an image model that can', |
| 655 | JSON.stringify(rec)); |
| 656 | |
| 657 | // The picture, under a name that spells NO image extension. |
| 658 | // |
| 659 | // It is the defect's own case — the one `taskWantsVision`'s doc comment admits it |
| 660 | // cannot know, "a worker that discovers an image for itself" — and the only fixture |
| 661 | // that could tell the defect from the fix if the spelling rule ran at all. It does |
| 662 | // not (see the header), so today BOTH kinds of task go to the text model; check 7 |
| 663 | // is the one that measures the other kind. This premise is asserted rather than |
| 664 | // assumed all the same, because the rule can be made reachable, and on the day it is |
| 665 | // a fixture that spelled `shot.png` would start passing check 1 without anything |
| 666 | // having learned anything about capability. |
| 667 | const PIC = `diamonds/${dId}/capture`; |
| 668 | const TASK = `@look ${PIC}`; |
| 669 | await put(PIC, PNG); |
| 670 | check(!IMAGE_EXT.test(TASK), |
| 671 | 'and the task names no image, so nothing but a CAPABILITY signal could move this ' |
| 672 | + 'worker off the text model — the case the defect is about', |
| 673 | `${String(IMAGE_EXT)} against ${JSON.stringify(TASK)}`); |
| 674 | |
| 675 | // ── 1, 3, 4, 5. One worker, shown a picture it cannot see ──── |
| 676 | clearMockLog(); |
| 677 | const mark = Date.now(); |
| 678 | // The per-Diamond turn counts BEFORE anything is dispatched. Without a baseline the |
| 679 | // "both against this Diamond" half of check 3 is satisfied by the daimon's own two |
| 680 | // turns, which are on this Diamond whatever the workers do — the clause would be |
| 681 | // there and would be measuring nothing, which is the shape of defect K. |
| 682 | const base = await spendSince(mark); |
| 683 | await steerDiamond(s, `@tools spawn_agent {"name":"looker","task":"${TASK}"}`); |
| 684 | const ran = await until(async () => { |
| 685 | const r = await runFor(TASK); |
| 686 | return !!r && ['done', 'error', 'stopped'].includes(r.status); |
| 687 | }); |
| 688 | await p.waitForTimeout(1500); |
| 689 | await shot(s, 'vision-1-looker'); |
| 690 | |
| 691 | const run = await runFor(TASK); |
| 692 | const legs = legsOf(TASK); |
| 693 | check(ran && legs.length > 0, |
| 694 | 'the worker ran and reached the model (nothing below is true of a worker that never ran)', |
| 695 | legLine(TASK)); |
| 696 | |
| 697 | // ── 1. The work moved to the image model ───────────────────── |
| 698 | // |
| 699 | // "leg 1 carries neither" as DESIGN_VISION.md §7 puts it is not quite right and is |
| 700 | // not asserted: leg 1 MUST carry the picture exactly once, because being refused is |
| 701 | // how the app learns the model is blind. What must be true is that the picture was |
| 702 | // refused there and accepted on the image model. |
| 703 | const leg1 = legs[0] || { model: '', images: 0, refused: 0 }; |
| 704 | const leg2 = legs[1] || null; |
| 705 | check(!!leg2 && leg2.model === VISION_MODEL && leg2.images > 0 |
| 706 | && leg1.model === TEXT_MODEL && leg1.refused > 0, |
| 707 | 'THE WORK MOVED TO THE IMAGE MODEL — a second leg on the Diamond\'s image model ' |
| 708 | + 'carrying the picture, after the text model refused it', |
| 709 | legLine(TASK)); |
| 710 | |
| 711 | // ── 3. Each leg billed to the model that spent it ──────────── |
| 712 | const money = await spendSince(mark); |
| 713 | const onText = money.entries.filter((e) => e.m === TEXT_MODEL); |
| 714 | const onVision = money.entries.filter((e) => e.m === VISION_MODEL); |
| 715 | // Every turn charged since the dispatch, and to whom. `bump` drops an empty id, so a |
| 716 | // leg billed to nobody — which is what a `finally` reading the SELECTION rather than |
| 717 | // the run used to do — shows up as a Diamond that grew by less than the ledger did. |
| 718 | const grew = Object.keys(money.diamonds) |
| 719 | .filter((k) => (money.diamonds[k] || 0) > (base.diamonds[k] || 0)); |
| 720 | const mine = (money.diamonds[dId] || 0) - (base.diamonds[dId] || 0); |
| 721 | check(onText.length === 1 && onVision.length === 1 |
| 722 | && mine === money.entries.length && grew.length === 1 && grew[0] === dId, |
| 723 | 'AND EACH LEG IS BILLED TO THE MODEL THAT SPENT IT, both against this Diamond', |
| 724 | `ledger since dispatch: ${JSON.stringify(money.entries)}; ` |
| 725 | + `${mine} of those ${money.entries.length} turn(s) went to this Diamond` |
| 726 | + (grew.filter((k) => k !== dId).length |
| 727 | ? `; also charged: ${grew.filter((k) => k !== dId).join(', ')}` : '')); |
| 728 | |
| 729 | // ── 4. Disclosed on screen, without hovering ───────────────── |
| 730 | const tile = await tileFor(TASK); |
| 731 | const names = (txt, m) => txt.includes(m) || txt.includes(m.split('/').pop()); |
| 732 | const seen = tile.found ? tile.said.join(' · ') : ''; |
| 733 | check(tile.found && !!tile.visible && names(seen, VISION_MODEL) && names(seen, TEXT_MODEL), |
| 734 | 'AND IT IS DISCLOSED ON SCREEN WITHOUT HOVERING — the tile\'s own visible text ' |
| 735 | + 'names both the model it left and the model it moved to', |
| 736 | !tile.found |
| 737 | ? `no tile whose task is ${JSON.stringify(TASK)} (${tile.cards} tile(s) in the pane)` |
| 738 | : (tile.visible ? '' : `the whole tile is not visible: ${tile.why}; `) |
| 739 | + `visible text: ${JSON.stringify(seen)}` |
| 740 | + (tile.hidden.length ? `; drawn but not visible: ${JSON.stringify(tile.hidden)}` : '') |
| 741 | + `; only on hover: ${JSON.stringify(tile.titles)}`); |
| 742 | |
| 743 | // ── 5. The second leg is a RESUME, not a restart ───────────── |
| 744 | // |
| 745 | // WHILE THE FIX IS ABSENT THIS CHECK CANNOT SPEAK FOR ITSELF. There is no second leg |
| 746 | // at all, so it fails with "there was no second leg to look at" — which is honest, |
| 747 | // and is a consequence of check 1 rather than an independent measurement. It only |
| 748 | // starts distinguishing restart from resume once something moves the worker. Said |
| 749 | // here rather than left for a reader to work out from a green run that never was. |
| 750 | // |
| 751 | // Restart and resume differ in one observable thing: what the last user message of |
| 752 | // the new session is. `resume()` seeds `[{user: task}, {assistant: its own text}]` |
| 753 | // and runs the turn on a NUDGE; a restart runs it on the task again. So the last |
| 754 | // user message being the task IS the restart. The assistant seed is asserted beside |
| 755 | // it, because a nudge with nothing carried forward is a restart with extra words. |
| 756 | const first2 = leg2 && leg2.reqs[0]; |
| 757 | const msgs2 = (first2 && first2.messages) || []; |
| 758 | const lastUser = [...msgs2].reverse().find((m) => m.role === 'user'); |
| 759 | const lastUserText = lastUser |
| 760 | ? (typeof lastUser.content === 'string' ? lastUser.content |
| 761 | : (lastUser.content || []).map((x) => x.text || '').join(' ')) |
| 762 | : ''; |
| 763 | // An assistant message with words in it, not a particular form of words: a restart |
| 764 | // carries NONE, so "there is one" is the whole discriminator, and matching the |
| 765 | // mock's phrasing would only make this brittle against the fixture. |
| 766 | const seed2 = msgs2.filter((m) => m.role === 'assistant' |
| 767 | && typeof m.content === 'string' && m.content.trim()); |
| 768 | check(!!first2 && lastUserText.trim() !== TASK && seed2.length > 0, |
| 769 | 'AND THE SECOND LEG IS A RESUME — seeded with the worker\'s own earlier words and ' |
| 770 | + 'carried on by a nudge, not started again from the task', |
| 771 | !first2 ? 'there was no second leg to look at' |
| 772 | : `last user message ${JSON.stringify(lastUserText.slice(0, 90))}; ` |
| 773 | + `earlier words carried: ${JSON.stringify(seed2.map((m) => m.content.slice(0, 60)))}`); |
| 774 | |
| 775 | // ── 2. It moved ONCE ───────────────────────────────────────── |
| 776 | // |
| 777 | // Its own Diamond, because the case is a Diamond whose IMAGE model is itself blind. |
| 778 | // Without the once-per-run guard this is a worker moved from a blind model to a |
| 779 | // blind model for as long as the pool will let it. |
| 780 | const dId2 = await newDiamond('Vision Ping Pong'); |
| 781 | if (!dId2) throw new Error('no second Diamond came into focus'); |
| 782 | await setModels(dId2, { |
| 783 | provider: prov, model: 'mock/fast', |
| 784 | workerProvider: prov, workerModel: TEXT_MODEL, |
| 785 | visionProvider: prov, visionModel: BLIND_VISION, |
| 786 | }); |
| 787 | const PIC2 = `diamonds/${dId2}/capture`; |
| 788 | const TASK2 = `@look ${PIC2}`; |
| 789 | await put(PIC2, PNG); |
| 790 | clearMockLog(); |
| 791 | await steerDiamond(s, `@tools spawn_agent {"name":"pingpong","task":"${TASK2}"}`); |
| 792 | await until(async () => { |
| 793 | const r = await runFor(TASK2); |
| 794 | return !!r && ['done', 'error', 'stopped'].includes(r.status); |
| 795 | }); |
| 796 | await p.waitForTimeout(1500); |
| 797 | await shot(s, 'vision-2-pingpong'); |
| 798 | const legs2 = legsOf(TASK2); |
| 799 | const tile2 = await tileFor(TASK2); |
| 800 | const run2 = await runFor(TASK2); |
| 801 | // THE LEG COUNT CANNOT SEE THE FAILURE THIS CHECK EXISTS FOR, and that was measured |
| 802 | // rather than argued. A repeat move goes to the SAME image model, and `legsOf` groups |
| 803 | // consecutive requests by model, so every repeat merges into leg 2 and the count stays |
| 804 | // at two. Under `--break loop` the app made 2,101 requests and was refused 700 times, |
| 805 | // ~83k tokens and about five cents in sixty seconds, and this check was green for all |
| 806 | // of it. |
| 807 | // |
| 808 | // So the once-ness is measured where it actually shows: the blind image model is handed |
| 809 | // the picture ONCE. Unbroken that is 1 refusal on leg 2; the runaway made 700. The |
| 810 | // terminal status is asserted beside it because a ping-pong does not merely cost money, |
| 811 | // it never ends -- and a check that waits for a terminal state it never reaches would |
| 812 | // otherwise report on a half-finished run. |
| 813 | const moved2 = legs2.length === 2 |
| 814 | && legs2[0].model === TEXT_MODEL && legs2[1].model === BLIND_VISION; |
| 815 | const once2 = legs2.length > 1 && legs2[1].refused <= 2 && legs2[1].reqs.length <= 12; |
| 816 | const ended2 = !!run2 && ['done', 'error', 'stopped'].includes(run2.status); |
| 817 | check(moved2 && once2 && ended2, |
| 818 | 'AND IT MOVED ONCE — an image model that is itself blind is tried once and not ' |
| 819 | + 'again, and the run ENDS', |
| 820 | `${legLine(TASK2)}; status ${run2 ? run2.status : '(no run)'}; ` |
| 821 | + `the tile says ${JSON.stringify(tile2.found ? tile2.said.join(' · ') : '(no tile)')}`); |
| 822 | |
| 823 | // ── 6. A task with no picture is untouched ─────────────────── |
| 824 | // |
| 825 | // The check that stops "always send workers to the image model" from passing every |
| 826 | // one of the five above. Same Diamond as check 1, so the image model is configured |
| 827 | // and available — it simply must not be used. |
| 828 | await selectDiamond(dId); |
| 829 | const back = await p.evaluate(() => { |
| 830 | const f = window.DaimondAttach && window.DaimondAttach.focus(); |
| 831 | return (f && f.kind === 'diamond') ? String(f.id) : ''; |
| 832 | }); |
| 833 | check(back === dId, |
| 834 | 'the first Diamond — the one with a SIGHTED image model — is in focus again, so ' |
| 835 | + 'the check below is about routing and not about a Diamond with nowhere to go', |
| 836 | `focus is ${back || '(nothing)'}, wanted ${dId}`); |
| 837 | const TASK3 = '@text there is nothing here to look at'; |
| 838 | clearMockLog(); |
| 839 | await steerDiamond(s, `@tools spawn_agent {"name":"reader","task":"${TASK3}"}`); |
| 840 | await until(async () => { |
| 841 | const r = await runFor(TASK3); |
| 842 | return !!r && ['done', 'error', 'stopped'].includes(r.status); |
| 843 | }); |
| 844 | await p.waitForTimeout(1200); |
| 845 | await shot(s, 'vision-3-noplicture'); |
| 846 | const legs3 = legsOf(TASK3); |
| 847 | const tile3 = await tileFor(TASK3); |
| 848 | const said3 = tile3.found ? tile3.said.join(' · ') : ''; |
| 849 | check(legs3.length === 1 && legs3[0].model === TEXT_MODEL && legs3[0].images === 0 |
| 850 | && !names(said3, VISION_MODEL), |
| 851 | 'BUT A TASK CARRYING NO PICTURE IS UNTOUCHED — one leg, on the text model, and ' |
| 852 | + 'the tile says nothing about an image model', |
| 853 | `${legLine(TASK3)}; the tile says ${JSON.stringify(said3)}`); |
| 854 | |
| 855 | // ── 7. The spelling rule, through the door dispatch uses ───── |
| 856 | // |
| 857 | // A task that NAMES a picture should reach the image model with no capability signal |
| 858 | // at all — DESIGN_VISION.md §5 keeps `IMAGE_EXT` on exactly that promise, as "a |
| 859 | // cheap hint that saves the first leg whenever the daimon happens to spell the |
| 860 | // filename". It saves nothing: `dispatch` supplies a model on every call, so |
| 861 | // `routeFor` never consults the task. `verify_diamondmodels` proves the rule through |
| 862 | // `routeForDiamond`, which passes `supplied: false` and is called by nothing under |
| 863 | // `www/` at all. |
| 864 | const TASK4 = '@text compare shots/rail.png against the mockup'; |
| 865 | clearMockLog(); |
| 866 | await steerDiamond(s, `@tools spawn_agent {"name":"speller","task":"${TASK4}"}`); |
| 867 | await until(async () => { |
| 868 | const r = await runFor(TASK4); |
| 869 | return !!r && ['done', 'error', 'stopped'].includes(r.status); |
| 870 | }); |
| 871 | await p.waitForTimeout(1200); |
| 872 | const legs4 = legsOf(TASK4); |
| 873 | const run4 = await runFor(TASK4); |
| 874 | // `sees` is asserted beside the model, not instead of it: landing on the image model |
| 875 | // is what a RE-ROUTE also does, and this check is about the first leg being routed |
| 876 | // there. Without it the check would pass on the very failure it was written for. |
| 877 | const routed4 = !!legs4[0] && legs4[0].model === VISION_MODEL && !!(run4 && run4.sees); |
| 878 | check(routed4, |
| 879 | 'AND A TASK THAT NAMES A PICTURE REACHES THE IMAGE MODEL through the door ' |
| 880 | + 'dispatch actually uses — the check the other six assume', |
| 881 | `${legLine(TASK4)}; the app recorded sees=${run4 ? run4.sees : '(no run)'}` |
| 882 | // The diagnosis belongs to the FAILURE, not to the line. Printed always, it |
| 883 | // made a green check assert that the defect was still live. |
| 884 | + (routed4 ? '' : ' — if sees is false, `supplied` was true on this dispatch and' |
| 885 | + ' `taskWantsVision` was never asked')); |
| 886 | |
| 887 | // Context for whoever reads the failures: what the app itself thought it was doing. |
| 888 | console.log('\n what the app recorded for the looker: ' |
| 889 | + JSON.stringify(run ? { model: run.model, provider: run.provider, sees: run.sees, |
| 890 | status: run.status, text: String(run.text || '').slice(0, 160) } : null)); |
| 891 | } finally { |
| 892 | await s.close(); |
| 893 | } |
| 894 | |
| 895 | console.log(bad ? `\n${bad} check(s) FAILED` : '\nall checks passed'); |
| 896 | process.exit(bad ? 1 : 0); |