Oregami
Repositories/oxedyne/daimond

oxedyne/daimond/dev/verify_vision.mjs

46.2 KiB, 1 run

created by r2519314175:791, which is this file's identity for as long as the history lasts, whatever it is later renamed to

download · who wrote it · its history

1// verify_vision.mjs — a worker that is shown a picture ends up on a model that can
2// see it, is billed for both halves, and says so on screen.
3//
4// WRITTEN BEFORE THE FIX, AND EXPECTED TO BE RED. Defect AA in
5// dev/DEFECTS_20260821.md and the design in dev/DESIGN_VISION.md; nothing in §7 of
6// that design is implemented at the time of writing. A verifier written after a fix
7// asserts what the fix happens to do. This one asserts what was wanted, so the checks
8// below are the specification and their failures are the defect, quoted.
9//
10// THE DEFECT. `Workers.routeFor` decides which model a worker runs on with
11// `sees = !supplied && taskWantsVision(task)`, and `taskWantsVision` is nothing but
12// `IMAGE_EXT.test(task)` — four extensions, matched against the TASK TEXT. So a worker
13// that will certainly be shown a picture runs on the text model unless the daimon
14// happened to spell a filename into the task; it then reads the image, the provider
15// refuses it, and `LlmClient` quietly takes the pictures out and carries on answering
16// about something it cannot see. That is what cost A$17 on 2026-08-20, and the only
17// disclosure of it was a `title` attribute.
18//
19// FOUND WHILE BUILDING THIS, AND WORSE THAN THE DEFECT AS WRITTEN: the spelling rule
20// does not run at all. `routeFor`'s first line is
21//
22// var sees = !supplied && taskWantsVision(task);
23//
24// and `supplied` is `!!(pick && pick.model)` where `pick` is `dispatch`'s fifth
25// argument. `dispatch` has exactly two callers — the daimon's, which passes
26// `diamondWorkerModel(diamondId)`, and the chat's, which passes
27// `chatWorkerModel(chat)` — and neither can return a pair without a model. So
28// `supplied` is TRUE on every dispatch that exists, `sees` is always false, and
29// `taskWantsVision`, `IMAGE_EXT` and `diamondVisionModel` decide nothing: `vm` is
30// resolved, passed in, and never read. A task that DOES spell `shot.png` goes to the
31// text model exactly like one that does not.
32//
33// It was not noticed because `dev/verify_diamondmodels.mjs:500-540` tests the rule
34// through `Workers.routeForDiamond`, which hardcodes `supplied` to FALSE — a path
35// production never takes. `routeForDiamond` is called by nothing under `www/` at all;
36// its only callers are those two lines of that verifier, and its own doc comment says
37// so: "What a verifier drives, and what a future settings preview would." The rule is
38// correct, tested twice over, and reachable only from the test harness. BUILT BUT
39// UNREACHABLE, and "true of the code, wrong about the reader", in one object.
40//
41// The measurement, rather than the argument: `--break always` was first written as
42// `var sees = !supplied;` and changed nothing at all — no check moved, the worker still
43// ran on the text model, and `run.sees` was still false. A patch that cannot change the
44// answer proves the term it patched was already dead.
45//
46// CHECK 7 BELOW IS THE ONE THAT WOULD HAVE CAUGHT THIS, and it is why there are seven
47// rather than the six DESIGN_VISION.md asks for. It puts a task naming a picture
48// through the door `dispatch` actually opens, rather than through the rule's own
49// method. The two green tests that cover this today both drive doors production never
50// opens, which is exactly how a rule stays correct and dead for as long as it likes.
51//
52// SEVEN PROPERTIES — six from DESIGN_VISION.md §7 "Proof", and one the building of it
53// turned up:
54//
55// 1. THE WORK MOVED TO THE IMAGE MODEL. The second leg names the Diamond's image
56// model and carries a picture; the first leg never reached that model, and the
57// picture it did carry was refused.
58// 2. AND IT MOVED ONCE. A Diamond whose IMAGE model is itself blind does not
59// ping-pong: two legs, and no third.
60// 3. AND EACH LEG IS BILLED TO THE MODEL THAT SPENT IT, both against one Diamond.
61// 4. AND IT IS DISCLOSED ON SCREEN, WITHOUT HOVERING. The check that makes 1-3 mean
62// something: a re-route nobody can see is the silent fallback again with more
63// machinery behind it.
64// 5. AND THE SECOND LEG IS A RESUME, not a restart from the top.
65// 6. BUT A TASK CARRYING NO PICTURE IS UNTOUCHED. The check that stops "always
66// switch to the image model" from passing 1-5.
67// 7. AND A TASK THAT NAMES A PICTURE REACHES THE IMAGE MODEL THROUGH THE DOOR
68// `dispatch` ACTUALLY USES. Not one of the design's six, and not fixed by the
69// design either: §7 adds a capability re-route and never touches `supplied`, so
70// this one stays red after that fix lands. That is the point of it — a check that
71// survives the approved design is the check that says the design is incomplete.
72//
73// ── THE BREAKS, AND WHICH OF THEM CAN BE PROVED TODAY ─────────────────────────
74//
75// House discipline is that a check is proved by reddening it on purpose first. Four
76// of these six breaks patch lines the fix has not written yet, so they cannot be
77// applied and REFUSE, naming the line they wanted. That is deliberate: the break table
78// is the other half of the specification, and a break that silently patched nothing
79// would be worse than one that says it could not.
80//
81// TWO GATES, not one, and the second was earned. `bill-after`'s anchor — `recordSpend`
82// in the worker's `finally` — is in the file TODAY, so the anchor test passed it and it
83// would have patched in an `if (run._reroute)` that is false for ever: a run identical
84// to a plain one, reported under a break name. So a break may also declare what its
85// PATCH depends on, and is refused when that is missing.
86//
87// node dev/verify_vision.mjs --break spelling 1 (and 2-5) — routing by IMAGE_EXT
88// alone, which is the tree as it
89// stands: patches NOTHING and is
90// simply a plain run under a name.
91// node dev/verify_vision.mjs --break always 6 — every worker is routed to the
92// image model. LIVE TODAY: it patches
93// `routeFor`, which exists.
94// node dev/verify_vision.mjs --break loop 2 — REFUSES until the fix lands.
95// node dev/verify_vision.mjs --break bill-after 3 — REFUSES until the fix lands.
96// node dev/verify_vision.mjs --break silent 4 — REFUSES until the fix lands.
97// node dev/verify_vision.mjs --break restart 5 — REFUSES until the fix lands.
98//
99// Check 7 has no break of its own, because it is red on the tree as shipped: a break
100// for it would be a patch that made it GREEN, which is a fix and not a break.
101//
102// One break does not redden one check. `always` reddens 6 AND 1, because routing is
103// upstream of everything; `spelling` reddens 1 through 5, and 7 with them.
104// DESIGN_VISION.md's table says otherwise and is wrong about it — see the report
105// accompanying this file.
106//
107// ── WHAT THIS FILE CANNOT YET DISTINGUISH ────────────────────────────────────
108//
109// Check 4 looks for a DISCLOSURE on the worker tile and the fix has not written one,
110// so it cannot name a class. It therefore asserts a property no class name can dodge —
111// the tile's own visible text names both the model the worker left and the model it
112// moved to — and it prints the tile's whole visible text and every `title` attribute
113// in it when it fails, so the failure reads as "the tile says nothing; the only
114// mention of the image model is a tooltip" rather than "selector not found". If an
115// implementer discloses the move by naming ONLY the destination model, this check will
116// redden on a fix that is arguably correct: naming the destination alone does not tell
117// a reader that anything moved, which is why it is asked for, but it is a judgement
118// and it is written down here rather than hidden in a selector.
119//
120// Check 5 asserts the SHAPE of the second leg's first request — that its last user
121// message is not the bare task, and that it carries the worker's own earlier words.
122// The design's stronger sentence, "a worker that already wrote a file must not write
123// it twice", is not asserted, and cannot honestly be: whether a model repeats work it
124// can see it already did is a property of the model, and this one is a script. What IS
125// app-level is restart-versus-resume, and that is what is measured.
126//
127// AND UNTIL SOMETHING MOVES A WORKER, CHECK 5 CANNOT SPEAK FOR ITSELF. There is no
128// second leg to look at, so it reports that, which is honest but is a consequence of
129// check 1 rather than a measurement of its own. It becomes a real check the moment the
130// re-route exists, and not before.
131//
132// eval "$(bash dev/world.sh 5 --up)"
133// node dev/verify_vision.mjs
134//
135// Needs dev/serve.mjs and dev/mockllm.mjs — the mock with model-name blindness
136// (`mock/blind` refuses a picture, `mock/eyes` takes it, at ONE endpoint). The
137// preflight refuses if the mock answering this world is an older one, because every
138// check below would then be measuring a fixture that is not there. No gateway, no wasm
139// rebuild: the branch under test is JavaScript.
140
141import fs from 'node:fs';
142import path from 'node:path';
143import { fileURLToPath } from 'node:url';
144import { open, steerDiamond, scratch, shot, mockLog, clearMockLog, MOCK } from './harness.mjs';
145
146const HERE = path.dirname(fileURLToPath(import.meta.url));
147const WWW = path.join(HERE, '..', 'www');
148const SRC = fs.readFileSync(path.join(WWW, 'js/daimond.js'), 'utf8');
149const MOCK_LOG_PATH = process.env.DAIMOND_MOCK_LOG || path.join(HERE, 'mockllm.log');
150
151// The two blind models and the sighted one. `mock/blind-too` is never listed by the
152// mock's catalogue and does not need to be: blindness is a property of the NAME, and
153// `DaimondModels.resolve` will resolve any id under a provider that holds a key.
154const TEXT_MODEL = 'mock/blind';
155const VISION_MODEL = 'mock/eyes';
156const BLIND_VISION = 'mock/blind-too';
157
158let bad = 0;
159const check = (pass, name, detail) => {
160 if (!pass) bad++;
161 console.log((pass ? ' ok ' : ' FAIL ') + name + (detail != null ? ' — ' + detail : ''));
162};
163
164// ── The breaks ───────────────────────────────────────────────────────
165//
166// Each is an anchor in js/daimond.js and what to put in its place, served over the
167// real file through `page.route` so the page under test is the shipped page with one
168// line changed. An anchor that is not in the file EXACTLY ONCE stops the run: a break
169// that patches nothing would produce a green page under a red name, which is the
170// failure this whole file exists to avoid.
171const BREAKS = {
172 // The tree as it stands. There is nothing to patch, because the defect is the
173 // current behaviour: routing by spelling and no re-route at all.
174 spelling: {
175 anchor: 'var sees = !supplied && taskWantsVision(task);',
176 patched: 'var sees = !supplied && taskWantsVision(task);',
177 why: 'the routing rule as shipped — this break patches NOTHING and is the '
178 + 'defect itself, run under a name so the table in DESIGN_VISION.md has a '
179 + 'row that can be executed',
180 },
181 // The plausible over-correction: send every worker to the image model. Checks 1-5
182 // would pass and the bill would double for the nineteen workers in twenty that
183 // never look at anything.
184 always: {
185 anchor: 'var sees = !supplied && taskWantsVision(task);',
186 patched: 'var sees = true;',
187 why: 'every worker is routed to the image model, whatever it was asked to do',
188 // `!supplied` was the first spelling of this break and it changed NOTHING, which
189 // is how the finding below was made: `supplied` is true on every dispatch there
190 // is, so the whole right-hand side of that line is already dead. The break has to
191 // overwrite the lot.
192 },
193 // ── Below here the anchors belong to code that does not exist yet ──
194 //
195 // Each names the line DESIGN_VISION.md §7 item 5 asks for. They will refuse until
196 // that line is written, and the refusal names it, so the implementer is told what
197 // this file will patch rather than being left to guess.
198 loop: {
199 anchor: 'if (run._reroute) return;',
200 patched: 'if (false) return;',
201 why: 'the once-per-run guard on the re-route is dropped, so a blind image '
202 + 'model is moved to again and again',
203 },
204 'bill-after': {
205 // ITS ANCHOR IS ALREADY IN THE FILE — `recordSpend`'s call in the worker
206 // `finally` is live today — so the anchor test alone lets this break through and
207 // it then patches in an `if (run._reroute)` that is false for ever. It would run
208 // clean, look exactly like a plain run, and report check 3 red under a break name
209 // that had done nothing. `needs` is the second gate: the break is refused until
210 // the field its patch depends on exists.
211 needs: 'run._reroute',
212 anchor: 'recordSpend(run.model, _pt, _ct, _ca, _cost, run.provider, run.diamondId || \'\');',
213 patched: 'if (run._reroute) { run.model = run._reroute.to.model; run.provider = run._reroute.to.provider; }\n\t\t\t\t'
214 + 'recordSpend(run.model, _pt, _ct, _ca, _cost, run.provider, run.diamondId || \'\');',
215 why: 'the model is swapped BEFORE the spend is recorded, so the wasted first '
216 + 'leg is billed to the image model — the same class of lie as defect A',
217 },
218 silent: {
219 anchor: 'run.reroutedFrom',
220 patched: 'run.__no_such_field',
221 why: 'the tile stops knowing it was re-routed, so the disclosure falls back '
222 + 'to the tooltip it was before',
223 },
224 restart: {
225 anchor: 'Workers.reroute',
226 patched: 'Workers.__restartFromTheTop',
227 why: 'the second leg is built fresh and run from the task again, repeating '
228 + 'every side effect the first leg had already had',
229 },
230};
231
232const BREAK = (() => {
233 const eq = process.argv.find(a => a.startsWith('--break='));
234 if (eq) return eq.split('=')[1];
235 const i = process.argv.indexOf('--break');
236 return i > 0 ? String(process.argv[i + 1] || '') : '';
237})();
238if (BREAK && !BREAKS[BREAK]) {
239 console.error(`unknown break '${BREAK}'; one of: ${Object.keys(BREAKS).join(', ')}`);
240 process.exit(2);
241}
242if (BREAK) {
243 const b = BREAKS[BREAK];
244 // An anchor that is present is not the same as a break that can bite: a patch may
245 // depend on a field the fix has not written. Both gates, and the same refusal.
246 if (b.needs && !SRC.includes(b.needs)) {
247 console.error(`\n REFUSED --break ${BREAK} would patch nothing.`);
248 console.error(` Its anchor is in the file, but the patch depends on ${JSON.stringify(b.needs)},`);
249 console.error(' which is not. Applying it would produce a run indistinguishable from a');
250 console.error(' plain one, under a break name — the exact thing this file exists to stop.');
251 console.error(` What it is for: ${b.why}.\n`);
252 process.exit(2);
253 }
254 const n = SRC.split(b.anchor).length - 1;
255 if (n !== 1) {
256 console.error(`\n REFUSED --break ${BREAK} cannot be applied.`);
257 console.error(` It patches: ${JSON.stringify(b.anchor)}`);
258 console.error(` which is in www/js/daimond.js ${n} time(s), not once.`);
259 console.error(` What it is for: ${b.why}.`);
260 console.error(' If this is one of the four breaks that belong to the vision fix,');
261 console.error(' the fix has not landed yet and there is nothing to break. Running');
262 console.error(' the file with no --break measures the defect, which is the point');
263 console.error(' of it today.\n');
264 process.exit(2);
265 }
266}
267
268// ── The routing rule, read out of the app rather than copied ─────────
269//
270// Check 1's whole premise is that its task names NO image, so today's rule sends the
271// worker to the text model and only a capability signal could move it. A copy of the
272// regex here would go stale the moment somebody widened the real one — which §5 of the
273// design argues against and somebody will eventually try anyway — and this file would
274// then be proving something about a rule the app no longer uses.
275const IMAGE_EXT = (() => {
276 const m = /var IMAGE_EXT = (\/.*\/[a-z]*);/.exec(SRC);
277 if (!m) {
278 console.error('IMAGE_EXT is not in www/js/daimond.js where this file reads it. '
279 + 'Check 1 cannot state its own premise, so nothing below would mean anything.');
280 process.exit(2);
281 }
282 return new Function('return ' + m[1])();
283})();
284
285// ── Preflight: is the mock this run reads the mock this run drives, and is it
286// the one that knows how to be blind? ───────────────────────────────
287//
288// Both halves, because either alone is the failure this suite has been burned by.
289// A mock writing to a log nobody reads makes every check below report "the model was
290// never called"; a mock from before this fixture existed answers 200 to a picture on
291// `mock/blind`, so the worker never learns anything, no event is ever emitted, and six
292// checks fail with a story about the app that is entirely about the fixture.
293{
294 const before = mockLog().length;
295 let why = '';
296 try {
297 const r = await fetch(MOCK, {
298 method: 'POST',
299 headers: { 'content-type': 'application/json' },
300 body: JSON.stringify({ model: 'mock/fast', stream: false,
301 messages: [{ role: 'user', content: 'vision-preflight-probe' }] }),
302 });
303 if (!r.ok) why = `the mock answered ${r.status}`;
304 else {
305 await r.text();
306 if (mockLog().length <= before) why = 'the probe was not written to the log';
307 }
308 } catch (e) {
309 why = 'the mock could not be reached: ' + e.message;
310 }
311 if (!why) {
312 // The fixture itself: a picture on a blind model must come back 400, in the
313 // provider's own words about `image_url`, because that is the shape
314 // `LlmClient::stream_turn` learns blindness from.
315 try {
316 const r = await fetch(MOCK, {
317 method: 'POST',
318 headers: { 'content-type': 'application/json' },
319 body: JSON.stringify({ model: TEXT_MODEL, stream: false, messages: [{
320 role: 'user', content: [
321 { type: 'text', text: 'look' },
322 { type: 'image_url', image_url: { url: 'data:image/png;base64,AAAA' } },
323 ] }] }),
324 });
325 const body = await r.text();
326 if (r.status !== 400 || !/image_url/.test(body)) {
327 why = `${TEXT_MODEL} answered ${r.status} to a picture instead of refusing it — `
328 + 'this mock predates the vision fixture in dev/mockllm.mjs';
329 }
330 } catch (e) {
331 why = 'the blindness probe could not be sent: ' + e.message;
332 }
333 }
334 if (why) {
335 console.log(' REFUSED ' + why);
336 console.log(` driving: ${MOCK}`);
337 console.log(` reading: ${MOCK_LOG_PATH}`);
338 console.log(' Nothing below could measure the app: it would measure the fixture.');
339 console.log(' `bash dev/world.sh N --down` then --up, from THIS worktree.');
340 process.exit(2);
341 }
342}
343
344// ── Reading the wire ─────────────────────────────────────────────────
345
346/// Whether a logged request is a WORKER's turn on `task` rather than the daimon's own
347/// turn about it.
348///
349/// The daimon's transcript quotes the task inside its `spawn_agent` call, so every
350/// round of the dispatching turn mentions it. A worker's turn is the one where the task
351/// IS a user message — and a RESUMED worker's session is seeded with that same user
352/// message, which is exactly why it is the right test: both legs of one worker match,
353/// and the daimon never does.
354const isWorkerTurn = (e, task) => (e.messages || []).some(
355 (m) => m.role === 'user' && typeof m.content === 'string' && m.content.trim() === task);
356
357/// Every request one worker's task produced, oldest first.
358const turnsFor = (task) => mockLog().filter((e) => isWorkerTurn(e, task));
359
360/// The legs of one worker: consecutive requests grouped by the model they named.
361///
362/// A leg is a session on one model. Two legs is a re-route; one is the defect; three
363/// or more is a ping-pong. Grouping by RUNS of the same model rather than by the set of
364/// models is what tells the last two apart.
365const legsOf = (task) => {
366 const out = [];
367 for (const e of turnsFor(task)) {
368 const last = out[out.length - 1];
369 if (last && last.model === e.model) last.reqs.push(e);
370 else out.push({ model: e.model, reqs: [e] });
371 }
372 return out.map((l) => ({
373 model: l.model,
374 reqs: l.reqs,
375 images: l.reqs.reduce((n, e) => n + (e.images || 0), 0),
376 refused: l.reqs.filter((e) => e.refusedImages).length,
377 }));
378};
379
380const legLine = (task) => legsOf(task)
381 .map((l, i) => `leg ${i + 1}: ${l.model}, ${l.reqs.length} req, ${l.images} image(s)`
382 + (l.refused ? `, ${l.refused} refused` : ''))
383 .join(' | ') || 'no worker turn reached the mock at all';
384
385/// Wait for a condition, polling.
386const until = async (fn, ms = 60000) => {
387 const t0 = Date.now();
388 while (Date.now() - t0 < ms) {
389 if (await fn()) return true;
390 await new Promise((r) => setTimeout(r, 300));
391 }
392 return false;
393};
394
395// ── The session ──────────────────────────────────────────────────────
396
397const s = await open({
398 name: 'vision',
399 profile: scratch('pw', 'vision' + (BREAK ? '-' + BREAK : '')),
400 // The damaged file in place of the real one, registered before `goto`.
401 route: (BREAK && BREAKS[BREAK].patched !== BREAKS[BREAK].anchor) ? (async (page) => {
402 const body = SRC.replace(BREAKS[BREAK].anchor, BREAKS[BREAK].patched);
403 await page.route('**/js/daimond.js', (r) => r.fulfill({
404 status: 200, contentType: 'application/javascript', body,
405 }));
406 }) : null,
407});
408const p = s.page;
409if (BREAK) {
410 console.log(`\n*** RUNNING UNDER --break ${BREAK}: ${BREAKS[BREAK].why}.`);
411 if (BREAKS[BREAK].patched === BREAKS[BREAK].anchor) {
412 console.log('*** This break patches nothing: it IS the shipped behaviour.');
413 }
414 console.log('*** Failures below are the point ***\n');
415}
416
417/// A 2x3 PNG, byte for byte the one dev/verify_fileview.mjs uses, so what the app
418/// sniffs is a real picture and not a signature with filler behind it.
419const PNG = [
420 0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a, 0x00, 0x00, 0x00, 0x0d,
421 0x49, 0x48, 0x44, 0x52, 0x00, 0x00, 0x00, 0x02, 0x00, 0x00, 0x00, 0x03,
422 0x08, 0x02, 0x00, 0x00, 0x00, 0x36, 0x88, 0x49, 0xd6, 0x00, 0x00, 0x00,
423 0x10, 0x49, 0x44, 0x41, 0x54, 0x78, 0xda, 0x63, 0xf8, 0xcf, 0x00, 0x04,
424 0xff, 0x19, 0x50, 0x28, 0x00, 0x3e, 0xd6, 0x05, 0xfb, 0xb6, 0xd6, 0xf9,
425 0xda, 0x00, 0x00, 0x00, 0x00, 0x49, 0x45, 0x4e, 0x44, 0xae, 0x42, 0x60,
426 0x82,
427];
428
429/// Put bytes in the workspace through the engine's own door, as dev/verify_fileview
430/// does: it applies the path jail and the per-account namespace, and a hand-rolled
431/// OPFS walk would write somewhere the worker's fence does not reach.
432const put = (file, bytes) => p.evaluate(async ({ file, bytes }) => {
433 const m = await import('/pkg/oxedyne_daimond.js');
434 const app = new m.DaimondApp('http://127.0.0.1/v1/chat/completions', '', 'none', 4096, '', true);
435 await app.write_bytes(file, new Uint8Array(bytes));
436}, { file, bytes });
437
438/// Shut the Admin drawer if it is open.
439///
440/// It opens over the rail, so the `+` buttons sit UNDER it and a click on one is
441/// swallowed by `.admin-drawer-head` — which Playwright reports as a timeout on a
442/// button it can plainly see. It matters on the SECOND run of a fixed profile, because
443/// the drawer's state is remembered: run one is fresh and green, run two dies on a
444/// button, and the difference is not in the app.
445///
446/// THE HARNESS ALREADY KNEW. `harness.newChat` has closed this drawer, with the reason
447/// written beside it, since long before this file existed; this file reached past the
448/// harness for the rail and re-met the trap on its own. Reach for the harness first —
449/// it is the same failure as writing a helper without searching for the machinery that
450/// already does the job, and that one was walked past twice more the same night.
451const closeAdmin = async () => {
452 const btn = p.locator('#admin-close');
453 if (await btn.isVisible().catch(() => false)) {
454 await btn.click({ force: true });
455 await p.waitForTimeout(300);
456 }
457};
458
459/// Make a Diamond through the real dialog and answer its id.
460const newDiamond = async (name) => {
461 await closeAdmin();
462 await p.click('#new-diamond-btn');
463 await p.waitForSelector('.dlg-input', { timeout: 8000 });
464 await p.fill('.dlg-input', name);
465 await p.click('.dlg-ok');
466 await p.waitForTimeout(1200);
467 return p.evaluate(() => {
468 const f = window.DaimondAttach && window.DaimondAttach.focus();
469 return (f && f.kind === 'diamond') ? String(f.id) : '';
470 });
471};
472
473/// Put a Diamond back in focus, the way a person does: by clicking its rail box.
474const selectDiamond = async (id) => {
475 await closeAdmin();
476 await p.$$eval('.diamond-box', (els, want) => {
477 for (const e of els) if (e.dataset.id === want) { e.click(); return; }
478 }, id);
479 await p.waitForTimeout(900);
480};
481
482/// Give a Diamond its three models.
483///
484/// Written straight into the record `setDiamondModel` owns, rather than driven through
485/// the tile dialog's "Workers, images" pulldown. That pulldown is not what this file is
486/// about, and dev/verify_diamondmodels.mjs already drives it; what matters here is that
487/// the record downstream reads is the one the user's choice produces, which is why the
488/// shape is written out in full and read back.
489const setModels = (id, rec) => p.evaluate(({ id, rec }) => {
490 const all = JSON.parse(localStorage.getItem('daimond-diamond-models') || '{}');
491 all[id] = rec;
492 localStorage.setItem('daimond-diamond-models', JSON.stringify(all));
493 return all[id];
494}, { id, rec });
495
496/// The worker run the app persisted for `task`, as it sees it.
497const runFor = (task) => p.evaluate((task) => {
498 let box = {};
499 try { box = JSON.parse(localStorage.getItem('daimond-workers') || '{}'); } catch (e) { return null; }
500 const runs = (box && box.runs) || [];
501 return runs.find((r) => (r.task || '').trim() === task) || null;
502}, task);
503
504/// Install the visibility predicate on the page, once, as `window.__visionShown`.
505///
506/// VISIBILITY IS ASSERTED PROPERLY, and this is the part of the file that exists
507/// because `verify_view.mjs:271` went vacuous (defect K): a chip scrolled out of a
508/// horizontal scroller still returns a rect with area, and content under
509/// `content-visibility: hidden` keeps its last layout. So an element counts as shown
510/// only when its CENTRE lies inside the intersection of every clipping ancestor's box
511/// and the viewport, AND a hit test at that centre lands on it or inside it. Both, not
512/// either: containment says it is not scrolled away, the hit test says nothing is drawn
513/// over it, and neither says anything about the other.
514///
515/// One installation shared by the checks and by the instrument's own self-test, so what
516/// the self-test proves is the predicate the checks then use and not a copy of it.
517const installShown = () => p.evaluate(() => {
518 // The intersection of the viewport with every ancestor that clips.
519 const clipRect = (el) => {
520 let r = { l: 0, t: 0, r: window.innerWidth, b: window.innerHeight };
521 for (let a = el.parentElement; a; a = a.parentElement) {
522 const cs = getComputedStyle(a);
523 if (cs.contentVisibility === 'hidden') return { l: 0, t: 0, r: -1, b: -1 };
524 if (!/auto|scroll|hidden|clip/.test(cs.overflowX + ' ' + cs.overflowY)) continue;
525 const q = a.getBoundingClientRect();
526 r = { l: Math.max(r.l, q.left), t: Math.max(r.t, q.top),
527 r: Math.min(r.r, q.right), b: Math.min(r.b, q.bottom) };
528 }
529 return r;
530 };
531 window.__visionShown = (el) => {
532 if (!el) return { ok: false, why: 'no element' };
533 const cs = getComputedStyle(el);
534 if (cs.display === 'none') return { ok: false, why: 'display:none' };
535 if (cs.visibility === 'hidden') return { ok: false, why: 'visibility:hidden' };
536 if (Number(cs.opacity) === 0) return { ok: false, why: 'opacity:0' };
537 // Chrome's own answer, where it has one: it knows about content-visibility and
538 // about ancestors this walk would have to guess at.
539 if (typeof el.checkVisibility === 'function'
540 && !el.checkVisibility({ checkOpacity: true, checkVisibilityCSS: true,
541 contentVisibilityAuto: true })) {
542 return { ok: false, why: 'checkVisibility() says no' };
543 }
544 const b = el.getBoundingClientRect();
545 if (b.width < 1 || b.height < 1) return { ok: false, why: 'no area' };
546 const cx = b.left + b.width / 2, cy = b.top + b.height / 2;
547 const c = clipRect(el);
548 if (cx < c.l || cx > c.r || cy < c.t || cy > c.b) {
549 return { ok: false, why: `centre (${Math.round(cx)},${Math.round(cy)}) is outside its `
550 + `clip rect (${Math.round(c.l)},${Math.round(c.t)})-(${Math.round(c.r)},${Math.round(c.b)})`
551 + ' — it has a rect with area but nobody can see it' };
552 }
553 const hit = document.elementFromPoint(cx, cy);
554 if (!hit || !(hit === el || el.contains(hit) || hit.contains(el))) {
555 return { ok: false, why: 'the hit test at its centre lands on '
556 + (hit ? '<' + hit.tagName.toLowerCase() + ' class="' + hit.className + '">' : 'nothing') };
557 }
558 return { ok: true, why: '' };
559 };
560});
561
562/// The worker tile for `task`, and everything the checks read off it.
563const tileFor = (task) => p.evaluate((task) => {
564 const shown = window.__visionShown;
565 const card = [...document.querySelectorAll('#agents-list .acard')]
566 .find((c) => ((c.querySelector('.atask') || {}).textContent || '').trim() === task);
567 if (!card) return { found: false, cards: document.querySelectorAll('#agents-list .acard').length };
568 // Every leaf in the tile that carries words of its own, so "the tile says X" can be
569 // traced to the one element that says it rather than to an innerText soup.
570 const said = [...card.querySelectorAll('*')]
571 .filter((e) => e.children.length === 0 && (e.textContent || '').trim())
572 .map((e) => ({ cls: e.className || '', text: (e.textContent || '').trim(), vis: shown(e) }));
573 return {
574 found: true,
575 cls: card.className,
576 visible: shown(card).ok ? card.innerText : '',
577 why: shown(card).why,
578 said: said.filter((x) => x.vis.ok).map((x) => x.text),
579 hidden: said.filter((x) => !x.vis.ok).map((x) => x.text + ' [' + x.vis.why + ']'),
580 // Everything only a mouse would ever find, which is what the app offers today.
581 titles: [...card.querySelectorAll('[title]')]
582 .map((e) => (e.className || e.tagName) + ': ' + e.getAttribute('title')),
583 };
584}, task);
585
586/// Every ledger entry and Diamond-turn count since `mark`.
587const spendSince = (mark) => p.evaluate((mark) => {
588 let entries = [];
589 try { entries = JSON.parse(localStorage.getItem('daimond-ledger') || '[]'); } catch (e) { entries = []; }
590 let sig = { diamonds: {} };
591 try { sig = window.DaimondSignals ? window.DaimondSignals.snapshot() : sig; } catch (e) { /* absent */ }
592 return {
593 entries: entries.filter((e) => e && e.t >= mark).map((e) => ({ m: e.m, u: e.u, p: e.p, c: e.c })),
594 diamonds: Object.keys(sig.diamonds || {}).reduce((o, k) => {
595 o[k] = (sig.diamonds[k] || {}).turns || 0; return o;
596 }, {}),
597 };
598}, mark);
599
600try {
601 // ── The instrument's own self-test ───────────────────────────
602 //
603 // Before anything is claimed about the app: the visibility predicate check 4 rests
604 // on must reject a thing that has a rect with area and is nonetheless invisible.
605 // This is defect K made deliberately and measured, so that a green check 4 later
606 // cannot be the vacuity that check was written to escape.
607 await installShown();
608 const control = await p.evaluate(() => {
609 const wrap = document.createElement('div');
610 wrap.style.cssText = 'position:fixed;left:10px;top:10px;width:80px;height:20px;'
611 + 'overflow:hidden;z-index:99998';
612 const away = document.createElement('span');
613 away.textContent = 'scrolled out of a clipped scroller';
614 away.style.cssText = 'display:inline-block;margin-left:600px;white-space:nowrap';
615 const here = document.createElement('span');
616 here.textContent = 'plainly there';
617 here.style.cssText = 'display:inline-block;position:fixed;left:10px;top:60px;'
618 + 'background:#000;color:#fff;z-index:99999';
619 wrap.appendChild(away);
620 document.body.appendChild(wrap);
621 document.body.appendChild(here);
622 const shown = window.__visionShown;
623 const b = away.getBoundingClientRect();
624 const out = {
625 awayRects: away.getClientRects().length,
626 awayArea: b.width * b.height,
627 away: shown(away),
628 here: shown(here),
629 };
630 wrap.remove(); here.remove();
631 return out;
632 });
633 check(control.awayRects > 0 && control.awayArea > 0
634 && !control.away.ok && control.here.ok,
635 'INSTRUMENT: the visibility predicate rejects a chip scrolled out of a clipped '
636 + 'scroller (a rect with area that nobody can see) and accepts one that is drawn',
637 `${control.awayRects} rect(s), area ${Math.round(control.awayArea)}; `
638 + `hidden → ${control.away.ok ? 'ACCEPTED, which is the vacuity of defect K' : control.away.why}; `
639 + `visible → ${control.here.ok ? 'accepted' : 'REJECTED: ' + control.here.why}`);
640
641 // ── The fixture ──────────────────────────────────────────────
642 const prov = await p.evaluate(() => {
643 const d = window.DaimondModels.getDefault();
644 return d.provider;
645 });
646 const dId = await newDiamond('Vision Routing');
647 if (!dId) throw new Error('no Diamond came into focus after the New Diamond dialog');
648 const rec = await setModels(dId, {
649 provider: prov, model: 'mock/fast',
650 workerProvider: prov, workerModel: TEXT_MODEL,
651 visionProvider: prov, visionModel: VISION_MODEL,
652 });
653 check(rec.workerModel === TEXT_MODEL && rec.visionModel === VISION_MODEL,
654 'the Diamond has a text worker model that cannot see and an image model that can',
655 JSON.stringify(rec));
656
657 // The picture, under a name that spells NO image extension.
658 //
659 // It is the defect's own case — the one `taskWantsVision`'s doc comment admits it
660 // cannot know, "a worker that discovers an image for itself" — and the only fixture
661 // that could tell the defect from the fix if the spelling rule ran at all. It does
662 // not (see the header), so today BOTH kinds of task go to the text model; check 7
663 // is the one that measures the other kind. This premise is asserted rather than
664 // assumed all the same, because the rule can be made reachable, and on the day it is
665 // a fixture that spelled `shot.png` would start passing check 1 without anything
666 // having learned anything about capability.
667 const PIC = `diamonds/${dId}/capture`;
668 const TASK = `@look ${PIC}`;
669 await put(PIC, PNG);
670 check(!IMAGE_EXT.test(TASK),
671 'and the task names no image, so nothing but a CAPABILITY signal could move this '
672 + 'worker off the text model — the case the defect is about',
673 `${String(IMAGE_EXT)} against ${JSON.stringify(TASK)}`);
674
675 // ── 1, 3, 4, 5. One worker, shown a picture it cannot see ────
676 clearMockLog();
677 const mark = Date.now();
678 // The per-Diamond turn counts BEFORE anything is dispatched. Without a baseline the
679 // "both against this Diamond" half of check 3 is satisfied by the daimon's own two
680 // turns, which are on this Diamond whatever the workers do — the clause would be
681 // there and would be measuring nothing, which is the shape of defect K.
682 const base = await spendSince(mark);
683 await steerDiamond(s, `@tools spawn_agent {"name":"looker","task":"${TASK}"}`);
684 const ran = await until(async () => {
685 const r = await runFor(TASK);
686 return !!r && ['done', 'error', 'stopped'].includes(r.status);
687 });
688 await p.waitForTimeout(1500);
689 await shot(s, 'vision-1-looker');
690
691 const run = await runFor(TASK);
692 const legs = legsOf(TASK);
693 check(ran && legs.length > 0,
694 'the worker ran and reached the model (nothing below is true of a worker that never ran)',
695 legLine(TASK));
696
697 // ── 1. The work moved to the image model ─────────────────────
698 //
699 // "leg 1 carries neither" as DESIGN_VISION.md §7 puts it is not quite right and is
700 // not asserted: leg 1 MUST carry the picture exactly once, because being refused is
701 // how the app learns the model is blind. What must be true is that the picture was
702 // refused there and accepted on the image model.
703 const leg1 = legs[0] || { model: '', images: 0, refused: 0 };
704 const leg2 = legs[1] || null;
705 check(!!leg2 && leg2.model === VISION_MODEL && leg2.images > 0
706 && leg1.model === TEXT_MODEL && leg1.refused > 0,
707 'THE WORK MOVED TO THE IMAGE MODEL — a second leg on the Diamond\'s image model '
708 + 'carrying the picture, after the text model refused it',
709 legLine(TASK));
710
711 // ── 3. Each leg billed to the model that spent it ────────────
712 const money = await spendSince(mark);
713 const onText = money.entries.filter((e) => e.m === TEXT_MODEL);
714 const onVision = money.entries.filter((e) => e.m === VISION_MODEL);
715 // Every turn charged since the dispatch, and to whom. `bump` drops an empty id, so a
716 // leg billed to nobody — which is what a `finally` reading the SELECTION rather than
717 // the run used to do — shows up as a Diamond that grew by less than the ledger did.
718 const grew = Object.keys(money.diamonds)
719 .filter((k) => (money.diamonds[k] || 0) > (base.diamonds[k] || 0));
720 const mine = (money.diamonds[dId] || 0) - (base.diamonds[dId] || 0);
721 check(onText.length === 1 && onVision.length === 1
722 && mine === money.entries.length && grew.length === 1 && grew[0] === dId,
723 'AND EACH LEG IS BILLED TO THE MODEL THAT SPENT IT, both against this Diamond',
724 `ledger since dispatch: ${JSON.stringify(money.entries)}; `
725 + `${mine} of those ${money.entries.length} turn(s) went to this Diamond`
726 + (grew.filter((k) => k !== dId).length
727 ? `; also charged: ${grew.filter((k) => k !== dId).join(', ')}` : ''));
728
729 // ── 4. Disclosed on screen, without hovering ─────────────────
730 const tile = await tileFor(TASK);
731 const names = (txt, m) => txt.includes(m) || txt.includes(m.split('/').pop());
732 const seen = tile.found ? tile.said.join(' · ') : '';
733 check(tile.found && !!tile.visible && names(seen, VISION_MODEL) && names(seen, TEXT_MODEL),
734 'AND IT IS DISCLOSED ON SCREEN WITHOUT HOVERING — the tile\'s own visible text '
735 + 'names both the model it left and the model it moved to',
736 !tile.found
737 ? `no tile whose task is ${JSON.stringify(TASK)} (${tile.cards} tile(s) in the pane)`
738 : (tile.visible ? '' : `the whole tile is not visible: ${tile.why}; `)
739 + `visible text: ${JSON.stringify(seen)}`
740 + (tile.hidden.length ? `; drawn but not visible: ${JSON.stringify(tile.hidden)}` : '')
741 + `; only on hover: ${JSON.stringify(tile.titles)}`);
742
743 // ── 5. The second leg is a RESUME, not a restart ─────────────
744 //
745 // WHILE THE FIX IS ABSENT THIS CHECK CANNOT SPEAK FOR ITSELF. There is no second leg
746 // at all, so it fails with "there was no second leg to look at" — which is honest,
747 // and is a consequence of check 1 rather than an independent measurement. It only
748 // starts distinguishing restart from resume once something moves the worker. Said
749 // here rather than left for a reader to work out from a green run that never was.
750 //
751 // Restart and resume differ in one observable thing: what the last user message of
752 // the new session is. `resume()` seeds `[{user: task}, {assistant: its own text}]`
753 // and runs the turn on a NUDGE; a restart runs it on the task again. So the last
754 // user message being the task IS the restart. The assistant seed is asserted beside
755 // it, because a nudge with nothing carried forward is a restart with extra words.
756 const first2 = leg2 && leg2.reqs[0];
757 const msgs2 = (first2 && first2.messages) || [];
758 const lastUser = [...msgs2].reverse().find((m) => m.role === 'user');
759 const lastUserText = lastUser
760 ? (typeof lastUser.content === 'string' ? lastUser.content
761 : (lastUser.content || []).map((x) => x.text || '').join(' '))
762 : '';
763 // An assistant message with words in it, not a particular form of words: a restart
764 // carries NONE, so "there is one" is the whole discriminator, and matching the
765 // mock's phrasing would only make this brittle against the fixture.
766 const seed2 = msgs2.filter((m) => m.role === 'assistant'
767 && typeof m.content === 'string' && m.content.trim());
768 check(!!first2 && lastUserText.trim() !== TASK && seed2.length > 0,
769 'AND THE SECOND LEG IS A RESUME — seeded with the worker\'s own earlier words and '
770 + 'carried on by a nudge, not started again from the task',
771 !first2 ? 'there was no second leg to look at'
772 : `last user message ${JSON.stringify(lastUserText.slice(0, 90))}; `
773 + `earlier words carried: ${JSON.stringify(seed2.map((m) => m.content.slice(0, 60)))}`);
774
775 // ── 2. It moved ONCE ─────────────────────────────────────────
776 //
777 // Its own Diamond, because the case is a Diamond whose IMAGE model is itself blind.
778 // Without the once-per-run guard this is a worker moved from a blind model to a
779 // blind model for as long as the pool will let it.
780 const dId2 = await newDiamond('Vision Ping Pong');
781 if (!dId2) throw new Error('no second Diamond came into focus');
782 await setModels(dId2, {
783 provider: prov, model: 'mock/fast',
784 workerProvider: prov, workerModel: TEXT_MODEL,
785 visionProvider: prov, visionModel: BLIND_VISION,
786 });
787 const PIC2 = `diamonds/${dId2}/capture`;
788 const TASK2 = `@look ${PIC2}`;
789 await put(PIC2, PNG);
790 clearMockLog();
791 await steerDiamond(s, `@tools spawn_agent {"name":"pingpong","task":"${TASK2}"}`);
792 await until(async () => {
793 const r = await runFor(TASK2);
794 return !!r && ['done', 'error', 'stopped'].includes(r.status);
795 });
796 await p.waitForTimeout(1500);
797 await shot(s, 'vision-2-pingpong');
798 const legs2 = legsOf(TASK2);
799 const tile2 = await tileFor(TASK2);
800 const run2 = await runFor(TASK2);
801 // THE LEG COUNT CANNOT SEE THE FAILURE THIS CHECK EXISTS FOR, and that was measured
802 // rather than argued. A repeat move goes to the SAME image model, and `legsOf` groups
803 // consecutive requests by model, so every repeat merges into leg 2 and the count stays
804 // at two. Under `--break loop` the app made 2,101 requests and was refused 700 times,
805 // ~83k tokens and about five cents in sixty seconds, and this check was green for all
806 // of it.
807 //
808 // So the once-ness is measured where it actually shows: the blind image model is handed
809 // the picture ONCE. Unbroken that is 1 refusal on leg 2; the runaway made 700. The
810 // terminal status is asserted beside it because a ping-pong does not merely cost money,
811 // it never ends -- and a check that waits for a terminal state it never reaches would
812 // otherwise report on a half-finished run.
813 const moved2 = legs2.length === 2
814 && legs2[0].model === TEXT_MODEL && legs2[1].model === BLIND_VISION;
815 const once2 = legs2.length > 1 && legs2[1].refused <= 2 && legs2[1].reqs.length <= 12;
816 const ended2 = !!run2 && ['done', 'error', 'stopped'].includes(run2.status);
817 check(moved2 && once2 && ended2,
818 'AND IT MOVED ONCE — an image model that is itself blind is tried once and not '
819 + 'again, and the run ENDS',
820 `${legLine(TASK2)}; status ${run2 ? run2.status : '(no run)'}; `
821 + `the tile says ${JSON.stringify(tile2.found ? tile2.said.join(' · ') : '(no tile)')}`);
822
823 // ── 6. A task with no picture is untouched ───────────────────
824 //
825 // The check that stops "always send workers to the image model" from passing every
826 // one of the five above. Same Diamond as check 1, so the image model is configured
827 // and available — it simply must not be used.
828 await selectDiamond(dId);
829 const back = await p.evaluate(() => {
830 const f = window.DaimondAttach && window.DaimondAttach.focus();
831 return (f && f.kind === 'diamond') ? String(f.id) : '';
832 });
833 check(back === dId,
834 'the first Diamond — the one with a SIGHTED image model — is in focus again, so '
835 + 'the check below is about routing and not about a Diamond with nowhere to go',
836 `focus is ${back || '(nothing)'}, wanted ${dId}`);
837 const TASK3 = '@text there is nothing here to look at';
838 clearMockLog();
839 await steerDiamond(s, `@tools spawn_agent {"name":"reader","task":"${TASK3}"}`);
840 await until(async () => {
841 const r = await runFor(TASK3);
842 return !!r && ['done', 'error', 'stopped'].includes(r.status);
843 });
844 await p.waitForTimeout(1200);
845 await shot(s, 'vision-3-noplicture');
846 const legs3 = legsOf(TASK3);
847 const tile3 = await tileFor(TASK3);
848 const said3 = tile3.found ? tile3.said.join(' · ') : '';
849 check(legs3.length === 1 && legs3[0].model === TEXT_MODEL && legs3[0].images === 0
850 && !names(said3, VISION_MODEL),
851 'BUT A TASK CARRYING NO PICTURE IS UNTOUCHED — one leg, on the text model, and '
852 + 'the tile says nothing about an image model',
853 `${legLine(TASK3)}; the tile says ${JSON.stringify(said3)}`);
854
855 // ── 7. The spelling rule, through the door dispatch uses ─────
856 //
857 // A task that NAMES a picture should reach the image model with no capability signal
858 // at all — DESIGN_VISION.md §5 keeps `IMAGE_EXT` on exactly that promise, as "a
859 // cheap hint that saves the first leg whenever the daimon happens to spell the
860 // filename". It saves nothing: `dispatch` supplies a model on every call, so
861 // `routeFor` never consults the task. `verify_diamondmodels` proves the rule through
862 // `routeForDiamond`, which passes `supplied: false` and is called by nothing under
863 // `www/` at all.
864 const TASK4 = '@text compare shots/rail.png against the mockup';
865 clearMockLog();
866 await steerDiamond(s, `@tools spawn_agent {"name":"speller","task":"${TASK4}"}`);
867 await until(async () => {
868 const r = await runFor(TASK4);
869 return !!r && ['done', 'error', 'stopped'].includes(r.status);
870 });
871 await p.waitForTimeout(1200);
872 const legs4 = legsOf(TASK4);
873 const run4 = await runFor(TASK4);
874 // `sees` is asserted beside the model, not instead of it: landing on the image model
875 // is what a RE-ROUTE also does, and this check is about the first leg being routed
876 // there. Without it the check would pass on the very failure it was written for.
877 const routed4 = !!legs4[0] && legs4[0].model === VISION_MODEL && !!(run4 && run4.sees);
878 check(routed4,
879 'AND A TASK THAT NAMES A PICTURE REACHES THE IMAGE MODEL through the door '
880 + 'dispatch actually uses — the check the other six assume',
881 `${legLine(TASK4)}; the app recorded sees=${run4 ? run4.sees : '(no run)'}`
882 // The diagnosis belongs to the FAILURE, not to the line. Printed always, it
883 // made a green check assert that the defect was still live.
884 + (routed4 ? '' : ' — if sees is false, `supplied` was true on this dispatch and'
885 + ' `taskWantsVision` was never asked'));
886
887 // Context for whoever reads the failures: what the app itself thought it was doing.
888 console.log('\n what the app recorded for the looker: '
889 + JSON.stringify(run ? { model: run.model, provider: run.provider, sees: run.sees,
890 status: run.status, text: String(run.text || '').slice(0, 160) } : null));
891} finally {
892 await s.close();
893}
894
895console.log(bad ? `\n${bad} check(s) FAILED` : '\nall checks passed');
896process.exit(bad ? 1 : 0);