oxedyne/daimond/dev/verify_toolroad.mjs
24.7 KiB, 1 run
created by r2519314175:751, which is this file's identity for as long as the history lasts, whatever it is later renamed to
download · who wrote it · its history
| 1 | // verify_toolroad.mjs — a TOOL call that dies on the road never reaches the model. |
| 2 | // |
| 3 | // THE DEFECT, as the owner met it on a real iPhone on 2026-08-28: |
| 4 | // |
| 5 | // "on ios the retry mechanism is 'succeeding' but the overall ux is failing, the response |
| 6 | // started 'I can't get through to the web right now to look this up ....'" |
| 7 | // |
| 8 | // The provider retry ladder (src/llm.rs, eight attempts over up to 120 s) does work across an iOS |
| 9 | // freeze, and `dev/verify_predrop.mjs` proves it. It was beside the point. THE LADDER GUARDS THE |
| 10 | // CALL TO THE MODEL; IT DOES NOT GUARD THE CALLS THE MODEL MAKES. A home-screen PWA is put in the |
| 11 | // back/forward cache on every app switch, so every in-flight request dies — and a `web_fetch` |
| 12 | // that dies was handed back to the model AS A TOOL RESULT SAYING IT FAILED. The model then did |
| 13 | // the reasonable thing with a failed tool: it apologised and answered around it. |
| 14 | // |
| 15 | // WHAT MADE IT SERIOUS. That apology was not a transient. `ToolRegistry::dispatch_unbilled` |
| 16 | // caught every tool `Err` and returned `MessageContent::text(error_line(…))`, so the failure |
| 17 | // became an ordinary result, the model answered, and THE TURN COMPLETED NORMALLY. It travelled |
| 18 | // the SUCCESS path, not an error path. So `captureSession` (www/js/daimond.js) stored the whole |
| 19 | // exchange — the failed tool result and the apology — into `chat.session.msgs`, which is the |
| 20 | // model's own conversation and is replayed on every later turn and folded into the summary. A |
| 21 | // moment of platform behaviour became a durable false fact in the record. |
| 22 | // |
| 23 | // THE CATEGORY ERROR AT THE ROOT, which is what the code now says out loud: a local failure is |
| 24 | // not information the model should reason about. "The host refused you", "that page is 404", |
| 25 | // "the API returned an error" are RESULTS and the model should adapt to them. "Your user's phone |
| 26 | // went to sleep and the fetch never left the device" is an infrastructure event, and handing it |
| 27 | // over as though it were a fact about the world is what produces the apology. |
| 28 | // |
| 29 | // ── WHAT IS CHECKED ────────────────────────────────────────────────────────── |
| 30 | // |
| 31 | // 1. THE INSTRUMENT FIRED. The tool's request really did meet a rejected fetch. |
| 32 | // 2. THE MODEL WAS NEVER TOLD. Not one request in the whole sitting carries a tool message |
| 33 | // naming the failure. Read off dev/mockllm-N.log, which is what the model was ACTUALLY |
| 34 | // shown, rather than off the screen. |
| 35 | // 3. AND THE TURN CAME BACK BADGED, with a Continue — the machinery lane/n-drop built for a |
| 36 | // dead provider call, reached now by a dead tool call. |
| 37 | // 4. AND THE LADDER WAS CLIMBED, and said so while it climbed. |
| 38 | // 5. AND THE STORED SESSION CARRIES NO APOLOGY AND NO DANGLING TOOL CALL. |
| 39 | // 6. THE CONTROL: a tool that fails for a REMOTE reason still reaches the model, unchanged. |
| 40 | // Without this, "nothing reaches the model" would pass by breaking every tool. |
| 41 | // 7. THE REGRESSION GUARD — see below. |
| 42 | // 8. The two spellings of the mark, Rust and JavaScript, are the same string. |
| 43 | // 9. THE LADDER PARKS WHILE THE PAGE IS FROZEN, and resumes when it comes back. |
| 44 | // |
| 45 | // ── CHECK 7 IS A REGRESSION GUARD AND CANNOT FAIL TODAY ────────────────────── |
| 46 | // |
| 47 | // It types a SECOND prompt into a sitting whose tool call died, and asserts the provider is not |
| 48 | // handed an assistant turn bearing a tool_use block with no matching result — which every |
| 49 | // provider rejects outright, taking the whole conversation with it. |
| 50 | // |
| 51 | // It is written because that property is INVISIBLE on the restore path. Press Continue, or |
| 52 | // reload, and `restore_session` runs `pair_up` (src/protocol.rs), which strips the dangling call |
| 53 | // for you. Only the LIVE session, carried on by typing again in the same sitting, can hold one — |
| 54 | // and only `Agent::abandon_round` prevents that. A future refactor that dropped the in-memory |
| 55 | // repair would pass every other check in this tree and break exactly this. It is load-bearing |
| 56 | // rather than redundant, which is the reason to keep it green rather than to delete it as a |
| 57 | // check that never fails. |
| 58 | // |
| 59 | // ── HOW THE FAILURE IS PRODUCED ────────────────────────────────────────────── |
| 60 | // |
| 61 | // `window.fetch` is replaced in the page with one that rejects with `new TypeError('Load |
| 62 | // failed')` — WebKit's own wording, verbatim — for `/api/web/fetch` and only while armed. The |
| 63 | // same instrument `dev/verify_predrop.mjs` uses for the provider call, pointed one layer down. |
| 64 | // This is a simulation of Safari and is honest about it: what it proves is that the APP treats a |
| 65 | // dead tool fetch correctly. Whether iOS produces that sentence in the field is a question for a |
| 66 | // device, and the owner has one. |
| 67 | // |
| 68 | // PROVED AGAINST BROKEN CODE FIRST: |
| 69 | // |
| 70 | // node dev/verify_toolroad.mjs --break swallow # the catch as it stood: SEVEN checks fail |
| 71 | // node dev/verify_toolroad.mjs --wording alien # a browser nobody has met: all still pass |
| 72 | // node dev/verify_toolroad.mjs # and then, clean |
| 73 | // |
| 74 | // `--break swallow` restores `dispatch_unbilled`'s behaviour of 2026-08-27 exactly — every tool |
| 75 | // error becoming a sentence — by disarming the mark at the point it is set, which is the only |
| 76 | // half of this that lives in a file a verifier can patch. Measured 2026-08-28 in world 21: seven |
| 77 | // checks fail, and the two that matter most read |
| 78 | // |
| 79 | // FAIL THE MODEL WAS NEVER TOLD THE FETCH DIED — 1 message(s), first: "Error: Load failed" |
| 80 | // FAIL THE STORED SESSION CARRIES NO APOLOGY AND NO DEAD TOOL RESULT — 1 message(s) of 4 |
| 81 | // |
| 82 | // which is the defect itself, reproduced: the failure reaching the model, and reaching the |
| 83 | // durable record it is replayed from. |
| 84 | // |
| 85 | // `--wording alien` is NOT a break; it is the argument for having done this with a mark rather |
| 86 | // than with a regex. It rejects with a sentence no classifier in this tree has ever seen, and |
| 87 | // everything must still pass — because what is being tested for is Daimond's own mark and not |
| 88 | // the browser's prose. Run it after changing anything about how the road is recognised. |
| 89 | // |
| 90 | // A THIRD PROPERTY IS NOT PROVED HERE AND IS PROVED IN RUST INSTEAD, because it cannot be |
| 91 | // reached from a browser: that a REFUSAL is never read as a road failure however it is worded. |
| 92 | // See `test_a_refusal_is_not_a_road_failure_however_it_is_worded` in src/tools.rs, which fails |
| 93 | // on a pattern classifier with "an answer from the far end was classified as the road: Daimond |
| 94 | // Hands refused that." That is the check that makes the whole design safe, and a regex over tool |
| 95 | // results would not survive it. |
| 96 | // |
| 97 | // eval "$(bash dev/world.sh 21 --up)" |
| 98 | // node dev/verify_toolroad.mjs |
| 99 | // |
| 100 | // Needs dev/serve.mjs and the mock. No gateway: `/api/web/fetch` is stubbed here. |
| 101 | import fs from 'node:fs'; |
| 102 | import path from 'node:path'; |
| 103 | import { fileURLToPath } from 'node:url'; |
| 104 | import { open, newChat, scratch, shot, storedChats } from './harness.mjs'; |
| 105 | |
| 106 | const HERE = path.dirname(fileURLToPath(import.meta.url)); |
| 107 | const LOG = process.env.DAIMOND_MOCK_LOG || path.join(HERE, 'mockllm.log'); |
| 108 | |
| 109 | const BREAK = (() => { |
| 110 | const i = process.argv.indexOf('--break'); |
| 111 | return i > 0 ? String(process.argv[i + 1] || '') : ''; |
| 112 | })(); |
| 113 | |
| 114 | // The lines that decide it, each quoted whole so a move breaks this file loudly rather than |
| 115 | // letting it patch nothing and report a pass. |
| 116 | const ANCHORS = { |
| 117 | // gateway.js: the mark itself. Disarming it is exactly the old behaviour — an unmarked |
| 118 | // rejection is not classifiable as the road, so it becomes a tool result as it always did. |
| 119 | swallow: ['www/js/gateway.js', |
| 120 | '\t\t\treturn real.apply(window, arguments).catch(function (e) { throw roadMark(e); });\n', |
| 121 | '\t\t\treturn real.apply(window, arguments);\n'], |
| 122 | }; |
| 123 | if (BREAK && !ANCHORS[BREAK]) { |
| 124 | console.error(`unknown break '${BREAK}'; one of: ${Object.keys(ANCHORS).join(', ')}`); |
| 125 | process.exit(2); |
| 126 | } |
| 127 | |
| 128 | let bad = 0; |
| 129 | const check = (pass, name, detail) => { |
| 130 | if (!pass) bad++; |
| 131 | console.log((pass ? ' ok ' : ' FAIL ') + name + (detail ? ' — ' + detail : '')); |
| 132 | }; |
| 133 | |
| 134 | const SRC = {}; |
| 135 | for (const [k, [rel, from]] of Object.entries(ANCHORS)) { |
| 136 | const file = path.join(HERE, '..', rel); |
| 137 | SRC[rel] = SRC[rel] || fs.readFileSync(file, 'utf8'); |
| 138 | if (SRC[rel].split(from).length !== 2) { |
| 139 | console.error(`the line the '${k}' break patches is not in ${rel} exactly once; ` |
| 140 | + 'the anchor has moved and the break would patch nothing'); |
| 141 | process.exit(2); |
| 142 | } |
| 143 | } |
| 144 | |
| 145 | // ── 8. One mark, two spellings ─────────────────────────────────────── |
| 146 | // |
| 147 | // The engine tests for this string EXACTLY (`ROAD_MARK`, src/tools.rs) and the page writes it |
| 148 | // (`roadMark`, www/js/gateway.js). Neither can see the other, so the coupling is asserted here: |
| 149 | // a change to one of them fails this file rather than quietly stopping the classification. |
| 150 | const RUST_MARK = (() => { |
| 151 | const s = fs.readFileSync(path.join(HERE, '..', 'src/tools.rs'), 'utf8'); |
| 152 | const m = s.match(/pub const ROAD_MARK: &str = "([^"]+)";/); |
| 153 | return m ? m[1] : ''; |
| 154 | })(); |
| 155 | const JS_MARK = (() => { |
| 156 | const s = SRC['www/js/gateway.js']; |
| 157 | const m = s.match(/var ROAD_MARK = '([^']+)';/); |
| 158 | return m ? m[1] : ''; |
| 159 | })(); |
| 160 | check(!!RUST_MARK && RUST_MARK === JS_MARK, |
| 161 | 'THE MARK IS ONE STRING in both halves — src/tools.rs and www/js/gateway.js', |
| 162 | `rust ${JSON.stringify(RUST_MARK)} / js ${JSON.stringify(JS_MARK)}`); |
| 163 | |
| 164 | // WebKit's own sentence for a fetch that never got a response, verbatim. |
| 165 | // https://trackjs.com/javascript-errors/load-failed/ |
| 166 | const WEBKIT_WORDING = 'Load failed'; |
| 167 | |
| 168 | /// A sentence for a dead fetch that NOTHING in this tree recognises. |
| 169 | /// |
| 170 | /// Neither `CLIENT_ROAD` nor `BROWSER_ROAD` (www/js/daimond.js) matches a word of it, and it is |
| 171 | /// deliberately plausible: every engine words this differently and the next one to appear will |
| 172 | /// word it differently again. Under `--wording alien` every check must still pass, which is only |
| 173 | /// possible because the classification is done on Daimond's own mark. |
| 174 | const ALIEN_WORDING = 'The operation was interrupted before completion'; |
| 175 | |
| 176 | const WORDING = (() => { |
| 177 | const i = process.argv.indexOf('--wording'); |
| 178 | return (i > 0 && String(process.argv[i + 1] || '') === 'alien') ? ALIEN_WORDING : WEBKIT_WORDING; |
| 179 | })(); |
| 180 | |
| 181 | const s = await open({ |
| 182 | name: 'toolroad', |
| 183 | profile: scratch('pw', 'toolroad' + (BREAK ? '-' + BREAK : '')), |
| 184 | route: async (page) => { |
| 185 | if (BREAK) { |
| 186 | const [rel, from, to] = ANCHORS[BREAK]; |
| 187 | const body = SRC[rel].replace(from, to); |
| 188 | await page.route('**/' + rel.split('/').slice(1).join('/'), (r) => r.fulfill({ |
| 189 | status: 200, contentType: 'application/javascript', body, |
| 190 | })); |
| 191 | } |
| 192 | // The gateway route the Web panel's `fetch` posts to. There is no gateway in this world, |
| 193 | // so it is answered here — and this is the ANSWER case, which check 6 needs. |
| 194 | await page.route('**/api/web/fetch', (r) => r.fulfill({ |
| 195 | status: 200, contentType: 'application/json', |
| 196 | body: JSON.stringify({ ok: true, url: 'https://example.test/', title: 'Example', |
| 197 | text: 'the page said this', bytes: 18 }), |
| 198 | })); |
| 199 | // The Safari failure, armed from the test rather than from the mock: what is being |
| 200 | // simulated is the BROWSER's behaviour, not the far end's, so it belongs in the browser. |
| 201 | // Installed before any script runs — which also means `guardFetch` in js/gateway.js |
| 202 | // wraps THIS, exactly as it wraps the real one. |
| 203 | await page.addInitScript((wording) => { |
| 204 | const real = window.fetch.bind(window); |
| 205 | window.__loadFailCount = 0; |
| 206 | window.fetch = function (input, init) { |
| 207 | const url = typeof input === 'string' ? input |
| 208 | : (input && input.url) || String(input || ''); |
| 209 | if (window.__loadFail && /\/api\/web\/fetch/.test(url)) { |
| 210 | window.__loadFailCount++; |
| 211 | // A TypeError with no response and no status, which is the whole of what a |
| 212 | // page gets when a fetch dies before its headers. |
| 213 | return Promise.reject(new TypeError(wording)); |
| 214 | } |
| 215 | return real(input, init); |
| 216 | }; |
| 217 | }, WORDING); |
| 218 | }, |
| 219 | }); |
| 220 | const { page: p } = s; |
| 221 | if (BREAK) console.log(`\n*** RUNNING UNDER --break ${BREAK}: failures below are the point ***\n`); |
| 222 | if (WORDING !== WEBKIT_WORDING) { |
| 223 | console.log(`\n*** a browser that says ${JSON.stringify(WORDING)}: everything must still pass ***\n`); |
| 224 | } |
| 225 | |
| 226 | /// Every line of the mock log, as records. |
| 227 | /// |
| 228 | /// LINES AND NOT BYTES. The obvious mark is `statSync(LOG).size`, and it is wrong: the log holds |
| 229 | /// the system prompt, which is full of multi-byte characters, so a BYTE offset used to slice a |
| 230 | /// decoded STRING lands past where it should and takes the next record's opening with it. The |
| 231 | /// first version of this file did exactly that, and the regression guard read "0 requests" for a |
| 232 | /// turn that had plainly made one. The log is one JSON object per line, so a line count is exact. |
| 233 | const records = () => { |
| 234 | let raw = ''; |
| 235 | try { raw = fs.readFileSync(LOG, 'utf8'); } catch (e) { return []; } |
| 236 | return raw.split('\n').filter(Boolean) |
| 237 | .map((l) => { try { return JSON.parse(l); } catch (e) { return null; } }); |
| 238 | }; |
| 239 | |
| 240 | /// How many requests the model has been sent so far, as a mark to read forward from. |
| 241 | const logMark = () => records().length; |
| 242 | |
| 243 | /// Every request the model was shown since `from`, parsed. |
| 244 | const shown = (from) => records().slice(from).filter(Boolean); |
| 245 | |
| 246 | /// Every message of every request in that span, flattened. |
| 247 | const shownMsgs = (from) => { |
| 248 | const out = []; |
| 249 | for (const r of shown(from)) { |
| 250 | const ms = (r && r.body && r.body.messages) || (r && r.messages) || []; |
| 251 | for (const m of ms) out.push(m); |
| 252 | } |
| 253 | return out; |
| 254 | }; |
| 255 | |
| 256 | /// What the thread holds, and what is offered about the last answer in it. |
| 257 | const thread = () => p.evaluate(() => { |
| 258 | const inter = [...document.querySelectorAll('#chat-output .chat-msg.interrupted')].pop() || null; |
| 259 | return { |
| 260 | interrupted: !!inter, |
| 261 | continues: !!(inter && inter.querySelector('.turn-interrupted button')), |
| 262 | badge: inter ? ((inter.querySelector('.ti-label') || {}).textContent || '') : '', |
| 263 | errors: [...document.querySelectorAll('#chat-output .chat-msg-error, #chat-output .error-log')].length, |
| 264 | text: (document.getElementById('chat-output') || {}).textContent || '', |
| 265 | tries: window.__loadFailCount || 0, |
| 266 | }; |
| 267 | }); |
| 268 | |
| 269 | /// The MODEL's own conversation as the app has STORED it — not the screen transcript. |
| 270 | /// |
| 271 | /// `chat.session.msgs` is what `captureSession` writes at the end of every turn, and it is what |
| 272 | /// `restore_session` feeds back to the engine on the next one. It is the durable record, and it |
| 273 | /// is the thing the apology was getting into. |
| 274 | const storedSession = async () => { |
| 275 | const chats = await storedChats(s); |
| 276 | if (!chats || !chats.length) return null; |
| 277 | let best = null; |
| 278 | for (const c of chats) { |
| 279 | if (c && c.session && Array.isArray(c.session.msgs) && c.session.msgs.length) { |
| 280 | if (!best || (c.updated || 0) >= (best.updated || 0)) best = c; |
| 281 | } |
| 282 | } |
| 283 | return best ? { msgs: best.session.msgs } : null; |
| 284 | }; |
| 285 | |
| 286 | /// Wait for the composer to offer Send again — the app's own "the turn is over". |
| 287 | const settle = async (label, timeout = 200000) => { |
| 288 | const t0 = Date.now(); |
| 289 | while (Date.now() - t0 < timeout) { |
| 290 | const busy = await p.evaluate(() => { |
| 291 | const b = document.getElementById('chat-send'); |
| 292 | return !!b && (b.classList.contains('stop') || b.disabled); |
| 293 | }); |
| 294 | if (!busy) return { ended: true, ms: Date.now() - t0 }; |
| 295 | await p.waitForTimeout(250); |
| 296 | } |
| 297 | console.log(` note the ${label} turn was still running after ${timeout} ms`); |
| 298 | return { ended: false, ms: Date.now() - t0 }; |
| 299 | }; |
| 300 | |
| 301 | /// Does any message here carry the failure, in any of the shapes it could take? |
| 302 | /// |
| 303 | /// Deliberately broad. The point is not that one particular sentence is absent but that NOTHING |
| 304 | /// about a dead fetch reached the model, so this looks for the browser's wording, the app's mark |
| 305 | /// and the tool layer's error opening alike. |
| 306 | const carriesFailure = (m) => { |
| 307 | const c = m && m.content; |
| 308 | const s = typeof c === 'string' ? c : JSON.stringify(c || ''); |
| 309 | return s.indexOf(WORDING) !== -1 |
| 310 | || /Load failed|daimond-road|Error: .*fetch|could not reach that page/i.test(s); |
| 311 | }; |
| 312 | |
| 313 | try { |
| 314 | await newChat(s); |
| 315 | |
| 316 | // ── 1-5. The road goes under a tool call ───────────────────── |
| 317 | const mark0 = logMark(); |
| 318 | await p.evaluate(() => { window.__loadFail = true; }); |
| 319 | await p.fill('#chat-input', '@tool web_fetch {"url":"https://example.test/"}'); |
| 320 | await p.click('#chat-send'); |
| 321 | // Sampled while the ladder is still climbing, which is the only window there is. |
| 322 | await p.waitForTimeout(2000); |
| 323 | const caption = await p.evaluate(() => { |
| 324 | const el = document.querySelector('.chat-spinner-say'); |
| 325 | return el ? (el.textContent || '') : ''; |
| 326 | }); |
| 327 | |
| 328 | const end1 = await settle('tool-road'); |
| 329 | const t1 = await thread(); |
| 330 | await shot(s, 'toolroad-interrupted'); |
| 331 | |
| 332 | check(t1.tries > 0, |
| 333 | 'THE INSTRUMENT FIRED — the tool really did meet a rejected fetch', |
| 334 | `${t1.tries} attempt(s) refused with ${JSON.stringify(WORDING)}`); |
| 335 | |
| 336 | // 2. THE CHECK THIS FILE EXISTS FOR. |
| 337 | const msgs1 = shownMsgs(mark0); |
| 338 | const told = msgs1.filter(carriesFailure); |
| 339 | check(told.length === 0, |
| 340 | 'THE MODEL WAS NEVER TOLD THE FETCH DIED — no request carries the failure', |
| 341 | told.length |
| 342 | ? `${told.length} message(s), first: ${JSON.stringify(String(told[0].content).slice(0, 120))}` |
| 343 | : `${msgs1.length} message(s) shown to the model, none of them the failure`); |
| 344 | |
| 345 | // And the apology it would have written is not on screen either, which is the symptom the |
| 346 | // owner actually reported. Matched on the shape rather than on one model's words. |
| 347 | check(!/can.?t get through|could not reach the web|unable to (access|reach)/i.test(t1.text), |
| 348 | 'AND NO APOLOGY WAS WRITTEN for a failure that was never the world\'s', |
| 349 | JSON.stringify(t1.text.replace(/\s+/g, ' ').slice(0, 120))); |
| 350 | |
| 351 | // 4. The ladder was climbed, and said so. |
| 352 | check(t1.tries > 1, |
| 353 | 'THE LADDER WAS CLIMBED — a read is tried again rather than reported', |
| 354 | `${t1.tries} attempt(s)`); |
| 355 | check(/connection dropped|trying/i.test(caption), |
| 356 | 'and the app said so while it climbed, in its own voice', |
| 357 | JSON.stringify(caption.slice(0, 90))); |
| 358 | |
| 359 | // 3. And the turn came back the way a dropped provider call does. |
| 360 | check(end1.ended, 'the turn ends rather than hanging on the road', `${end1.ms} ms`); |
| 361 | check(t1.interrupted && t1.continues, |
| 362 | 'THE TURN IS HANDED BACK, badged, with a Continue', |
| 363 | t1.interrupted ? '' : `${t1.errors} error line(s), nothing marked interrupted`); |
| 364 | check(/connection dropped/i.test(t1.badge), |
| 365 | 'and the badge says the CONNECTION DROPPED, not that a tool failed', |
| 366 | JSON.stringify(t1.badge.slice(0, 90))); |
| 367 | |
| 368 | // 5. And nothing false is in the model's own stored conversation. |
| 369 | const sess1 = await storedSession(); |
| 370 | if (!sess1) { |
| 371 | console.log(' note no probe for the stored session; check 5 is skipped'); |
| 372 | } else { |
| 373 | const dirty = sess1.msgs.filter(carriesFailure); |
| 374 | check(dirty.length === 0, |
| 375 | 'THE STORED SESSION CARRIES NO APOLOGY AND NO DEAD TOOL RESULT', |
| 376 | dirty.length ? `${dirty.length} message(s) of ${sess1.msgs.length}` : `${sess1.msgs.length} kept`); |
| 377 | // pair_up's rule, asserted on the stored list: no assistant turn may carry a call that |
| 378 | // nothing answers. |
| 379 | const dangling = sess1.msgs.filter((m, i) => { |
| 380 | if (m.role !== 'assistant' || !(m.tool_calls || []).length) return false; |
| 381 | const answered = new Set(); |
| 382 | for (let j = i + 1; j < sess1.msgs.length && sess1.msgs[j].role === 'tool'; j++) { |
| 383 | answered.add(sess1.msgs[j].tool_call_id); |
| 384 | } |
| 385 | return m.tool_calls.some((tc) => !answered.has(tc.id)); |
| 386 | }); |
| 387 | check(dangling.length === 0, |
| 388 | 'AND NO ASSISTANT TURN CARRIES A CALL NOTHING ANSWERED', |
| 389 | `${dangling.length} dangling`); |
| 390 | } |
| 391 | |
| 392 | // ── 9. THE LADDER PARKS WHILE THE PAGE IS FROZEN ───────────── |
| 393 | // |
| 394 | // This is the one genuinely new mechanism and the only part of it that can be measured from |
| 395 | // here. A request already in flight CANNOT be parked -- the promise belongs to the browser, |
| 396 | // the page's JavaScript is not running while the page is frozen, and by the time anything of |
| 397 | // ours runs again the request has already been rejected. What can be parked is the moment |
| 398 | // BEFORE a request, and that is what `Agent::over_the_road` does: it sleeps its backoff and |
| 399 | // then waits for `document.visibilityState` to say `visible` again. |
| 400 | // |
| 401 | // MEASURED BY THE SHAPE OF THE COUNT, which is falsifiable without a break: with the park, |
| 402 | // attempts stall at one while the page is hidden and resume when it comes back. Without it, |
| 403 | // all eight are spent into a dead page within a few seconds and the count would already be at |
| 404 | // its ceiling by the first sample below. The two are not close. |
| 405 | // |
| 406 | // `visibilityState` is overridden rather than the page really being backgrounded, because a |
| 407 | // headless browser has no app to switch to. What is simulated is the one signal the engine |
| 408 | // reads, and it reads it the same way whoever set it. |
| 409 | await newChat(s); |
| 410 | await p.evaluate(() => { |
| 411 | Object.defineProperty(document, 'visibilityState', |
| 412 | { get: () => 'hidden', configurable: true }); |
| 413 | window.__loadFail = true; |
| 414 | window.__loadFailCount = 0; |
| 415 | }); |
| 416 | await p.fill('#chat-input', '@tool web_fetch {"url":"https://example.test/frozen"}'); |
| 417 | await p.click('#chat-send'); |
| 418 | await p.waitForTimeout(8000); |
| 419 | const whileHidden = await p.evaluate(() => window.__loadFailCount || 0); |
| 420 | check(whileHidden > 0 && whileHidden <= 2, |
| 421 | 'THE LADDER PARKS WHILE THE PAGE IS FROZEN rather than spending itself into a dead page', |
| 422 | `${whileHidden} attempt(s) in 8 s hidden — the whole budget is ${8} attempts`); |
| 423 | // And it comes back when the page does. |
| 424 | await p.evaluate(() => { |
| 425 | Object.defineProperty(document, 'visibilityState', |
| 426 | { get: () => 'visible', configurable: true }); |
| 427 | }); |
| 428 | await p.waitForTimeout(4000); |
| 429 | const afterShow = await p.evaluate(() => window.__loadFailCount || 0); |
| 430 | check(afterShow > whileHidden, |
| 431 | 'AND IT RESUMES WHEN THE PAGE COMES BACK, rather than having given up in the dark', |
| 432 | `${whileHidden} attempt(s) hidden, ${afterShow} after the restore`); |
| 433 | await settle('frozen'); |
| 434 | await shot(s, 'toolroad-frozen'); |
| 435 | |
| 436 | // ── 7. The regression guard ────────────────────────────────── |
| 437 | // |
| 438 | // Carry on by TYPING, not by pressing Continue and not by reloading — the one path on which |
| 439 | // `pair_up` does not run for you. What must not happen is the provider being handed an |
| 440 | // assistant turn whose tool_use block nothing answers. |
| 441 | await p.evaluate(() => { window.__loadFail = false; }); |
| 442 | const mark1 = logMark(); |
| 443 | await p.fill('#chat-input', '@text carrying on'); |
| 444 | await p.click('#chat-send'); |
| 445 | const end2 = await settle('second-prompt'); |
| 446 | const msgs2 = shown(mark1); |
| 447 | let dangled = 0, requests = 0; |
| 448 | for (const r of msgs2) { |
| 449 | const ms = (r && r.body && r.body.messages) || (r && r.messages) || []; |
| 450 | if (!ms.length) continue; |
| 451 | requests++; |
| 452 | for (let i = 0; i < ms.length; i++) { |
| 453 | const m = ms[i]; |
| 454 | if (m.role !== 'assistant' || !(m.tool_calls || []).length) continue; |
| 455 | const answered = new Set(); |
| 456 | for (let j = i + 1; j < ms.length && ms[j].role === 'tool'; j++) { |
| 457 | answered.add(ms[j].tool_call_id); |
| 458 | } |
| 459 | if (m.tool_calls.some((tc) => !answered.has(tc.id))) dangled++; |
| 460 | } |
| 461 | } |
| 462 | check(end2.ended && requests > 0 && dangled === 0, |
| 463 | 'REGRESSION GUARD: typing again in the same sitting sends a LEGAL conversation', |
| 464 | `${requests} request(s), ${dangled} carrying an unanswered tool call`); |
| 465 | |
| 466 | // ── 6. THE CONTROL ─────────────────────────────────────────── |
| 467 | // |
| 468 | // A tool that fails for a REMOTE reason must still reach the model, unchanged. Without this, |
| 469 | // every check above would pass on a build that simply stopped reporting tool failures at all |
| 470 | // — which is a far worse app than the one being fixed. |
| 471 | await p.route('**/api/web/fetch', (r) => r.fulfill({ |
| 472 | status: 502, contentType: 'application/json', |
| 473 | body: JSON.stringify({ ok: false, error: 'that page refused the gateway (502)' }), |
| 474 | })); |
| 475 | await newChat(s); |
| 476 | const mark2 = logMark(); |
| 477 | await p.fill('#chat-input', '@tool web_fetch {"url":"https://example.test/gone"}'); |
| 478 | await p.click('#chat-send'); |
| 479 | const end3 = await settle('remote-refusal'); |
| 480 | const t3 = await thread(); |
| 481 | const msgs3 = shownMsgs(mark2); |
| 482 | const heard = msgs3.filter((m) => { |
| 483 | const c = m && m.content; |
| 484 | const s2 = typeof c === 'string' ? c : JSON.stringify(c || ''); |
| 485 | return m.role === 'tool' && /refused the gateway|502/i.test(s2); |
| 486 | }); |
| 487 | check(end3.ended, 'a remote refusal ends too', `${end3.ms} ms`); |
| 488 | check(heard.length > 0, |
| 489 | 'THE CONTROL: A REMOTE REFUSAL STILL REACHES THE MODEL — only the road is withheld', |
| 490 | heard.length ? '' : `${msgs3.length} message(s) shown, none carrying the far end's answer`); |
| 491 | check(!t3.interrupted, |
| 492 | 'and it is NOT offered back as an interrupted turn: the far end answered', |
| 493 | t3.interrupted ? 'a 502 was badged as a dropped connection' : ''); |
| 494 | await shot(s, 'toolroad-remote'); |
| 495 | } finally { |
| 496 | await s.close(); |
| 497 | } |
| 498 | |
| 499 | console.log(bad ? `\n${bad} check(s) FAILED` : '\nall checks passed'); |
| 500 | process.exit(bad ? 1 : 0); |