Oregami
Repositories/oxedyne/daimond

oxedyne/daimond/dev/verify_toolroad.mjs

24.7 KiB, 1 run

created by r2519314175:751, which is this file's identity for as long as the history lasts, whatever it is later renamed to

download · who wrote it · its history

1// verify_toolroad.mjs — a TOOL call that dies on the road never reaches the model.
2//
3// THE DEFECT, as the owner met it on a real iPhone on 2026-08-28:
4//
5// "on ios the retry mechanism is 'succeeding' but the overall ux is failing, the response
6// started 'I can't get through to the web right now to look this up ....'"
7//
8// The provider retry ladder (src/llm.rs, eight attempts over up to 120 s) does work across an iOS
9// freeze, and `dev/verify_predrop.mjs` proves it. It was beside the point. THE LADDER GUARDS THE
10// CALL TO THE MODEL; IT DOES NOT GUARD THE CALLS THE MODEL MAKES. A home-screen PWA is put in the
11// back/forward cache on every app switch, so every in-flight request dies — and a `web_fetch`
12// that dies was handed back to the model AS A TOOL RESULT SAYING IT FAILED. The model then did
13// the reasonable thing with a failed tool: it apologised and answered around it.
14//
15// WHAT MADE IT SERIOUS. That apology was not a transient. `ToolRegistry::dispatch_unbilled`
16// caught every tool `Err` and returned `MessageContent::text(error_line(…))`, so the failure
17// became an ordinary result, the model answered, and THE TURN COMPLETED NORMALLY. It travelled
18// the SUCCESS path, not an error path. So `captureSession` (www/js/daimond.js) stored the whole
19// exchange — the failed tool result and the apology — into `chat.session.msgs`, which is the
20// model's own conversation and is replayed on every later turn and folded into the summary. A
21// moment of platform behaviour became a durable false fact in the record.
22//
23// THE CATEGORY ERROR AT THE ROOT, which is what the code now says out loud: a local failure is
24// not information the model should reason about. "The host refused you", "that page is 404",
25// "the API returned an error" are RESULTS and the model should adapt to them. "Your user's phone
26// went to sleep and the fetch never left the device" is an infrastructure event, and handing it
27// over as though it were a fact about the world is what produces the apology.
28//
29// ── WHAT IS CHECKED ──────────────────────────────────────────────────────────
30//
31// 1. THE INSTRUMENT FIRED. The tool's request really did meet a rejected fetch.
32// 2. THE MODEL WAS NEVER TOLD. Not one request in the whole sitting carries a tool message
33// naming the failure. Read off dev/mockllm-N.log, which is what the model was ACTUALLY
34// shown, rather than off the screen.
35// 3. AND THE TURN CAME BACK BADGED, with a Continue — the machinery lane/n-drop built for a
36// dead provider call, reached now by a dead tool call.
37// 4. AND THE LADDER WAS CLIMBED, and said so while it climbed.
38// 5. AND THE STORED SESSION CARRIES NO APOLOGY AND NO DANGLING TOOL CALL.
39// 6. THE CONTROL: a tool that fails for a REMOTE reason still reaches the model, unchanged.
40// Without this, "nothing reaches the model" would pass by breaking every tool.
41// 7. THE REGRESSION GUARD — see below.
42// 8. The two spellings of the mark, Rust and JavaScript, are the same string.
43// 9. THE LADDER PARKS WHILE THE PAGE IS FROZEN, and resumes when it comes back.
44//
45// ── CHECK 7 IS A REGRESSION GUARD AND CANNOT FAIL TODAY ──────────────────────
46//
47// It types a SECOND prompt into a sitting whose tool call died, and asserts the provider is not
48// handed an assistant turn bearing a tool_use block with no matching result — which every
49// provider rejects outright, taking the whole conversation with it.
50//
51// It is written because that property is INVISIBLE on the restore path. Press Continue, or
52// reload, and `restore_session` runs `pair_up` (src/protocol.rs), which strips the dangling call
53// for you. Only the LIVE session, carried on by typing again in the same sitting, can hold one —
54// and only `Agent::abandon_round` prevents that. A future refactor that dropped the in-memory
55// repair would pass every other check in this tree and break exactly this. It is load-bearing
56// rather than redundant, which is the reason to keep it green rather than to delete it as a
57// check that never fails.
58//
59// ── HOW THE FAILURE IS PRODUCED ──────────────────────────────────────────────
60//
61// `window.fetch` is replaced in the page with one that rejects with `new TypeError('Load
62// failed')` — WebKit's own wording, verbatim — for `/api/web/fetch` and only while armed. The
63// same instrument `dev/verify_predrop.mjs` uses for the provider call, pointed one layer down.
64// This is a simulation of Safari and is honest about it: what it proves is that the APP treats a
65// dead tool fetch correctly. Whether iOS produces that sentence in the field is a question for a
66// device, and the owner has one.
67//
68// PROVED AGAINST BROKEN CODE FIRST:
69//
70// node dev/verify_toolroad.mjs --break swallow # the catch as it stood: SEVEN checks fail
71// node dev/verify_toolroad.mjs --wording alien # a browser nobody has met: all still pass
72// node dev/verify_toolroad.mjs # and then, clean
73//
74// `--break swallow` restores `dispatch_unbilled`'s behaviour of 2026-08-27 exactly — every tool
75// error becoming a sentence — by disarming the mark at the point it is set, which is the only
76// half of this that lives in a file a verifier can patch. Measured 2026-08-28 in world 21: seven
77// checks fail, and the two that matter most read
78//
79// FAIL THE MODEL WAS NEVER TOLD THE FETCH DIED — 1 message(s), first: "Error: Load failed"
80// FAIL THE STORED SESSION CARRIES NO APOLOGY AND NO DEAD TOOL RESULT — 1 message(s) of 4
81//
82// which is the defect itself, reproduced: the failure reaching the model, and reaching the
83// durable record it is replayed from.
84//
85// `--wording alien` is NOT a break; it is the argument for having done this with a mark rather
86// than with a regex. It rejects with a sentence no classifier in this tree has ever seen, and
87// everything must still pass — because what is being tested for is Daimond's own mark and not
88// the browser's prose. Run it after changing anything about how the road is recognised.
89//
90// A THIRD PROPERTY IS NOT PROVED HERE AND IS PROVED IN RUST INSTEAD, because it cannot be
91// reached from a browser: that a REFUSAL is never read as a road failure however it is worded.
92// See `test_a_refusal_is_not_a_road_failure_however_it_is_worded` in src/tools.rs, which fails
93// on a pattern classifier with "an answer from the far end was classified as the road: Daimond
94// Hands refused that." That is the check that makes the whole design safe, and a regex over tool
95// results would not survive it.
96//
97// eval "$(bash dev/world.sh 21 --up)"
98// node dev/verify_toolroad.mjs
99//
100// Needs dev/serve.mjs and the mock. No gateway: `/api/web/fetch` is stubbed here.
101import fs from 'node:fs';
102import path from 'node:path';
103import { fileURLToPath } from 'node:url';
104import { open, newChat, scratch, shot, storedChats } from './harness.mjs';
105
106const HERE = path.dirname(fileURLToPath(import.meta.url));
107const LOG = process.env.DAIMOND_MOCK_LOG || path.join(HERE, 'mockllm.log');
108
109const BREAK = (() => {
110 const i = process.argv.indexOf('--break');
111 return i > 0 ? String(process.argv[i + 1] || '') : '';
112})();
113
114// The lines that decide it, each quoted whole so a move breaks this file loudly rather than
115// letting it patch nothing and report a pass.
116const ANCHORS = {
117 // gateway.js: the mark itself. Disarming it is exactly the old behaviour — an unmarked
118 // rejection is not classifiable as the road, so it becomes a tool result as it always did.
119 swallow: ['www/js/gateway.js',
120 '\t\t\treturn real.apply(window, arguments).catch(function (e) { throw roadMark(e); });\n',
121 '\t\t\treturn real.apply(window, arguments);\n'],
122};
123if (BREAK && !ANCHORS[BREAK]) {
124 console.error(`unknown break '${BREAK}'; one of: ${Object.keys(ANCHORS).join(', ')}`);
125 process.exit(2);
126}
127
128let bad = 0;
129const check = (pass, name, detail) => {
130 if (!pass) bad++;
131 console.log((pass ? ' ok ' : ' FAIL ') + name + (detail ? ' — ' + detail : ''));
132};
133
134const SRC = {};
135for (const [k, [rel, from]] of Object.entries(ANCHORS)) {
136 const file = path.join(HERE, '..', rel);
137 SRC[rel] = SRC[rel] || fs.readFileSync(file, 'utf8');
138 if (SRC[rel].split(from).length !== 2) {
139 console.error(`the line the '${k}' break patches is not in ${rel} exactly once; `
140 + 'the anchor has moved and the break would patch nothing');
141 process.exit(2);
142 }
143}
144
145// ── 8. One mark, two spellings ───────────────────────────────────────
146//
147// The engine tests for this string EXACTLY (`ROAD_MARK`, src/tools.rs) and the page writes it
148// (`roadMark`, www/js/gateway.js). Neither can see the other, so the coupling is asserted here:
149// a change to one of them fails this file rather than quietly stopping the classification.
150const RUST_MARK = (() => {
151 const s = fs.readFileSync(path.join(HERE, '..', 'src/tools.rs'), 'utf8');
152 const m = s.match(/pub const ROAD_MARK: &str = "([^"]+)";/);
153 return m ? m[1] : '';
154})();
155const JS_MARK = (() => {
156 const s = SRC['www/js/gateway.js'];
157 const m = s.match(/var ROAD_MARK = '([^']+)';/);
158 return m ? m[1] : '';
159})();
160check(!!RUST_MARK && RUST_MARK === JS_MARK,
161 'THE MARK IS ONE STRING in both halves — src/tools.rs and www/js/gateway.js',
162 `rust ${JSON.stringify(RUST_MARK)} / js ${JSON.stringify(JS_MARK)}`);
163
164// WebKit's own sentence for a fetch that never got a response, verbatim.
165// https://trackjs.com/javascript-errors/load-failed/
166const WEBKIT_WORDING = 'Load failed';
167
168/// A sentence for a dead fetch that NOTHING in this tree recognises.
169///
170/// Neither `CLIENT_ROAD` nor `BROWSER_ROAD` (www/js/daimond.js) matches a word of it, and it is
171/// deliberately plausible: every engine words this differently and the next one to appear will
172/// word it differently again. Under `--wording alien` every check must still pass, which is only
173/// possible because the classification is done on Daimond's own mark.
174const ALIEN_WORDING = 'The operation was interrupted before completion';
175
176const WORDING = (() => {
177 const i = process.argv.indexOf('--wording');
178 return (i > 0 && String(process.argv[i + 1] || '') === 'alien') ? ALIEN_WORDING : WEBKIT_WORDING;
179})();
180
181const s = await open({
182 name: 'toolroad',
183 profile: scratch('pw', 'toolroad' + (BREAK ? '-' + BREAK : '')),
184 route: async (page) => {
185 if (BREAK) {
186 const [rel, from, to] = ANCHORS[BREAK];
187 const body = SRC[rel].replace(from, to);
188 await page.route('**/' + rel.split('/').slice(1).join('/'), (r) => r.fulfill({
189 status: 200, contentType: 'application/javascript', body,
190 }));
191 }
192 // The gateway route the Web panel's `fetch` posts to. There is no gateway in this world,
193 // so it is answered here — and this is the ANSWER case, which check 6 needs.
194 await page.route('**/api/web/fetch', (r) => r.fulfill({
195 status: 200, contentType: 'application/json',
196 body: JSON.stringify({ ok: true, url: 'https://example.test/', title: 'Example',
197 text: 'the page said this', bytes: 18 }),
198 }));
199 // The Safari failure, armed from the test rather than from the mock: what is being
200 // simulated is the BROWSER's behaviour, not the far end's, so it belongs in the browser.
201 // Installed before any script runs — which also means `guardFetch` in js/gateway.js
202 // wraps THIS, exactly as it wraps the real one.
203 await page.addInitScript((wording) => {
204 const real = window.fetch.bind(window);
205 window.__loadFailCount = 0;
206 window.fetch = function (input, init) {
207 const url = typeof input === 'string' ? input
208 : (input && input.url) || String(input || '');
209 if (window.__loadFail && /\/api\/web\/fetch/.test(url)) {
210 window.__loadFailCount++;
211 // A TypeError with no response and no status, which is the whole of what a
212 // page gets when a fetch dies before its headers.
213 return Promise.reject(new TypeError(wording));
214 }
215 return real(input, init);
216 };
217 }, WORDING);
218 },
219});
220const { page: p } = s;
221if (BREAK) console.log(`\n*** RUNNING UNDER --break ${BREAK}: failures below are the point ***\n`);
222if (WORDING !== WEBKIT_WORDING) {
223 console.log(`\n*** a browser that says ${JSON.stringify(WORDING)}: everything must still pass ***\n`);
224}
225
226/// Every line of the mock log, as records.
227///
228/// LINES AND NOT BYTES. The obvious mark is `statSync(LOG).size`, and it is wrong: the log holds
229/// the system prompt, which is full of multi-byte characters, so a BYTE offset used to slice a
230/// decoded STRING lands past where it should and takes the next record's opening with it. The
231/// first version of this file did exactly that, and the regression guard read "0 requests" for a
232/// turn that had plainly made one. The log is one JSON object per line, so a line count is exact.
233const records = () => {
234 let raw = '';
235 try { raw = fs.readFileSync(LOG, 'utf8'); } catch (e) { return []; }
236 return raw.split('\n').filter(Boolean)
237 .map((l) => { try { return JSON.parse(l); } catch (e) { return null; } });
238};
239
240/// How many requests the model has been sent so far, as a mark to read forward from.
241const logMark = () => records().length;
242
243/// Every request the model was shown since `from`, parsed.
244const shown = (from) => records().slice(from).filter(Boolean);
245
246/// Every message of every request in that span, flattened.
247const shownMsgs = (from) => {
248 const out = [];
249 for (const r of shown(from)) {
250 const ms = (r && r.body && r.body.messages) || (r && r.messages) || [];
251 for (const m of ms) out.push(m);
252 }
253 return out;
254};
255
256/// What the thread holds, and what is offered about the last answer in it.
257const thread = () => p.evaluate(() => {
258 const inter = [...document.querySelectorAll('#chat-output .chat-msg.interrupted')].pop() || null;
259 return {
260 interrupted: !!inter,
261 continues: !!(inter && inter.querySelector('.turn-interrupted button')),
262 badge: inter ? ((inter.querySelector('.ti-label') || {}).textContent || '') : '',
263 errors: [...document.querySelectorAll('#chat-output .chat-msg-error, #chat-output .error-log')].length,
264 text: (document.getElementById('chat-output') || {}).textContent || '',
265 tries: window.__loadFailCount || 0,
266 };
267});
268
269/// The MODEL's own conversation as the app has STORED it — not the screen transcript.
270///
271/// `chat.session.msgs` is what `captureSession` writes at the end of every turn, and it is what
272/// `restore_session` feeds back to the engine on the next one. It is the durable record, and it
273/// is the thing the apology was getting into.
274const storedSession = async () => {
275 const chats = await storedChats(s);
276 if (!chats || !chats.length) return null;
277 let best = null;
278 for (const c of chats) {
279 if (c && c.session && Array.isArray(c.session.msgs) && c.session.msgs.length) {
280 if (!best || (c.updated || 0) >= (best.updated || 0)) best = c;
281 }
282 }
283 return best ? { msgs: best.session.msgs } : null;
284};
285
286/// Wait for the composer to offer Send again — the app's own "the turn is over".
287const settle = async (label, timeout = 200000) => {
288 const t0 = Date.now();
289 while (Date.now() - t0 < timeout) {
290 const busy = await p.evaluate(() => {
291 const b = document.getElementById('chat-send');
292 return !!b && (b.classList.contains('stop') || b.disabled);
293 });
294 if (!busy) return { ended: true, ms: Date.now() - t0 };
295 await p.waitForTimeout(250);
296 }
297 console.log(` note the ${label} turn was still running after ${timeout} ms`);
298 return { ended: false, ms: Date.now() - t0 };
299};
300
301/// Does any message here carry the failure, in any of the shapes it could take?
302///
303/// Deliberately broad. The point is not that one particular sentence is absent but that NOTHING
304/// about a dead fetch reached the model, so this looks for the browser's wording, the app's mark
305/// and the tool layer's error opening alike.
306const carriesFailure = (m) => {
307 const c = m && m.content;
308 const s = typeof c === 'string' ? c : JSON.stringify(c || '');
309 return s.indexOf(WORDING) !== -1
310 || /Load failed|daimond-road|Error: .*fetch|could not reach that page/i.test(s);
311};
312
313try {
314 await newChat(s);
315
316 // ── 1-5. The road goes under a tool call ─────────────────────
317 const mark0 = logMark();
318 await p.evaluate(() => { window.__loadFail = true; });
319 await p.fill('#chat-input', '@tool web_fetch {"url":"https://example.test/"}');
320 await p.click('#chat-send');
321 // Sampled while the ladder is still climbing, which is the only window there is.
322 await p.waitForTimeout(2000);
323 const caption = await p.evaluate(() => {
324 const el = document.querySelector('.chat-spinner-say');
325 return el ? (el.textContent || '') : '';
326 });
327
328 const end1 = await settle('tool-road');
329 const t1 = await thread();
330 await shot(s, 'toolroad-interrupted');
331
332 check(t1.tries > 0,
333 'THE INSTRUMENT FIRED — the tool really did meet a rejected fetch',
334 `${t1.tries} attempt(s) refused with ${JSON.stringify(WORDING)}`);
335
336 // 2. THE CHECK THIS FILE EXISTS FOR.
337 const msgs1 = shownMsgs(mark0);
338 const told = msgs1.filter(carriesFailure);
339 check(told.length === 0,
340 'THE MODEL WAS NEVER TOLD THE FETCH DIED — no request carries the failure',
341 told.length
342 ? `${told.length} message(s), first: ${JSON.stringify(String(told[0].content).slice(0, 120))}`
343 : `${msgs1.length} message(s) shown to the model, none of them the failure`);
344
345 // And the apology it would have written is not on screen either, which is the symptom the
346 // owner actually reported. Matched on the shape rather than on one model's words.
347 check(!/can.?t get through|could not reach the web|unable to (access|reach)/i.test(t1.text),
348 'AND NO APOLOGY WAS WRITTEN for a failure that was never the world\'s',
349 JSON.stringify(t1.text.replace(/\s+/g, ' ').slice(0, 120)));
350
351 // 4. The ladder was climbed, and said so.
352 check(t1.tries > 1,
353 'THE LADDER WAS CLIMBED — a read is tried again rather than reported',
354 `${t1.tries} attempt(s)`);
355 check(/connection dropped|trying/i.test(caption),
356 'and the app said so while it climbed, in its own voice',
357 JSON.stringify(caption.slice(0, 90)));
358
359 // 3. And the turn came back the way a dropped provider call does.
360 check(end1.ended, 'the turn ends rather than hanging on the road', `${end1.ms} ms`);
361 check(t1.interrupted && t1.continues,
362 'THE TURN IS HANDED BACK, badged, with a Continue',
363 t1.interrupted ? '' : `${t1.errors} error line(s), nothing marked interrupted`);
364 check(/connection dropped/i.test(t1.badge),
365 'and the badge says the CONNECTION DROPPED, not that a tool failed',
366 JSON.stringify(t1.badge.slice(0, 90)));
367
368 // 5. And nothing false is in the model's own stored conversation.
369 const sess1 = await storedSession();
370 if (!sess1) {
371 console.log(' note no probe for the stored session; check 5 is skipped');
372 } else {
373 const dirty = sess1.msgs.filter(carriesFailure);
374 check(dirty.length === 0,
375 'THE STORED SESSION CARRIES NO APOLOGY AND NO DEAD TOOL RESULT',
376 dirty.length ? `${dirty.length} message(s) of ${sess1.msgs.length}` : `${sess1.msgs.length} kept`);
377 // pair_up's rule, asserted on the stored list: no assistant turn may carry a call that
378 // nothing answers.
379 const dangling = sess1.msgs.filter((m, i) => {
380 if (m.role !== 'assistant' || !(m.tool_calls || []).length) return false;
381 const answered = new Set();
382 for (let j = i + 1; j < sess1.msgs.length && sess1.msgs[j].role === 'tool'; j++) {
383 answered.add(sess1.msgs[j].tool_call_id);
384 }
385 return m.tool_calls.some((tc) => !answered.has(tc.id));
386 });
387 check(dangling.length === 0,
388 'AND NO ASSISTANT TURN CARRIES A CALL NOTHING ANSWERED',
389 `${dangling.length} dangling`);
390 }
391
392 // ── 9. THE LADDER PARKS WHILE THE PAGE IS FROZEN ─────────────
393 //
394 // This is the one genuinely new mechanism and the only part of it that can be measured from
395 // here. A request already in flight CANNOT be parked -- the promise belongs to the browser,
396 // the page's JavaScript is not running while the page is frozen, and by the time anything of
397 // ours runs again the request has already been rejected. What can be parked is the moment
398 // BEFORE a request, and that is what `Agent::over_the_road` does: it sleeps its backoff and
399 // then waits for `document.visibilityState` to say `visible` again.
400 //
401 // MEASURED BY THE SHAPE OF THE COUNT, which is falsifiable without a break: with the park,
402 // attempts stall at one while the page is hidden and resume when it comes back. Without it,
403 // all eight are spent into a dead page within a few seconds and the count would already be at
404 // its ceiling by the first sample below. The two are not close.
405 //
406 // `visibilityState` is overridden rather than the page really being backgrounded, because a
407 // headless browser has no app to switch to. What is simulated is the one signal the engine
408 // reads, and it reads it the same way whoever set it.
409 await newChat(s);
410 await p.evaluate(() => {
411 Object.defineProperty(document, 'visibilityState',
412 { get: () => 'hidden', configurable: true });
413 window.__loadFail = true;
414 window.__loadFailCount = 0;
415 });
416 await p.fill('#chat-input', '@tool web_fetch {"url":"https://example.test/frozen"}');
417 await p.click('#chat-send');
418 await p.waitForTimeout(8000);
419 const whileHidden = await p.evaluate(() => window.__loadFailCount || 0);
420 check(whileHidden > 0 && whileHidden <= 2,
421 'THE LADDER PARKS WHILE THE PAGE IS FROZEN rather than spending itself into a dead page',
422 `${whileHidden} attempt(s) in 8 s hidden — the whole budget is ${8} attempts`);
423 // And it comes back when the page does.
424 await p.evaluate(() => {
425 Object.defineProperty(document, 'visibilityState',
426 { get: () => 'visible', configurable: true });
427 });
428 await p.waitForTimeout(4000);
429 const afterShow = await p.evaluate(() => window.__loadFailCount || 0);
430 check(afterShow > whileHidden,
431 'AND IT RESUMES WHEN THE PAGE COMES BACK, rather than having given up in the dark',
432 `${whileHidden} attempt(s) hidden, ${afterShow} after the restore`);
433 await settle('frozen');
434 await shot(s, 'toolroad-frozen');
435
436 // ── 7. The regression guard ──────────────────────────────────
437 //
438 // Carry on by TYPING, not by pressing Continue and not by reloading — the one path on which
439 // `pair_up` does not run for you. What must not happen is the provider being handed an
440 // assistant turn whose tool_use block nothing answers.
441 await p.evaluate(() => { window.__loadFail = false; });
442 const mark1 = logMark();
443 await p.fill('#chat-input', '@text carrying on');
444 await p.click('#chat-send');
445 const end2 = await settle('second-prompt');
446 const msgs2 = shown(mark1);
447 let dangled = 0, requests = 0;
448 for (const r of msgs2) {
449 const ms = (r && r.body && r.body.messages) || (r && r.messages) || [];
450 if (!ms.length) continue;
451 requests++;
452 for (let i = 0; i < ms.length; i++) {
453 const m = ms[i];
454 if (m.role !== 'assistant' || !(m.tool_calls || []).length) continue;
455 const answered = new Set();
456 for (let j = i + 1; j < ms.length && ms[j].role === 'tool'; j++) {
457 answered.add(ms[j].tool_call_id);
458 }
459 if (m.tool_calls.some((tc) => !answered.has(tc.id))) dangled++;
460 }
461 }
462 check(end2.ended && requests > 0 && dangled === 0,
463 'REGRESSION GUARD: typing again in the same sitting sends a LEGAL conversation',
464 `${requests} request(s), ${dangled} carrying an unanswered tool call`);
465
466 // ── 6. THE CONTROL ───────────────────────────────────────────
467 //
468 // A tool that fails for a REMOTE reason must still reach the model, unchanged. Without this,
469 // every check above would pass on a build that simply stopped reporting tool failures at all
470 // — which is a far worse app than the one being fixed.
471 await p.route('**/api/web/fetch', (r) => r.fulfill({
472 status: 502, contentType: 'application/json',
473 body: JSON.stringify({ ok: false, error: 'that page refused the gateway (502)' }),
474 }));
475 await newChat(s);
476 const mark2 = logMark();
477 await p.fill('#chat-input', '@tool web_fetch {"url":"https://example.test/gone"}');
478 await p.click('#chat-send');
479 const end3 = await settle('remote-refusal');
480 const t3 = await thread();
481 const msgs3 = shownMsgs(mark2);
482 const heard = msgs3.filter((m) => {
483 const c = m && m.content;
484 const s2 = typeof c === 'string' ? c : JSON.stringify(c || '');
485 return m.role === 'tool' && /refused the gateway|502/i.test(s2);
486 });
487 check(end3.ended, 'a remote refusal ends too', `${end3.ms} ms`);
488 check(heard.length > 0,
489 'THE CONTROL: A REMOTE REFUSAL STILL REACHES THE MODEL — only the road is withheld',
490 heard.length ? '' : `${msgs3.length} message(s) shown, none carrying the far end's answer`);
491 check(!t3.interrupted,
492 'and it is NOT offered back as an interrupted turn: the far end answered',
493 t3.interrupted ? 'a 502 was badged as a dropped connection' : '');
494 await shot(s, 'toolroad-remote');
495} finally {
496 await s.close();
497}
498
499console.log(bad ? `\n${bad} check(s) FAILED` : '\nall checks passed');
500process.exit(bad ? 1 : 0);