Files
OmniRoute/tests/unit/combo-input-bound-failure-8375.test.ts
Hernan Javier Ardila Sanchez 3be0a5d290 fix: combo input-bound, Responses->Chat image strip, qwen-web toolCalling, empty-response exhaustion (#8476)
* test(tail): retire stale i18n __MISSING__ repro + fix qianfan website URL

Base-red slice 6, rebased onto the advanced release/v3.8.49 (91fd5f9). The oauth
grok-cli #7610 guard was already fixed on the base by #8027 (it reads the warning
from grokCliAuthJson.ts) — dropped from this slice to avoid a conflicting duplicate.
Remaining two, still red on the current base:

- i18n #7258: the "focused repro" asserted zh-TW.json STILL carries raw __MISSING__:
  placeholders. That backlog was filled (the "no locale has a raw __MISSING__: leaf"
  invariant is the durable guard); retired the now-inverted repro.
- qianfan: Baidu renamed the product page (product/wenxinworkshop -> product-s/
  qianfan_home); updated the expected website URL.

Validated (clean env): i18n 4/0, qianfan 5/0; oauth-modal-grok 2/0 already green on base.

* fix(resilience): short-circuit combo on input-bound failures (context_length_exceeded) (#8375)

isInputBoundRequestFailure() predicate detects deterministic input-bound
errors (context_length_exceeded/context_window_exceeded). The combo loop
propagates the original 400 immediately instead of burning MAX_GLOBAL_ATTEMPTS
retrying identical oversized inputs against every account.

Test: combo-input-bound-failure-8375.test.ts (1 test, 2 assertions)

* fix(resilience): add early-exit in combo dispatcher for input-bound failures (#8375)

When isInputBoundRequestFailure detects context_length_exceeded,
the combo loop returns {ok:false, response} immediately instead of
re-dispatching the oversized request.

Test: node --import tsx/esm --test tests/unit/combo-input-bound-failure-8375.test.ts
- 1 test, 2 assertions, 0 fail

* fix(translator): strip input_image from tool outputs in Responses->Chat downgrade (#8459)

toolOutputContentToString() extracts input_text/output_text parts and
replaces input_image with a placeholder instead of JSON.stringify'ing
the content-part array (which embedded raw ~52KB base64 as inert text).

Applied to both function_call_output and custom_tool_call_output branches.

Existing translator tests: 88/88 pass.
New tests: 4/4 pass.

* fix(providers): set qwen-web toolCalling to false — web-cookie provider has no native function calling (#8437)

qwen-web is a web-cookie provider that emulates tools via synthetic system
prompt text and <tool> XML parsing, never sending a native tools[] field
upstream. The filterTargetsByRequestCompatibility gate filters out non-tool-
calling targets when the request carries tools, but qwen-web's registry entry
had toolCalling=true, so the filter let it through and a tool-using session
failing over to qwen-web would silently degrade to text-only chat with
'Tool X does not exists' errors.

Sibling web-cookie providers (chatgpt-web, yuanbao-web, claude-web, etc.)
all correctly set toolCalling: false — qwen-web was an outlier introduced
in PR #7874.

Verification:
- LSP diagnostics: clean
- Pattern matches chatgpt-web, yuanbao-web, and other web-cookie providers

* fix(backend): empty upstream response mislabeled as exhausted_connection (#8397)

isEmptyContentFailure guard only matched '/empty content/i' but the actual
error text from detectMalformedNonStream is 'returned an empty response
(no usable choices/output)' — which lacks the word 'content'. Expanded
regex to also match '/empty response/i' so these transient upstream glitches
don't get classified as connection-level exhaustion in combo diagnostics.

Test: 28 existing combo-target-exhaustion tests pass (no new test needed)

* test(#8397): add regression test for empty-response 502 not marking provider/connection exhausted

* test(qwen-web): align registry snapshot with toolCalling:false

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(resilience): scope #8375 input-bound short-circuit to homogeneous remainders

The isInputBoundFailure short-circuit (context_length_exceeded /
context_window_exceeded) fired unconditionally on the first target, aborting
the whole combo even when later targets are a different model with a larger
context window — regressing the intentional heterogeneous-combo fallback that
isContextOverflow400 (#6637) protects. Reproduced with a 2-target combo
(small-context model fails, larger-context model would have succeeded): the
combo never reached target 2.

Scope the short-circuit to remainders where every remaining target shares the
same modelStr as the one that just failed — the "retrying will fail
identically" premise for context_length_exceeded only holds within a
homogeneous same-model pool.

Rebaselines open-sse/services/combo.ts's frozen file-size cap (3642->3679)
for this PR's own combo.ts growth (config/quality/file-size-baseline.json).

Refs #8375

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: Probe Test <probe@example.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: ikelvingo <im.kelvinwong@gmail.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
2026-07-26 03:53:32 -03:00

83 lines
2.9 KiB
TypeScript

/**
* #8375 — A combo whose first target returns `context_length_exceeded` for an
* oversized input must propagate the 400 immediately instead of re-dispatching
* the identical oversized request against other accounts of the same model.
*
* Without this fix:
* - The 400 `context_length_exceeded` is request-scoped and deterministic for
* the same input — every account of the same model will reject it identically.
* - `isRequestScopedUpstreamFailure()` correctly classifies it, but the combo
* loop never acts on that classification to short-circuit.
* - The combo retries MAX_GLOBAL_ATTEMPTS=30 times, burning all attempts, and
* returns a misleading 503 "Maximum combo retry limit reached".
*
* Fix: new `isInputBoundRequestFailure()` predicate that detects input-bound
* deterministic errors. When it fires, the combo returns `{ ok: false, response }`
* from `executeTarget`, which the outer loop treats as fatal — stopping the
* combo and propagating the original 400.
*/
import test from "node:test";
import assert from "node:assert/strict";
import fs from "node:fs";
import os from "node:os";
import path from "node:path";
const TEST_DATA_DIR = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-combo-8375-"));
process.env.DATA_DIR = TEST_DATA_DIR;
process.env.API_KEY_SECRET = process.env.API_KEY_SECRET || "combo-8375-test-secret";
const { handleComboChat } = await import("../../open-sse/services/combo.ts");
const noop = () => {};
const log = { info: noop, warn: noop, debug: noop, error: noop };
function contextLengthExceededResponse() {
return new Response(
JSON.stringify({
error: {
code: "context_length_exceeded",
message:
"Input exceeds the context window for nvidia/z-ai/glm-5.2: estimated 159324 input tokens, limit 128000.",
},
}),
{ status: 400, headers: { "Content-Type": "application/json" } }
);
}
function makeCombo(models: string[]) {
return {
name: "test-combo-8375",
strategy: "priority",
models: models.map((m) => ({ model: m })),
};
}
test("#8375 combo stops at the first context_length_exceeded instead of re-dispatching", async () => {
const modelsCalled: string[] = [];
const handleSingleModel = async (_body: unknown, modelStr: string) => {
modelsCalled.push(modelStr);
return contextLengthExceededResponse();
};
const result = await handleComboChat({
body: { model: "test", messages: [{ role: "user", content: "hi" }] },
combo: makeCombo(["nvidia/z-ai/glm-5.2", "nvidia/z-ai/glm-5.2", "nvidia/z-ai/glm-5.2"]),
handleSingleModel,
log,
settings: {},
allCombos: [],
});
// The guard must short-circuit after the FIRST target — never reach #2 or #3.
assert.equal(
modelsCalled.length,
1,
`input-bound 400 must stop the combo at target 1, but it tried: ${modelsCalled.join(", ")}`
);
assert.equal(
result.status,
400,
"the combo must surface the original 400 to the client, not a 503"
);
});