Files
OmniRoute/tests/unit/combo-empty-content-failover-5085.test.ts
Hernan Javier Ardila Sanchez 3be0a5d290 fix: combo input-bound, Responses->Chat image strip, qwen-web toolCalling, empty-response exhaustion (#8476)
* test(tail): retire stale i18n __MISSING__ repro + fix qianfan website URL

Base-red slice 6, rebased onto the advanced release/v3.8.49 (91fd5f9). The oauth
grok-cli #7610 guard was already fixed on the base by #8027 (it reads the warning
from grokCliAuthJson.ts) — dropped from this slice to avoid a conflicting duplicate.
Remaining two, still red on the current base:

- i18n #7258: the "focused repro" asserted zh-TW.json STILL carries raw __MISSING__:
  placeholders. That backlog was filled (the "no locale has a raw __MISSING__: leaf"
  invariant is the durable guard); retired the now-inverted repro.
- qianfan: Baidu renamed the product page (product/wenxinworkshop -> product-s/
  qianfan_home); updated the expected website URL.

Validated (clean env): i18n 4/0, qianfan 5/0; oauth-modal-grok 2/0 already green on base.

* fix(resilience): short-circuit combo on input-bound failures (context_length_exceeded) (#8375)

isInputBoundRequestFailure() predicate detects deterministic input-bound
errors (context_length_exceeded/context_window_exceeded). The combo loop
propagates the original 400 immediately instead of burning MAX_GLOBAL_ATTEMPTS
retrying identical oversized inputs against every account.

Test: combo-input-bound-failure-8375.test.ts (1 test, 2 assertions)

* fix(resilience): add early-exit in combo dispatcher for input-bound failures (#8375)

When isInputBoundRequestFailure detects context_length_exceeded,
the combo loop returns {ok:false, response} immediately instead of
re-dispatching the oversized request.

Test: node --import tsx/esm --test tests/unit/combo-input-bound-failure-8375.test.ts
- 1 test, 2 assertions, 0 fail

* fix(translator): strip input_image from tool outputs in Responses->Chat downgrade (#8459)

toolOutputContentToString() extracts input_text/output_text parts and
replaces input_image with a placeholder instead of JSON.stringify'ing
the content-part array (which embedded raw ~52KB base64 as inert text).

Applied to both function_call_output and custom_tool_call_output branches.

Existing translator tests: 88/88 pass.
New tests: 4/4 pass.

* fix(providers): set qwen-web toolCalling to false — web-cookie provider has no native function calling (#8437)

qwen-web is a web-cookie provider that emulates tools via synthetic system
prompt text and <tool> XML parsing, never sending a native tools[] field
upstream. The filterTargetsByRequestCompatibility gate filters out non-tool-
calling targets when the request carries tools, but qwen-web's registry entry
had toolCalling=true, so the filter let it through and a tool-using session
failing over to qwen-web would silently degrade to text-only chat with
'Tool X does not exists' errors.

Sibling web-cookie providers (chatgpt-web, yuanbao-web, claude-web, etc.)
all correctly set toolCalling: false — qwen-web was an outlier introduced
in PR #7874.

Verification:
- LSP diagnostics: clean
- Pattern matches chatgpt-web, yuanbao-web, and other web-cookie providers

* fix(backend): empty upstream response mislabeled as exhausted_connection (#8397)

isEmptyContentFailure guard only matched '/empty content/i' but the actual
error text from detectMalformedNonStream is 'returned an empty response
(no usable choices/output)' — which lacks the word 'content'. Expanded
regex to also match '/empty response/i' so these transient upstream glitches
don't get classified as connection-level exhaustion in combo diagnostics.

Test: 28 existing combo-target-exhaustion tests pass (no new test needed)

* test(#8397): add regression test for empty-response 502 not marking provider/connection exhausted

* test(qwen-web): align registry snapshot with toolCalling:false

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(resilience): scope #8375 input-bound short-circuit to homogeneous remainders

The isInputBoundFailure short-circuit (context_length_exceeded /
context_window_exceeded) fired unconditionally on the first target, aborting
the whole combo even when later targets are a different model with a larger
context window — regressing the intentional heterogeneous-combo fallback that
isContextOverflow400 (#6637) protects. Reproduced with a 2-target combo
(small-context model fails, larger-context model would have succeeded): the
combo never reached target 2.

Scope the short-circuit to remainders where every remaining target shares the
same modelStr as the one that just failed — the "retrying will fail
identically" premise for context_length_exceeded only holds within a
homogeneous same-model pool.

Rebaselines open-sse/services/combo.ts's frozen file-size cap (3642->3679)
for this PR's own combo.ts growth (config/quality/file-size-baseline.json).

Refs #8375

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: Probe Test <probe@example.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: ikelvingo <im.kelvinwong@gmail.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
2026-07-26 03:53:32 -03:00

204 lines
6.9 KiB
TypeScript

/**
* #5085 — A multi-leg combo whose first leg returns a 502 "Provider returned
* empty content" must FAIL OVER to the next leg within the same request, not
* surface the 502 to the caller. An empty completion is a fake-success failure
* (HTTP 200 with no content → rewritten to 502 in chatCore), and for a combo
* whose whole purpose is resilience it should behave like any other transient
* leg failure and advance.
*
* Reproduction: a `priority` combo with two legs on DIFFERENT providers. Leg 1
* returns the empty-content 502 (exactly the body buildErrorBody produces in
* chatCore's isEmptyContentResponse branch); leg 2 returns a healthy 200. The
* combo must try leg 2 and surface its 200 — never return the leg-1 502.
*/
import test from "node:test";
import assert from "node:assert/strict";
import fs from "node:fs";
import os from "node:os";
import path from "node:path";
const TEST_DATA_DIR = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-combo-5085-"));
process.env.DATA_DIR = TEST_DATA_DIR;
process.env.API_KEY_SECRET = process.env.API_KEY_SECRET || "combo-5085-test-secret";
const { handleComboChat } = await import("../../open-sse/services/combo.ts");
const noop = () => {};
const log = { info: noop, warn: noop, debug: noop, error: noop };
// Mirrors chatCore's empty-content branch: buildErrorBody(502, "Provider returned empty content").
function emptyContent502() {
return new Response(
JSON.stringify({ error: { message: "Provider returned empty content", type: "bad_gateway" } }),
{ status: 502, headers: { "Content-Type": "application/json" } }
);
}
function healthy200(model: string) {
return new Response(
JSON.stringify({
id: "ok",
object: "chat.completion",
model,
choices: [
{
index: 0,
message: { role: "assistant", content: "hello from " + model },
finish_reason: "stop",
},
],
}),
{ status: 200, headers: { "Content-Type": "application/json" } }
);
}
function makeCombo(models: string[]) {
return {
name: "test-combo-5085",
strategy: "priority",
models: models.map((m) => ({ model: m })),
};
}
test("#5085 combo fails over to the next leg when leg 1 returns empty-content 502", async () => {
const modelsCalled: string[] = [];
const handleSingleModel = async (_body: unknown, modelStr: string) => {
modelsCalled.push(modelStr);
// First leg (different provider) returns the empty-content 502; second leg is healthy.
if (modelsCalled.length === 1) return emptyContent502();
return healthy200(modelStr);
};
const result = await handleComboChat({
body: { model: "test", messages: [{ role: "user", content: "hi" }] },
combo: makeCombo(["nvidia/minimaxai/minimax-m3", "openai/gpt-4o-mini"]),
handleSingleModel,
log,
settings: {},
allCombos: [],
});
assert.equal(
modelsCalled.length,
2,
`empty-content 502 on leg 1 must advance to leg 2, but tried: ${modelsCalled.join(", ")}`
);
assert.equal(
result.status,
200,
"the combo must surface the healthy second leg's 200, not the leg-1 empty-content 502"
);
});
// ── Precise unit test of the exhaustion classifier ──────────────────────────
// The integration test above already passes because the empty-content leg and
// the healthy leg are on different providers. The real defect is provider-level:
// an empty-content 502 is currently classified as a CONNECTION-level failure
// (502 ∈ CONNECTION_LEVEL_ERROR_STATUSES), which marks the whole provider/
// connection exhausted and skips every REMAINING SAME-PROVIDER leg (#1731v2).
// An empty completion arrived on a HEALTHY connection (HTTP 200, no content) and
// must not be treated as a bad connection.
const { applyComboTargetExhaustion } =
await import("../../open-sse/services/combo/targetExhaustion.ts");
function makeTarget(provider: string, modelStr: string, connectionId: string | null = null) {
return {
kind: "model" as const,
stepId: "s",
executionKey: "e",
modelStr,
provider,
providerId: provider,
connectionId,
weight: 1,
label: null,
};
}
function freshSets() {
return {
exhaustedProviders: new Set<string>(),
exhaustedConnections: new Set<string>(),
transientRateLimitedProviders: new Set<string>(),
};
}
test("#5085 empty-content 502 must NOT mark the provider/connection exhausted (model-level, not connection-level)", () => {
const sets = freshSets();
const providerExhausted = applyComboTargetExhaustion(
makeTarget("nvidia", "nvidia/minimaxai/minimax-m3"),
{
result: { status: 502, headers: new Headers() },
fallbackResult: { reason: "server_error" },
errorText: "Provider returned empty content",
rawModel: "minimaxai/minimax-m3",
isTokenLimitBreach: false,
allAccountsRateLimited: false,
sets,
log,
tag: "COMBO",
exhaustedLogLevel: "info",
}
);
assert.equal(providerExhausted, false, "empty-content is not a quota exhaustion");
assert.equal(
sets.exhaustedProviders.has("nvidia"),
false,
"empty-content 502 must NOT mark the whole provider exhausted — remaining same-provider legs must still be tried"
);
});
test("#8397 empty-response 502 (no usable choices/output) must NOT mark provider/connection exhausted", () => {
const sets = freshSets();
const providerExhausted = applyComboTargetExhaustion(
makeTarget("nvidia", "nvidia/minimaxai/minimax-m3"),
{
result: { status: 502, headers: new Headers() },
fallbackResult: { reason: "server_error" },
errorText: "upstream returned an empty response without usable output",
rawModel: "minimaxai/minimax-m3",
isTokenLimitBreach: false,
allAccountsRateLimited: false,
sets,
log,
tag: "COMBO",
exhaustedLogLevel: "info",
}
);
assert.equal(providerExhausted, false, "empty-response is not a quota exhaustion");
assert.equal(
sets.exhaustedProviders.has("nvidia"),
false,
"empty-response 502 must NOT mark the whole provider exhausted"
);
assert.equal(
sets.exhaustedConnections.size,
0,
"empty-response 502 must NOT mark any connection exhausted"
);
});
test("#5085 a real connection-level 502 (gateway error) STILL marks the provider exhausted", () => {
const sets = freshSets();
applyComboTargetExhaustion(makeTarget("nvidia", "nvidia/minimaxai/minimax-m3"), {
result: { status: 502, headers: new Headers() },
fallbackResult: { reason: "server_error" },
errorText: "Bad gateway: upstream connection reset",
rawModel: "minimaxai/minimax-m3",
isTokenLimitBreach: false,
allAccountsRateLimited: false,
sets,
log,
tag: "COMBO",
exhaustedLogLevel: "info",
});
assert.equal(
sets.exhaustedProviders.has("nvidia"),
true,
"a genuine gateway 502 must still mark the provider connection-exhausted (#1731v2 preserved)"
);
});