Files
OmniRoute/tests/unit/t23-t24-fallback-resilience.test.ts
Diego Rodrigues de Sa e Souza 36abd86929 fix(ci): clear the 08-08 base-red layers — dead-code, prod crash in chat.ts, Responses payload regression, born-red stdio test, gate drifts (#9757)
* fix(ci): drop unused RadarReferrals type export — dead-code ratchet back to 227 baseline

The radar referral-links feature (#9697) exported the inferred type
RadarReferrals from feedSchema.ts but nothing imports it (the singular
RadarReferral is the consumed type). knip counts it as a new dead export,
pushing the dead-code ratchet to 228 > 227 and failing Fast Quality Gates
on every PR born after the merge. RadarReferralsSchema itself stays — it
is used by RadarFeedSchema.

Refs #9737

* fix(ci): clear the 08-08 base-red layer — prod crash in chat.ts, Responses API payload regression, born-red stdio test, gate drifts

Six independent base-reds from the 08-07 evening merge batch, each verified
against the pure release/v3.8.50 tip:

- src/sse/handlers/chat.ts: #9467's squash carried a refactor hunk that
  renamed the all-rate-limited breaker guard to an UNDEFINED variable
  (isAllRateLimited) — a production ReferenceError on the all-accounts-429
  path (chat.ts is outside typecheck:core scope, so only tests caught it).
  Restore credentials?.allRateLimited. Guard: chat-rate-limit-body-lock (2/2),
  also un-breaks batch_api and chat-combo-live-test.
- open-sse/utils/stream.ts: #9315 switched providerPayload summaries to the
  accumulated responseBody, but in passthrough paths that body is synthesized
  in chat-completion shape — Responses API lost its `response` object in the
  dashboard payload. Keep the events-derived summary for OPENAI_RESPONSES
  only. Guard: stream-utils + stream-collector-9315 suites (51/51).
- tests/unit/mcp-stdio-json-purity.test.ts: born red — the full CLI chain
  takes ~10s (2x tsx import + DB init) and the test slept a fixed 4s. Poll
  for the first stdout line with a 60s deadline instead.
- tests/unit/plugins-route-error-sanitization.test.ts: register #9445's new
  marketplace/install route in PLUGIN_ROUTES (route already sanitizes) (33/33).
- tests/unit/provider-models-route-codex.test.ts: realign pinned GPT-5.6
  input limit to #9432's deliberate 272000→922000 bump (7/7).
- lint: fix 11 no-explicit-any errors in repro-9630 + specialty-9293 tests,
  prune 1 orphaned suppression, allowlist the opencode-ai devDependency
  (#8869, publisher-verified), and reword a doc line the fabricated-docs
  gate misread as an env var.

Gates re-verified locally: lint:json --max-warnings 0 exit 0, dead-code 227,
typecheck:core clean, check:deps OK, check:fabricated-docs OK.

Refs #9737

* fix(ci): clear the third 08-08 base-red layer — invalid ru rule pack, stale event pin, orphaned UI repro test, pack/mutation/file-size drifts

Follow-up to the previous layer: the serial fast-gates chain unmasked one
more stratum after file-size/dead-code went green, all verified against the
merged release/v3.8.50 tip:

- compression rules ru/ultra.json (#9581): two rules shipped
  minIntensity "notes", which is not a valid CavemanIntensity
  (lite|full|ultra) — loading ANY language pack list threw and killed the
  rtk-loader suite. Mapped both to "ultra" (they are the most aggressive
  punctuation/case rules, matching the en pack tiers). 2/2.
- plugins-welcome-banner-e2e: #9668 added the onStreamComplete builtin
  event (real emission path via runOnStreamCompleteHooks) and missed this
  pinned-list sibling. 35/35.
- tests/unit/free-pool-frontend-repro (#9046): landed as .tsx with
  node:test semantics — no runner collects tests/unit/*.tsx, so it NEVER
  ran (test-discovery NEW-orphan). It contains zero JSX; renamed to .test.ts
  so the unit runner's existing glob collects it. 5/5 (first real run).
- pack-policy: allow + require bin/mcpStdioConsoleGuard.mjs (#9281) — it is
  preloaded via node --import by bin/mcp-server.mjs, so a published artifact
  without it crashes 'omniroute --mcp' at startup.
- stryker.conf.json: add 5 covering unit tests from the batch (#8779/#9204/
  #9330/#9630/openrouter-passthrough) to tap.testFiles (--strict drift).
- file-size-baseline: consolidate the base-drift rebaseline for the 12
  files grown by the 08-06..08-08 batches (#9616's entries never reached the
  base; measured on this branch's tree — this PR's own source edits add zero
  lines to any frozen file).

Local battery: file-size/deps/test-discovery/mutation/pack-policy/dead-code/
duplication/docs-all/secrets/vuln/workflows ratchets all exit 0; full lint
gate --max-warnings 0 exit 0.

Refs #9737

* fix(types): clear the 3 uncovered open-sse-typecheck regressions + realign combo skip-code siblings

Fourth base-red layer unmasked by the serial gates. The other 4 typecheck
regressions (codex.ts, kiro.ts, tierResolver.test.ts, translator/index.ts)
already have dedicated open [TS7] PRs (#9748/#9753/#9742/#9747) — not
duplicated here. This commit covers only what no open PR owns:

- devin-agentic/serializer.ts TS2367: drop the dead 'role === "system"'
  branch — the guard above already narrows role to user|assistant (system
  throws unsupported_role). Devin suites 104/104.
- raycast.ts TS2416: the buildHeaders 'override' never matched the base
  signature (2nd param is the signed payload string, not the stream
  boolean) — renamed to a private buildRaycastRequestHeaders helper so a
  polymorphic buildHeaders(credentials, true) call can never bind here.
- modelMetadataRegistry.ts TS2352: PricingByProvider → nested-record cast
  now goes through unknown (shape is runtime-guarded by findInsensitive).
- combo-routing-engine.test.ts: realign 2 pre-dispatch-skip expectations to
  #9630's deliberate ALL_TARGETS_SKIPPED contract (87/87).

Refs #9737

* fix(ci): clear the fifth 08-08 base-red layer — reasoning-placeholder contract sweep, GPT-5.6 limits sweep, vi key parity

The 08-08 merges (#9610 reasoning replay, #9432 GPT-5.6 limits, #9630 combo
skip codes, #9336 provider key links) each changed a contract and left
sibling tests pinning the old one. Full grep sweep per contract, not just
the shard that happened to go red:

- reasoning placeholder (#9573/#9610): the fix DELIBERATELY removed
  NON_ANTHROPIC_THINKING_PLACEHOLDER injection on cache miss — the model
  echoed the placeholder as its own reasoning (empty stop) and re-poisoned
  cache + client history; DeepSeek's 400 is specific to an EMPTY STRING, not
  an absent field. Realigned reasoning-cache (2 cases, renamed to describe
  omission) + tool-request-sanitization (1 case + dead import). 60/60.
- GPT-5.6 Codex limits (#9432, 272000 -> 1050000 ctx / 922000 input):
  realigned vscode-token-routes-gpt56 (2) + vscode-token-routes (3). 43/43
  together with t23-t24.
- combo skip codes (#9630): t23-t24-fallback-resilience T24 now expects
  ALL_TARGETS_SKIPPED like the combo-routing-engine siblings.
- vi.json key parity: #9336 added providers.getApiKey/getApiKeyDescription
  to en.json without syncing vi (the only locale with a parity gate).
  Translated both; providers block reordered to match en key order. 5/5.
- pack-artifact-policy.test.ts: sibling of this PR's own required-paths
  change (bin/mcpStdioConsoleGuard.mjs). 10/10.
- combo-routing-engine.test.ts: dropped the 6 comment lines added in the
  previous commit so the frozen test file-size stays at its baseline (the
  rationale lives in that commit message, not the test body).

Gates: file-size, test-discovery, mutation-test-coverage, pack-policy,
open-sse-typecheck, dead-code all exit 0.

Refs #9737

* fix(translator): keep the reasoning_content placeholder for Xiaomi MiMo — #9610 traded one live 400 for another

The xiaomi-mimo replay test (9router#1321) went red on the base after #9610
removed the NON_ANTHROPIC_THINKING_PLACEHOLDER injection globally. That test
is NOT stale — it guards a documented upstream 400 ('Param Incorrect: The
reasoning_content in the thinking mode must be passed back to the API'), so
realigning it would have masked a reintroduced production bug.

Two real bugs conflict here:
- #9573: forwarding the placeholder makes the model continue its chain of
  thought FROM that text (echo -> empty stop) and re-poisons cache/history.
- 9router#1321/#1337: omitting reasoning_content on a plain replay turn makes
  Xiaomi MiMo reject the request outright.

#9610's evidence for omitting is provider-specific — it verified that
deepseek-v4-flash accepts an ABSENT field. It does not extend to MiMo. So the
omission stays for every provider #9610 covered, and the placeholder survives
the cache miss only for xiaomi-mimo (new requiresReasoningContentPresence
predicate next to isReasoningOnlyReplayTarget). The echo that comes back is
still stripped on the way in by isInternalReasoningPlaceholder(), so #9573's
cache/history poisoning stays fixed for MiMo too.

Both contracts now hold simultaneously: xiaomi-mimo replay + reasoning-cache +
tool-request-sanitization 61/61; placeholder-strip/responses/translator/combo
regression sweep 168/168. Gates: file-size, open-sse-typecheck, dead-code,
mutation-test-coverage exit 0; typecheck:core clean.

A live check on the VPS (Hard Rule #18 path 2) is the only way to confirm the
DeepSeek half of #9610's empirical claim; flagging it in the PR rather than
widening this fix on speculation.

Refs #9737

* test(translator): pin the reasoning-placeholder provider scope so neither half of the conflict can silently re-break

#9610 removed the placeholder globally on the strength of ONE provider's
observed behavior (deepseek-v4-flash accepting an absent reasoning_content),
which re-opened the MiMo 400 (9router#1321). The previous commit scoped the
placeholder to xiaomi-mimo; this pins BOTH directions in one test so the next
global edit fails loudly instead of trading the bugs again:

- xiaomi-mimo plain replay turn, cache miss -> reasoning_content present
  (narrowing the scope away from MiMo re-opens 9router#1321)
- deepseek plain replay turn, cache miss -> reasoning_content absent
  (widening it back to DeepSeek re-opens the #9573 echo bug)

Guard verified by mutation: forcing requiresReasoningContentPresence() to
return true makes the DeepSeek half fail (1 pass / 1 fail), and the file was
restored from the pre-probe copy before committing.

Also checked kimi-coding/kimi-coding-apikey, the other strict-contract entries
in REASONING_REPLAY_PROVIDERS: their originating PR (#7673) fixes capture and
replay of REAL reasoning and documents no 400 on an absent field, so they stay
out of the placeholder scope — evidence-scoped, not speculatively widened.

Reasoning suites together: 87/87. Gates: file-size, test-discovery,
mutation-test-coverage, dead-code exit 0; eslint clean.

Refs #9737

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-08 09:08:45 -03:00

198 lines
6.5 KiB
TypeScript

import test from "node:test";
import assert from "node:assert/strict";
const { checkFallbackError } = await import("../../open-sse/services/accountFallback.ts");
const { handleComboChat } = await import("../../open-sse/services/combo.ts");
const { resetAllCircuitBreakers } = await import("../../src/shared/utils/circuitBreaker.ts");
test.beforeEach(() => {
resetAllCircuitBreakers();
});
function createLog() {
const entries = [];
return {
info: (tag: string, msg: string) => entries.push({ level: "info", tag, msg }),
warn: (tag: string, msg: string) => entries.push({ level: "warn", tag, msg }),
error: (tag: string, msg: string) => entries.push({ level: "error", tag, msg }),
debug: (tag: string, msg: string) => entries.push({ level: "debug", tag, msg }),
entries,
};
}
function createStatusSequenceHandler(sequence) {
let idx = 0;
return async () => {
const step = sequence[idx++] || { status: 200 };
if (step.status === 200) {
return new Response(JSON.stringify({ ok: true }), { status: 200 });
}
return new Response(
JSON.stringify({
error: { message: step.message || `Error ${step.status}` },
}),
{
status: step.status,
headers: step.headers || { "content-type": "application/json" },
}
);
};
}
test("T23: 429 with long Retry-After uses real reset cooldown instead of short exponential backoff", () => {
const headers = new Headers({ "retry-after": "3600" });
const result = checkFallbackError(429, "Rate limit exceeded", 2, null, "groq", headers);
assert.equal(result.shouldFallback, true);
assert.equal(result.reason, "rate_limit_exceeded");
assert.equal(result.newBackoffLevel, 0);
assert.ok(result.cooldownMs > 3_590_000);
});
test("T24: combo awaits short 503 cooldown before falling through to next model", async () => {
const log = createLog();
const result = await handleComboChat({
body: {},
combo: {
name: "t24-short-cooldown",
strategy: "priority",
// Cross-provider targets: a 503 marks the failing provider's remaining same-provider
// targets for skip (#1731v2), so the fallthrough target must be a DIFFERENT provider
// for this cooldown-wait test to exercise the fall-through-to-next-model path.
models: [
{ model: "groq/model-a", weight: 0 },
{ model: "openai/model-b", weight: 0 },
],
config: { fallbackDelayMs: 2000, maxRetries: 1 },
},
// Two transient failures on first model, then success on fallback model.
handleSingleModel: createStatusSequenceHandler([
{ status: 503 },
{ status: 503 },
{ status: 200 },
]),
isModelAvailable: () => true,
log,
settings: null,
allCombos: null,
});
assert.equal(result.ok, true);
// checkFallbackError returns COOLDOWN_MS.transient (5000ms) for a plain 503.
// fallbackDelayMs=2000, cooldownMs=5000 ≤ MAX_FALLBACK_WAIT_MS(5000) → fallbackWaitMs=2000ms.
// The combo MUST emit a debug log before waiting, proving the wait behavior is wired.
const waitLog = log.entries.find((e) => e.msg.includes("Waiting") && e.msg.includes("fallback"));
assert.ok(waitLog, "combo must emit a debug wait-before-fallback log for short 503 cooldowns");
});
test("T24: combo skips wait when 503 cooldown is long (>5s)", async () => {
const log = createLog();
const result = await handleComboChat({
body: {},
combo: {
name: "t24-long-cooldown",
strategy: "priority",
// Cross-provider targets (see t24-short-cooldown): the fall-through target must be a
// different provider so the #1731v2 same-provider skip doesn't short-circuit it.
models: [
{ model: "groq/model-a", weight: 0 },
{ model: "openai/model-b", weight: 0 },
],
config: { fallbackDelayMs: 2000, maxRetries: 1 },
},
handleSingleModel: createStatusSequenceHandler([
{
status: 503,
message: "rate limit exceeded",
headers: { "content-type": "application/json", "retry-after": "120" },
},
{
status: 503,
message: "rate limit exceeded",
headers: { "content-type": "application/json", "retry-after": "120" },
},
{ status: 200 },
]),
isModelAvailable: () => true,
log,
settings: null,
allCombos: null,
});
assert.equal(result.ok, true);
const waitLog = log.entries.find((e) => e.msg.includes("Waiting") && e.msg.includes("fallback"));
assert.equal(waitLog, undefined);
});
test("T24: all inactive accounts return 503 service_unavailable (not 406)", async () => {
const result = await handleComboChat({
body: {},
combo: {
name: "t24-all-inactive",
strategy: "priority",
models: [
{ model: "groq/model-a", weight: 0 },
{ model: "groq/model-b", weight: 0 },
],
},
handleSingleModel: async () => {
throw new Error("handleSingleModel should not be called when all models are unavailable");
},
isModelAvailable: () => false,
log: createLog(),
settings: null,
allCombos: null,
});
assert.equal(result.status, 503);
const body = (await result.json()) as any;
assert.equal(body.error?.code, "ALL_TARGETS_SKIPPED");
});
test("combo falls through 400s and reaches the next model", async () => {
const calls = [];
const sequence = [
{ status: 429, message: "No capacity available for model gemini-3.1-pro-preview" },
{ status: 400, message: "bad request" },
{ status: 200 },
];
const result = await handleComboChat({
body: {},
combo: {
name: "t24-provider-scoped-400",
strategy: "priority",
models: [
{ model: "free/gemini-3.1-pro-preview", weight: 0 },
{ model: "aio/gemini-3.1-pro-preview-thinking-high", weight: 0 },
{ model: "openrouter/google/gemini-3.1-pro-preview", weight: 0 },
],
config: { maxRetries: 0 },
},
handleSingleModel: async (_body, modelStr) => {
calls.push(modelStr);
const step = sequence[calls.length - 1] || { status: 200 };
if (step.status === 200) {
return new Response(JSON.stringify({ ok: true }), { status: 200 });
}
return new Response(JSON.stringify({ error: { message: step.message } }), {
status: step.status,
headers: { "content-type": "application/json" },
});
},
isModelAvailable: () => true,
log: createLog(),
settings: null,
allCombos: null,
});
assert.equal(result.ok, true);
assert.deepEqual(calls, [
"free/gemini-3.1-pro-preview",
"aio/gemini-3.1-pro-preview-thinking-high",
"openrouter/google/gemini-3.1-pro-preview",
]);
});