Files
OmniRoute/tests/e2e/system-failover.test.ts
Diego Rodrigues de Sa e Souza 6248699ce5 Release/v3.8.0 — full changelog with 660+ commits (#2419)
* fix(cli-tools): guard modelId type before calling indexOf

E2E shakedown v3.8.0: cli-tools quebrava com TypeError quando dynamicModels
continha entradas sem .id (objeto retornado diretamente em vez de string).

* fix(offline): avoid SSR/CSR hydration mismatch on navigator.onLine

Replace useState+lazy-initializer with useSyncExternalStore so the server
snapshot (() => false) and client snapshot (() => navigator.onLine) are
declared separately. React hydrates with the server value and switches to
the real online status client-side without a mismatch.

* chore(i18n): add missing en.json keys for translator, cli-tools, memory, onboarding

Adds 58 missing keys identified by the new dashboard audit script:
- cliTools: 18 custom CLI builder keys (CustomCliCard)
- translator: 24 keys covering stream transformer, live monitor, test bench
- memory: 12 health/pagination/dialog keys
- onboarding.tier: 8 keys for the tier tour walkthrough

Also adds scripts/i18n/audit-dashboard-pages.mjs which scans all dashboard
pages, reports t() calls referencing missing en.json keys, and flags
candidate hardcoded JSX/attribute strings.

* chore(i18n): replace hardcoded UI text with t() calls across dashboard (round 1)

Subagents refactored 8 high-impact dashboard pages, replacing 81 of the
407 hardcoded English/PT strings flagged by the audit with proper
useTranslations() lookups. Added 73 corresponding keys to en.json across
the home, apiManager, providers, settings, and usage namespaces.

Pages affected:
- BudgetTab (27 → 0)
- HomePageClient (2 → 0)
- RoutingTab (25 → 7)
- ResilienceTab (38 → 18)
- SystemStorageTab (42 → 21)
- providers/[id] (17 → 15)
- ApiManagerPageClient (14 → 13)
- OneproxyTab (13 → 10)

Also adds two helper scripts:
- scripts/i18n/extract-keys-from-diff.mjs — extracts new keys from git diff
- scripts/i18n/merge-keys.mjs — merges a pending-keys JSON into en.json

Remaining hardcoded strings will be addressed in follow-up rounds.

* chore(i18n): replace hardcoded UI text with t() calls across dashboard (round 2)

Continues round 1 (commit 8d34f4c65). Round-2 subagents refactored
additional dashboard pages, replacing 77 more hardcoded strings with
useTranslations() lookups. Added 79 corresponding keys to en.json
across the a2aDashboard, agents, analytics, apiManager, cliTools,
common, and settings namespaces.

Pages affected:
- a2a/page (new useTranslations + 6 keys)
- agent-skills/page (new useTranslations + 9 keys)
- AutoRoutingAnalyticsTab (new useTranslations + 6 keys)
- AppearanceTab (8 → 6 remaining)
- OneproxyTab (10 → 0)
- ResilienceTab (18 → 0 missing key)
- RoutingTab (7 → 0 missing key)
- VisionBridgeSettingsTab (new useTranslations + 6 keys)
- CopilotToolCard (7 → 0 missing key)
- ApiManagerPageClient (13 → 0 missing key)
- gamification/admin (new useTranslations + 7 keys)

Hardcoded total: 326 → 249. Real missing keys: 0 (the 6 still flagged
are false positives in exampleTemplates.tsx where t is passed as a
parameter — keys exist at translator.templatePayloads.*).

* chore(i18n): replace hardcoded UI text with t() calls across dashboard (round 3)

Round-3 subagents and manual edits refactored 9 more dashboard pages
(plus 2 small extras), replacing ~80 hardcoded strings with
useTranslations() lookups. Added 79 corresponding keys to en.json
across analytics, cloudAgents, combos, common, health, settings, and
usage namespaces.

Pages affected:
- analytics/ComboHealthTab (new useTranslations + 15 keys)
- analytics/CompressionAnalyticsTab (new useTranslations + 11 keys)
- settings/SystemStorageTab (21 → 0 missing key)
- tokens/page (new useTranslations + 13 keys)
- usage/BudgetTab (9 missing fixed)
- health/page (manual: 6 keys)
- cloud-agents/page (manual: 3 keys)
- combos/page (manual: 1 key)

Hardcoded total: 249 → 164. Real missing keys: 0 (6 remaining are
exampleTemplates.tsx false positives).

Also adds scripts/i18n/build-pending-from-missing.mjs which reads
_audit.json and locates English values from HEAD to rebuild
_pending-keys.json after race-condition resets between subagent edits.

* chore(i18n): localize remaining dashboard settings labels

Replace hardcoded labels in compression and resilience settings with
translation lookups to continue the dashboard i18n cleanup.

Add the v3.8.0 dashboard shakedown runbook to document the manual
smoke-test process and known dev environment pitfalls.

* chore(i18n): replace hardcoded UI text with t() calls across dashboard (round 4)

Round-4 subagent + manual key-resolution refactored remaining strings in
3 high-traffic settings/API tabs, plus extracted English values for
keys that were already added as t() calls but lost during the previous
en.json race-condition resets.

Pages affected:
- api-manager/ApiManagerPageClient (7 → 0 missing key)
- settings/CompressionSettingsTab (8 → 0 missing key)
- settings/MemorySkillsTab (8 → 0 missing key)
- settings/ResilienceTab (4 more keys recovered)

Hardcoded total: 164 → 140. Real missing keys: 0 (6 remaining are the
exampleTemplates.tsx false positives — t passed as parameter).

* chore(i18n): replace hardcoded UI text with t() calls across dashboard (round 5)

Round-5 agent began processing the remaining smaller dashboard files.
Added 5 more keys to en.json for providers/[id]/page.tsx OAuth flow
labels and the cross-OS auto-detection hint.

Pages affected:
- providers/[id]/page.tsx (5 keys)

Hardcoded total: 140 → 136. Real missing keys: 0.

* chore(i18n): resolve last 2 missing providers/[id] keys

Adds providerDetailMyClaudeAccountPlaceholder and
providerDetailPathAutoDetected — the final user-visible labels in the
providers/[id] page that the round-5 subagent rewrote to t() calls
without yet adding to en.json.

Real missing keys: 0 (6 remaining are exampleTemplates.tsx false
positives — t is passed as a parameter so the audit cannot resolve the
namespace; keys do exist at translator.templatePayloads.*).

* chore(i18n): replace hardcoded UI text with t() calls across dashboard (round 6 — 10 parallel agents)

Round-6 dispatched 10 parallel subagents covering all 57 remaining
dashboard files. Each agent worked on a disjoint file set to avoid
en.json race conditions. Added ~60 new i18n keys across 9 namespaces
covering small UI labels, table headers, search placeholders, and
empty-state messages.

Major changes:
- analytics: SearchAnalyticsTab, ProviderUtilizationTab, DiversityScoreCard, CompressionAnalyticsTab (new useTranslations + keys)
- batch: BatchDetailModal, BatchListTab, FileDetailModal, FilesListTab (new useTranslations + keys)
- settings: CliproxyapiSettingsTab, PayloadRulesTab, ModelCooldownsCard, AppearanceTab, PricingTab (mostly new useTranslations)
- endpoint: TokenSaverCard, ApiEndpointsTab, EndpointPageClient
- cache: CachePerformance, IdempotencyLayer, ReasoningCacheTab, MediaPageClient, page
- combos: IntelligentComboPanel, page
- playground: ChatPlayground, SearchPlayground
- providers: ProviderCard
- onboarding: TierFlowDiagram
- changelog: ChangelogViewer
- home: ProviderTopology, TierCoverageWidget, BootstrapBanner, BadgeToast
- usage: BudgetTab, BudgetTelemetryCards, QuotaTable
- quotaShare: QuotaSharePageClient
- profile: page
- leaderboard: page
- skills: page

Hardcoded total: 131 → 60. Real missing keys: 0 plus 1 false-positive
for combos.modePack (lookup via prop-passed t).

* chore(i18n): finalize round-6 keys for batch/cache/endpoint/usage

Adds the remaining keys produced by parallel agents A4, A6, A8, A9:
- common: batch-related labels (BatchDetailModal, BatchListTab,
  FileDetailModal, FilesListTab, page) + profile/leaderboard
- cache: hit rate, latency, retry, avg chars
- endpoint: token saver, API endpoints, copy URL, cloud/local labels
- usage: noSpend, activeSessions, quotaAlerts, budget timing
- skills: install/marketplace/filter
- proxyRegistry/quotaShare/mcpDashboard: misc labels

Hardcoded total: 60 → 48. Real missing keys: 0 (modePack remaining is a
false positive — combos.modePack exists but the audit can't resolve it
since IntelligentComboPanel receives t as a prop).

* fix(playground): dedupe filteredModels to avoid duplicate React key warning

The /v1/models endpoint can return the same model id twice (e.g., when a
model is listed by both an alias and its canonical provider), which made
the <Select> emit two <option> elements with the same key — triggering
"Encountered two children with the same key, codex/gpt-5.5".

Replace the chained filter + map with a single pass that skips ids
already added.

* fix(playground): guard against non-string model ids before .split/.startsWith

The /v1/models endpoint can include synthetic entries (combos, locals,
in-progress imports) with a null/undefined id. The playground used to
call m.id.split("/") in the provider-discovery loop, which threw on the
first non-string entry; the surrounding .catch(() => {}) silently
swallowed the error, so the provider/model/account dropdowns ended up
empty even though /v1/models returned thousands of valid entries.

- Skip entries without a string id before split/startsWith.
- Log the rejection in the .catch handler so future regressions are
  visible in DevTools instead of silently emptying the UI.

* fix(playground): guard ChatPlayground filteredModels for non-string ids

Same root cause as commit 49fe356b9: ChatPlayground filtered models
with m.id.startsWith(...) which crashed on null/undefined ids returned
by /v1/models (synthetic combo entries). Apply the same defensive guard
and dedupe used in the parent page.

* fix(claude): drop orphan tool_result after fixToolAdjacency strip (discussion #2410)

Discussion #2410 reports Claude returning 400 for sequences like:
  assistant: tool_use(id=X)
  user: <plain text>           ← breaks adjacency
  user: tool_result(id=X)

The previous round added `fixToolAdjacency` (commit 44d9abac9) which
correctly strips the orphan tool_use from the assistant message. But
that left the now-unmatched tool_result intact, so the upstream
rejected the request with:

  messages.N.content.M: unexpected `tool_use_id` found in `tool_result`
  blocks: X. Each tool_result block must have a corresponding tool_use
  block in the previous message.

Fix: after running `fixToolAdjacency`, re-run `fixToolPairs` to drop
the orphaned tool_result blocks. All three call sites updated:
  - contextManager.purifyHistory (both inside the binary-search loop
    and the final pass)
  - BaseExecutor message-prep (Claude path)
  - claudeCodeCompatible request signer

Also tightens an unrelated dynamic-key access in
readNestedString (claudeCodeCompatible) to satisfy the prototype-
pollution scanner triggered by the post-tool semgrep hook.

* fix(mitm): point runtime manager re-export to js entrypoint

Use the emitted `.js` path for the runtime manager re-export so dynamic
runtime loading resolves correctly outside the Turbopack alias handling.

* docs: add AgentRouter setup guide (#2422)

Integrated into release/v3.8.0 — AgentRouter setup guide docs.

* feat: add new feature on combos - falloverBeforeRetry (#2417)

Integrated into release/v3.8.0 — falloverBeforeRetry for per-model quota skipping in combos.

* feat(batch): implement 10 feature requests harvested  (#2414)

Integrated into release/v3.8.0 — batch of 10 feature requests: llama.cpp local provider, upstream error exposure, Termux detection, providers rotate CLI, t3.chat web skeleton, Zed Docker integration, Kiro multi-account OAuth isolation, auto-combo cost blending, auto-combo context filter, combo provider-level exhaustion tracking (#1731). Conflicts with #2417 (falloverBeforeRetry) resolved.

* fix(gamification): resolve SQL bug, auth gap, pagination, and anomaly scoring (#2421)

Integrated into release/v3.8.0 — 6 critical gamification bug fixes: SQL SELECT in checkActionCountBadges, federation auth enforcement, leaderboard pagination offset, real z-score computation, addXp level calculation, and barrel index.ts

* docs(changelog): add post-release entries for #2414 #2417 #2421 #2422

- feat(batch): T3-Chat-Web executor, exhaustedProviders set (#1731), Zed Docker
- feat(combos): falloverBeforeRetry + setTry loop (#2417 — @hartmark)
- fix(gamification): SQL SELECT bug, federation auth, pagination, z-score (#2421 — @oyi77)
- docs: AgentRouter setup guide (#2422 — @leninejunior)

* fix(security): resolve CodeQL random/password-hash alerts and sync docs & tests

---------

Co-authored-by: diegosouzapw <diego.souza.pw@gmail.com>
Co-authored-by: Lenine Júnior <lenine@engrene.com.br>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
2026-05-20 02:05:50 -03:00

700 lines
24 KiB
TypeScript

import test from "node:test";
import assert from "node:assert/strict";
import fs from "node:fs";
import fsp from "node:fs/promises";
import os from "node:os";
import path from "node:path";
import net from "node:net";
import { spawn } from "node:child_process";
import { fileURLToPath } from "node:url";
import { MockUpstreamServer, buildCompletion, buildError } from "./helpers/mockUpstreamServer.ts";
const TEST_DATA_DIR = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-system-failover-"));
const DASHBOARD_PORT = await getFreePort();
const REPO_ROOT = fileURLToPath(new URL("../..", import.meta.url));
process.env.DATA_DIR = TEST_DATA_DIR;
process.env.DISABLE_SQLITE_AUTO_BACKUP = "true";
process.env.API_KEY_SECRET = process.env.API_KEY_SECRET || "system-failover-secret-123456";
process.env.REQUIRE_API_KEY = "false";
const core = await import("../../src/lib/db/core.ts");
const providersDb = await import("../../src/lib/db/providers.ts");
const combosDb = await import("../../src/lib/db/combos.ts");
const settingsDb = await import("../../src/lib/db/settings.ts");
const accountFallback = await import("../../open-sse/services/accountFallback.ts");
function resetConnectionCooldowns() {
accountFallback.clearAllModelLockouts();
const db = core.getDbInstance() as any;
db.prepare(
`UPDATE provider_connections
SET rate_limited_until = NULL,
test_status = 'active',
backoff_level = 0,
last_error = NULL,
last_error_type = NULL,
last_error_source = NULL,
error_code = NULL,
last_error_at = NULL
WHERE rate_limited_until IS NOT NULL
OR test_status != 'active'`
).run();
db.pragma("wal_checkpoint(TRUNCATE)");
}
function getFreePort() {
return new Promise<number>((resolve, reject) => {
const server = net.createServer();
server.once("error", reject);
server.listen(0, "127.0.0.1", () => {
const address = server.address();
if (!address || typeof address === "string") {
server.close();
reject(new Error("Failed to allocate a free port"));
return;
}
const { port } = address;
server.close((closeError) => {
if (closeError) reject(closeError);
else resolve(port);
});
});
});
}
function sleep(ms: number) {
return new Promise((resolve) => setTimeout(resolve, ms));
}
async function seedProvider(label: string, apiKey: string, baseUrl: string) {
const providerId = `openai-compatible-sys-${label}`;
await providersDb.createProviderNode({
id: providerId,
type: "openai-compatible",
name: `System ${label}`,
prefix: label,
apiType: "chat",
baseUrl,
});
await providersDb.createProviderConnection({
provider: providerId,
authType: "apikey",
name: `conn-${label}`,
apiKey,
isActive: true,
testStatus: "active",
providerSpecificData: { baseUrl, apiType: "chat" },
});
return { providerId, model: `${label}/test-model`, apiKey };
}
function createServerProcess(dataDir: string, port: number) {
const stdoutLines: string[] = [];
const stderrLines: string[] = [];
let exitInfo: { code: number | null; signal: NodeJS.Signals | null } | null = null;
const child = spawn(process.execPath, ["scripts/dev/run-next-playwright.mjs", "dev"], {
cwd: REPO_ROOT,
env: {
...process.env,
DATA_DIR: dataDir,
PORT: String(port),
DASHBOARD_PORT: String(port),
API_PORT: String(port),
HOST: "127.0.0.1",
REQUIRE_API_KEY: "false",
API_KEY_SECRET: process.env.API_KEY_SECRET || "system-failover-secret-123456",
DISABLE_SQLITE_AUTO_BACKUP: "true",
INITIAL_PASSWORD: "",
NEXT_TELEMETRY_DISABLED: "1",
OMNIROUTE_DISABLE_BACKGROUND_SERVICES: "true",
OMNIROUTE_DISABLE_TOKEN_HEALTHCHECK: "true",
OMNIROUTE_DISABLE_LOCAL_HEALTHCHECK: "true",
OMNIROUTE_HIDE_HEALTHCHECK_LOGS: "true",
OMNIROUTE_E2E_BOOTSTRAP_MODE: "open",
},
stdio: ["ignore", "pipe", "pipe"],
});
child.once("exit", (code, signal) => {
exitInfo = { code, signal };
});
child.stdout.on("data", (chunk) => {
const lines = String(chunk).split(/\r?\n/).filter(Boolean);
stdoutLines.push(...lines);
if (stdoutLines.length > 200) stdoutLines.splice(0, stdoutLines.length - 200);
});
child.stderr.on("data", (chunk) => {
const lines = String(chunk).split(/\r?\n/).filter(Boolean);
stderrLines.push(...lines);
if (stderrLines.length > 200) stderrLines.splice(0, stderrLines.length - 200);
});
return {
child,
stdoutLines,
stderrLines,
baseUrl: `http://127.0.0.1:${port}`,
get exitInfo() {
return exitInfo;
},
};
}
async function waitForServer(
baseUrl: string,
logs: {
stdoutLines: string[];
stderrLines: string[];
exitInfo?: { code: number | null; signal: NodeJS.Signals | null } | null;
}
) {
const startedAt = Date.now();
let lastError = "";
while (Date.now() - startedAt < 120_000) {
if (logs.exitInfo) {
throw new Error(
[
`OmniRoute exited before it became ready (code=${logs.exitInfo.code}, signal=${logs.exitInfo.signal})`,
"--- stdout ---",
...logs.stdoutLines.slice(-40),
"--- stderr ---",
...logs.stderrLines.slice(-40),
].join("\n")
);
}
try {
const response = await fetch(`${baseUrl}/api/monitoring/health`, {
signal: AbortSignal.timeout(5_000),
});
if (response.ok) return;
lastError = `HTTP ${response.status}`;
} catch (error: any) {
lastError = error instanceof Error ? error.message : String(error);
}
await sleep(500);
}
throw new Error(
[
`Timed out waiting for OmniRoute to start: ${lastError}`,
"--- stdout ---",
...logs.stdoutLines.slice(-40),
"--- stderr ---",
...logs.stderrLines.slice(-40),
].join("\n")
);
}
async function stopProcess(child: ReturnType<typeof spawn>) {
if (child.killed) return;
child.kill("SIGTERM");
const exited = await Promise.race([
new Promise<boolean>((resolve) => child.once("exit", () => resolve(true))),
sleep(5_000).then(() => false),
]);
if (!exited && !child.killed) {
child.kill("SIGKILL");
await new Promise<void>((resolve) => child.once("exit", () => resolve()));
}
}
async function postChat(
baseUrl: string,
model: string,
content: string,
extraHeaders?: Record<string, string>
) {
const response = await fetch(`${baseUrl}/api/v1/chat/completions`, {
method: "POST",
headers: { "Content-Type": "application/json", ...extraHeaders },
body: JSON.stringify({
model,
stream: false,
messages: [{ role: "user", content }],
}),
signal: AbortSignal.timeout(30_000),
});
const text = await response.text();
const json = text ? JSON.parse(text) : {};
return { response, json };
}
async function resetBreakers(url: string) {
await fetch(`${url}/api/resilience/reset`, {
method: "POST",
signal: AbortSignal.timeout(5_000),
});
}
const serverA = new MockUpstreamServer();
const serverB = new MockUpstreamServer();
let app:
| {
child: ReturnType<typeof spawn>;
stdoutLines: string[];
stderrLines: string[];
baseUrl: string;
}
| undefined;
const TOKEN_A = "sk-sys-a";
const TOKEN_B = "sk-sys-b";
const TOKEN_A2 = "sk-sys-a2";
const TOKEN_B2 = "sk-sys-b2";
test.before(async () => {
const baseUrlA = await serverA.start();
const baseUrlB = await serverB.start();
serverA.configureToken(TOKEN_A, {
defaultResponse: buildCompletion("server A ok", { model: "sys-a/test-model" }),
});
serverA.configureToken(TOKEN_A2, {
defaultResponse: buildCompletion("server A2 ok", { model: "sys-a2/test-model" }),
});
serverB.configureToken(TOKEN_B, {
defaultResponse: buildCompletion("server B ok", { model: "sys-b/test-model" }),
});
serverB.configureToken(TOKEN_B2, {
defaultResponse: buildCompletion("server B2 ok", { model: "sys-b2/test-model" }),
});
const provA = await seedProvider("sys-a", TOKEN_A, baseUrlA);
const provB = await seedProvider("sys-b", TOKEN_B, baseUrlB);
const provA2 = await seedProvider("sys-a2", TOKEN_A2, baseUrlA);
const provB2 = await seedProvider("sys-b2", TOKEN_B2, baseUrlB);
await combosDb.createCombo({
name: "sys-priority",
strategy: "priority",
config: { maxRetries: 0, retryDelayMs: 0 },
models: [provA.model, provB.model],
});
await combosDb.createCombo({
name: "sys-priority-v2",
strategy: "priority",
config: { maxRetries: 0, retryDelayMs: 0 },
models: [provA2.model, provB2.model],
});
await combosDb.createCombo({
name: "sys-priority-fobr",
strategy: "priority",
config: { maxRetries: 0, retryDelayMs: 0, failoverBeforeRetry: true },
models: [provA.model, provB.model],
});
await combosDb.createCombo({
name: "sys-priority-setretry",
strategy: "priority",
config: {
maxRetries: 0,
retryDelayMs: 0,
failoverBeforeRetry: true,
maxSetRetries: 1,
setRetryDelayMs: 500,
},
models: [provA.model, provB.model],
});
await combosDb.createCombo({
name: "sys-same-server",
strategy: "priority",
config: { maxRetries: 0, retryDelayMs: 0 },
models: [provA.model, provA2.model],
});
await combosDb.createCombo({
name: "sys-same-server-fobr",
strategy: "priority",
config: { maxRetries: 0, retryDelayMs: 0, failoverBeforeRetry: true },
models: [provA.model, provA2.model],
});
await combosDb.createCombo({
name: "sys-single-provider",
strategy: "priority",
config: { maxRetries: 0, retryDelayMs: 0 },
models: ["sys-a/modelA", "sys-a/modelB"],
});
await combosDb.createCombo({
name: "sys-single-provider-fobr",
strategy: "priority",
config: { maxRetries: 0, retryDelayMs: 0, failoverBeforeRetry: true },
models: ["sys-a/modelA", "sys-a/modelB"],
});
await settingsDb.updateSettings({
resilienceSettings: {
requestQueue: {
autoEnableApiKeyProviders: true,
requestsPerMinute: 120,
minTimeBetweenRequestsMs: 0,
concurrentRequests: 4,
maxWaitMs: 2_000,
},
connectionCooldown: {
oauth: { baseCooldownMs: 500, useUpstreamRetryHints: true, maxBackoffSteps: 3 },
apikey: { baseCooldownMs: 200, useUpstreamRetryHints: false, maxBackoffSteps: 0 },
},
providerBreaker: {
oauth: { failureThreshold: 3, resetTimeoutMs: 2_000 },
apikey: { failureThreshold: 2, resetTimeoutMs: 1_500 },
},
waitForCooldown: {
enabled: false,
maxRetries: 0,
maxRetryWaitSec: 0,
},
},
requestRetry: 0,
maxRetryIntervalSec: 0,
requireLogin: false,
setupComplete: true,
});
core.closeDbInstance();
app = createServerProcess(TEST_DATA_DIR, DASHBOARD_PORT);
await waitForServer(app.baseUrl, app);
const warmup = await postChat(app.baseUrl, "sys-b/test-model", "warm up");
assert.equal(warmup.response.status, 200, JSON.stringify(warmup.json));
serverB.resetState(TOKEN_B);
});
test.after(async () => {
if (app) await stopProcess(app.child);
await serverA.stop();
await serverB.stop();
core.closeDbInstance();
await fsp.rm(TEST_DATA_DIR, { recursive: true, force: true });
});
test("primary healthy: request routes to Server A only", async () => {
assert.ok(app);
serverA.resetState(TOKEN_A);
serverB.resetState(TOKEN_B);
const result = await postChat(app.baseUrl, "sys-priority", "healthy primary");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.choices[0].message.content, "server A ok");
assert.equal(result.json.model, "sys-a/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 1);
assert.equal(serverB.getState(TOKEN_B).hits, 0);
});
test("500 Internal Server Error: combo falls back to Server B", async () => {
assert.ok(app);
serverA.resetState(TOKEN_A, [buildError(500, "Internal Server Error")]);
serverB.resetState(TOKEN_B);
const result = await postChat(app.baseUrl, "sys-priority", "500 fallback");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.choices[0].message.content, "server B ok");
assert.equal(result.json.model, "sys-b/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 1);
assert.equal(serverB.getState(TOKEN_B).hits, 1);
});
test("503 Service Unavailable: combo falls back to Server B", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(503, "Service Unavailable")]);
serverB.resetState(TOKEN_B);
const result = await postChat(app.baseUrl, "sys-priority", "503 fallback");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.choices[0].message.content, "server B ok");
assert.equal(result.json.model, "sys-b/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 1);
assert.equal(serverB.getState(TOKEN_B).hits, 1);
});
test("both servers fail (500): request returns a 5xx error to the client", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(500, "A is down")]);
serverB.resetState(TOKEN_B, [buildError(500, "B is down")]);
const result = await postChat(app.baseUrl, "sys-priority", "both down");
assert.ok(result.response.status >= 500, `expected 5xx, got ${result.response.status}`);
});
test("combo fallback to Server B survives sequential 503 failures from Server A", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(503, "transient blip"), buildError(503, "second blip")]);
serverB.resetState(TOKEN_B);
const first = await postChat(app.baseUrl, "sys-priority", "seq 503 attempt 1");
assert.equal(first.response.status, 200, JSON.stringify(first.json));
assert.equal(first.json.choices[0].message.content, "server B ok");
assert.equal(first.json.model, "sys-b/test-model");
// Wait for the 200ms apikey cooldown to expire so the second request also
// goes through the full A→B fallback path rather than skipping A entirely.
await sleep(250);
const second = await postChat(app.baseUrl, "sys-priority", "seq 503 attempt 2");
assert.equal(second.response.status, 200, JSON.stringify(second.json));
assert.equal(second.json.choices[0].message.content, "server B ok");
assert.equal(second.json.model, "sys-b/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 2);
assert.equal(serverB.getState(TOKEN_B).hits, 2);
});
test("429 with Retry-After and wait-for-cooldown: primary retries then falls back to B", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [
buildError(429, "rate limited, retry after 1s", { "Retry-After": "1" }),
]);
serverB.resetState(TOKEN_B);
const patchRes = await fetch(`${app.baseUrl}/api/resilience`, {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
connectionCooldown: {
apikey: { useUpstreamRetryHints: true, baseCooldownMs: 200 },
},
waitForCooldown: { enabled: true, maxRetries: 1, maxRetryWaitSec: 2 },
}),
signal: AbortSignal.timeout(10_000),
});
});
test("failoverBeforeRetry enabled: upstream error triggers immediate failover to next target", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(429, "rate limited")]);
serverB.resetState(TOKEN_B);
const result = await postChat(app.baseUrl, "sys-priority-fobr", "test failover before retry");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.choices[0].message.content, "server B ok");
assert.equal(result.json.model, "sys-b/test-model");
// With failoverBeforeRetry=true, A should be hit exactly ONCE (no intra-URL retry)
assert.equal(serverA.getState(TOKEN_A).hits, 1);
assert.equal(serverB.getState(TOKEN_B).hits, 1);
});
test("failoverBeforeRetry disabled: 429 triggers executor intra-URL retry, succeeds on retry", async () => {
assert.ok(app);
// Full resilience reset so A isn't blocked by residual breaker/cooldown state
const patchRes = await fetch(`${app.baseUrl}/api/resilience`, {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
connectionCooldown: {
apikey: { useUpstreamRetryHints: false, baseCooldownMs: 0, maxBackoffSteps: 0 },
oauth: { useUpstreamRetryHints: false, baseCooldownMs: 0, maxBackoffSteps: 0 },
},
waitForCooldown: { enabled: false, maxRetries: 0, maxRetryWaitSec: 0 },
}),
signal: AbortSignal.timeout(10_000),
});
assert.equal(patchRes.status, 200);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
// Wait for any residual cooldowns to expire
await sleep(300);
// One 429, then the default (200) on retry
serverA.resetState(TOKEN_A, [buildError(429, "rate limited")]);
serverB.resetState(TOKEN_B);
const result = await postChat(app.baseUrl, "sys-priority", "test failover before retry disabled");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
// With maxRetries=0 the combo does not retry A — it fails over to B immediately.
assert.equal(result.json.model, "sys-b/test-model");
// With failoverBeforeRetry=false, A should be hit TWICE (initial + 1 intra-URL retry).
// This contrasts with failoverBeforeRetry=true where A is hit exactly ONCE.
assert.equal(serverA.getState(TOKEN_A).hits, 2);
});
test("maxSetRetries: both A and B fail first pass, A 429 again, B 200 on retry", async () => {
assert.ok(app);
// Reset to a clean resilience slate: disable cooldowns, disable waitForCooldown,
// reset breakers, and clear connection rate_limited_until from previous tests.
const patchRes = await fetch(`${app.baseUrl}/api/resilience`, {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
connectionCooldown: {
apikey: { useUpstreamRetryHints: false, baseCooldownMs: 0, maxBackoffSteps: 0 },
oauth: { useUpstreamRetryHints: false, baseCooldownMs: 0, maxBackoffSteps: 0 },
},
waitForCooldown: { enabled: false, maxRetries: 0, maxRetryWaitSec: 0 },
}),
signal: AbortSignal.timeout(10_000),
});
assert.equal(patchRes.status, 200);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
// Set try 0: A 429, B 500 — both fail
// Set try 1: A 429, B 200 — B succeeds
serverA.resetState(TOKEN_A, [buildError(429, "rate limited"), buildError(429, "rate limited")]);
serverB.resetState(TOKEN_B, [
buildError(500, "server error"),
buildCompletion("server B ok on retry", { model: "sys-b/test-model" }),
]);
const result = await postChat(app.baseUrl, "sys-priority-setretry", "test max set retries", {
"x-internal-test": "combo-health-check",
});
// Set try 0: A 429, B 500 → both fail
// Set try 1: A 429, B 200 → B succeeds
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.choices[0].message.content, "server B ok on retry");
assert.equal(result.json.model, "sys-b/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 2);
assert.equal(serverB.getState(TOKEN_B).hits, 2);
// Restore defaults so other tests are not affected
await fetch(`${app.baseUrl}/api/resilience`, {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
connectionCooldown: {
apikey: { useUpstreamRetryHints: false, baseCooldownMs: 200, maxBackoffSteps: 0 },
oauth: { useUpstreamRetryHints: true, baseCooldownMs: 500, maxBackoffSteps: 3 },
},
}),
signal: AbortSignal.timeout(10_000),
});
});
test("same server failoverBeforeRetry disabled: first model 429 retried before trying second", async () => {
assert.ok(app);
// Full resilience reset — previous test restored defaults and may have left A cooldown
const patchRes = await fetch(`${app.baseUrl}/api/resilience`, {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
connectionCooldown: {
apikey: { useUpstreamRetryHints: false, baseCooldownMs: 0, maxBackoffSteps: 0 },
oauth: { useUpstreamRetryHints: false, baseCooldownMs: 0, maxBackoffSteps: 0 },
},
waitForCooldown: { enabled: false, maxRetries: 0, maxRetryWaitSec: 0 },
}),
signal: AbortSignal.timeout(10_000),
});
assert.equal(patchRes.status, 200);
await sleep(300);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(429, "rate limited")]);
serverA.resetState(TOKEN_A2);
const result = await postChat(
app.baseUrl,
"sys-same-server",
"test same server failover disabled"
);
assert.equal(result.response.status, 200, JSON.stringify(result.json));
// With maxRetries=0 the combo fails over to A2 rather than retrying A.
assert.equal(result.json.model, "sys-a2/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 2);
});
test("same server failoverBeforeRetry enabled: first model 429 skipped to second immediately", async () => {
assert.ok(app);
await sleep(300);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(429, "rate limited")]);
serverA.resetState(TOKEN_A2);
const result = await postChat(
app.baseUrl,
"sys-same-server-fobr",
"test same server failover enabled"
);
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.model, "sys-a2/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 1);
assert.equal(serverA.getState(TOKEN_A2).hits, 1);
});
test("single provider, modelA 500: combo fails over to modelB", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(500, "model A error")]);
const result = await postChat(app.baseUrl, "sys-single-provider", "test modelA 500");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.model, "sys-a/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 2);
});
test("single provider, modelA 503: combo fails over to modelB", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(503, "Service Unavailable")]);
const result = await postChat(app.baseUrl, "sys-single-provider", "test modelA 503");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.model, "sys-a/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 2);
});
test("single provider, modelA 429 with fobr: immediate failover to modelB", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(429, "rate limited")]);
const result = await postChat(app.baseUrl, "sys-single-provider-fobr", "test fobr");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.model, "sys-a/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 2);
});
test("single provider, modelA 500 with fobr: modelA retry", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(500, "Oops!")]);
const result = await postChat(app.baseUrl, "sys-single-provider-fobr", "test fobr");
assert.equal(result.response.status, 200, JSON.stringify(result.json));
assert.equal(result.json.model, "sys-a/test-model");
assert.equal(serverA.getState(TOKEN_A).hits, 2);
});
test("single provider, both models fail: request returns 5xx to client", async () => {
assert.ok(app);
await resetBreakers(app.baseUrl);
resetConnectionCooldowns();
serverA.resetState(TOKEN_A, [buildError(500, "modelA down"), buildError(500, "modelB down")]);
const result = await postChat(app.baseUrl, "sys-single-provider", "both down");
assert.ok(result.response.status >= 500, `expected 5xx, got ${result.response.status}`);
assert.equal(serverA.getState(TOKEN_A).hits, 2);
});