Files
OmniRoute/scripts/test/combo-live-vps.mjs
Diego Rodrigues de Sa e Souza 7b139fdb5e Release v3.8.38 (#5078)
* chore(release): open v3.8.38 development cycle

* fix(executors): strip client_metadata for cerebras and mistral (#4727)

Integrated into release/v3.8.38 (leva 5)

* fix(codebuddy): only send reasoning params when client requests reasoning (#5019)

Integrated into release/v3.8.38 (leva 5)

* fix(sse): keep streaming for forceStream providers when client requests JSON (#5021)

Integrated into release/v3.8.38 (leva 5)

* fix(sse): guard non-JSON SSE lines and duplicate [DONE] (#4937)

Integrated into release/v3.8.38 (leva 5)

* feat(blackbox): refresh provider model catalog (#4935)

Integrated into release/v3.8.38 (leva 5)

* fix(sse): dedupe case-variant Anthropic version/beta headers (#4846)

Integrated into release/v3.8.38 (leva 5)

* feat(sse): Kiro inline <thinking> stream splitter (#4911)

Integrated into release/v3.8.38 (leva 5)

* feat(cursor): parse Composer DeepSeek-style inline tool calls (#4912)

Integrated into release/v3.8.38 (leva 5)

* feat(proxy): auth-less host:port batch import (#4938)

Integrated into release/v3.8.38 (leva 5)

* fix(oauth): support Kiro IDC (organization) token import (#4944)

Integrated into release/v3.8.38 (leva 5)

* fix(translator): preserve cache_control for DashScope OpenAI-compat providers (port from 9router#2069) (#5013)

Integrated into release/v3.8.38 (leva 5)

* fix(tts): resolve Gemini TTS models from catalog (#4934)

Integrated into release/v3.8.38 (leva 5)

* fix(sse): don't cool down the connection on a self-inflicted upstream timeout (504) (#5064)

Integrated into release/v3.8.38 (leva 5)

* fix(sse): robust Anthropic /v1/messages streaming — real ping keepalive + client-disconnect guard (#5063)

Integrated into release/v3.8.38 (leva 5)

* feat(video): add Alibaba DashScope (wan2.7-t2v) provider (#5051)

Integrated into release/v3.8.38 (leva 5)

* fix: preserve model hidden flags (isHidden) across model sync (#5086)

Integrated into release/v3.8.38 (leva 5)

* fix(models): derive model discovery config from registry modelsUrl (#5087)

Integrated into release/v3.8.38 (leva 5)

* fix(compression): replace fileURLToPath(import.meta.url) with runtime anchors for standalone bundle (#5089)

Integrated into release/v3.8.38 (leva 5)

* feat(cc): add summarized thinking display toggle (#5055)

Integrated into release/v3.8.38 (leva 5)

* Harden selected API error responses (#5032)

Integrated into release/v3.8.38 (leva 5)

* chore(quality): rebaseline file-size for leva 5 PR batch drift

6 frozen files grew from merged leva-5 PRs (cursor #4912, kiro #4911,
videoGeneration #5051, default #4727, base #4846, chat #5064); all covered
by per-PR tests. See _rebaseline_2026_06_26_leva5 in the baseline.

* feat(compression): compression playground (Play + Compare tabs) in the studio (#5080)

Integrated into release/v3.8.38

* fix(combo): fail over on empty-content 502 instead of exhausting the provider (#5085) (#5104)

* fix(dashboard): surface detailed credential-validation error in add-connection modal (#5088) (#5106)

* feat(providers): allow local/private provider URLs by default with scoped metadata-safe guard (#5066) (#5107)

* fix(diagnostics): treat non-streaming Claude messages shape as valid output (#5108) (#5116)

* fix(db): translate pt-BR SQLite driver-fallback log lines to English (#5103) (#5115)

* fix(sse): repair release base-reds — malformed-response false positives + header casing + stale tests (#5117)

Repairs the release/v3.8.38 base-reds; unblocks #5078.

* chore(quality): rebaseline file-size for responseSanitizer (#5117) + AddApiKeyModal drift

* fix(translator): forward image tool_result blocks as image_url (#5100)

Base-reds fixed (#5117); image tool_result→image_url. Integrated into release/v3.8.38.

* fix(responses): default text.format for openai-compatible responses providers (#5101)

Base-reds fixed (#5117); default text.format + file-size rebaseline. Integrated into release/v3.8.38.

* feat(dashboard): expose Fusion judgeModel + fusionTuning in the combo editor (#5074)

Base-reds fixed (#5117); Fusion editor + file-size rebaseline. Integrated into release/v3.8.38.

* feat(quota): add opt-in Codex/Claude auto-ping keepalive (#5102)

Base-reds fixed (#5117); auto-ping keepalive + file-size rebaseline. Integrated into release/v3.8.38.

* test(release): relocate 2 orphan test files into the collected flat tests/unit dir (#5120)

Unblocks Lint (test-discovery) on #5078. Integrated into release/v3.8.38.

* fix(translator): preserve reasoning-replay reasoning_content + repair 3 release-green test reds (#5122)

Repairs 3 release-green test reds + test-masking; unblocks #5078.

* test(golden): redact live Node version from provider translate-path snapshot (#5125)

Final golden unblock for #5078.

* test(golden): redact OmniRoute app version from translate-path snapshot (#5126)

Coverage shard golden unblock for #5078.

* Ignore disconnect races during in-band stream error handling (#5007)

Integrated into release/v3.8.38

* Track final connection IDs in failover logs (#5016)

Integrated into release/v3.8.38

* fix(sse): convert Gemini body to OpenAI format in antigravity MITM handler (#4845)

Integrated into release/v3.8.38 (rebased on tip, CHANGELOG re-injected)

* feat(providers): add ZenMux Free session-cookie provider (#5105)

Integrated into release/v3.8.38 (rebased on tip, CHANGELOG re-injected)

* feat(dashboard): click-to-edit model alias in provider page (#5119)

Integrated into release/v3.8.38 (rebased on tip, i18n scope verified, CHANGELOG re-injected)

* feat(mcp): web-session robustness — cookie dedup (PR6) + browser-pool observability (PR7) (#3368) (#5121)

Integrated into release/v3.8.38 (rebased on tip; cookie-dedup branch extracted to findExistingCookieConnection helper → complexity-neutral; CHANGELOG added)

* fix(usage): dedupe request-usage logging and debounce stats (#4940)

Integrated into release/v3.8.38 (rebased on tip; DB-handle hang was stale-base artifact — resetDbInstance already closes the handle, test green 5/5; file-size drift consolidated at release; CHANGELOG re-injected)

* fix(dashboard): key model visibility toggle on canonical providerId (#5091)

Integrated into release/v3.8.38 (retargeted main→release; .tsx visibility-key test green 2/2)

* chore(deps): bump actions/cache from 5.0.5 to 6.0.0 (#5112)

Integrated into release/v3.8.38 (retargeted main→release; workflow-only actions/cache bump — unit failures were stale main base-reds)

* fix(streaming): harden long OpenAI-compatible SSE streams (#5124)

Integrated into release/v3.8.38 (rebased on tip; streamHandler conflict with #5007 disconnect-guard resolved — both coexist, stream-handler 22/22 green)

* feat: Add Grok Build (xAI) provider with OAuth import-token flow (#5020)

Integrated into release/v3.8.38 (rebased on tip; Hard Rule #11 fix — Grok public client_id now via resolvePublicCred(grok_id), 3 literals removed; grok-oauth 7/7 + check:public-creds green)

* feat(providers): add Factory (factory.ai) as a subscription gateway provider (#5065)

Integrated into release/v3.8.38 (rebased on tip; added factory registry test for PR Test Policy + fixed check:env-doc-sync phantom FACTORY_API_KEY; factory loads in PROVIDERS, no Zod issue — that flag was a false positive)

* chore(test): reconcile golden snapshot + apikey count for new providers

#5020 (grok-cli), #5065 (factory), #5105 (zenmux-free) added providers but did
not regenerate tests/snapshots/provider/translate-path.json (now +3 entries) nor
bump the APIKEY_PROVIDERS count (159->160 for the factory gateway). Test-only
reconciliation; no production change.

* fix(resilience): harden quota and model lockout edge cases (#5093)

Integrated into release/v3.8.38 (rebased on tip). TRUST-BUT-VERIFY: dropped the PR's 0dd7df641 'fix unit gates' commit which reverted #5122 reasoning-replay (preserveReasoningContent) + re-introduced #4849 O(n^2) growth, and restored 5 tests it had realigned. Kept only the 3 declared resilience fixes (quota cutoff guard, gemini MIME, model-lockout maxCooldownMs); 23/23 green.

* Hydrate quota cache and scope auto combo candidates (#5015)

Integrated into release/v3.8.38 (rebased on tip). Kept core quota-cache hydration + auto-combo candidate scoping + combos UI; dropped out-of-scope toolCloaking refactor (conflicted with #4813 stripEnumDescriptions — took tip) and the unrelated sse-auth test split. Added quota-cache-hydrate-5015 regression test (Rule #18); combo-account-allowlist 8/8 + hydration 2/2 green.

* chore(quality): reconcile complexity + file-size baselines for v3.8.38 owner-PR batch

complexity 1972->1978 (+6) and file-size providers.ts 1093->1107 / usageHistory.ts
934->983 — drift from the /review-prs merge batch (#4845/#5105/#5020/#4940/#5093/
#5015 + #5121 cookie-dedup helper extraction). check:complexity/check:file-size do
not run on the PR->release fast-path, so the branch accrued unmeasured; all legit
feature/fix growth, not regression. See per-key justifications in each baseline.

* fix(security): exact-host Anthropic baseUrl check (CodeQL js/incomplete-url-substring-sanitization #674) (#5130)

The anthropic-compatible Bearer-fallback gate decided whether a configured baseUrl
targeted the official api.anthropic.com host via a substring `.includes("api.anthropic.com")`.
A look-alike upstream such as `https://api.anthropic.com.evil.test` or
`https://evil.test/?x=api.anthropic.com` matched the substring and was wrongly treated as
official, suppressing the Bearer fallback meant for third-party gateways
(CodeQL #674, js/incomplete-url-substring-sanitization, high).

Replace the substring test with an exported `isOfficialAnthropicBaseUrl()` helper that
parses the URL and compares the hostname for exact equality. Empty baseUrl stays official;
scheme-less hosts are parsed with an assumed https://; an unparseable baseUrl falls back to
third-party (Bearer emitted) as the safer default. Behavior for legitimate official/third-party
baseUrls is unchanged.

Adds tests/unit/anthropic-official-baseurl-host.test.ts covering official, look-alike,
scheme-less, and unparseable inputs plus a static guard that the substring pattern is gone.

* fix(proxy): repair one-click Deno & Cloudflare relay deployments (#5128) (#5132)

* fix(services): embed WS proxy honours LIVE_WS_HOST; reject empty messages early (#5110) (#5133)

* fix(api): resolve /v1/models/{id} case-insensitively (#5082) (#5135)

* fix(providers): add MiniMax M3 & Nemotron 3 Ultra to Cline catalog (#3321) (#5136)

* fix(proxy): make SOCKS5 handshake timeout tunable via SOCKS_HANDSHAKE_TIMEOUT_MS (#5109) (#5137)

* feat(sidebar): add support for colored menu icons (#3812)

Integrated into release/v3.8.38 (recreated on tip — fork had unrelated history; added getSidebarIconAccent regression test, Rule #18). Clean 2-file UI feature.

* fix(providers): complete grok-cli OAuth wiring + zenmux-free web-session metadata

Base-red repair for #5020 (grok-cli) and #5105 (zenmux-free), surfaced by the
full CI on the release PR (#5078) — the PR->release fast-path does not run the
oauth-providers-config / web-session-credentials / provider-consistency gates.

- grok-cli: register in OAUTH_PROVIDERS (providers.ts canonical list, fixes
  check:provider-consistency), add OAUTH_PROVIDER_IDS.GROK_CLI + GROK_CLI_CONFIG
  in oauth constants (provider config now sourced there, not a local literal),
  align oauth-providers-config.test.ts (EXPECTED_PROVIDER_KEYS + config map).
- zenmux-free: declare its web-session credential requirement (full Cookie header)
  in WEB_SESSION_CREDENTIAL_REQUIREMENTS.

Local: oauth-providers-config 27/27, web-session-credentials 4/4, grok-cli-oauth
7/7, check:provider-consistency OK, +115 OAUTH_PROVIDERS tests green.

* Fix resilience settings page response mapping (#5139)

Integrated into release/v3.8.38. Thanks @rdself for the fix and the regression test.

* fix(kiro): retire claude-sonnet-4.5 from catalog + pin 400 model-unavailable test (#5140)

Extracted the real change from #5140 (the bot PR regenerated the entire
freeModelCatalog.data.ts + touched package-lock.json; only the targeted
edits are kept here):
- remove claude-sonnet-4.5 from the Kiro registry entry
- remove the matching kiro free-model catalog row
- pin Kiro's verbatim 400 "Invalid model..." to isModelUnavailableError

Closes #4484

* fix(sidebar): drop orphan `settings` accent color (typecheck:core red) (#5142)

SIDEBAR_ICON_ACCENTS is typed Partial<Record<HideableSidebarItemId, string>>,
but `settings` is not a hideable item id (only `settings-general`,
`settings-appearance`, … and `context-settings` exist; there is no item with
`id: "settings"`), so the accent was unreachable. It broke `typecheck:core`
on the release tip ("'settings' does not exist in type …", introduced by
#3812 colored menu icons). Removing the orphan key restores a clean
typecheck:core (rc=0).

* feat: salvage batch 2 — diagnostics null-guard (#5096) + observed quota reset windows (#5025) (#5141)

* fix(diagnostics): null-guard content blocks in detectMalformedNonStream

A null (or non-object) entry in a Claude-native `content` array made the
non-stream classifier throw `TypeError: Cannot read properties of null
(reading 'type')`, crashing the malformed-response detection path. Guard
before type-asserting each block: a null/non-object block is simply skipped.

Two regression tests added (null block among valid blocks → null; only-null
blocks → empty_choices).

Salvaged from closed PR #5096 (base-stale; only the defensive guard — the
Claude-shape recognition it also carried already landed via #5108).

Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>

* feat(quota): persist observed provider quota reset windows

Adds `provider_quota_reset_events` (migration 108) + `db/quotaResetEvents.ts`
to record real upstream weekly-quota window transitions whenever a quota
refresh shows the reset rolling to a new cycle (different day, later resetAt).
`apiKeyUsageLimits` now prefers the observed window start over the inferred
`resetAt − 7d`, falling back to snapshot inference when no event is recorded
yet. `quotaCache.setQuotaCache` records the transition opportunistically.

`recordProviderQuotaResetEventIfChanged` only fires for the primary weekly
window (not daily/sonnet), is idempotent (INSERT OR IGNORE on the unique
window key), and no-ops when the reset didn't actually roll. 4 unit tests
(tests/unit/lib/quota-reset-events.test.ts).

Salvaged from closed PR #5025 (which bundled this with two unrelated
features + a colliding migration 104). Renumbered to 108; module re-exported
from localDb (Rule #2).

Co-authored-by: Witroch4 <175152067+Witroch4@users.noreply.github.com>

---------

Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Witroch4 <175152067+Witroch4@users.noreply.github.com>

* docs(i18n): sync 3.8.38 CHANGELOG section to 41 mirrors (unblock docs-accuracy) (#5144)

The root CHANGELOG [3.8.38] section grew with this cycle's merged PRs, but the
docs/i18n/<lang>/CHANGELOG.md mirrors were not re-synced — drifting >25% in body
size and failing check:docs-sync (the "Docs accuracy" fast-gate step) for every
open PR against the release.

Ran scripts/release/sync-changelog-i18n.mjs 3.8.38 3.8.37 to copy the root
[3.8.38] section into all 41 mirrors. check:docs-all now passes (exit 0).

Sections are copied verbatim; the per-language translation pass runs at release
time via i18n:run — this only restores the size-sync the gate enforces.

* feat(compression): pure per-step fidelity checker (4 invariants, fail-open)

* feat(compression): fidelityGate config + rejected breakdown fields

* feat(compression): wire per-step fidelity gate into stacked pipeline (opt-in)

* feat(compression): preview route accepts fidelityGate flag (playground)

* feat(compression): playground fidelity-gate toggle + lane rejection display

* docs(compression): note fidelityGate advanced thresholds are intentionally API-omitted

* refactor(compression): extract fidelity-gate step helpers to shrink strategySelector (file-size gate)

bodyToText and gateAdvance moved to fidelityGateStep.ts; StackAccumulator exported.
strategySelector: 889->854 (-35). Residual +6 vs pre-Milestone-B frozen 848 is the
irreducible StackOptions.fidelityGate field + two stacked-loop dispatch reads + import.
Baseline updated to 854 with justification. No cycle introduced (import type only).
940 compression tests pass; typecheck clean.

* test(usage): wire usageHistoryDedup under unit runner brace-list (#5145)

Integrated into release/v3.8.38.

* feat: salvage batch from closed stale PRs (#5038, #5057, #5076) (#5138)

Integrated into release/v3.8.38.

* test(combo): deterministic routing-decision matrix for all 17 strategies (#5146)

Integrated into release/v3.8.38.

* feat(compression): fuzzy near-duplicate dedup (session-dedup 2nd pass + playground toggle) (#5143)

Integrated into release/v3.8.38.

* chore(quality): rebaseline file-size for sidebarVisibility.ts + chat.ts drift (#5147)

Mid-cycle drift on release/v3.8.38 from already-merged PRs that the fast-path
(PR->release skips check:file-size) let accumulate without a bump:

- src/shared/constants/sidebarVisibility.ts 1100->1198 (#3812 colored menu
  icons, per-item accent map; #5142 dropped one orphan, net still above frozen)
- src/sse/handlers/chat.ts 1560->1575 (#5064 self-inflicted-timeout cooldown
  skip + #5124 long OpenAI-compatible SSE hardening + #5110 embed-WS
  LIVE_WS_HOST honour / early empty-message reject)

Each covered by its own PR tests; structural shrink of chat.ts tracked in #3501.
Unblocks the Fast Quality Gates for PRs targeting release/v3.8.38.

* chore(release): finalize v3.8.38 CHANGELOG + cycle reconciliation

- Reconcile [3.8.38]: +18 bullets (compression fidelity-gate/fuzzy-dedup #5143,
  quota keepalive #5102, web-session robustness #5121, MiniMax/Nemotron #5136,
  model-visibility #5091, failover logs #5016, disconnect races #5007, sidebar
  orphan #5142, SRE playbooks salvage #5138, new Security #5130 + Maintenance roll-up)
- Credit salvaged-PR authors (@JxnLexn / @KooshaPari / @herjarsa / @Witroch4)
- Remove phantom bullet for CLOSED-not-merged #5092 (setup aggregator never landed)
- Fix isHidden bullet PR citation #4389 -> #5086 (@herjarsa)
- Back-fill forgotten v3.8.36 bullet: #5026 crypto.randomUUID ID-gen (@hamsa0x7)
- Sync 41 i18n CHANGELOG mirrors; README What's New -> v3.8.38
- Rebaseline cycle drift: eslint 3987->4002, cognitive 833->841, dead-exports
  345->346, cyclomatic 1978->1980 (file-size handled by #5147)

* fix(i18n): add missing English UI labels (#5153)

Integrated into release/v3.8.38

* Preserve non-stream reasoning fields for compatible clients (#5155)

Integrated into release/v3.8.38

* feat(compression): ionizer engine — lossy JSON-array sampling reversible via CCR (#5148)

Integrated into release/v3.8.38

* test(combo): gated live smoke for combo strategies (in-process + VPS HTTP) (#5151)

Integrated into release/v3.8.38

* test: refresh release expectations to match current code (#5150)

Integrated into release/v3.8.38 (test-only base-red alignment extracted from #5150)

---------

Co-authored-by: Éder Costa <eder.almeida.costa@gmail.com>
Co-authored-by: José Victor Ferreira <root@josevictor.me>
Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: fulorgnas <46461624+fulorgnas@users.noreply.github.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: Jan Leon <Jan.gaschler@gmail.com>
Co-authored-by: R. Beltran <rbeltran8000@gmail.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: Ramel Tecnologia - Rafa Martins <146174365+rafacpti23@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Witroch4 <175152067+Witroch4@users.noreply.github.com>
2026-06-27 09:07:12 -03:00

657 lines
25 KiB
JavaScript

/**
* scripts/test/combo-live-vps.mjs
*
* Phase-3 VPS HTTP scenario driver for OmniRoute combo routing.
* Exercises 6 strategies (priority / round-robin / weighted / cost-optimized /
* fusion / auto) against the live server at 192.168.0.15:20128.
*
* Usage:
* node scripts/test/combo-live-vps.mjs
* node scripts/test/combo-live-vps.mjs --only=round-robin
* node scripts/test/combo-live-vps.mjs --failover # runs all 7 base scenarios + real failover
*
* Safety rules:
* - Only creates/deletes __live_test__* combos
* - Always cleans up in `finally` blocks
* - Never stops services or touches other data
*
* Exit code: 0 if all scenarios PASS or SKIP; non-zero only on real FAIL.
* Task 8 (--failover) can be appended after the main() call at the bottom.
*/
import { chat, createCombo, deleteCombo, listHealthyProviders, nonce } from "./_vpsClient.mjs";
// ---------------------------------------------------------------------------
// Candidate list — broad to maximise coverage across volatile provider health.
// listHealthyProviders() probes each with a real chat() call.
// ---------------------------------------------------------------------------
const BROAD_CANDIDATES = [
"groq/llama-3.1-8b-instant",
"groq/llama-3.3-70b-versatile",
"minimax/MiniMax-M3",
"minimax/minimax-m3",
"kimi-coding-apikey/moonshot-v1-8k",
"openrouter/openai/gpt-3.5-turbo",
"cerebras/llama-3.3-70b",
"cerebras/llama3.1-8b",
"deepseek/deepseek-chat",
"ollama-cloud/glm-5.2",
"glm/glm-4-flash",
"gemini/gemini-2.0-flash",
];
// Known approximate input cost ($/M tokens) from OmniRoute's default-pricing constants.
// Used only to identify cheap vs pricey pairs for cost-optimized scenario.
// Models absent from this map are treated as unknown cost (Infinity in the server's
// sortModelsByCost, i.e. sorted last — effectively "most expensive").
const KNOWN_INPUT_COST = {
"groq/llama-3.1-8b-instant": 0, // inference-hosts.ts: price=0 (free tier)
"groq/llama-3.3-70b-versatile": 0, // inference-hosts.ts: price=0 (free tier)
"cerebras/llama3.1-8b": 0, // inference-hosts.ts: price=0
"cerebras/llama-3.3-70b": 0, // inference-hosts.ts: price=0
"deepseek/deepseek-chat": 0, // inference-hosts.ts: price=0
"minimax/MiniMax-M3": 0.5, // regional.ts: $0.5/M input
"minimax/minimax-m3": 0.5, // regional.ts: $0.5/M input
"kimi-coding-apikey/moonshot-v1-8k": 1, // not in pricing table → Infinity on server → treated pricey
};
// ---------------------------------------------------------------------------
// CLI args
// ---------------------------------------------------------------------------
const onlyArg = process.argv.find((a) => a.startsWith("--only="));
const onlyScenario = onlyArg ? onlyArg.slice(7) : null;
// --failover: opt-in flag that appends a real-failover scenario (broken primary → healthy fallback)
const failoverFlag = process.argv.includes("--failover");
// ---------------------------------------------------------------------------
// Result tracking
// ---------------------------------------------------------------------------
let exitCode = 0;
const summary = [];
function pass(name, detail = "") {
summary.push({ name, result: "PASS" });
console.log(`PASS [${name}]${detail ? ": " + detail : ""}`);
}
function skip(name, reason) {
summary.push({ name, result: "SKIP" });
console.log(`SKIP [${name}]: ${reason}`);
}
function fail(name, reason, err = null) {
summary.push({ name, result: "FAIL" });
console.error(`FAIL [${name}]: ${reason}`);
if (err) console.error(" caused by:", err?.message ?? String(err));
exitCode = 1;
}
async function runScenario(name, fn) {
if (onlyScenario && onlyScenario !== name) return;
try {
await fn();
} catch (err) {
fail(name, `unexpected error: ${err?.message ?? String(err)}`, err);
}
}
// ---------------------------------------------------------------------------
// Helper: split "provider/model" → { providerId, modelPart }
// For multi-segment paths like "openrouter/openai/gpt-3.5-turbo":
// providerId = "openrouter", modelPart = "openai/gpt-3.5-turbo"
// ---------------------------------------------------------------------------
function splitProviderModel(full) {
const idx = full.indexOf("/");
if (idx < 0) return { providerId: full, modelPart: full };
return { providerId: full.slice(0, idx), modelPart: full.slice(idx + 1) };
}
// ---------------------------------------------------------------------------
// STEP 0: Cache probe
// Verifies that a sqlite-inserted combo is immediately routable.
// getComboByName() in combos.ts does a direct DB read with no TTL cache,
// so the combo should be visible as soon as the SSH INSERT commits.
// If not (unexpected caching behaviour), we poll up to 12s before blocking.
// ---------------------------------------------------------------------------
async function step0CacheProbe(healthy) {
console.log("\n=== STEP 0: sqlite → routable cache probe ===");
const probe = healthy[0];
const probeName = "__live_test__probe";
let id;
let blocked = false;
try {
id = createCombo({ name: probeName, strategy: "priority", models: [probe] });
console.log(` created combo id=${id} model=${probe}`);
// Attempt immediate chat — expect instant visibility (getComboByName bypasses cache)
let r = await chat(probeName, { maxTokens: 4 });
if (r.status === 200 && r.text) {
console.log(
` chat() → ${r.status} model=${r.model} text="${r.text.slice(0, 40)}"`
);
console.log(" PROBE RESULT: immediately routable — getComboByName bypasses in-memory cache");
} else {
console.log(
` Immediate chat → status=${r.status} text=${r.text ?? "(empty)"}`
);
console.log(" Polling up to 12 s (TTL cache unexpectedly active)...");
let resolved = false;
for (let i = 0; i < 6; i++) {
await new Promise((res) => setTimeout(res, 2000));
r = await chat(probeName, { maxTokens: 4 });
if (r.status === 200 && r.text) {
console.log(` Resolved after ~${(i + 1) * 2}s: model=${r.model}`);
console.log(" PROBE RESULT: visible after TTL expiry");
resolved = true;
break;
}
}
if (!resolved) {
console.error(" PROBE RESULT: BLOCKER — combo not routable within 12s");
blocked = true;
}
}
} finally {
if (id) {
try {
deleteCombo(probeName);
console.log(` cleanup: ${probeName} deleted`);
} catch (e) {
console.error(` cleanup error: ${e?.message}`);
}
}
}
if (blocked) {
throw new Error("BLOCKER: sqlite-inserted combo not routable within 12s — cannot run scenarios");
}
}
// ---------------------------------------------------------------------------
// Scenario 1: priority
// Two healthy models. Single chat() call → 200 + non-empty text.
// ---------------------------------------------------------------------------
async function scenarioPriority(healthy) {
const name = "__live_test__priority";
if (healthy.length < 1) {
skip("priority", "no healthy providers");
return;
}
const models = healthy.slice(0, 2);
let id;
try {
id = createCombo({ name, strategy: "priority", models });
const r = await chat(name, { maxTokens: 16 });
if (r.status !== 200) {
fail("priority", `status=${r.status} raw=${JSON.stringify(r.raw)?.slice(0, 120)}`);
return;
}
if (!r.text) {
fail("priority", "response text is empty");
return;
}
pass("priority", `status=200 model=${r.model} text="${r.text.slice(0, 40)}"`);
} finally {
if (id) try { deleteCombo(name); } catch { /* best effort */ }
}
}
// ---------------------------------------------------------------------------
// Scenario 2: round-robin
// ≥2 healthy models. 5 calls with unique nonces → at least 2 distinct
// response.model values (each call must return 200).
// ---------------------------------------------------------------------------
async function scenarioRoundRobin(healthy) {
const name = "__live_test__round-robin";
if (healthy.length < 2) {
skip("round-robin", `need ≥2 healthy providers, found ${healthy.length}`);
return;
}
const models = healthy.slice(0, Math.min(3, healthy.length));
let id;
try {
id = createCombo({ name, strategy: "round-robin", models });
const served = new Set();
for (let i = 0; i < 5; i++) {
const r = await chat(name, { maxTokens: 16 });
if (r.status !== 200) {
fail("round-robin", `call ${i + 1} returned status=${r.status}`);
return;
}
served.add(r.model);
}
if (served.size < 2) {
fail(
"round-robin",
`only 1 distinct model across 5 calls: [${[...served].join(", ")}] — round-robin not distributing`
);
return;
}
pass("round-robin", `${served.size} distinct models across 5 calls: [${[...served].join(", ")}]`);
} finally {
if (id) try { deleteCombo(name); } catch { /* best effort */ }
}
}
// ---------------------------------------------------------------------------
// Scenario 3: weighted
// Two healthy models with equal weight (50/50). 8 calls → both appear at
// least once (loose statistical check — P(only 1 model in 8 calls) ≈ 0.8%).
// ---------------------------------------------------------------------------
async function scenarioWeighted(healthy) {
const name = "__live_test__weighted";
if (healthy.length < 2) {
skip("weighted", `need ≥2 healthy providers, found ${healthy.length}`);
return;
}
const [m1, m2] = healthy.slice(0, 2);
const { providerId: p1, modelPart: mp1 } = splitProviderModel(m1);
const { providerId: p2, modelPart: mp2 } = splitProviderModel(m2);
let id;
try {
id = createCombo({
name,
strategy: "weighted",
models: [
{ providerId: p1, model: mp1, weight: 50 },
{ providerId: p2, model: mp2, weight: 50 },
],
});
const tally = {};
for (let i = 0; i < 8; i++) {
const r = await chat(name, { maxTokens: 16 });
if (r.status !== 200) {
fail("weighted", `call ${i + 1} returned status=${r.status}`);
return;
}
tally[r.model] = (tally[r.model] ?? 0) + 1;
}
const distinct = Object.keys(tally);
if (distinct.length < 2) {
fail(
"weighted",
`only 1 distinct model across 8 calls: ${JSON.stringify(tally)} — weighted routing not distributing`
);
return;
}
pass("weighted", `8 calls distribution: ${JSON.stringify(tally)}`);
} finally {
if (id) try { deleteCombo(name); } catch { /* best effort */ }
}
}
// ---------------------------------------------------------------------------
// Scenario 4: cost-optimized
// Find a cheap+pricey healthy pair (based on KNOWN_INPUT_COST).
// Insert pricey first (position 0) so cost-optimized must reorder.
// Assert: the served model matches the cheap provider.
//
// OmniRoute's sortModelsByCost uses getPricingForModel which merges
// default-pricing constants. groq models have price=0; minimax M3 has $0.5.
// Cost-optimized sorts ascending → groq (0) before minimax (0.5).
// ---------------------------------------------------------------------------
async function scenarioCostOptimized(healthy) {
const name = "__live_test__cost-optimized";
// Partition healthy into cheap (price=0) and pricey (price>0 or unknown>0)
const cheap = healthy.filter(
(m) => KNOWN_INPUT_COST[m] !== undefined && KNOWN_INPUT_COST[m] === 0
);
const pricey = healthy.filter(
(m) => KNOWN_INPUT_COST[m] !== undefined && KNOWN_INPUT_COST[m] > 0
);
if (cheap.length === 0 || pricey.length === 0) {
skip(
"cost-optimized",
`no distinguishable cheap+pricey pair among healthy=[${healthy.join(", ")}]`
);
return;
}
const cheapModel = cheap[0];
const priceyModel = pricey[0];
let id;
try {
// Insert pricey first — cost-optimized should reorder to serve cheapModel first
id = createCombo({
name,
strategy: "cost-optimized",
models: [priceyModel, cheapModel],
});
const r = await chat(name, { maxTokens: 16 });
if (r.status !== 200) {
fail("cost-optimized", `status=${r.status}`);
return;
}
if (!r.text) {
fail("cost-optimized", "empty response text");
return;
}
// Verify the cheap model was served (response.model contains cheap provider's model name)
const cheapProvider = splitProviderModel(cheapModel).providerId;
// Both direct match and provider-substring match are accepted since OmniRoute
// returns the raw upstream model name (e.g. "llama-3.1-8b-instant" not "groq/...")
const cheapModelPart = splitProviderModel(cheapModel).modelPart.toLowerCase();
const servedModel = (r.model ?? "").toLowerCase();
const isChapModel =
servedModel === cheapModelPart ||
servedModel.includes(cheapModelPart) ||
servedModel === cheapModel.toLowerCase() ||
// fallback: check provider header if available (response.model could be bare name)
servedModel.includes(cheapProvider.toLowerCase());
if (!isChapModel) {
fail(
"cost-optimized",
`expected cheaper model (${cheapModel}, price=${KNOWN_INPUT_COST[cheapModel]}) but got ${r.model} — cost-optimized may not have reordered`
);
return;
}
pass(
"cost-optimized",
`cheaper model served: ${r.model} (cheap=${cheapModel}@$${KNOWN_INPUT_COST[cheapModel]}, pricey=${priceyModel}@$${KNOWN_INPUT_COST[priceyModel]})`
);
} finally {
if (id) try { deleteCombo(name); } catch { /* best effort */ }
}
}
// ---------------------------------------------------------------------------
// Scenario 5: fusion
// Panel of 2-3 healthy models fan out in parallel; judge synthesizes 1 answer.
// Cost guard: panel ≤3, max_tokens=16, one call.
// ---------------------------------------------------------------------------
async function scenarioFusion(healthy) {
const name = "__live_test__fusion";
if (healthy.length < 2) {
skip("fusion", `need ≥2 healthy providers, found ${healthy.length}`);
return;
}
// Panel: up to 3 distinct models
const panelModels = healthy.slice(0, Math.min(3, healthy.length));
// Judge: reuse first healthy model (cheap, already warmed up)
const judgeModel = panelModels[0];
let id;
try {
id = createCombo({
name,
strategy: "fusion",
models: panelModels,
config: {
judgeModel,
fusionTuning: { minPanel: 2 },
},
});
// Use a unique nonce in content to defeat semantic cache
const r = await chat(name, {
maxTokens: 16,
content: `hi ${nonce()} answer in one word`,
});
if (r.status !== 200) {
fail("fusion", `status=${r.status} raw=${JSON.stringify(r.raw)?.slice(0, 120)}`);
return;
}
if (!r.text) {
fail("fusion", "judge returned empty synthesized text");
return;
}
pass("fusion", `synthesized text="${r.text.slice(0, 60)}" served-model=${r.model}`);
} finally {
if (id) try { deleteCombo(name); } catch { /* best effort */ }
}
}
// ---------------------------------------------------------------------------
// Scenario 6: auto
// model="auto" → virtual auto-combo bypasses DB lookup entirely.
// Assert: status=200, non-empty text, response.model is a real model (not "auto").
// Also test "auto/fast" variant (skip if it 400s as unknown).
// ---------------------------------------------------------------------------
async function scenarioAuto() {
// 6a: bare "auto"
const r = await chat("auto", { maxTokens: 16 });
if (r.status !== 200) {
fail("auto", `status=${r.status} raw=${JSON.stringify(r.raw)?.slice(0, 120)}`);
return;
}
if (!r.text) {
fail("auto", "empty response text");
return;
}
const servedModel = r.model ?? "";
if (servedModel === "auto" || !servedModel) {
fail("auto", `response.model is still "auto" — pool was not resolved`);
return;
}
pass("auto", `status=200 resolved-model=${servedModel} text="${r.text.slice(0, 40)}"`);
// 6b: "auto/fast" variant
const r2 = await chat("auto/fast", { maxTokens: 16 });
if (r2.status === 400 || r2.status === 404) {
skip("auto/fast", `variant not recognised (${r2.status})`);
} else if (r2.status !== 200 || !r2.text) {
fail("auto/fast", `status=${r2.status} text=${r2.text ?? "(empty)"}`);
} else {
pass("auto/fast", `status=200 model=${r2.model} text="${r2.text.slice(0, 40)}"`);
}
}
// ---------------------------------------------------------------------------
// Scenario 7: failover (opt-in via --failover)
//
// Approach: CROSS-PROVIDER BOGUS-MODEL (deterministic, no SSH crypto required).
//
// Why cross-provider?
// A same-provider bogus-model fails silently: when target[0] (bogus) returns a
// 404, OmniRoute calls recordProviderCooldown(providerA, undefined) — a provider-
// wide key. The combo then pre-screens target[1] (same providerA, real model) and
// finds providerA in cooldown → skips it → combo returns the 404 from target[0].
// Using DIFFERENT providers avoids this: target[0] puts providerA in cooldown,
// target[1] on providerB is unaffected, providerB serves the healthy response.
//
// Mechanism:
// - Target[0]: <providerA>/__nonexistent_model_xyz__ → upstream 404 → providerA cooldown
// - Target[1]: <providerB>/<realModel> → not in cooldown → 200 served
// - config.maxRetries:0, retryDelayMs:0 → immediate fallover, no retry delay
// - Bogus model can never return 200 → any 200 proves fallover to the real target
//
// If only 1 distinct provider is healthy: SKIP (BROKEN-CONNECTION approach needed; see
// comment below for how to implement it using encrypted wrong API key via SSH sqlite).
//
// BROKEN-CONNECTION approach (for future reference or if cross-provider is unavailable):
// 1. SSH to read STORAGE_ENCRYPTION_KEY from /root/.omniroute/.env
// 2. Encrypt a wrong API key using scryptSync(key,"omniroute-field-encryption-v1",32)+AES-256-GCM
// 3. INSERT broken provider_connection row into provider_connections table via SSH sqlite
// 4. Combo: [{providerId:glm, model:glm/glm-4-flash, connectionId:brokenConnId}, realModel]
// 5. Broken conn → 401 → recordProviderCooldown("glm",brokenConnId) → key "glm:brokenConnId"
// 6. Target[1] on different provider → unaffected → 200
// 7. Finally: DELETE combo AND broken connection
// ---------------------------------------------------------------------------
async function scenarioFailover(healthy) {
const name = "__live_test__failover";
if (healthy.length < 1) {
skip("failover", "no healthy providers — cannot run failover scenario");
return;
}
// Group healthy providers by providerId to find two distinct providers
const byProvider = new Map();
for (const m of healthy) {
const { providerId } = splitProviderModel(m);
if (!byProvider.has(providerId)) byProvider.set(providerId, []);
byProvider.get(providerId).push(m);
}
const distinctProviders = [...byProvider.keys()];
if (distinctProviders.length < 2) {
// Cannot use cross-provider approach with only one provider.
// Same-provider bogus-model fails: 404 → recordProviderCooldown(provider, undefined) →
// provider-wide cooldown blocks target[1] on the same provider.
skip(
"failover",
`need ≥2 distinct healthy providers for cross-provider approach; ` +
`found only 1 [${distinctProviders.join(", ")}]. ` +
`Implement BROKEN-CONNECTION approach to run failover with a single provider.`
);
return;
}
// CROSS-PROVIDER BOGUS-MODEL:
// target[0] = <providerA>/__nonexistent_model_xyz__ (will get 404, puts providerA in cooldown)
// target[1] = <providerB>/<realHealthyModel> (different provider, not in cooldown)
const bogusProvider = distinctProviders[0];
const realModel = byProvider.get(distinctProviders[1])[0];
const bogusModel = `${bogusProvider}/__nonexistent_model_xyz__`;
const { modelPart: realModelPart } = splitProviderModel(realModel);
let id;
try {
id = createCombo({
name,
strategy: "priority",
models: [bogusModel, realModel],
config: { maxRetries: 0, retryDelayMs: 0 },
});
console.log(
` failover combo (CROSS-PROVIDER BOGUS-MODEL):\n` +
` [0] ${bogusModel} ← broken primary (will 404)\n` +
` [1] ${realModel} ← healthy fallback (different provider)\n` +
` strategy=priority, maxRetries=0`
);
const r = await chat(name, { maxTokens: 16 });
// Assertions:
// 1. status=200: the combo succeeded — ONLY possible if it fell over to target[1],
// because target[0] (bogus model) can NEVER return 200 from the upstream.
// 2. Non-empty text: real LLM content was returned (not an empty error body).
// 3. Served model is NOT the bogus one (belt-and-suspenders; 200 already proves it).
// 4. Served model matches the real healthy target (positive proof of which model served).
if (r.status !== 200) {
fail(
"failover",
`expected status=200 after fallover (broken ${bogusModel}${realModel}), ` +
`got status=${r.status}. ` +
`raw=${JSON.stringify(r.raw)?.slice(0, 200)}`
);
return;
}
if (!r.text) {
fail("failover", `status=200 but empty text — fallover may have served a no-content response`);
return;
}
const bogusModelPart = "__nonexistent_model_xyz__";
const servedModel = (r.model ?? "").toLowerCase();
// Negative proof: bogus model did NOT serve (should never happen, but guard anyway)
if (servedModel.includes(bogusModelPart.toLowerCase())) {
fail(
"failover",
`bogus model string "${bogusModelPart}" appears in served model field "${r.model}" — ` +
`impossible 200 from a non-existent model; something is wrong`
);
return;
}
// Positive proof: served model matches the real healthy target
const realModelLower = realModelPart.toLowerCase();
const servedMatchesReal =
servedModel === realModelLower ||
servedModel.includes(realModelLower) ||
servedModel === realModel.toLowerCase();
const proofNote = servedMatchesReal
? ""
: ` [NOTE: served="${r.model}" ≠ expected="${realModelPart}" ` +
`— upstream alias likely; 200+text from bogus-primary is the failover proof]`;
pass(
"failover",
`CROSS-PROVIDER BOGUS-MODEL: broken primary (${bogusModel}) → ` +
`fallover → served ${r.model} (real target: ${realModel}) ` +
`text="${r.text.slice(0, 40)}"${proofNote}`
);
} finally {
if (id) {
try {
deleteCombo(name);
console.log(` cleanup: ${name} deleted`);
} catch (e) {
console.error(` cleanup error: ${e?.message}`);
}
}
}
}
// ---------------------------------------------------------------------------
// Main
// ---------------------------------------------------------------------------
async function main() {
console.log("=== OmniRoute combo-live-vps scenario driver ===");
if (onlyScenario) console.log(`Filtering to scenario: ${onlyScenario}`);
console.log();
// --- Health probe ---
console.log("Probing healthy providers (broad candidate list)...");
const healthy = await listHealthyProviders(BROAD_CANDIDATES);
console.log(` healthy (${healthy.length}/${BROAD_CANDIDATES.length}): [${healthy.join(", ")}]`);
if (healthy.length === 0) {
console.error("FATAL: no healthy providers — cannot run any scenarios");
process.exit(1);
}
// --- STEP 0: cache probe ---
if (!onlyScenario) {
// Run STEP 0 unconditionally unless --only is passed (targeted run skips housekeeping)
try {
await step0CacheProbe(healthy);
} catch (err) {
console.error(`\nBLOCKER: ${err.message}`);
process.exit(1);
}
}
console.log("\n=== Scenarios ===\n");
await runScenario("priority", () => scenarioPriority(healthy));
await runScenario("round-robin", () => scenarioRoundRobin(healthy));
await runScenario("weighted", () => scenarioWeighted(healthy));
await runScenario("cost-optimized", () => scenarioCostOptimized(healthy));
await runScenario("fusion", () => scenarioFusion(healthy));
await runScenario("auto", () => scenarioAuto());
// --- Scenario 7: failover (opt-in) ---
// Only runs when --failover is passed. Uses a bogus-model primary to force a real
// combo failover to the healthy secondary, proving the priority fallover path works
// against the live server without stopping any service.
if (failoverFlag) {
console.log("\n=== Failover scenario (--failover) ===\n");
await runScenario("failover", () => scenarioFailover(healthy));
}
// --- Summary ---
console.log("\n=== Summary ===");
for (const { name, result } of summary) {
console.log(` ${result.padEnd(4)} [${name}]`);
}
const passed = summary.filter((s) => s.result === "PASS").length;
const skipped = summary.filter((s) => s.result === "SKIP").length;
const failed = summary.filter((s) => s.result === "FAIL").length;
console.log(`\n ${passed} PASS ${skipped} SKIP ${failed} FAIL`);
process.exit(exitCode);
}
main().catch((err) => {
console.error("FATAL:", err?.message ?? err);
process.exit(1);
});