Files
OmniRoute/tests/unit/combo-scoring-inspector.test.ts
Diego Rodrigues de Sa e Souza d259d9fcba fix(ci): clear base-reds on release/v3.8.50 (round 3) (#10213)
* fix(ci): clear base-reds on release/v3.8.50 (round 3)

- CHANGELOG.md: restore the top [Unreleased] section dropped by the #10189
  reconcile (docs-sync gate: first section must be Unreleased)
- env-doc-sync: document CONDUCTOR_ORCHESTRATOR_TOKEN + CONDUCTOR_SPOKESPERSON_URL
  in .env.example/ENVIRONMENT.md; allowlist the CI-only GITHUB_STEP_SUMMARY and
  TS7_BASE_REF (ts7 ratchet signals); drop a stray merge artifact line
- providers: restore the audited chatanywhere metadata entry that base-reds
  round 2 dropped together with its duplicate — the provider was half-wired
  (registry+endpoint without APIKEY metadata), which is what the wave3 test
  catches; re-pin providers-constants-split at the measured 228
- docs counts: 338 -> 339 (today's +2 void-ai/helixmind, -1 Puter) via
  gen:provider-reference + README/AGENTS/llm.txt/package.json/diagrams/i18n mirrors
- file-size ratchet: annotated rebaseline for the two pre-existing drifts
  (ModelSelectModal 1138, gateways 1250) following the 2026-08-11 precedent

Refs #9985

* fix(ci): base-reds round 3b — stale sibling tests + mode-pack weight contract

- check-docs-counts-sync.test.ts: drop the imports/subtests of the four helpers
  #10196 removed from the gate script (readMcpFactsFromSource, listLocalizedDocs,
  makeRequiredCountsValidator, checkFreeTierInventory) — the new-API tests that
  #10196 added stay; the file now loads again under the node runner
- quota-connection-recovery.test.ts: convert from vitest APIs to node:test —
  the file lives in tests/unit/*.test.ts (node-runner glob) and the vitest
  runtime crashes when imported outside vitest, killing the whole shard entry
- modePacks.ts: re-normalize all six mode packs to sum 1.0 — #8940 added
  sessionAvailability: 0.05 to every pack without rebalancing (1.05 total);
  ratios preserved exactly (÷1.05), so post-normalizeScoringWeights behavior
  is unchanged; restores the declared sum-to-1.0 contract the 4235 test pins

Refs #9985

* fix(ci): base-reds round 3c — vitest siblings, weights default, secrets FP, mutation tap

- DistributeProxiesButton.test.tsx: wrap renders in NextIntlClientProvider —
  #9245 localized the component (useTranslations) and left the test without
  the intl context, failing all 14 cases
- scoring.ts: re-normalize DEFAULT_WEIGHTS to sum 1.0 (same #8940 class as the
  mode packs — sessionAvailability added without rebalancing; ratios preserved)
- .gitleaks.toml: generalize the kimi sponsor-banner localStorage-key allowlist
  to -v\d+ — #10200 bumped v1→v2 and the stale regex regressed the secrets
  ratchet with a false positive
- stryker.conf.json: register 6 covering unit tests in tap.testFiles (4 modules)
  so their mutant kills count — unblocks check:mutation-test-coverage --strict

Refs #9985

* fix(ci): base-reds round 3d — inspector factor gap, stale registry/gap tests, i18n key sync

- comboScoringInspector: add cacheAffinity/sessionAvailability/connectionDensity
  to FACTOR_KEYS + the factor-key type — calculateScore() weighs them but the
  breakdown omitted them, so the explained contributions never summed to the
  reported score (inspector bug, red on the pure tip)
- combo-scoring-inspector.test: make the explicit-weights override sum-neutral
  (±0.05 shift) so it stays valid for any DEFAULT_WEIGHTS values — the hardcoded
  override only summed to 1.0 against the pre-#8940 defaults, which is also why
  explicit weights silently fell back to 'default' on the tip
- unorouter-registry.test: align to the canonical .com host (api.unorouter.ai
  301-redirects there, verified live) and to wave4's live model discovery
  (passthrough, no static seed) — the .ai/auto-model expectations were stale
- check-migration-numbering.test: 147 left KNOWN_GAPS when
  147_api_keys_model_access_mode.sql landed — assert absent (same as 143)
- i18n: sync-ui pass — 35,914 missing UI keys stamped as __MISSING__ placeholders
  across 42 locales (mechanical; greens the pt-BR key-presence integrity test;
  coverage pct unchanged by design — translation is a separate workstream)

Refs #9985

* fix(ci): base-reds round 3e — 2 real defects + 14 stale sibling tests (waves A-E)

Real defects fixed:
- src/lib/db/apiKeys.ts: #9313's empty-allowlist early return bypassed the group
  permission check, silently disabling group deny rules (#8817) for every key
  without a per-key allowlist; fall-through restored, restricted+[] deny-all kept
- open-sse/utils/proxyFetch.ts: #10032 re-appended the raw transport error to the
  propagated message, reintroducing the proxy user:password leak #9837 closed;
  new redactProxyDetailsInMessage() keeps the reason, redacts URL/credentials
- .github/workflows/quality.yml: #10134 added the TS7 ratchet as a separate
  blocking step AFTER the aggregated gates — the exact #8542 masking mechanism;
  folded into the non-fail-fast loop (still blocking, still PR-only) ⚠️ CI edit,
  gate-strengthening — explicit owner sign-off requested on the PR
- src/i18n/messages/ko.json: 3 machine-mistranslation regressions caught by the
  #8244 glossary checker (장애인→비활성화됨, 양말5://→socks5://, 비클로드→Claude가 아닌)

Stale sibling tests aligned to deliberately-moved contracts (each cites its mover):
request-log-detail-layout + -stream (#9245 intl provider), repro-8542 pin update,
quality-rail-gate-membership (#10134 shape), agentSkills-routes 45→46 (#9058),
cloudflare-ai-catalog-8717 (#8804 supersedes #8808), executor-xai (#9994),
vision-bridge-claude-wire (#9463 minimax→openai), sse-auth forced-pin (#8893),
tls-proxy-context (strengthened leak guards), rate-limit-local-error-classification
(#9164/#9342), minimax-thinking-signature (#9463), codebuddy-cn (#9723 +1 test),
github-copilot-custom-model (#9050), providers-g4f-batch3 (#9584),
synced-capability-warmup (#9199, stricter), sidebar-tools-group (#8221),
oauth-modal-grok-cli-paste (#9245); agentSkills/catalog.ts comment 45→46;
file-size rebaseline for proxyFetch (+19, annotated)

Refs #9985

* fix(ci): base-reds round 3f — waves F-J: 9 more real defects + stale sibling sweep

Real production defects fixed (all red on the pure tip, each with its origin):
- routeGuard.ts: #8949 accidentally DELETED the /api/providers/[id]/login
  local-only pattern — the route spawns a browser, so the loopback gate for a
  process-spawning route was gone (Hard Rules #15/#17); restored (314 guard
  tests green)
- agentSkills generator: #9058's category dispatch gave the config category an
  empty body, wiping skills/config-codex-cli/SKILL.md at the #10131 sync;
  fixed + SKILL.md regenerated via the official generator
- imageRegistry: #9982 broke same-provider bare aliasing (antigravity preview
  id sent upstream unresolved); new resolveSameProviderBareAlias() keeps the
  fal cross-provider fix intact
- imageRegistry: #9982's prefix strip handed the bare nano-banana ids to fal-ai,
  violating the pinned 2026-07-31 operator decision (adobe-firefly owns them);
  fal entries made prefix-only (dispatch already re-prefixes)
- mediaGeneration/fal.ts: the missing-credential 401 guard was lost when #10198
  deleted the superseded falHandler — tests were hitting the live network
- bottleneckPatch/rateLimitManager: #9041's merge clobbered #9604, resurrecting
  the Bottleneck v2.19.5 heartbeat bug (reservoir never refills); patched the
  library defect at the root and re-aligned chat-rate-limit-body-lock to the
  working reservoir contract
- processSupervisor.mjs: #9761 regressed the Node spawn to bare "node" (the
  #9156 launchd bug) and dropped #9209's ipv4first args; both restored
- openai-responses/pureHelpers: #9423's Agent null-sentinel was unreachable on
  the schemaless JSON-string path; gate extended
- i18n en.json: #8222's regen reverted the #9976 unclosed-tag fix and #8559's
  combo-cooldown copy; #9038 shipped 40 t() calls with no messages (runtime
  MISSING_MESSAGE); all restored/added + official sync-ui stamps, and vi's
  zero-marker policy re-established via the sanctioned translation backend

Stale sibling tests aligned (movers cited inline): chat-helpers (#9447),
executor-antigravity (#9351), video-fal-grok (#9982), visionBridge (#9759),
web-session-credentials (#8974), production-build-module-integrity (positive
anchor added), agentSkills-generator/skillManifestsLint/skills-injection/
agentSkillTools-mcp/listCapabilities-a2a (#9058), memory-settings (#10010),
model-catalog-policy-invalidation (#8906), model-alias-seed (#9485),
reactive-context-compaction (#8949), combo-provider-wildcard (broken upsert
helper), oauth-google-loopback (43-locale resurrected-key removal)

Validation: 501/501 across the 47 touched test files; typecheck:core, lint,
file-size, docs-sync all green.

Refs #9985

* fix(ci): base-reds round 3g — wave K/L: 4 more real defects + stale alignments

Real defects:
- base/reasoningEffort.ts: the stale duplicate cherry-pick #9612 re-added the
  codex minimal→low rewrite that #9883 had deliberately removed (OMP minimal
  passthrough); block removed again
- cursorImages.ts: #9840 wired prepareCursorImageForWire (sharp re-encode,
  fail-closed) into the SHARED resolveCursorImages, breaking zai-web and
  conol-web image uploads (HTTP 400 'undecodable'); new prepareForWire opt-out,
  Cursor default path unchanged (8 cursor suites green)
- modelCapabilities/snapshot: catalog prepare still issued 323 per-model reads
  of model_context_overrides + max_input_tokens overrides, violating #9199's
  bulk-load contract; both now resolve from the snapshot single pass
- v1-models-discovery-conformance: re-pinned to the bounded 30s SWR window
  (#9199/#10198) — the old 'stale-first regardless of age' contract is gone

Stale tests aligned (movers cited inline): codex-tools-strict-default (#9828
redundant-oneOf strip), devin-providers (#9245 i18n), db-migrationrunner-
constants-split (147→151 renumber #8228), gitlab-duo-oauth-setup (#9245),
chatcore-extracted-modules (#9161 outbound-protocol keying)

compression-api CI failures were cascade artifacts of codex-tools-strict-default
failing in the same force-exit shard process — no own defect (171/171 local).

Refs #9985

* fix(test): compression-api — register both describes before the runner starts

The DATA_DIR setup + route/db top-level awaits sat BETWEEN the two describes;
under --test-force-exit (the CI unit-runner flag) the process exits once the
already-registered tests finish, so on slow CI machines the whole second
describe died as 'Promise resolution is still pending' — the recurring
CI-only shard-2 failure that never reproduced locally without the flag.
Moved to the top of the file; 10/10 under --test-force-exit locally.

Refs #9985

* fix(quality): freeze modelCapabilities.ts at 1006 (annotated) — snapshot routing growth

Refs #9985

* fix(quality): move the modelCapabilities freeze into the frozen map (nested schema)

Refs #9985

* fix(i18n): translate all 39,718 pending UI keys across 42 locales (owner-approved)

Mass-translated every __MISSING__ placeholder via the official i18n:sync-ui
--translate-markers pipeline (operator backend), restoring i18nUiCoverage to the
100 baseline (was 89.9 after the merge-storm UI landings + the 42 keys #9038
never shipped).

Post-pass repairs, all caught by the existing gates:
- glossary: retired renderings the machine reintroduced normalized again
  (提供商→提供者 zh-CN/zh-TW, 鏈接→連結, 文檔→文件, 調用→呼叫, 供應商→提供者,
  響應→回應, 不活躍→未啟用 zh-TW; 클로드→Claude, 옴니루트→OmniRoute ko);
  DATA_DIR forbidden rendering avoided via 数据文件夹 rephrase
- ICU integrity: 120 values with renamed/dropped {params} repaired (39
  positional renames, 81 reset to the en source — functional over fluent)

Validation: glossary/pt-BR/vi/deno-relay/settings-keys/value-drift/google-
loopback suites 76/76; placeholder diff en×42 locales = 0; worst-locale
coverage = 100.0%.

Refs #9985

---------

Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-13 00:02:25 -03:00

525 lines
18 KiB
TypeScript

import test from "node:test";
import assert from "node:assert/strict";
import fs from "node:fs";
import os from "node:os";
import path from "node:path";
import { makeManagementSessionRequest } from "../helpers/managementSession.ts";
const TEST_DATA_DIR = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-scoring-inspector-"));
const ORIGINAL_DATA_DIR = process.env.DATA_DIR;
const ORIGINAL_INITIAL_PASSWORD = process.env.INITIAL_PASSWORD;
const ORIGINAL_JWT_SECRET = process.env.JWT_SECRET;
process.env.DATA_DIR = TEST_DATA_DIR;
const core = await import("../../src/lib/db/core.ts");
const combosDb = await import("../../src/lib/db/combos.ts");
const providersDb = await import("../../src/lib/db/providers.ts");
const settingsDb = await import("../../src/lib/db/settings.ts");
const quotaSnapshotsDb = await import("../../src/lib/db/quotaSnapshots.ts");
const callLogs = await import("../../src/lib/usage/callLogs.ts");
const comboMetrics = await import("../../open-sse/services/comboMetrics.ts");
const inspector = await import("../../src/lib/usage/comboScoringInspector.ts");
const comboHealthDashboard = await import("../../src/lib/usage/comboHealthDashboard.ts");
const route = await import("../../src/app/api/usage/combo-scoring-inspector/route.ts");
const { normalizeComboStep } = await import("../../src/lib/combos/steps.ts");
const { lockModel, clearAllModelLockouts } =
await import("../../open-sse/services/accountFallback.ts");
const { resetAllCircuitBreakers } = await import("../../src/shared/utils/circuitBreaker.ts");
const { DEFAULT_WEIGHTS } = await import("../../open-sse/services/autoCombo/scoring.ts");
const { MODE_PACKS } = await import("../../open-sse/services/autoCombo/modePacks.ts");
async function resetStorage() {
comboMetrics.resetAllComboMetrics();
clearAllModelLockouts();
resetAllCircuitBreakers();
core.resetDbInstance();
fs.rmSync(TEST_DATA_DIR, { recursive: true, force: true });
fs.mkdirSync(TEST_DATA_DIR, { recursive: true });
}
async function enableManagementAuth() {
process.env.INITIAL_PASSWORD = "combo-scoring-password";
await settingsDb.updateSettings({ requireLogin: true, password: "" });
}
async function seedAutoCombo(comboOverrides: Record<string, unknown> = {}) {
const comboInput = {
name: "combo-scoring-auto",
strategy: "auto",
...comboOverrides,
models: [
{
kind: "model",
providerId: "openai",
model: "openai/gpt-4o-mini",
connectionId: "scoring-conn-fast",
label: "Fast healthy target",
},
{
kind: "model",
providerId: "anthropic",
model: "anthropic/claude-3-5-haiku",
connectionId: "scoring-conn-slow",
label: "Slow low quota target",
},
],
};
const combo = await combosDb.createCombo(comboInput);
const firstStep = normalizeComboStep(comboInput.models[0], {
comboName: comboInput.name,
index: 0,
});
const secondStep = normalizeComboStep(comboInput.models[1], {
comboName: comboInput.name,
index: 1,
});
for (let index = 0; index < 4; index += 1) {
await callLogs.saveCallLog({
id: `scoring-fast-${index}`,
timestamp: new Date(Date.now() - index * 60_000).toISOString(),
method: "POST",
path: "/v1/chat/completions",
status: 200,
model: "openai/gpt-4o-mini",
requestedModel: comboInput.name,
provider: "openai",
connectionId: "scoring-conn-fast",
duration: 100,
tokens: { prompt_tokens: 100, completion_tokens: 100 },
comboName: comboInput.name,
comboStepId: firstStep.id,
comboExecutionKey: firstStep.id,
});
}
for (let index = 0; index < 4; index += 1) {
await callLogs.saveCallLog({
id: `scoring-slow-${index}`,
timestamp: new Date(Date.now() - index * 60_000).toISOString(),
method: "POST",
path: "/v1/chat/completions",
status: index === 0 ? 503 : 200,
model: "anthropic/claude-3-5-haiku",
requestedModel: comboInput.name,
provider: "anthropic",
connectionId: "scoring-conn-slow",
duration: 1_000,
tokens: { prompt_tokens: 100, completion_tokens: 100 },
comboName: comboInput.name,
comboStepId: secondStep.id,
comboExecutionKey: secondStep.id,
});
}
quotaSnapshotsDb.saveQuotaSnapshot({
provider: "openai",
connection_id: "scoring-conn-fast",
window_key: "daily",
remaining_percentage: 90,
is_exhausted: 0,
next_reset_at: null,
window_duration_ms: 86_400_000,
raw_data: null,
});
quotaSnapshotsDb.saveQuotaSnapshot({
provider: "anthropic",
connection_id: "scoring-conn-slow",
window_key: "daily",
remaining_percentage: 20,
is_exhausted: 0,
next_reset_at: null,
window_duration_ms: 86_400_000,
raw_data: null,
});
return { combo, firstStep, secondStep };
}
test.beforeEach(async () => {
await resetStorage();
});
test.after(async () => {
await resetStorage();
fs.rmSync(TEST_DATA_DIR, { recursive: true, force: true });
if (ORIGINAL_DATA_DIR === undefined) delete process.env.DATA_DIR;
else process.env.DATA_DIR = ORIGINAL_DATA_DIR;
if (ORIGINAL_INITIAL_PASSWORD === undefined) delete process.env.INITIAL_PASSWORD;
else process.env.INITIAL_PASSWORD = ORIGINAL_INITIAL_PASSWORD;
if (ORIGINAL_JWT_SECRET === undefined) delete process.env.JWT_SECRET;
else process.env.JWT_SECRET = ORIGINAL_JWT_SECRET;
});
test("scoring inspector ranks targets and explains score contributions", async () => {
const { combo, firstStep } = await seedAutoCombo();
const response = await inspector.buildComboScoringInspectorResponse({
range: "24h",
horizon: "7d",
comboId: String(combo.id),
taskType: "coding",
});
assert.equal(response.method, "read_only_recompute");
assert.equal(response.combos.length, 1);
assert.equal(response.combos[0].strategy, "auto");
assert.equal(response.combos[0].weightSource, "default");
assert.equal(response.combos[0].modePack, null);
assert.equal(response.combos[0].targets.length, 2);
assert.equal(response.combos[0].selectedExecutionKey, firstStep.id);
assert.equal(response.combos[0].targets[0].executionKey, firstStep.id);
assert.ok(response.combos[0].targets[0].score > response.combos[0].targets[1].score);
const contributionSum = response.combos[0].targets[0].factors.reduce(
(sum, factor) => sum + factor.contribution,
0
);
assert.ok(Math.abs(contributionSum - response.combos[0].targets[0].score) < 0.02);
assert.ok(response.combos[0].targets[0].factors.some((factor) => factor.key === "quota"));
assert.ok(
response.combos[0].targets[0].factors.some((factor) => factor.source === "combo_health")
);
});
test("scoring inspector reports mode packs over explicit auto weights", async () => {
const combo = await combosDb.createCombo({
name: "combo-scoring-mode-pack",
strategy: "auto",
models: ["openai/gpt-4o-mini"],
config: {
auto: {
modePack: "ship-fast",
weights: { ...DEFAULT_WEIGHTS },
},
},
});
const response = await inspector.buildComboScoringInspectorResponse({
range: "24h",
horizon: "7d",
comboId: String(combo.id),
combos: [combo],
skipAutopilot: true,
});
assert.equal(response.combos.length, 1);
assert.equal(response.combos[0].weightSource, "mode_pack");
assert.equal(response.combos[0].modePack, "ship-fast");
assert.deepEqual(response.combos[0].weights, MODE_PACKS["ship-fast"]);
});
test("scoring inspector reports valid explicit auto weights", async () => {
// Shift 0.05 from costInv to latencyInv so the override stays sum-normalized no matter
// what the DEFAULT_WEIGHTS values are (validateWeights requires sum ≈ 1.0 — the previous
// hardcoded override only summed to 1.0 against the pre-#8940 default values).
const explicitWeights = {
...DEFAULT_WEIGHTS,
latencyInv: DEFAULT_WEIGHTS.latencyInv + 0.05,
costInv: DEFAULT_WEIGHTS.costInv - 0.05,
};
const combo = await combosDb.createCombo({
name: "combo-scoring-explicit-weights",
strategy: "auto",
models: ["openai/gpt-4o-mini"],
autoConfig: { weights: explicitWeights },
});
const response = await inspector.buildComboScoringInspectorResponse({
range: "24h",
horizon: "7d",
comboId: String(combo.id),
combos: [combo],
skipAutopilot: true,
});
assert.equal(response.combos.length, 1);
assert.equal(response.combos[0].weightSource, "explicit");
assert.equal(response.combos[0].modePack, null);
assert.deepEqual(response.combos[0].weights, explicitWeights);
});
test("scoring inspector marks non-auto combos as explanatory recompute", async () => {
const combo = await combosDb.createCombo({
name: "combo-scoring-priority",
strategy: "priority",
models: ["openai/gpt-4o-mini"],
});
const response = await inspector.buildComboScoringInspectorResponse({
range: "24h",
horizon: "7d",
comboId: String(combo.id),
});
assert.equal(response.combos.length, 1);
assert.equal(
response.combos[0].warnings.some((warning) => warning.includes("not auto")),
true
);
});
test("scoring inspector skipAutopilot avoids rebuilding autopilot report", async () => {
const options: Parameters<typeof inspector.buildComboScoringInspectorResponse>[0] = {
range: "24h",
horizon: "7d",
healthResponse: {
timeRange: "24h",
combos: [
{
comboId: "combo-skip-autopilot",
comboName: "combo-skip-autopilot",
strategy: "auto",
models: [],
cost: { totalUsd: 0, avgPerRequestUsd: 0, byModel: [] },
quotaHealth: { providers: [], worstRemainingPct: 0 },
usageSkew: { modelDistribution: [], giniCoefficient: 0 },
performance: { avgLatencyMs: 0, successRate: 0, totalRequests: 0 },
targetHealth: [],
},
],
},
forecastResponse: {
asOf: "2024-01-01T00:00:00.000Z",
timeRange: "24h",
horizon: "7d",
method: "linear_history",
combos: [
{
comboId: "combo-skip-autopilot",
comboName: "combo-skip-autopilot",
strategy: "auto",
targets: [],
history: {
requests: 0,
inputTokens: 0,
outputTokens: 0,
cacheReadTokens: 0,
cacheCreationTokens: 0,
reasoningTokens: 0,
totalTokens: 0,
costUsd: 0,
avgDailyCostUsd: 0,
daysWithTraffic: 0,
windowDays: 1,
},
forecast: {
projectedRequests: 0,
projectedTokens: 0,
projectedCostUsd: 0,
},
quotaRisk: {
level: "unknown",
projectedWorstRemainingPct: null,
timeToExhaustDays: null,
worstTargetExecutionKey: null,
},
confidence: "no_data",
dataQuality: {
pricingCoveragePct: 0,
quotaCoverage: "none",
notes: [],
},
},
],
},
skipAutopilot: true,
};
// #7087: this used to assert that `options.combos` is never even READ when
// health/forecast/skipAutopilot are all supplied (via a getter that throws on
// access) — but that was pinning the exact bug: resolveConfiguredCombos() must
// check `options.combos` FIRST, unconditionally, so a caller-supplied `combos`
// array (e.g. from buildComboHealthDashboardResponse()) is never silently
// discarded just because health/forecast/autopilot were already resolved too.
// This test's own `options` genuinely has no `combos` (undefined), which still
// exercises the "no combos supplied + already have health signals -> skip the
// getCombos() DB round-trip" branch — no getter/trap needed for that.
const response = await inspector.buildComboScoringInspectorResponse(options);
assert.equal(response.combos.length, 1);
assert.equal(
response.combos[0].warnings.includes("Combo has no inspectable execution targets."),
true
);
});
test("scoring inspector includes resilience skip reasons for cooldowns and model lockouts", async () => {
const cooldownConnection = (await providersDb.createProviderConnection({
provider: "openai",
authType: "apikey",
name: "cooldown account",
apiKey: "test-key",
rateLimitedUntil: new Date(Date.now() + 60_000).toISOString(),
testStatus: "unavailable",
lastErrorType: "rate_limit",
errorCode: 429,
})) as { id: string };
const lockedConnection = (await providersDb.createProviderConnection({
provider: "github",
authType: "apikey",
name: "locked account",
apiKey: "test-key",
})) as { id: string };
lockModel("github", lockedConnection.id, "github/gpt-4o", "rate_limited", 60_000);
const response = await inspector.buildComboScoringInspectorResponse({
range: "24h",
horizon: "7d",
skipAutopilot: true,
healthResponse: {
timeRange: "24h",
combos: [
{
comboId: "combo-resilience",
comboName: "combo-resilience",
strategy: "auto",
models: [],
cost: { totalUsd: 0, avgPerRequestUsd: 0, byModel: [] },
quotaHealth: { providers: [], worstRemainingPct: 0 },
usageSkew: { modelDistribution: [], giniCoefficient: 0 },
performance: { avgLatencyMs: 0, successRate: 0, totalRequests: 0 },
targetHealth: [
{
executionKey: "openai-cooldown",
stepId: "openai-cooldown",
model: "openai/gpt-4o-mini",
provider: "openai",
connectionId: cooldownConnection.id,
label: "Cooldown target",
requests: 0,
successRate: 0,
avgLatencyMs: 0,
lastStatus: null,
lastUsedAt: null,
quotaRemainingPct: null,
quotaIsExhausted: null,
quotaTrend: null,
quotaScope: "connection",
},
{
executionKey: "github-lockout",
stepId: "github-lockout",
model: "github/gpt-4o",
provider: "github",
connectionId: lockedConnection.id,
label: "Locked target",
requests: 0,
successRate: 0,
avgLatencyMs: 0,
lastStatus: null,
lastUsedAt: null,
quotaRemainingPct: null,
quotaIsExhausted: null,
quotaTrend: null,
quotaScope: "connection",
},
],
},
],
},
forecastResponse: {
asOf: "2026-05-22T00:00:00.000Z",
timeRange: "24h",
horizon: "7d",
method: "linear_history",
combos: [],
},
});
const targets = response.combos[0].targets;
const cooldownTarget = targets.find((target) => target.executionKey === "openai-cooldown");
const lockoutTarget = targets.find((target) => target.executionKey === "github-lockout");
assert.equal(cooldownTarget?.signals.resilience.targetState, "skipped");
assert.equal(
cooldownTarget?.signals.resilience.skipReasons.some(
(reason) => reason.code === "connection_cooldown" && reason.retryAfterMs! > 0
),
true
);
assert.equal(lockoutTarget?.signals.resilience.targetState, "skipped");
assert.equal(
lockoutTarget?.signals.resilience.skipReasons.some((reason) => reason.code === "model_lockout"),
true
);
});
test("scoring inspector route requires auth, validates query, and returns 404", async () => {
await enableManagementAuth();
const { combo } = await seedAutoCombo();
const unauthenticated = await route.GET(
new Request(`http://localhost/api/usage/combo-scoring-inspector?comboId=${combo.id}`)
);
assert.equal(unauthenticated.status, 401);
const invalid = await route.GET(
await makeManagementSessionRequest(
"http://localhost/api/usage/combo-scoring-inspector?range=bad"
)
);
assert.equal(invalid.status, 400);
const missing = await route.GET(
await makeManagementSessionRequest(
"http://localhost/api/usage/combo-scoring-inspector?comboId=11111111-1111-4111-8111-111111111111"
)
);
assert.equal(missing.status, 404);
const authenticated = await route.GET(
await makeManagementSessionRequest(
`http://localhost/api/usage/combo-scoring-inspector?range=24h&horizon=7d&taskType=coding&comboId=${combo.id}`
)
);
assert.equal(authenticated.status, 200);
const body = await authenticated.json();
assert.equal(body.method, "read_only_recompute");
assert.equal(body.combos.length, 1);
assert.equal(body.combos[0].targets.length, 2);
});
// #7087 follow-up: resolveConfiguredCombos() used to unconditionally return `[]`
// whenever healthResponse + forecastResponse + (skipAutopilot || autopilotReport)
// were ALL supplied, regardless of whether options.combos was also populated. That
// is exactly the call shape comboHealthDashboard.ts::buildComboHealthDashboardResponse
// always uses (it fetches combos once, then threads combos/healthResponse/
// forecastResponse/autopilotReport into buildComboScoringInspectorResponse
// together) -- so through the real dashboard integration, `combosById`/`combosByName`
// (which resolveInspectorWeights() reads to report a combo's real configured
// modePack/weights) always came back empty, and every combo's weightSource silently
// fell back to "default" no matter what was actually configured. This test drives
// the real buildComboHealthDashboardResponse() end-to-end (not
// buildComboScoringInspectorResponse() directly, which the PR's own tests already
// covered) with a combo configured for modePack "ship-fast".
test("scoring inspector reports the configured weight source through the dashboard integration", async () => {
const { combo } = await seedAutoCombo({
config: { auto: { modePack: "ship-fast" } },
});
const dashboard = await comboHealthDashboard.buildComboHealthDashboardResponse({
range: "24h",
horizon: "7d",
comboId: String(combo.id),
taskType: "coding",
combos: [combo],
});
assert.equal(dashboard.errors.scoring, undefined);
assert.ok(dashboard.scoring, "expected a scoring inspector response, not null");
assert.equal(dashboard.scoring?.combos.length, 1);
assert.equal(
dashboard.scoring?.combos[0].weightSource,
"mode_pack",
"dashboard's own combos should drive the reported weight source, not silently fall back to default"
);
assert.equal(dashboard.scoring?.combos[0].modePack, "ship-fast");
assert.deepEqual(dashboard.scoring?.combos[0].weights, MODE_PACKS["ship-fast"]);
});