handleComboChat/handleRoundRobinCombo tracked lastStatus (first-write-wins),
lastError (last-write-wins), and earliestRetryAfter (global MIN across all
targets) independently, so the final unavailableResponse() could surface a
status/message pair from two different failing targets and decorate a
config-class error (e.g. Antigravity's 422 missing_project_id, which carries
no retryAfter of its own) with an unrelated target's long reset window.
- lastStatus now overwrites on every failure (last-write-wins), matching
lastError, so status and message always come from the same target.
- the "(reset after ...)" decoration is only applied when the surfaced
status is itself rate-limit-class (429/503) — see the new
open-sse/services/combo/unavailableRetryGate.ts leaf module (both
combo.ts and chat.ts are already over their file-size baseline, so the
gate logic lives in a new module and combo.ts only wires it in).
Co-authored-by: ikelvingo <im.kelvinwong@gmail.com>
reverseModelsDevProviders() only matched MODELS_DEV_PROVIDER_MAP entries
by the canonical OmniRoute provider id, but the map's RHS for the OAuth
CLI providers (codex/claude) only lists their alias (cx/cc), never the
canonical id itself. Since the models.dev sync job writes
model_capabilities rows under openai/cx and anthropic/cc (never
codex/claude), and the auto-combo gate canonicalizes a codex/... or
claude/... target's provider to "codex"/"claude" before the lookup,
the synced capability row was unreachable for those two providers.
Also probe the provider's alias (via the already-imported
PROVIDER_ID_TO_ALIAS) when scanning the map, so a canonical id like
"codex"/"claude" still matches entries keyed only by their alias.
Co-authored-by: ikelvingo <im.kelvinwong@gmail.com>
getMitmStatus() hard-wired dnsConfigured to a single Antigravity hostname
regex regardless of which agent was being diagnosed, and the diagnose route
never accepted an agentId to check against. Add checkDNSEntryForAgent()
reusing resolveHostsForAgent()'s existing per-target host resolution, thread
an optional agentId through getMitmStatus(), and have the diagnose route
parse ?agentId= and pass it through. Callers that omit agentId keep the
legacy Antigravity-only behavior unchanged.
Co-authored-by: ikelvingo <im.kelvinwong@gmail.com>
bin/cli/sqlite.mjs::loadSqlite() had no fallback beyond better-sqlite3, unlike
the real server's driver cascade (src/lib/db/adapters/driverFactory.ts::tryOpenSync,
which tries bun:sqlite -> better-sqlite3 -> node:sqlite). On machines without a
working better-sqlite3 native binary, every `omniroute doctor` DB check reported
a false FAIL even when the actual server was healthy via its own driver cascade.
openSqliteDatabase() now falls back to tryOpenSync() when better-sqlite3 fails
to import, reusing the same already-tested cascade the real server uses instead
of re-deriving a second one.
Co-authored-by: ikelvingo <im.kelvinwong@gmail.com>
CN-region Moonshot/Kimi API keys (issued on the domestic
platform.kimi.com/moonshot.cn account) belong to a completely separate
keyspace than the international platform.kimi.ai/api.moonshot.ai
account, so OmniRoute's hard-coded international base URL rejects them
with a generic "Invalid API key" 401.
Neither "kimi" (legacy id) nor "moonshot" (current user-facing id) was
in CONFIGURABLE_BASE_URL_PROVIDERS, so the Add-connection modal never
rendered a base-URL field for them and there was no supported way to
point a new connection at api.moonshot.cn. The underlying
resolveBaseUrl()/buildUrl() primitives already honor a
providerSpecificData.baseUrl override generically (same mechanism used
by siliconflow, xiaomi-mimo, etc.) -- this only exposes that existing
affordance for kimi/moonshot, defaulting to the unchanged
international host so existing users see no behavior change.
Co-authored-by: ikelvingo <im.kelvinwong@gmail.com>
* test(sse): repair two base-red gates on release/v3.8.49
`release/v3.8.49` is red at its own HEAD (36f8fd10) — `Quality Gates` and
`Release-Green (continuous)` both fail — so every PR opened against it inherits
four failing checks regardless of content. Two of those causes had no owner:
#8480/#8481/#8482 cover the compression ladder, the antigravity catalog, and the
resilience/translator regressions respectively, none of them these.
1. `chatcore-client-usage-buffer.test.ts` — 5/5 failing.
#8331/#8356 (e8719783e) inserted an `options` parameter between
`clientResponseFormat` and `deps` on `applyClientUsageBuffer()`. The five call
sites here still passed `deps` in the fourth position, so the injected spies
landed in the `options` slot and the real implementations ran — every
`calls.*.length` assertion saw 0. The `Parameters<typeof
applyClientUsageBuffer>[3]` cast in `makeDeps()` masked the type error, which is
why it reached the branch. Call sites updated to `(resp, body, format, {}, deps)`
and the cast repointed to `[4]`.
Two cases added for the parameter that caused this, since nothing at this layer
exercised it: `preserveContextBudgetInVisibleUsage` re-folds `context_budget_*`
into the visible fields for the Claude-Code path, and the default path keeps the
real unbuffered #8331 numbers.
2. `claude-to-openai-think-close-5123.test.ts` — 4 unsuppressed
`@typescript-eslint/no-explicit-any` errors, failing `npm run lint`.
Its suppression entry allows `count: 2` but the file had grown to four `(chunk:
any)` callbacks. Rather than raise the frozen count, the cause is fixed:
`collectChunks()` returns `StreamChunk[]` instead of `unknown[]`, narrowing once
at the boundary so all four callbacks need no annotation. The suppression entry
is then stale and removed, which ratchets the file to zero.
Test-only plus one allowlist deletion; no production code. Verified on a clean
worktree of the base commit: `npm run lint` clean (was 4 errors),
`chatcore-client-usage-buffer` 7/7 (was 0/5), `claude-to-openai-think-close-5123`
3/3.
* chore(ci): rebaseline five inherited file-size overages on release/v3.8.49
`check:file-size` fails on release/v3.8.49 at its own HEAD (36f8fd10), which is
what turns `Fast Quality Gates` red for every PR against this base:
tokenHealthCheck.ts 832 -> 841
chat.ts 1865 -> 1866
auth.ts 2475 -> 2486
accountFallback.ts 1941 -> 1960
combo.ts 3630 -> 3642
None of these files is touched by this PR. The growth was inherited from already
merged PRs that did not bump their entries, so there is no offending branch left
to fix — the same situation the baseline already records under
`_rebaseline_2026_07_02_5798_release_green`, and the procedure it documents is to
raise the frozen values with a justification note.
Raised to the current base values only. The files stay frozen and cannot grow
further; an in-flight PR that adds lines to them bumps its own entry as usual
(#8482 touches accountFallback.ts and combo.ts and will need that).
* chore(ci): register 11 covering unit tests in stryker tap.testFiles
`check:mutation-test-coverage --strict` fails on release/v3.8.49 at its own HEAD:
11 unit tests that cover mutated modules are absent from `stryker.conf.json`
`tap.testFiles`, so their mutant kills do not count. The gate does not self-heal —
it names the exact files to add.
open-sse/services/accountFallback.ts + 8247-accountfallback-model-unhealthy,
8248-accountfallback-nvidia-degraded,
model-lockout-exact-cooldown-cap,
repro-antigravity-404-family-cooldown-hijack
src/sse/services/auth.ts + 7993-noauth-proxy-routing,
8200-perplexity-web-401-cooldown,
sse-auth-antigravity-credits
src/server/authz/routeGuard.ts + authz/route-guard-vnc-session-local-only
open-sse/utils/error.ts + error-sensitive-redaction
open-sse/utils/publicCreds.ts + adobe-firefly
src/shared/utils/circuitBreaker.ts + 8332-combo-vision-fallback
None is a file this PR touches, and the gate reports the identical 11 on a worktree
carrying none of these base-red fixes. Additions only (11 insertions, 0 deletions);
the array stays sorted. This widens what the mutation run accounts for rather than
relaxing anything.
* chore(ci): re-measure accountFallback.ts file-size cap against the current base tip
The entry frozen in this PR (1960) was the value at 36f8fd10; the base has
since advanced to 1cafd328c and the file is 1966 there, so check:file-size
would still have been red on the merge commit. Re-measured to 1966.
Same inherited drift the note already documents: check:file-size does not run
on the PR->release fast path, so growth accrues unmeasured between release
rebaselines. The other four entries still match the current tip
(tokenHealthCheck 841, chat 1866, auth 2486, combo frozen 3642 >= 3640).
---------
Co-authored-by: backryun <busan011@ormbiz.co.kr>
`normalizeExecutorResult()` has always accepted `Response | { response, url, headers,
transformedBody }` — the bare arm is what the web/scraping executors return from their
error and passthrough paths, and `chatcore-upstream-timeouts.test.ts` already covers
that both shapes are handled. But `BaseExecutor.execute` has no explicit return type,
so TypeScript inferred it from the method's single `return` — the object shape alone.
Every override returning a bare `Response` was therefore reported as incompatible:
* 14 × TS2739 in `duckduckgo-web.ts`, whose `execute()` additionally pinned its own
signature to just the object shape while returning `errorResponse()` /
`processResponse()` (both `Response`) from 14 valid paths
* TS2416 in `felo-web.ts` and `gitlab.ts`, which declare `Promise<Response>`
Fix the declaration rather than the call sites: export `ExecutorExecuteResult` from
`base.ts` — the same union `normalizeExecutorResult()` accepts — and annotate
`BaseExecutor.execute` with it. `duckduckgo-web.ts` then drops its over-narrow
annotation, matching BaseExecutor and the ~38 other executors that let the return type
be inferred.
Two subclasses read `.response` straight off `super.execute()` and now narrow first:
* `github.ts` — the existing `!result.response` guard already meant "bare Response,
nothing to materialize"; it is now expressed as `result instanceof Response`, which
is the same branch for every input (bare / object / nullish)
* `pollinations.ts` — reads the status through both arms for its pool bookkeeping
Wrapping DuckDuckGo's 14 returns would have been the wrong fix: the values are already
correct, and `normalizeExecutorResult()` produces exactly `{ response, url: "",
headers: {}, transformedBody: null }` for them.
Validation: full tsc error-set diff against the base config — 335 -> 319, **zero new
errors** (line-number-agnostic diff is empty; the two `duckduckgo-web.ts` TS2345s that
appear to move are the same two pre-existing errors renumbered by added comments, and
are left for a later slice). `typecheck:core` clean, `check:type-coverage` 92.17% ->
94.17%, and 49 of the 50 existing test files importing a touched executor pass —
`plan3-p0.test.ts` fails identically with and without this change (it reads the
developer's real ~/.omniroute DB rather than a test-scoped DATA_DIR).
The new test pins the runtime behavior of the narrowing so a later simplification
cannot quietly drop the bare-Response arm.
`gemini-business.ts` built its upstream fetch options with `combineAbortSignals(...)`,
which is defined nowhere in the repository. The module imports `mergeAbortSignals`
from `./base.ts` on line 31 and never used it — a rename that was only half applied.
Because the call sits inside the fetch options object literal, the ReferenceError was
thrown while *constructing* the arguments, before `fetch()` ran, and the surrounding
try/catch turned it into `makeErrorResult(502, "Gemini Business network error: ...")`.
So every Gemini Business request failed with what reads like an upstream outage. The
provider is registered and reachable (`open-sse/executors/index.ts`), so this affects
the whole provider, not an edge case.
`mergeAbortSignals(primary, secondary)` requires two real signals while
`ExecuteInput.signal` is `AbortSignal | null | undefined`, so the call is guarded and
falls back to the timeout alone — the same shape huggingchat, grok-web, claude-web,
and ninerouter already use.
Why it went unnoticed: this file is only type-checked by `open-sse/tsconfig.json`,
whose runs abort at `TS5101` (the deprecated `baseUrl`) before any file is checked,
and `typecheck:core` covers a curated 26-file allowlist that excludes every executor.
Removing that config error is #8473; this bug is what the first full run surfaced.
TDD: the two new tests fail on the parent commit — `execute()` never reaches the
stubbed `fetch` — and pass with the fix. They also cover the null-signal path, since
that is where an unguarded `mergeAbortSignals` would throw next.
First slice of the TypeScript 7 migration split requested on #7697: resolve the
type diagnostics under `open-sse/tsconfig.json` in the lowest-risk modules, with
no toolchain change. 12 diagnostics across 8 files, all outside the hot path —
`chatCore.ts` and `stream.ts` are deliberately left for a later, standalone slice.
Fixes, by cause:
* `Transformer.cancel` (progressTracker, sseHeartbeat, and stream.ts's existing
handler) — the WHATWG Streams standard defines `transformer.cancel(reason)` and
Node implements it (verified on v24: cancelling the readable side invokes it),
but `lib.dom.d.ts` still omits it from `Transformer`, so every such handler was
TS2353. These handlers clear the heartbeat/progress intervals when an SSE client
disconnects, so deleting them to satisfy the checker would leak a timer per
abandoned stream. The interface is patched in `open-sse/types.d.ts` instead.
* `earlyStreamKeepalive` — `SettledHandler` was discriminated by `ok: true | false`.
This workspace compiles with `strictNullChecks: false`, where a boolean-literal
discriminant narrows the positive branch but not the negative one, so reading
`.error` off the rejected arm did not type-check (the two `.response` reads
elsewhere in the file did, which is why only one site errored). Retagged with a
string discriminant, which narrows both branches under the same settings.
* `toolCallShim` / `openai-responses` — assigning back to a property declared
`unknown` resets the `typeof` narrowing, so the following comparison no longer
saw a number/array. Both now read through a local. The `Read` limit clamp is
behavior-identical: its two branches are mutually exclusive at READ_MAX_LIMIT 2000.
* `sanitizeToolResultId` — takes `unknown` but forwards to a `string` parameter; a
non-string id previously reached `.replace()` and threw. Coerced instead.
* `openaiHelper` — `opts = {}` inferred `{}`; typed as `FilterToOpenAIFormatOptions`.
* `cursorAgentProtobuf` — `Buffer.alloc(0)` infers `Buffer<ArrayBuffer>` under
@types/node 26 while the decoded field is `Buffer<ArrayBufferLike>`; the locals
now use bare `Buffer`, matching `requestMetadata` a few lines above.
Validation: 335 -> 321 diagnostics with zero new errors (full tsc error-set diff
against the base config). typecheck:core clean, lint clean, check:type-coverage
92.17% -> 94.17%. All 114 existing test files that import a touched module pass;
`plan3-p0.test.ts` fails identically with and without this change (it reads the
developer's real ~/.omniroute DB instead of a test-scoped DATA_DIR).
The new test covers the three behavioral surfaces rather than the refactors the
existing keepalive/heartbeat suites already hold: that `transformer.cancel()`
really fires and can clear an interval, the id coercion, and the limit-clamp bounds.
TypeScript 6.x raises TS5101 on `open-sse/tsconfig.json`: `baseUrl` is
deprecated and stops functioning in TypeScript 7.0. It was paired with
`ignoreDeprecations: "5.0"`, which no longer silences it under TS 6 (the
compiler now demands "6.0").
Remove `baseUrl: ".."` and rewrite the `paths` mappings relative to the
tsconfig's own directory, which is how TypeScript resolves them with no
baseUrl set:
"@/*" ./src/* -> ../src/*
"@omniroute/open-sse" ./open-sse -> ../open-sse
"@omniroute/open-sse/*" ./open-sse/* -> ../open-sse/*
`ignoreDeprecations` goes with it — baseUrl was the only deprecated option
it was suppressing.
Verified by diffing the full tsc error set against the previous config (the
base run used `ignoreDeprecations: "6.0"` so compilation proceeds past the
config error, which otherwise aborts type-checking and masks everything):
zero new errors, 28 fewer. All 28 were in `electron/*.js`, which
`baseUrl: ".."` had been dragging into the open-sse program via
root-relative resolution. Scoping the program back to open-sse also moves
`check:type-coverage` from 92.17% to 94.01%; the ratchet direction is up so
the gate passes, and the baseline is deliberately left alone because the
gain is a measurement-scope change rather than new typing work.
The guard test asserts no tsconfig reintroduces `baseUrl` or
`ignoreDeprecations`, and that every `paths` target still resolves to a real
directory — the second half is the part that matters, since dropping baseUrl
silently changes what those mappings point at.
The openai-responses translator tracked tool calls in its own funcCallIds/
funcNames/funcArgsBuf bookkeeping without ever writing to the shared
state.toolCalls Map that stream.ts's completion-log summary builder reads.
Every openai->openai-responses translated stream with a tool call was
persisted with finish_reason "stop" and no tool_calls, even though the
actual SSE events sent to the client were correct.
* fix(providers): prevent health sweep from expiring devin-cli local credentials
* fix(auth): drop devin-cli from supportsTokenRefresh explicit set (#8407)
Root cause: listing "devin-cli" as refresh-capable made tokenHealthCheck
force-expire local CLI connections that never have a refresh token.
Remove it from the explicit set (keep windsurf) and drop the health-check
provider hardcode — the existing supportsTokenRefresh=false guard is enough.
* clicking a provider card hitting back loses scroll position
* clean up code
* Rename some funcitons
* minor UX updates
* add null-guard in highlight()
* Add providerCardHandle tests
* add highlight tests. refactored code into separate utility functions
* Minor change to skip a firstElementChild call - use a ref to access the Link inside ProviderCard.
Mirror the search/route.ts pattern for /v1/web/fetch: skip rate-limited
stubs instead of letting them short-circuit auto-select, walk the
fixed-priority pool (fill-first) with a request-time fallback on
retryable/quota upstream statuses (429 always; 402/403 for
Firecrawl/Tavily/TinyFish quota-style tiers), and return a proper 429
(with Retry-After) when the whole pool is exhausted instead of a
generic 400. Explicit-provider requests never silently fall back.
* fix(api): stop the 2000-token safety buffer from inflating usage.prompt_tokens in the client response (#8331)
* fix(sse): scope #8331's usage-buffer fix around Claude-Code-compatible providers
The #8331 fix correctly stopped folding the 2000-token context-window safety
margin into client-visible prompt_tokens/input_tokens/total_tokens for normal
API metering clients. But it also silently changed the response shape for
Claude-Code-compatible providers, whose own context accounting reads the
buffered number straight out of usage — regressing
tests/unit/cc-compatible-provider.test.ts (expected 2007, got 7).
Fold the computed context_budget_* fields back into the visible usage fields
for that one path only (applyClientUsageBuffer's new
preserveContextBudgetInVisibleUsage option, gated on the existing
isClaudeCodeCompatible flag in chatCore.ts). Every other caller keeps the
real, unbuffered #8331 numbers.
Live incident: streamed assistant text was losing the spaces BETWEEN words
(e.g. "Bilden är en riktig JPEG nu" -> "Bildenärenriktig JPEG nu") on the
Responses-API and Claude streaming paths.
stripInternalReasoningPlaceholder() (#8081/#8162) is called on every
individual delta.content chunk, and unconditionally called .trim() even when
its sentinel ("(prior reasoning summary unavailable)") was never present in
that chunk. Tokenizers commonly emit sub-word tokens with a leading space as
part of the token (e.g. " en", " riktig") -- each such chunk got its only
whitespace character (the inter-word space) silently trimmed away before
being appended to the accumulated message, while the words themselves stayed
intact. Punctuation-only chunks were largely unaffected, matching what was
observed live.
Only trims when the sentinel is actually present -- preserves the original
#8081 intent (collapse a placeholder-only chunk to "") without touching the
overwhelming majority of chunks that never contain it.
Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
Live incident: memory.vec.upsert.fail {"error":"no such table: vec_memories"}
recurred repeatedly right after restarts, even though ensureReady() is called
immediately beforehand. ensureReady()'s signature-check-then-maybe-recreate
logic (resetForSignature does DROP TABLE IF EXISTS + CREATE VIRTUAL TABLE) is
not synchronized against a concurrent caller's upsertVector/deleteVector -- a
second in-flight memory write that independently decides (from a stale read
of memory_vec_meta) it also needs to reset the table can drop it out from
under another write's insert. Confirmed live: memory_vec_meta showed
vec_loaded=0 for the entire session across many restarts, then flipped to 1
mid-investigation once one attempt finally completed without interruption --
consistent with an intermittent race, not a permanently broken path (verified
the underlying sqlite-vec extension and CREATE VIRTUAL TABLE statement work
correctly in isolation, both on the host and inside the production container).
Rather than chase the exact interleaving (every underlying SQLite call is
synchronous via better-sqlite3, so the race window is narrow and did not
reproduce under simple Promise.all stress tests), makes the write path
resilient to arriving after the table was dropped: on a "no such table"
error, recreate vec_memories from the last-known-good memory_vec_meta
dimension and retry once.
Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
release/v3.8.49 already ships the Gemini 3.6 Flash catalog entries
(AGY_PUBLIC_MODELS, ANTIGRAVITY_PUBLIC_MODELS, MODEL_SPECS with
supportsThinking: false — Antigravity still rejects client-supplied
thinking params) via #8013. What was still missing: the `ag` pricing
rows in DEFAULT_PRICING_OAUTH, so getPricingForModel("ag", id)
returned null for the three tiers and cost/quota calculations
silently fell back to $0.
Pricing: $1.50 input / $7.50 output / $0.15 cached per MTok (Google's
2026-07-21 announcement), matching the existing 3.5-flash schedule
shape. Thinking tokens billed at output rate.
Extends the existing pricing-ag-flash-tiers.test.ts (RED-first: all
three tiers failed the "non-null pricing row" assertion before this
change) rather than re-adding the already-shipped catalog/modelSpecs
entries.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* fix(providers): reconcile Kimi K3 vision when attachment contradicts modalities
Synced models.dev rows for kimi-coding*/k3 can ship attachment=false while
modalities_input still lists image/video. Prefer the modality signal (and
normalize at sync + resolve) so supportsVision, attachment, and exposed
modalities agree.
Closes#8250
* fix(providers): keep Kimi K3 static fallback text-only (#8250)
The Kimi K3 vision reconciliation added supportsVision=true directly
to the kimi-coding registry entry for id "k3". That entry is the
static/stable fallback catalog used when discovered capabilities are
unavailable, and it must stay text-only per #4071 — the vision fix is
already applied correctly on the discovered path via
MODEL_SPECS["kimi-k3"] (aliases: ["k3"]) and
modelCapabilities.ts::resolveVisionCapability.
Restores the invariants guarded by
tests/unit/kimi-k2.7-code-registration.test.ts ("Kimi Code k3 fallback
leaves discovered capabilities unset") and
tests/unit/catalog-updates-v3829-kimi-qwen.test.ts ("kmca stable
fallback only carries documented static capabilities").
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* fix(translator): cap thinking budget on explicit budget_tokens path
* fix(translator): stop dropping thinkingConfig on cap-0 reasoning_effort path
The thinking-budget-cap guard added in this branch skipped thinkingConfig
entirely whenever a model's thinkingBudgetCap was 0 (e.g. gemini-3-flash),
including on the reasoning_effort/budgetMap path. That regressed the
pre-#6943 native-defaults contract (thinkingBudget 0 / includeThoughts
false must still be present) and crashed callers that read
`.thinkingConfig.thinkingBudget` unconditionally
(translator-openai-to-gemini-defaults.test.ts).
Also restore includeThoughts:true on the Claude-format explicit
thinking.budget_tokens path (openai-to-gemini.ts's Claude-format field and
claude-to-gemini.ts's native thinking field): budget_tokens:0 there is the
client's dynamic-thinking sentinel (#6813), not an off-switch, and must
stay true even after the new capping — the cap must only clamp positive
explicit values, never flip the zero sentinel's semantics.
Updates two tests this branch added that encoded the incorrect
"omit thinkingConfig / includeThoughts:false for the 0 sentinel" behavior,
to match the pre-existing, still-required contracts above.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* fix(providers): revert scope-creep flip of Gemini 3.5/3.6 Flash supportsThinking
The thinking-budget-cap fix accidentally expanded 5 shorthand modelSpecs
entries (gemini-3.5-flash, gemini-3.5-flash-low, gemini-3.6-flash-high/
medium/low) into explicit objects setting supportsThinking:true and
thinkingBudgetCap:24576. That flip was unrelated to the two proven test
regressions (translator-openai-to-gemini-defaults.test.ts and
claude-to-gemini-budget-tokens-zero-6813.test.ts, which only exercise
gemini-3-flash-preview, gemini-3.1-pro and gemini-2.5-pro) and reopens a
deliberately closed path from #8013: Antigravity still rejects
client-supplied thinking params for these Gemini 3.5/3.6 Flash tier ids,
so supportsThinking must stay false (inherited from
GEMINI_35_FLASH_MODEL_SPEC).
Reverted all 5 entries back to the release shorthand
`{ ...GEMINI_35_FLASH_MODEL_SPEC }`. gemini-3-flash, gemini-3.1-pro and
gemini-2.5-pro (the models the regression tests actually exercise) were
already correctly specced in the release baseline and are untouched.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* test: hermetic TLS mock for notion thread-session suite + drop duplicated usage-analytics file
Root cause A (#8159): sendNotionInferenceRequest() in
open-sse/executors/notion-web.ts was migrated from fetch() to
tlsFetchNotion() (open-sse/services/notionTlsClient.ts, native
tls-client-node binary) to get past Notion's Cloudflare TLS
fingerprinting. #8159 updated the mock in the sibling
tests/unit/executor-notion-web.test.ts (installNotionTlsMock, wired
through __setTlsFetchOverrideForTesting) but never touched
tests/unit/executor-notion-web-thread-sessions.test.ts (split out
earlier by #7900) — its 3 execute()-driven tests still mocked
globalThis.fetch, which tlsFetchNotion() never calls once the native
TLS client loads successfully. Confirmed live: all 3 tests hit real
https://app.notion.com with a fake cookie and got a real 401
(~8.1-8.5s each here; on a network with blocked/slow egress this
would instead hang up to the client's ~190s timeout+grace per test —
a CI-hang risk).
Fix: replicate installNotionTlsMock verbatim from the sibling file
into executor-notion-web-thread-sessions.test.ts so the 3 tests mock
the TLS override point instead of global fetch. Suite is now fully
hermetic — 8/8 pass, no network I/O, total runtime 31.0s -> 8.0s.
Root cause B (#7700): tests/unit/usage-analytics-route-extra.test.ts
was created as a byte-for-byte duplicate of 10 of the 22 tests in
tests/unit/usage-analytics-route.test.ts. #7300 later fixed a fixture
bug in the retention-window boundary test ("does not double-count raw
and aggregated rows") in the main file only — reading
getUserDatabaseSettings().retention.usageHistory live instead of a
hardcoded 30-day cutoff (default retention is 365 days) — leaving the
duplicate copy on the stale hardcoded value, which now fails (1 !== 2).
Fix: delete the duplicate file. All 10 of its test names exist
verbatim in the main file (verified with comm -12) and that file
passes 22/22:
- does not double-count raw and aggregated rows
- does not persist guessed API key attribution
- does not throw Unknown named parameter on short range (needsAggregated=false)
- does not throw Unknown named parameter with apiKey filter on long range
- groups renamed API key usage by stable ID
- includes activityMap for heatmap
- includes cost by API key
- omits global aggregates when filtering by API key
- returns 500 on database errors
- returns weeklyPattern for the costs dashboard
No coverage loss — same production code, same assertions, one fewer
redundant file.
Validation:
- RED executor-notion-web-thread-sessions.test.ts: 5 pass / 3 fail
(401 !== 200, real network hit), 31.0s
- RED usage-analytics-route-extra.test.ts: 9 pass / 1 fail
(1 !== 2), 18.7s
- GREEN executor-notion-web-thread-sessions.test.ts: 8/8 pass, 8.0s,
hermetic (no network)
- GREEN executor-notion-web.test.ts (sibling, untouched): 37/37
pass, byte-identical diff
- GREEN usage-analytics-route.test.ts (untouched): 22/22 pass,
byte-identical diff
- npx eslint on the changed file: clean
- npm run typecheck:core: clean (exit 0)
Refs #8159
Refs #7300
Refs #7700
* chore(quality): register usage-analytics-route-extra deletion in test-masking allowlist
check:test-masking hard-flags any deleted test file without a
_deletedWithReplacement entry. The deletion is legitimate (100% duplicate
suite, coverage retained verbatim in tests/unit/usage-analytics-route.test.ts)
-- same registration pattern as the video-dashscope entry.
Refs #7700
Refs #7300
* test: realign catalog snapshot tests to current deliberate catalog state
Six catalog/snapshot tests drifted behind deliberate catalog changes that
were already validated by newer sibling tests. No production code touched;
every change aligns a stale snapshot to behavior already validated by newer
sibling tests.
Root causes (all confirmed against the current code before editing):
- tests/unit/providers-constants-split.test.ts: APIKEY_PROVIDERS grew from
187 to 195 entries via #8077 (clova-studio/internlm/ant-ling, regional),
#8161 (sarvam/plamo → regional, writer → frontier-labs) and #8170
(typhoon → regional, inception → frontier-labs). Family counts verified
to sum to 195 (gateways 60, frontier-labs 24, inference-hosts 28,
enterprise-cloud 17, regional 40, specialty-media 26) with no duplicates.
Updated the two assertions and extended the changelog comment.
- tests/unit/qianfan-provider.test.ts: the expected Baidu Qianfan website
URL was the pre-#8128 wenxinworkshop path. #8128/#6271 moved it to
https://cloud.baidu.com/product-s/qianfan_home, already locked by the
sibling regression test tests/unit/baidu-qianfan-website-urls-6271.test.ts.
- tests/unit/t31-t33-t34-t38-model-specs.test.ts and
tests/unit/auto-combo-credentialed-model-pool.test.ts: the Antigravity
catalog refactor (#8013) retired gemini-3-pro-preview/claude-sonnet-5 and
renamed the Gemini 3.5 Flash tiers (low/medium/high ->
extra-low/low/gemini-3-flash-agent), confirmed against
ANTIGRAVITY_PUBLIC_MODELS and tests/unit/antigravity-retired-public-models.test.ts.
Swapped the retired IDs for currently-registered ones
(gemini-3.6-flash-high, claude-sonnet-4-6, gemini-3-flash-agent,
gemini-3.5-flash-low/extra-low) and moved the wildcard-exclusion prefix
test from the now-2-tier "gemini-3.5-*" group to "gemini-3.6-*", which has
3 real tiers today (same >=3 semantics, just pointed at a prefix that
still has 3 members).
- tests/unit/model-alias-seed.test.ts: getModelInfo("gemini-3.1-pro") now
canonicalizes through ALIAS_TO_PROVIDER_ID["agy"] = "antigravity" (#8050),
the same pattern already applied to opencode -> opencode-zen. Updated the
expected provider id.
- tests/unit/video-dashscope.test.ts (deleted, 216 lines): #8266
reorganized the Alibaba video catalog so the flat wan2.7-t2v id no longer
exists under the plain "alibaba" provider (only the dated
wan2.7-t2v-2026-06-12 does); the flat id now lives only under
"qwen-cloud". All 6 tests in the file failed because they built requests
against alibaba/wan2.7-t2v, which the new allowlist now rejects with 400
("unsupported alibaba video model") - verified directly against
VIDEO_PROVIDERS in open-sse/config/videoRegistry.ts. Coverage already
exists and was confirmed passing pre-deletion in
tests/unit/alibaba-video-media.test.ts (including an explicit
"Alibaba rejects video models outside its own allowlist" case for this
exact id) and tests/unit/qwen-cloud-video-media.test.ts (covers the same
id under qwen-cloud). Note: the deleted file's DashScope upstream
error-path assertions (401 missing credentials, 502 missing task_id, 502
FAILED status, 504 poll timeout) don't have a byte-for-byte equivalent in
the two replacement files, though the shared dashscopeHandler.ts code
path they exercise remains covered by several sibling *-media.test.ts
files for the happy path and local validation.
- tests/unit/authz/spawn-capable-prefixes-client-safe.test.ts: #7892 added
/api/vnc-session to the SPAWN_CAPABLE_PREFIXES deny-list (Hard Rules
#15/#17 hardening). Bumped the expected length 10 -> 11 and added the
entry to the test's named list for documentation.
Refs #8013, #8050, #8266, #7892, #8128
* chore(quality): allowlist the video-dashscope.test.ts deletion with its replacements
check:test-masking (pr-test-policy CI gate) requires a _deletedWithReplacement
entry for any deleted test file, even when the deletion is a verified-legitimate
supersession. Documents the same #8266 rationale from the prior commit in the
machine-checked allowlist so the deletion is not flagged as unexplained masking.
Refs #8266
Two independent "code is right, bookkeeping lagged" base-reds:
1. #8233 made open-sse/executors/muse-spark-web.ts import
sanitizeErrorMessage from utils/error.ts (a real Rule #12 fix), but
left its KNOWN_MISSING_ERROR_HELPER allowlist entry in
scripts/check/check-error-helper.mjs in place. The gate's own
stale-allowlist enforcement (assertNoStale) correctly flagged the now
-obsolete entry: `npm run check:error-helper` failed with "1 entrada(s)
obsoleta(s)", and tests/unit/check-error-helper.test.ts's "the shipped
allowlist freezes exactly the known current violators" test expected
an empty Set. Removed the entry (kept the assertNoStale machinery and
the general scope-header comments untouched).
2. #8064 added the "compression-exclusions" sidebar item right after
"compression-studio" in COMPRESSION_CONTEXT_GROUP (deliberate,
complete feature) but didn't update two order-snapshot tests written
before that item existed:
- tests/unit/sidebar-visibility.test.ts expected the "omni-proxy"
section's flattened id list to end the compression block at
"compression-studio".
- tests/unit/ui/sidebar-engine-items.test.ts asserted "Studio must be
last" in COMPRESSION_CONTEXT_GROUP.
Updated both to the real, intentional order: Settings -> Combos ->
engines -> Studio -> Exclusions (Studio now second-to-last,
Exclusions last).
Validation (red -> green):
- check:error-helper gate: red ("1 entrada(s) obsoleta(s)") -> green
("OK (898 files scanned, 0 known-missing frozen)")
- tests/unit/check-error-helper.test.ts: 31/32 -> 32/32
- tests/unit/sidebar-visibility.test.ts: 6/7 -> 7/7
- tests/unit/ui/sidebar-engine-items.test.ts: 13/14 -> 14/14
Refs #8233
Refs #8064
Root cause (two independent causes):
1. PR #8219 (commit 2a865aaaa7) added CacheSettingsTab.tsx with 12
t("settings.*") calls whose keys were never created anywhere, not even
in en.json (the source of truth). The same PR added only 3 sidebar/header
keys (settingsCache, settingsCacheSubtitle, settingsCacheDescription) to
en+es, without running `npm run i18n:sync-ui` to propagate to the other
41 locales.
2. tests/unit/i18n-missing-placeholder-fallback.test.ts had a "#7258 repro"
test asserting the real zh-TW.json still carried raw __MISSING__:
placeholders — a premise invalidated by #8024, which completed the
Traditional Chinese translation to 100%.
What changed:
- Added the 12 missing settings.* keys to en.json, mirroring the sibling
requestBodyLimit* family (placeholders {min}/{max}/{value} match the
component exactly).
- Added real, natural translations for all 15 CacheSettingsTab-related keys
(12 settings.* + 3 sidebar/header) to pt-BR.json, vi.json and es.json (es
already had the 3 sidebar/header keys).
- Ran `npm run i18n:sync-ui` (official tool, no locale hand-edited) to stub
the remaining 39 locales with __MISSING__:<english>. This also discovered
17 pre-existing unrelated missing keys (compression-exclusions settings,
8 new-provider onboarding descriptions) never synced since #8031, and
pruned 3 dead orphaned zh-TW-only keys (codexSessionAffinity{Title,Desc,
Ttl}, superseded by the generic sessionAffinity* keys since #7274,
confirmed unused anywhere in src/) — verified programmatically as
+32/-0/~0 changed per stub locale, +32/-3/~0 changed for zh-TW.
- Rewrote the "#7258 repro" test to use a synthetic fixture (same style as
the sibling deepMergeFallback fixtures in the same file) instead of
depending on zh-TW.json's real, evolving translation-completeness state.
Proves the same behavior: collectPlaceholderLeaves() detects a raw
__MISSING__: leaf before deepMergeFallback (the fix) is exercised.
Validation: all 4 previously-red files green (23/23 assertions). Broader
sweep of 271 i18n-adjacent unit tests unaffected. i18n:check-ui-coverage
(42/42 locales >=80%, 99.7-100%) and i18n:check-glossary both pass.
typecheck:core clean.
Refs #8219
Refs #8024