* feat(kiro): register GPT-5.6 Sol/Terra/Luna model family
Kiro announced its first OpenAI-family models on 2026-07-14
(kiro.dev/changelog/models): GPT-5.6 Sol (flagship), Terra (balanced
mid-tier) and Luna (fastest/cheapest), all sharing a 272k context
window. Registers the three base model ids in the kiro provider
registry with contextLength/maxOutputTokens so getResolvedModelCapabilities()
resolves the real 272k window instead of falling back to the generic
default. OmniRoute derives the thinking/agentic synthetic variants and
per-account rate multipliers dynamically at discovery time
(open-sse/services/kiroModels.ts), so only the three base entries need
static registration here.
Co-authored-by: Edison42 <gn00742754@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/2596
* chore(changelog): fragment for #7209
* fix(kiro): add GPT-5.6 Sol/Terra/Luna pricing rows
The registry additions in this PR exposed three new Kiro model ids
without matching pricing rows, tripping the catalog invariant that
every Kiro registry model must resolve a non-zero pricing row
(tests/unit/catalog-updates-v3x.test.ts) — the models would have
billed at $0.00.
Reuses the shared GPT_5_6_{SOL,TERRA,LUNA}_PRICING tiers already used
by the codex and openai aliases.
---------
Co-authored-by: Edison42 <gn00742754@gmail.com>
* fix(cli): fast-path --version to skip full CLI bootstrap
`omniroute --version` ran the entire CLI bootstrap before printing the
version: the tsx/esm + polyfill imports, env-file loading, and
Commander's ~70-command registration (importing DB, providers, OAuth,
and other heavy modules). That took ~1.5s just to print a version
string.
Add isVersionFastPath() (bin/cli/utils/versionFastPath.mjs) and check it
at the very top of bin/omniroute.mjs, before any of that work runs. It
only trips for an unambiguous bare `--version`/`-V` invocation (no other
args), so it never changes behavior for real commands or for `--help`
(whose output is generated dynamically from every registered
subcommand, so it still needs full registration and is deliberately not
fast-pathed).
`--version` now returns in ~0.3s instead of ~1.5s locally.
Co-authored-by: Sutarto Jordan Chrisfivo <Jordannst@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2414
* chore(changelog): fragment for #7208
* fix(build): enforce bin/cli/utils/versionFastPath.mjs in the pack-artifact gate
bin/omniroute.mjs now imports ./cli/utils/versionFastPath.mjs on its boot path
(the --version fast-path). bin/cli/ is only an allowlist PREFIX, so the file
vanishing from the npm tarball would never fail the unexpected-paths check --
only PACK_ARTIFACT_REQUIRED_PATHS makes its absence loud (#7065 class).
Adds the required path and updates the hardcoded expectation in
pack-artifact-policy.test.ts, matching the existing data-dir.mjs /
storageKeyProvision.mjs entries. Fixes the red in
tests/unit/pack-artifact-entrypoint-closures.test.ts, which derives the
requirement from the entrypoint's own imports.
---------
Co-authored-by: Sutarto Jordan Chrisfivo <Jordannst@users.noreply.github.com>
* fix(translator): register openai response projection for gemini clients
The response-translator registry had an OpenAI -> Antigravity response
projection registered, but no OpenAI -> Gemini one. When a client request is
detected as Gemini format (body-shape match on `contents: [...]`, per
detectFormat()) and combo routing lands the request on an OpenAI-native
provider, translateResponse() fell through its hub-and-spoke path with no
`openai -> gemini` translator registered, so the raw OpenAI
`chat.completion.chunk` shape reached the client unchanged instead of the
shared Gemini `response.candidates[]` envelope.
Registers FORMATS.OPENAI -> FORMATS.GEMINI reusing the existing
openaiToAntigravityResponse projection — Gemini and Antigravity already
share the same wrapped `{ response: { candidates: [...] } }` envelope
elsewhere in the pipeline (see the unwrapGeminiChunk callers in
open-sse/utils/stream.ts, which treat FORMATS.GEMINI and FORMATS.ANTIGRAVITY
identically), so no new conversion logic is introduced.
Co-authored-by: W ARELIK <warelik@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2399
* chore(changelog): fragment for #7207
---------
Co-authored-by: W ARELIK <warelik@users.noreply.github.com>
* fix(translator): preserve Gemini thought parts as reasoning_content on the OpenAI request bridge
Gemini thinking-mode output marks internal reasoning with `part.thought === true`
inside a content's `parts` array. geminiToOpenAIRequest() ran every part (thought
or not) through the same text-part branch, so a thought part was merged straight
into the message's visible `content` — leaking private reasoning into whatever the
OpenAI pivot forwarded downstream, and hiding it from Reasoning Replay Cache (which
only ever inspects `reasoning_content`).
Add convertGeminiContentWithReasoning(): split out `thought: true` parts before
delegating to the existing convertGeminiContent(), then re-attach the joined
thought text as `reasoning_content` on the resulting message (skipping tool/
functionResponse messages, whose schema has no such field). Non-strict-provider
stripping and reasoning-replay injection in translator/index.ts are untouched —
this only fixes what reasoning_content gets populated with on this one inbound
hop.
Co-authored-by: W ARELIK <warelik@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2401
* chore(changelog): fragment for #7206
---------
Co-authored-by: W ARELIK <warelik@users.noreply.github.com>
tryAwsSsoCache() only resolved clientId/clientSecret via data.clientIdHash -> <hash>.json. Newer kiro-auth-token.json files instead carry a top-level clientId directly, so that lookup silently failed and left clientId/clientSecret null, sending the dashboard's Import Token POST down the non-IDC path. That path (KiroService.validateImportToken -> readCachedClientCredentials) picked a client registration by region + latest-expiry across ALL cached SSO client registrations, ignoring the token's actual clientId — on a machine with multiple stale registrations this returned a mismatched clientId/clientSecret pair, producing 'Bad credentials' on refresh.
Fix: resolve clientId/clientSecret by scanning the cache for a registration file whose own clientId matches the token's clientId (falling back to clientIdHash first, then a direct-match scan), and thread an optional clientId hint into readCachedClientCredentials()/validateImportToken() so an exact match always wins over the region/latest-expiry heuristic.
Reported-by: Asher (@XCrag) (https://github.com/decolua/9router/issues/1253)
* fix(providers): add MiniMax image-generation provider (port from 9router#2482)
MiniMax already had entries in the music/audio/video registries, but no entry at all in imageRegistry.ts and no dedicated provider handler under open-sse/handlers/imageGeneration/providers/. A MiniMax image-model request therefore fell through the format dispatch in imageGeneration.ts to a 404/unmatched-format response instead of reaching MiniMax's synchronous image_generation endpoint.
Registers a minimax image provider (format: minimax-image, models image-01/image-01-live) and a new handleMinimaxImageGeneration handler that POSTs to https://api.minimax.io/v1/image_generation and normalizes data.image_urls into the OpenAI-compatible images payload.
Reported-by: felipeleite (https://github.com/decolua/9router/issues/2482)
* refactor(providers): split KIE image catalog out of imageRegistry to respect file-size cap
imageRegistry.ts hit 805 lines after adding the MiniMax image provider (cap is
800). Extract the KIE image-model catalog (largest single provider entry,
~35 models) into its own semantic-family module,
providers/registry/kie/imageModels.ts, following the same pattern already
used for LMARENA_DIRECT_IMAGE_MODELS. imageRegistry.ts now imports
KIE_IMAGE_MODELS instead of inlining the list.
Also update minimax-media-servicekinds.test.ts: getRegistryMediaKinds derives
membership by design from every registry in MEDIA_KIND_REGISTRIES, including
IMAGE_PROVIDERS. Now that minimax is a key in IMAGE_PROVIDERS, it correctly
gains the "image" kind alongside tts/video/music — the same behavior already
asserted for openai in this file. The exact-match assertion is updated to
["image","music","tts","video"]; the other assertions (which only check
.includes for tts/video/music/llm) were already correct and untouched.
* fix(providers): extract minimax image-gen helpers to fix complexity ratchet
check:complexity-ratchets regressed 2056 -> 2058 (handleMinimaxImageGeneration:
complexity 25, max-lines-per-function 97). Split logging, upstream-error,
no-images, success and fetch-error branches into small named helpers so the
handler stays within the cyclomatic-complexity (15) and max-lines-per-function
(80) ratchets. No behavior change; existing minimax-image-provider-2482 and
minimax-media-servicekinds unit tests still pass.
* feat(providers): curated OpenRouter embeddings catalog + specialty merge in live discovery (#6976)
OpenRouter serves embeddings via a dedicated OpenAI-compatible
/api/v1/embeddings endpoint that is omitted from /v1/models, and the
embeddingRegistry entry for it was stale (3 legacy ids). Meanwhile
providerModelsConfig gives openrouter a live discovery config, so
buildApiDiscoveryResponse's success path returned only the live chat
catalog verbatim — the specialty (embeddings/rerank) static catalog was
only ever merged in on the no-config local_catalog fallback, so OpenRouter
embeddings never surfaced through model discovery.
Refreshed the curated openrouter embeddingRegistry lineup (ids verified
against https://openrouter.ai/docs/api/reference/embeddings and the
collections page) and added a scoped, additive merge
(mergeSpecialtyCatalogIntoLiveModels, allowlisted to openrouter) that folds
embeddings/rerank entries from getStaticModelsForProvider() into the live
discovery response, deduped by id. Scoped as an allowlist rather than a
blanket merge because some providers (e.g. Gemini) already return
embedding models directly from their live /v1/models endpoint, where a
blind merge would risk stale/duplicate entries.
* test(providers): type the models discovery payload instead of any (#6976)
no-explicit-any is an error under tests/ (#6218), so the 4 `any`
usages in the new discovery assertions failed the max-warnings-0
lint gate. Replace them with an explicit ModelsResponseBody shape —
type-only change, all 13 assertions unchanged and still passing.
* test(providers): type the openrouter merge assertion callback (#6976)
The new #6976 assertion added a 56th explicit `any` to this file,
one over the 55 frozen in config/quality/eslint-suppressions.json,
tripping the max-warnings-0 lint gate. Type the callback param
instead of raising the frozen count — the debt ratchet only decreases.
All 59 tests still pass.
* fix(dashboard): hide disabled provider connections from combo builder
The combos page's fetchData() only filtered available connections by
testStatus ("active"/"success"), so a connection the user had
explicitly disabled (isActive: false) could still show up in the
combo builder if it carried a stale testStatus from before it was
disabled.
Add filterActiveConnections() in src/shared/utils/connectionStatus.ts
and apply it ahead of the existing testStatus filter.
Co-authored-by: itolstov <attid0@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/2526
* chore(changelog): fragment for #6984
* fix(combos): keep combos page within frozen size cap
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* fix(combos): extract filterUsableConnections to shrink the combos god-file
The combos page only filtered provider connections on testStatus, so a
connection the user had explicitly disabled survived with a stale
"active"/"success" status. The isActive + testStatus gate now lives in
the shared connectionStatus util as filterUsableConnections(), which the
page calls in a single line.
This keeps src/app/(dashboard)/dashboard/combos/page.tsx BELOW its frozen
file-size cap (4653 vs 4655 congelado — the file shrinks by 2 lines vs the
release tip) without touching config/quality/file-size-baseline.json, as
the gate asks ("modularize/extraia (DRY) para encolher").
The regression test now exercises filterUsableConnections directly instead
of hand-mirroring the page's filter chain, so it guards the real code path.
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
* fix(combos): drop nullish entries in filterActiveConnections
`connection?.isActive !== false` evaluated to true for null/undefined
entries, so nullish elements survived the filter. Callers read properties
off the result — filterUsableConnections() reads `connection.testStatus`
— which would throw "TypeError: Cannot read properties of null".
Guard with an explicit truthiness check. Covered by a test that fails
against the previous predicate.
Reported-by: gemini-code-assist
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
---------
Co-authored-by: itolstov <attid0@gmail.com>
getStaticModelsForProvider() only defined literal catalogs for
linkup-search, ollama-search, and searchapi-search out of the 12 ids in
SEARCH_PROVIDERS. The other 9 (serper-search, brave-search,
perplexity-search, exa-search, tavily-search, google-pse-search,
youcom-search, searxng-search, zai-search) returned undefined and hit
the 400 "does not support models listing" tail in the models route
during the "Import Models" step.
Instead of adding 9 more one-off literal entries, generalize the class:
when a provider has no dedicated STATIC_MODEL_PROVIDERS entry, fall
back to a catalog derived from SEARCH_PROVIDERS[id].searchTypes (every
search-registry entry already declares this). Future search providers
added to searchRegistry.ts automatically get a usable catalog with zero
extra code.
Closes#7529
ResponsesWsSession.persistHistory() guarded on a single historyLogged
boolean set once for the lifetime of the WebSocket connection. When a
Codex client reuses one connection for multiple sequential
response.create turns, only the first terminal event was persisted to
call_logs — every subsequent turn's usage/history was silently
dropped. firstResponseBody had the same per-connection freeze (||=),
so even a hypothetical second log entry would still carry turn 1's
request body.
Replace the boolean with a Set keyed by the terminal event's
response.id (falling back to a session-scoped sentinel for
session-ending failure paths that don't carry a response id: prepare
failure, upstream error/close, connect failure), and track each
turn's own request body via currentRequestBody instead of freezing on
firstResponseBody. This logs exactly once per logical turn while
keeping session-ending failures logged exactly once, and each logged
call now carries its own terminal response id and request payload.
Regression test: tests/unit/responses-ws-proxy-multi-turn-history.test.ts
opens one WS connection, sends two response.create turns, and asserts
two distinct call-log entries land at the internal bridge, each with
its own response id and request body.
Closes#7388
* fix(sse): sanitize empty-signature thinking blocks + hoist strict-provider system messages (#6953, #7293)
#6953: prepareClaudeRequest's "preserve latest-assistant thinking verbatim"
guard (claudeHelper.ts, anti-400 for legitimate Anthropic replay) did not
distinguish a genuine Claude signature from an empty one fabricated by a
non-Anthropic leg (e.g. codex reasoning_content). It forwarded signature:""
verbatim to Anthropic, which always 400s ("Invalid signature in thinking
block"), permanently locking combo routing onto the non-Anthropic leg. The
response-side half of this bug (openai-to-claude.ts synthesizing the empty
signature in the first place) was already fixed by #6982/PR#6982; this PR
closes the remaining request-side half. Fix: the verbatim-preserve guard now
requires every thinking-ish block on the latest assistant message to carry a
non-empty signature/data; otherwise it falls through to the existing
sanitization path (redacted_thinking + DEFAULT_THINKING_CLAUDE_SIGNATURE)
already applied to older turns.
#7293: translateRequest() is the single outbound choke point every chat
request passes through, including same-format (OpenAI→OpenAI) passthrough
where none of the format-specific translators run. systemMessageMustBeFirst()
/ PROVIDERS_SYSTEM_MUST_BE_FIRST (src/lib/memory/injection.ts, #6135/PR#6225)
was only consulted by the memory injector, so a client-injected system
message landing mid-array (OpenCode/Kilo Code style clients, Discussion
#6129) reached strict providers (xiaomi-mimo) untouched and 400'd. Fix: a new
helper (open-sse/translator/helpers/strictSystemHoist.ts) hoists every system
message onto index 0 for strict providers, reusing systemMessageMustBeFirst()
as the single source of truth, merging (never dropping) multiple offenders in
original order, and no-op'ing (same array reference) for non-strict providers
and already-compliant requests to preserve prompt-cache prefix stability.
Both defects live in the same file cluster (openai-to-claude request-path
translator + its helpers), hence one PR for both issues per triage guidance.
Regression tests:
- tests/unit/repro-6953.test.ts — RED (actual signature:'' forwarded) → GREEN
- tests/unit/probe-7293-strict-system-hoist.test.ts — RED (system message
left at index 10 of 70) → GREEN, plus multi-offender merge, existing-leading
merge, non-strict-provider no-op, and already-compliant no-op cases.
Gates run: file-size, complexity, cognitive-complexity (both at/under
baseline), typecheck:core (clean), eslint on changed files (clean),
test:vitest (254/254 green), plus all directly relevant existing suites
(translator-claude-helper-thinking, translator-xiaomi-mimo-reasoning-replay,
memory-system-first-6135, dashscope-cache-control-openai-2069,
xiaomi-mimo-provider, role-normalizer, translation.golden,
translators.property, translator-helper-branches, translator-claude-to-openai,
translator-same-format-null-flush — all green).
Closes#6953Closes#7293
* chore(quality): prune the now-stale claudeHelper no-explicit-any suppression (#6953)
The #6953 fix removed the single `any` that config/quality/eslint-suppressions.json
still had frozen for open-sse/translator/helpers/claudeHelper.ts, so the entry became
stale and ESLint's stale-suppression enforcement failed the 'No new ESLint warnings'
gate — the gate went red because the code got better.
Pruned that one entry only (never a global --prune-suppressions: other entries are
other sessions' frozen debt).
Z.AI's glm-4.6v vision endpoint enforces a 32768 max_tokens ceiling
server-side and 400s when a client sends a larger explicit max_tokens (e.g.
a client defaulting to 65536). paramSupport.ts's STRIP_RULES already has a
working clampToModelMaxOutput/maxOutputCap mechanism (used today for a
VolcEngine Kimi rule) but had no entry for zai/glm + glm-4.6v.
Added two rules: "zai" uses a fixed maxOutputCap (glm-4.6v is only reachable
there as a custom model attached to the connection, so it is not in
PROVIDER_MODELS["zai"] and clampToModelMaxOutput would find no catalog
ceiling); "glm" uses clampToModelMaxOutput (glm-4.6v IS in the registry
catalog there, GLM_SHARED_MODELS, maxOutputTokens: 32768).
Also discovered and fixed a second, deeper bug the "glm" rule alone would
not have caught: GlmExecutor.execute() drives its own fetch flow
(executeTransport()/transformForTransport()) and never runs through
DefaultExecutor.execute()'s stripUnsupportedParams() call site — so a
STRIP_RULES clamp entry for provider "glm" was dead code until
transformForTransport() now calls stripUnsupportedParams() directly.
Regression tests: tests/unit/zai-glm-max-tokens-clamp-7364.test.ts (reused
from the triage plan-file's RED probe, sanity assertion updated to lock the
fix instead of the bug) and tests/unit/glm-executor-max-tokens-clamp-7364.test.ts
(proves the real GlmExecutor.transformForTransport wiring, not just the
STRIP_RULES entry in isolation).
Gates run: check-file-size, check-complexity, check-cognitive-complexity,
typecheck:core, eslint (suppressions), tests/unit/zai-glm-max-tokens-clamp-7364.test.ts,
tests/unit/glm-executor-max-tokens-clamp-7364.test.ts,
tests/unit/executors-strip-unsupported-params.test.ts,
tests/unit/nvidia-minimax-thinking-strip.test.ts, tests/unit/glm-executor.test.ts — all green.
Refs #7364
DefaultExecutor.buildUrl()'s "zai"/"glm-coding-apikey" case always returned
the Anthropic Messages URL, ignoring a per-model targetFormat override
(custom-model dropdown, #2905) that resolves to "openai" — e.g. for a vision
model like glm-4.6v. chatCore/executionCredentials.ts now threads the
resolved override onto providerSpecificData.targetFormat so buildUrl (via
the new default/zaiFormatOverride.ts helper, extracted to respect the
file-size ratchet) can route to the OpenAI-compatible endpoint instead.
Separately, custom-model id lookup (lookupCustomModelMeta in
src/sse/services/model.ts, getCustomModelRow in src/lib/db/models.ts) did an
exact, case-sensitive match, so a model saved as "glm-4.6v" was invisible
when looked up as "glm-4.6V". Both now fall back to a case-insensitive match
after the exact match fails.
Regression tests: tests/unit/zai-glm-target-format-override.test.ts (reused
from the triage plan-file's RED probe) and
tests/unit/zai-execution-credentials-target-format-7364.test.ts (production
wiring in executionCredentials.ts).
Gates run: check-file-size, check-complexity, check-cognitive-complexity,
typecheck:core, eslint (suppressions), tests/unit/zai-glm-target-format-override.test.ts,
tests/unit/zai-execution-credentials-target-format-7364.test.ts,
tests/unit/executor-default-base.test.ts, tests/unit/custom-model-target-format.test.ts,
tests/unit/chatcore-execution-credentials.test.ts, tests/unit/chatcore-target-format.test.ts,
tests/unit/model-resolver.test.ts, tests/unit/model-alias-provider-resolution.test.ts,
tests/unit/combo-custom-provider-resolution.test.ts — all green.
Refs #7364
detectBinary() in tool-detector.ts never checked process.platform and never
passed shell:true, so on native Windows an installed CLI (npm installs
claude/codex/opencode as .cmd shims) was reported as NOT installed:
1. execFileImpl(binary, ["--version"]) fails without shell:true for .cmd
shims (Node's CVE-2024-27980 hardening).
2. the `which` fallback doesn't exist on native Windows (no WSL/git-bash).
Both threw, both were swallowed by empty catches, detectBinary returned
{installed: false}. cliRuntime.ts::locateCommand already solved this for the
runtime-spawn path (#968) but never propagated here — re-drift, per the
issue title.
Exports locateCommand from cliRuntime.ts and reuses it (+ shouldUseShellForCommand,
+ getLookupEnv) for the win32 existence/path probe in tool-detector.ts, keeping
the --version probe local but shell-gated. Also routes the which fallback through
the injectable execFileImpl hook (it previously called the raw execFileAsync,
making it unmockable and prone to false-positives from a real system which).
Root cause: waitForImageViaWebSocket() only parsed the singular
update_content.message (object) / payload.message / data.message shapes
in the celsius WebSocket frames chatgpt.com uses to deliver async
image_gen results. Some chatgpt.com deployments deliver the completed
tool-role image_asset_pointer message inside update_content.messages[]
(a plural array of { message: {...} } wrappers) instead, which produced
zero candidates, so the listener idled out the timeout and the request
failed with the generic 'ChatGPT Web completed without returning image
markdown' 502 with no x_image_resolution_failed flag.
Fix: also read update_content.messages[] and push each wrapped message
into the same candidate pipeline used for the singular shape.
Regression test: tests/unit/chatgpt-web-async-image-ws-shapes-7357.test.ts
drives the real ChatGptWebExecutor.execute() end-to-end (real SSE
parsing, real pollForAsyncImage()/waitForImageViaWebSocket()), mocking
only the network edges (tlsFetchChatGpt + global WebSocket), and proves
the plural-array frame now resolves to image markdown instead of being
dropped.
Gates run: check-file-size (OK), check-complexity (OK, 2054 <= 2056
baseline), check-cognitive-complexity (OK, 889 <= 890 baseline),
typecheck:core (clean), eslint on changed files (clean), full
tests/unit/chatgpt-web.test.ts (89/89), chatgpt-web-image-silentdrop.test.ts,
chatgpt-web-tools-5240.test.ts, chatgpt-web-models-split.test.ts,
chatgpt-web-sha3-boringssl-5531.test.ts, chatgpt-web-handoff-resume.test.ts,
chatgpt-web-citations(-escape).test.ts all pass.
captureCurrentProviderRequest mirrors every Bedrock Converse request into the
pending-request log tracker right after openAIToBedrockConverse() builds it,
including the decoded image.source.bytes Uint8Array. sanitizePayloadPII() and
redactPayload() in src/lib/logPayloads.ts both gate their recursive walk on
Array.isArray(), which is false for typed arrays, so each image fell into the
generic-object branch and got enumerated one JS key per decoded byte (twice,
once per function) before any truncation bound applied. For 3x ~1MB images
this took ~4s of synchronous, event-loop-blocking work, matching the
reporter's "1-2 images OK, 3+ fails" threshold and their --stack-size
observation (data-width pressure, not call-depth).
Add an opaque-binary short-circuit (ArrayBuffer.isView) ahead of the
Array.isArray branch in both functions, returning a fixed-size placeholder
instead of recursing. Apply the same guard to cloneBoundedForLog() in
open-sse/utils/requestLogger.ts for defense-in-depth (same blind spot, only
accidentally safe today via its own key-count slice).
Regression test reproduces the exact reporter shape (3x 1MB images) through
the real openAIToBedrockConverse() converter and protectPayloadForLog(),
asserting completion well under the previous ~4s and that binary bytes are
never expanded into per-byte object keys.
resolveRequestedModel() only special-cased "auto" and the composer
"-fast" suffix; every pinned Claude/GPT id carrying an effort/reasoning
suffix (e.g. "claude-opus-4-8-high", "gpt-5.5-high") fell through and
was sent to cursor's server verbatim as model_id, with an empty
parameters array. Cursor has no route for the suffixed id -- it only
knows the base id plus an out-of-band ModelParameter -- so it accepted
the request but returned an empty turn.
Split the known effort suffixes (-low/-medium/-high/-xhigh/-max) off
the base id: Claude ids surface an {id:"effort", value} parameter,
GPT ids surface {id:"reasoning", value}, matching the real cursor-agent
client's wire format. encodeAgentRunRequest()'s ModelDetails fields
derive from the same resolved base id, so the #3714 pinned-model
ModelDetails envelope stays correct without further changes.
Updates the existing resolveRequestedModel test that locked in the
buggy verbatim pass-through, and the #3714 ModelDetails test to assert
against the base id. Adds a dedicated regression test file proving the
Claude/GPT split plus non-regression of the "-fast" toggle and
unsuffixed ids.
#7268: classifyProviderError() only inspected the response body for
model-unavailable wording on 400/403/404, so a 401 body like "Model X is
not supported" (free-tier/aggregator providers) fell through to a generic
UNAUTHORIZED classification. Because chatCore.ts only calls lockModel(...,
"model_not_found", ...) on the MODEL_NOT_FOUND branch, the broken model was
never locked out and auto-combo kept re-selecting it every request. Added a
shared containsModelUnavailableMessage() regex (bounded, ReDoS-safe) in
errorClassifier.ts, consulted by the 401 branch before falling back to
ACCOUNT_DEACTIVATED/UNAUTHORIZED, and reused by modelFamilyFallback.ts's
isModelUnavailableError() for the literal "<model> is not supported" phrasing.
#7387: applySessionStickiness() (combo-level session stickiness) only gated
a sticky pin's release on testStatus (credits_exhausted/banned/expired) and
rateLimitedUntil (#6692's fix). It never consulted isAccountQuotaExhausted()
(src/domain/quotaCache.ts) — the authoritative per-window (5h/weekly) quota
signal that src/sse/services/auth.ts and sessionAffinityPin.ts (the
provider-level pin) already gate on. A connection whose quota window was
depleted, but that hadn't yet received a hard failure severe enough to flip
testStatus/rateLimitedUntil, was re-promoted to position 0 on every request
regardless of routing strategy. Added isStickyConnectionQuotaExhausted(), a
dynamic-import seam (mirroring resolveConnectionHealth/resolveSaturation, no
new static edge from open-sse/ into src/domain/) with an injectable checker
for tests, gating the release condition alongside the existing checks.
Regression tests: tests/unit/repro-7268-401-model-not-supported-lockout.test.ts,
tests/unit/repro-7387-sticky-quota-exhausted.test.ts (both RED before, GREEN
after). Existing sticky/error-classifier suites re-run and stay green.
Closes#7268Closes#7387
- byProvider now resolves the internal provider id to its configured
display name via getProviderById() (fallback: raw id for providers
not in the static registry). Fixes the Usage page showing "codex"
instead of "OpenAI Codex".
- byModel's in-memory dedup key now uses the normalized model name
instead of the raw one, so the same logical model recorded under a
bare and a provider-prefixed spelling (e.g. "glm-5.2" vs
"z-ai/glm-5.2") merges into a single aggregated row instead of
appearing twice with the same displayed name.
- Introduces a local UsageRows type alias in route.ts to shrink the
repeated "as Array<Record<string, unknown>>" casts back under the
frozen file-size baseline once the file was touched.
* fix(dashboard): resolve costs page 500 from out-of-scope t() in TopListCard (#7272)
TopListCard (CostOverviewTab.tsx) referenced the bare identifier `t`
from an outer component's scope when rendering the zero-cost /
!hasCostData branch, throwing "ReferenceError: t is not defined" and
crashing /dashboard/costs?range=all&apiKeyIds=...&groupBy=model
whenever a filtered slice landed only $0-cost rows.
Extracted TopListCard into its own component file and threaded the
resolved legacyFreeLabel string in as a prop, mirroring the existing
CostBreakdownTable pattern in the same file. Also fixes a case of the
typecheck:core dashboard .tsx coverage gap tracked in #7033.
* test(dashboard): move TopListCard #7272 regression test to vitest UI project
The node:test unit runner cannot load TopListCard's import chain
(@/shared/components -> ProviderIcon -> @lobehub/icons ESM), which made
the "Impacted unit tests (TIA subset; blocking)" CI job red with
"Unexpected token 'export'" on tests/unit/costs-toplistcard-legacy-free-label-7272.test.ts.
Moved the regression test to tests/unit/ui/ and rewrote it against the
vitest UI project (test:vitest:ui, blocking in the test-vitest CI job),
which handles the ESM import chain natively. Verified the test still
fails with "ReferenceError: t is not defined" against the pre-#7272
TopListCard body and passes against the fixed component.
* fix(db): pre-init sql.js WASM ahead of any getDbInstance() consumer (#7288)
* fix(db): close sqljs preinit ordering gap without top-level await (#7288)
The previous fix added a top-level await barrier at the bottom of
src/lib/db/core.ts to guarantee sql.js pre-init before any consumer
reached getDbInstance(). That made core.ts an async ES module, which
broke esbuild's CJS require() bundling for every test file that does
require("../../src/lib/db/core.ts") (tsx's CJS require hook rejects
requiring a transitive dependency with a top-level await), and caused
unrelated tests running in the same node:test process to fail with
"Promise resolution is still pending but the event loop has already
resolved".
Move the fix to the real startup entrypoint instead: registerNodejs()
(src/instrumentation-node.ts) now awaits ensureDbReadyForBoot() before
ensureSecrets()/clearStaleCrashCooldowns()/getSettings()/initAuditLog(),
all of which reach getDbInstance() transitively. ensureDbInitialized()
is idempotent, so later getDbInstance() calls are free cache reads.
The driverFactory.ts error-surfacing improvements from the original
#7288 fix (logging swallowed sync-driver errors, surfacing the real
sql.js pre-init failure instead of the generic "not pre-initialized
yet" message) are unchanged.
Updated tests/unit/db-sqljs-preinit-ordering-gap-7288.test.ts to prove:
no top-level await in core.ts, the corrected call order in
registerNodejs() (source-order assertion), and the original
getDbInstance()-no-longer-throws-the-misleading-message behavior driven
via the same warm-up path ensureDbReadyForBoot() now guarantees ahead
of every other startup step.
* refactor(db): drop the orphaned preInitSqlJsIfSyncDriversUnavailable helper (#7288)
Moving the ordering guarantee to registerNodejs() left this exported
helper with zero production callers — its own docblock still claimed it
was 'Chamada no top level de core.ts', describing an architecture the
hotfix removed. It was also redundant: ensureDbInitialized() already does
tryOpenSync-then-preInitSqlJs on the real boot path (core.ts:1358).
Its two tests exercised the helper as a stand-in for the real warm-up
('Simulates the fixed ordering'), so they proved a simulation rather than
production behaviour. They now drive ensureDbReadyForBoot()/tryOpenSync()
directly. The ordering guard still fails against the unfixed
instrumentation-node.ts (verified) and the whole file is 4/4 green.
The live parts of the driverFactory change (logSwallowedDriverError,
getSqlJsPreInitError) are untouched — both have real callers.
validateResponseQuality() only recognized Claude SSE lifecycle events
(message_start/content_block_*/message_stop/message_delta.stop_reason).
An OpenAI-shape stream (choices[].delta) that emits some bytes (e.g. a
role-only delta) and then closes without ever carrying finish_reason
(and without a data: [DONE] sentinel) fell through to the generic
replay branch and was forwarded to the client as a success instead of
triggering combo failover.
Adds OpenAI-shape lifecycle tracking (hasChoicePayload/hasTerminalMarker)
parallel to the existing Claude tracking: when an OpenAI-shape chunk was
seen but the stream ends without finish_reason or [DONE], and no
recognized content was found, mark the response invalid so combo
failover retries a sibling target. Healthy OpenAI streams (finish_reason
present, or real content found) are unaffected — they exit the peek
loop before reaching this check, preserving the #3399/#3685
pass-through contract.
Regression test: tests/unit/combo-streaming-openai-no-finish-reason-7285.test.ts
Root cause: validateOpenAILikeProvider's chat-probe status handling
(src/lib/providers/validation/openaiFormat.ts) only special-cased
401/403, 404/405, and >=500 — every other status, including 429, fell
through to the unqualified return { valid: true, error: null }. For
permanently-throttled free tiers (e.g. opencode-zen, classified
'avoid' in freeTierCatalog.ts), the dashboard connection Test reported
green forever while real traffic hit 429 on every request, with no
signal to the user.
Fix mirrors the existing validateBedrockProvider 429 precedent in the
same file: keep valid:true (the key is accepted) but add a warning
field describing the rate limit, instead of an indistinguishable pass.
Regression test: tests/unit/issue-7284-connection-test-masks-429.test.ts
mocks a 404 /models probe followed by a 429 chat probe and asserts the
result now carries { valid: true, error: null, warning: <rate-limit string> }.
Gates run (all green):
- node --import tsx/esm --test tests/unit/issue-7284-connection-test-masks-429.test.ts
- node --import tsx/esm --test tests/unit/provider-validation-firepass-403.test.ts tests/unit/validation-format-validators-split.test.ts (existing tests of the touched area, unaffected)
- node scripts/check/check-file-size.mjs
- node scripts/check/check-complexity.mjs (2054/2056 baseline, unchanged)
- node scripts/check/check-cognitive-complexity.mjs (889/890 baseline, unchanged)
- npm run typecheck:core
- npx eslint --suppressions-location config/quality/eslint-suppressions.json <changed files>
Termux/Android's Node reports process.platform === 'android'.
playwright-core's serverRegistry.js throws 'Unsupported platform:
<platform>' from a top-level IIFE at require time, so merely
importing the playwright package crashed — no browser ever launched.
claudeTurnstileSolver.ts was the only playwright consumer in the
codebase with a static top-level import; every other call site
(browserPool.ts, inAppLoginService.ts) already lazy-loads it. That
static import is unconditionally reachable from the Next.js
instrumentation hook on every boot via open-sse/executors/index.ts,
so any unsupported platform crashed the whole server at startup
regardless of which provider was configured.
Fix: import type { Browser, Page } (erased at compile time) and move
the chromium binding to a lazy await import("playwright") inside
solveTurnstile(), matching the existing pattern.
chatCore.ts fed applyCompressionAsync's supportsVision option from
isVisionModelId() — the deliberately-conservative model-id fragment
heuristic in src/shared/constants/visionModels.ts — instead of the
authoritative getResolvedModelCapabilities().supportsVision used by every
other vision-aware path (e.g. the vision-bridge guardrail).
gpt-5.5 is registered with supportsVision:true in modelSpecs.ts, but the
fragment list has no gpt-5.x entry, so the heuristic wrongly returned
false. That false reached lite.ts's replaceImageUrls(), whose gate is
`supportsVision !== false`, silently stripping every image_url block
before the request ever reached the executor.
getResolvedModelCapabilities().supportsVision resolves to null (not
false) for genuinely unknown models, which the same !== false gate
already treats as preserve-by-default — matching the conservative
semantics used elsewhere and avoiding the #4071/#4012 class of bug
(blinding a model that can actually see).
Root cause: the Providers page model-name filter
(filterConfiguredProviderEntries) matched only against
getModelsByProviderId(...), the static curated model registry, never
the live/synced catalog for the connection. Aggregator providers
(openrouter, kilocode, theoldllm...) declare a single-entry static
placeholder (e.g. openrouter's {id:'auto',name:'Auto (Best Available)'}),
so searching for any real upstream model name could never match and the
whole provider silently disappeared from the list.
Fix: source the live/synced catalog (already persisted per-connection via
GET /api/synced-available-models, the same store the combo builder's model
picker already relies on) via a new useSyncedModelsByProvider hook, and
union it with the static registry inside the filter. An empty/never-synced
catalog falls back to the static-only match so already-correct static
providers are unaffected.
Regression test: tests/unit/provider-model-filter-live-catalog-7250.test.ts
reproduces the original bug (static-only match returns 0 results for a
real model name) and proves the fix (live catalog match returns 1),
plus non-regression coverage for the static-only fast path and unrelated
providers.
* fix(cli): Windows cert check/uninstall key off the real CA identity, not a hardcoded legacy host (#7275)
checkCertInstalledWindows()/uninstallCertWindows() queried the Windows Root
store by the literal legacy hostname daily-cloudcode-pa.googleapis.com
regardless of the certPath passed in (the check's param was even
underscore-prefixed/unused). It only worked because that hostname happens to
be the generated CA's own commonName today (generate.ts derives it from
ANTIGRAVITY_TARGET.hosts[0]) -- a coincidence, not a derivation, with no
shared symbol coupling the two. installCertWindows() was already correct.
Both functions now derive a SHA-1 thumbprint straight from the certPath file
via the new exported certutilThumbprint() helper (reusing the existing
getCertFingerprint() logic checkCertInstalledMac() already keys off), so the
Windows store lookup/delete always matches the real generated CA regardless
of any future rename/reorder of ANTIGRAVITY_TARGET.hosts in generate.ts. Same
anti-pattern class as #6338 (DNS side).
* test(cli): assert certutil argv precisely instead of scanning for the legacy host (#7275)
The two negative substring checks tripped CodeQL's
js/incomplete-url-substring-sanitization (high) by looking like URL
sanitization while actually asserting a certutil argv. Replaced with
assertions on the exact argv / extracted certId, which are strictly
stronger: pinning the value proves no other identity can be passed.
Both still fail against the unfixed install.ts (5/5 RED) and pass with
the fix (5/5 GREEN).
deepMergeFallback in src/i18n/request.ts only substituted the English
fallback value when a key was entirely undefined. Keys backfilled by
scripts/i18n/sync-ui-keys.mjs with the __MISSING__:<english> sentinel
exist on the target object, so they passed through untouched and were
rendered raw to the user (395 zh-TW keys, 337 pt-BR keys, systemic
across locales).
Now any target leaf that still carries the __MISSING__: prefix is
treated the same as an absent key, so the clean EN value wins.
Does not touch the underlying translation content (395/337 strings) -
that is a separate content workstream.
The PKCE callback server binds the SERVER's loopback (localhost:PORT). When
the operator drives the OAuth flow from a different machine (OmniRoute on a
remote host/VPS), the provider redirects the browser to the operator's OWN
localhost:PORT — the confirmation screen hangs forever with no explanation.
start-callback-server now inspects the request Host: on a non-loopback host it
returns { remoteHost, tunnelCommand, message } so the UI can show the
'ssh -L PORT:127.0.0.1:PORT' instruction (or steer to the paste/import flow)
instead of a silent hang. Loopback access is unaffected. The Host header is
spoofable, so this drives only a UI hint — never an auth decision.
Logic extracted to remoteOAuthHint.ts (keeps the god-route under its size
budget and makes it unit-testable). TDD: 4 tests covering loopback (no hint),
null host (fail-open), and remote host (correct tunnel command for both the
fixed 1455 and OS-assigned ports).
Closes#7523
Every non-streaming Codex chat request for a ChatGPT-account connection failed
instantly with [502]: Response body is already used (reset after 1m). The
streaming/playground path was unaffected.
Root cause: peekCodexSseTransientError (open-sse/executors/codex.ts) peeked the
SSE prefix with response.body.getReader(), then called reader.releaseLock() and
response.body.getReader() a SECOND time on the same already-disturbed body to
build the replacement stream. Re-acquiring a reader on a disturbed body throws
on undici ('Response body is already used'); chatCore's generic upstream-error
handling then stamped the TypeError with a default 60s cooldown, masking a pure
code defect as a rate limit (and tripping the codex circuit breaker).
Fix: keep the single reader already held; never touch response.body again.
TDD: a getReader spy that throws on the 2nd acquire reproduces the exact hazard
— 1 test RED against the release code, 2/2 GREEN with the fix; the replacement
body stays byte-identical to the upstream SSE. No regression across the codex
unit suite. Reproduced live on the VPS 2026-07-16.
POST /api/oauth/codex/import accepted a payload with an already-invalidated
refresh_token and persisted it as an 'active' connection that could never
work — the failure surfaced confusingly only on first real use, long after
the import looked successful.
Validate each record's refresh_token against OpenAI's OAuth endpoint before
persisting, reusing refreshCodexToken() (free exchange, no quota). On an
unrecoverable refresh error the record is rejected with a clear re-auth
message; a valid token imports as before, with any rotated tokens applied.
Bulk import still processes each record independently — one dead token no
longer blocks the valid ones.
Reproduced live 2026-07-16: a 2026-07-10 auth.json imported clean but its
refresh_token returned 401 refresh_token_invalidated. TDD: 3 tests RED against
the old route, 5/5 GREEN with the fix.
Closes#7522
The connection Test button always reported success for a Codex connection
backed by a ChatGPT account: the probe sent `gpt-5.3-codex`, a codex-only
model the ChatGPT-account backend rejects outright with a 400 — the same
status the probe treats as 'auth accepted, body invalid'. A bad token and a
good token both came back 400, so Test could never fail on a bad token.
Probe with `gpt-5.5` (confirmed served for ChatGPT-account sessions via live
VPS test 2026-07-16) instead; `input: []` still yields the intended 400 for a
good token, 401/403 for a bad one.
Live verification (VPS): gpt-5.3-codex, gpt-5.6-sol and the gpt-5*-codex ids all
return 'not supported when using Codex with a ChatGPT account'; gpt-5.5 and
gpt-5.6-terra answer normally on the same account.
Closes#7521
* fix(electron): materialize Turbopack hashed-module symlinks during packaging
Reconstructed onto release/v3.8.47 to drop unrelated main-drift (deps/electron/proxy
files belong to #6620, not this PR); keeps only the author's changes.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* fix(electron): actually enable materializeSymlinks on the electron standalone path
The option existed in assembleStandalone but no production callsite passed it,
so packaged builds still shipped absolute symlinks into the build machine's
worktree for Turbopack hashed externals (better-sqlite3-<hash>, sqlite-vec-<hash>)
— verified by dpkg -c on a freshly built .deb. One-line enablement on the
electron prepare path, which is exactly the surface #6724/#6594 report.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: huohua-dev <258873123+huohua-dev@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
With --depth=1 the base ref is grafted, so 'git diff base...HEAD' resolves a
wrong merge-base for PR branches that recently merged the release branch. The
three-dot diff then attributes ALREADY-MERGED sibling PRs' changes to the PR
under test, producing false high-signal reds (deleted test files / weakened
asserts that exist in no ref reachable from the PR).
Observed live on #7329: the job blamed it for tests/unit/ui/provider-plan-config.test.tsx
(deleted by an unrelated merged PR) and for #7106's antigravity files. Local
reproduction with full history returns PASS for the same head.
The job's checkout is already fetch-depth: 0, so the full base fetch only
updates the ref — negligible cost.
* fix(dashboard): align onboarding welcome feature cards vertically
* fix(dashboard): align onboarding tier content
* chore: scope onboarding PR to UI fix
* i18n(pt-BR): add onboarding.tier.flowCaption + afterSetup keys
The two new tier keys added to en.json were missing from pt-BR.json, tripping
the i18n-pt-br no-drift test (#6695). Add their pt-BR translations.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com>
* fix(dashboard): add vision-capability toggle for custom OpenAI-compatible models (port from 9router#1904)
detectVisionInput()/getCustomVisionCapabilityFields() already honoured an explicit
supportsVision flag on a custom-model record, but there was no way to set it: the
POST/PUT /api/provider-models Zod schema and updateCustomModel()/addCustomModel()
silently dropped the field, and the 'Custom Models' add/edit UI had no checkbox at
all. Self-hosted/local backends that don't self-report an image input modality
(OpenRouter-style architecture.input_modalities) therefore had no way to be flagged
vision-capable, so the vision tag never appeared and image inputs were rejected.
Reported-by: nguyenphi37 (https://github.com/decolua/9router/issues/1904)
* refactor(dashboard): extract providerCredentialText from providerPageHelpers to respect the file-size gate
providerPageHelpers.ts is a frozen god-file (cap 1053, split(\n).length metric)
and this PR's own +3 lines (the #1904 supportsVision field) pushed it to 1054,
failing check:file-size. Extract the cohesive providerText utility + the 4
web-session-credential label/hint/title helpers into a new leaf module
(providerCredentialText.ts), re-exported from providerPageHelpers.ts for
backward compatibility so all existing import sites keep working unchanged.
File now sits at 946 lines, well under the frozen cap.
* refactor(db): extract tri-state override helper to keep the complexity ratchet at baseline
The #1904 supportsVision override added a second copy of the "absent keeps /
null clears / else coerce" block already used by preserveOpenAIDeveloperRole,
pushing updateCustomModel to 84 lines and check:complexity to 2057 > 2056.
The file-size failure was masking this one: the gate exits on its first red,
so complexity never ran until providerPageHelpers was back under its cap.
Fold both blocks into applyTriStateBooleanOverride(). Behavior is unchanged —
updateCustomModel is back under max-lines-per-function and the global count
returns to the 2056 baseline (cognitive-complexity stays at 890).
* Use OpenAI chunks for early chat keepalives
* Update keepalive assertion to match chat completion chunk format
---------
Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com>
* fix(combo): detect empty content_block in streaming SSE peek (port from 9router#1382)
The bounded SSE peek in validateResponseQuality() treated ANY content_block_start/delta/stop event as proof of real output and stopped buffering immediately, without checking whether the block actually carried text/tool_use content. Some upstreams (reported: DeepSeek, GLM via claude→openai translation) can open and close a text content_block with empty text and no tool_use on tool-heavy requests — the gateway logged success and forwarded a client-visible empty completion, and combo routing never failed over to the next model.
Track real content separately from 'a content_block_* event was seen': a tool_use/redacted_thinking block start is self-evidently real signal, a text/thinking block start is not (real content only confirmed via a subsequent delta carrying non-empty text/thinking, or an input_json_delta streaming tool arguments). A completed lifecycle (message_start + message_delta/stop) that never produced real content now fails validateResponseQuality(), matching the existing content_filter empty-stream detection path (#3685).
Reported-by: heishen6 (https://github.com/decolua/9router/issues/1382)
* refactor(combo): extract SSE lifecycle applier to keep the complexity ratchets at baseline
The #1382 empty-content_block peek added a branchy switch inline in
parseAccumulatedSse, pushing check:complexity to 2057 > baseline 2056.
Move the switch to a module-level applySseLifecycleEvent() and hold the
four lifecycle booleans in a single SseLifecycleFlags object threaded
through it, so the closure no longer copies flags in and out per event.
The per-event predicates (content_block_start / content_block_delta /
message_delta) are split into small guard helpers, which keeps the
applier flat — cognitive complexity punishes nesting, and an earlier
switch-only extraction traded the cyclomatic ratchet for a cognitive
regression at 891 > 890.
Logic is unchanged; both ratchets are now green (complexity 2055,
cognitive-complexity 890) and the #1382 regression tests still pass.
A fusion combo fans every panel model out in parallel and buffers each
model's full response text in memory simultaneously. With the runtime heap
capped by Dockerfile's OMNIROUTE_MEMORY_MB (default 1024MB), a large panel
(reported: ~73 models via an 'auto' combo with strategy: fusion) with
sizable concurrent responses can exceed the heap ceiling and OOM-crash the
whole container instead of failing one request.
handleFusionChat now rejects panels above a configurable hard cap
(FUSION_DEFAULTS.maxPanel = 40, overridable per-combo via
fusionTuning.maxPanel) with a clean 400 before fan-out begins.
Reported-by: Phong Vu (@fontvu) (https://github.com/decolua/9router/issues/1905)
* fix(api): check Vercel SSO-protection PATCH response on relay deploy (port from 9router#1037)
The Vercel relay deploy route disabled Deployment Protection (SSO) by firing a PATCH request with .catch(() => {}) and never checking res.ok. When Vercel rejects or no-ops the PATCH (plan doesn't allow disabling protection, an under-scoped token, etc.), the relay was still saved and activated as a healthy proxy pool, and later requests routed through it failed with an undiagnosed 403 Access denied from Vercel's own deployment protection — indistinguishable from an upstream-provider rejection (e.g. Codex/ChatGPT edge-IP blocking).
Extract disableSsoProtection() to check the PATCH response and surface an ssoProtectionWarning in the deploy response when it fails, so the failure source can be diagnosed instead of silently masked.
Reported-by: Rico Aditya (@ricatix) (https://github.com/decolua/9router/issues/1037)
* refactor(api): extract vercel-deploy POST helpers to keep the cognitive-complexity ratchet at baseline
The SSO-protection check added to POST pushed its cognitive complexity from
15 to 21, regressing the cognitive-complexity ratchet (891 > baseline 890).
Extract two pure helpers with identical behavior:
- buildDeployErrorResponse(): the sanitized non-ok Vercel deploy response
- resolveSsoProtectionWarning(): the SSO PATCH check + warning string
POST now reads as a flat sequence of guards. No behavior change.
* fix(cli): remove MITM DNS spoof entries before killing server process (port from 9router#1809)
stopMitm() killed the spawned MITM server process first and only removed the /etc/hosts DNS-spoof entries afterward. During that window any client whose DNS still resolved a target host to 127.0.0.1 but whose MITM listener was already dead got connect ECONNREFUSED 127.0.0.1:443 — exactly the community-confirmed workaround (stop DNS before stopping the server) proves. Swap the two steps so DNS is always cleared first, mirroring the ordering already used by repairMitm() and handleExitCleanup().
Reported-by: dionisius95 (https://github.com/decolua/9router/issues/1809)
* refactor(mitm): extract repair planning out of manager to respect the file-size cap
The #1809 DNS-before-kill ordering fix pushed src/mitm/manager.ts to 813
lines, over the 800-line cap check:file-size enforces for non-frozen files.
Move the pure repair-planning pieces (collectManagedHosts, the RepairPlan
shape and its filesystem/cert/DNS sweep) into a sibling src/mitm/repair.ts.
The in-memory session bookkeeping repairMitm() owns — cached sudo password,
orphaned flag, PID file — deliberately stays in manager.ts, so the seam is
"plan the repair" vs "own the session".
manager.ts is now 731 lines; behavior is unchanged. The DNS-first ordering
fix and its regression guard (tests/unit/mitm-stop-dns-before-kill-1809.ts)
are untouched and still pass.
* fix(mitm): split stopMitm() DNS/kill steps to fix complexity ratchet regression
stopMitm()'s new DNS-before-kill ordering (#1809) pushed its cyclomatic
complexity to 18 (max 15), regressing the complexity ratchet from 2056 to
2057. Extract the DNS-removal step and the process-kill step (in-memory +
PID-file fallback) into two private helpers, mirroring the existing
performRepairSteps() extraction pattern in repair.ts. Behavior unchanged;
complexity back at 2056 (cognitive-complexity drops to 889, one under
baseline).
parseInnerCall only split arg segments on a newline between the arg name
and its value. Cursor's live Composer/Auto output has been observed using a
single space instead, so those segments were treated as one long
(space-containing) arg name with an empty value, silently no-opping
Write/tool calls for Composer/Auto models.
Fall back to splitting on the first whitespace boundary when no newline is
present in the segment.
Reported-by: way-art (https://github.com/decolua/9router/issues/1811)
* fix(cli): verify better-sqlite3 native binary is actually loadable (port from 9router#2493)
isBetterSqliteBinaryValid() only checked the .node file's magic bytes (ELF/Mach-O/PE header), never whether the binary was built for the ABI (NODE_MODULE_VERSION) of the Node runtime that loads it. A stale or foreign-ABI binary passed the check and then segfaulted the process on the first database call instead of triggering a rebuild via npmInstallRuntime(). The fix adds a real load probe (require() in a throwaway subprocess) after the magic-byte check, so an incompatible binary is now correctly reported as invalid and the runtime self-heal reinstalls it.
Reported-by: Manikandan (@mrprohack) (https://github.com/decolua/9router/issues/2493)
* chore(changelog): move #2493 entry to changelog.d fragment
Consistency with the repo's canonical changelog.d/fixes/ workflow (avoids
merge-storm re-conflicts from editing CHANGELOG.md directly).
* fix(executors): forward X-Session-ID/X-Title agent metadata headers (port from 9router#2413)
Custom agent clients (e.g. non-OpenCode providers) commonly send X-Session-ID and X-Title headers for upstream request tracking/attribution, but forwardOpencodeClientHeaders() only forwarded x-opencode-* keys plus User-Agent, silently dropping these for every client. Extends the existing case-insensitive allowlist forwarding path with x-session-id/x-title.
Reported-by: Atikur Rahman Chitholian (@chitholian) (https://github.com/decolua/9router/issues/2413)
* chore(changelog): move #2413 entry to changelog.d fragment
Consistency with the repo's canonical changelog.d/fixes/ workflow (avoids
merge-storm re-conflicts from editing CHANGELOG.md directly).
Root cause: validateOpenAICompatibleProvider's chat-completions probe fallback treated ANY 4xx other than 401/403/429/400 as a silent 'credentials valid' pass with no warning, so a bogus/non-standard model id (e.g. Featherless/OpenRouter vendor/model typos) went undetected at Check time. The first real request then hit the upstream 404 model_not_found and the per-model lockout, holding the model unavailable for the configured reset window with no prior indication anything was wrong.
User-visible effect: 'Check' now returns valid:true with an explicit warning (including the upstream error message when parseable) whenever the chat probe answers 404, so a bad model id is caught before it reaches production traffic and the lockout.
Reported-by: advane204f (https://github.com/decolua/9router/issues/2032)
Root cause: SmartCrusher's system-message guard only excluded role === "system", but Codex CLI (open-sse/executors/codex.ts) sends its instructions/tool-schema turn with role "developer" (the Responses-API equivalent of system used by newer models). Every other system-exclusion guard in this codebase also covers developer (roleNormalizer.ts, contextManager.ts, claudeUpstreamMessages.ts, etc.) except this one, so Headroom happily tabular-compacted JSON arrays embedded in the developer turn (e.g. an update_plan tool schema example), corrupting the instructions the model needs to call the plan tool and breaking Codex CLI plan mode.
Fix: extend the guard in crushMessages()/collectCompactableArrays() (smartcrusher.ts) to skip role === "developer" alongside role === "system".
Reported-by: SingCJ (https://github.com/decolua/9router/issues/2132)