* chore(release): open v3.8.45 development cycle
* chore(release): parallel-cycle flow — sync-next-cycle script + Hard Rule #21 semantics (#6203)
Integrated into release/v3.8.45
* perf(test): tsx/esm loader + tsx 4.23 + órfãos recuperados + CI via npm scripts (plano testes+CI, Pacote 1) (#6214)
* perf(test): tsx/esm loader, tsx 4.23 bump, orphan tests recovered, CI runs npm scripts
Pacote 1 (quick wins) do plano mestre testes+CI:
- Swap --import tsx -> --import tsx/esm on the 19 test scripts: the repo is pure
ESM and the CJS hook costs ~1.3s PER test process (2,462 processes/run).
Measured: bootstrap 2.3s -> 1.1s; real-suite A/B (tests/unit/db, 12 files)
22.2s -> 14.1s wall (-36%), 82/82 pass. Non-test scripts keep the full hook.
- Bump tsx ^4.22.3 -> ^4.23.0 (fix for privatenumber/tsx#809 startup regression;
helps module resolution on big graphs — hook cost unchanged, honest note).
- Recover 22 ORPHAN test files (tests/unit/feature-triage/*.test.mjs, 53 cases,
53/53 pass) that matched no glob and ran in NO CI job; drop the dead
'executors' dir from the braces glob.
- Single source of truth for the unit-suite invocation: new test:unit:ci:shard
(shard via TEST_SHARD env) called by ci.yml test-unit/node24/node26/coverage
and quality.yml fast-unit — closing two silent drifts: CI was NOT importing
setupPolyfill.ts, and the fast path glob OMITTED tests/unit/memory + usage.
quality.yml TIA step + ci.yml test-integration get the tsx/esm swap only.
Validation: full unit suite 21,153 tests / 21,135 pass / 13 skip (5 fails =
known load-flake family, 10/10 green rerun isolated); vitest 237/237; smoke
db+feature-triage 135/135. Record run 2 = ci.yml workflow_dispatch on the
stacked pacote-2 branch (clean runners).
* fix(test): dashboard UI tests keep full tsx hook; recover 15 more .mjs orphans; extend discovery gate to .mjs
Follow-up do dispatch de validacao (run 28720431562), que pegou 2 problemas reais:
1. tests/unit/dashboard/** (11 arquivos, 102 casos) importam componentes React cujo
grafo puxa @lobehub/icons — o build es/ dele faz require() interno de arquivos com
sintaxe ESM, que so funciona com o patch CJS do tsx (sem ele: SyntaxError
'Unexpected token export' no CI; local vira crawl de ~60s/arquivo). Esses 11
arquivos rodam agora numa 2a invocacao com --import tsx COMPLETO (mesmo shard),
e o resto da suite mantem tsx/esm (-50% bootstrap). Validado: 102/102.
2. check:test-discovery falhou porque ancorava textualmente os globs nos workflows —
COLLECTORS atualizado p/ o modelo fonte-unica (ancora = nome do script
test:unit:ci:shard nos workflows) + varredura ESTENDIDA a .test.mjs, que era o
ponto cego que deixou os orfaos apodrecerem. A extensao revelou +15 orfaos .mjs
(top-level + db/) alem dos 22 de feature-triage — TODOS religados via glob
tests/unit/**/*.test.mjs (171/171 pass). Um deles (encryption-error-handling)
codificava o contrato PRE-hardening (decrypt falho retornava ciphertext cru —
vazamento); alinhado ao contrato shipped (null + log) com comentario.
Gate: [test-discovery] OK — 2892 arquivos, 22 collectors, 60 orfaos congelados
(divida rastreada, shrink-only).
* ci: dedup heavy pipeline — compat to nightly, coverage folded into unit shards, i18n single job, draft-skip (#6215)
Pacote 2+3-ci do plano mestre testes+CI (aprovado 2026-07-04). O CI pesado rodava a
suite unit 4x por sync da release-PR (95 jobs, 208 min-maquina) e o ciclo v3.8.44
disparou 123 desses runs (88 cancelados, 0 uteis) porque a release-PR viva fica
aberta o ciclo inteiro.
- D2: matrizes Node 24/26 (build + 8 jobs de teste, ~28% do custo por run) saem do
per-sync e viram .github/workflows/nightly-compat.yml (diario, fail-fast off,
resolve a release ativa como o nightly-release-green, abre issue de tracking em
falha). ci.yml/ci-summary limpos das referencias.
- D3: a matrix Coverage Shard x8 (~18% do custo) e eliminada — o job test-unit roda
os MESMOS shards sob c8/NODE_V8_COVERAGE e sobe os artifacts coverage-shard-N; o
job de merge (test-coverage) so repontou needs (padrao do CI do nodejs/node).
timeout test-unit 15->25min pelo overhead de instrumentacao.
- D4: a matrix i18n de ~40 jobs de <1min (saturava sozinha os 20 slots de
concorrencia da conta Free) vira 1 job que itera os idiomas com ::group:: por
idioma e artifact unico com resultados nomeados por idioma (antes 40 result.txt
colidiam no merge-multiple do ci-summary).
- P3: jobs pesados pulam pull_requests DRAFT (predicado em 10 jobs-raiz; o resto
pula pela cadeia de needs; ci-summary segue rodando como sinal unico) — a skill
/generate-release ja abre a release-PR viva como draft e flipa ready no 0a.0a
(commit eb04fc5 no repo .agents/skills).
- C5 (CodeQL schedule) NAO incluido: bloqueado na acao do dono Settings -> CodeQL
Default->Advanced (documentado no proprio codeql.yml).
Validacao: js-yaml parse ok; check:workflows zizmor 156 < baseline 159 (ratchet
verde); validacao de execucao = workflow_dispatch deste ci.yml neste branch ate
package-artifact + electron-package-smoke verdes (registrada no PR).
* feat(quality): no-new-warnings por PR — ESLint bulk suppressions + lint-guard fork-condicional (Pacote 4) (#6218)
* feat(quality): no-new-warnings per PR via native ESLint bulk suppressions
Pacote 4 do plano mestre testes+CI (aprovado 2026-07-04). O ratchet de
eslintWarnings so rodava no CI pesado (release-PR) -> o drift acumulava invisivel
e explodia na release (+41/+37/+88 por ciclo, rebaselinado as cegas — historico
no proprio quality-baseline.json). Modelo novo (SonarSource Clean-as-You-Code +
ESLint bulk suppressions nativo >=9.24):
- config/quality/eslint-suppressions.json congela a divida existente por
arquivo+regra: 476 arquivos / 4.273 violacoes.
- npm run lint + lint-staged (pre-commit) + novo job lint-guard no quality.yml
rodam suppressions-aware: violacao NOVA fica vermelha NO PR que a introduz
(bulk suppressions ainda eleva estouros de baseline por arquivo a error).
- 3 regras warn promovidas a error em src/** (react-hooks/exhaustive-deps,
@next/next/no-img-element, import/no-anonymous-default-export) — divida
existente congelada, ocorrencia nova = erro imediato.
- collect-metrics mede sob o baseline congelado -> a metrica eslintWarnings
vira 'divida liquida nova' (~0 em regime); baseline apertado 4279->0 no mesmo
PR (exigencia do require-tighten). Aperto do ESTOQUE congelado: npx eslint .
--prune-suppressions na reconciliacao da release.
- Principio Zero: lint-guard usa continue-on-error para PR de FORK (report-only;
a campanha /green-prs aplica o fix via co-autoria) — bloqueante so para
branches internas, a origem real do drift.
Validacao: negativo (any novo em tests/) exit 1; negativo (img em src/, regra
promovida) exit 1; positivo escopado exit 0; baseline gerado por --suppress-all
no tip (tree inteiro passa por construcao); YAML js-yaml ok.
* fix(quality): clear the 6 residual warnings so lint-guard runs clean at --max-warnings 0
The committed baseline still let 6 warnings through the lint-guard gate:
5 now-unused inline eslint-disable directives (the file-level suppressions
made them redundant — removed via eslint --fix, suppressions regenerated to
absorb the re-exposed occurrences) and 1 anonymous default export in
tests/load/k6-soak.js (outside the src/** severity-override scope — named
the k6 scenario function instead).
Verified on the clean tree: lint-guard exit=0; any-canary (new 'const x: any'
in open-sse) exit=1 — the gate bites on NEW violations while the 4,273
frozen ones stay suppressed (476 files).
* fix(ci): lint-guard continue-on-error must be boolean on non-PR events
github.event.pull_request is undefined on workflow_dispatch — the bare property
expression made the job fail at plan time (run 28722888456: 4 jobs green, run red,
lint-guard never materialized). Guard with event_name check so the expression is
always boolean: PR de fork = report-only (Principio Zero), resto = blocking.
* docs(changelog): v3.8.45 bullets for the tests+quality+CI pipeline overhaul (#6214, #6215, #6218)
i18n CHANGELOG mirrors intentionally left to the release reconciliation
(release:sync-changelog-i18n), per cycle practice.
* fix(api): stabilize relay SSRF-guard binding for minified builds (#6149) (#6224)
* fix(mcp): forward extra context through static tool loops (#6178) (#6228)
* fix(services): 9Router embed route + pre-spawn port probe (#6205) (#6227)
* fix(backend): system-first memory injection for strict providers (#6135) (#6225)
* fix(auth): clear error for stale-key decryption failures (#6148) (#6226)
* fix(backend): record reasoning source for zero-metered reasoning models (#6187) (#6229)
* fix(providers): refresh stale NVIDIA NIM model registry (#6108) (#6223)
* fix(backend): distinct max_input_tokens for GPT-family models (#6191) (#6230)
* fix(oauth): extract keychain-import-only guard to restore file-size freeze (base-red) (#6158)
`src/app/api/oauth/[provider]/[action]/route.ts` grew to 959 lines, past its
frozen cap of 924 (`check:file-size` → Fast Quality Gates red on release/v3.8.44).
The growth came from #6054 (graceful 400 for keychain-import-only providers / zed):
a doc block, two Sets (KEYCHAIN_IMPORT_ONLY_PROVIDERS, OAUTH_FLOW_ACTIONS) and a
keychainImportOnlyResponse() helper, plus two duplicated guard blocks in GET/POST.
That is a cohesive, self-contained leaf, so extract it to a new
`keychainImportOnly.ts` exposing `keychainImportOnlyGuard(provider, action)`
(returns the 400 NextResponse or null). The two route callsites collapse to a
2-line guard each. route.ts: 959 -> 918 (< 924, freeze restored). No behavior
change.
Tests (Rule #8/#18):
- Existing tests/unit/oauth-keychain-import-only-6041.test.ts (route-level GET/POST
zed 400) still pass unchanged — behavior preserved.
- New tests/unit/oauth-keychain-import-only-guard.test.ts pins the extracted guard
in isolation (zed+flow -> 400, normal provider -> null, zed+non-flow -> null).
* fix(dashboard): stop model-test error freezing the page (React #31 object toast) (#6161)
Clicking 'test' on a provider model (e.g. a ClinePass flash model) could freeze
the entire dashboard. Root cause: POST /api/models/test returned an OBJECT in
`error` on the Zod-validation and invalid-JSON paths (`validation.error.format()`
/ a details object). The client does `notify.error(data.error)`, and
NotificationToast renders the message directly as a React child — an object throws
React #31 ('Objects are not valid as a React child'), crashing the tree = frozen
page instead of a toast.
Fixed in three layers (defense in depth):
1. Server (root cause): /api/models/test now returns a STRING `error` on every
path — flattens Zod issues to text, returns 'Invalid JSON body' for bad JSON.
2. Client: onTestModel funnels the response through extractApiErrorMessage() so any
object-shaped error is coerced to a string before notify.error.
3. Toast: NotificationToast coerces title/message via toToastText() — a resilient
catch-all so no future caller can freeze the page with a non-string.
Tests (Rule #18, both node:test / blocking suite):
- tests/unit/models-test-error-shape.test.ts — asserts STRING error on Zod-fail,
missing-field, and invalid-JSON (fails on the pre-fix route: 3/3 red -> green).
- tests/unit/notification-toast-coercion.test.ts — toToastText coercion matrix.
* fix(dashboard): remove the always-on Auto-Routing (combo) banner from the home page (#6164)
The blue "Auto-Routing Active — OmniRoute is automatically routing requests
using combo-based strategies" banner was rendered unconditionally on the home
page (`/home`, the default dashboard landing) — it did NOT reflect whether
auto-routing was actually active, and reappeared on every fresh browser / private
window / cleared localStorage (dismissal is stored per-browser). It added noise
to the landing page without conveying live state.
Remove it: drop the <AutoRoutingBanner /> usage + import from home/page.tsx and
delete the now-unused component and its test.
* fix(cline): force upstream streaming for Cline/ClinePass (streaming-only API) (#6165)
* fix(cline): force upstream streaming for Cline/ClinePass (streaming-only API)
Cline's API (api.cline.bot) only implements streaming (streamText). A
non-streaming request returns HTTP 500 "generateText is not implemented" (Claude
models) or HTTP 502 "empty response" (others). Live-verified on the VPS:
stream:true → works (STREAM_OK), stream:false → fails. This is why testing a Cline
model in the dashboard (the test button sends stream:false) failed.
Fix (reuses the existing isClaudeCodeCompatible mechanism, no new handler):
- Flag `cline` and `clinepass` registry entries with `forceStream: true`.
- In chatCore, OR `providerRequiresStreaming` into `upstreamStream` (line 1591)
so the upstream request always streams for these providers, while the client's
original `stream` intent still drives the response format. The existing
non-streaming branch (parseNonStreamingResponseBody) already accumulates the
upstream SSE and converts it back to JSON for stream:false clients — the same
path Claude-Code-compatible providers already use.
Tests (Rule #18): tests/unit/cline-force-stream.test.ts pins the registry flags +
resolveStreamFlag forcing behavior. Live VPS before/after recorded on the PR.
* fix(sse): cline forceStream must stream upstream only, keep client JSON
The #2081 wiring fed providerRequiresStreaming into resolveStreamFlag,
forcing the client-facing stream flag to true for forceStream providers.
That skips the if(!stream) branch that drains a forced upstream SSE and
converts it back to JSON, so a stream:false caller (model-test button,
plain JSON API) got STREAM_EARLY_EOF instead of a JSON body.
Keep providerRequiresStreaming only on upstreamStream (force upstream to
stream); leave the client-facing stream as the client sent it, so
readNonStreamingResponseBody accumulates the SSE into JSON. The promised
handleForcedSSEToJson (#2081 comment) was never implemented — this uses
the existing non-streaming SSE-buffering path (same as isClaudeCodeCompatible).
Live-verified on VPS: cline stream:true worked, stream:false failed.
* fix(providers): correct Kiro model catalog to real upstream ids (#6170)
* fix(providers): correct Kiro model catalog to real upstream ids
Kiro's API (generateAssistantResponse) returns 400 "Invalid model. Please
select a different model" for any id it does not recognize. The registry
exposed fabricated ids (copied from OmniRoute's own Anthropic catalog) that
Kiro never serves, so every call to them 400'd. Live-verified on the VPS:
Removed (400 Invalid model):
- auto-kiro (no "auto" model id — was sent verbatim upstream)
- claude-fable-5 (Kiro offers no Fable)
- claude-opus-4.8/4.7/4.6 (Kiro offers no Opus)
Corrected:
- claude-sonnet-4.6 -> claude-sonnet-4.5 (Kiro's Sonnet is 4.5; 4.5 -> 200)
Kept:
- claude-sonnet-5 (real Kiro model, plan-gated per account)
- claude-haiku-4.5, deepseek-3.2, glm-5, minimax-m2.5/m2.1,
qwen3-coder-next (all proven 200 on the VPS)
Aligns the free-model catalog and drops the orphaned auto-kiro price key.
Regression guard: tests/unit/kiro-catalog-real-models.test.ts (3/3).
Kiro cluster #6112/#6113/#6099.
* test(providers): align stale Kiro-catalog tests to the corrected upstream ids
The fabricated Kiro ids removed in the parent commit (claude-fable-5,
claude-opus-4.8/4.7/4.6, claude-sonnet-4.6) were still asserted as present by
three pre-existing tests, which encoded the bug:
- catalog-updates-v3x: now asserts Kiro does NOT expose Fable 5 / Opus (kept the
legit cc exposure) and guards the real claude-sonnet-4.5 pricing.
- model-family-fallback-notation: the dot-notation example moves from kiro/ to
anthropic/ (which genuinely serves Opus/Fable in dot notation) — coverage kept.
- provider-models-route: the Kiro local-catalog assertion now expects the real
Sonnet 5 / Sonnet 4.5 set and negatively guards the fabricated ids.
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
* fix(sse): surface ChatGPT-web image silent-drop as an accurate error (#6208)
When ChatGPT Web generates an image as an image_asset_pointer but the pointer
fails to resolve to a downloadable URL (unknown asset scheme, download 403/
expired, oversize), resolveImagePointers returned [] — indistinguishable from
'no image produced' — so the image-generation handler reported the misleading
502 'completed without returning image markdown'. The image genuinely existed
upstream; OmniRoute dropped it silently.
Fix: the executor flags x_image_resolution_failed when a pointer existed but
none resolved (and logs the unresolved asset scheme for follow-up), and the
handler surfaces a truthful 'generated but not retrievable' 502 instead of
'no image markdown'. Adds executorFactory DI for unit testing.
TDD: tests/unit/chatgpt-web-image-silentdrop.test.ts (red -> green), plus the
existing chatgpt-web / image-generation-handler suites stay green.
Reported via community triage (mesh escalated backlog).
* fix(dashboard): providers page data-timeout guard + live-ws standalone wiring (#6211)
* fix(dashboard): providers page data-timeout guard + live-ws standalone wiring
Captura de trabalho em progresso: timeout de dados na página de providers,
ajuste em ProviderLimits e instrumentation-node, com testes novos
(providers-page-data-timeout, live-ws-standalone-wiring).
* chore(quality): rebaseline ProviderLimits/index.tsx file-size (+6, #6211 data-timeout guard)
Cohesive fix growth from PR #6211's data-timeout guard on the quota page's two
first-paint fetches (1121->1127). The fast-path PR->release skips check:file-size,
so the bump lands with the PR. Justification recorded in file-size-baseline.json.
* fix(translator): strip reasoning param for nvidia z-ai/glm-5.2 (#6181)
* fix(translator): strip reasoning param for nvidia z-ai/glm-5.2
NVIDIA NIM OpenAI-compatible wrapper rejects the reasoning body field
and returns HTTP 400 "Unsupported parameter(s): `reasoning`".
Add a StripRule scoped to provider=nvidia + model /z-ai\/glm-5\.2/i.
Mirrors PR #6102 drop pattern (minimax-m2.7 thinking).
* docs(translator): tighten nvidia glm-5.2 strip-rule comment
* fix(translator): anchor glm-5.2 strip rule with word boundary
* fix: add nvidia to PROVIDER_TOOL_LIMITS (1536) to prevent tool truncation (#6177)
NVIDIA NIM API (nvidia/* models) silently truncates the tool list to 128
(the default MAX_TOOLS_LIMIT) because nvidia is not in PROVIDER_TOOL_LIMITS.
Tools beyond index 127 are dropped, causing agents to lose access to
critical tools like task, read, or high-index MCP tools.
Verified that NVIDIA NIM API supports up to 1536 tools by direct testing.
End-to-end confirmed: 198 tools sent, model successfully called tools at
indices 193, 195, and 197 (previously dropped by truncation to 128).
Follows the same pattern as #5563 (grok-cli: 200), integrated in v3.8.43.
* feat(provider): add Claude 5 Sonnet to Claude Web provider (#6200) (#6209)
* feat(provider): add Claude 5 Sonnet to Claude Web provider (#6200)
* test(providers): guard claude-web claude-sonnet-5 registry entry (#6209)
Adds the missing regression test the PR-test-policy gate requires: asserts the
claude-web registry exposes claude-sonnet-5 (Claude 5 Sonnet web) alongside the
existing 4.6 Sonnet / 4.5 Haiku entries. Fails on the release base (no entry).
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
---------
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
* fix(cli): detect POSIX auto-set HOSTNAME via os.hostname() to fix bind address (#6194) (#6195)
POSIX shells (bash/zsh) always set HOSTNAME to the machine name. The
.env loader uses first-wins semantics, so HOSTNAME=0.0.0.0 in .env is
silently ignored. This causes the server to bind to the LAN hostname
instead of 0.0.0.0, breaking localhost access and all internal
self-requests (ModelSync, HealthCheck, cloud sync).
The fix compares process.env.HOSTNAME against os.hostname(): when they
match, it's the POSIX auto-set signature and HOSTNAME is ignored.
OMNIROUTE_SERVER_HOST takes precedence as the dedicated escape hatch.
Backward compatibility is preserved: users who set HOSTNAME to a value
that doesn't match the machine name (e.g. Windows CMD/PowerShell users
with HOSTNAME in .env) will still have their value honoured.
Closes #6194
* feat(sse): surface Kiro adaptive-thinking reasoning as reasoning_content (#6213)
Kiro/CodeWhisperer streams Claude's reasoning as native `reasoningContentEvent`
frames when adaptive thinking is enabled, but the Kiro executor had no handler
for them, so `reasoning_effort` requests returned no reasoning. Wire it end to
end:
- translator (openai-to-kiro): enable Kiro thinking when the request carries
`reasoning_effort`, Anthropic `output_config.effort`, or a `thinking` block
(`{type:"enabled",budget_tokens}` mapped to a level; `{type:"adaptive"}`
defaults to `high`, matching Anthropic's documented default). Prepends the
Kiro `<thinking_mode>`/`<max_thinking_length>` prompt directive and sets
top-level `additionalModelRequestFields` ({output_config.effort,
thinking:{type:"adaptive"}, max_tokens}). Gated on `supportsReasoning`; drops
non-default temperature/top_p (rejected by adaptive-only Claude models).
- executor transformRequest: forward `additionalModelRequestFields` to AWS
(previously dropped by the strict top-level allowlist).
- executor stream loop: parse `reasoningContentEvent` (and reasoningText
variants) into the OpenAI reasoning_content channel.
Verified against the live CodeWhisperer stream: reasoningContentEvent frames are
returned, and larger effort/budget measurably deepens reasoning up to the model
cap. Unit tests cover the effort sources, forwarding, temp/top_p stripping, and
native reasoning-frame parsing.
* fix(chatcore): exempt opencode client from the default 128-tool truncation (#6193)
* fix(chatcore): exempt opencode client from the default 128-tool truncation
The default MAX_TOOLS_LIMIT (128) cap made truncateToolList blind-slice
tools.slice(0, 128), dropping opencode's built-in task tool and part of
its MCP tools when the inbound list exceeded 128 — so models routed
through OmniRoute could not launch subagents or reach all their tools.
Detect the opencode client (any x-opencode-* header, or 'opencode' in
the user-agent) and bypass ONLY the speculative 128 default. A known
provider ceiling (proactive PROVIDER_TOOL_LIMITS or a detected limit)
always wins and still truncates, even for opencode, so upstreams with
real hard limits (e.g. grok-cli 200) keep their 400-avoidance guard.
Non-opencode clients are unchanged.
- requestFormat.ts: add isOpencodeClient(headers, userAgent) + expose it
on resolveChatCoreRequestFormat.
- toolLimitDetector.ts: add getKnownToolLimit(); getEffectiveToolLimit
becomes getKnownToolLimit(provider) ?? DEFAULT_LIMIT (byte-identical
for existing callers).
- upstreamBody.ts: truncateToolList takes bypassDefaultToolLimit and
encodes the precedence; fix cosmetic debug-log count.
- chatCore.ts: thread the flag into prepareUpstreamBody.
- tests: extend tool-limit-detector unit tests.
* refactor(tools): accept nullable provider in tool-limit resolvers
Address PR review: widen getKnownToolLimit / getEffectiveToolLimit to
(provider: string | null | undefined) to match the call sites in
truncateToolList, and add unit assertions covering null/undefined
providers (getKnownToolLimit -> null, getEffectiveToolLimit -> 128).
---------
Co-authored-by: DKotsyuba <16292493+DKotsyuba@users.noreply.github.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
* fix(providers): refresh GitHub Copilot catalog (#6154)
* fix(providers): refresh github copilot catalog
Limit GitHub Copilot discovery to the curated supported model set and keep the provider cooldown panel client-safe by moving countdown formatting out of localDb.
* chore(quality): rebaseline providerPageHelpers.ts file-size (+13, #6154 copilot catalog)
The GitHub Copilot catalog refresh grows the provider-page model-section helper
(1021->1034). Fast-path PR->release skips check:file-size, so the bump lands with
the PR. Justification recorded in file-size-baseline.json.
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
---------
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
* chore(quality): rebaseline kiro-translator file-size debt from #6213
The #6213 kiro adaptive-thinking feature grew openai-to-kiro.ts (853->890) and its
test (1093->1234); the fast-path PR->release does not gate check:file-size on merge,
so the growth accumulated on the release tip. Rebaselined to keep the tip green.
Justification recorded in file-size-baseline.json.
* fix(doctor): resolve two false-positive WARNs (#6162) (#6163)
* fix(doctor): resolve two false-positive WARNs (#6162)
The `omniroute doctor` command reported two warnings on healthy installs
even though the underlying checks actually passed. Both came from the
doctor probing state that already worked; they looked like bugs but users
couldn't tell without manual digging.
Issue 1 — Server liveness HTTP 401
/api/health and /api/health/degradation both require the management
token. Doctor called them without auth → 401 → WARN, even when the
Next.js server was clearly alive and listening.
Fix: probe the configured health endpoint first; on 401/403, fall
back to a publicly served static asset (/favicon.ico) to confirm the
server is alive. WARN now only fires when both probes fail.
Issue 2 — CLI Tools '@/shared' import
tool-detector.ts (and 3 other cli-helper files) import @/shared/...
aliases that resolve via tsconfig.json paths. The CLI ships raw TS
source (no compile step) and runs through tsx, but tsx does not honor
tsconfig paths at runtime, and tsconfig-paths only hooks CJS
Module._resolveFilename while doctor uses ESM `import()`.
Fix: replace @/shared/... with relative imports in the 4 cli-helper
files. This is the same pattern these files already use for ./config-
generator/* imports. No new dependency, no architectural change, and
the fix doesn't regress Next.js itself which keeps using @/shared.
Verified on v3.8.43 (Node v24.17, Windows 11):
Before: 7 ok, 2 warning(s), 0 failure(s)
After: 8 ok, N warning(s), 0 failure(s)
where N accurately reflects which CLI tools are installed and
configured for OmniRoute (e.g. Hermes Agent installed but not
pointed at 20128 → 2 real warnings, not 1 false-positive).
Refs #6162
* fix(doctor): derive fallback URL from primary URL via new URL()
Per Gemini code-assist review feedback: the previous fallback constructed
the /favicon.ico URL from defaults (127.0.0.1:PORT) which ignored custom
host/port/protocol configurations supplied via:
- OMNIROUTE_DOCTOR_LIVENESS_URL
- OMNIROUTE_DOCTOR_HOST
- --liveness-url / --host CLI flags
Parse the primary URL with new URL() to preserve protocol, host, port, and
subpaths. The previous default-based fallback remains as a catch-all for
invalid primary URLs.
* test(doctor): add regression tests for #6162 fixes
Two new test files lock the fix and satisfy the PR Test Policy gate
("production code change without tests"):
- tests/unit/cli-helper-tool-detector-paths-6162.test.ts
Locks the @/shared → relative imports fix across all 4 cli-helper
files. Asserts (a) no @/shared alias remains in the cli-helper
sources, and (b) each file is importable at runtime via tsx/ESM,
which would have thrown "Cannot find package '@/shared'" before
the fix.
- tests/unit/cli-doctor-liveness-fallback-6162.test.ts
Locks the /favicon.ico fallback in doctor.mjs. Asserts the
fallback probe exists, derives its URL from the primary URL via
new URL() (per Gemini review feedback), and that the buggy
'Server responded with HTTP 401' WARN path is gone.
Both tests use only node:test + node:assert/strict so they slot into
the existing 'test' and 'test:unit' scripts with no extra config.
* test(doctor): fix primary.ok regex in fallback test
The earlier regex /primary\.ok\s*\?/ required a '?' immediately after,
but the actual doctor.mjs code uses a multi-line if-block:
if (primary.ok) {
return ok(...);
}
Use /\bprimary\.ok\b/ instead so the assertion matches the existing
branching.
---------
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
* fix(doubao-web): switch provider to Dola global (#6235)
* fix(doubao-web): switch provider to Dola global
* fix(doubao-web): use .dola.com cookie domain for s_v_web_id + rebaseline test
The Dola switch left s_v_web_id with a host-only "www.dola.com" domain, which fails
the token-source contract (domain must start with "." or "http") — the sibling
sessionid/ttwid cookies and the canonical cookieDomain already use ".dola.com", which
also matches www.dola.com. Also rebaselines web-cookie-providers-new.test.ts (850->890)
for the provider-switch regression cases.
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
---------
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
* fix(providers): register zed in OAuth PROVIDERS to fix Unknown provider error (#6041) (#6078)
Registers a minimal import_token entry for the existing Zed IDE keychain-import
provider so getProvider("zed") no longer throws "Unknown provider: zed" when the
UI probes the OAuth capability endpoint; generateAuthData returns { supported: false }.
Test runner fix: the regression test imported from "vitest" but lives in tests/unit/
(node:test territory, outside the vitest include globs) — it ran in no runner. Converted
to node:test + node:assert so it actually executes (8/8 green).
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
* fix(oauth): align zed in OAUTH_PROVIDER_IDS + config enum after #6078 merge
#6078 registered zed in the OAuth PROVIDERS registry but did not add it to the
constants PROVIDERS id map nor the oauth-providers-config enumeration test, leaving
that test red on the release tip (getProvider enumeration vs EXPECTED mismatch).
Adds ZED to the id constants + zed to EXPECTED_PROVIDER_KEYS/EXPECTED_CONFIG_BY_PROVIDER.
* fix(mitm): strip colons from macOS cert fingerprint before keychain match (#6134) (#6204)
fix(mitm): strip colons from macOS cert fingerprint before keychain match (#6134). Extracted testable macCertOutputHasFingerprint helper + regression guard. Thanks @rianonehub. Integrated into release/v3.8.45.
* docs(architecture): sync stale DB-layer counts (45+/55 → 95+/110+) in REPOSITORY_MAP, db-schema diagram and llm.txt (+42 i18n mirrors) (#6167)
docs(architecture): sync stale DB-layer counts (45+/55 → 95+/110+) across REPOSITORY_MAP, db-schema diagram, llm.txt + 42 i18n mirrors (#6167). Docs-only; check:docs-all passes locally on the reconstruction. Reds are pre-existing base-red drift on release/v3.8.45 (dast-smoke #6228; executor-kiro.test.ts eslint anys; changelog/package.json version drift) — none introduced by this PR. Integrated into release/v3.8.45.
* fix(api): count tool_use/tool_result/thinking blocks in count_tokens estimate (port from 9router#2337) (#6221)
fix(api): count tool_use/tool_result/thinking blocks in count_tokens estimate (#6221, port from 9router#2337). TDD-covered (6/6); typed the test casts to clear no-new-eslint. Reds are pre-existing base-red drift on release/v3.8.45 (dast-smoke #6228; executor-kiro.test.ts eslint anys; changelog/package.json version drift) — none introduced by this PR. Thanks @luweiCN. Integrated into release/v3.8.45.
* fix(antigravity): strip trailing assistant prefill turn for Vertex Claude models (#6114)
fix(antigravity): strip trailing assistant prefill for Vertex Claude models (#6114). TDD-covered (6/6), merged on TDD strength per owner. Reds are pre-existing base-red drift on release/v3.8.45. Thanks @anki1kr. Integrated into release/v3.8.45.
* fix(security): require management auth for mutable cloud routes (#6233) (#6233)
fix(security): require management auth for mutable cloud routes (#6233). Verified: 3 PR tests + full authz/route-guard suite 241/241 green. Thanks @vittoroliveira-dev. Integrated into release/v3.8.45.
* fix(dashboard): use connection.id (UUID) not connection.provider (category) in onboarding wizard href (issue #6144) (#6166)
refactor(dashboard): extract tested buildProviderDetailsHref helper for onboarding wizard (#6166). Behavioral #6144 fix already on tip via #6145; this lands the tested-helper hardening. Thanks @KooshaPari. Integrated into release/v3.8.45.
* feat(rankings): add 'Configured Only' filter to Free Provider Rankings page (#6245)
feat(rankings): add 'Configured Only' filter to Free Provider Rankings (#6245, closes #6150). 9/9 test green. Thanks @Iammilansoni. Integrated into release/v3.8.45.
* fix(i18n): add 118 missing Italian translations (#6212)
i18n(it): add 118 Italian translations (#6212). Audited net-additive (0 keys dropped, valid JSON). Thanks @serverless83. Integrated into release/v3.8.45.
* test(dashboard): realign #6145 onboarding-href guard to the #6166 helper refactor (#6270)
Realign the #6145 onboarding-href guard to the #6166 helper refactor (buildProviderDetailsHref). Test-only; unblocks the fast-path unit job across the open PR queue. Base-reds only (dast-smoke #6228, docs version-drift, executor-kiro anys). Integrated into release/v3.8.45.
* feat(providers): add Yuanbao (web) cookie-session provider (#6196) (#6256)
feat(providers): add Yuanbao (web) cookie-session provider (#6196). TDD-covered; base-reds only (dast-smoke #6228, docs version-drift, executor-kiro anys — #6145 guard fixed on tip via #6270). Integrated into release/v3.8.45.
* feat(providers): route built-in agentrouter through dynamic CC wire image (#6056) (#6255)
feat(providers): route built-in agentrouter through dynamic CC wire image (#6056). TDD-covered (agentrouter-cc-wire-image.test.ts). Base-reds only. Integrated into release/v3.8.45.
* feat(providers): bulk-add API keys for Cloudflare Workers AI (#6174) (#6254)
feat(providers): bulk-add API keys for Cloudflare Workers AI (#6174). Per-entry providerSpecificData (fixes shared-object reuse); TDD guard bulk-api-key-parser-cloudflare.test.ts. Base-reds only. Integrated into release/v3.8.45. (thanks @muflifadla38)
* feat(dashboard): routing/settings UX clarity — share %, Cloud Sync rename, base-URL override (#6147) (#6253)
feat(dashboard): routing/settings UX clarity (#6147) — effective share %, Cloud Sync→Remote Settings Sync rename, opt-in advanced base-URL override. TDD guard routing-settings-ux-6147.test.ts (6/6). Base-reds only. Integrated into release/v3.8.45.
* feat(combo): add option to disable session stickiness (#6168) (#6252)
feat(combo): add option to disable session stickiness (#6168) — per-combo/global, precedence config→settings→false (preserves #3825). TDD guard combo-disable-session-stickiness.test.ts (8/8). Base-reds only. Integrated into release/v3.8.45. (thanks @RCrushMe)
* feat(docker): OMNIROUTE_NO_SUDO env flag for root-less MITM cert trust (#6122) (#6249)
feat(docker): OMNIROUTE_NO_SUDO env flag for root-less MITM cert trust (#6122). resolveSudoSpawn strips sudo when set; argv-array spawn preserved (Hard Rule #13). TDD guard mitm-systemCommands-no-sudo.test.ts (5/5). Base-reds only. Integrated into release/v3.8.45. (thanks @powellnorma)
* feat(providers): add Requesty as an OpenAI-compatible gateway provider (#6120) (#6250)
feat(providers): add Requesty as an OpenAI-compatible gateway provider (#6120). Fixed the APIKEY_PROVIDERS count guard (167→168) the feature had missed. TDD guard requesty-provider.test.ts (4/4) + providers-constants-split (4/4). Base-reds only. Integrated into release/v3.8.45. (thanks @chirag127)
* fix(providers): remove deprecated MiMo v2 entries (#6248)
chore(providers): remove deprecated MiMo V2 catalog entries (superseded by V2.5); realign provider-catalog tests. Merged — thank you, @backryun! Base-reds only. Integrated into release/v3.8.45.
* fix(github-skills): add missing import, add unit tests, fix settings JSON parse (#6186)
feat(skills): GitHub skill-discovery subsystem — search/score/scan/import agent skills from GitHub, MCP tools (scope-gated) + /api/github-skills route (host-pinned api.github.com, sanitized errors). Merged — thank you, @Moseyuh333! Pre-merge fixes: passed toolDef.scopes so tools gate correctly under scope enforcement, routed the GET error path through sanitizeErrorMessage+500 (Hard Rule #12), and realigned the agent-skills count guards (catalog/routes/generator/mcp) for the new catalog entry. Base-reds only. Integrated into release/v3.8.45.
* Fix/5976 continued (#6216)
fix(combo): 5 streaming-path fixes (#5976) — locked-stream 500, error-frame-only-if-no-content, Gemini MALFORMED_RESPONSE→content_filter failover, correlationId substring, per-model-500 lockout skip + request-logger UI. Merged — thank you, @hartmark! Maintainer follow-up: releaseQualityClone cancels the abandoned quality-check tee branch (per-request memory) + regression test. All 41 combo/streaming quality tests green. Base-reds only. Integrated into release/v3.8.45.
* feat(dashboard): filter Free Provider Rankings by configured/available (#6150) (#6251)
feat(dashboard): configured-only / available-only filters on Free Provider Rankings (#6150) — server-side query params + tested lib helper; supersedes the client-side #6245 toggle with an available-only dimension. Lib logic 11/11 green; UI validated live on VPS. Base-reds only. Integrated into release/v3.8.45.
* ci: unblock test jobs from the Build gate (start at minute 0) (#6275)
test-unit x8, test-vitest, test-integration x2 and test-security all had
needs: build but never download the next-build artifact — the dependency
only serialized ~20min of Build wall-clock in front of every test run.
Switch them to needs: changes with the same skip condition Build uses
(docs-only PRs and drafts still skip), so the test chain
(test-unit -> test-coverage -> quality-gate -> sonarqube) now runs in
parallel with Build instead of after it.
Jobs that genuinely consume the artifact keep needs: build unchanged:
test-e2e x9, package-artifact, electron-package-smoke.
Expected effect on a full ci.yml run: critical path drops from
build + tests (~30min+) to max(build, tests) — roughly 15-20min saved
per run, no extra runner minutes beyond starting the same jobs earlier.
* ci(build): switch Next.js production build to Turbopack (1.9x faster) (#6273)
Next 16 ships Turbopack as the stable production bundler. Benchmarked on a
32-core box against the same tree: webpack 1035s -> Turbopack 539s (1.92x).
The build script already supports the switch via OMNIROUTE_USE_TURBOPACK=1
(scripts/build/build-next-isolated.mjs) and next.config.mjs already mirrors
the resolveAlias stubs in the turbopack block, so this only flips the CI env.
The webpack .build/next/cache actions/cache step is removed in the same
commit: Turbopack does not use the webpack cache dir (its persistent FS
cache is experimental and NOT enabled), so restoring ~0.5 GB per run would
be pure wasted download. Revert restores webpack + its cache together.
Validation: standalone output smoke-tested (server.js boots, health 200,
dashboard 307); the 428 build warnings are the known benign 'overly broad
file pattern' static-analysis notices for dynamic fs usage (covered at
runtime by outputFileTracingIncludes). Downstream e2e x9, package-artifact
and electron-package-smoke consume this artifact, so a green CI run here
validates the Turbopack artifact end-to-end. nightly-compat and npm-publish
stay on webpack until this PR proves out.
* feat(build): make Turbopack the default bundler for dev and build (#6283)
Turbopack (stable in Next 16) becomes the code default in the three entry
points that previously required an explicit OMNIROUTE_USE_TURBOPACK=1:
- scripts/build/build-next-isolated.mjs (production build)
- scripts/dev/run-next.mjs (dev server)
- scripts/dev/run-next-playwright.mjs (playwright dev runner)
OMNIROUTE_USE_TURBOPACK=0 remains the webpack escape hatch (Windows /
native-binding / bundler-compat issues), and only the documented '0'
opts out — junk values keep the default.
Benchmarked on this codebase (same tree, Next 16.2.9): webpack 1035s vs
Turbopack 539s on a 32-core box; ~20min vs 6min59s on ubuntu-latest.
Artifact validated end-to-end (standalone smoke + e2e/package-artifact/
electron-package-smoke CI jobs, Docker amd64+arm64 builds clean with the
v3.8.27 ImportTracer panic gone on 16.2.9).
TDD: tests/unit/build-bundler-default-turbopack.test.ts (new) +
run-next-playwright.test.ts extended with the unset-env default case;
both red before the flip, green after. ENVIRONMENT.md updated.
* feat(docker): build the image with Turbopack (v3.8.27 panic gone on Next 16.2.9) (#6285)
Flips the builder stage to OMNIROUTE_USE_TURBOPACK=1 and rewrites the stale
panic comment: the v3.8.27-era TurbopackInternalError ('entered unreachable
code: there must be a path to a root' in ImportTracer::get_traces) no longer
reproduces on Next 16.2.9.
Validation (2026-07-05, this exact Dockerfile, default
OMNIROUTE_BUILD_MEMORY_MB=4096, no overrides):
- amd64: docker build 659s, exit 0, zero panic/OOM strings in the full log,
container smoke-tested (/api/monitoring/health 200)
- arm64 (qemu): exit 0, zero panic strings
Webpack stays available as the escape hatch:
--build-arg / -e OMNIROUTE_USE_TURBOPACK=0. The V8 heap ceiling is kept:
Turbopack's compile is native Rust, but prerender/export still runs on V8.
* ci: opt-in self-hosted VPS runners for the release window (anti-queue) (#6284)
Adds the on-demand self-hosted runner plumbing for /generate-release:
- scripts/vps/release-runner-up.sh: starts the runner VM on Proxmox, waits
for >=1 'omni-release' runner to report online via the GitHub API, then
flips the USE_VPS_RUNNER repo variable to true. Any failure/timeout sets
it back to false and exits 1 so the caller falls back to hosted runners.
- scripts/vps/release-runner-down.sh: flips USE_VPS_RUNNER=false FIRST
(so no job gets scheduled onto a dying runner), then gracefully shuts
the VM down. Idempotent.
- ci.yml: build, test-unit x8 and test-vitest pick their runner
dynamically. Self-hosted is used ONLY when USE_VPS_RUNNER == 'true'
AND the event is own-origin (push/dispatch, or a PR whose head repo is
this repository). Fork PRs and the var's default/absent state always
fall back to ubuntu-latest.
Why: the Free plan caps hosted concurrency at 20 jobs; a release run
saturates it and queues. The benchmarked VPS (32-core, 4 runners) matches
hosted per-job times (unit ~8.5min/shard, build 9min turbo) but eliminates
the queue, which is the real release bottleneck (~30-50min on a busy day).
Scope is conservative: only the three job families benchmarked on the VPS;
e2e/electron/integration stay hosted (playwright/xvfb provisioning not
validated on the runner workspace yet).
* docs(changelog): restore v3.8.45/v3.8.44 sections eaten by the #6193 merge auto-resolve + CI-perf campaign bullets (#6273 #6275 #6283 #6284 #6285)
* fix(dashboard): null-guard connection in EditConnectionModal base-URL override (#6147) (#6287)
fix(dashboard): null-guard connection in EditConnectionModal (#6147) — fixes 'Cannot read properties of null (reading authType)' crash on every provider-detail page entry. TDD: connModals.test.tsx null-mount 9/9. Base-reds only. Integrated into release/v3.8.45.
* chore(release-green): clear test-masking + docs-all HARD reds for the v3.8.45 pre-flight
- test-masking: allowlist the 4 verified-legitimate assert reductions of the
cycle (#6248 MiMo V2 removal, #6170 Kiro catalog correction, #6154 Copilot
catalog refresh) and register the #6164 AutoRoutingBanner test deletion with
a real replacement guard (tests/unit/home-no-autorouting-banner.test.ts —
asserts the banner stays out of home/page.tsx and the component stays deleted)
- docs-sync: executors count 68 -> 73 in ARCHITECTURE.md + CODEBASE_DOCUMENTATION.md
- env-doc-sync: document OMNIROUTE_NO_SUDO (#6249/#6122) in .env.example +
docs/reference/ENVIRONMENT.md
* fix(quality): clear the cycle's 11 net-new ESLint errors + make validate-release-green suppressions-aware
- executor-kiro/save-call-log/call-logs-correlation tests: replace 15 'as any'
casts with typed shapes (net-new no-explicit-any errors from #6213/#6216);
prune the now-empty suppression entries so the frozen baseline stays exact
- github-skills + usage/call-logs routes: raw toLowerCase().includes() search
replaced by matchesSearch() (no-restricted-syntax — Turkish-safe search,
behavior covered by tests/unit/call-logs-correlation-substring.test.ts and
tests/unit/github-collector.test.ts)
- validate-release-green.mjs: run ESLint with --suppressions-location (match
the npm run lint contract — frozen debt is not a release red) and raise the
lint timeout 15->30min (a full pass takes ~14min alone; the 15min ceiling
expired under concurrent suite load and surfaced as 'could not parse eslint
json')
* fix(skills): generate the missing omni-github-skills registry entry + align catalog count tests
PR #6186 added omni-github-skills to the agent-skills catalog (API 22 -> 23)
but did not run the generator, so skills/omni-github-skills/SKILL.md never
existed and 6 integration assertions split between the old (42/43) and new
counts. Generated via scripts/skills/generate-agent-skills.mjs --apply and
aligned agent-skills-discovery to the real totals (43 = 23 API + 20 CLI;
handlers return 44 with config-codex-cli). 30/30 discovery+content tests green.
* fix(combo): restrict the #6216 empty-stream failover to truly empty bodies (restores #3399/#3685 contracts)
The 'streaming no recognized content' branch added by #6216 marked ANY
stream that ended without content deltas as invalid — sweeping in two
regression-guarded pass-through contracts: an empty stream terminated by an
explicit 'data: [DONE]' (#3399 context-cache protection) and an incomplete
Claude lifecycle (ping only, no message_start; #3685 — stream-readiness
timeout territory, not failover). Both unit guards were red on the branch
and green on main (86/86 vs 84/86).
The branch now fires only for a truly EMPTY body (zero bytes — the Gemini
HTTP-200-empty case that motivated #6216), tracked via sawAnyBytes. New
guard: '#5976 truly EMPTY streaming body (zero bytes) -> invalid for combo
failover'. 87/87 across both files.
Also in this pre-flight batch:
- agentSkillTools-mcp: api.have upper bound 22 -> 23 (the #6186 catalog
addition updated the totals but missed this bound)
- delete tests/unit/free-provider-rankings-configured-filter.test.ts:
#6251 (server-side configuredOnly/availableOnly) superseded the #6245
client-side toggle it pinned; replacement declared in the test-masking
allowlist (tests/unit/freeProviderRankings-filters.test.ts, 11/11)
* chore(quality): prune stale ESLint suppressions (4,273 -> 4,233)
Entries whose violations no longer exist (cleaned by cycle merges and the
pre-flight fixes) made 'npm run lint' exit 2 with 'suppressions left that do
not occur anymore'. Regenerated via --prune-suppressions; net-new policy
unchanged.
* fix(proxy): #6246 stop the v3.8.44 proxy IP-leak + over-deactivation regression (#6296)
Merged into release/v3.8.45. Reds pré-existentes classificados: dast-smoke (infra), check:file-size (drift de baseline de god-files congelados), e 3 testes de contagem agentSkills stale (43→44 já corrigidos no tip pelo #6186 — somem no squash sobre o tip). Núcleo do fix #6246 (proxy IP-leak + over-deactivation).
* fix(proxy): make "Test All" read-only + add bulk enable/disable (#6246) (#6299)
Merged into release/v3.8.45. Delta do par #6246 (Test-All read-only + bulk enable/disable), reconciliado com o núcleo #6296 já mergeado. CHANGELOG restaurado (ambos bullets do #6246 coexistem). Reds pré-existentes: dast-smoke (infra) + file-size drift.
* fix(resilience): evict sticky affinity on pinned-account failover (#6219) (#6231)
Merged into release/v3.8.45. Sticky affinity failover (#6219). Sincronizado com tip; CHANGELOG restaurado (base bullets preservados, net +1). Reds pré-existentes: dast-smoke + file-size drift.
* fix(sse): drop commentary-phase text in Responses passthrough (#6199) (#6232)
Merged into release/v3.8.45. Responses commentary-phase filter (#6199). Sincronizado; CHANGELOG restaurado net +1. Reds: dast-smoke + file-size drift.
* fix: bug-fix sweep — log path, AgentBridge DNS, opencode-go headers, GitLab Duo tools, M365 EDU (#6197 #6127 #6198 #5997 #6220 #6210) (#6234)
Merged into release/v3.8.45. Bug-fix sweep (#6197 log path, #6127/#6198 AgentBridge DNS, #6210 M365 EDU, #6220 GitLab Duo tools, #5997 opencode-go headers). 27 testes verdes. CHANGELOG restaurado net +5. Reds: dast-smoke + file-size drift.
* fix(docker): add id= to BuildKit cache mounts for strict builders (#6291)
Merged into release/v3.8.45. Dockerfile-only: explicit id= on BuildKit cache mounts (fixes strict-frontend parse error). Reds pré-existentes (dast-smoke/file-size drift) não relacionados a mudança de Dockerfile. Thanks @karimalsalah.
* fix(sse): strip zero-width markers from streamed tool-call arguments (follow-up to #5857) (#6292)
Merged into release/v3.8.45. Strip zero-width markers from streamed tool-call arguments (#5857 follow-up). Test 50/50. CHANGELOG reconciled net +1. Reds pré-existentes (dast-smoke/file-size drift). Thanks @DKotsyuba.
* ci(quality): merge-integrity fast-gates + pre-flight hermetic mode (#6300)
Merged into release/v3.8.45. CI merge-integrity fast-gates (changelog-integrity + agent-skills-sync) + pre-flight hermetic mode. O gate novo detectou 11 SKILL.md gerados fora de sync (drift pré-existente no tip: omni-api-keys endpoints, omni-github-skills do #6186) — regenerados neste PR, Merge-integrity GREEN no CI. Reds restantes base-red (dast-smoke/file-size drift/live-data flaky). Testes novos 17/17.
* fix(a2a): finish the #6186 catalog-count update — 3 hardcoded 22s left in production
#6186 added omni-github-skills (API 22 -> 23) and updated computeCoverage's
total, but left the old count hardcoded in the A2A layer: listCapabilities
metadata reported coverage.api.total 22 (type literal + value) and
SkillCoverageSchema pinned z.literal(22) — so the schema would REJECT the
correct runtime value. Aligned all three to 23 + the unit fixtures
(listCapabilities-a2a, agentSkills-schemas, 46/46 with agentSkillTools-mcp).
* fix(quality): type the 7 net-new 'as any' casts from #6292 (Lint red on the release tip)
#6292 rewrote/added zero-width-marker tests with 7 new 'as any' result casts,
pushing the file past its frozen suppression (84) — and when violations
exceed the suppressed count ESLint reports ALL of them, so the ci.yml Lint
job went red with 90 errors. The 7 new sites get typed shapes (delta /
arguments / choices / output accessors — same pattern as fecf888fd); the
pre-existing debt stays frozen at the exact new count (83). 50/50 tests
green, file lint-clean under the suppressions baseline.
* fix(api): Zod-validate POST /api/github-skills + document new gate envs + pin merge-integrity actions
Three latent heavy-CI reds surfaced by the VPS validation dispatch (the fast
path never runs these gates):
- t06 route-validation: POST /api/github-skills destructured request.json()
blind — a non-array 'targets' would .map-crash. Now validateBody(zod)
with defaults preserved (Hard Rule #7). Guard:
tests/unit/github-skills-route-validation.test.ts (4/4).
- env-doc-sync: document OMNIROUTE_SKIP_SYSTEM_TRUST (#6310) and the
changelog-integrity gate envs CHANGELOG_BASE_REF/ALLOW_CHANGELOG_REMOVALS
(#6300) in .env.example + ENVIRONMENT.md.
- zizmor ratchet: the new merge-integrity job's checkout/setup-node uses were
unpinned (+1 finding, 160 > baseline 159); pinned by SHA -> 158 (< baseline).
* fix(quality): clear the 2 remaining heavy-gate reds on the release tip
- check:error-helper: githubSkillTools.ts (#6186 wave) built MCP install
error results with raw err.message — routed through sanitizeErrorMessage()
(Hard Rule #12; 19/19 agentSkillTools tests green)
- check:mutation-test-coverage: tests/unit/combo-provider-cooldown-sibling.test.ts
(#6216) was missing from stryker.conf tap.testFiles — added so its mutant
kills count on nightly-mutation
* fix(security): 405 method-first for /api/keys/{id}/devices (dast-smoke QUERY check)
Schemathesis's newer unsupported-methods check (unpinned tool drift) sends
QUERY /api/keys/{id}/devices and demands 405 Method Not Allowed; the path had
no HIGH_RISK_METHOD_RULES entry, so the auth layer answered 401 first. Add
the devices rule (GET-only) so undocumented methods get a clean method-first
405 — same pattern as the v3.8.44 TRACE fix. TDD:
tests/unit/dast-method-not-allowed.test.ts gains the devices QUERY case (4/4).
* fix(mitm): test suite and CI must never mutate the OS trust store (OMNIROUTE_SKIP_SYSTEM_TRUST) (#6310)
Incident 2026-07-05 on the self-hosted release runner (VM 113): the
integration test 'POST /cert: installs trust when cert exists' exercised the
REAL install path, wrote a 105-byte fake PEM (FakeMITMCertForTestingOnly)
into /usr/local/share/ca-certificates and update-ca-certificates baked the
invalid entry into ca-certificates.crt — breaking ALL system TLS on the VM
(curl error 77, apt cert failures, and the intermittent gzip-corrupted
next-build artifacts that failed 6/9 e2e shards in run 28754447912). Hosted
runners are ephemeral, so the same mutation went unnoticed for months.
- installCert/uninstallCert: skip the OS dispatch under
OMNIROUTE_SKIP_SYSTEM_TRUST=1 — AFTER the input checks, so the #4546
environment-skip contract (missing file throws -> structured skip) and the
already-installed/not-installed early returns are preserved.
- installTproxyCa/uninstallTproxyCa: same guard, only when no run dep is
injected (DI'd tests keep exercising the full command sequence with mocks).
- tests/_setup/isolateDataDir.ts sets the env for every node:test process;
ci.yml/quality.yml/nightly-release-green.yml set it workflow-wide (e2e runs
the real app outside the test setup).
TDD: tests/unit/system-trust-test-guard.test.ts (guard exported to every test
process; guarded install resolves on a real file without touching the OS;
missing-file contract preserved). 82/82 across the affected cert/tproxy/
agent-bridge suites.
* ci(vps): hermetic nightly pre-flight on the release runner (descoped: e2e/integration/electron stay hosted) (#6305)
* ci(vps): extend the dynamic omni-release runner to e2e/integration/electron + nightly pre-flight
Completes the #6284 rollout to the jobs where the VPS pays the most:
- test-e2e (9 shards, 15-20min/shard hosted — dominated by setup, not tests),
test-integration (2 shards) and electron-package-smoke now pick the
self-hosted omni-release runner under the same gate: vars.USE_VPS_RUNNER
== 'true' AND own-origin (fork PRs never reach self-hosted).
- nightly-release-green (the release pre-flight) becomes runner-dynamic too:
on the VPS it runs in a clean env — no operator OMNIROUTE_API_KEY, no
local noauth CLIs — eliminating the machine-specific false positives that
dominated the 2026-07-05 pre-flight; passes --hermetic (no-op until the
#6300 validator lands, then belt-and-suspenders).
Validation plan (per operator request): release-runner-up.sh -> full ci.yml
workflow_dispatch on this branch exercising e2e/integration/electron on the
VPS (proves playwright --with-deps + xvfb on the runner) -> down.sh -> VM
off verified.
* fix(mitm): test suite and CI must never mutate the OS trust store (OMNIROUTE_SKIP_SYSTEM_TRUST)
Incident 2026-07-05 on the self-hosted release runner (VM 113): the
integration test 'POST /cert: installs trust when cert exists' exercised the
REAL install path, wrote a 105-byte fake PEM (FakeMITMCertForTestingOnly)
into /usr/local/share/ca-certificates and update-ca-certificates baked the
invalid entry into ca-certificates.crt — breaking ALL system TLS on the VM
(curl error 77, apt cert failures, and the intermittent gzip-corrupted
next-build artifacts that failed 6/9 e2e shards in run 28754447912). Hosted
runners are ephemeral, so the same mutation went unnoticed for months.
- installCert/uninstallCert: skip the OS dispatch under
OMNIROUTE_SKIP_SYSTEM_TRUST=1 — AFTER the input checks, so the #4546
environment-skip contract (missing file throws -> structured skip) and the
already-installed/not-installed early returns are preserved.
- installTproxyCa/uninstallTproxyCa: same guard, only when no run dep is
injected (DI'd tests keep exercising the full command sequence with mocks).
- tests/_setup/isolateDataDir.ts sets the env for every node:test process;
ci.yml/quality.yml/nightly-release-green.yml set it workflow-wide (e2e runs
the real app outside the test setup).
TDD: tests/unit/system-trust-test-guard.test.ts (guard exported to every test
process; guarded install resolves on a real file without touching the OS;
missing-file contract preserved). 82/82 across the affected cert/tproxy/
agent-bridge suites.
* ci(vps): descope — e2e/integration/electron stay on hosted runners; keep the hermetic nightly pre-flight dynamic
Validation verdict (runs 28754447912 + 28757670732, VM 113):
- e2e cannot run >1 per VM: both jobs bind port 20128 ('already used') — needs
a per-job port in the playwright runner before any VM rollout.
- concurrent ~1GB artifact downloads truncate on the VM uplink (gzip 'invalid
compressed data' — with 2 runners the e2e shard passed; corruption returned
at 4) — actions/download-artifact has no integrity retry here.
- integration shard 2 exceeded its 15-min timeout twice on the VM.
The VPS remains a win for whole-machine jobs: build/unit/vitest (already
dynamic via #6284) and nightly-release-green (single job, clean env, hermetic
pre-flight) — which this PR keeps.
* chore(quality): v3.8.45 cycle-close file-size rebaseline (Phase 0 drift absorption)
13 files grown by the cycle's merged PRs (#6216 streaming+request-logger UI;
#6251/#6253 dashboard UX) — legitimate merged-feature growth absorbed by the
release captain per the Phase 0 drift policy; all entries stay frozen (cannot
grow further). Justification key: _rebaseline_2026_07_06_v3845_release_close.
* chore(quality): v3.8.45 cycle-close cognitive/cyclomatic rebaseline (Phase 0 drift absorption)
cognitive 867->877 (+10), cyclomatic 2028->2035 (+7) — inherited cycle drift
measured by check:release-green (hermetic) on the release tip; the captain's
pre-flight fixes are gate/test/workflow changes (complexity-neutral).
Justification keys: _rebaseline_2026_07_06_v3845_release_close.
* docs(changelog): v3.8.45 reconciliation — fold Unreleased into the version section, 30 missing bullets, contributors hall
Phase 0a reconciliation (/generate-release): every commit since v3.8.44 now
has a bullet or a Maintenance rollup (82 commits; the only ref-less residues
are the bump/recording commits); #6193 bullet gains its PR ref; #6298
diagnosis credited (@subhansh-dev, landed via #6234); [3.8.44] header dated;
'### 🙌 Contributors' table injected (27 external + maintainer) + 42 i18n
mirrors resynced.
* chore(release): v3.8.45 — 2026-07-06
* fix(resilience): 502/503/504 keep the connection-unavailability path — only the exact 500 skips lockout (#5976 contract)
The #6216 branch used 'status >= 500', so a 503 on a per-model-quota /
openai-compatible provider returned cooldownMs 0 — no model lockout AND no
connection cooldown — hot-looping the failing upstream and breaking the
resilience-http-e2e 'priority combo falls back on 503' guard on the release
PR (the request after a 503 hit the same primary again). #6216's OWN unit
contract pins the narrow behavior ('Gemini 503 should NOT skip cooldown',
combo-provider-cooldown-sibling.test.ts) but computes the condition inline
instead of exercising auth.ts. Code aligned to the tested contract:
status === 500 skips (intermittent, not model-specific); 502/503/504 keep
the pre-#6216 model-lockout path. Validated: the failing integration test
flips red -> green locally (19s, deterministic before).
* fix(security): crypto-backed randomNumericId in doubao-web (CodeQL js/insecure-randomness)
The synthetic Dola device/web id was built from Math.random; CodeQL flags it
as insecure randomness in a security context. Not a secret, but
crypto.getRandomValues costs the same and closes alert #692 at the source.
Doubao executor tests 14/14 green. Alerts #693-695 (incomplete-url-substring
in unit-test asserts) dismissed as false positives per the v3.8.35 precedent
(Hard Rule #14).
* chore(quality): shave the #5976 fix comment back under the auth.ts file-size freeze (2447)
---------
Co-authored-by: Danny S <36470572+kanztu@users.noreply.github.com>
Co-authored-by: Luis Alejandro Vega <LuisAlejandroVega@redesprivadasvirtuales.com>
Co-authored-by: Milan Soni <123074437+Iammilansoni@users.noreply.github.com>
Co-authored-by: R. Beltran <rbeltran8000@gmail.com>
Co-authored-by: VXNCXNX <93332837+VXNCXNX@users.noreply.github.com>
Co-authored-by: Denis Kotsyuba <kocubads96@gmail.com>
Co-authored-by: DKotsyuba <16292493+DKotsyuba@users.noreply.github.com>
Co-authored-by: backryun <bakryun0718@proton.me>
Co-authored-by: Aris <arissunandar399@gmail.com>
Co-authored-by: Ankit <177378174+anki1kr@users.noreply.github.com>
Co-authored-by: Rian Priskanova <rian@evercore.technology>
Co-authored-by: Vittor Guilherme Borges de Oliveira <vittoroliveira.dev@gmail.com>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: serverless83 <35410475+serverless83@users.noreply.github.com>
Co-authored-by: Moseyuh333 <148680980+Moseyuh333@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: HenryHaniHannoush <karimmalsalah@gmail.com>
69 KiB
title, version, lastUpdated
| title | version | lastUpdated |
|---|---|---|
| OmniRoute Architecture | 3.8.40 | 2026-06-28 |
OmniRoute Architecture
🌐 Languages: 🇺🇸 English | 🇧🇷 Português (Brasil) | 🇪🇸 Español | 🇫🇷 Français | 🇮🇹 Italiano | 🇷🇺 Русский | 🇨🇳 中文 (简体) | 🇩🇪 Deutsch | 🇮🇳 हिन्दी | 🇹🇭 ไทย | 🇺🇦 Українська | 🇸🇦 العربية | 🇯🇵 日本語 | 🇻🇳 Tiếng Việt | 🇧🇬 Български | 🇩🇰 Dansk | 🇫🇮 Suomi | 🇮🇱 עברית | 🇭🇺 Magyar | 🇮🇩 Bahasa Indonesia | 🇰🇷 한국어 | 🇲🇾 Bahasa Melayu | 🇳🇱 Nederlands | 🇳🇴 Norsk | 🇵🇹 Português (Portugal) | 🇷🇴 Română | 🇵🇱 Polski | 🇸🇰 Slovenčina | 🇸🇪 Svenska | 🇵🇭 Filipino | 🇨🇿 Čeština
Last updated: 2026-06-28
Executive Summary
OmniRoute is a local AI routing gateway and dashboard built on Next.js.
It provides a single OpenAI-compatible endpoint (/v1/*) and routes traffic across multiple upstream providers with translation, fallback, token refresh, and usage tracking.
Core capabilities:
- OpenAI-compatible API surface for CLI/tools (237 providers, 73 executors)
- Request/response translation across provider formats
- Model combo fallback (multi-model sequence)
- Structured combo steps (
provider + model + connection) with runtime ordering bycompositeTiers - Account-level fallback (multi-account per provider)
- Quota preflight and quota-aware P2C account selection in the main chat path
- OAuth + API-key provider connection management (17 OAuth provider modules)
- Embedding generation via
/v1/embeddings(6 providers, 9 models) - Image generation via
/v1/images/generations(10+ providers, 20+ models) - Audio transcription via
/v1/audio/transcriptions(7 providers) - Text-to-speech via
/v1/audio/speech(10 providers) - Video generation via
/v1/videos/generations(ComfyUI + SD WebUI) - Music generation via
/v1/music/generations(ComfyUI) - Web search via
/v1/search(5 providers) - Moderations via
/v1/moderations - Reranking via
/v1/rerank - Think tag parsing (
<think>...</think>) for reasoning models - Response sanitization for strict OpenAI SDK compatibility
- Role normalization (developer→system, system→user) for cross-provider compatibility
- Structured output conversion (json_schema → Gemini responseSchema)
- Local persistence for providers, keys, aliases, combos, settings, pricing (26 DB modules)
- Usage/cost tracking and request logging
- Optional cloud sync for multi-device/state sync
- IP allowlist/blocklist for API access control
- Thinking budget management (passthrough/auto/custom/adaptive)
- Global system prompt injection
- Session tracking and fingerprinting
- Per-account enhanced rate limiting with provider-specific profiles
- Circuit breaker pattern for provider resilience
- Anti-thundering herd protection with mutex locking
- Signature-based request deduplication cache
- Domain layer: cost rules, fallback policy, lockout policy
- Context Relay: session handoff summaries for account rotation continuity
- Domain state persistence (SQLite write-through cache for fallbacks, budgets, lockouts, circuit breakers)
- Policy engine for centralized request evaluation (lockout → budget → fallback)
- Request telemetry with p50/p95/p99 latency aggregation
- Combo target telemetry and historical combo target health via
combo_execution_key/combo_step_id - Correlation ID (X-Request-Id) for end-to-end tracing
- Compliance audit logging with opt-out per API key
- Eval framework for LLM quality assurance
- Health dashboard with real-time provider circuit breaker status
- MCP Server (87 tools) with 3 transports (stdio/SSE/Streamable HTTP)
- A2A Server (JSON-RPC 2.0 + SSE) with skills and task lifecycle
- Memory system (extraction, injection, retrieval, summarization)
- Skills system (registry, executor, sandbox, built-in skills)
- MITM proxy with certificate management and DNS handling
- Prompt injection guard middleware
- Prompt compression pipeline with Caveman, RTK, stacked pipelines, compression combos, language packs, and analytics
- ACP (Agent Communication Protocol) registry
- Modular OAuth providers (16 individual modules under
src/lib/oauth/providers/) - Uninstall/full-uninstall scripts
- OAuth environment repair action
- WebSocket bridge for OpenAI-compatible WS clients (
/v1/ws) - Sync token management (issue/revoke, ETag-versioned config bundle download)
- GLM Thinking (
glmt) first-class provider preset - Hybrid token counting (provider-side
/messages/count_tokenswith estimation fallback) - Model alias auto-seeding (30+ cross-proxy dialect normalizations at startup)
- Safe outbound fetch with SSRF guard, private URL blocking, and configurable retry
- Cooldown-aware chat retries with configurable
requestRetryandmaxRetryIntervalSec - Runtime environment validation with Zod at startup
- Compliance audit v2 with pagination, provider CRUD events, and SSRF-blocked validation logging
Primary runtime model:
- Next.js app routes under
src/app/api/*implement both dashboard APIs and compatibility APIs - A shared SSE/routing core in
src/sse/*+open-sse/*handles provider execution, translation, streaming, fallback, and usage
Reference Diagrams
Canonical, version-controlled Mermaid sources for the v3.8.0 platform live in
docs/diagrams/. Two are reproduced below for orientation;
the rest are linked from their domain-specific guides.
Source: diagrams/request-pipeline.mmd
Source: diagrams/resilience-3layers.mmd — also linked from RESILIENCE_GUIDE.md and the
CLAUDE.mdresilience reference.
Scope and Boundaries
In Scope
- Local gateway runtime
- Dashboard management APIs
- Provider authentication and token refresh
- Request translation and SSE streaming
- Local state + usage persistence
- Optional cloud sync orchestration
Out of Scope
- Cloud service implementation behind
NEXT_PUBLIC_CLOUD_URL - Provider SLA/control plane outside local process
- External CLI binaries themselves (Claude CLI, Codex CLI, etc.)
Dashboard Surface (Current)
Main pages under src/app/(dashboard)/dashboard/:
/dashboard— quick start + provider overview/dashboard/endpoint— endpoint proxy + MCP + A2A + API endpoint tabs/dashboard/providers— provider connections and credentials/dashboard/combos— combo strategies, templates, step-based builder, model routing rules, manual persisted ordering/dashboard/auto-combo— Auto Combo Engine: scoring weights, mode packs, virtual factory presets, telemetry/dashboard/costs— cost aggregation and pricing visibility/dashboard/analytics— usage analytics, evaluations, combo target health/dashboard/limits— quota/rate controls/dashboard/cli-tools— CLI onboarding, runtime detection, config generation/dashboard/agents— detected ACP agents + custom agent registration/dashboard/cloud-agents— cloud-hosted agent tasks (Codex Cloud, Devin, Jules) and task lifecycle/dashboard/skills— A2A skill registry, sandbox execution, built-in skill catalog/dashboard/memory— persistent conversational memory inspection and retrieval/dashboard/webhooks— outbound webhook subscriptions, secret rotation, retry stats/dashboard/batch— batch job submission and progress/dashboard/cache— read-through and reasoning cache statistics, eviction controls/dashboard/playground— interactive chat playground against any configured combo/model/dashboard/changelog— in-app changelog viewer (rendersCHANGELOG.md)/dashboard/system— runtime diagnostics, version info, environment validation surface/dashboard/onboarding— first-run setup wizard for new installations/dashboard/media— image/video/music playground/dashboard/search-tools— search provider testing and history/dashboard/health— uptime, circuit breakers, rate limits, quota-monitored sessions/dashboard/logs— request/proxy/audit/console logs/dashboard/settings— system settings tabs (general, routing, combo defaults, etc.)/dashboard/context/caveman— Caveman compression rules, language packs, preview, and output mode/dashboard/context/rtk— RTK command-output filters, preview, and runtime safety settings/dashboard/context/combos— named compression pipelines assigned to routing combos/dashboard/translator— translator inspection and request format conversion preview/dashboard/audit— compliance audit log browser with pagination and structured metadata/dashboard/usage— per-request usage browser tied tousage_history/dashboard/compression— compression analytics, statistics, and pipeline assignment/dashboard/api-manager— API key lifecycle and model permissions
High-Level System Context
flowchart LR
subgraph Clients[Developer Clients]
C1[Claude Code]
C2[Codex CLI]
C3[OpenClaw / Droid / Cline / Continue / Roo]
C4[Custom OpenAI-compatible clients]
BROWSER[Browser Dashboard]
end
subgraph Router[OmniRoute Local Process]
API[V1 Compatibility API\n/v1/*]
DASH[Dashboard + Management API\n/api/*]
CORE[SSE + Translation Core\nopen-sse + src/sse]
DB[(storage.sqlite)]
UDB[(usage tables + log artifacts)]
end
subgraph Upstreams[Upstream Providers]
P1[OAuth Providers\nClaude/Codex/Gemini/Qwen/Qoder/GitHub/Kiro/Cursor/Antigravity]
P2[API Key Providers\nOpenAI/Anthropic/OpenRouter/GLM/Kimi/MiniMax\nDeepSeek/Groq/xAI/Mistral/Perplexity\nTogether/Fireworks/Cerebras/Cohere/NVIDIA]
P3[Compatible Nodes\nOpenAI-compatible / Anthropic-compatible]
end
subgraph Cloud[Optional Cloud Sync]
CLOUD[Cloud Sync Endpoint\nNEXT_PUBLIC_CLOUD_URL]
end
C1 --> API
C2 --> API
C3 --> API
C4 --> API
BROWSER --> DASH
API --> CORE
DASH --> DB
CORE --> DB
CORE --> UDB
CORE --> P1
CORE --> P2
CORE --> P3
DASH --> CLOUD
Core Runtime Components
1) API and Routing Layer (Next.js App Routes)
Main directories:
src/app/api/v1/*andsrc/app/api/v1beta/*for compatibility APIssrc/app/api/*for management/configuration APIs- Next rewrites in
next.config.mjsmap/v1/*to/api/v1/*
Important compatibility routes:
src/app/api/v1/chat/completions/route.tssrc/app/api/v1/messages/route.tssrc/app/api/v1/responses/route.tssrc/app/api/v1/models/route.ts— includes custom models withcustom: truesrc/app/api/v1/embeddings/route.ts— embedding generation (6 providers)src/app/api/v1/images/generations/route.ts— image generation (4+ providers incl. Antigravity/Nebius)src/app/api/v1/messages/count_tokens/route.tssrc/app/api/v1/providers/[provider]/chat/completions/route.ts— dedicated per-provider chatsrc/app/api/v1/providers/[provider]/embeddings/route.ts— dedicated per-provider embeddingssrc/app/api/v1/providers/[provider]/images/generations/route.ts— dedicated per-provider imagessrc/app/api/v1beta/models/route.tssrc/app/api/v1beta/models/[...path]/route.ts
Management domains:
- Auth/settings:
src/app/api/auth/*,src/app/api/settings/* - Providers/connections:
src/app/api/providers* - Provider nodes:
src/app/api/provider-nodes* - Custom models:
src/app/api/provider-models(GET/POST/DELETE) - Model catalog:
src/app/api/models/route.ts(GET) - Proxy config:
src/app/api/settings/proxy(GET/PUT/DELETE) +src/app/api/settings/proxy/test(POST) - OAuth:
src/app/api/oauth/* - Keys/aliases/combos/pricing:
src/app/api/keys*,src/app/api/models/alias,src/app/api/combos*,src/app/api/pricing - Usage:
src/app/api/usage/* - Sync/cloud:
src/app/api/sync/*,src/app/api/cloud/* - CLI tooling helpers:
src/app/api/cli-tools/* - IP filter:
src/app/api/settings/ip-filter(GET/PUT) - Thinking budget:
src/app/api/settings/thinking-budget(GET/PUT) - System prompt:
src/app/api/settings/system-prompt(GET/PUT) - Compression:
src/app/api/settings/compression,src/app/api/compression/*, andsrc/app/api/context/* - Sessions:
src/app/api/sessions(GET) - Rate limits:
src/app/api/rate-limits(GET) - Resilience:
src/app/api/resilience(GET/PATCH) — request queue, connection cooldown, provider breaker, wait-for-cooldown config - Resilience reset:
src/app/api/resilience/reset(POST) — reset provider breakers - Cache stats:
src/app/api/cache/stats(GET/DELETE) - Telemetry:
src/app/api/telemetry/summary(GET) - Budget:
src/app/api/usage/budget(GET/POST) - Fallback chains:
src/app/api/fallback/chains(GET/POST/DELETE) - Compliance audit:
src/app/api/compliance/audit-log(GET, with pagination + structured metadata) - Evals:
src/app/api/evals(GET/POST),src/app/api/evals/[suiteId](GET) - Policies:
src/app/api/policies(GET/POST) - Sync tokens:
src/app/api/sync/tokens(GET/POST),src/app/api/sync/tokens/[id](GET/DELETE) - Config bundle:
src/app/api/sync/bundle(GET, ETag-versioned snapshot of settings/providers/combos/keys) - WebSocket:
src/app/api/v1/ws/route.ts— Upgrade handler for OpenAI-compatible WS clients
2) SSE + Translation Core
Main flow modules:
- Entry:
src/sse/handlers/chat.ts - Core orchestration:
open-sse/handlers/chatCore.ts - Provider execution adapters:
open-sse/executors/* - Format detection/provider config:
open-sse/services/provider.ts - Model parse/resolve:
src/sse/services/model.ts,open-sse/services/model.ts - Account fallback logic:
open-sse/services/accountFallback.ts - Translation registry:
open-sse/translator/index.ts - Stream transformations:
open-sse/utils/stream.ts,open-sse/utils/streamHandler.ts - Usage extraction/normalization:
open-sse/utils/usageTracking.ts - Think tag parser:
open-sse/utils/thinkTagParser.ts - Embedding handler:
open-sse/handlers/embeddings.ts - Embedding provider registry:
open-sse/config/embeddingRegistry.ts - Image generation handler:
open-sse/handlers/imageGeneration.ts - Image provider registry:
open-sse/config/imageRegistry.ts - Response sanitization:
open-sse/handlers/responseSanitizer.ts - Role normalization:
open-sse/services/roleNormalizer.ts
Services (business logic):
- Account selection/scoring:
open-sse/services/accountSelector.ts - Context lifecycle management:
open-sse/services/contextManager.ts - IP filter enforcement:
open-sse/services/ipFilter.ts - Session tracking:
open-sse/services/sessionManager.ts - Request deduplication:
open-sse/services/signatureCache.ts - System prompt injection:
open-sse/services/systemPrompt.ts - Thinking budget management:
open-sse/services/thinkingBudget.ts - Wildcard model routing:
open-sse/services/wildcardRouter.ts - Rate limit management:
open-sse/services/rateLimitManager.ts - Circuit breaker:
src/shared/utils/circuitBreaker.ts - Context handoff:
open-sse/services/contextHandoff.ts— handoff summary generation and injection for context-relay strategy - Compression:
open-sse/services/compression/*— proactive compression before provider translation; includes Caveman rules, RTK filters, stacked pipelines, compression combos, stats, and validation - Codex quota fetcher:
open-sse/services/codexQuotaFetcher.ts— fetches Codex quota for context-relay handoff decisions - Cooldown-aware retry:
src/sse/services/cooldownAwareRetry.ts— per-model cooldown retries with configurablerequestRetry/maxRetryIntervalSec - Safe outbound fetch:
src/shared/network/safeOutboundFetch.ts— guarded provider/model fetch with SSRF guard, private-URL blocking, retry, and timeout - Outbound URL guard:
src/shared/network/outboundUrlGuard.ts— validates provider URLs against private/localhost CIDR ranges - Provider request defaults:
open-sse/services/providerRequestDefaults.ts— provider-levelmaxTokens,temperature,thinkingBudgetTokensdefaults - GLM provider constants:
open-sse/config/glmProvider.ts— shared GLM models, quota URLs, GLMT timeout/defaults - Antigravity upstream:
open-sse/config/antigravityUpstream.ts— base URL and discovery path constants - Codex client constants:
open-sse/config/codexClient.ts— versioned user-agent and client-version values - Model alias seed:
src/lib/modelAliasSeed.ts— seeds 30+ cross-proxy dialect aliases at startup
Domain layer modules:
- Cost rules/budgets:
src/domain/costRules.ts - Fallback policy:
src/domain/fallbackPolicy.ts - Combo resolver:
src/domain/comboResolver.ts - Lockout policy:
src/domain/lockoutPolicy.ts - Policy engine:
src/domain/policyEngine.ts— centralized lockout → budget → fallback evaluation - Error codes catalog:
src/shared/constants/errorCodes.ts - Request ID:
src/shared/utils/requestId.ts - Fetch timeout:
src/shared/utils/fetchTimeout.ts - Request telemetry:
src/shared/utils/requestTelemetry.ts - Compliance/audit:
src/lib/compliance/index.ts - Eval runner:
src/lib/evals/evalRunner.ts - Domain state persistence:
src/lib/db/domainState.ts— SQLite CRUD for fallback chains, budgets, cost history, lockout state, circuit breakers
OAuth provider modules (16 individual files under src/lib/oauth/providers/):
- Registry index:
src/lib/oauth/providers/index.ts - Individual providers:
claude.ts,codex.ts,gemini.ts,antigravity.ts,agy.ts,qoder.ts,qwen.ts,kimi-coding.ts,github.ts,kiro.ts,cursor.ts,kilocode.ts,cline.ts,windsurf.ts,gitlab-duo.ts,trae.ts - Thin wrapper:
src/lib/oauth/providers.ts— re-exports from individual modules
5) Embedded Services (v3.8.4)
OmniRoute can install, supervise, and route to locally-running AI tool processes called embedded services. Two are shipped in v3.8.4: 9Router and CLIProxyAPI.
Architecture layers:
- UI (
/dashboard/providers/services) — two-tab page with lifecycle controls, live log streaming, API key management, and (for 9Router) embedded native UI via an internal reverse proxy. - API (
/api/services/{name}/*) — 8 endpoints for 9Router, 7 for CLIProxyAPI, all classified LOCAL_ONLY (hard rule #17). A sharedGET /api/services/[name]/logsSSE endpoint serves both services. - Supervisor (
src/lib/services/) — genericServiceSupervisorclass wrapschild_process.spawn, holds a 5 MB ring buffer for SSE log streaming, a health probe loop, an atomic operation lock, and a SIGTERM→SIGKILL graceful shutdown.bootstrap.tswires all configured services at process start. - Provider/executor (
open-sse/executors/ninerouter.ts) — 9Router is exposed as a real provider. Models are prefixed9router/{sub}/{model}and synced every 5 min from 9Router's/v1/modelsendpoint.
Deep-dive: docs/frameworks/EMBEDDED-SERVICES.md
Major Subsystems (v3.8.0)
A. Auto Combo Engine
Auto Combo dynamically scores and picks routing targets at request time, rather than
relying on a static combo definition. It powers the auto/* model prefix family.
- Engine entry:
open-sse/services/autoCombo/(autoComboEngine.ts,scoringEngine.ts,virtualFactory.ts,modePacks.ts) - Resolver:
src/domain/comboResolver.ts(auto-detection ofauto/prefix) - Dashboard:
/dashboard/auto-combo - Telemetry:
auto_combo_decisionsSQLite table
Key capabilities:
- 17 routing strategies (priority, weighted, fill-first, round-robin, P2C, random,
least-used, cost-optimized, reset-aware, reset-window, headroom, strict-random,
auto, lkgp, context-optimized, context-relay, fusion, plus a fallback path) —
auto is the headline addition in v3.8.0;
fusion(panel fan-out + judge synthesis,open-sse/services/fusion.ts) is new in v3.8.36. - 9-factor scoring: cost, latency p95, success rate, quota headroom, lockout proximity, breaker state, recent failures, model availability, and tag affinity.
- Virtual factory materializes ephemeral combos when no matching named combo exists, sourcing candidates from healthy active provider connections.
- Auto prefixes:
auto/coding,auto/cheap,auto/fast,auto/offline,auto/smart,auto/lkgp— each backed by a tuned weight profile. - 4 mode packs: coding, fast, cheap, smart — shipped as preset weight configurations callable from the dashboard.
For full algorithmic detail (factor formulas, weight tuning), see
docs/routing/AUTO-COMBO.md.
B. Cloud Agents
Cloud Agents wraps third-party hosted code-agent platforms (Codex Cloud, Devin, Jules) behind a uniform DB-backed task lifecycle. All task creation/inspection endpoints require management authentication.
- Module root:
src/lib/cloudAgent/(baseAgent.ts,registry.ts,api.ts,types.ts,db.ts, plus per-agent subdirectories underagents/) - Per-agent implementations:
agents/codex/,agents/devin/,agents/jules/ - Public endpoints:
/api/v1/agents/tasks/*(list/create/get/cancel) - Management endpoints:
/api/cloud/*(provisioning, status, batch) - Dashboard:
/dashboard/cloud-agents - Storage:
cloud_agent_taskstable
For per-agent provisioning and OAuth specifics, see
docs/frameworks/CLOUD_AGENT.md.
C. Guardrails
The guardrails module is a hot-reloadable middleware layer that inspects requests and responses for PII, prompt injection, and unsafe vision content. Violations short-circuit the request with HTTP 503 plus a structured error code, allowing downstream callers to retry or branch.
- Module root:
src/lib/guardrails/(base.ts,registry.ts,piiMasker.ts,promptInjection.ts,visionBridge.ts,visionBridgeHelpers.ts) - Hot reload: registry watches for config changes and rebuilds the chain in place
- Wire-in points: chat handler entry, image generation handler, response sanitizer
- HTTP contract: violations surface as
503witherror.code = "GUARDRAIL_VIOLATION"
For ruleset authoring and threshold tuning, see
docs/security/GUARDRAILS.md.
D. Domain Layer
The src/domain/ namespace centralizes policy decisions so route handlers do not
have to assemble lockout/budget/fallback logic themselves.
- Policy engine:
src/domain/policyEngine.ts— single entry point for pre-execution evaluation (lockout → budget → fallback ordering) - Cost rules:
src/domain/costRules.ts - Fallback policy:
src/domain/fallbackPolicy.ts - Lockout policy:
src/domain/lockoutPolicy.ts - Tag-based routing:
src/domain/tagRouter.ts - Combo resolver:
src/domain/comboResolver.ts— resolves combo names, auto/* prefixes, and wildcard model targets to concrete execution plans - Connection/model rule joiner:
src/domain/connectionModelRules.ts - Model availability snapshots:
src/domain/modelAvailability.ts - Provider expiration tracking:
src/domain/providerExpiration.ts - Quota cache:
src/domain/quotaCache.ts - Degradation state:
src/domain/degradation.ts - Configuration audit:
src/domain/configAudit.ts - OmniRoute response metadata builder:
src/domain/omnirouteResponseMeta.ts - Assessment subsystem:
src/domain/assessment/— periodic evaluation jobs
E. Authorization Pipeline
The authorization pipeline classifies every incoming request and applies the appropriate policy chain before dispatch.
- Pipeline entry:
src/server/authz/pipeline.ts - Request classifier:
src/server/authz/classify.ts— distinguishes public compatibility routes from management routes - Public route inventory:
src/shared/constants/publicApiRoutes.ts - Policies:
src/server/authz/policies/— composable predicates (requireApiKey,requireManagement,requireFreshAuth, etc.) - Header utilities:
src/server/authz/headers.ts - Assertion helper:
src/server/authz/assertAuth.ts - Request context:
src/server/authz/context.ts
Public vs management routes are a hard boundary: agent/cooldown APIs and provider mutations require management auth (HTTP 401 if missing).
For the full route classification rules, see
docs/architecture/AUTHZ_GUIDE.md.
F. Workflow FSM and Task-Aware Router
A finite-state-machine driven router layered above combo selection to direct traffic based on the detected workflow stage (planning, execution, review) and background-task affinity.
- Workflow FSM:
open-sse/services/workflowFSM.ts - Task-aware router:
open-sse/services/taskAwareRouter.ts - Background task detector:
open-sse/services/backgroundTaskDetector.ts - Intent classifier:
open-sse/services/intentClassifier.ts
The FSM transitions feed into Auto Combo's scoring, biasing toward cheaper models for background/automation tasks and toward stronger models for interactive planning/review turns.
G. Provider-Specific Resilience
Several providers ship dedicated resilience and stealth modules that piggy-back on the global circuit breaker / connection cooldown / model lockout layers:
- Antigravity 429 engine:
open-sse/services/antigravity429Engine.ts(rotates identity, scrubs response headers, drives credits/version tracking viaantigravityCredits.ts,antigravityHeaderScrub.ts,antigravityHeaders.ts,antigravityIdentity.ts,antigravityObfuscation.ts,antigravityVersion.ts) - ModelScope quota policy:
open-sse/services/modelscopePolicy.ts - Claude Code CCH (Compatibility Channel Handshake):
open-sse/services/claudeCodeCCH.ts, plusclaudeCodeCompatible.ts,claudeCodeConstraints.ts,claudeCodeExtraRemap.ts,claudeCodeToolRemapper.ts - Claude Code fingerprint shaping:
open-sse/services/claudeCodeFingerprint.ts - Claude Code obfuscation:
open-sse/services/claudeCodeObfuscation.ts - ChatGPT TLS client:
open-sse/services/chatgptTlsClient.ts(curl-impersonate style for ChatGPT-Web sessions) - ChatGPT image cache:
open-sse/services/chatgptImageCache.ts
For the full stealth playbook and operational guidance, see
docs/security/STEALTH_GUIDE.md.
H. Webhooks, Reasoning Cache, Read Cache
- Webhooks — outbound dispatch for provider/account/task events.
- Dispatcher:
src/lib/webhookDispatcher.ts - Storage:
webhooksSQLite table (viasrc/lib/db/webhooks.ts) - Dashboard:
/dashboard/webhooks(subscriptions, secrets, retry history) - For event taxonomy and retry semantics, see
docs/frameworks/WEBHOOKS.md.
- Dispatcher:
- Reasoning Cache — replayable reasoning blocks for providers that emit
thinking tokens (Claude, GLMT, etc.) so consecutive turns can skip re-thinking.
- DB layer:
src/lib/db/reasoningCache.ts - Service layer:
open-sse/services/reasoningCache.ts - For replay semantics, see
docs/routing/REASONING_REPLAY.md.
- DB layer:
- Read Cache — short-lived response cache keyed by signature and used to
collapse identical retries from broken upstream SDKs.
- DB layer:
src/lib/db/readCache.ts - Stats endpoint:
GET /api/cache/stats, dashboard at/dashboard/cache
- DB layer:
3) Persistence Layer
Primary state DB (SQLite):
- Core infra:
src/lib/db/core.ts(better-sqlite3, migrations, WAL) - Re-export facade:
src/lib/localDb.ts(thin compatibility layer for callers) - file:
${DATA_DIR}/storage.sqlite(or$XDG_CONFIG_HOME/omniroute/storage.sqlitewhen set, else~/.omniroute/storage.sqlite) - entities (tables + KV namespaces): providerConnections, providerNodes, modelAliases, combos, apiKeys, settings, pricing, customModels, proxyConfig, ipFilter, thinkingBudget, systemPrompt
Usage persistence:
- facade:
src/lib/usageDb.ts(decomposed modules insrc/lib/usage/*) - SQLite tables in
storage.sqlite:usage_history,call_logs,proxy_logs - optional file artifacts remain for compatibility/debug (
${DATA_DIR}/log.txt,${DATA_DIR}/call_logs/,<repo>/logs/...) - legacy JSON files are migrated to SQLite by startup migrations when present
Domain State DB (SQLite):
src/lib/db/domainState.ts— CRUD operations for domain state- Tables (created in
src/lib/db/core.ts):domain_fallback_chains,domain_budgets,domain_cost_history,domain_lockout_state,domain_circuit_breakers - Write-through cache pattern: in-memory Maps are authoritative at runtime; mutations are written synchronously to SQLite; state is restored from DB on cold start
4) Auth + Security Surfaces
- Dashboard cookie auth:
src/proxy.ts,src/app/api/auth/login/route.ts - API key generation/verification:
src/shared/utils/apiKey.ts - Provider secrets persisted in
providerConnectionsentries - Outbound proxy support via
open-sse/utils/proxyFetch.ts(env vars) andopen-sse/utils/networkProxy.ts(configurable per-provider or global) - SSRF / outbound URL guard:
src/shared/network/outboundUrlGuard.ts— blocks private/loopback/link-local ranges for all provider calls - Runtime env validation:
src/lib/env/runtimeEnv.ts— Zod schema for all environment variables, surfaced as startup errors/warnings - Sync tokens:
src/lib/db/syncTokens.ts— scoped tokens for config bundle download endpoints; backed bysync_tokensSQLite table (migration024_create_sync_tokens.sql) - WebSocket handshake auth:
src/lib/ws/handshake.ts— validates WS upgrade requests via API key or session cookie
5) Cloud Sync
- Scheduler init:
src/lib/initCloudSync.ts,src/shared/services/initializeCloudSync.ts,src/shared/services/modelSyncScheduler.ts - Periodic task:
src/shared/services/cloudSyncScheduler.ts - Periodic task:
src/shared/services/modelSyncScheduler.ts - Control route:
src/app/api/sync/cloud/route.ts
Request Lifecycle (/v1/chat/completions)
sequenceDiagram
autonumber
participant Client as CLI/SDK Client
participant Route as /api/v1/chat/completions
participant Chat as src/sse/handlers/chat
participant Core as open-sse/handlers/chatCore
participant Model as Model Resolver
participant Auth as Credential Selector
participant Exec as Provider Executor
participant Prov as Upstream Provider
participant Stream as Stream Translator
participant Usage as usageDb
Client->>Route: POST /v1/chat/completions
Route->>Chat: handleChat(request)
Chat->>Model: parse/resolve model or combo
alt Combo model
Chat->>Chat: iterate combo models (handleComboChat)
end
Chat->>Auth: getProviderCredentials(provider)
Auth-->>Chat: active account + tokens/api key
Chat->>Core: handleChatCore(body, modelInfo, credentials)
Core->>Core: detect source format
Core->>Core: translate request to target format
Core->>Exec: execute(provider, transformedBody)
Exec->>Prov: upstream API call
Prov-->>Exec: SSE/JSON response
Exec-->>Core: response + metadata
alt 401/403
Core->>Exec: refreshCredentials()
Exec-->>Core: updated tokens
Core->>Exec: retry request
end
Core->>Stream: translate/normalize stream to client format
Stream-->>Client: SSE chunks / JSON response
Stream->>Usage: extract usage + persist history/log
Combo + Account Fallback Flow
flowchart TD
A[Incoming model string] --> B{Is combo name?}
B -- Yes --> C[Load combo models sequence]
B -- No --> D[Single model path]
C --> E[Try model N]
E --> F[Resolve provider/model]
D --> F
F --> G[Select account credentials]
G --> H{Credentials available?}
H -- No --> I[Return provider unavailable]
H -- Yes --> J[Execute request]
J --> K{Success?}
K -- Yes --> L[Return response]
K -- No --> M{Fallback-eligible error?}
M -- No --> N[Return error]
M -- Yes --> O[Mark account unavailable cooldown]
O --> P{Another account for provider?}
P -- Yes --> G
P -- No --> Q{In combo with next model?}
Q -- Yes --> E
Q -- No --> R[Return all unavailable]
Fallback decisions are driven by open-sse/services/accountFallback.ts using status codes and error-message heuristics. Combo routing adds one extra guard: provider-scoped 400s such as upstream content-block and role-validation failures are treated as model-local failures so later combo targets can still run.
OAuth Onboarding and Token Refresh Lifecycle
sequenceDiagram
autonumber
participant UI as Dashboard UI
participant OAuth as /api/oauth/[provider]/[action]
participant ProvAuth as Provider Auth Server
participant DB as localDb
participant Test as /api/providers/[id]/test
participant Exec as Provider Executor
UI->>OAuth: GET authorize or device-code
OAuth->>ProvAuth: create auth/device flow
ProvAuth-->>OAuth: auth URL or device code payload
OAuth-->>UI: flow data
UI->>OAuth: POST exchange or poll
OAuth->>ProvAuth: token exchange/poll
ProvAuth-->>OAuth: access/refresh tokens
OAuth->>DB: createProviderConnection(oauth data)
OAuth-->>UI: success + connection id
UI->>Test: POST /api/providers/[id]/test
Test->>Exec: validate credentials / optional refresh
Exec-->>Test: valid or refreshed token info
Test->>DB: update status/tokens/errors
Test-->>UI: validation result
Refresh during live traffic is executed inside open-sse/handlers/chatCore.ts via executor refreshCredentials().
Cloud Sync Lifecycle (Enable / Sync / Disable)
sequenceDiagram
autonumber
participant UI as Endpoint Page UI
participant Sync as /api/sync/cloud
participant DB as localDb
participant Cloud as External Cloud Sync
participant Claude as ~/.claude/settings.json
UI->>Sync: POST action=enable
Sync->>DB: set cloudEnabled=true
Sync->>DB: ensure API key exists
Sync->>Cloud: POST /sync/{machineId} (providers/aliases/combos/keys)
Cloud-->>Sync: sync result
Sync->>Cloud: GET /{machineId}/v1/verify
Sync-->>UI: enabled + verification status
UI->>Sync: POST action=sync
Sync->>Cloud: POST /sync/{machineId}
Cloud-->>Sync: remote data
Sync->>DB: update newer local tokens/status
Sync-->>UI: synced
UI->>Sync: POST action=disable
Sync->>DB: set cloudEnabled=false
Sync->>Cloud: DELETE /sync/{machineId}
Sync->>Claude: switch ANTHROPIC_BASE_URL back to local (if needed)
Sync-->>UI: disabled
Periodic sync is triggered by CloudSyncScheduler when cloud is enabled.
Data Model and Storage Map
erDiagram
SETTINGS ||--o{ PROVIDER_CONNECTION : controls
PROVIDER_NODE ||--o{ PROVIDER_CONNECTION : backs_compatible_provider
PROVIDER_CONNECTION ||--o{ USAGE_ENTRY : emits_usage
SETTINGS {
boolean cloudEnabled
number stickyRoundRobinLimit
boolean requireLogin
string password_hash
string fallbackStrategy
json rateLimitDefaults
json providerProfiles
}
PROVIDER_CONNECTION {
string id
string provider
string authType
string name
number priority
boolean isActive
string apiKey
string accessToken
string refreshToken
string expiresAt
string testStatus
string lastError
string rateLimitedUntil
json providerSpecificData
}
PROVIDER_NODE {
string id
string type
string name
string prefix
string apiType
string baseUrl
}
MODEL_ALIAS {
string alias
string targetModel
}
COMBO {
string id
string name
string[] models
}
API_KEY {
string id
string name
string key
string machineId
}
USAGE_ENTRY {
string provider
string model
number prompt_tokens
number completion_tokens
string connectionId
string timestamp
}
CUSTOM_MODEL {
string id
string name
string providerId
}
PROXY_CONFIG {
string global
json providers
}
IP_FILTER {
string mode
string[] allowlist
string[] blocklist
}
THINKING_BUDGET {
string mode
number customBudget
string effortLevel
}
SYSTEM_PROMPT {
boolean enabled
string prompt
string position
}
Physical storage files:
- primary runtime DB:
${DATA_DIR}/storage.sqlite - request log lines:
${DATA_DIR}/log.txt(compat/debug artifact) - structured call payload archives:
${DATA_DIR}/call_logs/ - optional translator/request debug sessions:
<repo>/logs/...
Deployment Topology
flowchart LR
subgraph LocalHost[Developer Host]
CLI[CLI Tools]
Browser[Dashboard Browser]
end
subgraph ContainerOrProcess[OmniRoute Runtime]
Next[Next.js Server\nPORT=20128]
Core[SSE Core + Executors]
MainDB[(storage.sqlite)]
UsageDB[(usage tables + log artifacts)]
end
subgraph External[External Services]
Providers[AI Providers]
SyncCloud[Cloud Sync Service]
end
CLI --> Next
Browser --> Next
Next --> Core
Next --> MainDB
Core --> MainDB
Core --> UsageDB
Core --> Providers
Next --> SyncCloud
Module Mapping (Decision-Critical)
Route and API Modules
src/app/api/v1/*,src/app/api/v1beta/*: compatibility APIssrc/app/api/v1/providers/[provider]/*: dedicated per-provider routes (chat, embeddings, images)src/app/api/providers*: provider CRUD, validation, testingsrc/app/api/provider-nodes*: custom compatible node managementsrc/app/api/provider-models: custom model management (CRUD)src/app/api/models/route.ts: model catalog API (aliases + custom models)src/app/api/oauth/*: OAuth/device-code flowssrc/app/api/keys*: local API key lifecyclesrc/app/api/models/alias: alias managementsrc/app/api/combos*: fallback combo managementsrc/app/api/pricing: pricing overrides for cost calculationsrc/app/api/settings/proxy: proxy configuration (GET/PUT/DELETE)src/app/api/settings/proxy/test: outbound proxy connectivity test (POST)src/app/api/usage/*: usage and logs APIssrc/app/api/sync/*+src/app/api/cloud/*: cloud sync and cloud-facing helperssrc/app/api/cli-tools/*: local CLI config writers/checkerssrc/app/api/settings/ip-filter: IP allowlist/blocklist (GET/PUT)src/app/api/settings/thinking-budget: thinking token budget config (GET/PUT)src/app/api/settings/system-prompt: global system prompt (GET/PUT)src/app/api/settings/compression: global compression settings (GET/PUT)src/app/api/compression/*: compression preview, rule metadata, and language packssrc/app/api/context/caveman/config: Caveman settings alias (GET/PUT)src/app/api/context/rtk/*: RTK config, filter catalog, test endpoint, and raw-output recoverysrc/app/api/context/combos*: compression combo CRUD and routing-combo assignmentssrc/app/api/context/analytics: compression analytics aliassrc/app/api/sessions: active session listing (GET)src/app/api/rate-limits: per-account rate limit status (GET)src/app/api/sync/tokens: sync token CRUD (GET/POST)src/app/api/sync/tokens/[id]: sync token get/delete (GET/DELETE)src/app/api/sync/bundle: config bundle download (GET, ETag versioning)src/app/api/v1/ws: WebSocket upgrade handler for OpenAI-compatible WS clients
Routing and Execution Core
src/sse/handlers/chat.ts: request parse, combo handling, account selection loopopen-sse/handlers/chatCore.ts: translation, executor dispatch, retry/refresh handling, stream setupopen-sse/executors/*: provider-specific network and format behavior
Translation Registry and Format Converters
open-sse/translator/index.ts: translator registry and orchestration- Request translators:
open-sse/translator/request/*(9 modules —antigravity-to-openai,claude-to-gemini,claude-to-openai,gemini-to-openai,openai-responses,openai-to-claude,openai-to-cursor,openai-to-gemini,openai-to-kiro) - Response translators:
open-sse/translator/response/*(8 modules —claude-to-openai,cursor-to-openai,gemini-to-claude,gemini-to-openai,kiro-to-openai,openai-responses,openai-to-antigravity,openai-to-claude) - Helpers:
open-sse/translator/helpers/*(8 modules —claudeHelper,geminiHelper,geminiToolsSanitizer,maxTokensHelper,openaiHelper,responsesApiHelper,schemaCoercion,toolCallHelper) - Format constants:
open-sse/translator/formats.ts - Bootstrap and registry:
open-sse/translator/bootstrap.ts,open-sse/translator/registry.ts - Image-format helpers:
open-sse/translator/image/
Persistence
src/lib/db/*: persistent config/state and domain persistence on SQLitesrc/lib/localDb.ts: compatibility re-export for DB modulessrc/lib/usageDb.ts: usage history/call logs facade on top of SQLite tables
Provider Executor Coverage (Strategy Pattern)
Each provider has a specialized executor extending BaseExecutor (in open-sse/executors/base.ts), which provides URL building, header construction, retry with exponential backoff, credential refresh hooks, and the execute() orchestration method.
| Executor | Provider(s) | Special Handling |
|---|---|---|
DefaultExecutor |
OpenAI, Claude, Gemini, Qwen, OpenRouter, GLM, Kimi, MiniMax, DeepSeek, Groq, xAI, Mistral, Perplexity, Together, Fireworks, Cerebras, Cohere, NVIDIA, etc. | Dynamic URL/header config per provider |
AntigravityExecutor |
Google Antigravity | Custom project/session IDs, Retry-After parsing, 429 obfuscation |
AzureOpenAIExecutor |
Azure OpenAI | Deployment-based routing, api-version query enforcement |
BlackboxWebExecutor |
Blackbox AI (web-mode) | Web-session reverse with TLS fingerprint emulation |
ChatGPTWebExecutor |
ChatGPT web | TLS client + session cookie management (chatgptTlsClient.ts) |
ClaudeIdentityExecutor |
Claude.ai (CCH path) | Constraint + tool-remap pipelines, fingerprint shaping |
CliProxyApiExecutor |
CLIProxyAPI-compatible providers | Custom auth and protocol handling |
CloudflareAiExecutor |
Cloudflare Workers AI | Account ID injection, Neurons-based usage tracking |
CodexExecutor |
OpenAI Codex | Injects system instructions, forces reasoning effort |
CommandCodeExecutor |
Command Code | OAuth + per-session header rotation |
CursorExecutor |
Cursor IDE | ConnectRPC protocol, Protobuf encoding, request signing via checksum |
DevinCliExecutor |
Devin CLI | Devin task lifecycle bridging via cloud agent module |
GithubExecutor |
GitHub Copilot | Copilot token refresh, VSCode-mimicking headers |
GitlabExecutor |
GitLab Duo | GitLab OAuth + project-scoped routing |
GlmExecutor |
Z.AI GLM (incl. glmt preset) |
Thinking-budget aware, GLMT preset constants |
GrokWebExecutor |
xAI Grok web | Web-session reverse, mode selection (think/standard) |
KieExecutor |
KIE | Custom token issuance with rotating session anchors |
KiroExecutor |
AWS CodeWhisperer/Kiro | AWS EventStream binary format → SSE conversion |
MuseSparkWebExecutor |
Muse Spark (web) | Web-session reverse with image-message bridging |
NlpCloudExecutor |
NLP Cloud | Provider-specific request body shape |
OpenCodeExecutor |
OpenCode | AI SDK compatible provider setup |
PerplexityWebExecutor |
Perplexity web | Web-session reverse for chat continuation |
PetalsExecutor |
Petals distributed inference | Decentralized swarm routing |
PollinationsExecutor |
Pollinations AI | No API key required, rate-limited requests |
PuterExecutor |
Puter | Browser-based provider integration |
QoderExecutor |
Qoder AI | PAT and OAuth support, multi-model free tier |
VertexExecutor |
Google Vertex AI | Service account auth, region-based endpoints |
WindsurfExecutor |
Windsurf (Codeium) | Codeium OAuth + session token refresh |
All other providers (including custom compatible nodes) use the DefaultExecutor.
Provider Compatibility Matrix
Note: The matrix below is a representative sample of the 237 registered providers in OmniRoute v3.8.0. For the canonical and continuously-updated list, refer to
docs/reference/PROVIDER_REFERENCE.md(auto-generated) or the source of truth atsrc/shared/constants/providers.ts(Zod-validated at load).
| Provider | Format | Auth | Stream | Non-Stream | Token Refresh | Usage API |
|---|---|---|---|---|---|---|
| Claude | claude | API Key / OAuth | ✅ | ✅ | ✅ | ⚠️ Admin only |
| Gemini | gemini | API Key / OAuth | ✅ | ✅ | ✅ | ⚠️ Cloud Console |
| Antigravity | antigravity | OAuth | ✅ | ✅ | ✅ | ✅ Full quota API |
| OpenAI | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Codex | openai-responses | OAuth | ✅ forced | ❌ | ✅ | ✅ Rate limits |
| GitHub Copilot | openai | OAuth + Copilot Token | ✅ | ✅ | ✅ | ✅ Quota snapshots |
| Cursor | cursor | Custom checksum | ✅ | ✅ | ❌ | ❌ |
| Kiro | kiro | AWS SSO OIDC | ✅ (EventStream) | ❌ | ✅ | ✅ Usage limits |
| Qwen | openai | OAuth | ✅ | ✅ | ✅ | ⚠️ Per request |
| Qoder | openai | OAuth / PAT | ✅ | ✅ | ✅ | ⚠️ Per request |
| Kilo Code | openai | OAuth | ✅ | ✅ | ✅ | ❌ |
| Cline | openai | OAuth | ✅ | ✅ | ✅ | ❌ |
| Kimi Coding | openai | OAuth | ✅ | ✅ | ✅ | ❌ |
| OpenRouter | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| GLM/Kimi/MiniMax | claude | API Key | ✅ | ✅ | ❌ | ❌ |
| DeepSeek | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Groq | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| xAI (Grok) | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Mistral | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Perplexity | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Together AI | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Fireworks AI | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Cerebras | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Cohere | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| NVIDIA NIM | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Cloudflare AI | openai | API Token + Acct ID | ✅ | ✅ | ❌ | ❌ |
| Pollinations | openai | None (no key) | ✅ | ✅ | ❌ | ❌ |
| Scaleway AI | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| LongCat | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Ollama Cloud | openai | API Key (optional) | ✅ | ✅ | ❌ | ❌ |
| HuggingFace | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Nebius | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| SiliconFlow | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Hyperbolic | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Vertex AI | gemini | Service Account | ✅ | ✅ | ✅ | ⚠️ Cloud Console |
| Puter | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Command Code | openai | OAuth | ✅ | ✅ | ✅ | ⚠️ Per request |
| Z.AI / GLM | openai | API Key / OAuth | ✅ | ✅ | ❌ | ❌ |
| GLMT (preset) | claude | API Key | ✅ | ✅ | ❌ | ⚠️ Per request |
| Kimi Coding | openai | OAuth / API Key | ✅ | ✅ | ✅ | ❌ |
| KIE | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Windsurf | openai | OAuth (Codeium) | ✅ | ✅ | ✅ | ⚠️ Per request |
| GitLab Duo | openai | OAuth (GitLab) | ✅ | ✅ | ✅ | ❌ |
| Devin CLI | openai | OAuth | ✅ | ✅ | ✅ | ✅ Task API |
| Codex Cloud | openai-responses | OAuth | ✅ | ❌ | ✅ | ✅ Rate limits |
| Jules | openai | OAuth | ✅ | ✅ | ✅ | ✅ Task API |
| AgentRouter | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| ChatGPT-Web | openai | Session cookie + TLS | ✅ | ✅ | ❌ | ❌ |
| Grok-Web | openai | Session cookie | ✅ | ✅ | ❌ | ❌ |
| Perplexity-Web | openai | Session cookie | ✅ | ✅ | ❌ | ❌ |
| BlackBox-Web | openai | Session cookie + TLS | ✅ | ✅ | ❌ | ❌ |
| Muse-Spark-Web | openai | Session cookie | ✅ | ✅ | ❌ | ❌ |
| ModelScope | openai | API Key | ✅ | ✅ | ❌ | ⚠️ Quota policy |
| BazaarLink | openai | API Key | ✅ | ✅ | ❌ | ❌ |
| Petals | openai | None | ✅ | ✅ | ❌ | ❌ |
| Qoder | openai | OAuth / PAT | ✅ | ✅ | ✅ | ⚠️ Per request |
| OpenCode (Go/Zen) | openai | OAuth | ✅ | ✅ | ✅ | ❌ |
| CLIProxyAPI | openai | Custom | ✅ | ✅ | ❌ | ❌ |
Format Translation Coverage
Detected source formats include:
openaiopenai-responsesclaudegemini
Target formats include:
- OpenAI chat/Responses
- Claude
- Gemini/Antigravity envelope
- Kiro
- Cursor
Translations use OpenAI as the hub format — all conversions go through OpenAI as intermediate:
Source Format → OpenAI (hub) → Target Format
Translations are selected dynamically based on source payload shape and provider target format.
Additional processing layers in the translation pipeline:
- Response sanitization — Strips non-standard fields from OpenAI-format responses (both streaming and non-streaming) to ensure strict SDK compliance
- Role normalization — Converts
developer→systemfor non-OpenAI targets; mergessystem→userfor models that reject the system role (GLM, ERNIE) - Think tag extraction — Parses
<think>...</think>blocks from content intoreasoning_contentfield - Structured output — Converts OpenAI
response_format.json_schemato Gemini'sresponseMimeType+responseSchema
Supported API Endpoints
| Endpoint | Format | Handler |
|---|---|---|
POST /v1/chat/completions |
OpenAI Chat | src/sse/handlers/chat.ts |
POST /v1/messages |
Claude Messages | Same handler (auto-detected) |
POST /v1/responses |
OpenAI Responses | open-sse/handlers/responsesHandler.ts |
POST /v1/embeddings |
OpenAI Embeddings | open-sse/handlers/embeddings.ts |
GET /v1/embeddings |
Model listing | API route |
POST /v1/images/generations |
OpenAI Images | open-sse/handlers/imageGeneration.ts |
GET /v1/images/generations |
Model listing | API route |
POST /v1/providers/{provider}/chat/completions |
OpenAI Chat | Dedicated per-provider with model validation |
POST /v1/providers/{provider}/embeddings |
OpenAI Embeddings | Dedicated per-provider with model validation |
POST /v1/providers/{provider}/images/generations |
OpenAI Images | Dedicated per-provider with model validation |
POST /v1/messages/count_tokens |
Claude Token Count | API route |
GET /v1/models |
OpenAI Models list | API route (chat + embedding + image + custom models) |
GET /api/models/catalog |
Catalog | All models grouped by provider + type |
POST /v1beta/models/*:streamGenerateContent |
Gemini native | API route |
GET/PUT/DELETE /api/settings/proxy |
Proxy Config | Network proxy configuration |
POST /api/settings/proxy/test |
Proxy Connectivity | Proxy health/connectivity test endpoint |
GET/POST/DELETE /api/provider-models |
Provider Models | Provider model metadata backing custom and managed available models |
Bypass Handler
The bypass handler (open-sse/utils/bypassHandler.ts) intercepts known "throwaway" requests from Claude CLI — warmup pings, title extractions, and token counts — and returns a fake response without consuming upstream provider tokens. This is triggered only when User-Agent contains claude-cli.
Request Logging and Artifacts
The older file-based request logger (open-sse/utils/requestLogger.ts) is retained only for
legacy compatibility. The current runtime contract uses:
APP_LOG_TO_FILE=truefor application and audit logs written under<repo>/logs/- SQLite-backed call log records in
call_logs ${DATA_DIR}/call_logs/YYYY-MM-DD/...artifacts when the call log pipeline is enabled
Failure Modes and Resilience
1) Account/Provider Availability
- connection cooldown on retryable upstream failures
- account fallback before failing request
- combo model fallback when current model/provider path is exhausted
2) Token Expiry
- pre-check and refresh with retry for refreshable providers
- 401/403 retry after refresh attempt in core path
3) Stream Safety
- disconnect-aware stream controller
- translation stream with end-of-stream flush and
[DONE]handling - usage estimation fallback when provider usage metadata is missing
4) Cloud Sync Degradation
- sync errors are surfaced but local runtime continues
- scheduler has retry-capable logic, but periodic execution currently calls single-attempt sync by default
5) Data Integrity
- SQLite schema migrations and auto-upgrade hooks at startup
- legacy JSON → SQLite migration compatibility path
6) SSRF / Outbound URL Guard
src/shared/network/outboundUrlGuard.tsblocks all private/loopback/link-local target URLs before they reach provider executors- Provider model discovery and validation routes use
src/shared/network/safeOutboundFetch.tswhich applies the guard before every outbound request - Guard errors surface as
URL_GUARD_BLOCKEDwith HTTP 422 and are logged to the compliance audit trail viaproviderAudit.ts
Observability and Operational Signals
Runtime visibility sources:
- console logs from
src/sse/utils/logger.ts - per-request usage aggregates in SQLite (
usage_history,call_logs,proxy_logs) - four-stage detailed payload captures in SQLite (
request_detail_logs) whensettings.detailed_logs_enabled=true - textual request status log in
log.txt(optional/compat) - optional application log files under
logs/whenAPP_LOG_TO_FILE=true - optional request artifacts under
${DATA_DIR}/call_logs/when the call log pipeline is enabled - dashboard usage endpoints (
/api/usage/*) for UI consumption
Detailed request payload capture stores up to four JSON payload stages per routed call:
- raw request received from the client
- translated request actually sent upstream
- provider response reconstructed as JSON; streamed responses are compacted to the final summary plus stream metadata
- final client response returned by OmniRoute; streamed responses are stored in the same compact summary form
Security-Sensitive Boundaries
- JWT secret (
JWT_SECRET) secures dashboard session cookie verification/signing - Initial password bootstrap (
INITIAL_PASSWORD) should be explicitly configured for first-run provisioning - API key HMAC secret (
API_KEY_SECRET) secures generated local API key format - Provider secrets (API keys/tokens) are persisted in local DB and should be protected at filesystem level
- Cloud sync endpoints rely on API key auth + machine id semantics
Environment and Runtime Matrix
Environment variables actively used by code:
- App/auth:
JWT_SECRET,INITIAL_PASSWORD - Storage:
DATA_DIR - Compatible node behavior:
ALLOW_MULTI_CONNECTIONS_PER_COMPAT_NODE - Optional storage base override (Linux/macOS when
DATA_DIRunset):XDG_CONFIG_HOME - Security hashing:
API_KEY_SECRET,MACHINE_ID_SALT - Logging:
APP_LOG_TO_FILE,APP_LOG_RETENTION_DAYS,CALL_LOG_RETENTION_DAYS - Sync/cloud URLing:
NEXT_PUBLIC_BASE_URL,NEXT_PUBLIC_CLOUD_URL - Outbound proxy:
HTTP_PROXY,HTTPS_PROXY,ALL_PROXY,NO_PROXYand lowercase variants - SOCKS5 feature flags:
ENABLE_SOCKS5_PROXY,NEXT_PUBLIC_ENABLE_SOCKS5_PROXY - Platform/runtime helpers (not app-specific config):
APPDATA,NODE_ENV,PORT,HOSTNAME
Known Architectural Notes
usageDbandlocalDbshare the same base directory policy (DATA_DIR->XDG_CONFIG_HOME/omniroute->~/.omniroute) with legacy file migration./api/v1/route.tsdelegates to the same unified catalog builder used by/api/v1/models(src/app/api/v1/models/catalog.ts) to avoid semantic drift.- Request logger writes full headers/body when enabled; treat log directory as sensitive.
- Cloud behavior depends on correct
NEXT_PUBLIC_BASE_URLand cloud endpoint reachability. - The
open-sse/directory is published as the@omniroute/open-ssenpm workspace package. Source code imports it via@omniroute/open-sse/...(resolved by Next.jstranspilePackages). File paths in this document still use the directory nameopen-sse/for consistency. - Charts in the dashboard use Recharts (SVG-based) for accessible, interactive analytics visualizations (model usage bar charts, provider breakdown tables with success rates).
- E2E tests use Playwright (
tests/e2e/), run vianpm run test:e2e. Unit tests use Node.js test runner (tests/unit/), run vianpm run test:unit. Source code undersrc/is TypeScript (.ts/.tsx); theopen-sse/workspace remains JavaScript (.js). - Settings page is organized into 7 tabs: General, Appearance, AI, Security, Routing, Resilience, Advanced. The Resilience page only configures request queue, connection cooldown, provider breaker, and wait-for-cooldown behavior; live breaker runtime state is shown on the Health page.
- Context Relay strategy (
context-relay) is split across two layers:combo.tsdecides if a handoff should be generated,chat.tsinjects the handoff after account resolution. Handoff data lives incontext_handoffsSQLite table. This split is intentional because onlychat.tsknows whether the actual account changed. - Proxy enforcement is now comprehensive:
tokenHealthCheck.tsresolves proxy per connection,/api/providers/validateusesrunWithProxyContext, andproxyFetch.tsusesundici.fetch()to maintain dispatcher compatibility on Node 22. - Node.js runtime policy detection:
/api/settings/require-loginreturnsnodeVersionandnodeCompatiblefields. The login page renders a warning banner when the runtime falls outside the supported secure Node.js lines.
Operational Verification Checklist
- Build from source:
npm run build - Build Docker image:
docker build -t omniroute . - Start service and verify:
GET /api/settingsGET /api/v1/models- CLI target base URL should be
http://<host>:20128/v1whenPORT=20128