* fix(quality): resolve net-new lint errors and allowlist #9343 assert rewrite Two `no-explicit-any` errors landed with #9407 and #9320 after the suppressions inventory was generated. Project policy is to fix new violations rather than freeze them, so both are typed instead: - #9407: `executor as unknown as Record<string, unknown>` - #9320: `(k: { name?: string })` Also allowlists the net-assert reduction in web-tools-translation-2820 (39->35). #9343 inverted the contract — bare JSON must no longer be promoted to tool_calls without an explicit <tool> envelope — so the tests were rewritten to assert non-promotion, which costs fewer asserts than validating a promoted object. More restrictive, not weaker. * fix(quality): raise integration ceiling to 40min and unpin codex-cli version in test The integration gate's 20min ceiling killed a healthy run: measured 22m08s hermetic on an idle 16-core box (935 tests across 112 files, strictly serial at --test-concurrency=1 because ~16 of them bind a port or share a DB). The "~3-10min" estimate in the code was stale by ~3x. 40min keeps the ceiling's real purpose — turning a genuine hang into a visible failure — without failing a long-but-healthy suite. Also fixes a base-red in chat-pipeline:564c204efebumped DEFAULT_CODEX_CLIENT_VERSION to 0.146.0 but the User-Agent assertion still pinned 0.144.1. The line two above already read the constant via getCodexClientVersion(); this one duplicated the literal. Deriving it from the same source stops the next bump from breaking the test again. * fix(ratelimit): re-arm Bottleneck reservoir heartbeat after updateSettings Bottleneck 2.19.5 (frozen upstream dependency, no release since 2019) has a bug in LocalDatastore#_startHeartbeat() (node_modules/bottleneck/lib/ LocalDatastore.js:29,56): the guard `if (this.heartbeat == null && ...)` only (re)creates the periodic reservoir-refresh interval the first time it runs. Every later call -- including the one updateSettings() itself triggers internally -- falls into the else branch and does clearInterval(this.heartbeat) WITHOUT resetting the reference back to null. Because the stale reference sticks around, every future _startHeartbeat() call keeps taking the same dead else branch: the periodic reservoir refresh is gone forever after the first manual updateSettings() call on a limiter. Every limiter created by this file starts with a live heartbeat (buildLimiterDefaults() always sets reservoirRefreshInterval/ reservoirRefreshAmount), so the very first updateFromHeaders() / updateFromResponseBody() / applyRequestQueueSettings() call against a limiter permanently kills its refresh. In production this wedges the request queue once the reservoir hits 0: an auto-enrolled apikey connection accumulates its default 60 requests, the reservoir zeroes, the queue freezes for ~120s, the watchdog fires a synthetic 502 (RATE_LIMIT_QUEUE_WEDGED), the connection cools down and gets excluded from weighted combo pools -- turning a configured 70/30 split into ~50/50. Add applyLimiterSettings(), a module-local wrapper around limiter.updateSettings() that nulls the stale heartbeat reference and re-invokes _startHeartbeat() afterward so it takes the "start a fresh interval" branch again. Route all 5 updateSettings() call sites through it (updateAllLimiterSettings, both updateFromHeaders() branches, loadPersistedLimits(), and updateFromResponseBody()). updateAllLimiterSettings is now async and awaited by its two callers (initializeRateLimits, applyRequestQueueSettings); the sync call sites use the existing trackAsyncOperation() fire-and-forget tracking pattern. tests/integration/combo-matrix/weighted.test.ts is the E2E proof: the "weighted: 70/30" case now passes with zero WEDGED/RATE_LIMIT_QUEUE/502 log lines across 200 sequential requests (previously the wedge/recovery cycle inflated its runtime and skewed the distribution toward ~50/50). Refs #8213 * fix(tests): remove stray TDD probes committed by accident inf4e93f339dThree TDD repro/probe test files landed on the release tip viaf4e93f339d(docs: add management authentication terminology guide, files from a worktree. Each file is a pre-fix TDD probe that belongs to a *different*, still-in-flight fix branch/PR and duplicates a file path that PR already owns and will properly update on merge: - tests/unit/authz/probe-9033-repro.test.ts: probe for #9033 (IP blacklist direct-connection bypass). 3/4 asserts fail against this tree (D1, D2, Bonus — all assert the not-yet-implemented target behavior); D3 passes (pre-existing behavior). Owned by PR #9385 (open, unmerged), which modifies this exact path. - tests/unit/repro-8522.test.ts: probe for #8522 (absolute file-size baseline reds innocent PRs on inherited drift). First test fails against this tree's evaluateFileSizes (still absolute-only); second (sanity: real growth still flags) passes. #8522 is actually CLOSED upstream — PR #9355 merged the real fix into release/v3.8.50 today (2026-08-05T15:53Z) modifying this exact path — but this branch's merge-base with release/v3.8.50 (6b0e11e378) predates that merge, so the fix has not synced into this tree yet. - tests/unit/repro-8956.test.ts: probe for #8956 (resolveProjectRoot stops at synthetic Next.js standalone package.json). First test fails against this tree; second (sanity: named package.json still resolves) passes. Owned by PR #9354 (open, unmerged), which modifies this exact path. Each deleted file's real implementation + passing version already exists in its owning PR and will land normally through that PR's own merge — deleting the premature copy here does not lose any coverage. No config/quality/test-masking-allowlist.json entry was added: the _deletedWithReplacement schema only supports `replacement` (a test file that must already exist in this tree's HEAD — none does, the real versions live in the unmerged sibling PRs above) or `sourceRemoved` (production files that must be absent from HEAD — they are not, none of the three issues are implemented in this tree). Neither shape fits an "owned by an in-flight sibling PR" deletion, so the CI test-masking gate will flag these 3 deletions for mandatory human review on this branch's next PR diff against release/v3.8.50 — flagged for the owner rather than inventing a new allowlist shape. Refs #9033, #8522, #8956, #7786 * fix(tests): align 8189-classifier-compat with #9276 always-mode semantics tests/unit/8189-classifier-compat-auto-narrow.test.ts was a test-sibling forgotten when #9276 (commit6b531fbacd) removed the unconditional `if (mode === "always") return true` branch from shouldDefaultAllowClassifier(). tests/unit/claude-classifier-compat.test.ts was updated in that same commit; this file was not. Old contract: 'always' mode short-circuited every Claude-format request unconditionally (operator opt-in was treated as sufficient on its own). New contract: 'always' now requires the same SECURITY_MONITOR_MARKER system-prompt text as 'auto' — the marker-optional behavior let a normal chat request through /v1/messages be silently swallowed by an operator's 'always' opt-in. The single 'always' test (1 assert, no-marker body expecting true) is replaced by two tests mirroring the depth already used for 'auto' mode in the same file: no-marker/false and marker-present/true. Net effect is +1 assert, not a reduction — the new pair verifies both directions of the narrowed contract instead of only the now-incorrect unconditional case. Before: 3/4 pass (the 'always' test failed: expected true, got false). After: 5/5 pass. Refs #9276 * fix(tests): align deepseek-web-tools-execute with #9343 tool envelope contract tests/unit/deepseek-web-tools-execute-2820.test.ts (executor level) was a test-sibling forgotten when #9343 (commitd969555417) hardened tool-call parsing: bare JSON with no explicit <tool>/<tool_call> envelope is never promoted to tool_calls anymore (previously it was, whenever a tools[] set was requested — a security gap allowing prose/code-fenced JSON echoed back by the model, or a copy-attack, to trigger real tool execution). Three siblings were updated in the same commit: web-tools-translation.test.ts and web-tools-translation-2820.test.ts (parseToolCallsFromText, the shared translator), and deepseek-web-tools-variants.test.ts (parseDeepSeekToolCalls, deepseek-specific parser) — all inverted their bare-JSON assertions to `toolCalls === null` + `content === text` (preserved verbatim, not stripped). This file calls the executor's execute() (full HTTP round trip through buildToolAwareResult), so it was not touched by that diff and kept asserting the old contract (finish_reason: "tool_calls", content: null). Verified against source (open-sse/executors/deepseek-web.ts buildToolAwareResult): when parseDeepSeekToolCalls returns toolCalls=null, hasCalls is false, so finish_reason is "stop", message.tool_calls is never set, and message.content is the parser's returned content — which for text with no <tool>/<tool_call> tag at all is the original string, unchanged (parseToolCallsFromText's early-return branch). The test now asserts exactly that shape, at the same executor level as the rest of the file's tool_calls that make sense at that level as the rest of the file's tool_calls Refs #9343 * fix(tests): align visionBridge tests with #8430 contract (partial — see note) Two test-siblings were forgotten when #8430 (commit7e55abbc41) hardened Vision Bridge's vision-model selection: getBestVisionModel() now validates that a candidate has a usable active connection (hasUsableCredentialsForModel, DB-backed) before returning it, instead of unconditionally returning the fixedModel or a hardcoded "openai/gpt-4o-mini" default. Three siblings were updated in the same commit (visionBridgeRouter.test.ts, the new repro-8430.test.ts, vision-bridge-preserve-on-failure-4012.test.ts); these two were not. tests/unit/guardrails/visionBridgeHelpers.callVisionModel.test.ts (8 failures, all "No vision-capable provider connected"): callVisionModel()'s `routerConfig` param only merges into getBestVisionModel's CONFIG argument, never its `deps` argument, so there is no way to inject a credentials stub through this function's public signature (unlike the guardrail class and getBestVisionModel itself, which do accept an injectable `hasUsableCredentials`). These tests exercise callVisionModel's own request/response handling, not credential routing (already covered elsewhere), so the fix seeds one real usable `provider_connections` row per provider the file exercises (openai, anthropic) via createProviderConnection in a test.before() hook, with resetDbInstance() in test.after() per the DB-handle-cleanup convention. All 8 now pass. tests/unit/guardrails/visionBridge.test.ts (7 failures): 1 of the 7 (VB-S03) is a genuine forgotten-contract case, fixed here — same semantic flip already applied to vision-bridge-preserve-on-failure-4012.test.ts: in the combo describe path, when EVERY describe call fails, the raw image is now replaced with an "(unavailable)" stub instead of preserved, because that path is only reached for confirmed non-vision targets. Assertions inverted to match (imagePart undefined, unavailable-stub present), same assert count, no weakening. *** THE OTHER 6 (VB-S12, VB-S12b, VB-S01, VB-S13, VB-S07, VB-S10) ARE DELIBERATELY LEFT FAILING. *** These are NOT a #8430 contract change — root- caused to what looks like a separate, unintentional regression: the ONE call to getBestVisionModel() in visionBridge.ts's whole-request-reroute path (line 244, `getBestVisionModel({ fixedModel: configuredModel })`) does not pass a `deps` second argument, so it always uses the real DB-backed hasUsableCredentialsForModel instead of this.deps.hasUsableCredentials — even though the two adjacent checks in the very same function (`checkCreds(model)` at line 226, `checkCreds(bestModel)` at line 246) DO honor the injectable override. In this suite's empty-but-readable isolated test DB, that real check deterministically returns `false` (not the indeterminate `null` the file's own createGuardrail() comment says these tests rely on: "Fail-open (null) so classic VB-S01/S07/S10 reroute tests keep working without a live credential DB"), so getBestVisionModel silently returns null, the reroute branch's `if (bestModel && ...)` guard never fires, and every test that expects a reroute observes a silent no-op instead. Evidence this is a source gap, not a test that needs updating: - The file's own pre-existing comment names VB-S01/S07/S10 as tests the `null` fail-open default is SUPPOSED to keep green. - VB-CRED-01/02 (the file's only two tests that actually inject a non-default hasUsableCredentials mock) both pass today, but neither one's assertions distinguish "mock honored" from "mock ignored, real check also says no" — they don't prove the threading works, they just don't happen to notice it's missing. - visionBridgeRouter.test.ts, repro-8430.test.ts, and vision-bridge-preserve-on-failure-4012.test.ts (22 tests, all green) all either call getBestVisionModel directly with explicit deps, or mock callVisionModel wholesale (bypassing getBestVisionModel entirely) — none of them exercises this exact call site through the guardrail's own deps. Per instructions, this was intentionally NOT "fixed" by weakening these 6 tests' assertions (that would mask the gap) or by seeding fake DB credentials to route around it (that would hide a real production DI inconsistency behind a test-only workaround) or by touching src/lib/guardrails/visionBridge.ts (a production behavior change outside a test-alignment task's scope, and Hard Rule #18 requires its own TDD/validation cycle). Flagging for the owner: the likely one-line fix is threading `{ hasUsableCredentials: this.deps.hasUsableCredentials }` as getBestVisionModel's second argument at visionBridge.ts:244, mirroring the two adjacent call sites in the same function. Before: 15 failures (7 + 8). After: 9 pass added (1 + 8), 6 still fail (unchanged, by design). Refs #8430 * feat(quality): add strayFromCommit deletion allowlist form to test-masking gate The deletion allowlist supported two shapes: replacement (test rewritten elsewhere) and sourceRemoved (feature deleted). Neither fits a third legitimate case surfaced today: test files that entered the repo BY ACCIDENT — commitf4e93f339d(#7786 docs) swept another session's worktree artifacts into the release, including TDD probes owned by open fix PRs (probe-9033-repro -> PR #9385, repro-8956 -> PR #9354, repro-8522 -> PR #9355). Those probes fail by design until their owning PR merges, so every unit run on the release tip broke on them. The new strayFromCommit form is verified, not trusted: the gate asks git which commit actually ADDED the file (git log --diff-filter=A) and only exempts the deletion when it matches the declared hash; a non-empty reason naming the owning PR/issue is mandatory. Also allowlists the deepseek-web-tools-execute assert reduction (23->21) fromed661f2126— same #9343 contract-inversion class as the existing web-tools-translation entry. Gate unit tests: 55/55 pass. Full gate vs main: OK. * fix(guardrails): pass credential deps to getBestVisionModel at reroute call site The individual-model reroute path in VisionBridgeGuardrail.preCall() calls getBestVisionModel({ fixedModel: configuredModel }) without its second `deps` argument, so the router always falls back to the real DB-backed hasUsableCredentialsForModel instead of an injected `deps.hasUsableCredentials` override. The two adjacent credential checks in the same function (the original-model check and the best-model check, both via the local `checkCreds` binding) already thread deps correctly — only this middle call, added in #8430, was left out. Pass the same resolved `checkCreds` used by those two adjacent checks as `getBestVisionModel`'s deps argument so all three credential checks in this reroute path stay consistent. Fixes 6 tests in tests/unit/guardrails/visionBridge.test.ts that depended on the injected hasUsableCredentials mock being honored on this path: VB-S12, VB-S12b, VB-S01, VB-S13, VB-S07, VB-S10. Refs #8430 * fix(quality): raise unit ceiling to 100min and align 2 more forgotten sibling tests Unit ceiling 45->100min: a hermetic-env measurement on the loaded devbox (load 7-26) was still inside invocation 1 of 3 at 76min when killed; contention factor 2-3x measured, no idle measurement exists. The pre-flight's real condition is exactly that contended one (unit runs in Promise.all with integration+vitest), and there 45min provably killed a healthy suite and fabricated a false base-red. The 45min value came from v3.8.43 as an estimate never validated by measurement. TODO in-code: re-tighten after an idle run on the .113 box. Also aligns the 6th and 7th occurrences of the same systemic pattern (behavior change merged updating only part of the sibling tests): - issue-7859-gemini-web-redirect-valid: #9407 refined ServiceLogin redirects to mean expired session; the #7859 regression coverage is preserved via a non-ServiceLogin public redirect variant. - provider-validation-specialty claude-web 429: #9406 inverted the contract (rate-limited session is unhealthy); the dedicated repro file owns the full contract, this sibling now matches it. Also carries the file-size rebaseline for #9323's base.ts growth (1578->1623, WAF retry + burst guard) and the eslintWarnings baseline tightened 5000->0 (real measured value with the TS7 suppressions in place — 5000 left the ratchet inert). Refs #9407, #9406, #9323 * fix(tests): restore the 3 TDD probes now owned by merged fixes and drop their stray allowlist entries The base advanced while this PR was open: the real fixes for the three issues behind the stray probes all merged into release/v3.8.50 — #9385 (issue 9033), #9355 (issue 8522) and #9354 (issue 8956). - probe-9033-repro / repro-8522: the base rewrote both probes into the regression tests of their merged fixes, so the delete side of the rebase conflict was dropped and the base versions kept. - repro-8956: #9354 only realigned one fixture line in auto-update.test.ts (package.json marker now needs a name field) and added no test for the new skip-synthetic behavior — the probe is the ONLY regression coverage of that merged fix (2/2 green on the base), so deleting it would remove real coverage. Restored. With no test-file deletions left in the PR diff, the three strayFromCommit allowlist entries are stale and removed. The strayFromCommit form support in check-test-masking.mjs stays (covered by its own fixtures). * fix(quality): rebaseline file-size for PR #9529 own growth The base sits exactly at the old frozen values, so the base-relative mode (#8522) does not cover this growth — it is this PR's own: - open-sse/services/rateLimitManager.ts 1060->1105: the applyLimiterSettings() helper that re-arms the reservoir heartbeat after updateSettings (Bottleneck 2.19.5 fix, TDD in ratelimit-reservoir-refresh.test.ts). - tests/integration/chat-pipeline.test.ts 1592->1598: codex User-Agent derived from getCodexClientVersion() instead of a pinned literal. - tests/unit/provider-validation-specialty.test.ts 2980->2985: new claude-web 429 -> valid:false coverage (#9406). * fix(docs): sync provider count to 291 in README and CLAUDE The live catalog counts 291 providers but README.md/CLAUDE.md still said 290, so the STRICT 'Docs Gates (fast-path)' check reds EVERY open PR against release/v3.8.50 (verified on #9537/#9539 as well — inherited base-red, not introduced by this PR). Updated all provider-count mentions including the section anchor. * fix(tests): align launch-codex 6312 guard with the async #9454 spawn contract #9454 made resolveCodexSpawn async (PATH-probes a native codex.exe before the .cmd shim) and updated its own tests, but left this older sibling calling the function synchronously — destructuring the Promise yields undefined and reds Unit fast-path (1/4) for EVERY open PR against the release (verified on #9537/#9539; inherited base-red). Realigned to the async contract with an injected probe; keeps the original #6312 fallback guard plus the only non-Windows codex coverage (now also asserting the probe never runs off Windows). * fix(translator): move state-mutating reasoning summary helper out of the pure leaf #9500 added buildResponsesReasoningSummaryDelta(state, ...) to pureHelpers.ts, but the function reads AND mutates stream state (reasoningSummaryIndex map) — violating the leaf contract declared in the file header ('no host imports, no stream state') and guarded by response-openai-responses-purehelpers-split.test.ts, which reds Unit fast-path (4/4) for every open PR (inherited base-red, verified on #9537/#9539). Moved verbatim to the host next to the other stream-state helpers (markResponsesReasoningDeltaEmitted); the host was its only consumer. Behavior unchanged: repro-9500-reasoning-separator 3/3 green, leaf/host architecture tests green. * fix(quality): rebaseline openai-responses.ts for the leaf-state relocation The #9500 helper moved from pureHelpers.ts into the host (previous commit) grows the host file 1174->1204 while the leaf shrinks by the same amount — net-zero LOC across the pair, but the per-file frozen ratchet only sees the growing side. * fix(tests): let the 9442 cert-mode test see past the harness trust-store guard tests/_setup/isolateDataDir.ts sets OMNIROUTE_SKIP_SYSTEM_TRUST=1 globally, which makes installCert() return before issuing any command — so the #9442 install-gap test captured nothing and could NEVER pass under npm run test:unit (it only passed invoked directly, harness-less; inherited base-red on Unit fast-path 3/4, verified on #9537/#9539). Clear the flag for this file only (restored in test.after): safe because every spawned command is a logging stub on PATH and OMNIROUTE_NO_SUDO=1 strips sudo, so nothing touches the real trust store. 6/6 under the CI harness including system-trust-test-guard. --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
54 KiB
OmniRoute agent guide
Single source of truth. This file holds ALL project rules, conventions, architecture notes and Hard Rules for every AI assistant working this repository (Claude Code, Gemini, Codex, Copilot, and any other agent).
CLAUDE.mdandGEMINI.mdonly add assistant-specific deltas and point back here. When a rule needs to change, change it HERE — never re-fork it into an assistant-specific file.
Quick Start
npm install # Install deps (auto-generates .env from .env.example)
npm run dev # Dev server at http://localhost:20128
npm run build # Production build (Next.js 16 standalone)
npm run build:release # Release build
npm run lint # ESLint (0 errors expected; warnings are pre-existing)
npm run typecheck:core # TypeScript check (should be clean)
npm run typecheck:noimplicit:core # Strict check (no implicit any)
npm run test:coverage # Unit tests + coverage gate (60/60/60/60 — statements/lines/functions/branches)
npm run check # lint + test combined
npm run check:cycles # Detect circular dependencies
npm run check:docs-all # Run after changing documentation (includes fabricated-docs validation)
Running Tests
Run the most focused test for changed code first:
# Single test file (Node.js native test runner — most tests)
node --import tsx/esm --test tests/unit/your-file.test.ts
# Vitest (MCP server, autoCombo, cache)
npm run test:vitest
# All suites
npm run test:all
Other suites: npm run test:e2e, npm run test:protocols:e2e, npm run test:ecosystem.
For full test matrix, see CONTRIBUTING.md → "Running Tests". For deep architecture, see the
Repository map and Reference Documentation sections below.
Project at a Glance
OmniRoute — unified AI proxy/router. One endpoint, 291 LLM providers, auto-fallback.
| Layer | Location | Purpose |
|---|---|---|
| API Routes | src/app/api/v1/ |
Next.js App Router — entry points |
| Handlers | open-sse/handlers/ |
Request processing (chat, embeddings, etc) |
| Executors | open-sse/executors/ |
Provider-specific HTTP dispatch |
| Translators | open-sse/translator/ |
Format conversion (OpenAI↔Claude↔Gemini) |
| Transformer | open-sse/transformer/ |
Responses API ↔ Chat Completions |
| Services | open-sse/services/ |
Combo routing, rate limits, caching, etc |
| Database | src/lib/db/ |
SQLite domain modules (130 migrations) |
| Domain/Policy | src/domain/ |
Policy engine, cost rules, fallback logic |
| MCP Server | open-sse/mcp-server/ |
104 tools (42 base + memory/skill/agentSkill/pool/notion/obsidian/gamification/plugin modules), 3 transports (stdio / SSE / Streamable HTTP), 31 scopes |
| A2A Server | src/lib/a2a/ |
JSON-RPC 2.0 agent protocol |
| Skills | src/lib/skills/ |
Extensible skill framework |
| Memory | src/lib/memory/ |
Persistent conversational memory |
Monorepo: src/ (Next.js 16 app), open-sse/ (streaming engine workspace), electron/ (desktop app), tests/, bin/ (CLI entry point).
Request Pipeline
Client → /v1/chat/completions (Next.js route)
→ CORS → Zod validation → auth? → policy check → prompt injection guard
→ handleChatCore() [open-sse/handlers/chatCore.ts]
→ cache check → rate limit → combo routing?
→ resolveComboTargets() → handleSingleModel() per target
→ translateRequest() → getExecutor() → executor.execute()
→ fetch() upstream → retry w/ backoff
→ response translation → SSE stream or JSON
→ If Responses API: responsesTransformer.ts TransformStream
API routes follow a consistent pattern: Route → CORS preflight → Zod body validation → Optional auth (extractApiKey/isValidApiKey) → API key policy enforcement → Handler delegation (open-sse). No global Next.js middleware — interception is route-specific.
Combo routing (open-sse/services/combo.ts): 19 public strategies (priority, weighted, fill-first, round-robin, p2c, random, least-used, cost-optimized, reset-aware, reset-window, headroom, strict-random, auto, lkgp, context-optimized, cache-optimized, context-relay, fusion, pipeline). Each target calls handleSingleModel() which wraps handleChatCore() with per-target error handling and circuit breaker checks. The fusion strategy is the exception: it fans out to a panel of models in parallel, then a judge model synthesizes one final answer (open-sse/services/fusion.ts). See docs/routing/AUTO-COMBO.md for the 13-factor Auto-Combo scoring + the full strategy table and docs/architecture/RESILIENCE_GUIDE.md for the 3 resilience layers.
Resilience Runtime State
OmniRoute has three related but distinct temporary-failure mechanisms. Keep their scope separate when debugging routing behavior. See the 3-layer resilience diagram (source: docs/diagrams/resilience-3layers.mmd) for an at-a-glance map.
Provider Circuit Breaker
Scope: whole provider, e.g. glm, openai, anthropic.
Purpose: stop sending traffic to a provider that is repeatedly failing at the upstream/service level, so one unhealthy provider does not slow down every request.
Implementation:
- Core class:
src/shared/utils/circuitBreaker.ts - Chat gate/execution wiring:
src/sse/handlers/chatHelpers.ts,src/sse/handlers/chat.ts - Runtime status API:
src/app/api/monitoring/health/route.ts - Shared wrappers:
open-sse/services/accountFallback.ts - Persisted state table:
domain_circuit_breakers
States:
CLOSED: normal traffic is allowed.OPEN: provider is temporarily blocked; callers get a provider-circuit-open response or combo routing skips to another target.HALF_OPEN: reset timeout has elapsed; allow a probe request. Success closes the breaker, failure opens it again.
Defaults (open-sse/config/constants.ts):
- OAuth providers: threshold
3, reset timeout60s. - API-key providers: threshold
5, reset timeout30s. - Local providers: threshold
2, reset timeout15s.
Only provider-level failure statuses should trip the provider breaker:
(408, 500, 502, 503, 504);
Do not trip the whole-provider breaker for normal account/key/model errors like most
401, 403, or 429 cases. Those usually belong to connection cooldown or model
lockout. A generic API-key provider 403 should be recoverable unless it is classified
as a terminal provider/account error.
The breaker uses lazy recovery, not a background timer. When OPEN expires, reads such
as getStatus(), canExecute(), and getRetryAfterMs() refresh the state to
HALF_OPEN, so dashboards and combo candidate builders do not keep excluding an
expired provider forever.
Connection Cooldown
Scope: one provider connection/account/key.
Purpose: temporarily skip one bad key/account while allowing other connections for the same provider to continue serving requests.
Implementation:
- Write/update path:
src/sse/services/auth.ts::markAccountUnavailable() - Account selection/filtering:
src/sse/services/auth.ts::getProviderCredentials... - Cooldown calculation:
open-sse/services/accountFallback.ts::checkFallbackError() - Settings:
src/lib/resilience/settings.ts
Important fields on provider connections:
rateLimitedUntil;
testStatus: "unavailable";
lastError;
lastErrorType;
errorCode;
backoffLevel;
During account selection, a connection is skipped while:
new Date(rateLimitedUntil).getTime() > Date.now();
Cooldowns are also lazy: when rateLimitedUntil is in the past, the connection becomes
eligible again. On successful use, clearAccountError() clears testStatus,
rateLimitedUntil, error fields, and backoffLevel.
Default connection cooldown behavior:
- OAuth base cooldown:
5s. - API-key base cooldown:
3s. - API-key
429should prefer upstream retry hints (Retry-After, reset headers, or parseable reset text) when available. - Repeated recoverable failures use exponential backoff:
baseCooldownMs * 2 ** failureIndex;
The anti-thundering-herd guard prevents concurrent failures on the same connection from
repeatedly extending the cooldown or double-incrementing backoffLevel.
Terminal states are not cooldowns. banned, expired, and credits_exhausted are
intended to stay unavailable until credentials/settings change or an operator resets
them. Do not overwrite terminal states with transient cooldown state.
Model Lockout
Scope: provider + connection + model.
Purpose: avoid disabling a whole connection when only one model is unavailable or quota-limited for that connection.
Examples:
- Per-model quota providers returning
429. - Local providers returning
404for one missing model. - Provider-specific mode/model permission failures such as selected Grok modes.
Model lockout lives in open-sse/services/accountFallback.ts and lets the same
connection continue serving other models.
Debugging Guidance
- If all keys for a provider are skipped, inspect both provider breaker state and each
connection's
rateLimitedUntil/testStatus. - If a provider appears permanently excluded after the reset window, check whether code
is reading raw
stateinstead of usinggetStatus()/canExecute(). - If one provider key fails but others should work, prefer connection cooldown over provider breaker.
- If only one model fails, prefer model lockout over connection cooldown.
- If a state should self-recover, it should have a future timestamp/reset timeout and a read path that refreshes expired state. Permanent statuses require manual credential or config changes.
Repository map
Read the nearest AGENTS.md and the linked deep-dive before making a non-trivial change.
| Area | Location | Start here |
|---|---|---|
| API routes | src/app/api/v1/ |
docs/architecture/ARCHITECTURE.md |
| Streaming request handling | open-sse/handlers/ |
docs/architecture/ARCHITECTURE.md |
| Provider execution and translation | open-sse/executors/, open-sse/translator/ |
docs/architecture/CODEBASE_DOCUMENTATION.md |
| Routing and resilience | open-sse/services/ |
open-sse/services/AGENTS.md, docs/routing/AUTO-COMBO.md |
| Database and migrations | src/lib/db/, db/migrations/ |
src/lib/db/AGENTS.md |
| Domain policy | src/domain/ |
docs/architecture/ARCHITECTURE.md |
| MCP and A2A | open-sse/mcp-server/, src/lib/a2a/ |
docs/frameworks/MCP-SERVER.md, docs/frameworks/A2A-SERVER.md |
| Agent features | src/lib/{acp,memory,skills,cloudAgent}/ |
docs/frameworks/AGENT_PROTOCOLS_GUIDE.md, docs/frameworks/SKILLS.md |
| Safety and governance | src/lib/{guardrails,compliance}/, src/server/authz/ |
docs/security/GUARDRAILS.md, docs/architecture/AUTHZ_GUIDE.md |
| Operations | src/mitm/, tunnel modules, electron/ |
docs/ops/TUNNELS_GUIDE.md, docs/guides/ELECTRON_GUIDE.md |
File placement & repo-root hygiene
- Test files: ALL unit tests, integration tests, ecosystem tests, or Vitest files MUST strictly be placed within the
tests/directory (e.g.,tests/unit/,tests/integration/). NEVER create test files in the project root (/). - Scripts and utilities: ALL maintenance, debugging, generation, or experimental scripts (
.cjs,.mjs,.js,.ts) MUST be placed strictly inside one of thescripts/subfolders (build/,dev/,check/,docs/,i18n/,ad-hoc/). One-shot or experimental code goes underscripts/ad-hoc/. NEVER dump loose scripts in the project root (/) or the top-levelscripts/folder.
The project root MUST ONLY contain:
- Configuration files (
vitest.config.ts,next.config.mjs,eslint.config.mjs,tsconfig*.json,playwright.config.ts,prettier.config.mjs,postcss.config.mjs,sonar-project.properties,fly.toml,docker-compose*.yml,Dockerfile) - Dependency files (
package.json,package-lock.json) - Documentation files (
README.md,CHANGELOG.md,LICENSE,AGENTS.md,CLAUDE.md,GEMINI.md,CONTRIBUTING.md,SECURITY.md,CODE_OF_CONDUCT.md,llm.txt,Tuto_Qdrant.md) - CI/CD files and ignore definitions (
.gitignore,.dockerignore,.npmignore,.npmrc,.node-version,.nvmrc,.env.example)
When creating any validation tests or one-off logic scripts, default to scripts/ad-hoc/ or tests/unit/ according to your goals. Do not pollute the / root context.
Key Conventions
Code Style
- 2 spaces, semicolons, double quotes, 100 char width, es5 trailing commas (enforced by lint-staged via Prettier) — run Prettier on changed files
- Imports: external → internal (
@/,@omniroute/open-sse) → relative - Naming: files=camelCase/kebab, components=PascalCase, constants=UPPER_SNAKE
- ESLint:
no-eval,no-implied-eval,no-new-func= error everywhere;no-explicit-any= error inopen-sse/andtests/(since #6218 — pre-existing violations are frozen inconfig/quality/eslint-suppressions.json, new ones must be fixed;npm run lintapplies the suppressions and is what CI runs) - TypeScript:
strict: false, target ES2022, module esnext, resolution bundler. Prefer explicit types.
Database
- Always go through
src/lib/db/domain modules — never write raw SQL in routes or handlers - Never add logic to
src/lib/localDb.ts(re-export layer only) - Never barrel-import from
localDb.ts— import specificdb/modules instead - DB singleton:
getDbInstance()fromsrc/lib/db/core.ts(WAL journaling) - Migrations:
src/lib/db/migrations/— versioned SQL files, idempotent, run in transactions
Error Handling
- try/catch with specific error types, log with pino context
- Never swallow errors in SSE streams — use abort signals for cleanup
- Return proper HTTP status codes (4xx/5xx)
Security
- Never use
eval(),new Function(), or implied eval - Validate all inputs with Zod schemas
- Encrypt credentials at rest (AES-256-GCM); never log SQLite encryption keys
- Sanitize user HTML with DOMPurify
- Upstream header denylist:
src/shared/constants/upstreamHeaders.ts— keep sanitize, Zod schemas, and unit tests aligned when editing - Public upstream credentials (Gemini/Antigravity/Windsurf-style OAuth client_id/secret + Firebase Web keys extracted from public CLIs): MUST be embedded via
resolvePublicCred()fromopen-sse/utils/publicCreds.ts— never as string literals. Seedocs/security/PUBLIC_CREDS.mdfor the mandatory pattern. - Error responses (HTTP / SSE / executor / MCP handler): MUST route through
buildErrorBody()orsanitizeErrorMessage()fromopen-sse/utils/error.ts— never put rawerr.stackorerr.messagein a response body. Seedocs/security/ERROR_SANITIZATION.md. - Shell commands built from variables: when calling
exec()/spawn()with a script that needs runtime values, pass them via theenvoption (shell-escaped automatically) — never string-interpolate untrusted/external paths into the script body. Reference:src/mitm/cert/install.ts::updateNssDatabases. - Secure-by-default libraries (tldrsec/awesome-secure-defaults): prefer Helmet.js, DOMPurify, ssrf-req-filter, safe-regex, Google Tink over custom implementations whenever adding new security-sensitive surfaces.
Documentation accuracy
Documentation must describe verified behavior, not plausible behavior.
- Before documenting an API name, endpoint, path, CLI command, or environment variable,
search for it:
rg -n "name" src/ open-sse/ bin/. If it has no source match, do not document it. - Measure mutable counts instead of writing them from memory: use
wc -l <file>or a directory-specific count command. - Copy code examples from working usage or run them. Prefer a source link such as
path/to/file.ts:lineto an invented signature. - Run
npm run check:docs-allfor edits underdocs/; it includes the fabricated-docs validation.
Common Modification Scenarios
Adding a New Provider
- Register in
src/shared/constants/providers.ts(Zod-validated at load) - Add executor in
open-sse/executors/if custom logic needed (extendBaseExecutor) - Add translator in
open-sse/translator/if non-OpenAI format - Add OAuth config in
src/lib/oauth/constants/oauth.tsif OAuth-based — if the upstream CLI ships a public client_id/secret, embed viaresolvePublicCred()(seedocs/security/PUBLIC_CREDS.md), never as a literal - Register models in
open-sse/config/providerRegistry.ts - Write tests in
tests/unit/(include the publicCreds shape assertion if you added a new embedded default)
Adding a New API Route
- Create directory under
src/app/api/v1/your-route/ - Create
route.tswithGET/POSThandlers - Follow pattern: CORS → Zod body validation → optional auth → handler delegation
- Handler goes in
open-sse/handlers/(import from there, not inline) - Error responses use
buildErrorBody()/errorResponse()fromopen-sse/utils/error.ts(auto-sanitized — never puterr.stackorerr.messageraw in the body). Seedocs/security/ERROR_SANITIZATION.md. - Add tests — including at least one assertion that error responses do not leak stack traces (
!body.error.message.includes("at /"))
Adding a New DB Module
- Create
src/lib/db/yourModule.ts— importgetDbInstancefrom./core.ts - Export CRUD functions for your domain table(s)
- Add migration in
src/lib/db/migrations/if new tables needed - Re-export from
src/lib/localDb.ts(add to the re-export list only) - Write tests
Adding a New MCP Tool
- Add tool definition in
open-sse/mcp-server/tools/with Zod input schema + async handler - Register in tool set (wired by
createMcpServer()) - Assign to appropriate scope(s)
- Write tests (tool invocation logged to
mcp_audittable)
Adding a New A2A Skill
- Create skill in
src/lib/a2a/skills/(5 already exist: smart-routing, quota-management, provider-discovery, cost-analysis, health-report) - Skill receives task context (messages, metadata) → returns structured result
- Register in
A2A_SKILL_HANDLERSinsrc/lib/a2a/taskExecution.ts - Expose in
src/app/.well-known/agent.json/route.ts(Agent Card) - Write tests in
tests/unit/ - Document in
docs/frameworks/A2A-SERVER.mdskill table
Adding a New Cloud Agent
- Create agent class in
src/lib/cloudAgent/agents/extendingCloudAgentBase(3 already exist: codex-cloud, devin, jules) - Implement
createTask,getStatus,approvePlan,sendMessage,listSources - Register in
src/lib/cloudAgent/registry.ts - Add OAuth/credentials handling if needed (
src/lib/oauth/providers/) - Tests + document in
docs/frameworks/CLOUD_AGENT.md
Adding a New Embedded Service
- Create installer in
src/lib/services/installers/{name}.tsmodeled onninerouter.ts(userunNpmfrominstallers/utils.ts— no shell interpolation, hard rule #13). - Register the service in
src/lib/services/bootstrap.ts(add toSERVICES[]array and extendbuildSpawnArgsFactory()). - Add a DB seed row for the new service in
src/lib/db/migrations/(version_managertable,status='not_installed',auto_start=0). - Create 7 API endpoints under
src/app/api/services/{name}/(_lib.ts,install,start,stop,restart,update,status,auto-start). All delegate errors throughcreateErrorResponse(). The sharedlogsendpoint is already wired via[name]/logs/route.ts. - Verify
/api/services/is inLOCAL_ONLY_API_PREFIXESinsrc/server/authz/routeGuard.ts; add a test assertingisLocalOnlyPath()returnstruefor the new prefix if you add one (hard rule #17). - Add a UI tab in
src/app/(dashboard)/dashboard/providers/services/tabs/reusingServiceStatusCard,ServiceLifecycleButtons,ServiceLogsPanel. - Document in
docs/frameworks/EMBEDDED-SERVICES.md(update §1 service table + §4 API reference) anddocs/openapi.yaml. - Write tests: unit (
tests/unit/services/), integration (tests/integration/services/, gated byRUN_SERVICES_INT=1), and updatedocs/ops/RELEASE_CHECKLIST.mdsmoke section.
Adding a New Guardrail / Eval / Skill / Webhook event
- Guardrail:
src/lib/guardrails/→ docs:docs/security/GUARDRAILS.md - Eval suite:
src/lib/evals/→ docs:docs/frameworks/EVALS.md - Skill (sandbox):
src/lib/skills/→ docs:docs/frameworks/SKILLS.md - Webhook event:
src/lib/webhookDispatcher.ts→ docs:docs/frameworks/WEBHOOKS.md
Reference Documentation
For any non-trivial change, read the matching deep-dive first:
| Area | Doc |
|---|---|
| Repo navigation | docs/architecture/REPOSITORY_MAP.md |
| Architecture | docs/architecture/ARCHITECTURE.md |
| Engineering reference | docs/architecture/CODEBASE_DOCUMENTATION.md |
| Auto-Combo (13-factor scoring, 19 strategies) | docs/routing/AUTO-COMBO.md |
| Resilience (3 mechanisms) | docs/architecture/RESILIENCE_GUIDE.md |
| Reasoning replay | docs/routing/REASONING_REPLAY.md |
| Skills framework | docs/frameworks/SKILLS.md |
| Memory system (FTS5 + Qdrant) | docs/frameworks/MEMORY.md |
| Cloud agents | docs/frameworks/CLOUD_AGENT.md |
| Guardrails (PII / injection / vision) | docs/security/GUARDRAILS.md |
| Public upstream credentials (Gemini/etc.) | docs/security/PUBLIC_CREDS.md |
| Error message sanitization | docs/security/ERROR_SANITIZATION.md |
| Evals | docs/frameworks/EVALS.md |
| Compliance / audit | docs/security/COMPLIANCE.md |
| Webhooks | docs/frameworks/WEBHOOKS.md |
| Authorization pipeline | docs/architecture/AUTHZ_GUIDE.md |
| Stealth (TLS / fingerprint) | docs/security/STEALTH_GUIDE.md |
| Agent protocols (A2A / ACP / Cloud) | docs/frameworks/AGENT_PROTOCOLS_GUIDE.md |
| MCP server | docs/frameworks/MCP-SERVER.md |
| A2A server | docs/frameworks/A2A-SERVER.md |
| API reference + OpenAPI | docs/reference/API_REFERENCE.md + docs/openapi.yaml |
| Provider catalog (auto-generated) | docs/reference/PROVIDER_REFERENCE.md |
| Tunnels | docs/ops/TUNNELS_GUIDE.md |
| Electron desktop app | docs/guides/ELECTRON_GUIDE.md |
| Release flow | docs/ops/RELEASE_CHECKLIST.md |
| Embedded services | docs/frameworks/EMBEDDED-SERVICES.md |
| Quality gates (~48 scripts, allowlist policy) | docs/architecture/QUALITY_GATES.md |
Testing
| What | Command |
|---|---|
| Unit tests | npm run test:unit |
| Single file | node --import tsx/esm --test tests/unit/your-file.test.ts |
| Vitest (MCP, autoCombo) | npm run test:vitest |
| E2E (Playwright) | npm run test:e2e |
| Protocol E2E (MCP+A2A) | npm run test:protocols:e2e |
| Ecosystem | npm run test:ecosystem |
| Coverage gate | npm run test:coverage (60/60/60/60 — statements/lines/functions/branches) |
| Coverage report | npm run coverage:report |
PR rule: If you change production code in src/, open-sse/, electron/, or bin/, you must include or update tests in the same PR.
Test layer preference: unit first → integration (multi-module or DB state) → e2e (UI/workflow only). Encode bug reproductions as automated tests before or alongside the fix.
Both test runners must pass: npm run test:unit (Node native — most tests) AND npm run test:vitest (MCP server, autoCombo, cache) cover non-overlapping files. Both are wired in CI (jobs test-unit and test-vitest) and must be green before merging. A PR where only one suite passes may silently ship broken MCP tools or routing regressions.
Bug fix / issue triage protocol (Hard Rule #18): Every fix for a reported issue must be validated by one of the following — no exceptions:
- TDD (preferred) — write a failing test reproducing the bug → fix it → confirm the test passes. The test becomes the permanent regression guard. Touch only the files the test proves need changing; nothing more.
- Real-environment test (when TDD is not possible) — deploy to the production VPS (
root@192.168.0.15) and run a documented live test. Record the exact command + result in the PR description. Applies to: OAuth upstream flows, Cloudflare/WS upstream behavior, UI-only regressions, hardware-dependent behavior. - "It worked locally without a test" does not count. A fix without a test or a VPS validation record is not a fix — it is a guess.
Why this matters: fixing bug A while opening bug B is worse than not fixing at all. The TDD/VPS gate enforces surgical scope — you touch only what the failing test proves is broken. Examples where this paid off: #3090 (claude-web 403), #3113 (WS HTTP fallback), #3052 (heap-guard auto-calibration).
Copilot coverage policy: When a PR changes production code and coverage is below 60% (statements/lines/functions/branches), do not just report — add or update tests, rerun the coverage gate, then ask for confirmation. Include commands run, changed test files, and final coverage result in the PR report.
Review focus
- Keep database operations in
src/lib/db/; do not issue raw SQL from routes. - Send provider requests through
open-sse/handlers/. - Keep MCP and A2A pages as tabs inside
/dashboard/endpoint. - Preserve SSE cleanup, rate-limit header parsing, Zod validation, and provider-schema validation.
- Treat Memory and Skills as cross-cutting changes that can affect MCP tools, the request pipeline, and A2A skills.
- Do not close a contributor pull request after using its code; merge it through GitHub so the contributor receives credit.
Planning & Research Artifacts
_tasks/ is a separate, isolated git repository that is gitignored by the main
repo (.gitignore → _tasks/). It is the canonical home for working artifacts —
plans, specs/designs, research, hand-offs — so they stay versioned in their own
repo instead of polluting the main OmniRoute tree.
Hard rule — never write planning / research output under docs/ or the repo root.
Whenever any plan/spec/research generator runs in this project (superpowers or otherwise),
save to _tasks/ using the filename convention:
| Artifact | Save here |
|---|---|
| Plans | _tasks/superpowers/plans/YYYY-MM-DD-<feature>.md |
| Specs / design | _tasks/superpowers/specs/YYYY-MM-DD-<topic>-design.md |
| Research | _tasks/research/… |
| Hand-offs | _tasks/hands-off/<YYYY-MM-DD>_<branch>_v<versão>_sess-<id>/ |
Commit those artifacts inside the _tasks/ repo (git -C _tasks …), never in the main repo.
Git Workflow
# Never commit directly to main
git checkout -b feat/your-feature
git commit -m "feat: describe your change"
git push -u origin feat/your-feature
Branch prefixes: feat/, fix/, refactor/, docs/, test/, chore/
Commit format (Conventional Commits): feat(db): add circuit breaker — scopes: db, sse, oauth, dashboard, api, cli, docker, ci, mcp, a2a, memory, skills
Husky hooks:
- pre-commit: lint-staged +
check-docs-sync+check:any-budget:t11+check:tracked-artifacts - pre-push: intentionally light (PATH/npm sanity only).
any-budget+tracked-artifactsalready run on pre-commit; re-running them on every push was pure double-pay. CI still enforces both. (Was Fase 6A.12 full pre-push gate; folded into pre-commit in #6716.)
Worktree isolation (MANDATORY for every development task)
Multiple sessions/agents work this repo in parallel. The main checkout is shared, so a
git checkout/branch switch in it silently discards another session's uncommitted work and
yanks the branch out from under whatever else is running (incidents: 2026-06-05, 2026-06-13).
Rule: never develop on the shared main checkout. Every task gets its own git worktree on its own dedicated branch, and you MUST confirm the base branch with the operator before creating it.
-
Ask first — which base branch? Before creating anything, ask the operator (unless they already told you) from which branch the new worktree/branch should be cut. Do NOT assume
mainor "whatever I'm on" — the answer is usually the activerelease/vX.Y.Z, but it can be another feature/release branch. Get the base explicitly. -
Create an isolated worktree + branch off that base (never reuse the main checkout). 🔴 MANDATORY PATH: every worktree lives under
.claude/worktrees/— and nowhere else. This is the single canonical location. It is gitignored AND in thetsconfig.json/.dockerignoreexcludes, so worktrees never leak into the build scope. Never use.worktrees/, repo-root, or any other path — a worktree outside.claude/worktrees/(a) escapes the build-scope excludes and poisonsnext build(thetsconfiginclude: **/*globs ~70× the codebase → OOM; incident 2026-06-25) and (b) scatters worktrees across two dirs.BASE_BRANCH="release/vX.Y.Z" # ← the branch the operator confirmed in step 1 TASK="feat/your-feature" # feat/ fix/ refactor/ docs/ test/ chore/ git fetch origin "$BASE_BRANCH" git worktree add ".claude/worktrees/${TASK##*/}" -b "$TASK" "origin/$BASE_BRANCH" cd ".claude/worktrees/${TASK##*/}" # Reuse the main checkout's node_modules to skip a per-worktree npm install. # HARD LINKS (`cp -al`), never a symlink: ~5s for the whole tree and near-zero extra # disk (the inodes are shared), and unlike a symlink it does not break the dev server. cp -al "$(git -C <main_checkout> rev-parse --show-toplevel)/node_modules" node_modulesNever
ln -snode_modules. Turbopack rejects a symlink that resolves outside the project root, sonpm run devdies with a FATAL panic (Symlink [project]/node_modules is invalid, it points out of the filesystem root) while typecheck, lint and the test runners all keep passing — the error names "filesystem root", not the worktree, so it reads like a Next/build bug and costs real time to trace (incident 2026-07-31, #9043). -
Work, commit, push, open the PR — all from inside the worktree. Never
git checkouta different branch inside a worktree another session might share. -
Tear down only your own worktree + branch when done, from the main checkout:
git worktree remove .claude/worktrees/<dir>thengit branch -D <task>. Never blanket-deletefix/*/feat/*— other sessions keep their own; delete only the branches you created, by name. -
Never touch another session's worktree, branch, or uncommitted changes. If
git worktree listshows worktrees you didn't create, leave them alone. End every session with the main checkout back on the branch it started on (the activerelease/vX.Y.Z, nevermain).
Base-green check (PRs must not be born red)
Before cutting a branch, merging the base into a PR branch, mass-retargeting PRs, or opening a
PR: check whether the base tip is green. The Release-Green (continuous) workflow
(.github/workflows/nightly-release-green.yml) publishes the verdict in a single deduplicated
issue titled 🔴 Release branch not green: <branch> (label base-red). One call replaces any
local suite run for this purpose:
gh issue list --repo diegosouzapw/OmniRoute --state open \
--search "Release branch not green: <base> in:title"
If the base is red: never treat the inherited failures as your branch's defect; never "fix" them
inside your feature branch (a base-red fix is its own freeze-gated fix/release-vX.Y.Z-basereds
PR); and if you must open a PR anyway, add ⚠️ base-red inherited: #<issue> to the PR body so
reviewers and CI babysitters do not chase ghosts.
Upstream contributions
This checkout is a fork of diegosouzapw/OmniRoute. Keep fork-only deployment and personal
automation changes out of upstream PRs.
Start upstream work from the active upstream default branch, not main:
git fetch upstream
git switch -c <branch-name> upstream/<default-branch>
Target that same release branch in the pull request. Stage only the intended files, run the
focused checks, and use a Conventional Commit message (for example, docs: slim AGENTS.md).
Environment
- Runtime: Node.js ≥22.0.0 <23 || ≥24.0.0 <27, ES Modules. This is the only supported runtime for the published
omnirouteCLI, the server, and the test suites (node:test+ vitest) —engines.nodeis authoritative and end users never need Bun. A best-effortbun:sqlitecompatibility path exists so a global Bun install (bun install -g omniroute) can start withoutbetter-sqlite3(driver adapter + Bun-aware process spawning); it is not a supported runtime — no support guarantees — and every Bun-specific runtime change MUST preserve the Node driver/fallback chain and ship a Bun test (test:bun:db) or an explicit reason why the path is Node-only. - Bun (build/dev script runner + compatibility smoke only): Bun
1.3.14is pinned as an exact devDependency (provisioned through the existingnpm civia the lockfile's@oven/bun-*platform binaries — nosetup-bun/ad-hoc install). It is used only to execute a small, allow-listed set of TypeScript gate/generator scripts (replacingnode --import tsxfor startup speed): the CI checkscheck:provider-consistency,check:compression-budget,check:known-symbols, and the non-CIgen:provider-reference,bench:compression— plus the focusedtest:bun:dbcompatibility smoke suite for the best-effortbun:sqlitepath. Do NOT widen Bun tonpm install, the build (build:cli*),check:pack-artifact, the supported published runtime, or the main test runners — those stay on Node. Any new Bun-invoking gate/generator script must be validated byte-identical against itsnode --import tsxoutput first. After pulling the lockfile change, runnpm installsobunresolves locally (a stalenode_moduleswill fail those scripts withbun: not found). - TypeScript: 6.0+, target ES2022, module esnext, resolution bundler
- Path aliases:
@/*→src/,@omniroute/open-sse→open-sse/,@omniroute/open-sse/*→open-sse/* - Default port: 20128 (API + dashboard on same port)
- Data directory:
DATA_DIRenv var, defaults to~/.omniroute/ - Key env vars:
PORT,JWT_SECRET,API_KEY_SECRET,INITIAL_PASSWORD,REQUIRE_API_KEY,APP_LOG_LEVEL - Setup:
cp .env.example .envthen generateJWT_SECRET(openssl rand -base64 48) andAPI_KEY_SECRET(openssl rand -hex 32)
Quality Gates & Ratchets
OmniRoute has ~48 quality-gate scripts (scripts/check/ + scripts/quality/) wired
across 9 gate-running jobs in .github/workflows/ci.yml (lint, quality-gate,
quality-extended, docs-sync-strict, i18n-ui-coverage, i18n, pr-test-policy,
test-vitest, sonarqube), plus the quality.yml fast-gates job (PR→release/**) and
3 nightly workflows (nightly-property, nightly-resilience, nightly-llm-security;
nightly-mutation once merged). Full inventory, per-job breakdown, and operational
procedures are in docs/architecture/QUALITY_GATES.md.
Quick reference:
- Gates in jobs
lint+docs-sync-strict: pass/fail policy gates — fix the violation or add an allowlist entry with a justification comment + tracking issue. - Gates in job
quality-gate: ratchet — metrics (ESLint warnings, code coverage, duplication, complexity) must not regress vsquality-baseline.json. Update vianpm run quality:ratchet -- --updatewhen a metric genuinely improves. - Job
test-vitestrunsnpm run test:vitest(MCP tools, autoCombo, cache) — blocking.test:vitest:uiis advisory until UI component tests are triaged.
Allowlist policy (short form): Fix the cause; use the allowlist only for pre-existing violations you cannot fix in the same PR. Add a comment with justification + issue number. Stale allowlist entries (suppressing a violation that no longer exists) will be caught by the stale-enforcement added in Fase 6A.3.
Hard Rules
- Never commit secrets or credentials
- Never add logic to
localDb.ts - Never use
eval()/new Function()/ implied eval - Never commit directly to
main - Never write raw SQL in routes — use
src/lib/db/modules - Never silently swallow errors in SSE streams
- Always validate inputs with Zod schemas
- Always include tests when changing production code
- Coverage must not regress below the baseline frozen in
quality-baseline.json(ratchet); absolute floor is 60% (statements/lines/functions/branches). Update the baseline vianpm run quality:ratchet -- --updateonly when coverage genuinely improves. Seedocs/architecture/QUALITY_GATES.md. - Never bypass Husky hooks (
--no-verify,--no-gpg-sign) without explicit operator approval. - Never embed public upstream OAuth client_id/secret or Firebase Web keys as string literals — always go through
resolvePublicCred()(open-sse/utils/publicCreds.ts). Seedocs/security/PUBLIC_CREDS.md. - Never return raw
err.stack/err.messagein HTTP / SSE / executor responses — always route throughbuildErrorBody()orsanitizeErrorMessage()(open-sse/utils/error.ts). Seedocs/security/ERROR_SANITIZATION.md. - Never string-interpolate external paths or runtime values into shell scripts passed to
exec()/spawn()— pass via theenvoption instead. Reference:src/mitm/cert/install.ts::updateNssDatabases. - Never dismiss a CodeQL / Secret-Scanning alert without (a) first checking the pattern docs above to see if the helper applies, and (b) recording the technical justification in the dismissal comment. Precedent:
js/stack-trace-exposureraised on callsites that already route throughsanitizeErrorMessage()is a known CodeQL limitation (custom sanitizers not recognized) — dismiss asfalse positivereferencingdocs/security/ERROR_SANITIZATION.md. - Never expose routes that spawn child processes (
/api/mcp/,/api/cli-tools/runtime/) withoutisLocalOnlyPath()classification insrc/server/authz/routeGuard.ts. Loopback enforcement happens unconditionally before any auth check — leaked JWT via tunnel cannot trigger process spawning. Seedocs/security/ROUTE_GUARD_TIERS.md. - Never credit or advertise an AI assistant, LLM, or automation account in any commit/PR metadata. Two forbidden forms, both equivalent — they route attribution to a bot account (or advertise AI authorship) and hide the real author (
diegosouzapw): (a)Co-Authored-Bytrailers naming an AI/bot (e.g. names containing "Claude", "GPT", "Copilot", "Bot"; emails atanthropic.com/openai.com/ bot-ownednoreply.github.comaddresses); (b) AI-generation footers or descriptions anywhere in a commit message, PR title/body, or CHANGELOG — e.g.🤖 Generated with [Claude Code], "Generated with Claude Code", "Made with ", or anyCo-authored-by: Claude/GPT/Copilotline. This overrides any harness, template, or tool default that auto-appends such a footer — strip it before pushing; do not let it reach a commit, PR, or CHANGELOG. Human collaborators — including upstream PR authors and issue reporters being ported into OmniRoute — MAY and SHOULD be credited with standardCo-authored-by: Name <email>trailers; the upstream-port workflows (/port-upstream-features,/port-upstream-issues) depend on this. - Never expose routes under
/api/services/or/dashboard/providers/services/*/embed/withoutisLocalOnlyPath()classification insrc/server/authz/routeGuard.ts. These routes can spawn child processes (npm install,node). Loopback enforcement happens unconditionally before any auth check — a leaked JWT via tunnel cannot trigger process spawning. Seedocs/security/ROUTE_GUARD_TIERS.md. - Every bug fix must be validated before shipping: a failing-then-passing unit/integration test (TDD) OR a documented live test on the production VPS (192.168.0.15). A fix without either is not merged. See Testing → "Bug fix / issue triage protocol" for the full decision tree.
- Never develop on the shared main checkout. Every development task runs in its own git worktree on its own dedicated branch, and you MUST confirm the base branch with the operator before creating the worktree/branch — never assume
mainor the currently checked-out branch. Agit checkoutin the shared checkout silently destroys other sessions' uncommitted work. Tear down only the worktrees/branches you created (by name, neverfix/*/feat/*wildcards), leave other sessions' worktrees untouched, and end on the branch you started on (the activerelease/vX.Y.Z, nevermain). See Git Workflow → "Worktree isolation". - PII redaction/sanitization is opt-in — never on by default. OmniRoute proxies for self-hosted/local LLMs where the operator owns the data, so mutating request/response payloads by default would silently corrupt legitimate traffic. The two data-mutating PII feature flags MUST keep
defaultValue: "false"insrc/shared/constants/featureFlagDefinitions.ts:PII_REDACTION_ENABLED(request-side) andPII_RESPONSE_SANITIZATION(response + streaming). All three application points —src/lib/guardrails/piiMasker.ts(request guardrail),src/lib/piiSanitizer.ts(response),src/lib/streamingPiiTransform.ts(SSE) — are gated on these flags; with both off thepii-maskerguardrail still runs but never mutates payloads (data passes through untouched). Flipping either default to"true"requires explicit operator approval. The regression guard istests/unit/pii-opt-in-default.test.ts(asserts both definition defaults + behavioral pass-through). Opt-in is per-operator via env or the settings/DB override (src/lib/db/featureFlags.ts), never a silent default. Seedocs/security/GUARDRAILS.md. - Release-freeze — the FROZEN release branch belongs to the release captain; development does NOT stop (parallel-cycle model, 2026-07-04).
/generate-releaseopens a marker issue labeledrelease-freezeat the start of reconciliation (Phase 0a), immediately cuts the next cycle's branchrelease/vX+1from the frozen tip (Phase 0a.0b — bump + living release PR + re-home of open PRs), and closes the freeze once the release PR squash-merges tomain. Before merging any PR, every campaign workflow (/review-prs,/review-group-prs,/merge-prs,/triage-fix-bugs,/implement-fix-bugs,/triage-features,/implement-features,/green-prs,/port-upstream-*) MUST checkgh issue list --repo diegosouzapw/OmniRoute --label release-freeze --state open— if a freeze is active: NEVER merge into the frozenrelease/vX.Y.Znamed in the freeze title; instead resolve the ACTIVE development branch (the highestrelease/v*by semver — normallyrelease/vX+1, announced in a freeze-issue comment) and retarget the PR there (gh pr edit <N> --base release/vX+1, then VERIFY withgh pr view <N> --json baseRefName— the edit fails silently) and merge normally. HOLD only when the highest release/v* branch IS the frozen one (the short window before 0a.0b completes, or a pre-parallel-cycle release) — in that case leave the PR ready and open, tell the operator, and resume when the next branch appears or the freeze lifts. The just-shipped fixes reachrelease/vX+1via the Phase 5 sync-back (scripts/release/sync-next-cycle.mjs); do not try to sync mid-release. This is a coordination signal, not a permission lock: the release captain and the campaign sessions share thediegosouzapwidentity, so a GitHub branch-protection lock cannot distinguish them — only this honored marker prevents the mid-release commit races that forced full CHANGELOG re-reconciliation in v3.8.40/v3.8.41 (a parallel campaign advancedrelease/vX.Y.Zby 34 commits mid-run). The release captain's own reconciliation/cycle-open pushes are exempt — they are the release. Fixes that must land during a freeze (a homologation finding) follow the post-merge read-only rule: land onmainfirst viafix/release-vX.Y.Z-*. ⛔ ONLY/generate-releasemay raise a release-freeze, and ONLY at its Phase 0a (start of generating a new version) — lifted at Phase 12c after the squash-merge tomain. No campaign, session, or agent may open arelease-freezemarker at any other time — a freeze is never a mid-development coordination tool. If a session ever believes a freeze is genuinely, unavoidably necessary outside the/generate-releaseflow, it MUST first ask the operator (diegosouzapw) in chat, explicitly alert "estou criando um freeze" and get an explicit yes — never open, extend, or re-open arelease-freezeautonomously. Conversely, do not close/lift an active/generate-releasefreeze to unblock campaign merges: it protects the captain's single clean CI run and auto-lifts at Phase 12c — closing it early re-triggers the exact commit race it prevents. Verify a freeze is legitimate before acting on it: an openrelease-freezewhose title/body references an OPEN release PR (gh pr view <N> --json state) is the authorized captain freeze — hold, don't touch. (Cycle-model proposal:_tasks/finished/release-flow/2026-07-04_proposta-ciclo-paralelo-v2.md.) - Cross-session safety — this repo is worked by MANY parallel sessions/agents at once; never step on another's in-flight work. Two absolute bans, both recurring incidents (this rule exists because they keep happening):
- (a) Never
git stash/git stash pop— ANYWHERE in this repo, including inside an isolated worktree, and including inside any subagent you dispatch.git stashoperates on the shared repository object store, not the per-worktree working tree — so a stash pushed or popped in one session can silently clobber or resurrect another parallel session's uncommitted changes. This is not hypothetical: 2026-07-02 a#5923quotaCache change leaked into the unrelated#2296worktree via a globalstash pop, and the same class reincided through a subagent. To compare working changes against a base ref without stashing, usegit show <ref>:<path>orgit diff <ref> -- <path>; to confirm a typecheck/lint error is pre-existing on the base, inspect the base ref directly (git show origin/release/vX.Y.Z:<path>) — never stash your tree away to "get it clean". Put this ban verbatim in the prompt of every subagent that touches git (agents don't inherit this file's context — the recurrence was a subagent). - (b) Never merge, push, rebase, or force-push a PR / branch / worktree that another session is actively working. An open PR whose head is a live fix worktree in
.claude/worktrees/you did not create (e.g.fix-5852/fix-5923carrying fresh commits, even when they share yourdiegosouzapwidentity), or any branch another session owns, is off-limits — HOLD, and let the owning session merge it. Before merging or pushing to any PR you did not create this session, rungit worktree listto check for a matching in-flight worktree and re-checkgh pr view <N> --json state,headRefOid. Only the owning session merges its own in-flight PR; mid-flight merges race the owner and re-trigger the exact commit/CHANGELOG races Rule #19 and Rule #21 guard against. (Reinforces Rule #19.)
- (a) Never
PII & Stream Sanitization Learnings
1. Regex Security (ReDoS)
All regex patterns matching variable-length strings (e.g. IPv6 address, credit cards) must use strictly bounded, non-overlapping sequences (e.g., limit occurrences with bounded ranges {1,7}) to prevent catastrophic backtracking when processing untrusted inputs.
2. SSE Snapshot Handling
When parsing streaming LLM responses (e.g. Responses API), check if a chunk represents a final snapshot (done or completed events). Snapshot text must be sanitized directly as a standalone string (bypassing rolling delta buffers) to prevent text duplication at the end of the stream.
3. Database Handles in Tests
Ensure that any unit tests that trigger database migrations or establish SQLite connections call resetDbInstance() and properly clean up/close all DB handles in a test.after(...) hook. Failure to release database connection handles will cause Node's native test runner to hang indefinitely.
Local development access
The dashboard is reachable at the operator's chosen URL/port (default http://localhost:20128). Credentials are operator-specific:
- Initial admin password is read from the
INITIAL_PASSWORDenv var on first install (defaults toCHANGEMEin.env.example; rotate immediately after first login). - Local VPS / shared dev environments: ask the operator for the URL and current credentials — they live in their personal vault, NOT in this repo.
Any credential observed in a previous version of this file was a non-production demo value; treat it as compromised and do not reuse it.