Hard Rule #18: every issue fix must have TDD (failing test → pass) or a documented live VPS test (192.168.0.15) when TDD is not possible. Also clarifies that test:unit + test:vitest must both pass (non-overlapping coverage).
28 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Quick Start
npm install # Install deps (auto-generates .env from .env.example)
npm run dev # Dev server at http://localhost:20128
npm run build # Production build (Next.js 16 standalone)
npm run lint # ESLint (0 errors expected; warnings are pre-existing)
npm run typecheck:core # TypeScript check (should be clean)
npm run typecheck:noimplicit:core # Strict check (no implicit any)
npm run test:coverage # Unit tests + coverage gate (40/40/40/40 — statements/lines/functions/branches)
npm run check # lint + test combined
npm run check:cycles # Detect circular dependencies
Running Tests
# Single test file (Node.js native test runner — most tests)
node --import tsx/esm --test tests/unit/your-file.test.ts
# Vitest (MCP server, autoCombo, cache)
npm run test:vitest
# All suites
npm run test:all
For full test matrix, see CONTRIBUTING.md → "Running Tests". For deep architecture, see AGENTS.md.
Project at a Glance
OmniRoute — unified AI proxy/router. One endpoint, 160+ LLM providers, auto-fallback.
| Layer | Location | Purpose |
|---|---|---|
| API Routes | src/app/api/v1/ |
Next.js App Router — entry points |
| Handlers | open-sse/handlers/ |
Request processing (chat, embeddings, etc) |
| Executors | open-sse/executors/ |
Provider-specific HTTP dispatch |
| Translators | open-sse/translator/ |
Format conversion (OpenAI↔Claude↔Gemini) |
| Transformer | open-sse/transformer/ |
Responses API ↔ Chat Completions |
| Services | open-sse/services/ |
Combo routing, rate limits, caching, etc |
| Database | src/lib/db/ |
SQLite domain modules (45+ files, 55 migrations) |
| Domain/Policy | src/domain/ |
Policy engine, cost rules, fallback logic |
| MCP Server | open-sse/mcp-server/ |
43 tools (30 base + 3 memory + 4 skills + 6 notion), 3 transports, ~13 scopes |
| A2A Server | src/lib/a2a/ |
JSON-RPC 2.0 agent protocol |
| Skills | src/lib/skills/ |
Extensible skill framework |
| Memory | src/lib/memory/ |
Persistent conversational memory |
Monorepo: src/ (Next.js 16 app), open-sse/ (streaming engine workspace), electron/ (desktop app), tests/, bin/ (CLI entry point).
Request Pipeline
Client → /v1/chat/completions (Next.js route)
→ CORS → Zod validation → auth? → policy check → prompt injection guard
→ handleChatCore() [open-sse/handlers/chatCore.ts]
→ cache check → rate limit → combo routing?
→ resolveComboTargets() → handleSingleModel() per target
→ translateRequest() → getExecutor() → executor.execute()
→ fetch() upstream → retry w/ backoff
→ response translation → SSE stream or JSON
→ If Responses API: responsesTransformer.ts TransformStream
API routes follow a consistent pattern: Route → CORS preflight → Zod body validation → Optional auth (extractApiKey/isValidApiKey) → API key policy enforcement → Handler delegation (open-sse). No global Next.js middleware — interception is route-specific.
Combo routing (open-sse/services/combo.ts): 14 strategies (priority, weighted, fill-first, round-robin, P2C, random, least-used, cost-optimized, reset-aware, strict-random, auto, lkgp, context-optimized, context-relay). Each target calls handleSingleModel() which wraps handleChatCore() with per-target error handling and circuit breaker checks. See docs/routing/AUTO-COMBO.md for the 9-factor Auto-Combo scoring and docs/architecture/RESILIENCE_GUIDE.md for the 3 resilience layers.
Resilience Runtime State
OmniRoute has three related but distinct temporary-failure mechanisms. Keep their scope separate when debugging routing behavior. See the 3-layer resilience diagram (source: docs/diagrams/resilience-3layers.mmd) for an at-a-glance map.
Provider Circuit Breaker
Scope: whole provider, e.g. glm, openai, anthropic.
Purpose: stop sending traffic to a provider that is repeatedly failing at the upstream/service level, so one unhealthy provider does not slow down every request.
Implementation:
- Core class:
src/shared/utils/circuitBreaker.ts - Chat gate/execution wiring:
src/sse/handlers/chatHelpers.ts,src/sse/handlers/chat.ts - Runtime status API:
src/app/api/monitoring/health/route.ts - Shared wrappers:
open-sse/services/accountFallback.ts - Persisted state table:
domain_circuit_breakers
States:
CLOSED: normal traffic is allowed.OPEN: provider is temporarily blocked; callers get a provider-circuit-open response or combo routing skips to another target.HALF_OPEN: reset timeout has elapsed; allow a probe request. Success closes the breaker, failure opens it again.
Defaults (open-sse/config/constants.ts):
- OAuth providers: threshold
3, reset timeout60s. - API-key providers: threshold
5, reset timeout30s. - Local providers: threshold
2, reset timeout15s.
Only provider-level failure statuses should trip the provider breaker:
(408, 500, 502, 503, 504);
Do not trip the whole-provider breaker for normal account/key/model errors like most
401, 403, or 429 cases. Those usually belong to connection cooldown or model
lockout. A generic API-key provider 403 should be recoverable unless it is classified
as a terminal provider/account error.
The breaker uses lazy recovery, not a background timer. When OPEN expires, reads such
as getStatus(), canExecute(), and getRetryAfterMs() refresh the state to
HALF_OPEN, so dashboards and combo candidate builders do not keep excluding an
expired provider forever.
Connection Cooldown
Scope: one provider connection/account/key.
Purpose: temporarily skip one bad key/account while allowing other connections for the same provider to continue serving requests.
Implementation:
- Write/update path:
src/sse/services/auth.ts::markAccountUnavailable() - Account selection/filtering:
src/sse/services/auth.ts::getProviderCredentials... - Cooldown calculation:
open-sse/services/accountFallback.ts::checkFallbackError() - Settings:
src/lib/resilience/settings.ts
Important fields on provider connections:
rateLimitedUntil;
testStatus: "unavailable";
lastError;
lastErrorType;
errorCode;
backoffLevel;
During account selection, a connection is skipped while:
new Date(rateLimitedUntil).getTime() > Date.now();
Cooldowns are also lazy: when rateLimitedUntil is in the past, the connection becomes
eligible again. On successful use, clearAccountError() clears testStatus,
rateLimitedUntil, error fields, and backoffLevel.
Default connection cooldown behavior:
- OAuth base cooldown:
5s. - API-key base cooldown:
3s. - API-key
429should prefer upstream retry hints (Retry-After, reset headers, or parseable reset text) when available. - Repeated recoverable failures use exponential backoff:
baseCooldownMs * 2 ** failureIndex;
The anti-thundering-herd guard prevents concurrent failures on the same connection from
repeatedly extending the cooldown or double-incrementing backoffLevel.
Terminal states are not cooldowns. banned, expired, and credits_exhausted are
intended to stay unavailable until credentials/settings change or an operator resets
them. Do not overwrite terminal states with transient cooldown state.
Model Lockout
Scope: provider + connection + model.
Purpose: avoid disabling a whole connection when only one model is unavailable or quota-limited for that connection.
Examples:
- Per-model quota providers returning
429. - Local providers returning
404for one missing model. - Provider-specific mode/model permission failures such as selected Grok modes.
Model lockout lives in open-sse/services/accountFallback.ts and lets the same
connection continue serving other models.
Debugging Guidance
- If all keys for a provider are skipped, inspect both provider breaker state and each
connection's
rateLimitedUntil/testStatus. - If a provider appears permanently excluded after the reset window, check whether code
is reading raw
stateinstead of usinggetStatus()/canExecute(). - If one provider key fails but others should work, prefer connection cooldown over provider breaker.
- If only one model fails, prefer model lockout over connection cooldown.
- If a state should self-recover, it should have a future timestamp/reset timeout and a read path that refreshes expired state. Permanent statuses require manual credential or config changes.
Key Conventions
Code Style
- 2 spaces, semicolons, double quotes, 100 char width, es5 trailing commas (enforced by lint-staged via Prettier)
- Imports: external → internal (
@/,@omniroute/open-sse) → relative - Naming: files=camelCase/kebab, components=PascalCase, constants=UPPER_SNAKE
- ESLint:
no-eval,no-implied-eval,no-new-func= error everywhere;no-explicit-any= warn inopen-sse/andtests/ - TypeScript:
strict: false, target ES2022, module esnext, resolution bundler. Prefer explicit types.
Database
- Always go through
src/lib/db/domain modules — never write raw SQL in routes or handlers - Never add logic to
src/lib/localDb.ts(re-export layer only) - Never barrel-import from
localDb.ts— import specificdb/modules instead - DB singleton:
getDbInstance()fromsrc/lib/db/core.ts(WAL journaling) - Migrations:
src/lib/db/migrations/— versioned SQL files, idempotent, run in transactions
Error Handling
- try/catch with specific error types, log with pino context
- Never swallow errors in SSE streams — use abort signals for cleanup
- Return proper HTTP status codes (4xx/5xx)
Security
- Never use
eval(),new Function(), or implied eval - Validate all inputs with Zod schemas
- Encrypt credentials at rest (AES-256-GCM)
- Upstream header denylist:
src/shared/constants/upstreamHeaders.ts— keep sanitize, Zod schemas, and unit tests aligned when editing - Public upstream credentials (Gemini/Antigravity/Windsurf-style OAuth client_id/secret + Firebase Web keys extracted from public CLIs): MUST be embedded via
resolvePublicCred()fromopen-sse/utils/publicCreds.ts— never as string literals. Seedocs/security/PUBLIC_CREDS.mdfor the mandatory pattern. - Error responses (HTTP / SSE / executor / MCP handler): MUST route through
buildErrorBody()orsanitizeErrorMessage()fromopen-sse/utils/error.ts— never put rawerr.stackorerr.messagein a response body. Seedocs/security/ERROR_SANITIZATION.md. - Shell commands built from variables: when calling
exec()/spawn()with a script that needs runtime values, pass them via theenvoption (shell-escaped automatically) — never string-interpolate untrusted/external paths into the script body. Reference:src/mitm/cert/install.ts::updateNssDatabases. - Secure-by-default libraries (tldrsec/awesome-secure-defaults): prefer Helmet.js, DOMPurify, ssrf-req-filter, safe-regex, Google Tink over custom implementations whenever adding new security-sensitive surfaces.
Common Modification Scenarios
Adding a New Provider
- Register in
src/shared/constants/providers.ts(Zod-validated at load) - Add executor in
open-sse/executors/if custom logic needed (extendBaseExecutor) - Add translator in
open-sse/translator/if non-OpenAI format - Add OAuth config in
src/lib/oauth/constants/oauth.tsif OAuth-based — if the upstream CLI ships a public client_id/secret, embed viaresolvePublicCred()(seedocs/security/PUBLIC_CREDS.md), never as a literal - Register models in
open-sse/config/providerRegistry.ts - Write tests in
tests/unit/(include the publicCreds shape assertion if you added a new embedded default)
Adding a New API Route
- Create directory under
src/app/api/v1/your-route/ - Create
route.tswithGET/POSThandlers - Follow pattern: CORS → Zod body validation → optional auth → handler delegation
- Handler goes in
open-sse/handlers/(import from there, not inline) - Error responses use
buildErrorBody()/errorResponse()fromopen-sse/utils/error.ts(auto-sanitized — never puterr.stackorerr.messageraw in the body). Seedocs/security/ERROR_SANITIZATION.md. - Add tests — including at least one assertion that error responses do not leak stack traces (
!body.error.message.includes("at /"))
Adding a New DB Module
- Create
src/lib/db/yourModule.ts— importgetDbInstancefrom./core.ts - Export CRUD functions for your domain table(s)
- Add migration in
src/lib/db/migrations/if new tables needed - Re-export from
src/lib/localDb.ts(add to the re-export list only) - Write tests
Adding a New MCP Tool
- Add tool definition in
open-sse/mcp-server/tools/with Zod input schema + async handler - Register in tool set (wired by
createMcpServer()) - Assign to appropriate scope(s)
- Write tests (tool invocation logged to
mcp_audittable)
Adding a New A2A Skill
- Create skill in
src/lib/a2a/skills/(5 already exist: smart-routing, quota-management, provider-discovery, cost-analysis, health-report) - Skill receives task context (messages, metadata) → returns structured result
- Register in
A2A_SKILL_HANDLERSinsrc/lib/a2a/taskExecution.ts - Expose in
src/app/.well-known/agent.json/route.ts(Agent Card) - Write tests in
tests/unit/ - Document in
docs/frameworks/A2A-SERVER.mdskill table
Adding a New Cloud Agent
- Create agent class in
src/lib/cloudAgent/agents/extendingCloudAgentBase(3 already exist: codex-cloud, devin, jules) - Implement
createTask,getStatus,approvePlan,sendMessage,listSources - Register in
src/lib/cloudAgent/registry.ts - Add OAuth/credentials handling if needed (
src/lib/oauth/providers/) - Tests + document in
docs/frameworks/CLOUD_AGENT.md
Adding a New Embedded Service
- Create installer in
src/lib/services/installers/{name}.tsmodeled onninerouter.ts(userunNpmfrominstallers/utils.ts— no shell interpolation, hard rule #13). - Register the service in
src/lib/services/bootstrap.ts(add toSERVICES[]array and extendbuildSpawnArgsFactory()). - Add a DB seed row for the new service in
src/lib/db/migrations/(version_managertable,status='not_installed',auto_start=0). - Create 7 API endpoints under
src/app/api/services/{name}/(_lib.ts,install,start,stop,restart,update,status,auto-start). All delegate errors throughcreateErrorResponse(). The sharedlogsendpoint is already wired via[name]/logs/route.ts. - Verify
/api/services/is inLOCAL_ONLY_API_PREFIXESinsrc/server/authz/routeGuard.ts; add a test assertingisLocalOnlyPath()returnstruefor the new prefix if you add one (hard rule #17). - Add a UI tab in
src/app/(dashboard)/dashboard/providers/services/tabs/reusingServiceStatusCard,ServiceLifecycleButtons,ServiceLogsPanel. - Document in
docs/frameworks/EMBEDDED-SERVICES.md(update §1 service table + §4 API reference) anddocs/reference/openapi.yaml. - Write tests: unit (
tests/unit/services/), integration (tests/integration/services/, gated byRUN_SERVICES_INT=1), and updatedocs/ops/RELEASE_CHECKLIST.mdsmoke section.
Adding a New Guardrail / Eval / Skill / Webhook event
- Guardrail:
src/lib/guardrails/→ docs:docs/security/GUARDRAILS.md - Eval suite:
src/lib/evals/→ docs:docs/frameworks/EVALS.md - Skill (sandbox):
src/lib/skills/→ docs:docs/frameworks/SKILLS.md - Webhook event:
src/lib/webhookDispatcher.ts→ docs:docs/frameworks/WEBHOOKS.md
Reference Documentation
For any non-trivial change, read the matching deep-dive first:
| Area | Doc |
|---|---|
| Repo navigation | docs/architecture/REPOSITORY_MAP.md |
| Architecture | docs/architecture/ARCHITECTURE.md |
| Engineering reference | docs/architecture/CODEBASE_DOCUMENTATION.md |
| Auto-Combo (9-factor scoring, 14 strategies) | docs/routing/AUTO-COMBO.md |
| Resilience (3 mechanisms) | docs/architecture/RESILIENCE_GUIDE.md |
| Reasoning replay | docs/routing/REASONING_REPLAY.md |
| Skills framework | docs/frameworks/SKILLS.md |
| Memory system (FTS5 + Qdrant) | docs/frameworks/MEMORY.md |
| Cloud agents | docs/frameworks/CLOUD_AGENT.md |
| Guardrails (PII / injection / vision) | docs/security/GUARDRAILS.md |
| Public upstream credentials (Gemini/etc.) | docs/security/PUBLIC_CREDS.md |
| Error message sanitization | docs/security/ERROR_SANITIZATION.md |
| Evals | docs/frameworks/EVALS.md |
| Compliance / audit | docs/security/COMPLIANCE.md |
| Webhooks | docs/frameworks/WEBHOOKS.md |
| Authorization pipeline | docs/architecture/AUTHZ_GUIDE.md |
| Stealth (TLS / fingerprint) | docs/security/STEALTH_GUIDE.md |
| Agent protocols (A2A / ACP / Cloud) | docs/frameworks/AGENT_PROTOCOLS_GUIDE.md |
| MCP server | docs/frameworks/MCP-SERVER.md |
| A2A server | docs/frameworks/A2A-SERVER.md |
| API reference + OpenAPI | docs/reference/API_REFERENCE.md + docs/reference/openapi.yaml |
| Provider catalog (auto-generated) | docs/reference/PROVIDER_REFERENCE.md |
| Release flow | docs/ops/RELEASE_CHECKLIST.md |
| Embedded services | docs/frameworks/EMBEDDED-SERVICES.md |
Testing
| What | Command |
|---|---|
| Unit tests | npm run test:unit |
| Single file | node --import tsx/esm --test tests/unit/file.test.ts |
| Vitest (MCP, autoCombo) | npm run test:vitest |
| E2E (Playwright) | npm run test:e2e |
| Protocol E2E (MCP+A2A) | npm run test:protocols:e2e |
| Ecosystem | npm run test:ecosystem |
| Coverage gate | npm run test:coverage (40/40/40/40 — statements/lines/functions/branches) |
| Coverage report | npm run coverage:report |
PR rule: If you change production code in src/, open-sse/, electron/, or bin/, you must include or update tests in the same PR.
Test layer preference: unit first → integration (multi-module or DB state) → e2e (UI/workflow only). Encode bug reproductions as automated tests before or alongside the fix.
Both test runners must pass: npm run test:unit (Node native — most tests) AND npm run test:vitest (MCP server, autoCombo, cache) cover non-overlapping files. Both must be green before merging. A PR where only one suite passes may silently ship broken MCP tools or routing regressions.
Bug fix / issue triage protocol (Hard Rule #18): Every fix for a reported issue must be validated by one of the following — no exceptions:
- TDD (preferred) — write a failing test reproducing the bug → fix it → confirm the test passes. The test becomes the permanent regression guard. Touch only the files the test proves need changing; nothing more.
- Real-environment test (when TDD is not possible) — deploy to the production VPS (
root@192.168.0.15) and run a documented live test. Record the exact command + result in the PR description. Applies to: OAuth upstream flows, Cloudflare/WS upstream behavior, UI-only regressions, hardware-dependent behavior. - "It worked locally without a test" does not count. A fix without a test or a VPS validation record is not a fix — it is a guess.
Why this matters: fixing bug A while opening bug B is worse than not fixing at all. The TDD/VPS gate enforces surgical scope — you touch only what the failing test proves is broken. Examples where this paid off: #3090 (claude-web 403), #3113 (WS HTTP fallback), #3052 (heap-guard auto-calibration).
Copilot coverage policy: When a PR changes production code and coverage is below 40% (statements/lines/functions/branches), do not just report — add or update tests, rerun the coverage gate, then ask for confirmation. Include commands run, changed test files, and final coverage result in the PR report.
Git Workflow
# Never commit directly to main
git checkout -b feat/your-feature
git commit -m "feat: describe your change"
git push -u origin feat/your-feature
Branch prefixes: feat/, fix/, refactor/, docs/, test/, chore/
Commit format (Conventional Commits): feat(db): add circuit breaker — scopes: db, sse, oauth, dashboard, api, cli, docker, ci, mcp, a2a, memory, skills
Husky hooks:
- pre-commit: lint-staged +
check-docs-sync+check:any-budget:t11 - pre-push:
npm run test:unit
Environment
- Runtime: Node.js ≥20.20.2 <21 || ≥22.22.2 <23 || ≥24 <25, ES Modules
- TypeScript: 5.9+, target ES2022, module esnext, resolution bundler
- Path aliases:
@/*→src/,@omniroute/open-sse→open-sse/,@omniroute/open-sse/*→open-sse/* - Default port: 20128 (API + dashboard on same port)
- Data directory:
DATA_DIRenv var, defaults to~/.omniroute/ - Key env vars:
PORT,JWT_SECRET,API_KEY_SECRET,INITIAL_PASSWORD,REQUIRE_API_KEY,APP_LOG_LEVEL - Setup:
cp .env.example .envthen generateJWT_SECRET(openssl rand -base64 48) andAPI_KEY_SECRET(openssl rand -hex 32)
Hard Rules
- Never commit secrets or credentials
- Never add logic to
localDb.ts - Never use
eval()/new Function()/ implied eval - Never commit directly to
main - Never write raw SQL in routes — use
src/lib/db/modules - Never silently swallow errors in SSE streams
- Always validate inputs with Zod schemas
- Always include tests when changing production code
- Coverage must stay ≥40% (statements, lines, functions, branches).
- Never bypass Husky hooks (
--no-verify,--no-gpg-sign) without explicit operator approval. - Never embed public upstream OAuth client_id/secret or Firebase Web keys as string literals — always go through
resolvePublicCred()(open-sse/utils/publicCreds.ts). Seedocs/security/PUBLIC_CREDS.md. - Never return raw
err.stack/err.messagein HTTP / SSE / executor responses — always route throughbuildErrorBody()orsanitizeErrorMessage()(open-sse/utils/error.ts). Seedocs/security/ERROR_SANITIZATION.md. - Never string-interpolate external paths or runtime values into shell scripts passed to
exec()/spawn()— pass via theenvoption instead. Reference:src/mitm/cert/install.ts::updateNssDatabases. - Never dismiss a CodeQL / Secret-Scanning alert without (a) first checking the pattern docs above to see if the helper applies, and (b) recording the technical justification in the dismissal comment. Precedent:
js/stack-trace-exposureraised on callsites that already route throughsanitizeErrorMessage()is a known CodeQL limitation (custom sanitizers not recognized) — dismiss asfalse positivereferencingdocs/security/ERROR_SANITIZATION.md. - Never expose routes that spawn child processes (
/api/mcp/,/api/cli-tools/runtime/) withoutisLocalOnlyPath()classification insrc/server/authz/routeGuard.ts. Loopback enforcement happens unconditionally before any auth check — leaked JWT via tunnel cannot trigger process spawning. Seedocs/security/ROUTE_GUARD_TIERS.md. - Never include
Co-Authored-Bytrailers that credit an AI assistant, LLM, or automation account (e.g. names containing "Claude", "GPT", "Copilot", "Bot"; emails atanthropic.com/openai.com/ bot-ownednoreply.github.comaddresses). Such trailers route attribution to the bot account on GitHub, hiding the real author (diegosouzapw) in PR history. Human collaborators — including upstream PR authors and issue reporters being ported into OmniRoute — MAY and SHOULD be credited with standardCo-authored-by: Name <email>trailers; the upstream-port workflows (/port-upstream-features,/port-upstream-issues) depend on this. - Never expose routes under
/api/services/or/dashboard/providers/services/*/embed/withoutisLocalOnlyPath()classification insrc/server/authz/routeGuard.ts. These routes can spawn child processes (npm install,node). Loopback enforcement happens unconditionally before any auth check — a leaked JWT via tunnel cannot trigger process spawning. Seedocs/security/ROUTE_GUARD_TIERS.md. - Every bug fix must be validated before shipping: a failing-then-passing unit/integration test (TDD) OR a documented live test on the production VPS (192.168.0.15). A fix without either is not merged. See Testing → "Bug fix / issue triage protocol" for the full decision tree.
PII & Stream Sanitization Learnings
1. Regex Security (ReDoS)
All regex patterns matching variable-length strings (e.g. IPv6 address, credit cards) must use strictly bounded, non-overlapping sequences (e.g., limit occurrences with bounded ranges {1,7}) to prevent catastrophic backtracking when processing untrusted inputs.
2. SSE Snapshot Handling
When parsing streaming LLM responses (e.g. Responses API), check if a chunk represents a final snapshot (done or completed events). Snapshot text must be sanitized directly as a standalone string (bypassing rolling delta buffers) to prevent text duplication at the end of the stream.
3. Database Handles in Tests
Ensure that any unit tests that trigger database migrations or establish SQLite connections call resetDbInstance() and properly clean up/close all DB handles in a test.after(...) hook. Failure to release database connection handles will cause Node's native test runner to hang indefinitely.