* chore(release): v3.7.9 — gemini-cli cloud code separation * chore(provider): Update Jina AI model catalog (#1874) Integrated into release/v3.7.9 * docs: update CHANGELOG for PR 1874 and retroactive credits * fix: resolve stream defaults and codex prompt mapping (#1873, #1872) * chore(compression): start caveman compression update * feat(compression): expand caveman compression and analytics pipeline Add caveman intensity levels, output mode instructions, validation, and preview diffs across the compression pipeline. Extend MCP and dashboard settings to support auto-trigger mode, system prompt preservation, MCP description compression, and caveman rule metadata. Record richer compression analytics with receipt fields, validation fallbacks, output mode data, and add the related database migration. Improve preservation handling for code, URLs, markdown, math, and other protected content while adding broad unit, integration, and golden-set coverage for caveman parity and compression behavior. * feat(compression): expose rule intensities and track usd savings Add estimated USD savings to compression analytics so saved tokens can be reported in cost terms alongside existing token metrics. Expose caveman rule intensity metadata for settings consumers and add a settings API route alias for rule lookup. Also preserve system prompts when aggressive compression falls back to lite mode. * feat(compression): RTK compression roadmap (#1889) * chore(rtk): initialize compression roadmap branch * feat(compression): add RTK engine and compression combos Introduce RTK command-aware tool-output compression alongside stacked RTK -> Caveman pipelines for mixed prompt contexts. Add engine registration, declarative RTK filter packs, language-aware Caveman rule loading, compression combo persistence and assignments, analytics grouped by engine/combo, and new MCP/API endpoints for configuration, previews, filters, and combo management. Expose the new capabilities in the dashboard with dedicated Context & Cache pages for Caveman, RTK, and compression combos, and update docs, i18n strings, migrations, and tests to cover the expanded compression surface. * feat(compression): expand RTK DSL, filter catalog, and recovery APIs Add RTK parity features across the compression pipeline, dashboard, and management APIs. This expands the built-in filter catalog, adds trust-gated custom filter loading, inline filter verification, code stripping, smarter detection, and optional redacted raw-output retention for authenticated recovery. Also extend Caveman with file-based multilingual rule packs, localized output-mode instructions, stricter preview/config schemas, engine registry metadata, analytics fields, and broad unit test coverage for RTK, rule loading, and stacked compression behavior. * fix(auth): protect oauth routes and health reset operations Require authenticated dashboard access for OAuth endpoints that can create or import provider connections when login enforcement is enabled. Move `/api/monitoring/health` to the readonly public route list so safe methods remain public while DELETE now returns 401 for anonymous requests. Also update Next.js native `.node` handling to avoid webpack parse failures from external packages such as ngrok and keytar, and add coverage for the new auth behavior. * build(compression): ship RTK rule and filter assets with app bundles Include compression JSON assets in Next output tracing, prepublish copies, and pack artifact policy checks so standalone and packaged builds can load RTK filters and caveman rule packs at runtime. Also harden compression runtime behavior by resolving alternate asset directories, scoping rule cache entries by source path, carrying RTK raw output pointers through stacked runs, degrading oversized preview diffs, and applying combo language/output mode defaults during chat routing. Add coverage for packaging rules, provider-scoped model parsing, smart truncate edge cases, raw output retention, and combo-driven compression behavior. * docs(workflows): update local repo paths to OmniRoute Replace outdated `/home/diegosouzapw/dev/proxys/9router` references with the current `OmniRoute` directory across deploy, release, and version bump workflow guides so local command examples match the renamed repository layout * feat(compression): complete RTK parity coverage * test(build): align next config assertions --------- Co-authored-by: diegosouzapw <diego.souza.pw@gmail.com> * feat(compression): expand caveman parity and MCP metadata compression Compress MCP registry and list metadata descriptions for tools, prompts, resources, and resource templates while keeping tool-response bodies unchanged. Expose those savings in compression status as `mcp_metadata_estimate` metadata rather than provider usage. Add Caveman rule-pack support for custom regex flags and match-specific replacement maps, update English rules for upstream parity, and tighten article and pleasantry handling. Also process RTK multipart text blocks independently so mixed media content compresses safely without duplicating output. * feat(compression): unify config validation and persist MCP savings Centralize compression config schemas across settings, preview, RTK, and combo APIs to enforce consistent validation for stacked pipelines and engine-specific options. Expand caveman and stacked compression behavior by applying default combos at runtime, surfacing validation and fallback metadata, and exposing aggressive and ultra adapter schemas for configuration UI and tests. Persist MCP description compression snapshots into analytics without counting them as provider usage, and extend the dashboard with the new RTK controls and localized labels. * fix(auth): require dashboard management auth for compression preview Block preview requests unless they come from a valid management session token so protected settings cannot be probed through the preview API. Add unit coverage for unauthenticated requests, invalid bearer tokens, and successful authenticated preview execution. * fix(compression): preserve stacked defaults and secure metadata routes Only apply saved default compression combos when they contain a stacked pipeline so seeded Caveman-only defaults do not replace the builtin stacked behavior. Also require management auth for compression language pack and rules metadata endpoints, and defer usage receipt attachment until compression analytics writes have completed to keep analytics records consistent. * fix(compression): align seeded standard savings combo with stacked default Update the seeded default compression combo to use the RTK then Caveman pipeline in both fresh installs and upgraded databases. Add a targeted migration and runtime guard that only rewrites the legacy seeded record when its original metadata and single-step pipeline still match, preserving user-customized default combos. Refresh docs and tests to reflect the stacked default and expanded RTK filter catalog. * docs(compression): document RTK+Caveman stacked savings ranges Refresh the compression docs and README to describe the default stacked pipeline in terms of eligible-context savings instead of the older generic token-saving range. Add upstream RTK and Caveman benchmark references, explain the multiplicative savings math behind the stacked default, and update feature summaries plus package metadata to match the revised positioning. * feat(image-gen): add NanoGPT image generation provider (#1899) Integrated into release/v3.7.9 * fix(codex): sanitize raw responses input (#1895) Integrated into release/v3.7.9 * Fix combo provider breaker profile handling (#1891) Integrated into release/v3.7.9 * fix(combos): align strategy contracts (#1892) Integrated into release/v3.7.9 * feat(proxy): move proxy configuration to dedicated System → Proxy page (#1907) Integrated into release/v3.7.9 * feat: add K/M/B/T cost shortener to prevent UI overflow (#1902) Integrated into release/v3.7.9 * feat(providers): implement bulk paste for extra API keys (#1916) Integrated into release/v3.7.9 * fix(migrations): treat duplicate-column ALTER as no-op (#1886) Integrated into release/v3.7.9 * fix(analytics): robust model pricing resolution, dark mode charts and SQL aggregation fixes (#1896) Integrated into release/v3.7.9 (migration renumbered to 044) * fix(oauth): per-connection mutex for rotating refresh tokens (#1885) Integrated into release/v3.7.9 * fix: resolve 3 bugs — Codex tool normalization (#1914), image gen proxy (#1904), zero-arg MCP tools (#1898) - fix(codex): flatten Chat Completions tool format to Responses format in normalizeCodexTools. Prevents 'Missing required parameter: tools[0].name' upstream errors when clients send {type:'function', function:{name,...}} instead of {type:'function', name,...}. - fix(proxy): add proxy-aware execution context to image generation route. Image requests now correctly use proxy settings from the connection's ProxyRegistry assignment, matching the pattern used by chat pipeline. - fix(translator): inject properties:{} into zero-argument MCP tool schemas during Anthropic→OpenAI translation. OpenAI strict mode requires explicit properties even for empty object schemas. Closes #1914, Closes #1904, Closes #1898 * chore(release): v3.7.9 — all changes in ONE commit * fix: allow local ollama provider connections (#1893) * fix(copilot): emit compatible reasoning text deltas (#1919) Integrated into release/v3.7.9 * fix(api-manager): show validation errors inline in modals, not behind backdrop (#1920) Integrated into release/v3.7.9 * docs: update changelog and pr body with merged prs * fix(providers): route agentrouter through anthropic endpoint headers Update the AgentRouter provider registry to use the Claude-compatible messages API and required Anthropic-style authentication headers. This bypasses unauthorized_client_error responses and exposes the supported model list through passthrough configuration. Also update the changelog and release PR notes to document the fix for #1921 * feat(logs): show compression tokens in request log UI (#1923) * docs: add PR #1923 to changelog --------- Co-authored-by: diegosouzapw <diego.souza.pw@gmail.com> Co-authored-by: backryun <bakryun0718@proton.me> Co-authored-by: Aculeasis <42580940+Aculeasis@users.noreply.github.com> Co-authored-by: Raxxoor <manker_lol@hotmail.com> Co-authored-by: Randi <55005611+rdself@users.noreply.github.com> Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com> Co-authored-by: Tubagus <54710482+0xtbug@users.noreply.github.com> Co-authored-by: smartenok-ops <smartenok@gmail.com> Co-authored-by: Gi99lin <74502520+Gi99lin@users.noreply.github.com> Co-authored-by: ivan-mezentsev <ivan@mezentsev.me> Co-authored-by: Andrew Munsell <andrew@wizardapps.net>
23 KiB
omniroute — Agent Guidelines
Project
Unified AI proxy/router — route any LLM through one endpoint. Multi-provider support with 160+ providers (OpenAI, Anthropic, Gemini, DeepSeek, Groq, xAI, Mistral, Fireworks, Cohere, NVIDIA, Cerebras, Pollinations, Puter, Cloudflare AI, HuggingFace, DeepInfra, SambaNova, Meta Llama API, Moonshot AI, AI21 Labs, Databricks, Snowflake, and many more) with MCP Server (37 tools), A2A v0.3 Protocol, and Electron desktop app.
Stack
- Runtime: Next.js 16 (App Router), Node.js
>=20.20.2 <21,>=22.22.2 <23, or>=24.0.0 <25, ES Modules ("type": "module") - Language: TypeScript 5.9 (
src/) + JavaScript (open-sse/,electron/) - Database: better-sqlite3 (SQLite) —
DATA_DIRconfigurable, default~/.omniroute/ - Streaming: SSE via
open-sseinternal workspace package - Styling: Tailwind CSS v4
- i18n: next-intl with 40+ languages
- Desktop: Electron (cross-platform: Windows, macOS, Linux)
- Schemas: Zod v4 for all API / MCP input validation
Build, Lint, and Test Commands
| Command | Description |
|---|---|
npm run dev |
Start Next.js dev server |
npm run build |
Production build (isolated) |
npm run start |
Run production build |
npm run build:cli |
Build CLI package |
npm run lint |
ESLint on all source files |
npm run typecheck:core |
TypeScript core type checking |
npm run typecheck:noimplicit:core |
Strict checking (no implicit any) |
npm run check |
Run lint + test |
npm run check:cycles |
Check for circular dependencies |
npm run electron:dev |
Run Electron app in dev mode |
npm run electron:build |
Build Electron app for current OS |
Running Tests
# All tests (unit + vitest + ecosystem + e2e)
npm run test:all
# Single test file (Node.js native test runner — most tests use this)
node --import tsx/esm --test tests/unit/your-file.test.ts
node --import tsx/esm --test tests/unit/plan3-p0.test.ts
node --import tsx/esm --test tests/unit/fixes-p1.test.ts
node --import tsx/esm --test tests/unit/security-fase01.test.ts
# Integration tests
node --import tsx/esm --test tests/integration/*.test.ts
# Vitest (MCP server, autoCombo)
npm run test:vitest
# E2E with Playwright
npm run test:e2e
# Protocol clients E2E (MCP transports, A2A)
npm run test:protocols:e2e
# Ecosystem compatibility tests
npm run test:ecosystem
# Coverage (see CONTRIBUTING.md)
npm run test:coverage
For authoritative coverage requirements, test execution, and PR gates, see CONTRIBUTING.md.
Code Style Guidelines
Formatting (Prettier — enforced via lint-staged)
2 spaces · semicolons required · double quotes (") · 100 char width · es5 trailing commas.
Always run prettier --write on changed files.
TypeScript
- Target: ES2022 · Module:
esnext· Resolution:bundler strict: false— prefer explicit types, don't rely on inference- Path aliases:
@/*→src/,@omniroute/open-sse→open-sse/,@omniroute/open-sse/*→open-sse/*
ESLint Rules
- Security (error, everywhere):
no-eval,no-implied-eval,no-new-func - Relaxed in
open-sse/andtests/:@typescript-eslint/no-explicit-any= warn - React hooks rules and
@next/next/no-assign-module-variabledisabled inopen-sse/andtests/
Naming
| Element | Convention | Example |
|---|---|---|
| Files | camelCase / kebab-case | chatCore.ts, tokenHealthCheck.ts |
| React components | PascalCase | Dashboard.tsx, ProviderCard.tsx |
| Functions/variables | camelCase | getHealth(), switchCombo() |
| Constants | UPPER_SNAKE | MAX_RETRIES, DEFAULT_TIMEOUT |
| Interfaces | PascalCase (I prefix optional) |
ProviderConfig |
| Enums | PascalCase (members too) | LogLevel.Error |
Imports
- Order: external → internal (
@/,@omniroute/open-sse) → relative (./,../) - No barrel imports from
localDb.ts— import from the specificdb/module instead
Error Handling
- try/catch with specific error types; always log with context (pino logger)
- Never silently swallow errors in SSE streams — use abort signals for cleanup
- Return proper HTTP status codes (4xx client, 5xx server)
Security
- NEVER commit API keys, secrets, or credentials
- Validate all user inputs with Zod schemas
- Auth middleware required on all API routes
- Never log SQLite encryption keys
- Sanitize user content (dompurify for HTML)
Architecture
Data Layer (src/lib/db/)
All persistence uses SQLite through domain-specific modules:
core.ts, providers.ts, models.ts, combos.ts, apiKeys.ts, settings.ts,
backup.ts, proxies.ts, prompts.ts, webhooks.ts, detailedLogs.ts,
domainState.ts, registeredKeys.ts, quotaSnapshots.ts, modelComboMappings.ts,
cliToolState.ts, encryption.ts, readCache.ts, secrets.ts, stateReset.ts,
contextHandoffs.ts, compression.ts.
Schema migrations live in db/migrations/ and run via migrationRunner.ts.
src/lib/localDb.ts is a re-export layer only — never add logic there.
DB Internals
core.ts:getDbInstance()returns a singletonbetter-sqlite3instance with WAL journaling.SCHEMA_SQLdefines 15 base tables. Helpers:rowToCamel,encryptConnectionFields.migrationRunner.ts: Applies versioned SQL files fromdb/migrations/inside transactions. Tracks applied migrations in_omniroute_migrationstable.- Migrations: 22 files (
001_initial_schema.sql→022_compression_settings.sql). Each migration is idempotent and runs in a transaction. - Domain modules import
getDbInstance()fromcore.tsfor all CRUD operations. Each module owns a specific table/set of tables (e.g.,providers.ts→provider_connections,combos.ts→combos). Encryption helpers protect sensitive fields at rest. localDb.tsre-exports all domain modules — consumers import from here for convenience.
API Route Layer (src/app/api/v1/)
Next.js App Router routes — each follows a consistent pattern:
Route → CORS preflight → Body validation (Zod) → Optional auth (extractApiKey/isValidApiKey)
→ API key policy enforcement (enforceApiKeyPolicy) → Handler delegation (open-sse)
| Route | Handler | Notes |
|---|---|---|
chat/completions/route.ts |
handleChat() |
+ prompt injection guard (clones request) |
responses/route.ts |
handleChat() (unified) |
Responses API format |
embeddings/route.ts |
handleEmbedding() |
Model listing + creation |
images/generations/route.ts |
handleImageGeneration() |
Model listing + creation |
audio/transcriptions/route.ts |
audio handler | Multipart form data |
audio/speech/route.ts |
TTS handler | Binary audio response |
videos/generations/route.ts |
video handler | ComfyUI/SD WebUI |
music/generations/route.ts |
music handler | ComfyUI workflows |
moderations/route.ts |
moderation handler | Content safety |
rerank/route.ts |
rerank handler | Document relevance |
search/route.ts |
search handler | Web search (5 providers) |
No global Next.js middleware file — interception is route-specific. Auth is optional
(controlled by REQUIRE_API_KEY env). Prompt injection guard is unique to chat completions.
Request Pipeline (open-sse/)
The open-sse/ workspace is the core streaming engine. Full request flow:
Client Request
→ src/app/api/v1/.../route.ts (Next.js route)
→ open-sse/handlers/chatCore.ts::handleChatCore()
→ Semantic/signature cache check
→ Rate limit check (rateLimitManager)
→ Combo routing? → open-sse/services/combo.ts::handleComboChat()
→ resolveComboTargets() → ordered ResolvedComboTarget[]
→ For each target: handleSingleModel() (wraps chatCore)
→ translateRequest() (open-sse/translator/)
→ Convert source format (e.g., OpenAI) → target format (e.g., Claude)
→ getExecutor() → provider-specific executor instance
→ executor.execute() (BaseExecutor → DefaultExecutor or provider-specific)
→ buildUrl() + buildHeaders() + transformRequest()
→ fetch() to upstream provider
→ Retry logic with exponential backoff
→ Response translation back to client format
→ If Responses API: responsesTransformer.ts TransformStream
→ SSE stream or JSON response to client
Handlers (open-sse/handlers/): chatCore.ts, responsesHandler.ts, embeddings.ts,
imageGeneration.ts, videoGeneration.ts, musicGeneration.ts, audioSpeech.ts,
audioTranscription.ts, moderations.ts, rerank.ts, search.ts.
Upstream headers: merged after default auth; same header name replaces executor value.
T5 intra-family fallback recomputes headers using only the fallback model id.
Forbidden header names: src/shared/constants/upstreamHeaders.ts — keep sanitize,
Zod schemas, and unit tests aligned when editing.
Provider Categories
- Free (4): Qoder AI, Qwen Code, Gemini CLI (deprecated), Kiro AI
- OAuth (8): Claude Code, Antigravity, Codex, GitHub Copilot, Cursor, Kimi Coding, Kilo Code, Cline
- API Key (120+): OpenAI, Anthropic, Gemini, DeepSeek, Groq, xAI, Mistral, Perplexity, Together, Fireworks, Cerebras, Cohere, NVIDIA, Nebius, SiliconFlow, Hyperbolic, HuggingFace, OpenRouter, Vertex AI, Cloudflare AI, Scaleway, AI/ML API, Pollinations, Puter, Longcat, Alibaba, Kimi, Minimax, Blackbox, Synthetic, Kilo Gateway, Z.AI, GLM, Deepgram, AssemblyAI, ElevenLabs, Cartesia, PlayHT, Inworld, NanoBanana, SD WebUI, ComfyUI, Ollama Cloud, Perplexity Search, Serper, Brave, Exa, Tavily, OpenCode Zen/Go, Bailian Coding Plan, DeepInfra, Vercel AI Gateway, Lambda AI, SambaNova, nScale, OVHcloud AI, Baseten, PublicAI, Moonshot AI, Meta Llama API, v0 (Vercel), Morph, Featherless AI, FriendliAI, LlamaGate, Galadriel, Weights & Biases Inference, Volcengine, AI21 Labs, Venice.ai, Codestral, Upstage, Maritalk, Xiaomi MiMo, Inference.net, NanoGPT, Predibase, Bytez, Heroku AI, Databricks, Snowflake Cortex, GigaChat (Sber), CrofAI, AgentRouter, ChatGPT Web, Baidu Qianfan, AWS Polly, RunwayML, GitLab Duo, Amazon Q, Empower, Poe, and many more.
- Self-Hosted (8+): LM Studio, vLLM, Lemonade, Llamafile, Triton, Docker Model Runner, Xinference, Oobabooga
- Custom: OpenAI-compatible (
openai-compatible-*) and Anthropic-compatible (anthropic-compatible-*) prefixes
Providers are registered in src/shared/constants/providers.ts with Zod validation at module load.
Executors (open-sse/executors/)
Provider-specific request executors: base.ts, default.ts, cursor.ts, codex.ts,
antigravity.ts, github.ts, gemini-cli.ts, kiro.ts, qoder.ts, vertex.ts,
cloudflare-ai.ts, opencode.ts, pollinations.ts, puter.ts.
Executor Internals
base.ts(BaseExecutor): Abstract base withbuildUrl(),buildHeaders(),transformRequest(), retry logic (exponential backoff), andexecute(). Subclasses override URL/header/transform methods for provider-specific behavior.default.ts(DefaultExecutor extends BaseExecutor): Handles most OpenAI-compatible providers. Reads provider config fromproviderRegistry.tsto resolve base URL, auth header format, and request transformations.getExecutor()(executors/index.ts): Factory that returns the correct executor instance based on provider ID. Provider-specific executors (Cursor, Codex, Vertex, etc.) override only what differs from the default.
Translator (open-sse/translator/)
Translates between API formats (OpenAI-format ↔ Anthropic, Gemini, etc.). Includes request/response translators with helpers for image handling.
Translator Internals
translator/index.ts: ExportstranslateRequest()and format constants. Called bychatCore.tsbefore executor dispatch.- Flow:
translateRequest(body, sourceFormat, targetFormat)→ detects source format (OpenAI, Anthropic, Gemini) → applies the matching translator module → returns transformed body ready for the target provider. - Response translation runs in reverse after upstream response, converting back to the client's expected format.
Transformer (open-sse/transformer/)
responsesTransformer.ts — transforms Responses API format to/from Chat Completions format.
Transformer Internals
createResponsesApiTransformStream(): Returns aTransformStreamthat converts Chat Completions SSE chunks (data: {"choices":[...]}) into Responses API SSE events (response.output_item.added,response.output_text.delta, etc.).- Used when the client sends a Responses API request: the request is internally converted to Chat Completions format, dispatched normally, and the response is piped through this transform stream before reaching the client.
Services (open-sse/services/)
36+ service modules including: combo.ts (routing engine), usage.ts, tokenRefresh.ts,
rateLimitManager.ts, accountFallback.ts, sessionManager.ts, wildcardRouter.ts,
autoCombo/, intentClassifier.ts, taskAwareRouter.ts, thinkingBudget.ts,
contextManager.ts, modelDeprecation.ts, modelFamilyFallback.ts,
emergencyFallback.ts, workflowFSM.ts, backgroundTaskDetector.ts, ipFilter.ts,
signatureCache.ts, volumeDetector.ts, contextHandoff.ts, compression/ (prompt
compression pipeline), and more.
Prompt Compression Pipeline (compression/)
Modular prompt compression that runs proactively before the existing reactive context manager.
strategySelector.ts: Selects compression mode based on config, compression combo assignments, combo overrides, auto-trigger thresholds, and defaults. Priority: assigned compression combo > combo override > auto-trigger > default mode > off.lite.ts: 5 lite-mode techniques:collapseWhitespace,dedupSystemPrompt,compressToolResults,removeRedundantContent,replaceImageUrls. Target: 10-15% savings at <1ms latency.caveman.ts/cavemanRules.ts: Caveman-style semantic condensation backed by built-in rules plus file-loaded language packs undercompression/rules/.engines/rtk/: Rule-based terminal/tool-output compression inspired by RTK patterns. Detects command output classes, applies JSON filter packs, deduplicates repeated lines, strips ANSI/code noise, and preserves errors/actionable context. The RTK JSON DSL supports replace, match-output short-circuit, strip/keep, per-line truncation, head/tail/max-line truncation, inline tests, trust-gated project/global custom filters, and optional redacted raw-output retention for authenticated recovery.engines/registry.ts: Registers engines (caveman,rtk) and powers stacked pipelines.stats.ts: Per-request compression stats tracking (original tokens, compressed tokens, savings %, techniques used, engine breakdown, compression combo id).types.ts:CompressionMode(off/lite/standard/aggressive/ultra/rtk/stacked),CompressionConfig,CompressionStats,CompressionResult.- DB settings in
src/lib/db/compression.ts, compression combos insrc/lib/db/compressionCombos.ts, API routes undersrc/app/api/settings/compression/,src/app/api/context/*, and preview/language-pack routes undersrc/app/api/compression/*.
Combo Routing Engine (combo.ts)
handleComboChat(): Entry point for combo-routed requests. Receives the combo config and iterates through targets in order until one succeeds or all fail.resolveComboTargets(): Expands a combo configuration into an ordered array ofResolvedComboTarget[], each specifying provider + model + account + credentials.- Strategies (13): priority, weighted, fill-first, round-robin, P2C, random, least-used, cost-optimized, strict-random, auto, lkgp, context-optimized, context-relay.
- Each target calls
handleSingleModel()which wrapshandleChatCore()with per-target error handling and circuit breaker checks.
Domain Layer (src/domain/)
Policy engine modules: policyEngine.ts, comboResolver.ts, costRules.ts,
degradation.ts, fallbackPolicy.ts, lockoutPolicy.ts, modelAvailability.ts,
providerExpiration.ts, quotaCache.ts, responses.ts, configAudit.ts.
MCP Server (open-sse/mcp-server/)
37 tools, 3 transports (stdio / SSE / Streamable HTTP). Scoped auth (10 scopes), Zod schemas.
Core tools (20): get_health, list_combos, get_combo_metrics, switch_combo, check_quota, route_request, cost_report, list_models_catalog, web_search, simulate_route, set_budget_guard, set_routing_strategy, set_resilience_profile, test_combo, get_provider_metrics, best_combo_for_task, explain_route, get_session_snapshot, db_health_check, sync_pricing.
Cache tools (2): cache_stats, cache_flush.
Compression tools (5): compression_status, compression_configure, set_compression_engine, list_compression_combos, compression_combo_stats.
1proxy tools (3): oneproxy_fetch, oneproxy_rotate, oneproxy_stats.
Memory tools (3): memory_search, memory_add, memory_clear.
Skill tools (4): skills_list, skills_enable, skills_execute, skills_executions.
MCP Internals
- Tool registration: Each tool is an object with
{ name, description, inputSchema: ZodSchema, handler: async (args) => {...} }. Zod validates inputs before the handler fires. createMcpServer()andstartMcpStdio()exported frommcp-server/index.ts.createMcpServer()wires all tool sets;startMcpStdio()launches the stdio transport.- Transports: stdio (CLI
omniroute --mcp), SSE (/api/mcp/sse), Streamable HTTP (/api/mcp/stream). All share the same tool/scope engine. - Scopes (10): Control which tool categories an API key can access. Enforcement happens before handler dispatch.
- Audit: Every tool invocation is logged to SQLite (
mcp_audittable) with tool name, args, success/failure, API key attribution, and timestamp.
A2A Server (src/lib/a2a/)
JSON-RPC 2.0, SSE streaming, Task Manager with TTL cleanup.
Agent Card at /.well-known/agent.json.
Skills: quotaManagement.ts, smartRouting.ts.
A2A Internals
taskManager.ts: State machine lifecycle for tasks:submitted → working → completed | failed | canceled. Tasks have TTL and are cleaned up automatically.- JSON-RPC methods:
message/send(sync),message/stream(SSE),tasks/get,tasks/cancel. Dispatched viaPOST /a2a. - Skills: Registered in a DB-backed registry. Each skill receives task context
(messages, metadata) and returns structured results.
quotaManagement.tssummarizes quota;smartRouting.tsrecommends routing decisions. - Agent Card:
/.well-known/agent.jsonexposes capabilities, skills, and metadata for client auto-discovery.
ACP Module (src/lib/acp/)
Agent Communication Protocol registry and manager.
Memory System (src/lib/memory/)
Extraction, injection, retrieval, summarization, and store modules for persistent conversational memory across sessions.
Skills System (src/lib/skills/)
Extensible skill framework: registry, executor, sandbox, built-in skills, custom skill support, interception, and injection.
Skills Internals
registry.ts: DB-backed skill registration and discovery. Skills have metadata (name, description, version, enabled status) stored in SQLite.executor.ts: Execution engine with configurable timeout and retry logic. Receives skill name + input, looks up the skill, runs it in the sandbox.sandbox.ts: Isolation layer for custom (user-provided) skills. Limits resource access and execution time.- Built-in skills: Ship with OmniRoute (e.g., quota management, routing). Located alongside the registry.
- Interception/Injection: Skills can intercept requests in the pipeline (pre/post processing) or inject context into prompts.
Compliance (src/lib/compliance/)
Policy index for compliance enforcement.
MITM Proxy (src/mitm/)
MITM proxy capability with certificate management, DNS handling, and target routing.
Middleware (src/middleware/)
Request middleware including promptInjectionGuard.ts.
Adding a New Provider
- Register in
src/shared/constants/providers.ts - Add executor in
open-sse/executors/(if custom logic needed) - Add translator in
open-sse/translator/(if non-OpenAI format) - Add OAuth config in
src/lib/oauth/constants/oauth.ts(if OAuth-based) - Add models in
open-sse/config/providerRegistry.ts
Subdirectory AGENTS.md Files
open-sse/AGENTS.md— Streaming engine, request pipeline, handlers, and executorssrc/lib/db/AGENTS.md— SQLite persistence, domain modules, migrationsopen-sse/services/AGENTS.md— Routing engine, combo resolution, strategy selection
Review Focus
- DB ops go through
src/lib/db/modules, never raw SQL in routes - Provider requests flow through
open-sse/handlers/ - MCP/A2A pages are tabs inside
/dashboard/endpoint, not standalone routes - No memory leaks in SSE streams (abort signals, cleanup)
- Rate limit headers must be parsed correctly
- All API inputs validated with Zod schemas
- Provider constants validated at module load via Zod (
src/shared/validation/providerSchema.ts) - Pricing data syncs from LiteLLM via
src/lib/pricingSync.ts - Memory/Skills are cross-cutting: affect MCP tools, request pipeline, and A2A skills
- ⛔ NEVER close a contributor's PR after using their code — always merge via GitHub so they get credit. See
.agents/workflows/review-prs.mdfor full policy.