* test(base): realign six suites with contracts that #9100/#8990/#9009 deliberately changed Continuing the base-red drain — every one of these reproduces on the pure tip. - tests/snapshots/provider/translate-path.json: regenerated via UPDATE_GOLDEN=1. The diff is ADDITION-ONLY — the unorouter block from #9009; no existing provider entry changed. 3/3. - tests/unit/provider-models-route.test.ts:ff012ff420added onboardUser as a bootstrap fallback next to loadCodeAssist; the mock now excludes it from the discovery-URL ledger like it already excluded loadCodeAssist, otherwise it consumed the injected 503 and the retry assertion misfired. 59/59. - tests/unit/responses-commentary-passthrough-6199.test.ts: #8990 (c996dc93c2) deliberately preserves `tools` on the TERMINAL response.completed snapshot (Codex CLI rebuilds its tool list from it); the assertion now pins the echoed tools instead of their absence. Still stripped on created/in_progress. 7/7. - tests/unit/vision-compression-authoritative-capability-7237.test.ts:68cb678780added the 'gpt-5' fragment, so the heuristic-vs-spec DRIFT this suite documented no longer exists; the cases now guard the agreement, keep a conservative-for-unknown-ids probe, and reproduce the strip-bug shape with an explicit false instead of deriving it. 4/4. - tests/unit/provider-limits-proxy-fail-closed.test.ts + tests/unit/image-generation-route.test.ts: #9100 made the proxy reachability probe NON-BLOCKING (optimistic dispatch; the probe aborts only in-flight requests — its own t14 sibling was updated to this exact pattern). Instant mocks therefore won the race and the PROXY_UNREACHABLE 503 became unobservable (a success or a generic 502). The mocks now stay in flight (never-resolving, so the aborted continuation cannot reach the restored real fetch), and the fail-closed proof is the settled rejection itself plus zero egress AFTER the fast-fail. Production fail-closed semantics are unchanged — the proxy dispatch path still throws; only the mock timing was stale. 3/3 and 20/20. Refs #9298 * fix(guardrails): forward the router deps seam through callVisionModel tests/unit/guardrails/vision-bridge-sse-and-reasoning.test.ts was 7/7 red on any clean box (CI shard 3/4): callVisionModel() called getBestVisionModel()/ getFallbackModels() WITHOUT the routers' existing VisionBridgeRouterDeps seam, so the credential check always hit the live connections DB — no vision-capable connection meant 'No vision-capable provider connected' before the mocked fetch was ever reached, and on a dev box auto-selection could swap the fixed model under the assertions. The routers already accepted deps; only the forwarding was missing. Added the optional 5th param (backward compatible — the sole production caller, visionBridge.ts, injects its own callVisionModel and is unaffected) and the suite now pins selection with hasUsableCredentials: async () => null (indeterminate → the fixed model is honored, DB untouched). 7/7. Sibling suites re-run green: vision-bridge-callmodel 2/2, visionBridge 25/25, visionBridgeHelpers.callVisionModel 8/8, visionBridgeRouter 10/10, vision-bridge-cc-no-reroute 8/8. Refs #9298 * fix(db,combo): clear the NEW base-reds the 08-06 merge batch introduced The tip moved while the first sweep PR (#9600) was in review, and three fresh base-reds landed with it — same classes as before, all reproduced on the pure tip9995bc4893: 1. ANOTHER migration collision: #9061 shipped 134_ccr_blocks.sql onto the slot 134_proxy_logs_egress_ip.sql (#9291) has held since 08-04. getMigrationFiles() throws on collision, so every DB-touching test died at bootstrap again. Renumbered to 139 (next free slot). No retroactive guard needed this time: both statements are IF NOT EXISTS, and no DB can have applied it as 134 — the runner refused to run at all while the collision existed. 2. BROKEN IMPORT killing the combo module graph: #8894 imported preferAntigravityConnectionsWithStoredProject from ../antigravityProjectPersistence.ts — a module that exists NOWHERE in the repo (it came from an unmerged sibling branch). Anything importing quotaStrategies.ts died with ERR_MODULE_NOT_FOUND. Implemented the helper in the real persistence module (antigravityProjectPersist.ts, #8491) with the semantics the call site needs — prefer connections that already carry a stored projectId, never emptying the pool — and pointed the import there. New regression suite tests/unit/antigravity-prefer-stored-project.test.ts (5/5), including an import-graph probe that reproduces the break shape. 3. Sibling-test drift from #9106 (gemini-3.1-pro-high now user-callable): its own suites were updated but provider-models-route.test.ts was not. Expected discovery list realigned; testFrozen 1784->1787 justified in the baseline (irreducible +2 after comment compression; gate counts split-newlines). Also regenerated tests/snapshots/provider/translate-path.json — addition-only: devin-cli-agentic, raycast, regolo (today's provider merges), zero removals. image-generation-route 20/20 (was import-dead), provider-models-route 59/59, antigravity-prefer-stored-project 5/5, provider-translate-path-golden 3/3. Refs #9298 * fix(changelog): convert the #9415 fragment to the required bullet shape Another base-red from the 08-06 batch:bd4407cb64landed changelog.d/features/9415-newapi-sub2api-aggregator-balance.md as YAML frontmatter + a prose paragraph. Every other fragment in changelog.d/ is a single markdown bullet, and both consumers enforce that — scripts/check/check-changelog-integrity.mjs:97 and the release aggregator (scripts/release/aggregate-changelog.mjs:57) reject anything that does not start with '- ', so 'Merge integrity (changelog + generated skills)' was red for every PR targeting the release branch. Rewritten as a bullet with the standard issue link, preserving the feature description (aggregator gateway toggle, /api/user/self balance read, dashboard badge, quota-preflight skip, NEWAPI_AGGREGATOR_BALANCE flag default off, quotaPerUnit override). Swept the rest of changelog.d/ — this was the only malformed fragment. check:changelog-integrity OK. Refs #9298 * fix(types,docs): clear the 5 typecheck errors and the fabricated env vars on the base Third pass over the base-reds, from the 2026-08-06T22:51Z verdict on #9298 — it reported "Typecheck (core)" with only the FIRST error; there are five, all on the pure tip9995bc4893. Two are real production defects. **Real bugs** - open-sse/services/compression/engines/ccr/index.ts:295 called enforceGlobalBudget(entry.bytes) against an (owner, bytes) signature. The `bytes` argument arrived undefined, so `ccrTotalBytes + undefined` is NaN, `NaN > MAX` is false (the eviction loop exits immediately) and `NaN <= MAX` is false (the re-admit is refused). The #9061 durable tier therefore NEVER repopulated its in-memory map: every retrieve after a restart or an eviction re-read from SQLite forever, and evictions could not prefer the owning principal. Fixed and pinned by a new case in tests/unit/ccr-durable-store-9061.test.ts (11/11) — verified failing against the buggy call and passing against the fix. - open-sse/services/combo/fusionPanel.ts:54 read `step.model` after #8894 widened ComboStep with ComboProviderWildcardStep (which carries modelPattern, not model), so a wildcard step in a fusion panel pushed `undefined` onto the panel. Now resolved through getComboModelString(), which already handles every step shape and returns null for the ones without a concrete model id. **Type-only** - accountSemaphore.ts:203 — isBypassed() returns a plain boolean and cannot narrow `number | null` (an `x is null | undefined` predicate would be unsound: 0 bypasses too). Added resolveActiveCap(), the narrowing companion isBypassed is now defined in terms of; the acquire path uses the narrowed value. - comboStructure.ts:140 — same #8894 widening: `prompt` only exists on a model step, so it is now read under a kind check. - firecrawlQuotaFetcher.ts:136 — the function returns full FirecrawlQuota objects but was annotated Promise<QuotaInfo | null>, which made the custom-base literal an excess-property error. Widened to the accurate type (FirecrawlQuota extends QuotaInfo, so callers are unaffected). **Fabricated docs (the "Docs sync + fabricated-docs (strict)" HARD failure)** docs/ops/VM_DEPLOYMENT_GUIDE.md recommended OMNIROUTE_MAX_POOL_SIZE and OMNIROUTE_DB_POOL_SIZE (#9471). Neither is read anywhere in the codebase. Replaced with the two knobs that do exist and are already documented in ENVIRONMENT.md: OMNIROUTE_MEMORY_MB and OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT. typecheck:core 5 errors -> 0. check:fabricated-docs + check:env-doc-sync OK. accountSemaphore 6/6, ccr-durable-store 11/11, ccr-protocol 9/9, combo-fusion-strategy 10/10, combo-fusion-comboref 5/5, combo-fusion-warn 4/4, firecrawl-executor 7/7, executor-firecrawl-fetch 4/4. Refs #9298 * fix(tests): type the #3440 vertex helpers instead of `any` (the 3 base ESLint errors) The "ESLint errors: 3 error(s)" HARD failure in the #9298 verdict is tests/unit/vertex-functioncall-id-3440.test.ts lines 32/41/50: the three find*(result: any) walkers. `@typescript-eslint/no-explicit-any` is an ERROR in tests/ (and open-sse/) since #6218, and this file landed on 2026-08-04 without a suppressions entry, so every run of `lint:json --max-warnings 0` failed. That step prints nothing on failure, which is why the gate looked like a silent crash across the open PRs. Replaced with a GeminiRequestLike interface describing exactly what the three walkers traverse (contents[].parts[]), so the assertions keep their meaning and nothing is cast away. eslint on the file: clean. Suite: 6/6. Refs #9298 * docs(proxy): use an RFC 5737 documentation IP in the proxy examples The #9298 verdict headlines its docs failure with `L810 [stale-version] 1.2.3: const removed = await failOneproxyProxy("1.2.3.4", 8080)`. That is a false positive: check-deprecated-versions.mjs matches `/\bv?[12]\.\d+\.\d+\b/`, and the example IP literal 1.2.3.4 contains "1.2.3". Swapped both occurrences in PROXY_GUIDE.md (and its pl mirror) for 203.0.113.7, from the RFC 5737 documentation range that exists precisely for examples — it cannot collide with a version pattern and is the correct thing to print in docs regardless. Drift count 64 -> 62; no gate threshold was touched. The gate that actually FAILED under "Docs sync + fabricated-docs (strict)" was check:fabricated-docs (the invented pool env vars), fixed in the previous commit; this one removes the misleading line the verdict quotes. * test(base): allowlist probeUtils and realign the #7849 suite to the replacement bound Two more base-reds, both visible only after the migration collision stopped killing the shards. **check-db-rules — src/lib/db/probeUtils.ts not classified** #9541 added probeUtils.ts (transient-error retry for the SQLite corruption probe). It is imported ONLY by src/lib/db/core.ts, exactly like its siblings schemaColumns / optimizationSettings / providerNodeSelect, so re-exporting it through localDb.ts would push callers toward the barrel-import anti-pattern the gate exists to prevent. Added to INTENTIONALLY_INTERNAL with that rationale. check-db-rules 22/22, check:db-rules exit 0. **session-dedup-memory-7849 — pinned a mechanism that was replaced**7f36b192f0(#7855 follow-up) swapped the shared "suffix work budget" for the MAX_SUFFIX_STARTS / MAX_TOTAL_BLOCK_BYTES guards and deleted both the budget and its SUFFIX_WORK_BUDGET_WARNING string. It updated session-dedup.test.ts but not this sibling, so 3 of its 4 cases asserted a warning that can no longer be emitted. Realigned to the contract that actually survives — which is the invariant #7849 was opened for, not the mechanism: - the pathological pair must stay BOUNDED (completes in <4s, body intact) — measured at ~280ms on the current guards; - it must FAIL OPEN — original body returned by identity, compressed false, stats null (the explanatory zero-savings stats belonged to the removed budget path, which skipped before producing any); - the 512 MiB child fixture must still exit 0 with the full engine chain (session-dedup, lite, rtk, headroom, caveman) — that IS the OOM guard — and session-dedup must still report its skip, now pinned by prefix since the reason string moved with the mechanism. No threshold was loosened and no case was deleted: 4/4 here, 8/8 on the sibling session-dedup.test.ts. Refs #9298 * docs(mcp): bump the tool count to 105 and realign two vitest count pins Three more base-reds from the same 08-06 batch, all count/contract drift that the merged PRs left in sibling files. **Docs Gates (fast-path) — 3 STRICT drifts** check:docs-counts measures the MCP tool set from live code: it is 105 now (#8925 added omniroute_create_combo), while README.md, AGENTS.md and docs/frameworks/MCP-SERVER.md still claimed 104. Updated all five occurrences (two of them inside SVG alt text). check:docs-all exits 0. **Vitest (fast-path) — 2 failures** - open-sse/mcp-server/__tests__/essentialTools.test.ts pinned 11 phase-1 tools; #8925 shipped omniroute_create_combo as phase 1, making it 12. Verified by enumerating MCP_ESSENTIAL_TOOLS directly. - tests/unit/autoCombo/provider-family-combos.test.ts pinned the auto/glm provider set to [auggie, glm, zai]. #8914 (Devin ACP bridge) added devin-cli-agentic, whose catalog (registry/devin/catalog.ts:90-93) advertises the glm-5-2* line — so it belongs in the family pool for exactly the reason the test's own comment gives for auggie: a no-auth backend that genuinely serves a family model is a legitimate member. Expected set updated, invariant unchanged. npm run test:vitest 36/36 files, 340/340 tests. Refs #9298 * fix(combo,usage,oauth): drain the base-reds the shard fix exposed With the migration collision and the broken import out of the way the four unit shards actually run, and a further layer of base-reds became visible on the pure tip9995bc4893. Three are production defects. **Production defects** - open-sse/services/combo/runtimeUnitCapacity.ts:58 called resolveComboTargets() WITHOUT the hidden-model snapshot, so it fell back to the default getHiddenModelsByProvider() — a fresh full key_value read PER nested combo-ref unit, on every request. #8878 threaded the snapshot through the other call sites and missed this one. Threaded it from executeRuntimeUnitCombo (and from the dispatchPrelude call site), restoring the one-snapshot-per-request invariant combo-hidden-leaf-routing.test.ts pins. 9/9. - open-sse/services/usage/firecrawl.ts silently ignored its own `apiKey` parameter:91bb6aa619moved the fetch to fetchFirecrawlQuota(connectionId, connection), which reads the key off the connection record, so any caller passing the key directly got "Firecrawl API key not available". The explicit key is now merged into the connection passed down. firecrawl-usage 8/8. - src/lib/oauth/constants/oauth.ts was missing a RAYCAST entry in PROVIDERS while src/lib/oauth/providers/index.ts registers `raycast` (#8895), so every consumer reading PROVIDERS did not know Raycast Pro exists. Also added its OAUTH_TEST_CONFIG entry (checkExpiry only — it is an `import_token` provider with refreshToken always null), which #8408's guard explicitly requires rather than grandfathering. oauth-providers-config 25/25, oauth-test-config-8408 2/2. **Count / contract drift from the same batch** - feature flags 45 -> 46, APIKEY_PROVIDERS 197 -> 198 (Raycast Pro #8895), unique MCP tools 107 -> 108. Each re-derived from the source of truth. - vi + pt-BR locales: translated the 8 keys #9415 added (providers.newApiAggregator* and providers.modelTestQuotaTooltip) instead of relaxing the parity guard. i18n-vi 5/5, i18n-pt-br 3/3. - login-bootstrap-route: #9491 added `authenticated` to the require-login payload so /login can redirect an active session; the three deepEqual bodies now carry it. 10/10. **Flaky-by-construction, made deterministic** tests/unit/chat-combo-live-test.test.ts asserted the early-keepalive frame with a 100ms mocked upstream while resolveKeepaliveThreshold() is 2000ms for openai/*. It only ever passed while unrelated handler latency happened to push the total past the threshold — incidental, not deterministic, and it stopped holding once the handler got faster. The mock now sleeps 2400ms so the slow path is guaranteed and the assertion means what it says. 5/5. typecheck:core exit 0. check:file-size (base-relative) OK. Refs #9298 * test(base): run the orphaned #8890 suite and realign three mechanism pins **check:test-discovery — a suite that had NEVER executed** #8890 landed open-sse/services/__tests__/fail-fast-concurrency-gate.test.ts into a directory no runner collects (only one explicit file from that folder is in vitest.mcp.config.ts), so it ran zero times since it merged. Wired it into the runner AND into check-test-discovery.mjs's mirrored collector list, which the gate keeps in sync deliberately. It passes 4/4 now that it actually runs — test:vitest goes 36 -> 37 files, 340 -> 344 tests. **check-db-rules-classification** — 37 -> 38 audited modules, adding probeUtils alongside the INTENTIONALLY_INTERNAL entry from the previous commit. **ratelimit-reservoir-refresh** — #9604 (rolling RPM leases) DELETED Bottleneck's fixed-window reservoir, so currentReservoir() is null and the poll for `reservoir === 2` could never settle. It updated several sibling suites but not this one. The pin on the removed mechanism is gone; what remains is the invariant the original Bottleneck heartbeat bug actually broke and that #9529 opened this test for — after a header-learned updateSettings() the limiter must keep admitting work, proven by racing a post-exhaustion request against a 5s timer. 1/1. **translator-openai-to-gemini** — #9568 (c9a3361e5a) made buildChangedToolNameMap emit IDENTITY entries too, because Gemini lowercases tool names in functionCall responses and the response translator needs a key to map them back. Any request carrying tools therefore carries `_toolNameMap` in the Antigravity envelope now. Expected key list updated and the map's contents asserted explicitly rather than left implicit. 45/45. Refs #9298 * fix(db): restore node-backed synced catalogs and realign the #8944 context hints **Production regression from #9294 (d69f521491)** lookupModelMeta moved from getSyncedAvailableModels(providerId) to getActiveSyncedCatalog(providerId). The new reader unions models only from rows in `provider_connections` with isActive = 1 — but a provider NODE lives in `provider_nodes` and NEVER has a connections row, so filtering by active connection ids silently dropped every node's synced catalog. The consequence was not just a missing list: lookupModelMeta reads that catalog for RUNTIME METADATA, so for openai-compatible nodes it took out - `supportedThinkingEfforts`, which is what splitSyncedEffortSuffix needs — so `<prefix>/<model>-high` stopped resolving to the base id and the effort was never derived (#7694), and - `contextWindow` / `maxInputTokens`, used by the combo context-window filter. getActiveSyncedCatalog now falls back to the provider-wide key_value set — the exact pre-#9294 source — when no active connection carries a catalog, and marks that fallback explicitly NON-authoritative. #9294's live-catalog gating is about what an active connection actually serves, so a node-backed catalog informs metadata while never being able to reject a model as unavailable. `available` therefore stays fail-open for nodes, as it was before. sync-reasoning-supported-efforts-7694 23/23 (was 21/2). live-model-catalog-reconciliation-8926 11/11 and combo-provider-wildcard 23/23 confirm #9294's own coverage is untouched. **#8944 sibling-test drift**714a315a1a("Treat context metadata as a routing hint") deliberately turned the context-window check from a HARD filter into an ordering hint: a catalog-too-small target is demoted, not removed, because a stale catalog entry must never delete the only target that could accept the request at runtime. The PR updated one case in this suite and left three asserting the old drop behaviour. Realigned to the new contract — the too-small target must lose the ordering to the fitting one while remaining present — and renamed them from "still rejects"/"still dropped" to "is demoted"/"ordered last" so the names stop describing the removed behaviour. 14/14. **file-size** tests/unit/translator-openai-to-gemini.test.ts testFrozen 1616 -> 1619: the frozen value sat exactly at the base size, so the 3 lines the previous commit's _toolNameMap alignment needs could not fit. Justified in the baseline. typecheck:core exit 0. Refs #9298 * chore(stryker): register the two covering suites missing from tap.testFiles check:mutation-test-coverage flags any unit test that covers a mutated module but is absent from stryker.conf.json tap.testFiles — without the entry its mutant kills do not count toward the module's score. - tests/unit/antigravity-prefer-stored-project.test.ts covers open-sse/services/combo/quotaStrategies.ts (added earlier in this PR). - tests/unit/executor-devin-cli-agentic-acp.test.ts covers src/sse/services/auth.ts — pre-existing drift, same gate, same fix. Inserted in alphabetical position only; the rest of the file is byte-identical (it is not prettier-formatted upstream and reformatting it is out of scope here). Refs #9298 * fix(db): drop the never-wired getSessionModelUsageCounts (knip regression) The dead-code ratchet only ran once the earlier Fast Quality Gates steps stopped failing, and it lands at 228 vs baseline 227. The extra symbol is src/lib/db/contextHandoffs.ts::getSessionModelUsageCounts, added by #8894 "for least-used strategy" and never wired: the least-used branch in applyStrategyOrdering.ts uses the pre-existing sortTargetsByUsage(), and the helper has no caller in src/, open-sse/ or tests/. It is the same incomplete-PR shape as that PR's import of a module which does not exist in the repo (fixed earlier in this branch). Removed rather than baselined — bumping the ratchet would loosen the gate, and removal is exactly the remedy the gate prescribes. Same treatment the Dario installer's never-wired uninstall() got in #9600. The implementation is recoverable froma598fbb090whenever someone actually wires a session-aware least-used strategy. check:dead-code 228 -> 227 (baseline untouched). check:db-rules exit 0. context-handoff 13/13, db-context-handoffs 7/7, service-context-handoff 11/11. Refs #9298 * fix(security): embed the Raycast signature secret via resolvePublicCred (HR#11) The secret-scan ratchet only ran once the earlier Fast Quality Gates steps stopped failing, and it lands at 1 finding vs baseline 0. The finding is open-sse/services/raycast.ts:19 — RAYCAST_DEFAULT_SIG_SECRET, a 64-hex request-signature secret that #8895 committed as a bare string literal. It is genuinely public (community-extracted from the Raycast macOS client; the SAME value ships to every install, it is not a per-user credential), which is exactly the category Hard Rule #11 governs: public upstream credentials MUST go through resolvePublicCred() (open-sse/utils/publicCreds.ts), never a literal — see docs/security/PUBLIC_CREDS.md. So the fix is the mandated pattern, not a .gitleaks.toml allowlist entry: added `raycast_sig_secret` to EMBEDDED_DEFAULTS as the XOR-masked byte sequence and resolved it with the existing RAYCAST_SIG_SECRET env override. The providerSpecificData.sigSecret override is untouched. Verified the decoded value is byte-identical to the literal it replaces. check:secrets secretFindings 1 -> 0. check:public-creds exit 0. publicCreds 12/12, raycast-auth 6/6, raycast-local-extract 1/1, trae-publiccred 3/3. typecheck:core exit 0. Refs #9298 --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
37 KiB
title, version, lastUpdated
| title | version | lastUpdated |
|---|---|---|
| OmniRoute MCP Server Documentation | 3.8.40 | 2026-06-28 |
OmniRoute MCP Server Documentation
Model Context Protocol server with 105 tools across routing, cache, compression, memory, skills, proxy, pool, and context source operations.
Source of truth:
open-sse/mcp-server/server.tscomputes 104 unique tools withcountUniqueMcpTools(): 42 canonical definitions (including the six CCR lifecycle tools and the agent-skills trio), plus memory (3), skills (4), GitHub skills (3), pool (6), gamification (8), plugins (8), Notion (6), Obsidian (22), and two RTK-only compression tools.
Installation
OmniRoute MCP is built-in. Start it with:
omniroute --mcp
Or via the open-sse transport:
# HTTP streamable transport (port 20130)
omniroute --dev # MCP auto-starts on /mcp endpoint
Transports
The MCP server exposes three transports, all backed by the same createMcpServer() factory:
| Transport | Where | When to use |
|---|---|---|
stdio |
open-sse/mcp-server/server.ts |
IDE integrations (Claude Desktop, Cursor, etc.) |
sse |
POST/GET /api/mcp/sse via httpTransport |
Browser/agent clients that need an event stream |
streamable-http |
POST/GET/DELETE /api/mcp/stream |
Multi-session HTTP clients (mcp-session-id header) |
The active HTTP transport (sse or streamable-http) is selected by the mcpTransport setting. Switching transports closes existing sessions on the other transport.
Remote access (manage-scope bypass)
/api/mcp/* is in the LOCAL_ONLY tier (src/server/authz/routeGuard.ts) — by default only loopback hosts (localhost, 127.0.0.1, ::1) can reach it. Since v3.8.2, non-loopback clients may connect if they present an Authorization: Bearer <api-key> whose key carries the manage scope. This is the only way to reach the remote MCP server through a tunnel, reverse proxy, or public hostname.
# Grant manage scope: open the dashboard API Keys page and toggle
# "Management Access" on the key, or POST scopes:["manage"] when creating.
# Then connect from a remote MCP client:
curl -i \
-H "Host: your-public-host.example" \
-H "Authorization: Bearer sk-…" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-03-26","capabilities":{},"clientInfo":{"name":"my-client","version":"0"}}}' \
https://your-public-host.example/api/mcp/stream
A non-manage key (or no Bearer) returns 403 LOCAL_ONLY. The sibling prefix /api/cli-tools/runtime/* is intentionally NOT bypassable — see Route Guard Tiers — Manage-scope carve-out.
IDE Configuration
See MCP Client Configuration for Claude Desktop, Cursor, Cline, and compatible MCP client setup.
Essential Tools (8) — Phase 1
| Tool | Scopes | Description |
|---|---|---|
omniroute_get_health |
read:health |
Uptime, memory, circuit breakers, rate limits, cache stats |
omniroute_list_combos |
read:combos |
All configured combos with strategies (optional metrics) |
omniroute_get_combo_metrics |
read:combos |
Performance metrics for a specific combo |
omniroute_switch_combo |
write:combos |
Activate or deactivate a combo |
omniroute_check_quota |
read:quota |
Quota used/total, percent remaining, reset time, token health |
omniroute_route_request |
execute:completions |
Send a chat completion through OmniRoute routing |
omniroute_cost_report |
read:usage |
Cost report by period (session/day/week/month) |
omniroute_list_models_catalog |
read:models |
Full model catalog with capabilities, status, pricing |
Phase 1 — Search
| Tool | Scopes | Description |
|---|---|---|
omniroute_web_search |
execute:search |
Web search through OmniRoute search gateway (Serper/Brave/Perplexity/Exa/Tavily/Google PSE/Linkup/SearchAPI/SearXNG) with failover |
Advanced Tools (11) — Phase 2
| Tool | Scopes | Description |
|---|---|---|
omniroute_simulate_route |
read:health, read:combos |
Dry-run routing simulation with fallback tree |
omniroute_set_budget_guard |
write:budget |
Session budget with degrade/block/alert action |
omniroute_set_routing_strategy |
write:combos |
Update combo strategy at runtime (priority/weighted/auto/etc.) |
omniroute_set_resilience_profile |
write:resilience |
Apply aggressive / balanced / conservative resilience preset |
omniroute_test_combo |
execute:completions, read:combos |
Live test of every provider in a combo using a real upstream call |
omniroute_get_provider_metrics |
read:health |
Per-provider metrics with p50/p95/p99 latency and circuit breaker state |
omniroute_best_combo_for_task |
read:combos, read:health |
Recommend combo by task type with budget/latency constraints |
omniroute_explain_route |
read:health, read:usage |
Explain why a request was routed to a provider (scoring factors + fallbacks) |
omniroute_get_session_snapshot |
read:usage |
Full session snapshot: cost, tokens, top models/providers, errors, budget guard |
omniroute_db_health_check |
read:health, write:resilience |
Diagnose (and optionally auto-repair) database drift like broken combo refs / orphan rows |
omniroute_sync_pricing |
pricing:write |
Sync pricing data from external sources (LiteLLM); supports dryRun |
Cache Tools (2)
| Tool | Scopes | Description |
|---|---|---|
omniroute_cache_stats |
read:cache |
Semantic cache, prompt-cache, and idempotency stats |
omniroute_cache_flush |
write:cache |
Flush cache globally or by signature/model |
Compression Tools (13)
| Tool | Scopes | Description |
|---|---|---|
omniroute_compression_status |
read:compression |
Compression settings, analytics summary, and cache-aware stats (includes analytics.mcpDescriptionCompression metadata) |
omniroute_compression_configure |
write:compression |
Configure compression mode, threshold, target ratio, system-prompt preservation, MCP description compression toggle |
omniroute_set_compression_engine |
write:compression |
Pick the active engine (off/caveman/rtk/stacked) and Caveman/RTK intensity |
omniroute_list_compression_combos |
read:compression |
List named compression combos and their engine pipelines |
omniroute_compression_combo_stats |
read:compression |
Analytics grouped by compression combo and engine |
omniroute_ccr_store |
write:compression |
Store caller-isolated content in the bounded in-memory CCR store and return a marker plus ccr:// reference |
omniroute_ccr_retrieve |
read:compression |
Retrieve CCR content in full or with head, tail, lines, grep, and stats modes |
omniroute_ccr_inspect |
read:compression |
Inspect caller-owned CCR metadata without returning content |
omniroute_ccr_list |
read:compression |
List paginated metadata for caller-owned CCR blocks |
omniroute_ccr_delete |
write:compression |
Delete a caller-owned CCR block |
omniroute_ccr_stats |
read:compression |
Report caller-scoped memory usage, lifecycle counters, and store limits |
omniroute_rtk_discover |
read:compression |
Discover recurring noise in opt-in RTK output samples |
omniroute_rtk_learn |
read:compression |
Generate a reviewable RTK filter draft from opt-in samples |
CCR entries are in-memory only and disappear on restart. Each block is limited to 2 MiB, each principal to 16 MiB, and the global store to 64 MiB. Entries default to a 24-hour TTL (maximum seven days). Full MCP retrieval is limited to 256 KiB; larger blocks remain available through the ranged and grep modes. Storage, retrieval, listing, inspection, deletion, and stats are isolated by the authenticated API-key principal. Audit records contain hashes and size metadata, never content.
omniroute_compression_status reports MCP description compression separately under
analytics.mcpDescriptionCompression. Those values are metadata-size estimates for MCP listable
descriptions (tools, prompts, resources, and resourceTemplates); they are not provider usage
receipts and are marked with source: "mcp_metadata_estimate".
MCP Accessibility Tree Filter (v3.8.0)
Separate from the compression tools above, OmniRoute includes a post-execution filter that compresses the tool results of MCP browser/accessibility tools before they are returned to the agent. This filter is not itself a tool — it runs transparently on any tool result that contains verbose accessibility-tree or browser-snapshot text (≥2000 chars).
Key behaviors:
- Collapses ≥30 consecutive repeated sibling lines into head + tail summary
- Preserves
[ref=eXX]anchors required by Playwright/computer-use - Hard-truncates oversized text (>50,000 chars) with a navigation hint
- Expected savings: 60–80% on browser snapshot payloads
Configuration: compression.mcpAccessibility in global settings (migration 056).
Implementation: open-sse/services/compression/engines/mcpAccessibility/.
Full docs: Compression Engines — MCP Accessibility Tree Filter.
See Compression Engines and RTK Compression for the runtime compression model behind these tools.
1Proxy Tools (3)
| Tool | Scopes | Description |
|---|---|---|
omniroute_oneproxy_fetch |
read:proxies |
Fetch free proxies from the 1proxy marketplace (protocol/country/quality/limit filters) |
omniroute_oneproxy_rotate |
read:proxies |
Get the next available proxy by strategy (random / quality / sequential) |
omniroute_oneproxy_stats |
read:proxies |
Pool stats, sync status, distribution by protocol and country |
Memory Tools (3)
Defined in open-sse/mcp-server/tools/memoryTools.ts. Auth/scope is enforced through the standard MCP scope pipeline.
| Tool | Scopes | Description |
|---|---|---|
omniroute_memory_search |
read:memory |
Search memories by query / type / API key with token-budget enforcement |
omniroute_memory_add |
write:memory |
Add a new memory entry (factual / episodic / procedural / semantic) |
omniroute_memory_clear |
write:memory |
Clear memories for an API key, optionally filtered by type or olderThan timestamp |
Skill Tools (4)
Defined in open-sse/mcp-server/tools/skillTools.ts. Backed by src/lib/skills/registry + src/lib/skills/executor.
| Tool | Scopes | Description |
|---|---|---|
omniroute_skills_list |
read:skills |
List registered skills with optional filtering by API key, name, or enabled state |
omniroute_skills_enable |
write:skills |
Enable or disable a specific skill by ID |
omniroute_skills_execute |
execute:skills |
Execute a skill with provided input and return the execution record |
omniroute_skills_executions |
read:skills |
List recent skill execution history |
Notion Context Source (6)
Defined in open-sse/mcp-server/tools/notionTools.ts. Token stored in key_value table via src/lib/db/notion.ts. REST client in src/lib/notion/api.ts. Settings API in src/app/api/settings/notion/route.ts. Dashboard UI in src/app/(dashboard)/dashboard/endpoint/components/NotionSourceCard.tsx.
Configure your Notion integration token from the Context Sources tab in the Endpoint dashboard, or via the REST API:
# Set token
curl -X POST http://localhost:20128/api/settings/notion \
-H "Content-Type: application/json" \
-d '{"token": "ntn_..."}'
# Check status
curl http://localhost:20128/api/settings/notion
# Disconnect
curl -X DELETE http://localhost:20128/api/settings/notion
| Tool | Scopes | Description |
|---|---|---|
notion_search |
read:notion |
Full-text search across all pages and databases |
notion_get_page |
read:notion |
Get a page by ID with its properties |
notion_list_block_children |
read:notion |
List the child blocks of a page or block |
notion_query_database |
read:notion |
Query a database with filters, sorts, and pagination |
notion_get_database |
read:notion |
Get database schema by ID |
notion_append_blocks |
write:notion |
Append children blocks to a parent block (max 100 per request) |
Agent Skill Catalog Tools (3)
Defined in open-sse/mcp-server/tools/agentSkillTools.ts. Backed by src/lib/agentSkills/catalog. These tools expose the 42-entry Agent Skills documentation catalog to MCP clients and external agents. Scope: read:catalog.
| Tool | Scopes | Description |
|---|---|---|
omniroute_agent_skills_list |
read:catalog |
List all 42 agent skills with optional category (api|cli) and area filters; returns metadata + coverage |
omniroute_agent_skills_get |
read:catalog |
Get full metadata + SKILL.md content for a single skill by canonical id |
omniroute_agent_skills_coverage |
read:catalog |
Coverage stats: how many of the 22 API and 20 CLI skills have SKILL.md files on the filesystem vs catalog totals |
See AGENT-SKILLS.md for the full catalog and how external agents consume it.
Related Frameworks (v3.8.0)
The MCP tool inventory above (104 unique tools, computed by countUniqueMcpTools()) is intentionally
scoped to runtime routing/cache/compression/memory/skills/proxy/context-source operations. Two adjacent
frameworks ship alongside the MCP server in v3.8.0 and are documented separately:
Cloud Agents
Cloud Agents are out-of-process AI coding agents (codex-cloud, devin, jules) wired into
OmniRoute through the same connection model used for LLM providers. They are exposed via
their own REST surface (/api/v1/agents/*) and are not part of the MCP tool catalog
— calling a Cloud Agent does not consume an MCP scope.
- Implementation:
src/lib/cloudAgent/(registry.ts,agents/codex-cloud.ts,agents/devin.ts,agents/jules.ts). - Lifecycle:
createTask,getStatus,approvePlan,sendMessage,listSources. - Documentation: docs/frameworks/CLOUD_AGENT.md.
Guardrails
Guardrails are pre/post-execution filters (vision-bridge, pii-masker, prompt-injection) applied inside the chat pipeline. They run before the MCP tool/route layer is reached and emit structured violations to the audit pipeline; they are not invoked as MCP tools.
- Implementation:
src/lib/guardrails/. - Documentation: docs/security/GUARDRAILS.md.
When debugging an MCP call that appears blocked, check both the MCP audit log
(scope_denied:* entries) and the guardrails audit trail — a request may be rejected by
a guardrail before it ever reaches the MCP scope enforcement layer.
REST API Endpoints
| Endpoint | Method | Description | Auth |
|---|---|---|---|
/api/mcp/status |
GET |
Server status: heartbeat, HTTP transport state, audit activity summary | Management (session/admin) |
/api/mcp/tools |
GET |
Tool catalog (name, description, scopes, phase, source endpoints) | Management |
/api/mcp/sse |
GET / POST |
SSE transport endpoint (gated by mcpEnabled + mcpTransport === "sse") |
API key + scopes |
/api/mcp/stream |
POST/GET/DELETE |
Streamable HTTP transport (uses mcp-session-id header; DELETE ends the session) |
API key + scopes |
/api/mcp/audit |
GET |
Audit log entries from mcp_tool_audit (filters: limit, offset, tool, success, apiKeyId) |
Management |
/api/mcp/audit/stats |
GET |
Aggregated audit stats (totalCalls, successRate, avgDurationMs, top tools) |
Management |
Source files: src/app/api/mcp/{status,tools,sse,stream,audit,audit/stats}/route.ts.
Both SSE and Streamable HTTP transports are blocked until the MCP server is enabled in Settings (mcpEnabled) and the appropriate mcpTransport is selected. If the wrong transport is configured the route returns HTTP 400 with a hint to switch settings.
Authentication & Scopes
MCP tools are authenticated through API key scopes. Scope enforcement is centralized in
open-sse/mcp-server/scopeEnforcement.ts. Each tool requires specific scopes:
| Scope | Tools |
|---|---|
read:health |
get_health, get_provider_metrics, simulate_route, explain_route, best_combo_for_task, db_health_check |
read:combos |
list_combos, get_combo_metrics, simulate_route, best_combo_for_task, test_combo |
write:combos |
switch_combo, set_routing_strategy |
read:quota |
check_quota |
read:usage |
cost_report, get_session_snapshot, explain_route |
read:models |
list_models_catalog |
execute:completions |
route_request, test_combo |
execute:search |
web_search |
write:budget |
set_budget_guard |
write:resilience |
set_resilience_profile, db_health_check |
pricing:write |
sync_pricing |
read:cache |
cache_stats |
write:cache |
cache_flush |
read:compression |
compression_status, list_compression_combos, compression_combo_stats |
write:compression |
compression_configure, set_compression_engine |
read:proxies |
oneproxy_fetch, oneproxy_rotate, oneproxy_stats |
read:notion |
notion_search, notion_list_databases, notion_get_database, notion_query_database, notion_read |
write:notion |
notion_append_blocks |
read:memory |
memory_search |
write:memory |
memory_add, memory_clear |
read:skills |
skills_list, skills_executions |
write:skills |
skills_enable |
execute:skills |
skills_execute |
read:catalog |
agent_skills_list, agent_skills_get, agent_skills_coverage |
Wildcard scopes are supported: read:* grants all read-scopes, * grants full access.
mcp:connect — narrow route capability (#7895)
Reaching the HTTP/SSE MCP transport (/api/mcp/*) from non-loopback requires the
/api/mcp/ LOCAL_ONLY carve-out (see docs/security/ROUTE_GUARD_TIERS.md). Historically
that carve-out only accepted a full manage/admin-scope API key — too broad for a
caller that only needs to talk MCP. src/shared/constants/managementScopes.ts now
exports MCP_CONNECT_SCOPE = "mcp:connect": an additive, narrow scope (same precedent as
SELF_USAGE_SCOPE) that authorizes ONLY the /api/mcp/ bypass in
src/server/authz/policies/management.ts — it grants no other management-route access
and is deliberately kept OUT of MANAGEMENT_API_KEY_SCOPES. A key holding manage/admin
still passes the carve-out unchanged; mcp:connect is a lower-privilege alternative for
remote MCP-only callers, checked via hasMcpConnectOrManageScope().
Per-key HTTP scope binding (#7895)
Over HTTP/SSE, open-sse/mcp-server/httpTransport.ts now resolves the caller's real
api_keys.scopes via resolveMcpCallerAuthInfo() (open-sse/mcp-server/httpAuthContext.ts)
and passes it to the MCP SDK's transport.handleRequest(req, { authInfo }), so
extra.authInfo.scopes reaching each tool call reflects the Bearer key's own scopes.
scopeEnforcement.ts's resolveCallerScopeContext() already prioritized authInfo over
the _meta and OMNIROUTE_MCP_SCOPES env fallback — this only populates that first,
highest-priority source, which was previously unfed over HTTP. When no API key resolves
(no header, invalid key), authInfo stays undefined and resolution falls through to the
existing meta/env chain unchanged. This does NOT flip OMNIROUTE_MCP_ENFORCE_SCOPES's
default — enforcement still has to be explicitly enabled; this change only makes the
per-key path take precedence once it is. stdio has no per-caller identity (see
mcpCallerIdentity.ts) and is unaffected — it stays on the _meta/env fallback chain.
Environment Variables
| Variable | Default | Purpose |
|---|---|---|
OMNIROUTE_BASE_URL |
http://localhost:20128 |
Base URL the MCP server uses when calling OmniRoute internal APIs |
OMNIROUTE_API_KEY |
(empty) | API key forwarded as Authorization: Bearer to internal API calls |
OMNIROUTE_MCP_ENFORCE_SCOPES |
false (only "true" enables it) |
When enabled, missing scopes deny tool calls and log scope_denied:<reason> in audit log |
OMNIROUTE_MCP_SCOPES |
(empty) | Comma-separated allowlist of scopes considered "available" by default (used when caller does not provide its own scopes) |
OMNIROUTE_MCP_COMPRESS_DESCRIPTIONS |
(unset = on) | When set to 0/false/off/no, disables MCP description compression at registration time |
OMNIROUTE_MCP_DESCRIPTION_COMPRESSION |
(unset = on) | Alternate alias for the same toggle as above |
MCP_TOOL_DENY |
(unset = no filter) | Comma-separated tool names to drop from tools/list (tool-cardinality reduction — see below) |
MCP_TOOL_ALLOW |
(unset = no filter) | Comma-separated tool names to keep exclusively (allow-list mode — see below) |
DATA_DIR |
~/.omniroute |
Heartbeat file is written to ${DATA_DIR}/runtime/mcp-heartbeat.json |
Description Compression
MCP tool, prompt, and resource registries can compress descriptions at registration/list time to reduce the metadata footprint exposed to clients (and therefore the prompt context cost). The implementation lives in open-sse/mcp-server/descriptionCompressor.ts and is wired into the MCP server via compressMcpRegistryMetadata inside createMcpServer().
- Compression runs over the description text using the Caveman ruleset (
getRulesForContext("all", "full")) with preserved-block extraction (code spans, fenced blocks, etc.) so structural content is not altered. - Toggle per-deployment via the
compression.mcpDescriptionCompressionEnabledvalue in thekey_valuesettings table (default: enabled) — exposed in the UI as Analytics → MCP description compression. - Toggle process-wide via either
OMNIROUTE_MCP_COMPRESS_DESCRIPTIONS=falseorOMNIROUTE_MCP_DESCRIPTION_COMPRESSION=false. - Realtime stats are surfaced via
omniroute_compression_statusunderanalytics.mcpDescriptionCompressionand taggedsource: "mcp_metadata_estimate"to disambiguate from real provider usage receipts.
Tool Cardinality Reduction (F4.3)
Description compression shrinks each tool's metadata; tool-cardinality reduction goes one step further by reducing how many tools are announced at all. Advertising fewer tools in the tools/list manifest cuts the per-request token cost the client's model pays for the tool catalog ("layer 5" compression). The implementation is a pure, stateless filter in open-sse/mcp-server/toolCardinality.ts (reduceToolManifest), wired into the registration loop in createMcpServer() (open-sse/mcp-server/server.ts).
Opt-in, off by default. The filter only runs when at least one of two environment variables is set; with neither set, all 105 tools are announced unchanged.
| Variable | Mode |
|---|---|
MCP_TOOL_DENY |
Blacklist — comma-separated tool names that are always dropped from tools/list |
MCP_TOOL_ALLOW |
Allow-list — comma-separated tool names; only these survive, everything else is dropped |
deny takes priority over allow. Names are comma-separated, trimmed, and empty entries are ignored. Examples:
# Drop two tools from the catalog
MCP_TOOL_DENY="omniroute_get_health,omniroute_list_combos" omniroute --mcp
# Announce only the routing + quota tools (allow-list mode)
MCP_TOOL_ALLOW="omniroute_route_request,omniroute_check_quota" omniroute --mcp
How filtered tools are removed: registration always succeeds; a tool the profile rejects is then .disable()d on the MCP SDK handle, so it never appears in tools/list but the wiring stays intact (clean enable/disable, no re-registration). The profile parser is readMcpToolProfileFromEnv(process.env), which returns null (no filtering) when both vars are empty.
The richer ToolProfile shape behind reduceToolManifest also supports scope-intersection filtering (allowScopes, with read:*-style wildcard matching) and a deterministic maxTools cap, but those two knobs need the full manifest at registration time and are not exposed through the environment variables today (a tools/list-level hook is a tracked follow-up). estimateManifestTokens() is available to compare the manifest token cost before and after reduction.
Runtime Heartbeat
The stdio transport persists liveness to ${DATA_DIR}/runtime/mcp-heartbeat.json every 5 seconds. The dashboard (/api/mcp/status) reads this file plus PID liveness to derive online. HTTP transports report state from in-process getMcpHttpStatus() instead (no file write).
The heartbeat snapshot contains:
{
"pid": 12345,
"startedAt": "2026-05-13T12:34:56.000Z",
"lastHeartbeatAt": "2026-05-13T12:35:01.000Z",
"version": "1.8.1",
"transport": "stdio",
"scopesEnforced": false,
"allowedScopes": [],
"toolCount": 43
}
Audit Logging
Every tool call is logged to the SQLite mcp_tool_audit table by open-sse/mcp-server/audit.ts:
- Tool name, arguments (hashed/truncated as per per-tool
auditLevel), result - Duration in ms, success/failure flag, error message (when applicable)
- API key hash, timestamp
- Scope denials are logged as
scope_denied:<reason>with the missing scope list
Use the dashboard or the /api/mcp/audit and /api/mcp/audit/stats REST endpoints to inspect recent calls.
Files
| File | Purpose |
|---|---|
open-sse/mcp-server/server.ts |
MCP server factory, stdio entry point, scoped tool registrations |
open-sse/mcp-server/httpTransport.ts |
SSE + Streamable HTTP transport (session management) |
open-sse/mcp-server/scopeEnforcement.ts |
Tool scope evaluation and caller resolution |
open-sse/mcp-server/audit.ts |
Tool call audit logging (mcp_tool_audit) |
open-sse/mcp-server/runtimeHeartbeat.ts |
stdio heartbeat writer (mcp-heartbeat.json) |
open-sse/mcp-server/descriptionCompressor.ts |
Description compression for tool / prompt / resource registries |
open-sse/mcp-server/schemas/tools.ts |
Zod schemas + tool registry (MCP_TOOLS, 34 entries) |
open-sse/mcp-server/tools/advancedTools.ts |
Phase 2 + cache + 1proxy tool handlers |
open-sse/mcp-server/tools/compressionTools.ts |
Compression tool handlers |
open-sse/mcp-server/tools/memoryTools.ts |
Memory tool definitions (3 tools) |
open-sse/mcp-server/tools/skillTools.ts |
Skill tool definitions (4 tools) |
open-sse/mcp-server/tools/notionTools.ts |
Notion context source tool definitions (6 tools) |
open-sse/mcp-server/tools/gamificationTools.ts |
Gamification tool definitions (8 tools) |
open-sse/mcp-server/tools/pluginTools.ts |
Plugin registration and management tools (8 tools) |
src/app/api/mcp/status/route.ts |
/api/mcp/status endpoint |
src/app/api/mcp/tools/route.ts |
/api/mcp/tools endpoint |
src/app/api/mcp/sse/route.ts |
/api/mcp/sse SSE transport route |
src/app/api/mcp/stream/route.ts |
/api/mcp/stream Streamable HTTP transport route |
src/app/api/mcp/audit/route.ts |
/api/mcp/audit audit log query |
src/app/api/mcp/audit/stats/route.ts |
/api/mcp/audit/stats aggregated audit metrics |
src/lib/notion/api.ts |
Notion REST API client (retry, timeout, error classification) |
src/lib/db/notion.ts |
Notion token persistence (key_value table) |
src/app/api/settings/notion/route.ts |
Notion settings API (GET/POST/DELETE) |
src/app/(dashboard)/dashboard/endpoint/components/NotionSourceCard.tsx |
Notion token management UI |
tests/unit/notion-api.test.ts |
Notion API client tests (7) |
tests/unit/notion-tools.test.ts |
Notion tools scope enforcement tests (10) |
tests/unit/db/notion.test.mjs |
Notion DB module tests (3) |