From 03410476d76673a0d3c4483111f061ce1f1ca752 Mon Sep 17 00:00:00 2001 From: Diego Rodrigues de Sa e Souza Date: Tue, 30 Jun 2026 18:38:20 -0300 Subject: [PATCH] docs(readme): refresh metrics, list 17 strategies, add Quota-Share + real provider logos - Unify provider count to 236; MCP tools 87->94; cloud agents 3->4 (+Cursor); compression 9->10 engines (+relevance) - Tests -> 21,000+ across 2,586 files; footer -> v3.8.43 - Raise lower bounds to real values: 90+ free, 80+ commands, 24+ CLIs - Language flag grid 33->43 (15/14/14, all locales) - List all 17 routing strategies; new Quota-Share section before Resilience - Real provider logos (lobe-icons + local agentrouter) in providers grid and Free Forever - Top Contributors: refreshed stats + add herjarsa; 280+ title; half-size avatars; contrib.rocks 100->200 - Acknowledgments: refreshed star counts; fix headroom repo rename --- README.md | 355 +++++++++++++++++++++++++++++++++--------------------- 1 file changed, 220 insertions(+), 135 deletions(-) diff --git a/README.md b/README.md index d17a351455..8e41eacd95 100644 --- a/README.md +++ b/README.md @@ -6,7 +6,7 @@ # ๐Ÿš€ OmniRoute โ€” The Free AI Gateway -### Never stop coding. Connect every AI tool to **236 providers** โ€” **50+ free** โ€” through one endpoint. +### Never stop coding. Connect every AI tool to **236 providers** โ€” **90+ free** โ€” through one endpoint. **Plug Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini. Auto-fallback.**
@@ -19,8 +19,15 @@
-[![231 AI Providers](https://img.shields.io/badge/231-AI_Providers-6C5CE7?style=for-the-badge)](#-231-ai-providers--50-free) -[![50+ Free](https://img.shields.io/badge/50%2B-Free_Tiers-00B894?style=for-the-badge)](#-231-ai-providers--50-free) +

+ +โญ Star the repo if OMNIROUTE helped you save money and make your work easier. [![Stars](https://img.shields.io/github/stars/diegosouzapw/OmniRoute?style=social)](https://github.com/diegosouzapw/OmniRoute) +

+ +diegosouzapw%2FOmniRoute | Trendshift + +[![236 AI Providers](https://img.shields.io/badge/236-AI_Providers-6C5CE7?style=for-the-badge)](#-236-ai-providers--90-free) +[![90+ Free](https://img.shields.io/badge/90%2B-Free_Tiers-00B894?style=for-the-badge)](#-236-ai-providers--90-free) [![1.6B Free Tokens/mo](https://img.shields.io/badge/1.6B-Free_Tokens%2Fmo-00B894?style=for-the-badge)](docs/reference/FREE_TIERS.md) [![Token Savings](https://img.shields.io/badge/up_to_95%25-Token_Savings-E17055?style=for-the-badge)](#%EF%B8%8F-save-1595-tokens--automatically) [![17 Strategies](https://img.shields.io/badge/17-Routing_Strategies-0984E3?style=for-the-badge)](#-combos--the-flagship) @@ -34,79 +41,78 @@ [![Telegram](https://img.shields.io/badge/Telegram-26A5E4?style=for-the-badge&logo=telegram&logoColor=white)](https://t.me/omnirouteOficial) [![WhatsApp Global](https://img.shields.io/badge/WhatsApp_Global-25D366?style=for-the-badge&logo=whatsapp&logoColor=white)](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) [![WhatsApp Brasil](https://img.shields.io/badge/WhatsApp_Brasil-25D366?style=for-the-badge&logo=whatsapp&logoColor=white)](https://chat.whatsapp.com/BTGJXIyjeNIIgExvTMGGhI) +[![Website](https://img.shields.io/badge/Website-omniroute.online-blue?logo=google-chrome&logoColor=white)](https://omniroute.online) **Questions, provider tips, roadmap & support โ†’ [Discord](https://discord.gg/EkzRkpzKYt) ยท [Telegram](https://t.me/omnirouteOficial) ยท WhatsApp [๐ŸŒ Global](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) / [๐Ÿ‡ง๐Ÿ‡ท Brasil](https://chat.whatsapp.com/BTGJXIyjeNIIgExvTMGGhI)**
-diegosouzapw%2FOmniRoute | Trendshift - -[![npm](https://img.shields.io/npm/v/omniroute?logo=npm&style=flat-square)](https://www.npmjs.com/package/omniroute) -[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE) -[![Node](https://img.shields.io/badge/node-%E2%89%A522.0.0-brightgreen?style=flat-square)](package.json) -[![Stars](https://img.shields.io/github/stars/diegosouzapw/OmniRoute?style=social)](https://github.com/diegosouzapw/OmniRoute) - -
+### ๐Ÿงฉ Available [![npm version](https://img.shields.io/npm/v/omniroute?color=cb3837&logo=npm)](https://www.npmjs.com/package/omniroute) ![NPM Monthly](https://img.shields.io/npm/dm/omniroute?label=npm/month&color=cb3837&logo=npm) [![Docker Hub](https://img.shields.io/docker/v/diegosouzapw/omniroute?label=Docker%20Hub&logo=docker&color=2496ED)](https://hub.docker.com/r/diegosouzapw/omniroute) +[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE) ![Docker Pulls](https://img.shields.io/docker/pulls/diegosouzapw/omniroute?label=docker%20pulls&logo=docker&color=2496ED) ![Electron Downloads](https://img.shields.io/github/downloads/diegosouzapw/omniroute/total?style=flat&label=electron%20downloads&logo=electron&color=47848F) -[![Website](https://img.shields.io/badge/Website-omniroute.online-blue?logo=google-chrome&logoColor=white)](https://omniroute.online) -
- -
- -[**๐Ÿš€ Quick Start**](#-quick-start) โ€ข [**๐ŸŽฏ Combos**](#-combos--the-flagship) โ€ข [**๐ŸŒ Providers**](#-231-ai-providers--50-free) โ€ข [**๐Ÿ”Œ CLI & MCP**](#-full-cli--a2a--mcp) โ€ข [**๐Ÿ—œ๏ธ Compression**](#%EF%B8%8F-save-1595-tokens--automatically) โ€ข [**๐ŸŒ Website**](https://omniroute.online) +[**๐Ÿš€ Quick Start**](#-quick-start) โ€ข [**๐ŸŽฏ Combos**](#-combos--the-flagship) โ€ข [**๐ŸŒ Providers**](#-236-ai-providers--90-free) โ€ข [**๐Ÿ”Œ CLI & MCP**](#-full-cli--a2a--mcp) โ€ข [**๐Ÿ—œ๏ธ Compression**](#%EF%B8%8F-save-1595-tokens--automatically) โ€ข [**๐ŸŒ Website**](https://omniroute.online) [๐Ÿ’ฅ The Promise](#-the-promise) โ€ข [๐Ÿค” Why](#-why-omniroute) โ€ข [๐Ÿ† What Sets Apart](#-what-sets-omniroute-apart) โ€ข [๐Ÿค– Compatible CLIs](#-compatible-clis--coding-agents) โ€ข [๐Ÿ–ฅ๏ธ Where It Runs](#%EF%B8%8F-where-omniroute-runs--anywhere) โ€ข [๐Ÿ”’ Private](#-private--local-first) โ€ข [๐ŸŽฌ In Action](#-omniroute-in-action) โ€ข [๐Ÿ“š Explore More](#-explore-more) โ€ข [๐Ÿ“ง Support](#-support--community)
- ๐ŸŒ Available in 41+ languages + ๐ŸŒ In 42+ languages + - - - - - - - - - - - - - - - - + + + + + - - - - - - + + + + + + + + + + + + + + + + + + + + + + + + +
๐Ÿ‡บ๐Ÿ‡ธ ๐Ÿ‡ง๐Ÿ‡ท๐Ÿ‡ต๐Ÿ‡น ๐Ÿ‡ช๐Ÿ‡ธ ๐Ÿ‡ซ๐Ÿ‡ท ๐Ÿ‡ฎ๐Ÿ‡น๐Ÿ‡ท๐Ÿ‡บ๐Ÿ‡จ๐Ÿ‡ณ๐Ÿ‡น๐Ÿ‡ผ ๐Ÿ‡ฉ๐Ÿ‡ช๐Ÿ‡ฏ๐Ÿ‡ต๐Ÿ‡ฐ๐Ÿ‡ท๐Ÿ‡ฎ๐Ÿ‡ณ
๐Ÿ‡น๐Ÿ‡ญ๐Ÿ‡ป๐Ÿ‡ณ๐Ÿ‡ฎ๐Ÿ‡ฉ๐Ÿ‡ฒ๐Ÿ‡พ๐Ÿ‡ต๐Ÿ‡ญ๐Ÿ‡ธ๐Ÿ‡ฆ๐Ÿ‡ฎ๐Ÿ‡ฑ๐Ÿ‡ฆ๐Ÿ‡ฟ๐Ÿ‡ณ๐Ÿ‡ฑ๐Ÿ‡ท๐Ÿ‡บ ๐Ÿ‡บ๐Ÿ‡ฆ ๐Ÿ‡ต๐Ÿ‡ฑ ๐Ÿ‡จ๐Ÿ‡ฟ๐Ÿ‡ธ๐Ÿ‡ฐ๐Ÿ‡ท๐Ÿ‡ด๐Ÿ‡ญ๐Ÿ‡บ
๐Ÿ‡ณ๐Ÿ‡ฑ ๐Ÿ‡ง๐Ÿ‡ฌ ๐Ÿ‡ฉ๐Ÿ‡ฐ ๐Ÿ‡ซ๐Ÿ‡ฎ ๐Ÿ‡ณ๐Ÿ‡ด ๐Ÿ‡ธ๐Ÿ‡ช๐Ÿ‡ญ๐Ÿ‡บ๐Ÿ‡ท๐Ÿ‡ด๐Ÿ‡ธ๐Ÿ‡ฐ๐Ÿ‡ต๐Ÿ‡น๐Ÿ‡จ๐Ÿ‡ณ๐Ÿ‡น๐Ÿ‡ผ๐Ÿ‡ฏ๐Ÿ‡ต๐Ÿ‡ฐ๐Ÿ‡ท๐Ÿ‡น๐Ÿ‡ญ๐Ÿ‡ป๐Ÿ‡ณ๐Ÿ‡ฎ๐Ÿ‡ฉ๐Ÿ‡ฒ๐Ÿ‡พ๐Ÿ‡ต๐Ÿ‡ญ
๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ง๐Ÿ‡ฉ๐Ÿ‡ต๐Ÿ‡ฐ๐Ÿ‡ฎ๐Ÿ‡ท๐Ÿ‡ธ๐Ÿ‡ฆ๐Ÿ‡ฎ๐Ÿ‡ฑ๐Ÿ‡น๐Ÿ‡ท๐Ÿ‡ฆ๐Ÿ‡ฟ๐Ÿ‡น๐Ÿ‡ฟ
@@ -144,12 +150,12 @@ ๐Ÿšซ Never hit limits
Auto-fallback across 236 providers in milliseconds. Quota out? Next provider takes over โ€” zero downtime. ๐Ÿ’ธ Save up to 95% tokens
RTK + Caveman stacked compression cuts 15โ€“95% of eligible tokens (~89% avg on tool-heavy sessions). - ๐Ÿ†“ $0 to start
50+ providers with a free tier, 11 free forever (Kiro, Qoder, Pollinations, LongCatโ€ฆ). No card needed. + ๐Ÿ†“ $0 to start
90+ providers with a free tier, 11 free forever (Kiro, Qoder, Pollinations, LongCatโ€ฆ). No card needed. - ๐Ÿ”Œ Every tool works
16+ coding agents โ€” Claude Code, Codex, Cursor, Cline, Copilot, Antigravity โ€” through one config. + ๐Ÿ”Œ Every tool works
24+ coding agents โ€” Claude Code, Codex, Cursor, Cline, Copilot, Antigravity โ€” through one config. ๐Ÿงฉ One endpoint
OpenAI โ†” Claude โ†” Gemini โ†” Responses API translation. Point any tool at /v1 and it just works. - ๐Ÿ›ก๏ธ Production-grade
Circuit breakers, TLS stealth, MCP (87 tools), A2A, memory, guardrails, evals. 14,965 tests. + ๐Ÿ›ก๏ธ Production-grade
Circuit breakers, TLS stealth, MCP (94 tools), A2A, memory, guardrails, evals. 21,000+ tests. @@ -223,21 +229,56 @@ No combo to create. Set your model to `auto` (or a variant) and OmniRoute builds ### ๐Ÿ”€ Or build your own โ€” 17 routing strategies -| Goal | Strategy / combo | -| --------------------------------------- | -------------------------------------------------- | -| ๐Ÿฅ‡ Drain my subscription before paying | `priority` / `fill-first` | -| โš–๏ธ Spread load across accounts | `round-robin` ยท `weighted` ยท `p2c` ยท `least-used` | -| ๐Ÿ’ธ Always cheapest viable model | `cost-optimized` ยท `auto/cheap` | -| ๐Ÿง  Hand off long context between models | `context-relay` ยท `context-optimized` | -| ๐ŸŽฒ Randomized / privacy routing | `random` ยท `strict-random` | -| ๐Ÿงฌ Fan out to a panel + judge synthesis | `fusion` | -| ๐Ÿ“Š Route by remaining quota headroom | `reset-window` ยท `headroom` | -| ๐Ÿค– Just make it smart | `auto` (9-factor scoring) ยท `lkgp` ยท `reset-aware` | +All **17** strategies โ€” mix & match per combo step: + +| # | Strategy | What it does | +| --- | ------------------- | ---------------------------------------------------------------- | +| 1 | `priority` | First-target ordered list โ€” drain each before the next ๐Ÿฅ‡ | +| 2 | `fill-first` | Fill each target's quota fully before moving on | +| 3 | `weighted` | Weighted random by per-target weight | +| 4 | `round-robin` | Cycle through targets in order | +| 5 | `p2c` | Power-of-two-choices random load balancing | +| 6 | `least-used` | Pick the target with the lowest current load | +| 7 | `random` | Uniform random pick (deduplicated) | +| 8 | `strict-random` | Random without de-duplicating repeats ๐ŸŽฒ | +| 9 | `cost-optimized` | Minimize $ per request from live catalog pricing ๐Ÿ’ธ | +| 10 | `headroom` | Pick the target with the most remaining quota | +| 11 | `reset-window` | Prefer the target whose quota window resets soonest | +| 12 | `reset-aware` | Rank by quota reset time โ€” short windows first ๐Ÿ“Š | +| 13 | `context-relay` | Hand off context across targets for long conversations ๐Ÿง  | +| 14 | `context-optimized` | Pick the best fit for the current context size | +| 15 | `lkgp` | Last-Known-Good Path โ€” sticky to the last successful target | +| 16 | `auto` | 9-factor live scoring across every connection ๐Ÿค– | +| 17 | `fusion` | Fan out to a panel of models + a judge synthesizes one answer ๐Ÿงฌ | The Auto-Combo engine scores every candidate on **9 factors** (health, quota, cost, latency, success rate, freshnessโ€ฆ) โ€” see [`docs/routing/AUTO-COMBO.md`](docs/routing/AUTO-COMBO.md). ## +### โš–๏ธ Quota-Share โ€” split one subscription across a team โœจ NEW + +> Running several keys against the **same upstream account** (one Codex Pro plan, one Kimi key, one GLM Coding seat)? A burst on one key can burn the whole 5-hour / hourly quota and lock everyone else out. **Quota-Share** distributes a provider's time-based quota **fairly** across the keys in a pool โ€” and it's _work-conserving_, so an idle member's slice is lent out instead of wasted. + +| Knob | What it controls | +| ------------------------ | ------------------------------------------------------------------------------- | +| โš–๏ธ **Allocation weight** | each key's slice of the pool โ€” e.g. `50 / 30 / 20` | +| ๐Ÿ“ **Dimensions** | track `%` ยท requests ยท tokens ยท `$`, per **5h / 7d / per-model** window | +| ๐Ÿšฆ **Policy** | `hard` (block over share) ยท `soft` (deprioritize) ยท `burst` (use idle headroom) | +| ๐Ÿงฑ **Cap** | absolute ceiling per key, independent of mode | + +``` +Pool "team-codex" ยท 1 Codex Pro account ยท 3 keys ยท 5-hour window + โ”œโ”€ alice weight 50 โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘ โ‰ค 50% of the shared 5h quota + โ”œโ”€ bob weight 30 โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘ โ‰ค 30% + โ””โ”€ ci-bot weight 20 โ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘ โ‰ค 20% +Generous mode (<50% pool used) โ†’ idle shares are lent out +Strict mode (โ‰ฅ50% pool used) โ†’ each key held to its fair share +``` + +Enforced in the hot path **before** the request leaves OmniRoute, with per-(key, model) caps + session stickiness for prompt-cache integrity. ๐Ÿ“– [Quota Sharing Engine](docs/routing/QUOTA_SHARE.md) + +## + ### ๐Ÿงฑ Resilience is built in (3 independent layers) | Layer | Scope | What it does | @@ -267,15 +308,15 @@ Result: 4 layers of fallback = zero downtime | Feature | OmniRoute | Other routers | | -------------------------------------- | ------------------------------------------------------------------- | ------------- | -| ๐ŸŒ Providers | **231** | 20โ€“100 | -| ๐Ÿ†“ Free providers | **50+ (11 free forever)** | 1โ€“5 | +| ๐ŸŒ Providers | **236** | 20โ€“100 | +| ๐Ÿ†“ Free providers | **90+ (11 free forever)** | 1โ€“5 | | ๐Ÿ”€ Routing strategies | **17** (priority, weighted, cost-optimized, context-relay, fusionโ€ฆ) | 1โ€“3 | | ๐Ÿ—œ๏ธ Token compression | **RTK + Caveman stacked (15โ€“95%)** | None / 20โ€“40% | -| ๐Ÿงฐ Built-in MCP server | **87 tools, 3 transports, 30 scopes** | Rare | +| ๐Ÿงฐ Built-in MCP server | **94 tools, 3 transports, 30 scopes** | Rare | | ๐Ÿค A2A agent protocol | **6 skills, JSON-RPC 2.0** | None | | ๐Ÿง  Memory (FTS5 + vector) | **Yes** | Rare | | ๐Ÿ›ก๏ธ Guardrails (PII, injection, vision) | **Yes** | Rare | -| โ˜๏ธ Cloud agents | **Codex, Devin, Jules** | None | +| โ˜๏ธ Cloud agents | **Codex, Cursor, Devin, Jules** | None | | ๐Ÿฅท TLS fingerprint stealth | **JA3/JA4 via wreq-js** | None | | ๐Ÿ–ฅ๏ธ Multi-platform | **Web ยท Desktop ยท Termux ยท PWA** | Web only | | ๐ŸŒ i18n | **42 locales** | 0โ€“4 | @@ -290,13 +331,15 @@ Result: 4 layers of fallback = zero downtime -> Recent highlights from **v3.8.20 โ†’ v3.8.41**. Full history in [`CHANGELOG.md`](CHANGELOG.md). +> Recent highlights from **v3.8.20 โ†’ v3.8.43**. Full history in [`CHANGELOG.md`](CHANGELOG.md). +- **๐Ÿ—œ๏ธ Compression hardening** โ€” a default-on **inflation guard** (discard the stacked result and send the verbatim original whenever compression would _grow_ the prompt), completed **Caveman rule packs** for German / French / Japanese (dedup + ultra) plus a new **Chinese (ๆ–‡่จ€ / wรฉnyรกn) input pack** with zh-vs-ja auto-detection, and **RTK filters for Gradle & .NET (`dotnet`)** build output. โ†’ [Compression](docs/compression/COMPRESSION_ENGINES.md) +- **๐Ÿ’ธ Honest flat-rate cost** โ€” subscription / coding-plan providers (ChatGPT Web, grok-web, the Minimax / Kimi / GLM / Alibaba Coding plans, Xiaomi MiMoโ€ฆ) now read **$0** in cost analytics instead of an inflated per-token estimate, while budget / quota / routing keep estimating unchanged. โ†’ [API Reference](docs/reference/API_REFERENCE.md) - **โš–๏ธ Quota-Share routing** โ€” a dedicated combo strategy that spreads load across accounts by _available quota_: Deficit-Round-Robin scheduling, per-connection `max_concurrent` with cooldown-wait queueing, multi-window usage buckets (5h / 7d / per-model), per-(key,model) caps, session stickiness for prompt-cache integrity, and proactive saturation from upstream token-usage headers. โ†’ [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) - **๐Ÿค– One-command CLI/agent setup** โ€” a dedicated `setup-*` command configures each coding tool to route through OmniRoute (Claude Code, Codex, Cline, Continue, Cursor, Roo Code, Kilo Code, Crush, Goose, Qwen Code, Aider, OpenCode); `omniroute launch` / `omniroute launch-codex` are zero-config launchers. โ†’ [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md) - **๐Ÿ›ฐ๏ธ Remote mode** โ€” drive a remote OmniRoute from any machine with scoped access tokens (`omniroute connect` / `omniroute contexts` / `omniroute tokens`), plus an `omniroute login antigravity` helper that runs Google "native/desktop" OAuth on your own machine and pastes a credential blob into a remote/VPS install (where the loopback redirect is unreachable). โ†’ [Remote Mode](docs/guides/REMOTE-MODE.md) - **๐Ÿงญ Smarter auto-routing** โ€” OpenRouter-style `auto/:` combos (e.g. `auto/coding:fast`, `auto/reasoning:pro`), a **Fusion** strategy (fan out to a panel of models in parallel, then synthesize via a judge), **task-aware routing** (best-fit connection per task type), per-request `X-Route-Model` override, live Arena-ELO + models.dev model intelligence, per-step account allowlists, provider-wildcard combo steps, nested combo-ref execution, sticky weighted selection, and `web_search`-aware routing. โ†’ [Auto-Combo](docs/routing/AUTO-COMBO.md) -- **๐Ÿ—œ๏ธ Pluggable compression** โ€” an async pipeline of **9 composable engines** with Compression Studios, an LLMLingua-2 ONNX engine and a heuristic/SLM two-tier **Ultra**, RTK, delegated Anthropic Context Editing, **Output Styles** (output-axis steering: terse-prose / less-code / terse-CJK), an **adaptive context-budget dial** (escalate only as far as needed to fit the context window), per-request `x-omniroute-compression` control, an opt-in offline eval harness, one-click **Headroom** proxy lifecycle management from the dashboard (Docker sidecar supported), a synthetic **compression playground** (Play lanes + A/B Compare with USD-capped fidelity verdicts), an opt-in **per-step fidelity gate** that rejects a lossy engine before it degrades the prompt, a **best-of-N candidate encoder** (GCF vs TOON โ€” keep whichever is shorter, with an A/B bytes/token table in the studio), **CCR ranged/grep/stats retrieval** (pull an exact byte/line slice or summary of a stored block instead of re-expanding it), and a unified panel with named profiles + an active-profile selector. โ†’ [Compression](docs/compression/COMPRESSION_ENGINES.md) +- **๐Ÿ—œ๏ธ Pluggable compression** โ€” an async pipeline of **10 composable engines** with Compression Studios, an LLMLingua-2 ONNX engine and a heuristic/SLM two-tier **Ultra**, RTK, delegated Anthropic Context Editing, **Output Styles** (output-axis steering: terse-prose / less-code / terse-CJK), an **adaptive context-budget dial** (escalate only as far as needed to fit the context window), per-request `x-omniroute-compression` control, an opt-in offline eval harness, one-click **Headroom** proxy lifecycle management from the dashboard (Docker sidecar supported), a synthetic **compression playground** (Play lanes + A/B Compare with USD-capped fidelity verdicts), an opt-in **per-step fidelity gate** that rejects a lossy engine before it degrades the prompt, a **best-of-N candidate encoder** (GCF vs TOON โ€” keep whichever is shorter, with an A/B bytes/token table in the studio), **CCR ranged/grep/stats retrieval** (pull an exact byte/line slice or summary of a stored block instead of re-expanding it), and a unified panel with named profiles + an active-profile selector. โ†’ [Compression](docs/compression/COMPRESSION_ENGINES.md) - **๐Ÿ•ต๏ธ Transparent MITM decrypt (TPROXY)** โ€” capture & translate traffic from CLIs that ignore proxy env vars, with a per-SNI certificate authority and a trust-store installer. โ†’ [MITM/TPROXY](docs/security/MITM-TPROXY-DECRYPT.md) - **๐Ÿ’ธ Cost telemetry everywhere** โ€” `X-OmniRoute-*` cost/usage headers on every endpoint (including media), a non-token cost engine, a cache-HIT `X-OmniRoute-Cost-Saved` header, and per-key USD spend quotas. โ†’ [API Reference](docs/reference/API_REFERENCE.md) - **๐Ÿง  Memory you control** โ€” opt-in int8 vector quantization (Qdrant + sqlite-vec), memory off by default, and a per-request `x-omniroute-no-memory` header. โ†’ [Memory](docs/frameworks/MEMORY.md) @@ -336,7 +379,7 @@ Result: 4 layers of fallback = zero downtime ๏ผ‹ also works with ยท Cline ยท Antigravity ยท Windsurf ยท AMP ยท Hermes ยท Qwen CLI ยท Roo ยท Continue ยท any OpenAI-compatible tool -๐Ÿ“– Per-tool setup for all 16+ tools โ†’ [`docs/reference/CLI-TOOLS.md`](docs/reference/CLI-TOOLS.md) ยท ๐Ÿงฉ OpenCode plugin โ†’ [`@omniroute/opencode-provider`](https://www.npmjs.com/package/@omniroute/opencode-provider) +๐Ÿ“– Per-tool setup for all 24+ tools โ†’ [`docs/reference/CLI-TOOLS.md`](docs/reference/CLI-TOOLS.md) ยท ๐Ÿงฉ OpenCode plugin โ†’ [`@omniroute/opencode-provider`](https://www.npmjs.com/package/@omniroute/opencode-provider) @@ -344,27 +387,60 @@ Result: 4 layers of fallback = zero downtime
-# ๐ŸŒ 231 AI Providers โ€” 50+ Free +# ๐ŸŒ 236 AI Providers โ€” 90+ Free
-> The most complete catalog of any open-source router: **236 providers**, **50+ with a free tier**, **11 free forever**. +> The most complete catalog of any open-source router: **236 providers**, **90+ with a free tier**, **11 free forever**.
+### ๐Ÿข Every major lab โ€” through one endpoint + + + + + + + + + + + + + + + + + + + + + + + + + + +
OpenAI
OpenAI
Anthropic
Anthropic
Gemini
Gemini
xAI Grok
xAI Grok
DeepSeek
DeepSeek
Mistral
Mistral
Qwen
Qwen
Meta Llama
Meta Llama
Groq
Groq
NVIDIA
NVIDIA
MiniMax
MiniMax
Cohere
Cohere
Perplexity
Perplexity
Hugging Face
HuggingFace
Together
Together
Fireworks
Fireworks
Cloudflare
Cloudflare
Baidu
Baidu
+ +โ€ฆand 220+ more โ€” every icon resolves live from the dashboard's provider catalog. ๐Ÿ“– [Provider Reference](docs/reference/PROVIDER_REFERENCE.md) + +
+ ### ๐Ÿ†“ Free Forever โ€” $0, no card - - - - + + + + - - - + + +
AgentRouter
GPT-5, Claude, Gemini
$100 free credits
Qoder AI
Kimi-K2, DeepSeek-R1
Unlimited FREE
Pollinations
GPT-5, Claude, Llama 4
No key needed
LongCat
LongCat-2.0
10M tokens one-time (KYC) ๐Ÿ”‘
AgentRouter
AgentRouter
GPT-5, Claude, Gemini
$100 free credits
Qoder AI
Qoder AI
Kimi-K2, DeepSeek-R1
Unlimited FREE
Pollinations
Pollinations
GPT-5, Claude, Llama 4
No key needed
LongCat
LongCat
LongCat-2.0
10M tokens one-time (KYC) ๐Ÿ”‘
Cloudflare AI
50+ models
10K neurons/day
NVIDIA NIM
129 models
~40 RPM free
Cerebras
Qwen3 235B
1M tokens/day
Cloudflare AI
Cloudflare AI
50+ models
10K neurons/day
NVIDIA NIM
NVIDIA NIM
129 models
~40 RPM free
Cerebras
Cerebras
Qwen3 235B
1M tokens/day
@@ -420,7 +496,7 @@ Result: 4 layers of fallback = zero downtime
-> OmniRoute isn't just a server โ€” it's a **full command-line cockpit** with **60+ commands**, plus open agent protocols so an AI agent can drive OmniRoute **by itself**. +> OmniRoute isn't just a server โ€” it's a **full command-line cockpit** with **80+ commands**, plus open agent protocols so an AI agent can drive OmniRoute **by itself**. ### โŒจ๏ธ A real CLI (not just `start`) @@ -460,7 +536,7 @@ Expose OmniRoute over **MCP** or **A2A** and any capable agent gets the keys to | Protocol | Endpoint | Use it for | | ------------------ | ----------------------------------------------- | ------------------------------------------------------ | | ๐Ÿงฐ **MCP (stdio)** | `omniroute --mcp` | Plug into Claude Desktop, Cursor, any MCP client | -| ๐ŸŒŠ **MCP (HTTP)** | `http://localhost:20128/api/mcp/stream` | Remote MCP โ€” **87 tools**, 30 scopes, full audit trail | +| ๐ŸŒŠ **MCP (HTTP)** | `http://localhost:20128/api/mcp/stream` | Remote MCP โ€” **94 tools**, 30 scopes, full audit trail | | ๐Ÿ“ก **MCP (SSE)** | `http://localhost:20128/api/mcp/sse` | Streaming MCP transport | | ๐Ÿค **A2A** | `http://localhost:20128/.well-known/agent.json` | Agent-to-agent, **JSON-RPC 2.0** + SSE, 6 skills | @@ -479,9 +555,9 @@ claude mcp add-server omniroute --type http --url http://localhost:20128/api/mcp -> **Why use many token when few token do trick?** Every request passes through OmniRoute's compression pipeline **transparently** โ€” no client changes. It's now a **stack of 9 composable engines** that run in order and mix & match per routing combo โ€” building on ideas from [RTK](https://github.com/rtk-ai/rtk), [Caveman](https://github.com/JuliusBrussee/caveman) (โญ 51K+), [LLMLingua-2](https://github.com/microsoft/LLMLingua), and [Troglodita](https://github.com/leninejunior/troglodita) (PT-BR). +> **Why use many token when few token do trick?** Every request passes through OmniRoute's compression pipeline **transparently** โ€” no client changes. It's now a **stack of 10 composable engines** that run in order and mix & match per routing combo โ€” building on ideas from [RTK](https://github.com/rtk-ai/rtk), [Caveman](https://github.com/JuliusBrussee/caveman) (โญ 78K+), [LLMLingua-2](https://github.com/microsoft/LLMLingua), and [Troglodita](https://github.com/leninejunior/troglodita) (PT-BR). -### ๐Ÿงฑ The 9-engine stack +### ๐Ÿงฑ The 10-engine stack Engines run in pipeline order; each is independently toggleable and configurable per combo: @@ -491,11 +567,12 @@ Engines run in pipeline order; each is independently toggleable and configurable | 2 | **CCR** | Archives large blocks behind retrieve markers, fetched on demand | | 3 | **RTK** | Smart tool-result filtering, dedup & truncation (command-aware) | | 4 | **Headroom** | Lossless tabular compaction of homogeneous JSON arrays (~30%+) | -| 5 | **Caveman** | Rule-based prose compression (~65โ€“75% on output) | -| 6 | **LLMLingua-2** | ML semantic pruning via MobileBERT ONNX โ€” code-safe, async | -| 7 | **Lite** | Whitespace + image-URL trimming (latency-light baseline) | -| 8 | **Aggressive** | Summarization + progressive aging of old turns | -| 9 | **Ultra** | Heuristic token pruning with an optional small-model (SLM) tier | +| 5 | **Relevance** | Extractive sentence scoring against the last user query | +| 6 | **Caveman** | Rule-based prose compression (~65โ€“75% on output) | +| 7 | **LLMLingua-2** | ML semantic pruning via MobileBERT ONNX โ€” code-safe, async | +| 8 | **Lite** | Whitespace + image-URL trimming (latency-light baseline) | +| 9 | **Aggressive** | Summarization + progressive aging of old turns | +| 10 | **Ultra** | Heuristic token pruning with an optional small-model (SLM) tier | Code blocks, URLs and structured data are **always preserved** byte-perfect. **One-click presets** combine the engines: @@ -529,7 +606,7 @@ Code blocks, URLs and structured data are **always preserved** byte-perfect. **O ### ๐Ÿ“– How it works โ€” pipeline, architecture & savings math ``` -Client (10,000 tok) โ”€โ”€โ–ถ OmniRoute Compression (9 engines) โ”€โ”€โ–ถ Provider (~1,080 tok, up to 95% saved) +Client (10,000 tok) โ”€โ”€โ–ถ OmniRoute Compression (10 engines) โ”€โ”€โ–ถ Provider (~1,080 tok, up to 95% saved) ``` Default stacked combo runs `RTK โ†’ Caveman`. When both act on the same tool/context payload, savings compound: @@ -544,7 +621,7 @@ Code blocks, URLs, JSON and structured data are **always protected** by the pres ### ๐ŸŽš๏ธ Beyond the engines โ€” output styles, the adaptive dial & per-request control -The 9 engines above shrink what goes **in**. Three more layers shape **how**, **when**, and what comes **out**: +The 10 engines above shrink what goes **in**. Three more layers shape **how**, **when**, and what comes **out**: - **๐Ÿช„ Output Styles** _(output-axis steering)_ โ€” inject deterministic, cache-safe response-shaping instructions; combinable, each at `lite` / `full` / `ultra` intensity. Adding a style is a one-line registry entry: - **Terse prose** โ€” drop filler / articles / hedging; keep technical substance exact. @@ -720,16 +797,16 @@ podman compose --profile base up -d **The $0 Free Stack โ€” combine into one unbreakable combo:** -| Provider | Prefix | Free models | Quota | -| ----------------- | ----------- | ----------------------------------------------- | ----------------- | -| **Kiro** | `kr/` | Claude Sonnet 4.5, Haiku 4.5, Opus 4.6 | 50 credits/mo | -| **Qoder** | `if/` | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 | โ™พ๏ธ Unlimited | -| **Qwen** | `qw/` | qwen3-coder-plus/flash/next | โ™พ๏ธ Unlimited | -| **Pollinations** | `pol/` | GPT-5, Claude, Gemini, DeepSeek, Llama 4 | No key needed | +| Provider | Prefix | Free models | Quota | +| ----------------- | ----------- | ----------------------------------------------- | ------------------ | +| **Kiro** | `kr/` | Claude Sonnet 4.5, Haiku 4.5, Opus 4.6 | 50 credits/mo | +| **Qoder** | `if/` | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 | โ™พ๏ธ Unlimited | +| **Qwen** | `qw/` | qwen3-coder-plus/flash/next | โ™พ๏ธ Unlimited | +| **Pollinations** | `pol/` | GPT-5, Claude, Gemini, DeepSeek, Llama 4 | No key needed | | **LongCat** | `lc/` | LongCat-2.0 | 10M one-time (KYC) | -| **Cloudflare AI** | `cf/` | 50+ models | 10K neurons/day | -| **NVIDIA NIM** | `nvidia/` | 129 models | ~40 RPM | -| **Cerebras** | `cerebras/` | Qwen3 235B, GPT-OSS 120B | 1M tok/day | +| **Cloudflare AI** | `cf/` | 50+ models | 10K neurons/day | +| **NVIDIA NIM** | `nvidia/` | 129 models | ~40 RPM | +| **Cerebras** | `cerebras/` | Qwen3 235B, GPT-OSS 120B | 1M tok/day | > ๐Ÿ’ก The dashboard "cost" is a **savings tracker**, not a bill โ€” OmniRoute never charges you. A "$290 total cost" using free models means **$290 saved**. @@ -780,7 +857,7 @@ Compression: aggressive (~50%) โ†’ double your free quota ยท Cost: $0/mo **Routing:** 15 strategies ยท task-aware smart routing ยท thinking budget controls ยท wildcard routing ยท system prompt injection. **Compatibility:** OpenAI โ†” Claude โ†” Gemini โ†” Responses API ยท auto OAuth refresh (PKCE, 8 providers) ยท multi-account round-robin ยท Batch + Files API ยท live OpenAPI 3.0. -**Protocols:** MCP (87 tools, 3 transports, 30 scopes) ยท A2A (JSON-RPC 2.0, SSE, 6 skills) ยท ACP ยท cloud agents (Codex, Devin, Jules). +**Protocols:** MCP (94 tools, 3 transports, 30 scopes) ยท A2A (JSON-RPC 2.0, SSE, 6 skills) ยท ACP ยท cloud agents (Codex, Cursor, Devin, Jules). **Plugins:** custom plugin marketplace (system-configured registry URL with SSRF-guarded fetch) ยท install / enable / disable ยท Notion + Obsidian knowledge-base integrations (WebDAV file server, vault search, note CRUD). **Embedded services:** one-click install & lifecycle management of local sidecar services (CLIProxy, NineRouter). **Quality & Ops:** built-in **Evals** (golden-set: exact/contains/regex/custom) ยท guardrails (PII, injection, vision) ยท health dashboard ยท p50/p95/p99 telemetry ยท webhooks ยท compliance audit. @@ -874,7 +951,7 @@ Compression: aggressive (~50%) โ†’ double your free quota ยท Cost: $0/mo - **Protocols**: MCP (stdio/HTTP) + A2A v0.3 (JSON-RPC 2.0 + SSE) - **Streaming**: Server-Sent Events (SSE) + WebSocket bridge (`/v1/ws`) - **Auth**: OAuth 2.0 (PKCE) + JWT + API Keys + MCP Scoped Authorization -- **Testing**: Node.js test runner + Vitest (**14,965 test cases** across 517 files โ€” unit, integration, E2E, security, ecosystem) +- **Testing**: Node.js test runner + Vitest (**21,000+ test cases** across 2,586 files โ€” unit, integration, E2E, security, ecosystem) - **Platforms**: Desktop (Electron), Android (Termux), PWA (any browser) - **CI/CD**: GitHub Actions (auto npm publish + Docker Hub on release) - **Website**: [omniroute.online](https://omniroute.online) @@ -937,7 +1014,7 @@ Compression: aggressive (~50%) โ†’ double your free quota ยท Cost: $0/mo | ------------------------------------------------- | --------------------------------------------------- | | [API Reference](docs/reference/API_REFERENCE.md) | All endpoints with examples | | [OpenAPI Spec](docs/openapi.yaml) | OpenAPI 3.0 specification | -| [MCP Server](open-sse/mcp-server/README.md) | 87 MCP tools, IDE configs, Python/TS/Go clients | +| [MCP Server](open-sse/mcp-server/README.md) | 94 MCP tools, IDE configs, Python/TS/Go clients | | [MCP Server Guide](docs/frameworks/MCP-SERVER.md) | MCP installation, transports, and tool reference | | [A2A Server](src/lib/a2a/README.md) | JSON-RPC 2.0 protocol, skills, streaming, task mgmt | | [A2A Server Guide](docs/frameworks/A2A-SERVER.md) | A2A agent card, tasks, skills, and streaming | @@ -951,7 +1028,7 @@ Compression: aggressive (~50%) โ†’ double your free quota ยท Cost: $0/mo | [Security Policy](SECURITY.md) | Vulnerability reporting and security practices | | [i18n Guide](docs/guides/I18N.md) | 40+ language support, translation workflow, RTL | | [Release Checklist](docs/ops/RELEASE_CHECKLIST.md) | Pre-release validation steps | -| [Coverage Plan](docs/ops/COVERAGE_PLAN.md) | Test coverage strategy and 14,965 test suite | +| [Coverage Plan](docs/ops/COVERAGE_PLAN.md) | Test coverage strategy and 21,000+ test suite |
@@ -965,23 +1042,23 @@ Compression: aggressive (~50%) โ†’ double your free quota ยท Cost: $0/mo - oyi77
+ oyi77
oyi77

- ๐Ÿฅ‡ 190 commits โ€ข +72K lines
+ ๐Ÿฅ‡ 189 commits โ€ข +155K lines
Analytics engine, SQL aggregations,
proxy marketplace, test coverage
- Chris Staley
+ Chris Staley
Chris Staley

- ๐Ÿฅˆ 72 commits โ€ข +5.7K lines
+ ๐Ÿฅˆ 70 commits โ€ข +5.7K lines
SSE stream hardening, Responses API,
Gemini pagination, test regression fixes
- zenobit
+ zenobit
zenobit

๐Ÿฅ‰ 62 commits โ€ข +24K lines
@@ -989,20 +1066,28 @@ Compression: aggressive (~50%) โ†’ double your free quota ยท Cost: $0/mo - R.D. & Randi
+ R.D. & Randi
R.D. & Randi

- ๐Ÿ… 107 commits โ€ข +28K lines
+ ๐Ÿ… 108 commits โ€ข +30K lines
Endpoints page, tunnel integrations,
Docker workflows, A2A status, compression UI
- benzntech
+ benzntech
benzntech

- ๐Ÿ… 20 commits โ€ข +7.5K lines
+ ๐Ÿ… 22 commits โ€ข +7.5K lines
Electron desktop app, auto-updater,
release build workflows, cross-platform CI
+ + + herjarsa
+ herjarsa +

+ ๐Ÿ… 21 commits โ€ข +6K lines
+ Zero-latency combos, vision-bridge auto-routing,
catalog context-length, resilience 429 hints
+ @@ -1016,11 +1101,11 @@ Compression: aggressive (~50%) โ†’ double your free quota ยท Cost: $0/mo
-## ๐Ÿ‘ฅ Contributors +## ๐Ÿ‘ฅ 280+ Contributors
-[![Contributors](https://contrib.rocks/image?repo=diegosouzapw/OmniRoute&max=100&columns=20&anon=1)](https://github.com/diegosouzapw/OmniRoute/graphs/contributors) +[![Contributors](https://contrib.rocks/image?repo=diegosouzapw/OmniRoute&max=200&columns=20&anon=1)](https://github.com/diegosouzapw/OmniRoute/graphs/contributors) ### How to Contribute @@ -1085,36 +1170,36 @@ OmniRoute stands on the shoulders of giants. It started as a fork of **[9router] | Project | โญ | How it inspired OmniRoute | | ------------------------------------------------------------------------------- | ----: | ------------------------------------------------------------------------------------------------------------------------------------- | -| **[9router](https://github.com/decolua/9router)** ยท decolua | 17.9k | The original project this fork is built on โ€” extended here with multi-modal APIs and a full TypeScript rewrite. | -| **[CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI)** ยท router-for-me | 37.8k | The Go implementation that inspired this JavaScript / TypeScript port. | -| **[LiteLLM](https://github.com/BerriAI/litellm)** ยท BerriAI | 50.8k | The AI gateway whose public pricing dataset feeds our cost-tracking sync and whose provider-normalization model informed our routing. | +| **[9router](https://github.com/decolua/9router)** ยท decolua | 19.0k | The original project this fork is built on โ€” extended here with multi-modal APIs and a full TypeScript rewrite. | +| **[CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI)** ยท router-for-me | 38.8k | The Go implementation that inspired this JavaScript / TypeScript port. | +| **[LiteLLM](https://github.com/BerriAI/litellm)** ยท BerriAI | 52.1k | The AI gateway whose public pricing dataset feeds our cost-tracking sync and whose provider-normalization model informed our routing. | ### ๐Ÿ—œ๏ธ Context & token compression โ€” engines -| Project | โญ | How it inspired OmniRoute | -| ---------------------------------------------------------------------------- | ----: | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **[Caveman](https://github.com/JuliusBrussee/caveman)** ยท JuliusBrussee | 74.5k | The viral "why use many token when few token do trick" project โ€” its caveman-speak philosophy powers our standard compression mode and 30+ filler/condensation rules. | -| **[RTK โ€“ Rust Token Killer](https://github.com/rtk-ai/rtk)** ยท rtk-ai | 63.6k | High-performance command-output compression โ€” inspired our RTK engine, JSON filter DSL, raw-output recovery and the stacked RTK โ†’ Caveman pipeline. | -| **[headroom](https://github.com/chopratejas/headroom)** ยท chopratejas | 33.6k | Reversible context-compression (SmartCrusher) โ€” inspired our `headroom` engine and the `ccr` retrieve-marker pattern. | -| **[LLMLingua](https://github.com/microsoft/LLMLingua)** ยท Microsoft | 6.3k | Prompt-compression research (LLMLingua / LLMLingua-2) โ€” inspired our async, code-safe, fail-open `llmlingua` engine. | -| **[llmlingua-2-js](https://github.com/atjsh/llmlingua-2-js)** ยท atjsh | 27 | The JS/ONNX port (MobileBERT / XLM-RoBERTa) used as the worker-thread backend for our LLMLingua engine. | -| **[Troglodita](https://github.com/leninejunior/troglodita)** ยท Lenine Jรบnior | 15 | PT-BR token compression โ€” powers our pt-BR language pack: pleonasm reduction and filler removal tuned for Brazilian-Portuguese grammar. | -| **[ponytail](https://github.com/DietrichGebert/ponytail)** ยท DietrichGebert | 51.4k | The viral "lazy senior dev" YAGNI-coder skill โ€” inspired our **less-code** Output Style: smallest-working-change steering that cuts _generated_ code (the output-axis sibling to Caveman's terse prose). | +| Project | โญ | How it inspired OmniRoute | +| ----------------------------------------------------------------------------- | ----: | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **[Caveman](https://github.com/JuliusBrussee/caveman)** ยท JuliusBrussee | 78.2k | The viral "why use many token when few token do trick" project โ€” its caveman-speak philosophy powers our standard compression mode and 30+ filler/condensation rules. | +| **[RTK โ€“ Rust Token Killer](https://github.com/rtk-ai/rtk)** ยท rtk-ai | 67.3k | High-performance command-output compression โ€” inspired our RTK engine, JSON filter DSL, raw-output recovery and the stacked RTK โ†’ Caveman pipeline. | +| **[headroom](https://github.com/headroomlabs-ai/headroom)** ยท headroomlabs-ai | 54.5k | Reversible context-compression (SmartCrusher) โ€” inspired our `headroom` engine and the `ccr` retrieve-marker pattern. | +| **[LLMLingua](https://github.com/microsoft/LLMLingua)** ยท Microsoft | 6.4k | Prompt-compression research (LLMLingua / LLMLingua-2) โ€” inspired our async, code-safe, fail-open `llmlingua` engine. | +| **[llmlingua-2-js](https://github.com/atjsh/llmlingua-2-js)** ยท atjsh | 28 | The JS/ONNX port (MobileBERT / XLM-RoBERTa) used as the worker-thread backend for our LLMLingua engine. | +| **[Troglodita](https://github.com/leninejunior/troglodita)** ยท Lenine Jรบnior | 16 | PT-BR token compression โ€” powers our pt-BR language pack: pleonasm reduction and filler removal tuned for Brazilian-Portuguese grammar. | +| **[ponytail](https://github.com/DietrichGebert/ponytail)** ยท DietrichGebert | 68.8k | The viral "lazy senior dev" YAGNI-coder skill โ€” inspired our **less-code** Output Style: smallest-working-change steering that cuts _generated_ code (the output-axis sibling to Caveman's terse prose). | ### ๐Ÿงฉ Compact formats, token research & code-aware tooling | Project | โญ | How it inspired OmniRoute | | ---------------------------------------------------------------------------------------------- | ----: | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **[TOON](https://github.com/toon-format/toon)** ยท toon-format | 24.6k | Token-Oriented Object Notation โ€” its columnar, header-plus-rows model shaped our tabular compaction stage. | -| **[GCF โ€“ Graph Compact Format](https://github.com/blackwell-systems/gcf)** ยท Blackwell Systems | 11 | Schema-aware "JSON for LLMs" notation โ€” co-inspired our lossless homogeneous-array compaction with `[N rows]` markers. | -| **[token-optimizer-mcp](https://github.com/ooples/token-optimizer-mcp)** ยท ooples | 409 | Brotli/SQLite cache + per-session context-delta โ€” inspired our `session-dedup` engine. | -| **[token-savior](https://github.com/Mibayy/token-savior)** ยท Mibayy | 993 | Bash-output compaction + MCP profiles โ€” inspired our compression bail-out discipline and MCP tool-manifest reduction. | -| **[token-saver](https://github.com/ppgranger/token-saver)** ยท ppgranger | 103 | Content-aware, per-file-type output compression with failure-aware bail-out โ€” validated our per-type dispatch and minimum-gain skip. | -| **[token-optimizer](https://github.com/alexgreensh/token-optimizer)** ยท alexgreensh | 1.4k | "Find the ghost tokens" โ€” its offload + recoverable-handle pattern informed our CCR offload thinking. | -| **[TokenMizer](https://github.com/Shweta-Mishra-ai/tokenmizer)** ยท Shweta-Mishra-ai | 1 | A session-graph + cross-turn line-dedup blueprint that informed our session-dedup design. | +| **[TOON](https://github.com/toon-format/toon)** ยท toon-format | 24.7k | Token-Oriented Object Notation โ€” its columnar, header-plus-rows model shaped our tabular compaction stage. | +| **[GCF โ€“ Graph Compact Format](https://github.com/blackwell-systems/gcf)** ยท Blackwell Systems | 14 | Schema-aware "JSON for LLMs" notation โ€” co-inspired our lossless homogeneous-array compaction with `[N rows]` markers. | +| **[token-optimizer-mcp](https://github.com/ooples/token-optimizer-mcp)** ยท ooples | 421 | Brotli/SQLite cache + per-session context-delta โ€” inspired our `session-dedup` engine. | +| **[token-savior](https://github.com/Mibayy/token-savior)** ยท Mibayy | 1.0k | Bash-output compaction + MCP profiles โ€” inspired our compression bail-out discipline and MCP tool-manifest reduction. | +| **[token-saver](https://github.com/ppgranger/token-saver)** ยท ppgranger | 110 | Content-aware, per-file-type output compression with failure-aware bail-out โ€” validated our per-type dispatch and minimum-gain skip. | +| **[token-optimizer](https://github.com/alexgreensh/token-optimizer)** ยท alexgreensh | 1.5k | "Find the ghost tokens" โ€” its offload + recoverable-handle pattern informed our CCR offload thinking. | +| **[TokenMizer](https://github.com/Shweta-Mishra-ai/tokenmizer)** ยท Shweta-Mishra-ai | 2 | A session-graph + cross-turn line-dedup blueprint that informed our session-dedup design. | | **[OmniCompress](https://github.com/jessefreitas/OmniCompress)** ยท jessefreitas | 2 | Rust columnar-JSON + content-addressed retrieve + cross-message dedup โ€” validated our `headroom`/`ccr`/`session-dedup` engine design and the cache-stable "compressed form is position-independent" invariant. | -| **[mcp-compressor](https://github.com/atlassian-labs/mcp-compressor)** ยท Atlassian Labs | 80 | MCP tool-schema/description compression โ€” informed our MCP tool-manifest cardinality reduction. | -| **[RepoMapper](https://github.com/pdavis68/RepoMapper)** ยท pdavis68 | 182 | Aider-style repo-map ranking โ€” informed our repo-map / retrieval-ranking exploration. | +| **[mcp-compressor](https://github.com/atlassian-labs/mcp-compressor)** ยท Atlassian Labs | 89 | MCP tool-schema/description compression โ€” informed our MCP tool-manifest cardinality reduction. | +| **[RepoMapper](https://github.com/pdavis68/RepoMapper)** ยท pdavis68 | 181 | Aider-style repo-map ranking โ€” informed our repo-map / retrieval-ranking exploration. | | **[quiet-shell-mcp](https://github.com/mrsimpson/quiet-shell-mcp)** ยท mrsimpson | 4 | Declarative shell-output reduction over MCP โ€” validated our declarative bash-output compaction. | | **[ts-morph](https://github.com/dsherret/ts-morph)** ยท David Sherret | 6.1k | TypeScript Compiler API toolkit โ€” inspired our parser-based comment removal that preserves string, template and regex literals. | @@ -1122,27 +1207,27 @@ OmniRoute stands on the shoulders of giants. It started as a fork of **[9router] | Project | โญ | How it inspired OmniRoute | | ------------------------------------------------------------------ | ----: | ------------------------------------------------------------------------------------------------------------------- | -| **[Mem0](https://github.com/mem0ai/mem0)** ยท mem0ai | 58.9k | Universal memory layer โ€” its proxy-as-write/read-boundary model shaped our memory architecture. | -| **[Letta (MemGPT)](https://github.com/letta-ai/letta)** ยท letta-ai | 23.4k | Stateful agents with tiered memory โ€” inspired our Context Control & Recovery (CCR) tiered model. | +| **[Mem0](https://github.com/mem0ai/mem0)** ยท mem0ai | 59.8k | Universal memory layer โ€” its proxy-as-write/read-boundary model shaped our memory architecture. | +| **[Letta (MemGPT)](https://github.com/letta-ai/letta)** ยท letta-ai | 23.6k | Stateful agents with tiered memory โ€” inspired our Context Control & Recovery (CCR) tiered model. | | **[WFGY](https://github.com/onestardao/WFGY)** ยท onestardao | 1.8k | The ProblemMap taxonomy of 16 recurring RAG/LLM failure modes โ€” the shared vocabulary in our troubleshooting guide. | ### ๐Ÿ›ฐ๏ธ Traffic inspection, MITM & transparent proxy | Project | โญ | How it inspired OmniRoute | | --------------------------------------------------------------------------------- | ---: | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **[llm-interceptor](https://github.com/chouzz/llm-interceptor)** ยท chouzz | 46 | MITM interception/analysis of coding-assistant โ†” LLM traffic โ€” our Traffic Inspector ports its SSE merge, conversation normalization, host passthrough and secret masking (MIT). | -| **[ProxyBridge](https://github.com/InterceptSuite/ProxyBridge)** ยท InterceptSuite | 5.1k | Transparent per-process proxy routing โ€” inspired our crash-safe MITM teardown, socket idle-timeouts, `/proc` process attribution and TPROXY capture. | +| **[llm-interceptor](https://github.com/chouzz/llm-interceptor)** ยท chouzz | 48 | MITM interception/analysis of coding-assistant โ†” LLM traffic โ€” our Traffic Inspector ports its SSE merge, conversation normalization, host passthrough and secret masking (MIT). | +| **[ProxyBridge](https://github.com/InterceptSuite/ProxyBridge)** ยท InterceptSuite | 5.3k | Transparent per-process proxy routing โ€” inspired our crash-safe MITM teardown, socket idle-timeouts, `/proc` process attribution and TPROXY capture. | ### ๐Ÿ“š Model data, observability & UI | Project | โญ | How it inspired OmniRoute | | -------------------------------------------------------------------------- | ----: | -------------------------------------------------------------------------------------------------------------------------- | -| **[models.dev](https://github.com/anomalyco/models.dev)** ยท SST / OpenCode | 5.1k | Open database of AI model specs, pricing and capabilities โ€” synced natively into our model catalog. | -| **[React Flow / xyflow](https://github.com/xyflow/xyflow)** ยท xyflow | 37.1k | The node-based graph library powering our real-time Compression Studio and Combo/Routing Studio. | -| **[LangGraph](https://github.com/langchain-ai/langgraph)** ยท LangChain | 35.1k | LangGraph Studio's live workflow-graph visualization inspired our Studios' real-time cascade view. | -| **[Langfuse](https://github.com/langfuse/langfuse)** ยท Langfuse | 29.3k | Its trace โ†’ span โ†’ generation observability model shaped our Compression Studio waterfall. | +| **[models.dev](https://github.com/anomalyco/models.dev)** ยท SST / OpenCode | 5.6k | Open database of AI model specs, pricing and capabilities โ€” synced natively into our model catalog. | +| **[React Flow / xyflow](https://github.com/xyflow/xyflow)** ยท xyflow | 37.4k | The node-based graph library powering our real-time Compression Studio and Combo/Routing Studio. | +| **[LangGraph](https://github.com/langchain-ai/langgraph)** ยท LangChain | 36.1k | LangGraph Studio's live workflow-graph visualization inspired our Studios' real-time cascade view. | +| **[Langfuse](https://github.com/langfuse/langfuse)** ยท Langfuse | 30.1k | Its trace โ†’ span โ†’ generation observability model shaped our Compression Studio waterfall. | | **[Kiali](https://github.com/kiali/kiali)** ยท Kiali | 3.6k | Istio service-mesh observability โ€” inspired our circuit-breaker badges and error-edge visuals in the Routing/Combo Studio. | -| **[lobe-icons](https://github.com/lobehub/lobe-icons)** ยท LobeHub | 2.1k | AI/LLM brand logos that render the provider icons across our dashboard. | +| **[lobe-icons](https://github.com/lobehub/lobe-icons)** ยท LobeHub | 2.2k | AI/LLM brand logos that render the provider icons across our dashboard. | ### ๐Ÿ›ก๏ธ Security @@ -1168,7 +1253,7 @@ MIT License - see [LICENSE](LICENSE) for details. **[โฌ† Back to top](#-omniroute)** ยท Built with โค๏ธ for the open-source AI community. -OmniRoute v3.8.24 ยท Node โ‰ฅ22.0.0 ยท MIT License ยท omniroute.online +OmniRoute v3.8.43 ยท Node โ‰ฅ22.0.0 ยท MIT License ยท omniroute.online