-
-๐ **Available in:** ๐บ๐ธ [English](README.md) | ๐ง๐ท [Portuguรชs (Brasil)](docs/i18n/pt-BR/README.md) | ๐ช๐ธ [Espaรฑol](docs/i18n/es/README.md) | ๐ซ๐ท [Franรงais](docs/i18n/fr/README.md) | ๐ฎ๐น [Italiano](docs/i18n/it/README.md) | ๐ท๐บ [ะ ัััะบะธะน](docs/i18n/ru/README.md) | ๐จ๐ณ [ไธญๆ (็ฎไฝ)](docs/i18n/zh-CN/README.md) | ๐ฉ๐ช [Deutsch](docs/i18n/de/README.md) | ๐ฎ๐ณ [เคนเคฟเคจเฅเคฆเฅ](docs/i18n/in/README.md) | ๐น๐ญ [เนเธเธข](docs/i18n/th/README.md) | ๐บ๐ฆ [ะฃะบัะฐัะฝััะบะฐ](docs/i18n/uk-UA/README.md) | ๐ธ๐ฆ [ุงูุนุฑุจูุฉ](docs/i18n/ar/README.md) | ๐ฆ๐ฟ [Azษrbaycan dili](docs/i18n/az/README.md) | ๐ฏ๐ต [ๆฅๆฌ่ช](docs/i18n/ja/README.md) | ๐ป๐ณ [Tiแบฟng Viแปt](docs/i18n/vi/README.md) | ๐ง๐ฌ [ะัะปะณะฐััะบะธ](docs/i18n/bg/README.md) | ๐ฉ๐ฐ [Dansk](docs/i18n/da/README.md) | ๐ซ๐ฎ [Suomi](docs/i18n/fi/README.md) | ๐ฎ๐ฑ [ืขืืจืืช](docs/i18n/he/README.md) | ๐ญ๐บ [Magyar](docs/i18n/hu/README.md) | ๐ฎ๐ฉ [Bahasa Indonesia](docs/i18n/id/README.md) | ๐ฐ๐ท [ํ๊ตญ์ด](docs/i18n/ko/README.md) | ๐ฒ๐พ [Bahasa Melayu](docs/i18n/ms/README.md) | ๐ณ๐ฑ [Nederlands](docs/i18n/nl/README.md) | ๐ณ๐ด [Norsk](docs/i18n/no/README.md) | ๐ต๐น [Portuguรชs (Portugal)](docs/i18n/pt/README.md) | ๐ท๐ด [Romรขnฤ](docs/i18n/ro/README.md) | ๐ต๐ฑ [Polski](docs/i18n/pl/README.md) | ๐ธ๐ฐ [Slovenฤina](docs/i18n/sk/README.md) | ๐ธ๐ช [Svenska](docs/i18n/sv/README.md) | ๐ต๐ญ [Filipino](docs/i18n/phi/README.md) | ๐จ๐ฟ [ฤeลกtina](docs/i18n/cs/README.md)
+
+
[](https://www.npmjs.com/package/omniroute)
-

-
-
[](https://hub.docker.com/r/diegosouzapw/omniroute)


-[](https://github.com/diegosouzapw/OmniRoute/blob/main/LICENSE)
-
-
-
-[](https://github.com/diegosouzapw)
-[](https://github.com/diegosouzapw)
[](https://omniroute.online)
-[](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t)
----
-
-## ๐ผ๏ธ Main Dashboard
-
-
-

-
-
----
-
-## ๐ธ Dashboard Preview
-
-
-Click to see dashboard screenshots
-
-| Page | Screenshot |
-| -------------- | ------------------------------------------------- |
-| **Providers** |  |
-| **Combos** |  |
-| **Analytics** |  |
-| **Health** |  |
-| **Translator** |  |
-| **Settings** |  |
-| **CLI Tools** |  |
-| **Usage Logs** |  |
-| **Endpoints** |  |
-
---
-### ๐ค Free AI Provider for your favorite coding agents
+
+๐ Table of Contents
-_Connect any AI-powered IDE or CLI tool through OmniRoute โ free API gateway for unlimited coding._
+
-
+- [๐ฅ The Promise](#-the-promise)
+- [๐ค Why OmniRoute?](#-why-omniroute)
+- [๐ฏ Combos โ The Flagship](#-combos--the-flagship)
+- [๐ What Sets OmniRoute Apart](#-what-sets-omniroute-apart)
+- [๐ค Compatible CLIs & Coding Agents](#-compatible-clis--coding-agents)
+- [๐ 177 AI Providers โ 50+ Free](#-177-ai-providers--50-free)
+- [๐ฅ๏ธ Where OmniRoute Runs](#%EF%B8%8F-where-omniroute-runs--anywhere)
+- [๐ Private & Local-First](#-private--local-first)
+- [๐ Full CLI + A2A & MCP](#-full-cli--a2a--mcp)
+- [๐๏ธ Save 15โ95% Tokens](#%EF%B8%8F-save-1595-tokens--automatically)
+- [โก Quick Start](#-quick-start)
+- [๐ฌ OmniRoute in Action](#-omniroute-in-action)
+- [๐ Explore More](#-explore-more)
+- [๐ง Support & Community](#-support--community)
-๐ก All agents connect via http://localhost:20128/v1 or http://cloud.omniroute.online/v1 โ one config, unlimited models and quota
+
---
-## ๐บ OmniRoute in Action โ Video Guides
+## ๐ฅ The Promise
-
+> One endpoint. **177 providers.** Never stop building โ and let OmniRoute pick the cheapest one that works.
-
-
-
-
- ๐ง๐ท Portuguรชs
- Guia completo do OmniRoute
- |
-
-
-
-
- ๐บ๐ธ English
- Complete OmniRoute walkthrough
- |
-
-
-
-
- ๐ท๐บ ะ ัััะบะธะน
- ะะพะปะฝะพะต ััะบะพะฒะพะดััะฒะพ ะฟะพ OmniRoute
- |
+ ๐ซ Never hit limits Auto-fallback across 177 providers in milliseconds. Quota out? Next provider takes over โ zero downtime. |
+ ๐ธ Save up to 95% tokens RTK + Caveman stacked compression cuts 15โ95% of eligible tokens (~89% avg on tool-heavy sessions). |
+ ๐ $0 to start 50+ providers with a free tier, 11 free forever (Kiro, Qoder, Pollinations, LongCatโฆ). No card needed. |
+
+
+ ๐ Every tool works 16+ coding agents โ Claude Code, Codex, Cursor, Cline, Copilot, Antigravity โ through one config. |
+ ๐งฉ One endpoint OpenAI โ Claude โ Gemini โ Responses API translation. Point any tool at /v1 and it just works. |
+ ๐ก๏ธ Production-grade Circuit breakers, TLS stealth, MCP (37 tools), A2A, memory, guardrails, evals. 4,690+ tests. |
-
-
-> ๐ฌ **Made a video about OmniRoute?** We'd love to feature it here! Open an [issue](https://github.com/diegosouzapw/OmniRoute/issues/new) or [discussion](https://github.com/diegosouzapw/OmniRoute/discussions) with the link and we'll add it to this showcase.
-
---
## ๐ค Why OmniRoute?
-**One endpoint. 207+ providers. Never stop building.**
+**Stop juggling 10 dashboards, dead API keys, and surprise bills.**
-Stop juggling 10 dashboards, dead API keys, and surprise bills. OmniRoute
-routes every request through the cheapest viable provider โ automatically.
-
-### The 3-tier fallback (zero-downtime AI)
+| โ The daily pain | โ
How OmniRoute fixes it |
+| ------------------------------------------------------ | ----------------------------------------------------------------------------- |
+| ๐ Subscription quota expires unused every month | **Maximize subscriptions** โ track quota, use every token before reset |
+| ๐ Rate limits stop you mid-coding | **4-tier auto-fallback** โ Subscription โ API โ Cheap โ Free, in milliseconds |
+| ๐ฅ Tool outputs (`git diff`, `grep`, logs) burn tokens | **RTK + Caveman compression** โ save 15โ95% eligible tokens per request |
+| ๐ธ Expensive APIs ($20โ50/mo per provider) | **Cost-optimized routing** โ auto-route to the cheapest viable model |
+| ๐งฐ Each AI tool wants its own setup | **One endpoint, every tool, one dashboard** |
+| ๐ AI blocked in your country | **3-level proxy** + TLS fingerprint stealth โ use AI from anywhere |
```
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
-โ Your IDE / CLI / App โ
-โ (Claude Code, Cursor, Cline, Copilot, โฆ) โ
-โโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
- โ http://localhost:20128/v1
- โผ
+โ Your IDE / CLI (Claude Code, Cursor, Clineโฆ) โ
+โโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
+ โ http://localhost:20128/v1
+ โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
-โ OmniRoute (Smart Router) โ
-โ โข RTK token saver (47 specialized filters) โ
-โ โข Caveman terse-mode (3 levels + SHARED_BOUNDARIES) โ
-โ โข Auto-fallback combos (14 strategies) โ
-โ โข Circuit breaker ยท TLS fingerprint stealth (JA3/JA4) โ
-โ โข Memory ยท MCP server ยท A2A ยท Guardrails ยท Evals โ
-โโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
- โ
- โโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโ
- โผ Tier 1 โผ Tier 2 โผ Tier 3
-SUBSCRIPTION CHEAP FREE
-(Claude Code, (DeepSeek $0.27, (Kiro, OpenCode,
- Codex, Copilot, GLM $0.60, Gemini CLI,
- Cursor, Antigravity) MiniMax $0.20) Vertex $300cr)
-
- quota exhausted? budget hit? always available
- โ falls to Tier 2 โ falls to Tier 3
+โ OmniRoute โ Smart Router โ
+โ RTK + Caveman compression ยท 14 routing strategies โ
+โ Circuit breakers ยท TLS stealth ยท MCP ยท A2A ยท Guardrails โ
+โโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
+ โโโโโโโโโโโโโโโฌโโโโโดโโโโโโโโโฌโโโโโโโโโโโโโโ
+ โผ Tier 1 โผ Tier 2 โผ Tier 3 โผ Tier 4
+ SUBSCRIPTION API KEY CHEAP FREE
+ Claude Code, DeepSeek, GLM $0.5, Kiro, Qoder,
+ Codex, Copilot Groq, xAI MiniMax $0.2 Pollinations
+ quota out? โโโโถ budget hit? โโถ budget hit? โโถ always on
```
-### Why this matters
-
-- โ **Subscription quota wasted** every month? OmniRoute uses every token before expiry.
-- โ **Rate limits stop you mid-flow?** Auto-fallback to the next provider in milliseconds.
-- โ **Tool outputs burn tokens?** RTK compresses `git diff`, logs, and grep results 30-50%.
-- โ **Paying $50/mo across 5 providers?** Route to the cheapest viable model automatically.
-- โ **Each AI tool wants its own setup?** One endpoint, every tool, one dashboard.
-
-### What sets OmniRoute apart
-
-| Feature | OmniRoute | Other routers |
-| ----------------------------------- | ------------------------------------------------------------- | ------------- |
-| Providers | **207+** | 20-100 |
-| Combo strategies | **14** (priority, weighted, cost-optimized, context-relay, โฆ) | 1-3 |
-| Token compression (RTK) | **47 specialized filters** | None |
-| Built-in MCP server | **37 tools, 3 transports, 13 scopes** | Rare |
-| A2A agent protocol | **5 skills, JSON-RPC 2.0** | None |
-| Memory (FTS5 + vector) | **Yes** | Rare |
-| Guardrails (PII, injection, vision) | **Yes** | Rare |
-| Cloud agent integrations | Codex, Devin, Jules | None |
-| Circuit breaker per provider | **3-state, lazy recovery** | Rare |
-| TLS fingerprint stealth | **JA3/JA4 via wreq-js** | None |
-| Eval framework | **Built-in** | Rare |
-| CLI (no Electron required) | **Yes** + system tray | Varies |
-| i18n | **40+ locales** | 0-4 |
-
-See [`docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md`](docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md) for a detailed comparison vs LiteLLM, OpenRouter, and Portkey.
-
---
-**Also solves:**
+## ๐ฏ Combos โ The Flagship
-โ
**Prompt Compression** โ auto-compress prompts & tool outputs, save 15-95% eligible tokens per request with RTK+Caveman stacked mode
-โ
**Maximize subscriptions** โ track quota, use every bit before reset
-โ
**Auto fallback** โ Subscription โ Cheap โ Free, zero downtime
-โ
**Multi-account** โ round-robin between accounts per provider
-โ
**Format translation** โ OpenAI โ Claude โ Gemini โ Responses API, any tool works
-โ
**3-level proxy** โ bypass geo-blocks with global, per-provider, and per-key proxies
-โ
**10 multi-modal APIs** โ chat, images, video, music, audio, search in one endpoint
-โ
**MCP + A2A** โ 37 MCP tools + agent-to-agent protocol, production-ready
-โ
**Universal** โ works with Claude Code, Codex, Gemini CLI, Cursor, Cline, OpenClaw, any CLI tool
+> A **combo** is a chain of models OmniRoute routes across **automatically**. Quota runs out, a provider fails, or costs spike โ the combo silently slides to the next model. **This is what makes OmniRoute unbreakable.** ๐ก๏ธ
----
+### โก Zero-config โ just use `auto`
-## ๐ง Support
+No combo to create. Set your model to `auto` (or a variant) and OmniRoute builds a virtual combo from your connected providers, scored live:
-> ๐ฌ **Join our community!** [WhatsApp Group](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) โ Get help, share tips, and stay updated.
+| Model ID | What it optimizes for |
+| -------------- | -------------------------------------------------------------- |
+| `auto` | ๐ฏ Balanced default (LKGP โ sticks to your last good provider) |
+| `auto/coding` | ๐งโ๐ป Quality-first weights for code generation |
+| `auto/fast` | โก Lowest latency first |
+| `auto/cheap` | ๐ฐ Cheapest per token first |
+| `auto/offline` | ๐ Most quota / rate-limit headroom first |
+| `auto/smart` | ๐ญ Quality-first + 10% exploration to discover better models |
-- **Website**: [omniroute.online](https://omniroute.online)
-- **GitHub**: [github.com/diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute)
-- **Issues**: [github.com/diegosouzapw/OmniRoute/issues](https://github.com/diegosouzapw/OmniRoute/issues)
-- **WhatsApp**: [Community Group](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t)
-- **Contributing**: See [CONTRIBUTING.md](CONTRIBUTING.md), open a PR, or pick a `good first issue`
+### ๐ Or build your own โ 14 routing strategies
-### ๐ Reporting a Bug?
+| Goal | Strategy / combo |
+| --------------------------------------- | -------------------------------------------------- |
+| ๐ฅ Drain my subscription before paying | `priority` / `fill-first` |
+| โ๏ธ Spread load across accounts | `round-robin` ยท `weighted` ยท `p2c` ยท `least-used` |
+| ๐ธ Always cheapest viable model | `cost-optimized` ยท `auto/cheap` |
+| ๐ง Hand off long context between models | `context-relay` ยท `context-optimized` |
+| ๐ฒ Randomized / privacy routing | `random` ยท `strict-random` |
+| ๐ค Just make it smart | `auto` (9-factor scoring) ยท `lkgp` ยท `reset-aware` |
-When opening an issue, please run the system-info command and attach the generated file:
+
The Auto-Combo engine scores every candidate on **9 factors** (health, quota, cost, latency, success rate, freshnessโฆ) โ see [`docs/routing/AUTO-COMBO.md`](docs/routing/AUTO-COMBO.md).
-```bash
-npm run system-info
+### ๐งฑ Resilience is built in (3 independent layers)
+
+| Layer | Scope | What it does |
+| -------------------------- | ----------------- | -------------------------------------------------------------------------- |
+| ๐ **Circuit breaker** | whole provider | Stops hammering a provider that's failing upstream; auto-probes to recover |
+| ๐ค **Connection cooldown** | one account / key | Skips a rate-limited key while other keys keep serving |
+| ๐ฏ **Model lockout** | provider + model | Quarantines just one quota-limited model, not the whole connection |
+
+```
+Combo: "always-on" Strategy: priority
+ 1. cc/claude-opus-4-7 โ subscription (use it fully)
+ 2. cx/gpt-5.5 โ second subscription
+ 3. glm/glm-5.1 โ cheap backup ($0.5/1M)
+ 4. kr/claude-sonnet-4.5 โ FREE, unlimited (never fails)
+Result: 4 layers of fallback = zero downtime
```
-This generates a `system-info.txt` with your Node.js version, OmniRoute version, OS details, installed CLI tools (qoder, gemini, claude, codex, antigravity, droid, etc.), Docker/PM2 status, and system packages โ everything we need to reproduce your issue quickly. Attach the file directly to your GitHub issue.
+
๐ [Auto-Combo Engine](docs/routing/AUTO-COMBO.md) ยท [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md)
---
-## ๐ ๏ธ Supported CLI Tools
+## ๐ What Sets OmniRoute Apart
-OmniRoute works seamlessly with **16+ AI coding tools** โ one config, all tools:
+| Feature | OmniRoute | Other routers |
+| -------------------------------------- | ----------------------------------------------------------- | ------------- |
+| ๐ Providers | **177** | 20โ100 |
+| ๐ Free providers | **50+ (11 free forever)** | 1โ5 |
+| ๐ Routing strategies | **14** (priority, weighted, cost-optimized, context-relayโฆ) | 1โ3 |
+| ๐๏ธ Token compression | **RTK + Caveman stacked (15โ95%)** | None / 20โ40% |
+| ๐งฐ Built-in MCP server | **37 tools, 3 transports, 13 scopes** | Rare |
+| ๐ค A2A agent protocol | **5 skills, JSON-RPC 2.0** | None |
+| ๐ง Memory (FTS5 + vector) | **Yes** | Rare |
+| ๐ก๏ธ Guardrails (PII, injection, vision) | **Yes** | Rare |
+| โ๏ธ Cloud agents | **Codex, Devin, Jules** | None |
+| ๐ฅท TLS fingerprint stealth | **JA3/JA4 via wreq-js** | None |
+| ๐ฅ๏ธ Multi-platform | **Web ยท Desktop ยท Termux ยท PWA** | Web only |
+| ๐ i18n | **40+ locales** | 0โ4 |
+
๐ Detailed comparison vs LiteLLM, OpenRouter & Portkey โ [`docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md`](docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md)
+
+---
+
+## ๐ค Compatible CLIs & Coding Agents
+
+_One config โ `http://localhost:20128/v1` โ and **every** AI IDE or CLI runs on free & low-cost models._
+
+
-๐ Full setup for each tool: [`docs/CLI-TOOLS.md`](docs/CLI-TOOLS.md)
+
+๏ผ also works with ยท Cline ยท Antigravity ยท Windsurf ยท AMP ยท Hermes ยท Qwen CLI ยท Roo ยท Continue ยท any OpenAI-compatible tool
+
-### ๐งฉ OpenCode integration โ `@omniroute/opencode-provider`
-
-[](https://www.npmjs.com/package/@omniroute/opencode-provider)
-[](https://www.npmjs.com/package/@omniroute/opencode-provider)
-[](https://bundlephobia.com/package/@omniroute/opencode-provider)
-[](./@omniroute/opencode-provider/LICENSE)
-
-Schema-valid generator for [`opencode.json`](https://opencode.ai/config.json). Emits a `provider.omniroute` entry that delegates the runtime to [`@ai-sdk/openai-compatible`](https://www.npmjs.com/package/@ai-sdk/openai-compatible) โ every OpenCode request flows through OmniRoute's `/v1` surface and benefits from Auto-Combo routing, circuit breakers, key policies, and observability.
-
-**Two integration paths:**
-
-```bash
-# Path 1 โ CLI generator (ships with OmniRoute)
-omniroute config opencode \
- --baseUrl http://localhost:20128 \
- --apiKey "$OMNIROUTE_API_KEY"
-```
-
-```bash
-# Path 2 โ npm package (for scripted / programmatic setups)
-npm install --save-dev @omniroute/opencode-provider
-```
-
-```ts
-import { writeFileSync } from "node:fs";
-import { buildOmniRouteOpenCodeConfig } from "@omniroute/opencode-provider";
-
-const config = buildOmniRouteOpenCodeConfig({
- baseURL: "http://localhost:20128",
- apiKey: process.env.OMNIROUTE_API_KEY ?? "sk_omniroute",
-});
-
-writeFileSync("opencode.json", JSON.stringify(config, null, 2));
-```
-
-Resulting `opencode.json` (excerpt):
-
-```jsonc
-{
- "$schema": "https://opencode.ai/config.json",
- "provider": {
- "omniroute": {
- "npm": "@ai-sdk/openai-compatible",
- "name": "OmniRoute",
- "options": { "baseURL": "http://localhost:20128/v1", "apiKey": "sk_omniroute" },
- "models": {
- "claude-opus-4-5-thinking": { "name": "claude-opus-4-5-thinking" },
- "gemini-3.1-pro-high": { "name": "gemini-3.1-pro-high" },
- // โฆ
- },
- },
- },
-}
-```
-
-๐ฆ Package: [`@omniroute/opencode-provider`](https://www.npmjs.com/package/@omniroute/opencode-provider) ยท ๐ Full guide: [`docs/frameworks/OPENCODE.md`](docs/frameworks/OPENCODE.md) ยท ๐ Source: [`@omniroute/opencode-provider/`](./@omniroute/opencode-provider)
+
๐ Per-tool setup for all 16+ tools โ [`docs/CLI-TOOLS.md`](docs/CLI-TOOLS.md) ยท ๐งฉ OpenCode plugin โ [`@omniroute/opencode-provider`](https://www.npmjs.com/package/@omniroute/opencode-provider)
---
-## ๐ Supported Providers โ 160+
+## ๐ 177 AI Providers โ 50+ Free
-### ๐ OAuth Providers
+> The most complete catalog of any open-source router: **177 providers**, **50+ with a free tier**, **11 free forever**.
-
-
- Claude Code Anthropic OAuth |
- Antigravity Google OAuth |
- Codex OpenAI OAuth |
- GitHub Copilot GitHub OAuth |
- Cursor Cursor OAuth |
-
-
- Kimi Coding Moonshot OAuth |
- Kilo Code Kilo OAuth |
- Cline Cline OAuth |
- |
-
-
-
-### ๐ Free Providers (No Cost)
+### ๐ Free Forever โ $0, no card
+
๐ข Kiro AI Claude Sonnet/Haiku Unlimited FREE |
๐ข Qoder AI Kimi-K2, DeepSeek-R1 Unlimited FREE |
- ๐ข Pollinations GPT-5, Claude, Llama 4 No API key needed |
- ๐ข Qwen Code Qwen3 Coder Plus Unlimited FREE |
+ ๐ข Pollinations GPT-5, Claude, Llama 4 No key needed |
+ ๐ข LongCat Flash-Lite 50M tokens/day ๐ฅ |
- ๐ข LongCat AI Flash-Lite 50M tokens/day |
๐ข Cloudflare AI 50+ models 10K neurons/day |
- ๐ข Puter AI GPT-4.1, Claude Rate-limited free |
- ๐ข NVIDIA NIM Llama, Mistral 1K req/day free |
+ ๐ข Gemini CLI gemini-3-flash 180K/mo free |
+ ๐ข NVIDIA NIM 129 models ~40 RPM free |
+ ๐ข Cerebras Qwen3 235B 1M tokens/day |
+
-### ๐ API Key Providers (120+)
+### ๐ OAuth (11) ยท ๐ API Key (122) ยท ๐ Self-Hosted (10)
-
-
- | OpenAI |
- Anthropic |
- Gemini |
- DeepSeek |
- Groq |
- xAI (Grok) |
-
-
- | Mistral |
- OpenRouter |
- GLM |
- Kimi |
- MiniMax |
- Fireworks |
-
-
- | Together AI |
- Cerebras |
- Cohere |
- NVIDIA |
- Perplexity |
- SiliconFlow |
-
-
- | Nebius |
- HuggingFace |
- DeepInfra |
- SambaNova |
- Vertex AI |
- Azure OpenAI |
-
-
- | AWS Bedrock |
- Snowflake |
- Databricks |
- Venice.ai |
- AI21 Labs |
- Meta Llama |
-
-
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
-...and 90+ more providers
+โฆand 150+ more providers (full catalog)
-Alibaba ยท Amazon Q ยท AssemblyAI ยท Baidu Qianfan ยท Baseten ยท Black Forest Labs ยท Blackbox ยท Brave Search ยท Bytez ยท CablyAI ยท Cartesia ยท ChatGPT Web ยท Chutes.ai ยท Clarifai ยท Codestral ยท CrofAI ยท DataRobot ยท Deepgram ยท ElevenLabs ยท Empower ยท Exa Search ยท Fal.ai ยท Featherless AI ยท FenayAI ยท FriendliAI ยท Galadriel ยท GigaChat ยท GitLab Duo ยท GLHF Chat ยท GoAPI ยท Heroku AI ยท Hyperbolic ยท IBM watsonx ยท Inference.net ยท Inworld ยท Jina AI ยท Kilo Gateway ยท Lambda AI ยท LaoZhang ยท Linkup Search ยท LlamaGate ยท Maritalk ยท Modal ยท Moonshot AI ยท Morph ยท Muse Spark ยท NanoBanana ยท NanoGPT ยท NLP Cloud ยท Nous Research ยท Novita AI ยท nScale ยท OCI ยท Ollama Cloud ยท OVHcloud ยท PiAPI ยท PlayHT ยท Poe ยท Predibase ยท PublicAI ยท Qwen Code ยท Recraft ยท Reka ยท Runway ยท SAP ยท Scaleway ยท SearchAPI ยท SearXNG ยท Serper ยท Stability AI ยท Synthetic ยท Tavily ยท TheB.AI ยท Topaz ยท Upstage ยท v0 (Vercel) ยท Vercel AI Gateway ยท Volcengine ยท Voyage AI ยท W&B Inference ยท Xiaomi MiMo ยท You.com ยท Z.AI ยท + OpenAI/Anthropic-compatible custom endpoints
+
+
+Alibaba ยท Amazon Q ยท AssemblyAI ยท Baidu Qianfan ยท Baseten ยท Black Forest Labs ยท Blackbox ยท Brave Search ยท Bytez ยท CablyAI ยท Cartesia ยท ChatGPT Web ยท Chutes.ai ยท Clarifai ยท Codestral ยท CrofAI ยท DataRobot ยท Deepgram ยท ElevenLabs ยท Empower ยท Exa Search ยท Fal.ai ยท Featherless AI ยท FriendliAI ยท Galadriel ยท GigaChat ยท GitLab Duo ยท GLHF Chat ยท GoAPI ยท Heroku AI ยท Hyperbolic ยท IBM watsonx ยท Inference.net ยท Inworld ยท Jina AI ยท Kilo Gateway ยท Lambda AI ยท Linkup Search ยท LlamaGate ยท Maritalk ยท Modal ยท Moonshot AI ยท Morph ยท NanoGPT ยท NLP Cloud ยท Nous Research ยท Novita AI ยท nScale ยท OCI ยท Ollama Cloud ยท OVHcloud ยท PiAPI ยท PlayHT ยท Poe ยท Predibase ยท PublicAI ยท Recraft ยท Reka ยท Runway ยท SAP ยท SambaNova ยท Scaleway ยท SearchAPI ยท SearXNG ยท Serper ยท SiliconFlow ยท Snowflake ยท Stability AI ยท Synthetic ยท Tavily ยท TheB.AI ยท Topaz ยท Upstage ยท v0 (Vercel) ยท Vercel AI Gateway ยท Venice.ai ยท Volcengine ยท Voyage AI ยท W&B Inference ยท You.com ยท Z.AI ยท LM Studio ยท vLLM ยท Llamafile ยท oobabooga ยท ComfyUI ยท NVIDIA Triton ยท + OpenAI/Anthropic-compatible custom endpoints
+
+๐ Full machine-readable catalog โ [`docs/reference/PROVIDER_REFERENCE.md`](docs/reference/PROVIDER_REFERENCE.md)
-### ๐ Self-Hosted
+---
-
-
- | LM Studio |
- Ollama |
- vLLM |
- Llamafile |
- Docker Model Runner |
-
-
- | NVIDIA Triton |
- XInference |
- oobabooga |
- ComfyUI |
- SD WebUI |
-
-
+## ๐ฅ๏ธ Where OmniRoute Runs โ Anywhere
+
+> Same app, your machine, your rules. From a global npm install to **your phone** via Termux.
+
+| Platform | Install | Highlights |
+| ------------------------- | -------------------------------------------- | --------------------------------------------------------- |
+| ๐ฆ **npm (global)** | `npm install -g omniroute` | One command, any OS |
+| ๐ณ **Docker** | `docker run โฆ diegosouzapw/omniroute` | Multi-arch **AMD64 + ARM64** |
+| ๐ฅ๏ธ **Desktop (Electron)** | `npm run electron:build` | Native window + system tray โ **Windows / macOS / Linux** |
+| ๐ช **ARM** | native `arm64` | Raspberry Pi, ARM servers, Apple Silicon |
+| ๐ฑ **Android (Termux)** | `pkg install nodejs-lts && npx -y omniroute` | Runs **on your phone**, 24/7, no root |
+| ๐ฒ **PWA** | "Add to Home Screen" | Fullscreen, offline, installable from browser |
+| ๐งฉ **OpenCode plugin** | `@omniroute/opencode-provider` | Native OpenCode integration |
+| ๐ ๏ธ **From source** | `npm install && npm run dev` | Hack on it, contribute |
+
+
๐ [Docker Guide](docs/DOCKER_GUIDE.md) ยท [Desktop](electron/README.md) ยท [Termux](docs/TERMUX_GUIDE.md) ยท [PWA](docs/PWA_GUIDE.md) ยท [OpenCode](docs/frameworks/OPENCODE.md)
---
-## ๐ค AI Agent Skills
+## ๐ Private & Local-First
-Drop-in markdown manifests that let any AI agent consume OmniRoute via one fetch.
+> Your keys, your machine, your data. OmniRoute is a **local proxy** โ it never phones home.
-Tell your agent (Claude Desktop, ChatGPT, Cursor, Cline, etc.):
+- ๐ **Runs 100% on your hardware** โ npm, Docker, desktop, or your phone. No OmniRoute cloud sits in the request path.
+- ๐ **Credentials encrypted at rest** โ API keys & OAuth tokens sealed with **AES-256-GCM**.
+- ๐ซ **Zero telemetry by default** โ your prompts go only to the providers _you_ choose, nowhere else.
+- ๐ก๏ธ **Hardened gateway** โ API-key scoping, IP filtering, rate limits, prompt-injection guard, loopback-only process routes.
+- ๐ **MIT licensed & fully open-source** โ audit every line, self-host forever.
-> "Fetch this URL and use OmniRoute according to its instructions:
-> `https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute/SKILL.md`"
-
-10 skills available โ see [skills/README.md](./skills/README.md).
+
๐ [Authorization](docs/architecture/AUTHZ_GUIDE.md) ยท [Guardrails](docs/security/GUARDRAILS.md) ยท [Compliance](docs/security/COMPLIANCE.md)
---
-## ๐ How It Works
+## ๐ Full CLI + A2A & MCP
-```
-โโโโโโโโโโโโโโโ
-โ Your CLI โ (Claude Code, Codex, Gemini CLI, OpenClaw, Cursor, Cline...)
-โ Tool โ
-โโโโโโโโฌโโโโโโโ
- โ http://localhost:20128/v1
- โ
-โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
-โ OmniRoute (Smart Router) โ
-โ โข ๐๏ธ Prompt Compression (save 15-95% eligible) โ
-โ โข Format translation (OpenAI โ Claude โ Gemini) โ
-โ โข Quota tracking + Embeddings + Images โ
-โ โข Auto token refresh + Rate limit management โ
-โโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
- โ
- โโโ [Tier 1: SUBSCRIPTION] Claude Code, Codex, Gemini CLI
- โ โ quota exhausted
- โโโ [Tier 2: API KEY] DeepSeek, Groq, xAI, Mistral, NVIDIA NIM, etc.
- โ โ budget limit
- โโโ [Tier 3: CHEAP] GLM ($0.6/1M), MiniMax ($0.2/1M)
- โ โ budget limit
- โโโ [Tier 4: FREE] Qoder, Qwen, Kiro (unlimited)
+> OmniRoute isn't just a server โ it's a **full command-line cockpit** with **60+ commands**, plus open agent protocols so an AI agent can drive OmniRoute **by itself**.
-Result: Never stop coding, minimal cost + 15-95% eligible token savings
+### โจ๏ธ A real CLI (not just `start`)
+
+```bash
+omniroute # serve gateway + dashboard (port 20128)
+omniroute chat # interactive TUI chat client (slash: /model /combo /skill /memory)
+omniroute setup # guided first-run wizard
+omniroute doctor # diagnose providers, ports, native deps
```
+
+
+`providers` ยท `oauth` ยท `keys` ยท `combo` ยท `nodes` ยท `models` ยท `cache` ยท `compression` ยท `cost` ยท `usage` ยท `quota` ยท `health` ยท `resilience` ยท `telemetry` ยท `logs` ยท `audit` ยท `mcp` ยท `a2a` ยท `cloud` ยท `memory` ยท `skills` ยท `eval` ยท `tunnel` ยท `backup` ยท `sync` ยท `webhooks` ยท `policy` ยท `pricing` ยท `translator` ยท `simulate` โฆ
+
+
+
+### ๐ค Connect an agent โ and it controls OmniRoute itself
+
+Expose OmniRoute over **MCP** or **A2A** and any capable agent gets the keys to the whole gateway โ routing, providers, combos, cache, compression, memory โ autonomously.
+
+| Protocol | Endpoint | Use it for |
+| ------------------ | ----------------------------------------------- | ------------------------------------------------------ |
+| ๐งฐ **MCP (stdio)** | `omniroute --mcp` | Plug into Claude Desktop, Cursor, any MCP client |
+| ๐ **MCP (HTTP)** | `http://localhost:20128/api/mcp/stream` | Remote MCP โ **37 tools**, 13 scopes, full audit trail |
+| ๐ก **MCP (SSE)** | `http://localhost:20128/api/mcp/sse` | Streaming MCP transport |
+| ๐ค **A2A** | `http://localhost:20128/.well-known/agent.json` | Agent-to-agent, **JSON-RPC 2.0** + SSE, 5 skills |
+
+```bash
+# Give Claude Code the full OmniRoute toolset over MCP:
+claude mcp add-server omniroute --type http --url http://localhost:20128/api/mcp/stream
+```
+
+
๐ [MCP Server](docs/frameworks/MCP-SERVER.md) ยท [A2A Server](docs/frameworks/A2A-SERVER.md) ยท [Agent Protocols](docs/frameworks/AGENT_PROTOCOLS_GUIDE.md)
+
---
-## ๐๏ธ Prompt Compression โ Save 15-95% Eligible Tokens Automatically
+## ๐๏ธ Save 15โ95% Tokens โ Automatically
-> **Why use many token when few token do trick?** OmniRoute's built-in compression pipeline reduces token usage before requests reach the provider. It combines ideas from [RTK - Rust Token Killer](https://github.com/rtk-ai/rtk) and [Caveman](https://github.com/JuliusBrussee/caveman) (โญ 51K+).
+> **Why use many token when few token do trick?** Every request passes through OmniRoute's compression pipeline **transparently** โ no client changes. It stacks ideas from [RTK](https://github.com/rtk-ai/rtk) and [Caveman](https://github.com/JuliusBrussee/caveman) (โญ 51K+).
-### How It Works
+| Mode | Savings | Best for |
+| ------------------------------ | ---------- | --------------------------- |
+| ๐ชถ **Lite** | ~15% | Always-on safe default |
+| ๐ชจ **Standard (Caveman)** | ~30% | Daily coding |
+| โก **Aggressive** | ~50% | Long tool-heavy sessions |
+| ๐ฅ **Ultra** | ~75% | Maximum savings |
+| ๐งฐ **RTK** | 60โ90% | Shell/test/build/git output |
+| ๐ **Stacked (RTK โ Caveman)** | **78โ95%** | Mixed prompts + tool logs |
-Every request passes through the compression pipeline **transparently** โ no client changes needed:
+**Real example โ Standard mode:**
-```
-โโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ
-โ Client sends โโโโโโถโ OmniRoute Compression โโโโโโถโ Provider โ
-โ full prompt โ โ Pipeline (7 options) โ โ receives โ
-โ (10,000 tok) โ โ โ โ compressed โ
-โ โ โ ๐ชถ Lite ........... ~15% โ โ (~1,080 tok)โ
-โ โ โ ๐ชจ Standard ....... ~30% โ โ โ
-โ โ โ โก Aggressive ..... ~50% โ โ ๐ฐ up to 95%โ
-โ โ โ ๐ฅ Ultra .......... ~75% โ โ โ
-โ โ โ ๐งฐ RTK ............ 60-90% โ โ โ
-โ โ โ ๐ Stacked ........ 78-95% โ โ โ
-โโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ
-```
-
-### 7 Compression Options
-
-| Mode | Savings | Technique | Best For |
-| ------------------------- | ------- | ----------------------------------------------------------------------------------------------- | -------------------------------------- |
-| **Off** | 0% | No compression | When you need exact prompts |
-| **๐ชถ Lite** | ~15% | Whitespace collapse, dedup system prompts, image URL shortening | Always-on safe default |
-| **๐ชจ Standard (Caveman)** | ~30% | 30+ regex rules: filler removal, context condensation, structural compression, multi-turn dedup | Daily coding with Claude/Codex |
-| **โก Aggressive** | ~50% | All standard + progressive message aging + tool result summarization + LLM-based compression | Long sessions with many tool calls |
-| **๐ฅ Ultra** | ~75% | All aggressive + heuristic token pruning + stopword removal + score-based filtering | Maximum savings when tokens are scarce |
-| **๐งฐ RTK** | 60-90% | 49 command-aware filters, RTK-style JSON DSL, verify gate, trust-gated custom filters | Shell/test/build/git output in agents |
-| **๐ Stacked** | 78-95% | RTK first, then Caveman input condensation; ~89% with upstream average math | Mixed prompts with tool logs + prose |
-
-### RTK + Caveman Savings Math
-
-These numbers are based on the upstream project READMEs under `_references/_outros`:
-
-| Source | Upstream claim used by OmniRoute docs |
-| ------- | ------------------------------------------------------------------------------------------------------------------- |
-| Caveman | `~75%` fewer output tokens; benchmark average `65%` output savings, range `22-87%`; `~46%` input compression tool |
-| RTK | `60-90%` command-output token savings; sample session `~118,000 -> ~23,900` tokens, which is `79.7%` saved (`~80%`) |
-
-For the default stacked compression combo, OmniRoute runs:
-
-```txt
-RTK -> Caveman
-```
-
-When both engines can act on the same tool/context payload, the savings compound:
-
-```txt
-combined = 1 - (1 - RTK savings) * (1 - Caveman input savings)
-average = 1 - (1 - 0.80) * (1 - 0.46) = 89.2%
-range = 1 - (1 - 0.60..0.90) * (1 - 0.46) = 78.4-94.6%
-```
-
-Caveman output mode is separate from prompt compression. When enabled for responses, use Caveman's
-own upstream output numbers: `65%` average, `~75%` headline, `22-87%` observed range. Total bill
-savings depend on the prompt/output mix, but coding-agent sessions are often tool-context heavy, so
-the `RTK -> Caveman` combo is the best default for maximum context savings.
-
-### Before & After (Standard/Caveman Mode)
-
-**๐ฃ๏ธ Before compression (69 tokens):**
-
-> "The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I would recommend using useMemo to memoize the object."
-
-**๐ชจ After compression (19 tokens):**
-
-> "New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo."
-
-**Same answer. 72% less tokens. Zero accuracy loss.**
-
-### Architecture
-
-```
-Request Body
- โ
- โโ strategySelector.ts โโโ Picks mode (config / combo override / auto-trigger)
- โ
- โโ lite.ts โโโโโโโโโโโโโโโ Whitespace, dedup, image URLs, redundant content
- โโ caveman.ts โโโโโโโโโโโโ 30+ regex rules via cavemanRules.ts
- โ โโ preservation.ts โโโ Protects code blocks, URLs, JSON from compression
- โโ engines/rtk/ โโโโโโโโโโ Command detection + JSON DSL filters + raw-output recovery
- โโ engines/registry.ts โโโ Shared engine registry for caveman, RTK, and stacked
- โโ aggressive.ts โโโโโโโโโ Summarizer + tool result compressor + progressive aging
- โ โโ summarizer.ts โโโโโ Rule-based message summarization
- โ โโ toolResultCompressor.ts โโ file/grep/shell/JSON/error compression
- โ โโ progressiveAging.ts โโโโ Older messages โ shorter summaries
- โโ ultra.ts โโโโโโโโโโโโโโ Heuristic token scoring + pruning
- โโ ultraHeuristic.ts โ Stopword detection, score thresholds, force-preserve
-```
-
-### Configuration
-
-```
-Dashboard โ Context & Cache โ Caveman / RTK / Compression Combos
-```
-
-Or per-combo override:
-
-```json
-{
- "comboOverrides": {
- "my-coding-combo": "standard",
- "my-cheap-combo": "ultra"
- }
-}
-```
-
-Auto-trigger: set `autoTriggerTokens` to automatically enable compression when a request exceeds a token threshold.
-
-Compression combos can also assign a named compression pipeline to routing combos, so a coding combo can use RTK + Caveman while a paid subscription combo stays on lite mode.
-
-> ๐ชจ **Fun fact:** The standard/caveman mode is inspired by [Caveman](https://github.com/JuliusBrussee/caveman) โ the viral project that reports 65% average output-token savings while keeping technical accuracy. OmniRoute takes this further with a **7-option pipeline** and a default `RTK -> Caveman` combo that can reach ~89% average savings on eligible tool/context payloads.
-
-๐ **Full compression documentation:** [`docs/COMPRESSION_GUIDE.md`](docs/COMPRESSION_GUIDE.md) โข [`docs/RTK_COMPRESSION.md`](docs/RTK_COMPRESSION.md) โข [`docs/COMPRESSION_ENGINES.md`](docs/COMPRESSION_ENGINES.md) โข [`docs/COMPRESSION_RULES_FORMAT.md`](docs/COMPRESSION_RULES_FORMAT.md) โข [`docs/COMPRESSION_LANGUAGE_PACKS.md`](docs/COMPRESSION_LANGUAGE_PACKS.md)
-
----
-
-## ๐ฏ What OmniRoute Solves
-
-> **Every developer using AI tools faces these problems daily.** OmniRoute solves them all.
-
-| # | Problem | OmniRoute Solution |
-| --- | ---------------------------------------- | ----------------------------------------------------------------------------------------------- |
-| ๐ธ | Subscription quota expires mid-coding | **Smart 4-Tier Fallback** โ auto-routes Subscription โ API Key โ Cheap โ Free |
-| ๐ | Each provider has a different API format | **Format Translation** โ unified endpoint translates OpenAI โ Claude โ Gemini โ Responses |
-| ๐ | AI providers block my country/region | **3-Level Proxy** โ global, per-provider, and per-key proxy with TLS fingerprint spoofing |
-| ๐ | Can't afford AI subscriptions | **11 Free Providers** โ Kiro, Qoder, Pollinations, LongCat, Cloudflare AI, NVIDIA NIM... |
-| ๐ | Gateway is exposed without protection | **API Key Management** โ scoping, rotation, IP filtering, rate limiting, prompt injection guard |
-| ๐ | Provider went down, lost coding flow | **Circuit Breakers** โ auto-failover with cooldown, retry, anti-thundering herd |
-| ๐ง | Configuring each CLI tool is tedious | **CLI Tools Dashboard** โ one-click setup for Claude Code, Codex, Cursor, OpenClaw, Kilo |
-| ๐ | Managing OAuth tokens is hell | **Auto Token Refresh** โ OAuth PKCE for 8 providers, multi-account, LAN/remote fix |
-| ๐ | Don't know how much I'm spending | **Cost Analytics** โ per-token tracking, budget limits, usage stats per API key |
-| ๐ | Can't diagnose errors in AI calls | **Unified Logs** โ 4-tab dashboard (request, proxy, audit, console) + p50/p95/p99 telemetry |
+> **Before (69 tokens):** _"The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I would recommend using useMemo to memoize the object."_
+>
+> **After (19 tokens):** _"New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo."_
+>
+> **Same answer. 72% fewer tokens. Zero accuracy loss.** โ
-๐ See all 31 problems OmniRoute solves
+๐ How it works โ pipeline, architecture & savings math
-| # | Problem | Solution |
-| --- | --------------------------------------------- | -------------------------------------------------------------------------------------------------- |
-| 11 | Deploying/maintaining is complex | npm global, Docker multi-arch, Electron, Termux โ deploy anywhere |
-| 12 | Interface is English-only | 40+ languages with RTL support |
-| 13 | Need more than chat (images, audio, video) | 10 multi-modal APIs: embeddings, images, video, music, TTS, STT, moderation, rerank, search, batch |
-| 14 | No way to test/compare models | LLM Evals, Translator Playground, Chat Tester, Live Monitor |
-| 15 | Need to scale without losing performance | Semantic cache, request dedup, rate limit detection, queue & pacing |
-| 16 | Want to control model behavior globally | System prompt injection, thinking budget, wildcard routing |
-| 17 | Need MCP tools as first-class features | 29 MCP tools, 3 transports (stdio/SSE/HTTP), 10 scopes, audit trail |
-| 18 | Need A2A orchestration | JSON-RPC 2.0 + SSE streaming, task lifecycle, sync + stream paths |
-| 19 | Need real MCP process health | Runtime heartbeat, PID tracking, UI status cards |
-| 20 | Need auditable MCP execution | SQLite-backed audit with filters, pagination, stats |
-| 21 | Need scoped MCP permissions | 10 granular scopes per integration |
-| 22 | Need operational controls without redeploying | Combo switches, resilience tuning, breaker resets from dashboard |
-| 23 | Need A2A task lifecycle visibility | Task listing/filtering, drill-down, cancellation |
-| 24 | Need active stream metrics | Active stream counters, per-state counts, A2A dashboard cards |
-| 25 | Need standard agent discovery | Agent Card at `/.well-known/agent.json` |
-| 26 | Need protocol discoverability | Consolidated Endpoints page with Proxy, MCP, A2A, API tabs |
-| 27 | Need E2E protocol validation | Real MCP SDK + A2A client flows in `test:protocols:e2e` |
-| 28 | Need unified observability | Health + audit + telemetry across OpenAI, MCP, and A2A layers |
-| 29 | Need one runtime for proxy + tools + agents | OpenAI proxy + MCP + A2A in one stack with shared auth/resilience |
-| 30 | Need agentic workflows without glue-code | Unified endpoint, protocol UIs, production-ready foundations |
-| 31 | Long sessions crash with context limits | Proactive context compression, structural integrity guards, multi-layer dropping |
+
+
+```
+Client (10,000 tok) โโโถ OmniRoute Compression (7 options) โโโถ Provider (~1,080 tok, up to 95% saved)
+```
+
+Default stacked combo runs `RTK โ Caveman`. When both act on the same tool/context payload, savings compound:
+
+```txt
+combined = 1 โ (1 โ RTK) ร (1 โ Caveman_input)
+average = 1 โ (1 โ 0.80) ร (1 โ 0.46) = 89.2%
+range = 78.4 โ 94.6%
+```
+
+Code blocks, URLs, JSON and structured data are **always protected** by the preservation engine. Auto-trigger compression by token threshold, or assign a compression pipeline per routing combo.
+
+๐ [`COMPRESSION_GUIDE.md`](docs/COMPRESSION_GUIDE.md) ยท [`RTK_COMPRESSION.md`](docs/RTK_COMPRESSION.md) ยท [`COMPRESSION_ENGINES.md`](docs/COMPRESSION_ENGINES.md)
-๐ **Deep dives:** [Resilience Guide](docs/RESILIENCE_GUIDE.md) โข [Proxy Guide](docs/PROXY_GUIDE.md) โข [Setup Guide](docs/SETUP_GUIDE.md) โข [Compression Guide](docs/COMPRESSION_GUIDE.md)
-
---
-## ๐ Start Free โ Zero Configuration Cost
-
-> Setup AI coding in minutes at **$0/month**. Connect these free accounts and use the built-in **Free Stack** combo.
-
-| Step | Action | Providers Unlocked |
-| ---- | -------------------------------------------------- | ------------------------------------------------------------------ |
-| 1 | Connect **Kiro** (AWS Builder ID OAuth) | Claude Sonnet 4.5, Haiku 4.5 โ **unlimited** |
-| 2 | Connect **Qoder** (Google OAuth) | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1... โ **unlimited** |
-| 3 | Connect **Qwen** (Device Code) | qwen3-coder-plus, qwen3-coder-flash... โ **unlimited** |
-| 4 | Connect **Gemini CLI** (Google OAuth) | gemini-3-flash, gemini-2.5-pro โ **180K/mo free** |
-| 5 | `/dashboard/combos` โ **Free Stack ($0)** template | Round-robin all free providers automatically |
-
-**Point any IDE/CLI to:** `http://localhost:20128/v1` ยท API Key: `any-string` ยท Done.
-
-> **Optional extra coverage (also free):** Groq API key (30 RPM free), NVIDIA NIM (40 RPM free, 70+ models), Cerebras (1M tok/day), LongCat API key (50M tokens/day!), Cloudflare Workers AI (10K Neurons/day, 50+ models).
-
## โก Quick Start
-### 1) Install and run
+**1) Install & run**
```bash
npm install -g omniroute
omniroute
```
-Dashboard opens at `http://localhost:20128` ยท API at `http://localhost:20128/v1`.
+Dashboard at `http://localhost:20128` ยท API at `http://localhost:20128/v1`.
-### 2) Connect providers
+**2) Connect a FREE provider (no signup)**
-1. Dashboard โ **Providers** โ connect at least one provider (OAuth or API key)
-2. Dashboard โ **Endpoints** โ create an API key
-3. Dashboard โ **Combos** โ set your fallback chain (optional)
+Dashboard โ **Providers** โ connect **Kiro AI** (free Claude unlimited) or **OpenCode Free** (no auth) โ done.
-### 3) Point your coding tool
+**3) Point your coding tool**
```txt
Base URL: http://localhost:20128/v1
-API Key: [copy from Endpoint page]
-Model: if/kimi-k2-thinking (or any provider/model)
+API Key: [copy from Dashboard โ Endpoints]
+Model: auto (zero-config smart routing โ or any provider/model)
```
-Works with Claude Code, Codex CLI, Gemini CLI, Cursor, Cline, OpenClaw, OpenCode, and any OpenAI-compatible tool.
-
-
-๐ฆ More install methods (Docker, source, Arch, Void, pnpm)
-
-**Docker:**
+**4) Verify it's working**
```bash
-docker run -d --name omniroute --restart unless-stopped -p 20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest
+curl http://localhost:20128/v1/models -H "Authorization: Bearer YOUR_KEY"
```
-**From source:**
+You should see your connected models listed. ๐ That's it โ start coding, and OmniRoute auto-routes & falls back for you.
+
+
+๐ฆ More install methods โ Docker, source, pnpm, Arch
+
+
+
+**๐ณ Docker**
+
+```bash
+docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \
+ -p 20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest
+```
+
+**๐ ๏ธ From source**
```bash
cp .env.example .env && npm install
-PORT=20128 DASHBOARD_PORT=20129 NEXT_PUBLIC_BASE_URL=http://localhost:20129 npm run dev
+PORT=20128 npm run dev
```
-**pnpm:** `pnpm install -g omniroute && pnpm approve-builds -g && omniroute`
-
-**Arch Linux (AUR):** `yay -S omniroute-bin && systemctl --user enable --now omniroute.service`
-
-**MCP:** `omniroute --mcp` (stdio transport)
-
-**CLI options:** `omniroute setup`, `omniroute doctor`, `omniroute providers available`, `omniroute providers list`, `omniroute --port 3000`, `omniroute --no-open`, `omniroute --help`
-
-**Split-port mode:** `PORT=20128 DASHBOARD_PORT=20129 omniroute`
-
-**Uninstall:** `npm run uninstall` (keeps data) or `npm run uninstall:full` (removes everything)
-
-๐ Full details: [Setup Guide](#-setup-guide) ยท [Docker](#-docker) ยท [Void Linux template](#-quick-start)
-
-
-
----
-
-## ๐ณ Docker
-
-OmniRoute is available as a public Docker image on [Docker Hub](https://hub.docker.com/r/diegosouzapw/omniroute).
-
-**Quick run:**
+**๐ฆ pnpm**
```bash
-docker run -d \
- --name omniroute \
- --restart unless-stopped \
- --stop-timeout 40 \
- -p 20128:20128 \
- -v omniroute-data:/app/data \
- diegosouzapw/omniroute:latest
+pnpm install -g omniroute && pnpm approve-builds -g && omniroute
```
-**With environment file:**
+**๐ง Arch Linux (AUR)**
```bash
-# Copy and edit .env first
-cp .env.example .env
-
-docker run -d \
- --name omniroute \
- --restart unless-stopped \
- --stop-timeout 40 \
- --env-file .env \
- -p 20128:20128 \
- -v omniroute-data:/app/data \
- diegosouzapw/omniroute:latest
+yay -S omniroute-bin && systemctl --user enable --now omniroute.service
```
-**Using Docker Compose:**
-
-```bash
-# Base profile (no CLI tools)
-docker compose --profile base up -d
-
-# CLI profile (Claude Code, Codex, OpenClaw built-in)
-docker compose --profile cli up -d
-```
-
-Dashboard support for Docker deployments now includes a one-click **Cloudflare Quick Tunnel** on `Dashboard โ Endpoints`. The first enable downloads `cloudflared` only when needed, starts a temporary tunnel to your current `/v1` endpoint, and shows the generated `https://*.trycloudflare.com/v1` URL directly below your normal public URL. Endpoint tunnel panels, including Cloudflare, Tailscale, and ngrok, can be shown or hidden from `Settings โ Appearance` without changing active tunnel state.
-
-Notes:
-
-- Quick Tunnel URLs are temporary and change after every restart.
-- Quick Tunnels are not auto-restored after an OmniRoute or container restart. Re-enable them from the dashboard when needed.
-- Managed install currently supports Linux, macOS, and Windows on `x64` / `arm64`.
-- Managed Quick Tunnels default to HTTP/2 transport to avoid noisy QUIC UDP buffer warnings in constrained container environments. Set `CLOUDFLARED_PROTOCOL=quic` or `auto` if you want a different transport.
-- Docker images bundle system CA roots and pass them to managed `cloudflared`, which avoids TLS trust failures when the tunnel bootstraps inside the container.
-- SQLite runs in WAL mode. `docker stop` should be allowed to finish so OmniRoute can checkpoint the latest changes back into `storage.sqlite`.
-- The bundled Compose files already set a 40s stop grace period. If you run the image directly, keep `--stop-timeout 40` (or similar) so manual stops do not cut off shutdown cleanup.
-- Set `CLOUDFLARED_BIN=/absolute/path/to/cloudflared` if you want OmniRoute to use an existing binary instead of downloading one.
-
-**Using Docker Compose with Caddy (HTTPS Auto-TLS):**
-
-OmniRoute can be securely exposed using Caddy's automatic SSL provisioning. Ensure your domain's DNS A record points to your server's IP.
-
-```yaml
-services:
- omniroute:
- image: diegosouzapw/omniroute:latest
- container_name: omniroute
- restart: unless-stopped
- volumes:
- - omniroute-data:/app/data
- environment:
- - PORT=20128
- - NEXT_PUBLIC_BASE_URL=https://your-domain.com
-
- caddy:
- image: caddy:latest
- container_name: caddy
- restart: unless-stopped
- ports:
- - "80:80"
- - "443:443"
- command: caddy reverse-proxy --from https://your-domain.com --to http://omniroute:20128
-
-volumes:
- omniroute-data:
-```
-
-| Image | Tag | Size | Description |
-| ------------------------ | -------- | ------ | --------------------- |
-| `diegosouzapw/omniroute` | `latest` | ~250MB | Latest stable release |
-| `diegosouzapw/omniroute` | `3.7.8` | ~250MB | Current version |
-
-๐ **Full Docker documentation:** [`docs/DOCKER_GUIDE.md`](docs/DOCKER_GUIDE.md) โ Compose profiles, Caddy HTTPS, Cloudflare tunnels, and more.
-
----
-
-## ๐ฑ Multi-Platform โ Run Anywhere
-
-> OmniRoute runs on **Web**, **Desktop (Electron)**, **Android (Termux)**, and as a **Progressive Web App (PWA)**.
-
-| Platform | Install | Highlights |
-| -------------- | -------------------------------------------- | -------------------------------------------------------------------------- |
-| ๐ฅ๏ธ **Desktop** | `npm run electron:build` | Native window, system tray, auto-start, offline mode โ Windows/macOS/Linux |
-| ๐ฑ **Android** | `pkg install nodejs-lts && npx -y omniroute` | ARM native, no root, 24/7 via Termux:Boot โ your phone is an AI server |
-| ๐ฒ **PWA** | "Add to Home Screen" in browser | Fullscreen, offline page, service worker caching โ Android/iOS/Desktop |
-
-
-๐ฅ๏ธ Desktop App details
-
-- Native Electron app with system tray, auto-start, native notifications
-- One-click install: NSIS (Windows), DMG (macOS), AppImage (Linux)
-- Dev: `npm run electron:dev` ยท Build: `npm run electron:build`
-- ๐ Full docs: [`electron/README.md`](electron/README.md)
-
-
-
-
-๐ฑ Android (Termux) details
-
-```bash
-pkg update && pkg install nodejs-lts python build-essential git
-npx -y omniroute@latest
-```
-
-Access from any device on the same network: `http://PHONE_IP:20128/v1`
-
-- ๐ Full guide: [`docs/TERMUX_GUIDE.md`](docs/TERMUX_GUIDE.md)
-
-
-
-
-๐ฒ PWA details
-
-- **Android (Chrome):** โฎ โ "Add to Home screen"
-- **iOS (Safari):** Share โ "Add to Home Screen"
-- **Desktop (Chrome/Edge):** Install icon in address bar
-- ๐ Full docs: [`docs/PWA_GUIDE.md`](docs/PWA_GUIDE.md)
+๐ [Docker Guide](docs/DOCKER_GUIDE.md) โ Compose profiles, Caddy HTTPS, Cloudflare tunnels.
---
-## ๐ Bypass Geographic Blocks โ Use AI From Any Country
+## ๐ฌ OmniRoute in Action
-> ๐ท๐บ ๐จ๐ณ ๐ฎ๐ท ๐จ๐บ ๐น๐ท **In Russia, China, Iran, or any blocked region?** OmniRoute's 3-level proxy system solves this completely.
+
+
+
+
+ 
+ ๐ง๐ท Portuguรชs Guia completo
+ |
+
+ 
+ ๐บ๐ธ English Complete walkthrough
+ |
+
+ 
+ ๐ท๐บ ะ ัััะบะธะน ะะพะปะฝะพะต ััะบะพะฒะพะดััะฒะพ
+ |
+
+
+
-| Level | Badge | Configure In | Use Case |
-| ------------------ | ----- | ------------------ | ------------------------------- |
-| **Global** | ๐ข | Settings โ Proxy | All traffic through one proxy |
-| **Per-Provider** | ๐ก | Provider โ Proxy | Only specific providers proxied |
-| **Per-Connection** | ๐ต | Connection โ Proxy | Each API key uses its own proxy |
-
-**What gets proxied:** API requests โ
โข OAuth flows โ
โข Connection tests โ
โข Token refresh โ
โข Model sync โ
-
-**Protocols:** HTTP/HTTPS, SOCKS5 (`ENABLE_SOCKS5_PROXY=true`), Authenticated proxies
-
-### ๐ 1proxy โ Free Proxy Marketplace
-
-> Contributed by [@oyi77](https://github.com/oyi77) โ [#1847](https://github.com/diegosouzapw/OmniRoute/pull/1847)
-
-No proxy? Use the built-in **1proxy** integration for **hundreds of free, validated proxies** worldwide:
-
-- One-click sync (up to 500 proxies) โข Quality scores (0-100) โข Country filter โข Auto-rotation (quality/random/sequential) โข Auto-degradation โข Circuit breaker
-
-### Anti-Detection
-
-- ๐ **TLS Fingerprint Spoofing** โ browser-like TLS via `wreq-js`
-- ๐ **CLI Fingerprint Matching** โ matches native CLI binary signatures
-- ๐ **Proxy IP Preservation** โ stealth + IP masking simultaneously
-
-๐ **Full proxy documentation:** [`docs/PROXY_GUIDE.md`](docs/PROXY_GUIDE.md)
+> ๐ฌ **Made a video about OmniRoute?** Open an [issue](https://github.com/diegosouzapw/OmniRoute/issues/new) or [discussion](https://github.com/diegosouzapw/OmniRoute/discussions) with the link โ we'll feature it here.
---
----
-
-## ๐ฐ Pricing at a Glance
-
-| Tier | Provider | Cost | Quota Reset | Best For |
-| ------------------- | --------------------------- | ------------------------- | ---------------- | --------------------------------- |
-| **๐ณ SUBSCRIPTION** | Claude Code (Pro) | $20/mo | 5h + weekly | Already subscribed |
-| | Codex (Plus/Pro) | $20-200/mo | 5h + weekly | OpenAI users |
-| | Gemini CLI | **FREE** | 180K/mo + 1K/day | Everyone! |
-| | GitHub Copilot | $10-19/mo | Monthly | GitHub users |
-| **๐ API KEY** | NVIDIA NIM | **FREE** (dev forever) | ~40 RPM | 70+ open models |
-| | Cerebras | **FREE** (1M tok/day) | 60K TPM / 30 RPM | World's fastest |
-| | Groq | **FREE** (30 RPM) | 14.4K RPD | Ultra-fast Llama/Gemma |
-| | DeepSeek V3.2 | $0.27/$1.10 per 1M | None | Best price/quality reasoning |
-| | xAI Grok-4 Fast | **$0.20/$0.50 per 1M** ๐ | None | Fastest + tool calling, ultralow |
-| | xAI Grok-4 (standard) | $0.20/$1.50 per 1M ๐ | None | Reasoning flagship from xAI |
-| | Mistral | Free trial + paid | Rate limited | European AI |
-| | OpenRouter | Pay-per-use | None | 100+ models aggr. |
-| | AgentRouter ๐ | Pay-per-use | None | $200 free credits at signup |
-| **๐ฐ CHEAP** | GLM-5 (via Z.AI) ๐ | $0.5/1M | Daily 10AM | 128K output, newest flagship |
-| | GLM-4.7 | $0.6/1M | Daily 10AM | Budget backup |
-| | MiniMax M2.5 ๐ | $0.3/1M input | 5-hour rolling | Reasoning + agentic tasks |
-| | MiniMax M2.1 | $0.2/1M | 5-hour rolling | Cheapest option |
-| | Kimi K2.5 (Moonshot API) ๐ | Pay-per-use | None | Direct Moonshot API access |
-| | Kimi K2 | $9/mo flat | 10M tokens/mo | Predictable cost |
-| **๐ FREE** | Qoder | **$0** | Unlimited | 5 models unlimited |
-| | Qwen | **$0** | Unlimited | 4 models unlimited |
-| | Kiro | **$0** | Unlimited | Claude Sonnet/Haiku (AWS Builder) |
-| | LongCat Flash-Lite ๐ | **$0** (50M tok/day ๐ฅ) | 1 RPS | Largest free quota on Earth |
-| | Pollinations AI ๐ | **$0** (no key needed) | 1 req/15s | GPT-5, Claude, DeepSeek, Llama 4 |
-| | Cloudflare Workers AI ๐ | **$0** (10K Neurons/day) | ~150 resp/day | 50+ models, global edge |
-| | Scaleway AI ๐ | **$0** (1M tokens total) | Rate limited | EU/GDPR, Qwen3 235B, Llama 70B |
-
-> ๐ **New models added (Mar 2026):** Grok-4 Fast family at $0.20/$0.50/M (benchmarked at 1143ms โ 30% faster than Gemini 2.5 Flash), GLM-5 via Z.AI with 128K output, MiniMax M2.5 reasoning, DeepSeek V3.2 updated pricing, Kimi K2.5 via Moonshot direct API.
-
-**๐ก See the full [$0 Free Stack (11 providers)](#-free-models--11-providers-0-forever) below.**
-
-> ๐ก **Understanding Dashboard Costs:**
->
-> The "cost" displayed in the Usage Analytics page is **for tracking and comparison purposes only**.
-> OmniRoute itself **never charges you anything** โ it's free, open-source software running on your machine.
-> If your dashboard shows "$290 total cost" while using free models, that's how much you **saved** compared to paid API pricing.
-> Think of it as a **savings tracker**, not a bill.
-
----
-
-## ๐ Free Models โ 11 Providers, $0 Forever
-
-> Combine all free providers into one unbreakable combo โ OmniRoute auto-routes between them when quota runs out.
-
-| Provider | Prefix | Free Models | Quota |
-| ----------------- | ----------- | ------------------------------------------------------------- | -------------------- |
-| **Kiro** | `kr/` | Claude Sonnet 4.5, Haiku 4.5, Opus 4.6 | 50 CREDITS per month |
-| **Qoder** | `if/` | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1, minimax-m2.1 | โพ๏ธ Unlimited |
-| **Qwen** | `qw/` | qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next | โพ๏ธ Unlimited |
-| **Pollinations** | `pol/` | GPT-5, Claude, Gemini, DeepSeek, Llama 4, Mistral | No key needed |
-| **LongCat** | `lc/` | LongCat-Flash-Lite | 50M tokens/day ๐ฅ |
-| **Gemini CLI** | `gc/` | gemini-3-flash, gemini-2.5-pro | 180K tok/mo |
-| **Cloudflare AI** | `cf/` | 50+ models (Llama, Gemma, Mistral, Whisper) | 10K Neurons/day |
-| **Groq** | `groq/` | Llama 3.3 70B, Qwen3 32B, Kimi K2 | 14.4K RPD |
-| **NVIDIA NIM** | `nvidia/` | 129 models (DeepSeek, Llama, GLM, Kimi) | ~40 RPM |
-| **Cerebras** | `cerebras/` | Qwen3 235B, GPT-OSS 120B, Llama 3.1 | 1M tok/day |
-| **Scaleway** | `scw/` | Qwen3 235B, Llama 70B, DeepSeek V3 | 1M tokens (EU) |
+## ๐ Explore More
-๐ 25+ more free providers โ Groq, Cerebras, Mistral, GitHub Models, OpenRouter, and more
+๐ฐ Pricing at a glance & the $0 Free Stack (11 providers)
-**Also free (API Key required):**
-Mistral (1B tok/month) ยท OpenRouter (35+ `:free` models) ยท GitHub Models (GPT-5, 45+ models) ยท
-Cohere (1K calls/month) ยท Z.AI/GLM (permanent free Flash models) ยท SiliconFlow (1K RPM, 50K TPM) ยท
-Kilo Code (~200 req/hr auto-router) ยท HuggingFace ($0.10/mo credits) ยท Ollama Cloud (400+ models) ยท
-LLM7.io (30+ models) ยท Kluster AI ยท IBM watsonx (300K tok/month) ยท OpenCode Zen ยท Vercel AI Gateway ($5/mo)
+
-**Trial credits (one-time):**
-Baseten ($30) ยท NLP Cloud ($15) ยท AI21 ($10) ยท Upstage ($10) ยท SambaNova ($5) ยท Modal ($5/mo) ยท
-Fireworks ($1) ยท Nebius ($1) ยท Inference.net ($1 + $25 survey) ยท Hyperbolic ($1) ยท Novita ($0.50)
+| Tier | Example | Cost |
+| --------------------------- | ---------------------------------------- | ---------- |
+| ๐ณ **Subscription** | Claude Code Pro / Codex / Copilot | $10โ200/mo |
+| ๐ **API Key (free tiers)** | NVIDIA NIM, Cerebras, Groq | **FREE** |
+| ๐ฐ **Cheap** | GLM-5 $0.5/1M ยท MiniMax M2.5 $0.3/1M | pennies |
+| ๐ **Free Forever** | Kiro, Qoder, Qwen, Pollinations, LongCat | **$0** |
-**China-based (free tiers):**
-ModelScope ยท Tencent Hunyuan ยท Volcengine ยท ChatAnywhere ยท InternAI ยท Bigmodel
+**The $0 Free Stack โ combine into one unbreakable combo:**
-**Combined capacity: ~31,000+ RPD ยท ~32B+ tokens/month ยท 500+ models ยท $0**
+| Provider | Prefix | Free models | Quota |
+| ----------------- | ----------- | ----------------------------------------------- | ----------------- |
+| **Kiro** | `kr/` | Claude Sonnet 4.5, Haiku 4.5, Opus 4.6 | 50 credits/mo |
+| **Qoder** | `if/` | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 | โพ๏ธ Unlimited |
+| **Qwen** | `qw/` | qwen3-coder-plus/flash/next | โพ๏ธ Unlimited |
+| **Pollinations** | `pol/` | GPT-5, Claude, Gemini, DeepSeek, Llama 4 | No key needed |
+| **LongCat** | `lc/` | LongCat-Flash-Lite | 50M tokens/day ๐ฅ |
+| **Cloudflare AI** | `cf/` | 50+ models | 10K neurons/day |
+| **NVIDIA NIM** | `nvidia/` | 129 models | ~40 RPM |
+| **Cerebras** | `cerebras/` | Qwen3 235B, GPT-OSS 120B | 1M tok/day |
+
+> ๐ก The dashboard "cost" is a **savings tracker**, not a bill โ OmniRoute never charges you. A "$290 total cost" using free models means **$290 saved**.
+
+๐ Complete free directory โ [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md) โ 25+ providers, quotas, base URLs.
-๐ **Complete free provider directory:** [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md) โ 25+ providers, quotas, base URLs, model tables, and OmniRoute combo setup.
+
+๐ฏ Use Cases โ ready-made combo playbooks
----
+
-## ๐๏ธ Free Transcription Combo
-
-> Transcribe any audio/video for **$0** โ Deepgram leads with $200 free, AssemblyAI $50 fallback, Groq Whisper as unlimited emergency backup.
-
-| Provider | Free Credits | Best Model | Rate Limit |
-| ----------------- | ---------------------- | -------------------------------------------- | ---------------------------- |
-| ๐ข **Deepgram** | **$200 free** (signup) | `nova-3` โ best accuracy, 30+ languages | No RPM limit on free credits |
-| ๐ต **AssemblyAI** | **$50 free** (signup) | `universal-3-pro` โ chapters, sentiment, PII | No RPM limit on free credits |
-| ๐ด **Groq** | **Free forever** | `whisper-large-v3` โ OpenAI Whisper | 30 RPM (rate limited) |
-
----
-
-**Suggested combo in `/dashboard/combos`:**
+**$0 forever:**
```
-Name: free-transcription
-Strategy: Priority
-Nodes:
- [1] deepgram/nova-3 โ uses $200 free first
- [2] assemblyai/universal-3-pro โ fallback when Deepgram credits run out
- [3] groq/whisper-large-v3 โ free forever, emergency fallback
+1. kr/claude-sonnet-4.5 (Kiro โ unlimited)
+2. if/kimi-k2-thinking (Qoder โ unlimited)
+3. pol/gpt-5 (Pollinations โ no key)
+4. lc/longcat-flash-lite (50M tok/day backup)
+Compression: aggressive (~50%) โ double your free quota ยท Cost: $0/mo
```
-Then in `/dashboard/media` โ **Transcription** tab: upload any audio or video file โ select your combo endpoint โ get transcription in supported formats.
+**24/7 no interruptions:** chain 2 subscriptions โ cheap โ free for 5 layers of fallback.
+**Blocked region:** free providers + global/per-provider proxy โ access AI from any country.
+**Max savings:** subscription + cheap backup + `ultra` compression (~75%) โ ~$150โ300/mo saved for heavy users.
-## ๐ก Key Features
-
-> **4,690+ automated tests** across 517 test files. Not just a relay โ a full operational platform.
-
-| Feature | Why It Matters |
-| ---------------------------------------------------------------------------------------------------- | -------------------------------- |
-| ๐ง **Smart 4-Tier Fallback** โ Subscription โ API โ Cheap โ Free | Never stop coding, zero downtime |
-| ๐ **Format Translation** โ OpenAI โ Claude โ Gemini โ Responses API | Works with ANY CLI tool |
-| ๐๏ธ **Prompt Compression** โ 7 options including Caveman, RTK, and stacked pipelines | Save 15-95% eligible tokens |
-| ๐ค **MCP Server** โ 37 tools, 3 transports (stdio/SSE/HTTP), 10 scopes | IDE/agent tool integration |
-| ๐ก๏ธ **Resilience Engine** โ circuit breakers, cooldowns, TLS spoofing, anti-thundering herd | Auto-recovery from any failure |
-| ๐ต **10 Multi-Modal APIs** โ chat, embed, images, video, music, TTS, STT, moderation, rerank, search | One endpoint for everything |
-| ๐ **3-Level Proxy** โ global, per-provider, per-key + 1proxy free marketplace | Access AI from any country |
-| ๐ **Full Observability** โ unified logs, p50/p95/p99 telemetry, cost tracking, budget controls | Know exactly what's happening |
+
-๐ Complete feature list โ 30+ capabilities
+๐ Bypass geo-blocks โ 3-level proxy + stealth
-**Routing & Intelligence**
+
-- 13 balancing strategies (priority, weighted, round-robin, P2C, cost-optimized, context-relay...)
-- Task-aware smart routing (coding/vision/analysis) ยท Context relay session handoffs
-- Thinking budget controls (passthrough/auto/custom) ยท Wildcard routing ยท System prompt injection
+๐ท๐บ ๐จ๐ณ ๐ฎ๐ท ๐จ๐บ ๐น๐ท In a blocked region? OmniRoute's **3-level proxy** (Global / Per-Provider / Per-Connection) proxies API requests, OAuth flows, connection tests, token refresh & model sync.
-**Translation & Compatibility**
+- **Protocols:** HTTP/HTTPS, SOCKS5, authenticated proxies
+- **๐ 1proxy marketplace** โ hundreds of free validated proxies, quality scores, auto-rotation
+- **Anti-detection** โ TLS fingerprint spoofing (`wreq-js`), CLI fingerprint matching, proxy IP preservation
-- Auto token refresh (OAuth PKCE for 8 providers) ยท Multi-account round-robin
-- Responses API โ full `/v1/responses` for Codex ยท Batch API with Files API
-- OpenAPI 3.0 live spec + Try-It UI
+๐ [`docs/PROXY_GUIDE.md`](docs/PROXY_GUIDE.md)
-**Protocols**
+
-- A2A Server โ JSON-RPC 2.0, SSE streaming, task lifecycle, skills
-- ACP โ CLI agent discovery (14 agents + custom)
+
+โจ Full feature list โ 30+ capabilities (memory, evals, observability)
-**Platform**
+
-- Desktop (Electron) ยท Android (Termux) ยท PWA ยท Docker (AMD64 + ARM64)
-- Cloudflare / Tailscale / ngrok tunnels ยท 40+ languages with RTL
-- Semantic + signature cache (two-tier) ยท Request idempotency + deduplication
+**Routing:** 14 strategies ยท task-aware smart routing ยท thinking budget controls ยท wildcard routing ยท system prompt injection.
+**Compatibility:** OpenAI โ Claude โ Gemini โ Responses API ยท auto OAuth refresh (PKCE, 8 providers) ยท multi-account round-robin ยท Batch + Files API ยท live OpenAPI 3.0.
+**Protocols:** MCP (37 tools, 3 transports, 13 scopes) ยท A2A (JSON-RPC 2.0, SSE, skills) ยท ACP ยท cloud agents (Codex, Devin, Jules).
+**Quality & Ops:** built-in **Evals** (golden-set: exact/contains/regex/custom) ยท guardrails (PII, injection, vision) ยท health dashboard ยท p50/p95/p99 telemetry ยท webhooks ยท compliance audit.
+**AI Agent Skills:** drop-in markdown manifests โ point any agent at `skills/omniroute/SKILL.md`. 10 skills available.
-**Observability**
+๐ [MCP Server](open-sse/mcp-server/README.md) ยท [A2A Server](src/lib/a2a/README.md) ยท [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) ยท [Features Gallery](docs/FEATURES.md)
-- Health dashboard โ uptime, breakers, cache, lockouts
-- Evaluation framework โ golden set testing ยท Webhooks ยท Compliance audit
+
-**v3.6+ Highlights:**
-V1 WebSocket Bridge ยท Sync Tokens & Config Bundle ยท GLM Thinking (glmt) ยท Hybrid Token Counting ยท
-Safe Outbound Fetch ยท Wait For Cooldown ยท Runtime Env Validation ยท Vision Bridge ยท
-Grok-4 Fast ยท GLM-5 via Z.AI ยท MiniMax M2.5 ยท toolCalling flag ยท
-Multilingual Intent Detection ยท Benchmark-Driven Fallbacks ยท Request Deduplication
+
+๐ Setup, env vars & FAQ
-**Architecture Examples:**
+
-```txt
-Combo: "my-coding-stack" Format Translation:
- 1. cc/claude-opus-4-7 CLI โ OpenAI format
- 2. nvidia/llama-3.3-70b OmniRoute โ translates
- 3. glm/glm-4.7 Provider โ native format
- 4. if/kimi-k2-thinking
-```
+| Env var | Default | Purpose |
+| ----------------- | -------------- | -------------------------------- |
+| `PORT` | `20128` | API + dashboard port |
+| `REQUIRE_API_KEY` | `false` | Require API key for all requests |
+| `DATA_DIR` | `~/.omniroute` | Database & config storage |
-๐ [MCP Server README](open-sse/mcp-server/README.md) ยท [A2A Server README](src/lib/a2a/README.md) ยท [Resilience Guide](docs/RESILIENCE_GUIDE.md) ยท [Features Gallery](docs/FEATURES.md)
+**Will I be charged by OmniRoute?** No โ it's free, open-source software on your machine. You only pay paid providers directly. OmniRoute has no billing system.
+**Are FREE providers really unlimited?** Yes โ Kiro, Qoder, Pollinations, LongCat, Cloudflare. No catch.
+**Will compression hurt quality?** No โ it only compresses the **input**; code, URLs, JSON are always protected.
+**Does it work where AI is blocked?** Yes โ 3-level proxy + 1proxy marketplace reach all 177 providers.
+
+๐ [User Guide](docs/USER_GUIDE.md) ยท [API Reference](docs/API_REFERENCE.md) ยท [Environment Config](docs/ENVIRONMENT.md)
+
+
+
+
+๐ Troubleshooting
+
+
+
+| Problem | Quick fix |
+| ----------------------------------------- | ------------------------------------------------------------- |
+| "Language model did not provide messages" | Provider quota exhausted โ use a combo fallback |
+| Rate limiting (429) | Add fallback: `cc/claude โ glm/glm-4.7 โ if/kimi-k2-thinking` |
+| OAuth token expired | Auto-refreshed; if stuck, delete + re-auth in Providers |
+| `unsupported_country_region_territory` | Configure proxy in Settings โ Proxy |
+| Docker SQLite locks | Use `--stop-timeout 40` for clean WAL checkpoint |
+| Node runtime errors | Use Node `>=20.20.2 <21`, `>=22.22.2 <23`, or `>=24 <25` |
+
+๐ **Reporting a bug?** Run `npm run system-info` and attach `system-info.txt`. ๐ [`docs/TROUBLESHOOTING.md`](docs/TROUBLESHOOTING.md)
+
+
+
+
+๐ธ Dashboard screenshots
+
+
+
+| Page | Screenshot | Page | Screenshot |
+| ---------- | ------------------------------------------------- | ---------- | --------------------------------------------- |
+| Providers |  | Combos |  |
+| Analytics |  | Health |  |
+| Translator |  | Settings |  |
+| CLI Tools |  | Usage Logs |  |
---
-## ๐ฏ Use Cases โ Ready-Made Combo Playbooks
+## ๐ง Support & Community
-### Case 0: "I want zero-config, auto-routing NOW"
+> ๐ฌ **Join our WhatsApp groups** โ get help, share tips, stay updated:
+> ยท [**๐ International**](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) ยท [**๐ง๐ท Portuguรชs**](https://chat.whatsapp.com/CeGCxdFzqBe5Uki288wOvf)
-**Problem:** Don't want to create combos manually. Just want AI routing to work immediately.
-
-```bash
-# No combo creation needed! Use auto/ prefix directly:
-model: "auto" # Default LKGP routing across all connected providers
-model: "auto/coding" # Quality-first weights for code generation
-model: "auto/fast" # Low-latency routing (fastest provider first)
-model: "auto/cheap" # Cost-optimized (cheapest per token)
-model: "auto/offline" # High availability (most quota available)
-model: "auto/smart" # Best discovery (10% exploration rate)
-```
-
-**How it works:**
-
-1. Add providers in Dashboard โ Providers (OAuth or API key)
-2. Use `auto/` prefix in any AI tool โ **no combo creation needed**
-3. OmniRoute dynamically builds a virtual combo from your active connections
-4. Routes using LKGP (Last Known Good Provider) + 6-factor scoring
-5. Session stickiness ensures consistent provider selection
-
-**Dashboard indicator:** A blue banner at the top shows "Auto-Routing Active" with a link to `/dashboard/combos` for configuration.
-
-**Monthly cost:** $0 (uses your existing free providers) or whatever your connected providers cost
+- ๐ **Website**: [omniroute.online](https://omniroute.online)
+- ๐ **GitHub**: [github.com/diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute)
+- ๐ **Issues**: [report a bug](https://github.com/diegosouzapw/OmniRoute/issues) (attach `npm run system-info` output)
+- ๐ค **Contributing**: see [CONTRIBUTING.md](CONTRIBUTING.md) or pick a `good first issue`
---
-### Case 1: "I have a Claude Pro subscription"
-
-**Problem:** Quota expires unused, rate limits during heavy coding sessions.
-
-```
-Combo: "maximize-claude"
- 1. cc/claude-opus-4-7 (use subscription fully)
- 2. glm/glm-5.1 (cheap backup when quota out โ $0.5/1M)
- 3. kr/claude-sonnet-4.5 (free emergency fallback via Kiro)
-
-Compression: standard (caveman) โ saves 30% tokens = stretch quota further
-Monthly cost: $20 (subscription) + ~$3 (backup) = $23 total
-vs. $20 + hitting limits + lost productivity = frustration
-```
-
-### Case 2: "I want $0 forever"
-
-**Problem:** Can't afford subscriptions, need reliable AI for coding.
-
-```
-Combo: "free-forever"
- 1. kr/claude-sonnet-4.5 (Claude 4.5 free unlimited via Kiro)
- 2. if/kimi-k2-thinking (reasoning model free via Qoder)
- 3. pol/gpt-5 (GPT-5 free via Pollinations โ no key)
- 4. lc/longcat-flash-lite (50M tokens/day free backup)
-
-Compression: aggressive โ saves 50% tokens = double your free quota
-Monthly cost: $0
-Quality: Production-ready models + 50% token savings
-```
-
-### Case 3: "I need 24/7 coding, no interruptions"
-
-**Problem:** Deadlines, can't afford any downtime.
-
-```
-Combo: "always-on"
- 1. cc/claude-opus-4-7 (best quality โ subscription)
- 2. cx/gpt-5.5 (second subscription โ OpenAI)
- 3. glm/glm-5.1 (cheap, resets daily โ $0.5/1M)
- 4. minimax/MiniMax-M2.5 (cheapest paid โ $0.3/1M)
- 5. kr/claude-sonnet-4.5 (free unlimited โ never fails)
-
-Compression: lite โ saves 15% tokens passively, zero risk
-Result: 5 layers of fallback = zero downtime
-Monthly cost: $20-200 (subscriptions) + $5-10 (backup)
-```
-
-### Case 4: "I'm in a blocked region (Russia, China, Iran...)"
-
-**Problem:** AI providers block my country, VPNs are slow.
-
-```
-Combo: "unblocked-ai"
- 1. kr/claude-sonnet-4.5 (free via Kiro + proxy)
- 2. pol/deepseek-r1 (Pollinations โ no geo-block)
- 3. groq/llama-3.3-70b (Groq + proxy)
-
-Proxy: Global proxy set in Settings โ or per-provider proxy override
-Result: Access ALL providers from ANY country
-Monthly cost: $0 (free providers) + $0 (1proxy free marketplace)
-```
-
-### Case 5: "I want maximum token savings"
-
-**Problem:** Token costs are eating my budget, need to squeeze every token.
-
-```
-Combo: "ultra-saver"
- 1. cc/claude-opus-4-7 (subscription โ best quality)
- 2. glm/glm-5.1 (cheap backup)
-
-Compression: ultra โ saves 75% tokens
-Result: 10K token prompt โ 2.5K tokens sent
-Montly savings: ~$150-300/month in token costs for heavy users
-```
-
-## ๐งช Evaluations (Evals)
-
-OmniRoute includes a built-in evaluation framework to test LLM response quality against a golden set. Access it via **Analytics โ Evals** in the dashboard.
-
-### Built-in Golden Set
-
-The pre-loaded "OmniRoute Golden Set" contains test cases for:
-
-- Greetings, math, geography, code generation
-- JSON format compliance, translation, markdown generation
-- Safety refusal (harmful content), counting, boolean logic
-
-### Evaluation Strategies
-
-| Strategy | Description | Example |
-| ---------- | ------------------------------------------------ | -------------------------------- |
-| `exact` | Output must match exactly | `"4"` |
-| `contains` | Output must contain substring (case-insensitive) | `"Paris"` |
-| `regex` | Output must match regex pattern | `"1.*2.*3"` |
-| `custom` | Custom JS function returns true/false | `(output) => output.length > 10` |
-
----
-
-## ๐ Setup Guide
-
-### Connect Your Coding Tool
-
-Point any OpenAI-compatible tool to OmniRoute:
-
-```txt
-Base URL: http://localhost:20128/v1
-API Key: [from Dashboard โ Endpoints]
-```
-
-| Tool | Config Location |
-| --------------- | ----------------------------------------------------------------------------------------- |
-| **Claude Code** | `claude mcp add-server omniroute --type http --url http://localhost:20128/api/mcp/stream` |
-| **Codex CLI** | `OPENAI_BASE_URL=http://localhost:20128/v1 OPENAI_API_KEY=your-key codex` |
-| **Cursor** | Settings โ Models โ Add Model โ Override Base URL |
-| **Cline** | Extension settings โ Custom API Base URL |
-| **OpenClaw** | `OPENAI_BASE_URL=http://localhost:20128/v1 openclaw` |
-| **Gemini CLI** | Uses native OAuth via OmniRoute โ connect in Providers |
-
-### Protocols (MCP + A2A)
-
-```bash
-# MCP (stdio transport)
-omniroute --mcp
-
-# A2A (JSON-RPC 2.0)
-curl http://localhost:20128/.well-known/agent.json
-```
-
-### Key Environment Variables
-
-| Variable | Default | Purpose |
-| -------------------- | -------------- | ----------------------------------------- |
-| `PORT` | `20128` | API and dashboard port |
-| `DASHBOARD_PORT` | โ | Separate dashboard port (split-port mode) |
-| `REQUIRE_API_KEY` | `false` | Require API key for all requests |
-| `DATA_DIR` | `~/.omniroute` | Database and config storage |
-| `REQUEST_TIMEOUT_MS` | `600000` | Upstream response timeout |
-
-
-๐ Full Setup Guide โ All CLI tools, protocols, and environment variables
-
-๐ **Complete documentation:**
-
-- [User Guide](docs/USER_GUIDE.md) โ Providers, combos, CLI integration
-- [API Reference](docs/API_REFERENCE.md) โ All endpoints with examples
-- [MCP Server](open-sse/mcp-server/README.md) โ 37 tools, IDE configs
-- [A2A Server](src/lib/a2a/README.md) โ JSON-RPC, skills, streaming
-- [Environment Config](docs/ENVIRONMENT.md) โ Complete `.env` reference
-- [VM Deployment](docs/VM_DEPLOYMENT_GUIDE.md) โ VM + nginx + Cloudflare
-
-
-
----
-
-## โ Frequently Asked Questions
-
-
-๐ Why does my dashboard show high costs if I'm using free models?
-
-The dashboard tracks your token usage and displays **estimated costs** as if you were using paid APIs directly. This is **not actual billing** โ it's a reference to show how much you're saving.
-
-**Example:**
-
-- **Dashboard shows:** "$290 total cost"
-- **Reality:** You're using Kiro + Qoder (FREE unlimited)
-- **Your actual cost:** **$0.00**
-- **What $290 means:** Amount you **saved** by using free models instead of paid APIs!
-
-The cost display is a "savings tracker" to help you understand your usage patterns and optimization opportunities.
-
-
-
-
-๐ณ Will I be charged by OmniRoute?
-
-**No.** OmniRoute is free, open-source software that runs on your own computer. It never charges you anything.
-
-**You only pay:**
-
-- โ
**Subscription providers** (Claude Code $20/mo, Codex $20-200/mo) โ Pay them directly on their websites
-- โ
**API key providers** (DeepSeek, xAI, etc.) โ Pay them directly, OmniRoute just routes your requests
-- โ **OmniRoute itself** โ **Never charges anything, ever**
-
-OmniRoute is a local proxy/router. It doesn't have your credit card, can't send invoices, and has no billing system. It's completely free software.
-
-
-
-
-๐ Are FREE providers really unlimited?
-
-**Yes!** The current FREE providers are genuinely free with **no hidden charges**:
-
-- **Kiro AI**: Free unlimited Claude Sonnet/Haiku via AWS Builder ID / Google / GitHub OAuth
-- **Qoder**: Free unlimited kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 via PAT token
-- **Pollinations AI**: No API key needed โ GPT-5, Claude, DeepSeek, Llama 4
-- **LongCat Flash-Lite**: 50M tokens/day โ largest free quota available
-- **Cloudflare Workers AI**: 10K Neurons/day โ 50+ models at the edge
-
-OmniRoute just routes your requests to them โ there's no "catch" or future billing.
-
-
-
-
-๐ฐ How do I minimize my actual AI costs?
-
-**Free-First Strategy:**
-
-1. **Start with 100% free combo:**
-
- ```
- 1. kr/claude-sonnet-4.5 (Kiro โ unlimited free)
- 2. if/kimi-k2-thinking (Qoder โ unlimited free)
- 3. pol/gpt-5 (Pollinations โ no key needed)
- ```
-
- **Cost: $0/month**
-
-2. **Enable Prompt Compression** โ even `lite` mode saves ~15% passively
-
-3. **Add cheap backup** only if you need it:
-
- ```
- 4. glm/glm-5.1 ($0.5/1M tokens)
- ```
-
- **Additional cost: Only pay for what you actually use**
-
-4. **Use subscription providers last** โ only if you already have them. OmniRoute helps maximize their value through quota tracking.
-
-**Result:** Most users can operate at **$0/month** using only free tiers!
-
-
-
-
-๐๏ธ Will compression affect response quality?
-
-**No.** Compression only affects the **input** (your prompt), not the model's response. Each mode has been designed to preserve technical accuracy:
-
-- **Lite** (~15%): Only whitespace/formatting โ zero semantic change
-- **Standard** (~30%): Removes filler words ("please", "I think", "basically") โ same meaning
-- **Aggressive** (~50%): Summarizes old messages + compresses tool outputs โ core context preserved
-- **Ultra** (~75%): Heuristic pruning โ use only when token budget is critical
-
-Code blocks, URLs, JSON, and structured data are **always protected** from compression via the preservation engine.
-
-
-
-
-๐ Does OmniRoute work in countries where AI is blocked?
-
-**Yes!** OmniRoute has a 3-level proxy system:
-
-1. **Global proxy** โ all requests go through your proxy
-2. **Per-provider proxy** โ different proxy per provider
-3. **Per-API-key proxy** โ different proxy per key
-
-Plus the **1proxy free marketplace** for community-shared proxies. Users in Russia, China, Iran, and other restricted regions can access all 160+ providers through OmniRoute's proxy infrastructure.
-
-See the [Proxy Guide](docs/PROXY_GUIDE.md) for setup instructions.
-
-
-
----
-
-## ๐ Troubleshooting
-
-| Problem | Quick Fix |
-| --------------------------------------------- | ------------------------------------------------------------------------------------ |
-| **"Language model did not provide messages"** | Provider quota exhausted โ check quota tracker, use combo fallback |
-| **Rate limiting (429)** | Add fallback combo: `cc/claude โ glm/glm-4.7 โ if/kimi-k2-thinking` |
-| **OAuth token expired** | Auto-refreshed by OmniRoute. If stuck: delete + re-auth in Providers |
-| **`unsupported_country_region_territory`** | Configure proxy in Settings โ Proxy (see [Proxy Guide](docs/PROXY_GUIDE.md)) |
-| **Docker SQLite locks** | Use `--stop-timeout 40` for clean WAL checkpoint on shutdown |
-| **Node.js runtime errors** | Use Node.js `>=20.20.2 <21`, `>=22.22.2 <23`, or `>=24.0.0 <25` (24 LTS recommended) |
-| **`system-info` for bug reports** | Run `npm run system-info` and attach `system-info.txt` to your issue |
-
-๐ **Full troubleshooting guide:** [`docs/TROUBLESHOOTING.md`](docs/TROUBLESHOOTING.md)
-
## ๐ ๏ธ Tech Stack
diff --git a/README_REDESIGN_MANUAL.md b/README_REDESIGN_MANUAL.md
new file mode 100644
index 0000000000..53a506d549
--- /dev/null
+++ b/README_REDESIGN_MANUAL.md
@@ -0,0 +1,177 @@
+# ๐ Manual de Repaginaรงรฃo do README โ OmniRoute
+
+> Documento de trabalho. Analisa o README atual, compara com o 9router e propรตe **3 layouts**
+> para a parte **acima** de `## ๐ ๏ธ Tech Stack`. Nรฃo รฉ parte da documentaรงรฃo publicada โ pode
+> ser apagado depois que o redesign for aplicado.
+
+---
+
+## 0. Escopo do redesign
+
+| Faixa | Linhas | Aรงรฃo |
+| --------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------------------- |
+| **Topo โ fim do Troubleshooting** | **1โ1468** | ๐ฏ **Redesenhar** (27 seรงรตes) |
+| `## ๐ ๏ธ Tech Stack` em diante | **1469โ1678** | ๐ **Manter intacto** (Tech Stack, Documentation, Contributors, Star History, StarMapper, Acknowledgments, License) |
+
+Estado atual do alvo: **27 seรงรตes / ~1.468 linhas**. O 9router cobre o mesmo terreno em **~7 seรงรตes enxutas**.
+
+---
+
+## 1. Diagnรณstico โ o que estรก quebrado
+
+### 1.1 Hero sobrecarregado (linhas 1โ60)
+
+- **Dois blocos de badges** separados (linhas 9โ13 e 41โ56): npm, license, node, stars, Trendshift, depois npm version/week/month/year, Docker, electron, license (de novo), contributions, streak, website, whatsapp.
+- **Trendshift duplicado**: linha 13 **e** linha 25.
+- **Promo AgentRouter gigante** (badge `for-the-badge` + subtรญtulo "Limited offer") logo no topo โ parece anรบncio antes de explicar o produto.
+- **Mural de 40+ idiomas** (linha 35): uma parede de bandeiras a 35 linhas do topo.
+- **Parรกgrafo-descriรงรฃo denso** (linha 7) com 6 nรบmeros em negrito numa sรณ frase โ nada "brilha", tudo compete pela atenรงรฃo.
+- **Resultado**: a primeira tela รฉ uma muralha de badges/links/promo. Falta um soco visual de "o que รฉ + por que me importo".
+
+### 1.2 Nรบmeros inconsistentes (corrรณi credibilidade)
+
+| Mรฉtrica | Onde diverge | Valores conflitantes |
+| -------------------------- | ------------------------------------------------------------------------------------------------------ | -------------------- |
+| **Provedores totais** | hero (l.7) `160+` ยท Why (l.214) `207+` ยท tabela vs alternativas (l.260) `207+` ยท header (l.406) `160+` | **160+ vs 207+** |
+| **MCP tools / scopes** | hero/Also-solves `37 tools` ยท tabela 31 problemas (l.708) `29 tools, 10 scopes` | **37/13 vs 29/10** |
+| **Estratรฉgias de routing** | hero (l.7) `13` ยท tabela Why (l.261) `14` | **13 vs 14** |
+| **Provedores free** | "11 Free Providers" (l.689, l.1024) ยท tabela visual (l.426) mostra `8` ยท About `50+` | **8 vs 11 vs 50+** |
+
+### 1.3 Conteรบdo duplicado / redundante
+
+- **DOIS diagramas de fallback** que atรฉ divergem: "Why OmniRoute?" mostra **3 tiers** (l.221โ246); "How It Works" mostra **4 tiers** (l.531โ555).
+- **TRรS listas de benefรญcios sobrepostas**: "Why this matters / Also solves" (l.248โ289) + "What OmniRoute Solves" (l.680, tabela + 31 problemas) + "Key Features" (l.1091). Dizem quase a mesma coisa trรชs vezes.
+- **Compressรฃo mencionada 5ร**: hero, "Also solves", diagrama "How It Works", seรงรฃo dedicada, e "Key Features".
+
+### 1.4 Ordem de seรงรตes problemรกtica
+
+- `## ๐ง Support` na **linha 292** โ cedo demais (antes de mostrar CLI tools, provedores ou features).
+- `## โก Quick Start` sรณ na **linha 746** โ o dev rola ~745 linhas de marketing antes de descobrir como instalar. Para uma ferramenta de dev, isso รฉ invertido.
+- `## ๐ค AI Agent Skills` (l.516) espremida entre Providers e How It Works โ quebra o fluxo.
+
+### 1.5 Muros de texto
+
+- **Docker** (l.807โ896): parรกgrafos longos de notas (Cloudflare tunnel, WAL, stop-timeoutโฆ) que deveriam estar colapsados.
+- **Compressรฃo** (l.559โ677): excelente conteรบdo, mas longo e tรฉcnico demais para a primeira dobra de um README.
+- **FAQ** (l.1345โ1454, ~110 linhas), **Use Cases** (l.1159โ1265, ~106 linhas), **Pricing** (l.980โ1021): grandes e abertos.
+
+### 1.6 Tamanho
+
+~1.468 linhas antes do Tech Stack. Meta realista de reduรงรฃo: **โ60% a โ75%** de conteรบdo visรญvel (o resto vai para `` ou links de docs, sem perder nada).
+
+---
+
+## 2. Correรงรตes transversais (aplicam-se a QUALQUER layout escolhido)
+
+### 2.1 Fixar nรบmeros canรดnicos (escolher um valor e usar em todo lugar)
+
+| Mรฉtrica | Valor canรดnico recomendado | Observaรงรฃo |
+| ---------------------- | -------------------------------------------------------------- | --------------------------------------------------------------- |
+| Provedores totais | **160+** | Combina com a About nova e o header da seรงรฃo. Aposentar `207+`. |
+| Provedores free | **11 grรกtis "forever" + 50+ com free-tier** | Alinhar a tabela visual (hoje 8) e os headers. |
+| MCP | **37 tools ยท 3 transports ยท 13 scopes** | Corrigir `29 tools / 10 scopes` da tabela dos 31 problemas. |
+| Estratรฉgias de routing | **14** | Corrigir o `13` do hero. |
+| Compressรฃo | **15โ95% (RTK+Caveman stacked), ~89% mรฉdio** | Jรก consistente โ manter. |
+| Provas sociais | **4.690+ testes ยท 517 arquivos ยท 40+ idiomas ยท 16+ CLI tools** | Manter como credibilidade. |
+
+### 2.2 Deduplicar
+
+- **Um รบnico** diagrama de fallback (unificar os dois; usar 4 tiers: Subscription โ API Key โ Cheap โ Free).
+- **Uma รบnica** tabela consolidada de capacidades (fundir "What Solves" + "Key Features" + "Also solves").
+- **Um รบnico** Trendshift no hero.
+
+### 2.3 Reordenar
+
+- `Quick Start` sobe para logo apรณs o pitch.
+- `Support` desce para o rodapรฉ (logo antes do Tech Stack).
+- Niche (Agent Skills, Transcription, Evals) viram `` ou blocos curtos no fim.
+
+### 2.4 Persuasรฃo honesta (promessas que cumprimos)
+
+Manter o tom vendedor, mas toda afirmaรงรฃo ancorada: "never hit limits" โ justificado pelo **auto-fallback**; "save up to 95%" โ รฉ o teto do **stacked**, citar a mรฉdia ~89%; "unlimited FREE" โ รฉ o que Kiro/Qoder/Qwen oferecem hoje (datado). Evitar superlativos sem lastro.
+
+---
+
+## 3. Princรญpios de design
+
+**Do 9router (o que copiar):** hero curtรญssimo (logo + 1 tagline + 1 linha "conecte X โ Y" + 5 badges + nav); bloco `โ problema / โ
soluรงรฃo` escaneรกvel; **um** diagrama; Quick Start em 3 passos logo no inรญcio.
+
+**Nossa vantagem (o que destacar, e o 9router nรฃo tem):** 160+ provedores (vs 40+) ยท RTK **+ Caveman stacked** 15โ95% (vs RTK 20โ40%) ยท MCP (37 tools) ยท A2A ยท Memory ยท Skills ยท Guardrails ยท Evals ยท 14 estratรฉgias ยท multi-plataforma (Web/Desktop/Termux/PWA) ยท 40+ idiomas ยท TLS stealth ยท cloud agents.
+
+---
+
+## 4. Os 3 layouts propostos
+
+### ๐
ฐ๏ธ Layout A โ Minimalista / Dev-first
+
+**Filosofia:** chegar a valor + instalaรงรฃo no menor nรบmero de linhas. Colapsar agressivamente. Meta **~350โ450 linhas**.
+
+```
+1. Hero slim (img + tรญtulo + 1 tagline + 1 linha "conecte XโY" + 5 badges + nav)
+ โโ Badges extras, 40+ idiomas, sponsor AgentRouter
+2. ๐ค Why โ โproblema / โ
soluรงรฃo (estilo 9router) + 5 linhas do comparativo
+ โโ tabela comparativa completa
+3. ๐ How It Works โ UM diagrama de fallback unificado
+4. โก Quick Start โ 3 passos โโ Docker/source/Arch/pnpm
+5. ๐ค Tools + Providers (grids enxutos) โโ "+130 provedores"
+6. ๐ก Capabilities โ UMA tabela consolidada โโ 31 problemas
+7. ๐๏ธ Compressรฃo โ tabela + before/after โโ arquitetura/math
+8. ๐ Resto (Pricing, Use Cases, Proxy, Platforms, FAQโฆ) โ ou links docs
+```
+
+**Prรณs:** lรช em 1 minuto, dev instala rรกpido. **Contras:** marketing/visual mais discreto.
+
+---
+
+### ๐
ฑ๏ธ Layout B โ Vitrine visual / Marketing-first
+
+**Filosofia:** encantar e vender a promessa, depois provar. Landing-page. Meta **~600โ700 linhas**.
+
+```
+1. Hero visual (screenshot grande + headline-promessa + faixa de 3 stats
+ "160+ provedores ยท 15โ95% economia ยท $0 pra comeรงar" + CTAs)
+2. ๐ฅ A Promessa โ grid 2ร3 de cards (Never hit limits / Save 95% / $0 / Todo tool / 1 endpoint / Self-host)
+3. ๐ค Why โ โ/โ
+4. ๐ vs Alternativas โ tabela comparativa VISรVEL (รฉ persuasiva)
+5. ๐ฌ Prova social โ vรญdeos + Trendshift + estrelas no alto
+6. ๐ค Funciona com suas ferramentas โ grid com โญ stars
+7. ๐ 160+ provedores โ grid visual, free destacado
+8. ๐๏ธ Economize 15โ95% โ nรบmeros headline + before/after
+9. โก Quick Start โ 3 passos
+10. ๐ Resto โ
+```
+
+**Prรณs:** impressiona, รณtimo pra conversรฃo/compartilhamento. **Contras:** mais longo; dev tรฉcnico pode achar "vendedor demais".
+
+---
+
+### ๐
ฒ Layout C โ Equilibrado / Hรญbrido โญ (recomendado)
+
+**Filosofia:** hero limpo-mas-confiante, caminho rรกpido ao valor, prova persuasiva, e `` disciplinado. Narrativa: Gancho โ Por quรช โ Prova โ Instale โ Capacidades โ Aprofundamentos (colapsados). Meta **~500โ600 linhas**.
+
+```
+1. Hero limpo (img + tรญtulo + 1 tagline + 1 linha "conecte XโY" + faixa de stats
+ + 5 badges + Trendshift รบnico + nav)
+ โโ Badges secundรกrios + 40+ idiomas + sponsor
+2. ๐ค Why OmniRoute? โ โ/โ
+ UM diagrama de 4 tiers
+3. ๐ What sets us apart โ top-8 do comparativo โโ tabela completa
+4. โก Quick Start โ 3 passos + callout Free Stack โโ outros mรฉtodos
+5. ๐ค Works with 16+ tools โ grid coding-agents + CLI fundidos
+6. ๐ 160+ provedores (50+ free) โ OAuth + Free + top API โโ "+130 mais"
+7. โจ Capabilities โ UMA tabela consolidada โโ 31 problemas
+8. ๐๏ธ Compressรฃo โ save 15โ95% (tabela + before/after) โโ arquitetura/math
+9. ๐ฆ Platforms & Deploy โ Multi-plataforma + Docker condensados โโ notas
+10. ๐ More โ Pricing, Use Cases, Proxy, Evals, Free Models, Transcription, FAQ,
+ Troubleshooting, Skills, Vรญdeos โ grupos colapsรกveis / links docs
+11. ๐ง Support (movido pro fim)
+```
+
+**Prรณs:** limpo + persuasivo + completo, ~โ
do tamanho, nada importante perdido (sรณ colapsado). **Contras:** exige mais cuidado de execuรงรฃo que o A.
+
+---
+
+## 5. Recomendaรงรฃo e prรณximos passos
+
+- **Recomendado: Layout C** โ equilibra a clareza do 9router com os nossos diferenciais, sem perder conteรบdo (sรณ colapsa).
+- Independente da escolha, aplicar **todas** as correรงรตes transversais da seรงรฃo 2 (nรบmeros canรดnicos, deduplicaรงรฃo, reordenaรงรฃo).
+- Prรณximo passo: vocรช escolhe o layout; eu reescrevo as linhas 1โ1468 do `README.md`, preservando 1469+ intactas, e depois propago para os READMEs traduzidos (`docs/i18n/*`) se desejar.
diff --git a/docs/screenshots/MainOmniRoute.png b/docs/screenshots/MainOmniRoute.png
index 9609587000..c9879cbd4b 100644
Binary files a/docs/screenshots/MainOmniRoute.png and b/docs/screenshots/MainOmniRoute.png differ
diff --git a/src/lib/cloudAgent/baseAgent.ts b/src/lib/cloudAgent/baseAgent.ts
index 26c2defb59..9f5689b90b 100644
--- a/src/lib/cloudAgent/baseAgent.ts
+++ b/src/lib/cloudAgent/baseAgent.ts
@@ -86,10 +86,10 @@ export abstract class CloudAgentBase {
}
protected generateTaskId(): string {
- return `task_${Date.now()}_${crypto.randomUUID().replace(/-/g, "").substring(0, 9)}`;
+ return `task_${Date.now()}_${Math.random().toString(36).substring(2, 11)}`;
}
protected generateActivityId(): string {
- return `act_${Date.now()}_${crypto.randomUUID().replace(/-/g, "").substring(0, 9)}`;
+ return `act_${Date.now()}_${Math.random().toString(36).substring(2, 11)}`;
}
}
diff --git a/tests/unit/antigravity-discovery-bootstrap.test.ts b/tests/unit/antigravity-discovery-bootstrap.test.ts
index 9f211e65e9..3d0aa82f98 100644
--- a/tests/unit/antigravity-discovery-bootstrap.test.ts
+++ b/tests/unit/antigravity-discovery-bootstrap.test.ts
@@ -152,7 +152,7 @@ describe("ensureAntigravityProjectAssigned", () => {
const mockFetch = async (url: string, _init?: RequestInit): Promise => {
hitUrls.push(url);
- if (new URL(url).hostname === "daily-cloudcode-pa.googleapis.com") {
+ if (url.includes("daily-cloudcode-pa.googleapis.com")) {
// First URL fails
return new Response("not found", { status: 404 });
}