diff --git a/README.md b/README.md
index f1741d865f..74e2f571bc 100644
--- a/README.md
+++ b/README.md
@@ -197,19 +197,27 @@ _Connect any AI-powered IDE or CLI tool through OmniRoute โ free API gateway f
## ๐ค Why OmniRoute?
-**Stop wasting money and hitting limits:**
+**Stop wasting money, tokens and hitting limits:**
--
Subscription quota expires unused every month
--
Rate limits stop you mid-coding
--
Expensive APIs ($20-50/month per provider)
--
Manual switching between providers
+โ Subscription quota expires unused every month
+โ Rate limits stop you mid-coding
+โ Tool outputs (`git diff`, `grep`, `ls`...) burn tokens fast
+โ Expensive APIs ($20-50/month per provider)
+โ Manual switching between providers
+โ Each provider has a different API format
+โ AI providers blocked in your country
-**OmniRoute solves this:**
+**OmniRoute solves all of this:**
-- โ
**Maximize subscriptions** - Track quota, use every bit before reset
-- โ
**Auto fallback** - Subscription โ API Key โ Cheap โ Free, zero downtime
-- โ
**Multi-account** - Round-robin between accounts per provider
-- โ
**Universal** - Works with Claude Code, Codex, Gemini CLI, Cursor, Cline, OpenClaw, any CLI tool
+โ
**Prompt Compression** โ auto-compress prompts & tool outputs, save 15-75% tokens per request
+โ
**Maximize subscriptions** โ track quota, use every bit before reset
+โ
**Auto fallback** โ Subscription โ API Key โ Cheap โ Free, zero downtime
+โ
**Multi-account** โ round-robin between accounts per provider
+โ
**Format translation** โ OpenAI โ Claude โ Gemini โ Responses API, any tool works
+โ
**3-level proxy** โ bypass geo-blocks with global, per-provider, and per-key proxies
+โ
**10 multi-modal APIs** โ chat, images, video, music, audio, search in one endpoint
+โ
**MCP + A2A** โ 29 MCP tools + agent-to-agent protocol, production-ready
+โ
**Universal** โ works with Claude Code, Codex, Gemini CLI, Cursor, Cline, OpenClaw, any CLI tool
---
@@ -236,6 +244,151 @@ This generates a `system-info.txt` with your Node.js version, OmniRoute version,
---
+## ๐ ๏ธ Supported CLI Tools
+
+OmniRoute works seamlessly with **16+ AI coding tools** โ one config, all tools:
+
+
+
+ Claude Code Anthropic |
+ Codex CLI OpenAI |
+ Gemini CLI Google |
+ Cursor IDE |
+ OpenClaw CLI |
+ Antigravity VS Code |
+
+
+ Cline Extension |
+ Continue Extension |
+ Kilo Code Extension |
+ Kiro AWS IDE |
+ OpenCode CLI |
+ Droid CLI |
+
+
+ AMP CLI |
+ Copilot GitHub |
+ Windsurf IDE |
+ Hermes CLI |
+ Qwen CLI Alibaba |
+ Custom Any tool |
+
+
+
+๐ Full setup for each tool: [`docs/CLI-TOOLS.md`](docs/CLI-TOOLS.md)
+
+---
+
+## ๐ Supported Providers โ 160+
+
+### ๐ OAuth Providers
+
+
+
+ Claude Code Anthropic OAuth |
+ Antigravity Google OAuth |
+ Codex OpenAI OAuth |
+ GitHub Copilot GitHub OAuth |
+ Cursor Cursor OAuth |
+
+
+ Kimi Coding Moonshot OAuth |
+ Kilo Code Kilo OAuth |
+ Cline Cline OAuth |
+ |
+
+
+
+### ๐ Free Providers (No Cost)
+
+
+
+ ๐ข Kiro AI Claude Sonnet/Haiku Unlimited FREE |
+ ๐ข Qoder AI Kimi-K2, DeepSeek-R1 Unlimited FREE |
+ ๐ข Pollinations GPT-5, Claude, Llama 4 No API key needed |
+ ๐ข Qwen Code Qwen3 Coder Plus Unlimited FREE |
+
+
+ ๐ข LongCat AI Flash-Lite 50M tokens/day |
+ ๐ข Cloudflare AI 50+ models 10K neurons/day |
+ ๐ข Puter AI GPT-4.1, Claude Rate-limited free |
+ ๐ข NVIDIA NIM Llama, Mistral 1K req/day free |
+
+
+
+### ๐ API Key Providers (120+)
+
+
+
+ | OpenAI |
+ Anthropic |
+ Gemini |
+ DeepSeek |
+ Groq |
+ xAI (Grok) |
+
+
+ | Mistral |
+ OpenRouter |
+ GLM |
+ Kimi |
+ MiniMax |
+ Fireworks |
+
+
+ | Together AI |
+ Cerebras |
+ Cohere |
+ NVIDIA |
+ Perplexity |
+ SiliconFlow |
+
+
+ | Nebius |
+ HuggingFace |
+ DeepInfra |
+ SambaNova |
+ Vertex AI |
+ Azure OpenAI |
+
+
+ | AWS Bedrock |
+ Snowflake |
+ Databricks |
+ Venice.ai |
+ AI21 Labs |
+ Meta Llama |
+
+
+
+
+...and 90+ more providers
+
+Alibaba ยท Amazon Q ยท AssemblyAI ยท Baidu Qianfan ยท Baseten ยท Black Forest Labs ยท Blackbox ยท Brave Search ยท Bytez ยท CablyAI ยท Cartesia ยท ChatGPT Web ยท Chutes.ai ยท Clarifai ยท Codestral ยท CrofAI ยท DataRobot ยท Deepgram ยท ElevenLabs ยท Empower ยท Exa Search ยท Fal.ai ยท Featherless AI ยท FenayAI ยท FriendliAI ยท Galadriel ยท GigaChat ยท GitLab Duo ยท GLHF Chat ยท GoAPI ยท Heroku AI ยท Hyperbolic ยท IBM watsonx ยท Inference.net ยท Inworld ยท Jina AI ยท Kilo Gateway ยท Lambda AI ยท LaoZhang ยท Linkup Search ยท LlamaGate ยท Maritalk ยท Modal ยท Moonshot AI ยท Morph ยท Muse Spark ยท NanoBanana ยท NanoGPT ยท NLP Cloud ยท Nous Research ยท Novita AI ยท nScale ยท OCI ยท Ollama Cloud ยท OVHcloud ยท PiAPI ยท PlayHT ยท Poe ยท Predibase ยท PublicAI ยท Qwen Code ยท Recraft ยท Reka ยท Runway ยท SAP ยท Scaleway ยท SearchAPI ยท SearXNG ยท Serper ยท Stability AI ยท Synthetic ยท Tavily ยท TheB.AI ยท Topaz ยท Upstage ยท v0 (Vercel) ยท Vercel AI Gateway ยท Volcengine ยท Voyage AI ยท W&B Inference ยท Xiaomi MiMo ยท You.com ยท Z.AI ยท + OpenAI/Anthropic-compatible custom endpoints
+
+
+
+### ๐ Self-Hosted
+
+
+
+ | LM Studio |
+ Ollama |
+ vLLM |
+ Llamafile |
+ Docker Model Runner |
+
+
+ | NVIDIA Triton |
+ XInference |
+ oobabooga |
+ ComfyUI |
+ SD WebUI |
+
+
+
+---
+
## ๐ How It Works
```
@@ -684,24 +837,7 @@ No proxy? Use the built-in **1proxy** integration for **hundreds of free, valida
> ๐ **New models added (Mar 2026):** Grok-4 Fast family at $0.20/$0.50/M (benchmarked at 1143ms โ 30% faster than Gemini 2.5 Flash), GLM-5 via Z.AI with 128K output, MiniMax M2.5 reasoning, DeepSeek V3.2 updated pricing, Kimi K2.5 via Moonshot direct API.
-**๐ก $0 Combo Stack โ The Complete Free Setup:**
-
-```
-# ๐ Ultimate Free Stack 2026 โ 11 Providers, $0 Forever
-Kiro (kr/) โ Claude Sonnet/Haiku UNLIMITED
-Qoder (if/) โ kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 UNLIMITED
-LongCat Lite (lc/) โ LongCat-Flash-Lite โ 50M tokens/day ๐ฅ
-Pollinations (pol/) โ GPT-5, Claude, DeepSeek, Llama 4 โ no key needed
-Qwen (qw/) โ qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next UNLIMITED
-Gemini (gemini/) โ Gemini 2.5 Flash โ 1,500 req/day free API key
-Cloudflare AI (cf/) โ Llama 70B, Gemma 3, Mistral โ 10K Neurons/day
-Scaleway (scw/) โ Qwen3 235B, Llama 70B โ 1M free tokens (EU)
-Groq (groq/) โ Llama/Gemma ultra-fast โ 14.4K req/day
-NVIDIA NIM (nvidia/) โ 70+ open models โ 40 RPM forever
-Cerebras (cerebras/) โ Llama/Qwen world-fastest โ 1M tok/day
-```
-
-**Zero cost. Never stops coding.** Configure this as one OmniRoute combo and all fallbacks happen automatically โ no manual switching ever.
+**๐ก See the full [$0 Free Stack (11 providers)](#-free-models--11-providers-0-forever) below.**
> ๐ก **Understanding Dashboard Costs:**
>
@@ -712,156 +848,45 @@ Cerebras (cerebras/) โ Llama/Qwen world-fastest โ 1M tok/day
---
-## ๐ Free Models โ What You Actually Get
+## ๐ Free Models โ 11 Providers, $0 Forever
-> All models below are **100% free with zero credit card required**. OmniRoute auto-routes between them when one quota runs out โ combine them all for an unbreakable $0 combo.
+> Combine all free providers into one unbreakable combo โ OmniRoute auto-routes between them when quota runs out.
-### ๐ต CLAUDE MODELS (via Kiro โ AWS Builder ID)
+| Provider | Prefix | Free Models | Quota |
+| ----------------- | ----------- | ------------------------------------------------------------- | ----------------- |
+| **Kiro** | `kr/` | Claude Sonnet 4.5, Haiku 4.5, Opus 4.6 | โพ๏ธ Unlimited |
+| **Qoder** | `if/` | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1, minimax-m2.1 | โพ๏ธ Unlimited |
+| **Qwen** | `qw/` | qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next | โพ๏ธ Unlimited |
+| **Pollinations** | `pol/` | GPT-5, Claude, Gemini, DeepSeek, Llama 4, Mistral | No key needed |
+| **LongCat** | `lc/` | LongCat-Flash-Lite | 50M tokens/day ๐ฅ |
+| **Gemini CLI** | `gc/` | gemini-3-flash, gemini-2.5-pro | 180K tok/mo |
+| **Cloudflare AI** | `cf/` | 50+ models (Llama, Gemma, Mistral, Whisper) | 10K Neurons/day |
+| **Groq** | `groq/` | Llama 3.3 70B, Qwen3 32B, Kimi K2 | 14.4K RPD |
+| **NVIDIA NIM** | `nvidia/` | 129 models (DeepSeek, Llama, GLM, Kimi) | ~40 RPM |
+| **Cerebras** | `cerebras/` | Qwen3 235B, GPT-OSS 120B, Llama 3.1 | 1M tok/day |
+| **Scaleway** | `scw/` | Qwen3 235B, Llama 70B, DeepSeek V3 | 1M tokens (EU) |
-| Model | Prefix | Limit | Rate Limit |
-| ------------------- | ------ | ------------- | --------------------- |
-| `claude-sonnet-4.5` | `kr/` | **Unlimited** | No reported daily cap |
-| `claude-haiku-4.5` | `kr/` | **Unlimited** | No reported daily cap |
-| `claude-opus-4.6` | `kr/` | **Unlimited** | Latest Opus via Kiro |
+
+๐ 25+ more free providers โ Groq, Cerebras, Mistral, GitHub Models, OpenRouter, and more
-### ๐ข QODER MODELS (Free PAT via qodercli)
+**Also free (API Key required):**
+Mistral (1B tok/month) ยท OpenRouter (35+ `:free` models) ยท GitHub Models (GPT-5, 45+ models) ยท
+Cohere (1K calls/month) ยท Z.AI/GLM (permanent free Flash models) ยท SiliconFlow (1K RPM, 50K TPM) ยท
+Kilo Code (~200 req/hr auto-router) ยท HuggingFace ($0.10/mo credits) ยท Ollama Cloud (400+ models) ยท
+LLM7.io (30+ models) ยท Kluster AI ยท IBM watsonx (300K tok/month) ยท OpenCode Zen ยท Vercel AI Gateway ($5/mo)
-| Model | Prefix | Limit | Rate Limit |
-| ------------------ | ------ | ------------- | --------------- |
-| `kimi-k2-thinking` | `if/` | **Unlimited** | No reported cap |
-| `qwen3-coder-plus` | `if/` | **Unlimited** | No reported cap |
-| `deepseek-r1` | `if/` | **Unlimited** | No reported cap |
-| `minimax-m2.1` | `if/` | **Unlimited** | No reported cap |
-| `kimi-k2` | `if/` | **Unlimited** | No reported cap |
+**Trial credits (one-time):**
+Baseten ($30) ยท NLP Cloud ($15) ยท AI21 ($10) ยท Upstage ($10) ยท SambaNova ($5) ยท Modal ($5/mo) ยท
+Fireworks ($1) ยท Nebius ($1) ยท Inference.net ($1 + $25 survey) ยท Hyperbolic ($1) ยท Novita ($0.50)
-> Recommended connection method: **Personal Access Token + `qodercli`**. Browser OAuth is
-> experimental and disabled by default unless `QODER_OAUTH_*` environment variables are configured.
+**China-based (free tiers):**
+ModelScope ยท Tencent Hunyuan ยท Volcengine ยท ChatAnywhere ยท InternAI ยท Bigmodel
-### ๐ก QWEN MODELS (Device Code Auth)
+**Combined capacity: ~31,000+ RPD ยท ~32B+ tokens/month ยท 500+ models ยท $0**
-| Model | Prefix | Limit | Rate Limit |
-| ------------------- | ------ | ------------- | ------------------- |
-| `qwen3-coder-plus` | `qw/` | **Unlimited** | No reported cap |
-| `qwen3-coder-flash` | `qw/` | **Unlimited** | No reported cap |
-| `qwen3-coder-next` | `qw/` | **Unlimited** | No reported cap |
-| `vision-model` | `qw/` | **Unlimited** | Multimodal (images) |
+
-### ๐ฃ GEMINI CLI (Google OAuth)
-
-| Model | Prefix | Limit | Rate Limit |
-| ------------------------ | ------ | --------------------------- | ------------- |
-| `gemini-3-flash-preview` | `gc/` | **180K tok/month** + 1K/day | Monthly reset |
-| `gemini-2.5-pro` | `gc/` | 180K/month (shared pool) | High quality |
-
-### โซ NVIDIA NIM (Free API Key โ build.nvidia.com)
-
-| Tier | Daily Limit | Rate Limit | Notes |
-| ---------- | ------------ | ----------- | ------------------------------------------------------ |
-| Free (Dev) | No token cap | **~40 RPM** | 70+ models; transitioning to pure rate limits mid-2025 |
-
-Popular free models: `moonshotai/kimi-k2.5` (Kimi K2.5), `z-ai/glm4.7` (GLM 4.7), `deepseek-ai/deepseek-v3.2` (DeepSeek V3.2), `nvidia/llama-3.3-70b-instruct`, `deepseek/deepseek-r1`
-
-### โช CEREBRAS (Free API Key โ inference.cerebras.ai)
-
-| Tier | Daily Limit | Rate Limit | Notes |
-| ---- | ----------------- | ---------------- | ------------------------------------------- |
-| Free | **1M tokens/day** | 60K TPM / 30 RPM | World's fastest LLM inference; resets daily |
-
-Available free: `llama-3.3-70b`, `llama-3.1-8b`, `deepseek-r1-distill-llama-70b`
-
-### ๐ด GROQ (Free API Key โ console.groq.com)
-
-| Tier | Daily Limit | Rate Limit | Notes |
-| ---- | ------------- | ---------------- | ----------------------------------------- |
-| Free | **14.4K RPD** | 30 RPM per model | No credit card; 429 on limit, not charged |
-
-Available free: `llama-3.3-70b-versatile`, `gemma2-9b-it`, `mixtral-8x7b`, `whisper-large-v3`
-
-### ๐ด LONGCAT AI (Free API Key โ longcat.chat) ๐
-
-| Model | Prefix | Daily Free Quota | Notes |
-| ----------------------------- | ------ | ----------------- | ----------------------- |
-| `LongCat-Flash-Lite` | `lc/` | **50M tokens** ๐ฅ | Largest free quota ever |
-| `LongCat-Flash-Chat` | `lc/` | 500K tokens | Multi-turn chat |
-| `LongCat-Flash-Thinking` | `lc/` | 500K tokens | Reasoning / CoT |
-| `LongCat-Flash-Thinking-2601` | `lc/` | 500K tokens | Jan 2026 version |
-| `LongCat-Flash-Omni-2603` | `lc/` | 500K tokens | Multimodal |
-
-> 100% free while in public beta. Sign up at [longcat.chat](https://longcat.chat) with email or phone. Resets daily 00:00 UTC.
-
-### ๐ข POLLINATIONS AI (No API Key Required) ๐
-
-| Model | Prefix | Rate Limit | Provider Behind |
-| ---------- | ------ | ---------- | ------------------ |
-| `openai` | `pol/` | 1 req/15s | GPT-5 |
-| `claude` | `pol/` | 1 req/15s | Anthropic Claude |
-| `gemini` | `pol/` | 1 req/15s | Google Gemini |
-| `deepseek` | `pol/` | 1 req/15s | DeepSeek V3 |
-| `llama` | `pol/` | 1 req/15s | Meta Llama 4 Scout |
-| `mistral` | `pol/` | 1 req/15s | Mistral AI |
-
-> โจ **Zero friction:** No signup, no API key. Add the Pollinations provider with an empty key field and it works immediately.
-
-### ๐ CLOUDFLARE WORKERS AI (Free API Key โ cloudflare.com) ๐
-
-| Tier | Daily Neurons | Equivalent Usage | Notes |
-| ---- | ------------- | --------------------------------------- | ----------------------- |
-| Free | **10,000** | ~150 LLM resp / 500s audio / 15K embeds | Global edge, 50+ models |
-
-Popular free models: `@cf/meta/llama-3.3-70b-instruct`, `@cf/google/gemma-3-12b-it`, `@cf/openai/whisper-large-v3-turbo` (free audio!), `@cf/qwen/qwen2.5-coder-15b-instruct`
-
-> Requires API Token + Account ID from [dash.cloudflare.com](https://dash.cloudflare.com). Store Account ID in provider settings.
-
-### ๐ฃ SCALEWAY AI (1M Free Tokens โ scaleway.com) ๐
-
-| Tier | Free Quota | Location | Notes |
-| ---- | ------------- | ------------ | ----------------------------------- |
-| Free | **1M tokens** | ๐ซ๐ท Paris, EU | No credit card needed within limits |
-
-Available free: `qwen3-235b-a22b-instruct-2507` (Qwen3 235B!), `llama-3.1-70b-instruct`, `mistral-small-3.2-24b-instruct-2506`, `deepseek-v3-0324`
-
-> EU/GDPR compliant. Get API key at [console.scaleway.com](https://console.scaleway.com).
-
-> **๐ก The Ultimate Free Stack (11 Providers, $0 Forever):**
->
-> ```
-> Kiro (kr/) โ Claude Sonnet/Haiku UNLIMITED
-> Qoder (if/) โ kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 UNLIMITED
-> LongCat Lite (lc/) โ LongCat-Flash-Lite โ 50M tokens/day ๐ฅ
-> Pollinations (pol/) โ GPT-5, Claude, DeepSeek, Llama 4 โ no key needed
-> Qwen (qw/) โ qwen3-coder models UNLIMITED
-> Gemini (gemini/) โ Gemini 2.5 Flash โ 1,500 req/day free
-> Cloudflare AI (cf/) โ 50+ models โ 10K Neurons/day
-> Scaleway (scw/) โ Qwen3 235B, Llama 70B โ 1M free tokens (EU)
-> Groq (groq/) โ Llama/Gemma โ 14.4K req/day ultra-fast
-> NVIDIA NIM (nvidia/) โ 70+ open models โ 40 RPM forever
-> Cerebras (cerebras/) โ Llama/Qwen world-fastest โ 1M tok/day
-> ```
-
----
-
-## ๐ Free API Provider Directory โ 25+ Providers, 500+ Models, $0
-
-> **We analyzed 6 community repositories** aggregating free LLM API providers and consolidated everything into one definitive reference. This is the most comprehensive free-tier directory available.
-
-| Provider | Best Free Model | RPM | RPD | Tokens | Speed |
-| ----------------- | ---------------- | ----- | ------------ | ------------- | --------- |
-| **Groq** | Llama 3.3 70B | 30 | 14,400 | 6K TPM | ๐ข Fast |
-| **Cerebras** | Qwen3 235B | 30 | 14,400 | 1M TPD | ๐ข Fast |
-| **Mistral AI** | Mistral Large 3 | 60 | Unlimited | 1B/month | ๐ก Medium |
-| **Google Gemini** | Gemini 2.5 Flash | 5โ15 | 20โ1,500 | 250K TPM | ๐ข Fast |
-| **NVIDIA NIM** | 129 models | 40 | โ | โ | ๐ก Medium |
-| **OpenRouter** | 35+ :free models | 20 | 50โ1,000 | โ | ๐ก Medium |
-| **GitHub Models** | GPT-4.1, GPT-5 | 10โ15 | 50โ150 | 8K/4K per req | ๐ก Medium |
-| **Cloudflare AI** | 50+ models | โ | 10K neurons | โ | ๐ก Medium |
-| **Pollinations** | Text+Image+Video | โ | Hourly reset | โ | ๐ก Medium |
-| **SiliconFlow** | Qwen3-8B | 1,000 | โ | 50K TPM | ๐ก Medium |
-
-**Combined free capacity across all providers: ~31,000+ RPD ยท ~32B+ tokens/month ยท 500+ models ยท $0 forever.**
-
-The full directory includes 25+ providers with detailed rate limits, base URLs, model tables, trial credit providers (Baseten $30, AI21 $10, SambaNova $5, etc.), China-specific platforms (ModelScope, Volcengine, Tencent Hunyuan), and step-by-step OmniRoute combo configuration.
-
-๐ **Complete free provider directory with all models, quotas, and integration guide:** [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md)
+๐ **Complete free provider directory:** [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md) โ 25+ providers, quotas, base URLs, model tables, and OmniRoute combo setup.
---
@@ -875,6 +900,8 @@ The full directory includes 25+ providers with detailed rate limits, base URLs,
| ๐ต **AssemblyAI** | **$50 free** (signup) | `universal-3-pro` โ chapters, sentiment, PII | No RPM limit on free credits |
| ๐ด **Groq** | **Free forever** | `whisper-large-v3` โ OpenAI Whisper | 30 RPM (rate limited) |
+---
+
**Suggested combo in `/dashboard/combos`:**
```
@@ -890,94 +917,67 @@ Then in `/dashboard/media` โ **Transcription** tab: upload any audio or video
## ๐ก Key Features
-OmniRoute v3.7+ is an operational platform, not just a relay proxy โ backed by **4,690+ automated tests** across 517 test files.
+> **4,690+ automated tests** across 517 test files. Not just a relay โ a full operational platform.
-| Category | Feature | Why It Matters |
-| -------------------- | -------------------------------------------------------------------------------- | ---------------------------------- |
-| ๐ง **Routing** | Smart 4-Tier Fallback (Subscription โ API โ Cheap โ Free) | Never stop coding, zero downtime |
-| | 13 Balancing Strategies + Custom Combos | Tailor routing to your exact needs |
-| | Task-Aware Smart Routing (coding/vision/analysis) | Right model for every task |
-| | Context Relay โ session handoffs during rotation | No lost context mid-conversation |
-| | Thinking Budget Controls (passthrough/auto/custom) | Control reasoning costs precisely |
-| ๐ **Translation** | OpenAI โ Claude โ Gemini โ Responses API | Works with ANY CLI tool |
-| | Auto Token Refresh (OAuth PKCE for 8 providers) | No manual re-login ever |
-| | Responses API โ full `/v1/responses` for Codex | First-class Codex compatibility |
-| ๐ต **Multi-Modal** | 10 APIs: chat, embed, images, video, music, TTS, STT, moderation, rerank, search | One endpoint for everything |
-| | Batch API โ asynchronous processing with Files API | Background bulk processing |
-| | OpenAPI 3.0 โ live auto-generated spec + Try-It UI | API-first development |
-| ๐ก๏ธ **Resilience** | Circuit Breakers + Connection Cooldown + Anti-Thundering Herd | Auto-recovery from failures |
-| | TLS Fingerprint Spoofing + CLI Fingerprint Matching | Stealth + anti-ban protection |
-| | Semantic + Signature Cache (two-tier) | Reduce costs + latency |
-| | Request Idempotency + Rate Limit Detection | No duplicate charges |
-| ๐ค **Protocols** | MCP Server โ 29 tools, 3 transports, 10 scopes | IDE/agent tool integration |
-| | A2A Server โ JSON-RPC 2.0, SSE streaming, task lifecycle | Agent-to-agent orchestration |
-| | ACP โ CLI agent discovery (14 agents + custom) | Universal agent onboarding |
-| ๐ **Observability** | Unified Logs (request/proxy/audit/console) + p50/p95/p99 | Full request telemetry |
-| | Health Dashboard โ uptime, breakers, cache, lockouts | Operational visibility |
-| | Cost Tracking + Budget Controls | Financial governance |
-| | Evaluation Framework โ golden set testing | Quality assurance |
-| โ๏ธ **Platform** | Desktop (Electron), Android (Termux), PWA | Run anywhere |
-| | Docker (AMD64 + ARM64) with Compose profiles | One-command deploy |
-| | Cloudflare / Tailscale / ngrok Tunnels | Instant public endpoint |
-| | 40+ languages with RTL support | Global accessibility |
-| ๐๏ธ **Compression** | 5-mode pipeline: off / lite / standard / aggressive / ultra | Save 15-75% tokens |
+| Feature | Why It Matters |
+| ---------------------------------------------------------------------------------------------------- | -------------------------------- |
+| ๐ง **Smart 4-Tier Fallback** โ Subscription โ API โ Cheap โ Free | Never stop coding, zero downtime |
+| ๐ **Format Translation** โ OpenAI โ Claude โ Gemini โ Responses API | Works with ANY CLI tool |
+| ๐๏ธ **Prompt Compression** โ 5-mode pipeline (off โ lite โ standard โ aggressive โ ultra) | Save 15-75% tokens automatically |
+| ๐ค **MCP Server** โ 29 tools, 3 transports (stdio/SSE/HTTP), 10 scopes | IDE/agent tool integration |
+| ๐ก๏ธ **Resilience Engine** โ circuit breakers, cooldowns, TLS spoofing, anti-thundering herd | Auto-recovery from any failure |
+| ๐ต **10 Multi-Modal APIs** โ chat, embed, images, video, music, TTS, STT, moderation, rerank, search | One endpoint for everything |
+| ๐ **3-Level Proxy** โ global, per-provider, per-key + 1proxy free marketplace | Access AI from any country |
+| ๐ **Full Observability** โ unified logs, p50/p95/p99 telemetry, cost tracking, budget controls | Know exactly what's happening |
-๐ What's New โ v3.6+ Highlights
+๐ Complete feature list โ 30+ capabilities
-- ๐ V1 WebSocket Bridge โ OpenAI-compatible WS at `/v1/ws`
-- ๐ Sync Tokens & Config Bundle โ versioned config sync with ETag
-- ๐ง GLM Thinking (glmt) โ 65K tokens, 24K thinking budget, Claude-compatible
-- ๐ข Hybrid Token Counting โ provider-side + estimation fallback
-- ๐ก๏ธ Safe Outbound Fetch โ SSRF protection on all provider calls
-- โณ Wait For Cooldown โ auto-retry after connection cooldowns
-- ๐ Runtime Env Validation โ Zod schemas at startup
-- ๐ Compliance Audit v2 โ pagination, auth events, SSRF logging
-- ๐ Webhooks โ event-driven with test firing and dashboard management
-- ๐๏ธ Vision Bridge โ image analysis guardrail before routing
-- โก Grok-4 Fast โ $0.20/$0.50/M, 30% faster than Gemini Flash
-- ๐ง GLM-5 via Z.AI โ 128K output, $0.5/1M
-- ๐ฎ MiniMax M2.5 โ reasoning + agentic at $0.3/1M
-- ๐ฏ toolCalling flag โ per-model tool capability in registry
-- ๐ Multilingual Intent Detection โ PT/ZH/ES/AR in AutoCombo
-- ๐ Benchmark-Driven Fallbacks โ real p95 latency feeds scoring
-- ๐ Request Deduplication โ content-hash dedup window
+**Routing & Intelligence**
-
+- 13 balancing strategies (priority, weighted, round-robin, P2C, cost-optimized, context-relay...)
+- Task-aware smart routing (coding/vision/analysis) ยท Context relay session handoffs
+- Thinking budget controls (passthrough/auto/custom) ยท Wildcard routing ยท System prompt injection
-
-๐ Feature Deep Dive โ Expanded Details
+**Translation & Compatibility**
-#### Smart fallback with practical cost control
+- Auto token refresh (OAuth PKCE for 8 providers) ยท Multi-account round-robin
+- Responses API โ full `/v1/responses` for Codex ยท Batch API with Files API
+- OpenAPI 3.0 live spec + Try-It UI
+
+**Protocols**
+
+- A2A Server โ JSON-RPC 2.0, SSE streaming, task lifecycle, skills
+- ACP โ CLI agent discovery (14 agents + custom)
+
+**Platform**
+
+- Desktop (Electron) ยท Android (Termux) ยท PWA ยท Docker (AMD64 + ARM64)
+- Cloudflare / Tailscale / ngrok tunnels ยท 40+ languages with RTL
+- Semantic + signature cache (two-tier) ยท Request idempotency + deduplication
+
+**Observability**
+
+- Health dashboard โ uptime, breakers, cache, lockouts
+- Evaluation framework โ golden set testing ยท Webhooks ยท Compliance audit
+
+**v3.6+ Highlights:**
+V1 WebSocket Bridge ยท Sync Tokens & Config Bundle ยท GLM Thinking (glmt) ยท Hybrid Token Counting ยท
+Safe Outbound Fetch ยท Wait For Cooldown ยท Runtime Env Validation ยท Vision Bridge ยท
+Grok-4 Fast ยท GLM-5 via Z.AI ยท MiniMax M2.5 ยท toolCalling flag ยท
+Multilingual Intent Detection ยท Benchmark-Driven Fallbacks ยท Request Deduplication
+
+**Architecture Examples:**
```txt
-Combo: "my-coding-stack"
- 1. cc/claude-opus-4-7
- 2. nvidia/llama-3.3-70b
- 3. glm/glm-4.7
+Combo: "my-coding-stack" Format Translation:
+ 1. cc/claude-opus-4-7 CLI โ OpenAI format
+ 2. nvidia/llama-3.3-70b OmniRoute โ translates
+ 3. glm/glm-4.7 Provider โ native format
4. if/kimi-k2-thinking
```
-When quota, rate, or health fails, OmniRoute automatically moves to the next candidate without manual switching.
-
-#### Prompt Compression โ Token Savings Breakdown
-
-```
-Without compression: 47K tokens sent to LLM
-With Lite: 40K tokens sent (15% saved โ safe, always-on)
-With Standard: 33K tokens sent (30% saved โ caveman-speak rules)
-With Aggressive: 24K tokens sent (50% saved โ aging + summarization)
-With Ultra: 12K tokens sent (75% saved โ heuristic pruning)
-```
-
-#### Format Translation โ Universal Compatibility
-
-- **OpenAI** โ **Claude** โ **Gemini** โ **Cursor** โ **Kiro** โ **Vertex** โ **Antigravity** โ **Ollama** โ **Responses**
-- Your CLI tool sends OpenAI format โ OmniRoute translates โ Provider receives native format
-
-> ๐ **[MCP Server README](open-sse/mcp-server/README.md)** โ Tool reference, IDE configs, and client examples
->
-> ๐ **[A2A Server README](src/lib/a2a/README.md)** โ Skills, JSON-RPC methods, streaming, and task lifecycle
+๐ [MCP Server README](open-sse/mcp-server/README.md) ยท [A2A Server README](src/lib/a2a/README.md) ยท [Resilience Guide](docs/RESILIENCE_GUIDE.md) ยท [Features Gallery](docs/FEATURES.md)
@@ -1349,31 +1349,6 @@ See the [Proxy Guide](docs/PROXY_GUIDE.md) for setup instructions.
---
-## ๐บ๏ธ Roadmap
-
-OmniRoute has **218+ features planned** across multiple development phases. Here are the key areas:
-
-| Category | Planned Features | Highlights |
-| ----------------------------- | ---------------- | ----------------------------------------------------------------------------------------------------- |
-| ๐ง **Routing & Intelligence** | 25+ | Lowest-latency routing, tag-based routing, quota preflight, quota-aware P2C, step-based combo routing |
-| ๐ **Security & Compliance** | 20+ | SSRF hardening, credential cloaking, rate-limit per endpoint, management key scoping |
-| ๐ **Observability** | 15+ | OpenTelemetry integration, real-time quota monitoring, combo target health, cost tracking per model |
-| ๐ **Provider Integrations** | 20+ | Dynamic model registry, connection cooldowns, multi-account Codex, Copilot quota parsing |
-| โก **Performance** | 15+ | Dual cache layer, prompt cache, response cache, streaming keepalive, batch API |
-| ๐ **Ecosystem** | 10+ | WebSocket API, config hot-reload, distributed config store, commercial mode |
-
-### ๐ Coming Soon
-
-- ๐ **OpenCode Integration** โ Native provider support for the OpenCode AI coding IDE
-- ๐ **TRAE Integration** โ Full support for the TRAE AI development framework
-- ๐ฐ **Lowest-Cost Strategy** โ Automatically select the cheapest available provider
-- ๐ **OpenTelemetry Integration** โ OTLP traces and metrics export for enterprise observability
-- ๐ **Config Hot-Reload** โ Apply settings changes without server restart
-
-> ๐ Full feature specifications available in [`docs/new-features/`](docs/new-features/) (217 detailed specs)
-
----
-
## โญ Top Contributors
> OmniRoute is shaped by a passionate open-source community. These individuals have made exceptional contributions that directly impact the quality, stability, and reach of the project. **Thank you.**