From 35d56358b89df2defb147a0c58e605648c5c8142 Mon Sep 17 00:00:00 2001 From: Antigravity Assistant Date: Fri, 1 May 2026 15:26:01 -0300 Subject: [PATCH] docs(readme): highlight supported tools and core platform benefits Refresh the README marketing and onboarding content to better explain why OmniRoute is useful in day-to-day AI coding workflows. Add a supported CLI tools section and expand the benefits overview to cover prompt compression, format translation, proxy routing, and multi-modal capabilities. --- README.md | 515 ++++++++++++++++++++++++++---------------------------- 1 file changed, 245 insertions(+), 270 deletions(-) diff --git a/README.md b/README.md index f1741d865f..74e2f571bc 100644 --- a/README.md +++ b/README.md @@ -197,19 +197,27 @@ _Connect any AI-powered IDE or CLI tool through OmniRoute โ€” free API gateway f ## ๐Ÿค” Why OmniRoute? -**Stop wasting money and hitting limits:** +**Stop wasting money, tokens and hitting limits:** -- Subscription quota expires unused every month -- Rate limits stop you mid-coding -- Expensive APIs ($20-50/month per provider) -- Manual switching between providers +โŒ Subscription quota expires unused every month +โŒ Rate limits stop you mid-coding +โŒ Tool outputs (`git diff`, `grep`, `ls`...) burn tokens fast +โŒ Expensive APIs ($20-50/month per provider) +โŒ Manual switching between providers +โŒ Each provider has a different API format +โŒ AI providers blocked in your country -**OmniRoute solves this:** +**OmniRoute solves all of this:** -- โœ… **Maximize subscriptions** - Track quota, use every bit before reset -- โœ… **Auto fallback** - Subscription โ†’ API Key โ†’ Cheap โ†’ Free, zero downtime -- โœ… **Multi-account** - Round-robin between accounts per provider -- โœ… **Universal** - Works with Claude Code, Codex, Gemini CLI, Cursor, Cline, OpenClaw, any CLI tool +โœ… **Prompt Compression** โ€” auto-compress prompts & tool outputs, save 15-75% tokens per request +โœ… **Maximize subscriptions** โ€” track quota, use every bit before reset +โœ… **Auto fallback** โ€” Subscription โ†’ API Key โ†’ Cheap โ†’ Free, zero downtime +โœ… **Multi-account** โ€” round-robin between accounts per provider +โœ… **Format translation** โ€” OpenAI โ†” Claude โ†” Gemini โ†” Responses API, any tool works +โœ… **3-level proxy** โ€” bypass geo-blocks with global, per-provider, and per-key proxies +โœ… **10 multi-modal APIs** โ€” chat, images, video, music, audio, search in one endpoint +โœ… **MCP + A2A** โ€” 29 MCP tools + agent-to-agent protocol, production-ready +โœ… **Universal** โ€” works with Claude Code, Codex, Gemini CLI, Cursor, Cline, OpenClaw, any CLI tool --- @@ -236,6 +244,151 @@ This generates a `system-info.txt` with your Node.js version, OmniRoute version, --- +## ๐Ÿ› ๏ธ Supported CLI Tools + +OmniRoute works seamlessly with **16+ AI coding tools** โ€” one config, all tools: + + + + + + + + + + + + + + + + + + + + + + + + + + +
Claude Code
Anthropic
Codex CLI
OpenAI
Gemini CLI
Google
Cursor
IDE
OpenClaw
CLI
Antigravity
VS Code
Cline
Extension
Continue
Extension
Kilo Code
Extension
Kiro
AWS IDE
OpenCode
CLI
Droid
CLI
AMP
CLI
Copilot
GitHub
Windsurf
IDE
Hermes
CLI
Qwen CLI
Alibaba
Custom
Any tool
+ +๐Ÿ“– Full setup for each tool: [`docs/CLI-TOOLS.md`](docs/CLI-TOOLS.md) + +--- + +## ๐ŸŒ Supported Providers โ€” 160+ + +### ๐Ÿ” OAuth Providers + + + + + + + + + + + + + + + +
Claude Code
Anthropic OAuth
Antigravity
Google OAuth
Codex
OpenAI OAuth
GitHub Copilot
GitHub OAuth
Cursor
Cursor OAuth
Kimi Coding
Moonshot OAuth
Kilo Code
Kilo OAuth
Cline
Cline OAuth
+ +### ๐Ÿ†“ Free Providers (No Cost) + + + + + + + + + + + + + + +
๐ŸŸข Kiro AI
Claude Sonnet/Haiku
Unlimited FREE
๐ŸŸข Qoder AI
Kimi-K2, DeepSeek-R1
Unlimited FREE
๐ŸŸข Pollinations
GPT-5, Claude, Llama 4
No API key needed
๐ŸŸข Qwen Code
Qwen3 Coder Plus
Unlimited FREE
๐ŸŸข LongCat AI
Flash-Lite
50M tokens/day
๐ŸŸข Cloudflare AI
50+ models
10K neurons/day
๐ŸŸข Puter AI
GPT-4.1, Claude
Rate-limited free
๐ŸŸข NVIDIA NIM
Llama, Mistral
1K req/day free
+ +### ๐Ÿ”‘ API Key Providers (120+) + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
OpenAIAnthropicGeminiDeepSeekGroqxAI (Grok)
MistralOpenRouterGLMKimiMiniMaxFireworks
Together AICerebrasCohereNVIDIAPerplexitySiliconFlow
NebiusHuggingFaceDeepInfraSambaNovaVertex AIAzure OpenAI
AWS BedrockSnowflakeDatabricksVenice.aiAI21 LabsMeta Llama
+ +
+...and 90+ more providers + +Alibaba ยท Amazon Q ยท AssemblyAI ยท Baidu Qianfan ยท Baseten ยท Black Forest Labs ยท Blackbox ยท Brave Search ยท Bytez ยท CablyAI ยท Cartesia ยท ChatGPT Web ยท Chutes.ai ยท Clarifai ยท Codestral ยท CrofAI ยท DataRobot ยท Deepgram ยท ElevenLabs ยท Empower ยท Exa Search ยท Fal.ai ยท Featherless AI ยท FenayAI ยท FriendliAI ยท Galadriel ยท GigaChat ยท GitLab Duo ยท GLHF Chat ยท GoAPI ยท Heroku AI ยท Hyperbolic ยท IBM watsonx ยท Inference.net ยท Inworld ยท Jina AI ยท Kilo Gateway ยท Lambda AI ยท LaoZhang ยท Linkup Search ยท LlamaGate ยท Maritalk ยท Modal ยท Moonshot AI ยท Morph ยท Muse Spark ยท NanoBanana ยท NanoGPT ยท NLP Cloud ยท Nous Research ยท Novita AI ยท nScale ยท OCI ยท Ollama Cloud ยท OVHcloud ยท PiAPI ยท PlayHT ยท Poe ยท Predibase ยท PublicAI ยท Qwen Code ยท Recraft ยท Reka ยท Runway ยท SAP ยท Scaleway ยท SearchAPI ยท SearXNG ยท Serper ยท Stability AI ยท Synthetic ยท Tavily ยท TheB.AI ยท Topaz ยท Upstage ยท v0 (Vercel) ยท Vercel AI Gateway ยท Volcengine ยท Voyage AI ยท W&B Inference ยท Xiaomi MiMo ยท You.com ยท Z.AI ยท + OpenAI/Anthropic-compatible custom endpoints + +
+ +### ๐Ÿ  Self-Hosted + + + + + + + + + + + + + + + + +
LM StudioOllamavLLMLlamafileDocker Model Runner
NVIDIA TritonXInferenceoobaboogaComfyUISD WebUI
+ +--- + ## ๐Ÿ”„ How It Works ``` @@ -684,24 +837,7 @@ No proxy? Use the built-in **1proxy** integration for **hundreds of free, valida > ๐Ÿ†• **New models added (Mar 2026):** Grok-4 Fast family at $0.20/$0.50/M (benchmarked at 1143ms โ€” 30% faster than Gemini 2.5 Flash), GLM-5 via Z.AI with 128K output, MiniMax M2.5 reasoning, DeepSeek V3.2 updated pricing, Kimi K2.5 via Moonshot direct API. -**๐Ÿ’ก $0 Combo Stack โ€” The Complete Free Setup:** - -``` -# ๐Ÿ†“ Ultimate Free Stack 2026 โ€” 11 Providers, $0 Forever -Kiro (kr/) โ†’ Claude Sonnet/Haiku UNLIMITED -Qoder (if/) โ†’ kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 UNLIMITED -LongCat Lite (lc/) โ†’ LongCat-Flash-Lite โ€” 50M tokens/day ๐Ÿ”ฅ -Pollinations (pol/) โ†’ GPT-5, Claude, DeepSeek, Llama 4 โ€” no key needed -Qwen (qw/) โ†’ qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next UNLIMITED -Gemini (gemini/) โ†’ Gemini 2.5 Flash โ€” 1,500 req/day free API key -Cloudflare AI (cf/) โ†’ Llama 70B, Gemma 3, Mistral โ€” 10K Neurons/day -Scaleway (scw/) โ†’ Qwen3 235B, Llama 70B โ€” 1M free tokens (EU) -Groq (groq/) โ†’ Llama/Gemma ultra-fast โ€” 14.4K req/day -NVIDIA NIM (nvidia/) โ†’ 70+ open models โ€” 40 RPM forever -Cerebras (cerebras/) โ†’ Llama/Qwen world-fastest โ€” 1M tok/day -``` - -**Zero cost. Never stops coding.** Configure this as one OmniRoute combo and all fallbacks happen automatically โ€” no manual switching ever. +**๐Ÿ’ก See the full [$0 Free Stack (11 providers)](#-free-models--11-providers-0-forever) below.** > ๐Ÿ’ก **Understanding Dashboard Costs:** > @@ -712,156 +848,45 @@ Cerebras (cerebras/) โ†’ Llama/Qwen world-fastest โ€” 1M tok/day --- -## ๐Ÿ†“ Free Models โ€” What You Actually Get +## ๐Ÿ†“ Free Models โ€” 11 Providers, $0 Forever -> All models below are **100% free with zero credit card required**. OmniRoute auto-routes between them when one quota runs out โ€” combine them all for an unbreakable $0 combo. +> Combine all free providers into one unbreakable combo โ€” OmniRoute auto-routes between them when quota runs out. -### ๐Ÿ”ต CLAUDE MODELS (via Kiro โ€” AWS Builder ID) +| Provider | Prefix | Free Models | Quota | +| ----------------- | ----------- | ------------------------------------------------------------- | ----------------- | +| **Kiro** | `kr/` | Claude Sonnet 4.5, Haiku 4.5, Opus 4.6 | โ™พ๏ธ Unlimited | +| **Qoder** | `if/` | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1, minimax-m2.1 | โ™พ๏ธ Unlimited | +| **Qwen** | `qw/` | qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next | โ™พ๏ธ Unlimited | +| **Pollinations** | `pol/` | GPT-5, Claude, Gemini, DeepSeek, Llama 4, Mistral | No key needed | +| **LongCat** | `lc/` | LongCat-Flash-Lite | 50M tokens/day ๐Ÿ”ฅ | +| **Gemini CLI** | `gc/` | gemini-3-flash, gemini-2.5-pro | 180K tok/mo | +| **Cloudflare AI** | `cf/` | 50+ models (Llama, Gemma, Mistral, Whisper) | 10K Neurons/day | +| **Groq** | `groq/` | Llama 3.3 70B, Qwen3 32B, Kimi K2 | 14.4K RPD | +| **NVIDIA NIM** | `nvidia/` | 129 models (DeepSeek, Llama, GLM, Kimi) | ~40 RPM | +| **Cerebras** | `cerebras/` | Qwen3 235B, GPT-OSS 120B, Llama 3.1 | 1M tok/day | +| **Scaleway** | `scw/` | Qwen3 235B, Llama 70B, DeepSeek V3 | 1M tokens (EU) | -| Model | Prefix | Limit | Rate Limit | -| ------------------- | ------ | ------------- | --------------------- | -| `claude-sonnet-4.5` | `kr/` | **Unlimited** | No reported daily cap | -| `claude-haiku-4.5` | `kr/` | **Unlimited** | No reported daily cap | -| `claude-opus-4.6` | `kr/` | **Unlimited** | Latest Opus via Kiro | +
+๐Ÿ“– 25+ more free providers โ€” Groq, Cerebras, Mistral, GitHub Models, OpenRouter, and more -### ๐ŸŸข QODER MODELS (Free PAT via qodercli) +**Also free (API Key required):** +Mistral (1B tok/month) ยท OpenRouter (35+ `:free` models) ยท GitHub Models (GPT-5, 45+ models) ยท +Cohere (1K calls/month) ยท Z.AI/GLM (permanent free Flash models) ยท SiliconFlow (1K RPM, 50K TPM) ยท +Kilo Code (~200 req/hr auto-router) ยท HuggingFace ($0.10/mo credits) ยท Ollama Cloud (400+ models) ยท +LLM7.io (30+ models) ยท Kluster AI ยท IBM watsonx (300K tok/month) ยท OpenCode Zen ยท Vercel AI Gateway ($5/mo) -| Model | Prefix | Limit | Rate Limit | -| ------------------ | ------ | ------------- | --------------- | -| `kimi-k2-thinking` | `if/` | **Unlimited** | No reported cap | -| `qwen3-coder-plus` | `if/` | **Unlimited** | No reported cap | -| `deepseek-r1` | `if/` | **Unlimited** | No reported cap | -| `minimax-m2.1` | `if/` | **Unlimited** | No reported cap | -| `kimi-k2` | `if/` | **Unlimited** | No reported cap | +**Trial credits (one-time):** +Baseten ($30) ยท NLP Cloud ($15) ยท AI21 ($10) ยท Upstage ($10) ยท SambaNova ($5) ยท Modal ($5/mo) ยท +Fireworks ($1) ยท Nebius ($1) ยท Inference.net ($1 + $25 survey) ยท Hyperbolic ($1) ยท Novita ($0.50) -> Recommended connection method: **Personal Access Token + `qodercli`**. Browser OAuth is -> experimental and disabled by default unless `QODER_OAUTH_*` environment variables are configured. +**China-based (free tiers):** +ModelScope ยท Tencent Hunyuan ยท Volcengine ยท ChatAnywhere ยท InternAI ยท Bigmodel -### ๐ŸŸก QWEN MODELS (Device Code Auth) +**Combined capacity: ~31,000+ RPD ยท ~32B+ tokens/month ยท 500+ models ยท $0** -| Model | Prefix | Limit | Rate Limit | -| ------------------- | ------ | ------------- | ------------------- | -| `qwen3-coder-plus` | `qw/` | **Unlimited** | No reported cap | -| `qwen3-coder-flash` | `qw/` | **Unlimited** | No reported cap | -| `qwen3-coder-next` | `qw/` | **Unlimited** | No reported cap | -| `vision-model` | `qw/` | **Unlimited** | Multimodal (images) | +
-### ๐ŸŸฃ GEMINI CLI (Google OAuth) - -| Model | Prefix | Limit | Rate Limit | -| ------------------------ | ------ | --------------------------- | ------------- | -| `gemini-3-flash-preview` | `gc/` | **180K tok/month** + 1K/day | Monthly reset | -| `gemini-2.5-pro` | `gc/` | 180K/month (shared pool) | High quality | - -### โšซ NVIDIA NIM (Free API Key โ€” build.nvidia.com) - -| Tier | Daily Limit | Rate Limit | Notes | -| ---------- | ------------ | ----------- | ------------------------------------------------------ | -| Free (Dev) | No token cap | **~40 RPM** | 70+ models; transitioning to pure rate limits mid-2025 | - -Popular free models: `moonshotai/kimi-k2.5` (Kimi K2.5), `z-ai/glm4.7` (GLM 4.7), `deepseek-ai/deepseek-v3.2` (DeepSeek V3.2), `nvidia/llama-3.3-70b-instruct`, `deepseek/deepseek-r1` - -### โšช CEREBRAS (Free API Key โ€” inference.cerebras.ai) - -| Tier | Daily Limit | Rate Limit | Notes | -| ---- | ----------------- | ---------------- | ------------------------------------------- | -| Free | **1M tokens/day** | 60K TPM / 30 RPM | World's fastest LLM inference; resets daily | - -Available free: `llama-3.3-70b`, `llama-3.1-8b`, `deepseek-r1-distill-llama-70b` - -### ๐Ÿ”ด GROQ (Free API Key โ€” console.groq.com) - -| Tier | Daily Limit | Rate Limit | Notes | -| ---- | ------------- | ---------------- | ----------------------------------------- | -| Free | **14.4K RPD** | 30 RPM per model | No credit card; 429 on limit, not charged | - -Available free: `llama-3.3-70b-versatile`, `gemma2-9b-it`, `mixtral-8x7b`, `whisper-large-v3` - -### ๐Ÿ”ด LONGCAT AI (Free API Key โ€” longcat.chat) ๐Ÿ†• - -| Model | Prefix | Daily Free Quota | Notes | -| ----------------------------- | ------ | ----------------- | ----------------------- | -| `LongCat-Flash-Lite` | `lc/` | **50M tokens** ๐Ÿ’ฅ | Largest free quota ever | -| `LongCat-Flash-Chat` | `lc/` | 500K tokens | Multi-turn chat | -| `LongCat-Flash-Thinking` | `lc/` | 500K tokens | Reasoning / CoT | -| `LongCat-Flash-Thinking-2601` | `lc/` | 500K tokens | Jan 2026 version | -| `LongCat-Flash-Omni-2603` | `lc/` | 500K tokens | Multimodal | - -> 100% free while in public beta. Sign up at [longcat.chat](https://longcat.chat) with email or phone. Resets daily 00:00 UTC. - -### ๐ŸŸข POLLINATIONS AI (No API Key Required) ๐Ÿ†• - -| Model | Prefix | Rate Limit | Provider Behind | -| ---------- | ------ | ---------- | ------------------ | -| `openai` | `pol/` | 1 req/15s | GPT-5 | -| `claude` | `pol/` | 1 req/15s | Anthropic Claude | -| `gemini` | `pol/` | 1 req/15s | Google Gemini | -| `deepseek` | `pol/` | 1 req/15s | DeepSeek V3 | -| `llama` | `pol/` | 1 req/15s | Meta Llama 4 Scout | -| `mistral` | `pol/` | 1 req/15s | Mistral AI | - -> โœจ **Zero friction:** No signup, no API key. Add the Pollinations provider with an empty key field and it works immediately. - -### ๐ŸŸ  CLOUDFLARE WORKERS AI (Free API Key โ€” cloudflare.com) ๐Ÿ†• - -| Tier | Daily Neurons | Equivalent Usage | Notes | -| ---- | ------------- | --------------------------------------- | ----------------------- | -| Free | **10,000** | ~150 LLM resp / 500s audio / 15K embeds | Global edge, 50+ models | - -Popular free models: `@cf/meta/llama-3.3-70b-instruct`, `@cf/google/gemma-3-12b-it`, `@cf/openai/whisper-large-v3-turbo` (free audio!), `@cf/qwen/qwen2.5-coder-15b-instruct` - -> Requires API Token + Account ID from [dash.cloudflare.com](https://dash.cloudflare.com). Store Account ID in provider settings. - -### ๐ŸŸฃ SCALEWAY AI (1M Free Tokens โ€” scaleway.com) ๐Ÿ†• - -| Tier | Free Quota | Location | Notes | -| ---- | ------------- | ------------ | ----------------------------------- | -| Free | **1M tokens** | ๐Ÿ‡ซ๐Ÿ‡ท Paris, EU | No credit card needed within limits | - -Available free: `qwen3-235b-a22b-instruct-2507` (Qwen3 235B!), `llama-3.1-70b-instruct`, `mistral-small-3.2-24b-instruct-2506`, `deepseek-v3-0324` - -> EU/GDPR compliant. Get API key at [console.scaleway.com](https://console.scaleway.com). - -> **๐Ÿ’ก The Ultimate Free Stack (11 Providers, $0 Forever):** -> -> ``` -> Kiro (kr/) โ†’ Claude Sonnet/Haiku UNLIMITED -> Qoder (if/) โ†’ kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 UNLIMITED -> LongCat Lite (lc/) โ†’ LongCat-Flash-Lite โ€” 50M tokens/day ๐Ÿ”ฅ -> Pollinations (pol/) โ†’ GPT-5, Claude, DeepSeek, Llama 4 โ€” no key needed -> Qwen (qw/) โ†’ qwen3-coder models UNLIMITED -> Gemini (gemini/) โ†’ Gemini 2.5 Flash โ€” 1,500 req/day free -> Cloudflare AI (cf/) โ†’ 50+ models โ€” 10K Neurons/day -> Scaleway (scw/) โ†’ Qwen3 235B, Llama 70B โ€” 1M free tokens (EU) -> Groq (groq/) โ†’ Llama/Gemma โ€” 14.4K req/day ultra-fast -> NVIDIA NIM (nvidia/) โ†’ 70+ open models โ€” 40 RPM forever -> Cerebras (cerebras/) โ†’ Llama/Qwen world-fastest โ€” 1M tok/day -> ``` - ---- - -## ๐ŸŒ Free API Provider Directory โ€” 25+ Providers, 500+ Models, $0 - -> **We analyzed 6 community repositories** aggregating free LLM API providers and consolidated everything into one definitive reference. This is the most comprehensive free-tier directory available. - -| Provider | Best Free Model | RPM | RPD | Tokens | Speed | -| ----------------- | ---------------- | ----- | ------------ | ------------- | --------- | -| **Groq** | Llama 3.3 70B | 30 | 14,400 | 6K TPM | ๐ŸŸข Fast | -| **Cerebras** | Qwen3 235B | 30 | 14,400 | 1M TPD | ๐ŸŸข Fast | -| **Mistral AI** | Mistral Large 3 | 60 | Unlimited | 1B/month | ๐ŸŸก Medium | -| **Google Gemini** | Gemini 2.5 Flash | 5โ€“15 | 20โ€“1,500 | 250K TPM | ๐ŸŸข Fast | -| **NVIDIA NIM** | 129 models | 40 | โ€” | โ€” | ๐ŸŸก Medium | -| **OpenRouter** | 35+ :free models | 20 | 50โ€“1,000 | โ€” | ๐ŸŸก Medium | -| **GitHub Models** | GPT-4.1, GPT-5 | 10โ€“15 | 50โ€“150 | 8K/4K per req | ๐ŸŸก Medium | -| **Cloudflare AI** | 50+ models | โ€” | 10K neurons | โ€” | ๐ŸŸก Medium | -| **Pollinations** | Text+Image+Video | โ€” | Hourly reset | โ€” | ๐ŸŸก Medium | -| **SiliconFlow** | Qwen3-8B | 1,000 | โ€” | 50K TPM | ๐ŸŸก Medium | - -**Combined free capacity across all providers: ~31,000+ RPD ยท ~32B+ tokens/month ยท 500+ models ยท $0 forever.** - -The full directory includes 25+ providers with detailed rate limits, base URLs, model tables, trial credit providers (Baseten $30, AI21 $10, SambaNova $5, etc.), China-specific platforms (ModelScope, Volcengine, Tencent Hunyuan), and step-by-step OmniRoute combo configuration. - -๐Ÿ“– **Complete free provider directory with all models, quotas, and integration guide:** [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md) +๐Ÿ“– **Complete free provider directory:** [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md) โ€” 25+ providers, quotas, base URLs, model tables, and OmniRoute combo setup. --- @@ -875,6 +900,8 @@ The full directory includes 25+ providers with detailed rate limits, base URLs, | ๐Ÿ”ต **AssemblyAI** | **$50 free** (signup) | `universal-3-pro` โ€” chapters, sentiment, PII | No RPM limit on free credits | | ๐Ÿ”ด **Groq** | **Free forever** | `whisper-large-v3` โ€” OpenAI Whisper | 30 RPM (rate limited) | +--- + **Suggested combo in `/dashboard/combos`:** ``` @@ -890,94 +917,67 @@ Then in `/dashboard/media` โ†’ **Transcription** tab: upload any audio or video ## ๐Ÿ’ก Key Features -OmniRoute v3.7+ is an operational platform, not just a relay proxy โ€” backed by **4,690+ automated tests** across 517 test files. +> **4,690+ automated tests** across 517 test files. Not just a relay โ€” a full operational platform. -| Category | Feature | Why It Matters | -| -------------------- | -------------------------------------------------------------------------------- | ---------------------------------- | -| ๐Ÿง  **Routing** | Smart 4-Tier Fallback (Subscription โ†’ API โ†’ Cheap โ†’ Free) | Never stop coding, zero downtime | -| | 13 Balancing Strategies + Custom Combos | Tailor routing to your exact needs | -| | Task-Aware Smart Routing (coding/vision/analysis) | Right model for every task | -| | Context Relay โ€” session handoffs during rotation | No lost context mid-conversation | -| | Thinking Budget Controls (passthrough/auto/custom) | Control reasoning costs precisely | -| ๐Ÿ”„ **Translation** | OpenAI โ†” Claude โ†” Gemini โ†” Responses API | Works with ANY CLI tool | -| | Auto Token Refresh (OAuth PKCE for 8 providers) | No manual re-login ever | -| | Responses API โ€” full `/v1/responses` for Codex | First-class Codex compatibility | -| ๐ŸŽต **Multi-Modal** | 10 APIs: chat, embed, images, video, music, TTS, STT, moderation, rerank, search | One endpoint for everything | -| | Batch API โ€” asynchronous processing with Files API | Background bulk processing | -| | OpenAPI 3.0 โ€” live auto-generated spec + Try-It UI | API-first development | -| ๐Ÿ›ก๏ธ **Resilience** | Circuit Breakers + Connection Cooldown + Anti-Thundering Herd | Auto-recovery from failures | -| | TLS Fingerprint Spoofing + CLI Fingerprint Matching | Stealth + anti-ban protection | -| | Semantic + Signature Cache (two-tier) | Reduce costs + latency | -| | Request Idempotency + Rate Limit Detection | No duplicate charges | -| ๐Ÿค– **Protocols** | MCP Server โ€” 29 tools, 3 transports, 10 scopes | IDE/agent tool integration | -| | A2A Server โ€” JSON-RPC 2.0, SSE streaming, task lifecycle | Agent-to-agent orchestration | -| | ACP โ€” CLI agent discovery (14 agents + custom) | Universal agent onboarding | -| ๐Ÿ“Š **Observability** | Unified Logs (request/proxy/audit/console) + p50/p95/p99 | Full request telemetry | -| | Health Dashboard โ€” uptime, breakers, cache, lockouts | Operational visibility | -| | Cost Tracking + Budget Controls | Financial governance | -| | Evaluation Framework โ€” golden set testing | Quality assurance | -| โ˜๏ธ **Platform** | Desktop (Electron), Android (Termux), PWA | Run anywhere | -| | Docker (AMD64 + ARM64) with Compose profiles | One-command deploy | -| | Cloudflare / Tailscale / ngrok Tunnels | Instant public endpoint | -| | 40+ languages with RTL support | Global accessibility | -| ๐Ÿ—œ๏ธ **Compression** | 5-mode pipeline: off / lite / standard / aggressive / ultra | Save 15-75% tokens | +| Feature | Why It Matters | +| ---------------------------------------------------------------------------------------------------- | -------------------------------- | +| ๐Ÿง  **Smart 4-Tier Fallback** โ€” Subscription โ†’ API โ†’ Cheap โ†’ Free | Never stop coding, zero downtime | +| ๐Ÿ”„ **Format Translation** โ€” OpenAI โ†” Claude โ†” Gemini โ†” Responses API | Works with ANY CLI tool | +| ๐Ÿ—œ๏ธ **Prompt Compression** โ€” 5-mode pipeline (off โ†’ lite โ†’ standard โ†’ aggressive โ†’ ultra) | Save 15-75% tokens automatically | +| ๐Ÿค– **MCP Server** โ€” 29 tools, 3 transports (stdio/SSE/HTTP), 10 scopes | IDE/agent tool integration | +| ๐Ÿ›ก๏ธ **Resilience Engine** โ€” circuit breakers, cooldowns, TLS spoofing, anti-thundering herd | Auto-recovery from any failure | +| ๐ŸŽต **10 Multi-Modal APIs** โ€” chat, embed, images, video, music, TTS, STT, moderation, rerank, search | One endpoint for everything | +| ๐ŸŒ **3-Level Proxy** โ€” global, per-provider, per-key + 1proxy free marketplace | Access AI from any country | +| ๐Ÿ“Š **Full Observability** โ€” unified logs, p50/p95/p99 telemetry, cost tracking, budget controls | Know exactly what's happening |
-๐Ÿ†• What's New โ€” v3.6+ Highlights +๐Ÿ“‹ Complete feature list โ€” 30+ capabilities -- ๐ŸŒ V1 WebSocket Bridge โ€” OpenAI-compatible WS at `/v1/ws` -- ๐Ÿ”‘ Sync Tokens & Config Bundle โ€” versioned config sync with ETag -- ๐Ÿง  GLM Thinking (glmt) โ€” 65K tokens, 24K thinking budget, Claude-compatible -- ๐Ÿ”ข Hybrid Token Counting โ€” provider-side + estimation fallback -- ๐Ÿ›ก๏ธ Safe Outbound Fetch โ€” SSRF protection on all provider calls -- โณ Wait For Cooldown โ€” auto-retry after connection cooldowns -- ๐Ÿ” Runtime Env Validation โ€” Zod schemas at startup -- ๐Ÿ“‹ Compliance Audit v2 โ€” pagination, auth events, SSRF logging -- ๐Ÿ”” Webhooks โ€” event-driven with test firing and dashboard management -- ๐Ÿ‘๏ธ Vision Bridge โ€” image analysis guardrail before routing -- โšก Grok-4 Fast โ€” $0.20/$0.50/M, 30% faster than Gemini Flash -- ๐Ÿง  GLM-5 via Z.AI โ€” 128K output, $0.5/1M -- ๐Ÿ”ฎ MiniMax M2.5 โ€” reasoning + agentic at $0.3/1M -- ๐ŸŽฏ toolCalling flag โ€” per-model tool capability in registry -- ๐ŸŒ Multilingual Intent Detection โ€” PT/ZH/ES/AR in AutoCombo -- ๐Ÿ“Š Benchmark-Driven Fallbacks โ€” real p95 latency feeds scoring -- ๐Ÿ” Request Deduplication โ€” content-hash dedup window +**Routing & Intelligence** -
+- 13 balancing strategies (priority, weighted, round-robin, P2C, cost-optimized, context-relay...) +- Task-aware smart routing (coding/vision/analysis) ยท Context relay session handoffs +- Thinking budget controls (passthrough/auto/custom) ยท Wildcard routing ยท System prompt injection -
-๐Ÿ“– Feature Deep Dive โ€” Expanded Details +**Translation & Compatibility** -#### Smart fallback with practical cost control +- Auto token refresh (OAuth PKCE for 8 providers) ยท Multi-account round-robin +- Responses API โ€” full `/v1/responses` for Codex ยท Batch API with Files API +- OpenAPI 3.0 live spec + Try-It UI + +**Protocols** + +- A2A Server โ€” JSON-RPC 2.0, SSE streaming, task lifecycle, skills +- ACP โ€” CLI agent discovery (14 agents + custom) + +**Platform** + +- Desktop (Electron) ยท Android (Termux) ยท PWA ยท Docker (AMD64 + ARM64) +- Cloudflare / Tailscale / ngrok tunnels ยท 40+ languages with RTL +- Semantic + signature cache (two-tier) ยท Request idempotency + deduplication + +**Observability** + +- Health dashboard โ€” uptime, breakers, cache, lockouts +- Evaluation framework โ€” golden set testing ยท Webhooks ยท Compliance audit + +**v3.6+ Highlights:** +V1 WebSocket Bridge ยท Sync Tokens & Config Bundle ยท GLM Thinking (glmt) ยท Hybrid Token Counting ยท +Safe Outbound Fetch ยท Wait For Cooldown ยท Runtime Env Validation ยท Vision Bridge ยท +Grok-4 Fast ยท GLM-5 via Z.AI ยท MiniMax M2.5 ยท toolCalling flag ยท +Multilingual Intent Detection ยท Benchmark-Driven Fallbacks ยท Request Deduplication + +**Architecture Examples:** ```txt -Combo: "my-coding-stack" - 1. cc/claude-opus-4-7 - 2. nvidia/llama-3.3-70b - 3. glm/glm-4.7 +Combo: "my-coding-stack" Format Translation: + 1. cc/claude-opus-4-7 CLI โ†’ OpenAI format + 2. nvidia/llama-3.3-70b OmniRoute โ†’ translates + 3. glm/glm-4.7 Provider โ†’ native format 4. if/kimi-k2-thinking ``` -When quota, rate, or health fails, OmniRoute automatically moves to the next candidate without manual switching. - -#### Prompt Compression โ€” Token Savings Breakdown - -``` -Without compression: 47K tokens sent to LLM -With Lite: 40K tokens sent (15% saved โ€” safe, always-on) -With Standard: 33K tokens sent (30% saved โ€” caveman-speak rules) -With Aggressive: 24K tokens sent (50% saved โ€” aging + summarization) -With Ultra: 12K tokens sent (75% saved โ€” heuristic pruning) -``` - -#### Format Translation โ€” Universal Compatibility - -- **OpenAI** โ†” **Claude** โ†” **Gemini** โ†” **Cursor** โ†” **Kiro** โ†” **Vertex** โ†” **Antigravity** โ†” **Ollama** โ†” **Responses** -- Your CLI tool sends OpenAI format โ†’ OmniRoute translates โ†’ Provider receives native format - -> ๐Ÿ“– **[MCP Server README](open-sse/mcp-server/README.md)** โ€” Tool reference, IDE configs, and client examples -> -> ๐Ÿ“– **[A2A Server README](src/lib/a2a/README.md)** โ€” Skills, JSON-RPC methods, streaming, and task lifecycle +๐Ÿ“– [MCP Server README](open-sse/mcp-server/README.md) ยท [A2A Server README](src/lib/a2a/README.md) ยท [Resilience Guide](docs/RESILIENCE_GUIDE.md) ยท [Features Gallery](docs/FEATURES.md)
@@ -1349,31 +1349,6 @@ See the [Proxy Guide](docs/PROXY_GUIDE.md) for setup instructions. --- -## ๐Ÿ—บ๏ธ Roadmap - -OmniRoute has **218+ features planned** across multiple development phases. Here are the key areas: - -| Category | Planned Features | Highlights | -| ----------------------------- | ---------------- | ----------------------------------------------------------------------------------------------------- | -| ๐Ÿง  **Routing & Intelligence** | 25+ | Lowest-latency routing, tag-based routing, quota preflight, quota-aware P2C, step-based combo routing | -| ๐Ÿ”’ **Security & Compliance** | 20+ | SSRF hardening, credential cloaking, rate-limit per endpoint, management key scoping | -| ๐Ÿ“Š **Observability** | 15+ | OpenTelemetry integration, real-time quota monitoring, combo target health, cost tracking per model | -| ๐Ÿ”„ **Provider Integrations** | 20+ | Dynamic model registry, connection cooldowns, multi-account Codex, Copilot quota parsing | -| โšก **Performance** | 15+ | Dual cache layer, prompt cache, response cache, streaming keepalive, batch API | -| ๐ŸŒ **Ecosystem** | 10+ | WebSocket API, config hot-reload, distributed config store, commercial mode | - -### ๐Ÿ”œ Coming Soon - -- ๐Ÿ”— **OpenCode Integration** โ€” Native provider support for the OpenCode AI coding IDE -- ๐Ÿ”— **TRAE Integration** โ€” Full support for the TRAE AI development framework -- ๐Ÿ’ฐ **Lowest-Cost Strategy** โ€” Automatically select the cheapest available provider -- ๐Ÿ“Š **OpenTelemetry Integration** โ€” OTLP traces and metrics export for enterprise observability -- ๐Ÿ”„ **Config Hot-Reload** โ€” Apply settings changes without server restart - -> ๐Ÿ“ Full feature specifications available in [`docs/new-features/`](docs/new-features/) (217 detailed specs) - ---- - ## โญ Top Contributors > OmniRoute is shaped by a passionate open-source community. These individuals have made exceptional contributions that directly impact the quality, stability, and reach of the project. **Thank you.**