Updates docs/guides/I18N.md with a 'Translation pipeline (recommended)' section documenting the npm run i18n:run / :check / :run:dry flow, required env vars, state file semantics, and the legacy-script deprecation notice. The legacy Quick-Reference row for the Python translator is replaced with the new npm script entry. Re-translates two sources end-to-end through the new pipeline: - CLAUDE.md (1 chunk, 23k chars) - docs/architecture/ARCHITECTURE.md (14 chunks, 74k chars, one timeout retry) Both translations now have a fresh language bar regenerated from config/i18n.json (41 locales), an H1 heading with the native language tag, and prose translated into Brazilian Portuguese while preserving markdown syntax, code blocks, command names, env var identifiers, and version numbers verbatim. .i18n-state.json records the SHA-256 hash for each source and target. A second invocation with no source changes correctly reports 'work units: 0 (skipped up-to-date: 2 of 2)' and `npm run i18n:check` exits 0 — confirming the hash-based incremental + drift detection paths both work as designed. - Elapsed: ~10 min total at concurrency=4 - Cost: ~75k chars output through the configured backend Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
🚀 OmniRoute
The Free AI Gateway — one endpoint, 177+ providers, zero downtime.
Auto-fallback to free models. Stop coding interruptions. Cut tokens 15-95%.
Website · Quick Start · Docs · Discord/WhatsApp
v3.8.0 · MIT · Production-ready · Self-hosted
⚡ The Pitch (60 seconds)
| 🎯 The problem | ✅ How OmniRoute solves it |
|---|---|
| Hit rate limits on Claude/GPT? | Auto-fallback across 177 providers — never see a 429 again |
| Bored of switching API keys? | One endpoint (localhost:20128) speaks OpenAI, Anthropic, Gemini, Claude Code, Cursor formats |
| Paying $200/mo for AI? | 11 free providers + intelligent routing → most users pay $0 |
| Tokens too expensive? | RTK + Caveman compression saves 15-95% on eligible payloads |
| Blocked region? | 4-level proxy (account/provider/combo/global) + 1proxy free marketplace |
| Want CLI agents free? | Plug Cursor, Cline, Codex, Claude Code, Aider, 15+ CLIs at OmniRoute |
📺 Watch in action: Video demo
🖼️ Dashboard Preview
⚡ Quick Start
# Run instantly (npx — no install needed)
npx -y omniroute@latest
# Or install globally
npm install -g omniroute && omniroute
# Or via Docker
docker run -d -p 20128:20128 diegosouzapw/omniroute:3.8.0
→ Open http://localhost:20128 → login with admin / CHANGEME → connect your first provider via OAuth or API key.
Point any OpenAI-compatible client at OmniRoute:
export OPENAI_BASE_URL=http://localhost:20128/v1
export OPENAI_API_KEY=or_<your-omniroute-key>
That's it. Cursor, Cline, Codex, Continue, Aider, and any SDK now work. → Detailed setup: docs/guides/SETUP_GUIDE.md
🌟 What's new in v3.8.0
- 🤖 Auto-Combo zero-config routing — just use
auto/coding,auto/cheap,auto/fast,auto/offline,auto/smart,auto/lkgpas model IDs - 🎯 Manifest-aware tier routing W1-W4 — automatic tier prioritization
- 🆕 Command Code provider + Z.AI quota labels + KIE video expansion
- 🔐 Windsurf + Devin CLI + GitLab Duo OAuth flows
- 🆓 9 new free providers: LLM7, Lepton, Kluster, UncloseAI, BazaarLink, Completions, Enally, FreeTheAi, AgentRouter ($200 credits)
- 🩺 Model Cooldowns dashboard with manual re-enable
- 🎨 Cursor full OpenAI parity (tools, streaming, sessions)
- 📌 Per-session sticky routing for Codex
- 🔊 Inworld TTS enhancements
- 🧠 Reasoning Replay Cache — fixes 400s on DeepSeek V4, Kimi K2, Qwen-Thinking, GLM
- 🔄 Reset-aware routing strategy (14th strategy)
- 🛠️ 20+ new CLI commands (
omniroute setup/doctor/providers/combos)
→ Full changelog: CHANGELOG.md
🎯 Why OmniRoute Wins
| OmniRoute v3.8 | LiteLLM | OpenRouter | |
|---|---|---|---|
| Providers | 177+ | ~50 | ~50 |
| Free providers | 11 | 0 | 0 |
| OAuth providers | 14 | 0 | 1 |
| Routing strategies | 14 | 3 | 1 |
| Auto routing | ✅ 9-factor scoring | ❌ | ❌ |
| Prompt compression | ✅ RTK + Caveman | ❌ | ❌ |
| MCP server | ✅ 37 tools | ❌ | ❌ |
| A2A protocol | ✅ v0.3 + 5 skills | ❌ | ❌ |
| Desktop app | ✅ Electron 41 | ❌ | ❌ |
| PWA | ✅ | ❌ | ❌ |
| Self-hosted | ✅ MIT | Limited | ❌ (cloud) |
| Pricing | $0 forever | OSS / Cloud paid | 10% fee + API |
→ Detailed comparison: docs/guides/FEATURES.md
🛠️ Compatible CLI Tools (17+)
All work out-of-the-box once you point OPENAI_BASE_URL at OmniRoute:
Claude family: Claude Code · Cline · Continue · Kilo Code · Kimi Coding OpenAI family: Codex CLI · Cursor · Aider · OpenClaw · Droid · AMP Google family: Gemini CLI · Antigravity · Jules Others: Windsurf · GitLab Duo · Devin CLI · Hermes · Amazon Q · Kiro · Qoder · Custom
→ Full setup: docs/reference/CLI-TOOLS.md
🌐 Providers (177+)
🆓 Free providers (11 — no API key or unlimited tier)
| Provider | Highlight |
|---|---|
| Kiro AI | 50 credits/month (Claude Sonnet/Haiku) |
| Qoder AI | Unlimited (Kimi-K2, Qwen3, DeepSeek-R1) |
| Gemini CLI | 180K tokens/month |
| Amazon Q | AWS Builder ID OAuth |
| LongCat | 50M tokens/day |
| Pollinations | No API key, GPT-5 + Claude |
| AgentRouter | $200 free credits |
| LLM7 · Lepton · Kluster · UncloseAI | New v3.8 free tiers |
⚠️ Qwen Code OAuth was discontinued on 2026-04-15 (use API key with alicode provider instead).
→ Curated guide: docs/reference/FREE_TIERS.md · Full catalog: docs/reference/PROVIDER_REFERENCE.md (auto-generated)
🔐 OAuth providers (14)
Claude Code · Codex · GitHub Copilot · Cursor · Antigravity · Gemini · Kimi Coding · Kilo Code · Cline · Qwen · Kiro · Qoder · Windsurf · GitLab Duo
🔑 API key providers (~123)
OpenAI · Anthropic · Google · Mistral · Cohere · DeepSeek · Groq · Together · Fireworks · Cerebras · SambaNova · NVIDIA NIM · Bedrock · Vertex · Azure · Cloudflare AI · 100+ more.
🏠 Self-hosted (10)
Ollama · LM Studio · vLLM · Llamafile · Lemonade · Petals · Triton · Docker Model Runner · Xinference · Oobabooga
🤖 Auto-Combo — Zero-Config Routing
Just use auto/<variant> as model ID. No combo setup needed.
# 6 variants + plain `auto`:
auto/coding # → optimized for coding tasks
auto/cheap # → minimize cost
auto/fast # → minimize latency
auto/offline # → prefer local providers
auto/smart # → prefer top-tier models
auto/lkgp # → Last-Known-Good-Path (sticky)
auto # → balanced default
How it picks: 9-factor scoring (health · quota · cost · latency · taskFit · stability · tierPriority · tierAffinity · specificityMatch) over a virtual candidate pool built from all enabled providers.
→ Full guide: docs/routing/AUTO-COMBO.md
🗜️ Prompt Compression — Save 15-95% Tokens
Two engines, stackable:
- Caveman — natural-language condensation (filler removal, hedging, repeated context). 30+ regex rules per language pack (en, es, pt-BR, de, fr, ja).
- RTK — terminal/shell/git/test output. 49 declarative filters.
Modes: off · lite · standard · aggressive · ultra · rtk · stacked (RTK→Caveman, max savings).
→ docs/compression/COMPRESSION_GUIDE.md · docs/compression/RTK_COMPRESSION.md · docs/compression/COMPRESSION_LANGUAGE_PACKS.md
🌍 Bypass Geographic Blocks
For users in Russia, China, Iran, Cuba, Turkey and other regions:
- 4-level outbound proxy — account / provider / combo / global scopes
- 1proxy free marketplace — auto-syncs working HTTP/SOCKS5 proxies
- Anti-detection — TLS fingerprinting (JA3/JA4), CCH headshakes, header sanitization
- Public tunnels — Cloudflare (Quick or Named), ngrok, Tailscale Funnel for OAuth callbacks
→ docs/ops/PROXY_GUIDE.md · docs/ops/TUNNELS_GUIDE.md · docs/security/STEALTH_GUIDE.md
📱 Multi-Platform
| Platform | Install | Doc |
|---|---|---|
| CLI / Server | npm install -g omniroute |
SETUP_GUIDE.md |
| Desktop (Win/Mac/Linux) | Electron installer from GitHub Releases | ELECTRON_GUIDE.md |
| PWA | Install from any modern browser | PWA_GUIDE.md |
| Android (Termux) | pkg install nodejs-lts && npm i -g omniroute |
TERMUX_GUIDE.md |
| Docker | docker compose up (base/cli/host/cliproxyapi profiles) |
DOCKER_GUIDE.md |
| VM / VPS | Generic Ubuntu/Debian + nginx + systemd | VM_DEPLOYMENT_GUIDE.md |
| Fly.io | fly deploy |
FLY_IO_DEPLOYMENT_GUIDE.md |
🧩 Extensibility
| System | What it does | Docs |
|---|---|---|
| 🧠 Skills | Built-in skills + marketplace + sandboxed custom skills (Docker) | docs/frameworks/SKILLS.md |
| 💾 Memory | Persistent conversational memory (SQLite FTS5 + Qdrant vector) | docs/frameworks/MEMORY.md |
| ☁️ Cloud Agents | Submit long tasks to Codex Cloud / Devin / Jules | docs/frameworks/CLOUD_AGENT.md |
| 🪝 Webhooks | HMAC-signed event delivery (request.completed, quota.exceeded, etc.) | docs/frameworks/WEBHOOKS.md |
| 🛡️ Guardrails | PII masker, prompt injection guard, vision bridge — hot-reload | docs/security/GUARDRAILS.md |
| 🧪 Evals | Suite-based regression testing (combos/models/cases/rubrics) | docs/frameworks/EVALS.md |
| 🔍 Compliance/Audit | audit_log table, retention, noLog opt-out, SSRF logging |
docs/security/COMPLIANCE.md |
| 🛡️ MCP Server | 37 tools, 3 transports (stdio/SSE/Streamable HTTP), ~13 scopes | docs/frameworks/MCP-SERVER.md |
| 🤝 A2A Protocol | v0.3 JSON-RPC, 5 skills (smart-routing, quota, discovery, cost, health) | docs/frameworks/A2A-SERVER.md |
📚 Documentation
Everything you need, organized by area.
🚀 Start here
SETUP_GUIDE.md— install + connect first providerUSER_GUIDE.md— end-user manual (modes, combos, CLIs, audio, ~1200 lines)FREE_TIERS.md— start free, no cardTROUBLESHOOTING.md— common issues + v3.8 known issues
🏛️ Architecture
ARCHITECTURE.md— high-level architectureCODEBASE_DOCUMENTATION.md— engineering referenceREPOSITORY_MAP.md— every directory and root fileFEATURES.md— full feature matrix
🔌 API & contracts
API_REFERENCE.md— endpoint referenceopenapi.yaml— OpenAPI 3.0 specPROVIDER_REFERENCE.md— full catalog (auto-generated)CLI-TOOLS.md— CLI integrations + internal CLIENVIRONMENT.md— all env vars
🎯 Routing & resilience
AUTO-COMBO.md— Auto-Combo (9-factor scoring, 14 strategies)RESILIENCE_GUIDE.md— circuit breaker + cooldown + lockoutREASONING_REPLAY.md— reasoning cache for DeepSeek/Kimi/QwenSTEALTH_GUIDE.md— TLS fingerprinting + obfuscation
🤖 Agent protocols
AGENT_PROTOCOLS_GUIDE.md— A2A vs ACP vs Cloud AgentsMCP-SERVER.md— Model Context Protocol serverA2A-SERVER.md— Agent-to-Agent protocolCLOUD_AGENT.md— Codex Cloud / Devin / Jules
🧠 Extensions
SKILLS.md— Skills frameworkMEMORY.md— Memory systemEVALS.md— Eval frameworkGUARDRAILS.md— PII / injection / visionWEBHOOKS.md— Webhook deliveryCOMPLIANCE.md— Audit + retentionAUTHZ_GUIDE.md— Authorization pipeline
🗜️ Compression
COMPRESSION_GUIDE.mdCOMPRESSION_ENGINES.mdCOMPRESSION_RULES_FORMAT.mdCOMPRESSION_LANGUAGE_PACKS.mdRTK_COMPRESSION.md
🚀 Deployment
DOCKER_GUIDE.mdVM_DEPLOYMENT_GUIDE.mdFLY_IO_DEPLOYMENT_GUIDE.mdELECTRON_GUIDE.mdPWA_GUIDE.mdTERMUX_GUIDE.mdTUNNELS_GUIDE.mdPROXY_GUIDE.md
📋 Operations
RELEASE_CHECKLIST.md— release flow with Claude Code skillsCOVERAGE_PLAN.md— test coverage state (current: 82.58%/82.58%/84.23%/75.22%)I18N.md— 30 supported localesUNINSTALL.md
🤝 Contributing & policy
CONTRIBUTING.md— contributor guideSECURITY.md— security policyCODE_OF_CONDUCT.mdCLAUDE.md— rules for Claude Code agentsAGENTS.md— rules for non-Claude agentsGEMINI.md— rules for Gemini agents
💡 Use Cases
| Scenario | Solution |
|---|---|
| "Claude Pro user, hit rate limit" | Combo: Claude → GLM → DeepSeek (auto-fallback) |
| "Want $0 forever" | auto/cheap → Kiro/Qoder/Pollinations fallback chain |
| "24/7 coding, no interruptions" | auto/lkgp (sticky to last-good) + Resilience |
| "Blocked region" | 1proxy free marketplace + Cloudflare Quick Tunnel |
| "Max token savings" | Stacked compression: rtk → caveman (78-95% on logs) |
| "Multi-agent system" | Expose OmniRoute as A2A node, route via smart-routing skill |
| "Long-running coding task" | Cloud Agents → Devin/Jules with management auth |
→ Detailed playbooks: USER_GUIDE.md · AUTO-COMBO.md
📡 Protocols supported
OmniRoute speaks all major AI protocols — clients don't need to change:
- OpenAI (Chat Completions, Responses, Embeddings, Images, Audio, Files, Batches, Rerank, Moderations)
- Anthropic Messages (Claude format, with thinking blocks + reasoning replay)
- Google Gemini (generateContent + Vertex)
- Claude Code (CLI-specific format with CCH + fingerprinting)
- Cursor (proprietary format with tool calls)
- Kiro (AWS Builder ID OAuth)
- MCP (Model Context Protocol — 37 tools, stdio/SSE/Streamable HTTP)
- A2A (Agent-to-Agent v0.3 JSON-RPC — agent card at
/.well-known/agent.json)
🏗️ Architecture (10-second tour)
Client → /v1/chat/completions → [CORS → Zod → Auth → Authz → Guardrails]
→ handleChatCore() → [Cache → Rate limit → Combo routing]
→ translateRequest → getExecutor → fetch upstream (with retry)
→ response translation → SSE stream or JSON
→ [Compliance audit] → response
Major pieces:
src/app/— Next.js 16 App Router (60+ API routes + 30 dashboard pages)src/lib/— 50+ domain modules (db, a2a, memory, skills, guardrails, evals, …)open-sse/— Streaming engine workspace (31 executors, 9+8+9 translators, 80+ services, 37-tool MCP server)src/domain/— Pure business logic (policies, fallback, cost rules)src/server/— Server-only (authz pipeline, cors)
→ Deep dive: docs/architecture/ARCHITECTURE.md · docs/architecture/CODEBASE_DOCUMENTATION.md
🌍 i18n
UI translated to 30 languages with full RTL support for Arabic and Hebrew.
🌐 English · Português · Español · Français · Deutsch · 中文 · 日本語 · 한국어 · العربية · हिन्दी · Русский · + 19 more
→ Adding a language: docs/guides/I18N.md
🤝 Community
- 🌐 Website: omniroute.online
- 📦 npm: omniroute
- 🐳 Docker Hub: diegosouzapw/omniroute
- 💬 WhatsApp (BR): Brazilian community group — see README link
- 🐛 Issues: GitHub Issues
- 💡 Discussions: GitHub Discussions
❤️ Contributing
We welcome PRs! Start with:
- Read
CONTRIBUTING.md— setup, conventional commits, testing - Pick an issue labeled
good first issue - Branch from
main(feat/*,fix/*,docs/*,refactor/*,test/*,chore/*) - Hooks will run lint + test on commit/push
Adding a provider? docs/architecture/ARCHITECTURE.md § Adding a New Provider
Adding an MCP tool? docs/frameworks/MCP-SERVER.md
Adding an A2A skill? docs/frameworks/A2A-SERVER.md § Adding a New Skill
🔒 Security
- Reporting: see
SECURITY.mdfor disclosure policy - Supported versions: 3.8.x (Active), 3.7.x (Security only)
- Secrets: never commit. Use
.env(auto-generated from.env.exampleon first install) or vaults - Encryption: credentials at rest with AES-256-GCM
- Authz: route-aware classification (
src/server/authz/) — seedocs/architecture/AUTHZ_GUIDE.md - Guardrails: PII masking, prompt injection detection — hot-reloadable
📄 License
MIT © 2025-2026 Diego Souza
Free forever. Self-hosted. No tracking. No cloud lock-in.
⬆ Back to top · Built with ❤️ for the open-source AI community.
OmniRoute v3.8.0 · Node ≥20.20.2 · MIT License






