diff --git a/README.md b/README.md index ed089adfec..13853697cc 100644 --- a/README.md +++ b/README.md @@ -1,11 +1,11 @@ - - + + - + @@ -13,7 +13,7 @@ - + @@ -23,1854 +23,471 @@ "@type": "SoftwareApplication", "name": "OmniRoute", "applicationCategory": "DeveloperApplication", - "operatingSystem": "Windows, macOS, Linux, Android, Termux", - "offers": { - "@type": "Offer", - "price": "0", - "priceCurrency": "USD" - }, - "description": "Unified AI proxy/router with 160+ providers, RTK+Caveman compression (15-95% savings), Auto-Combo intelligent routing to FREE & low-cost AI models.", + "operatingSystem": "Windows, macOS, Linux, Android (Termux)", + "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" }, + "description": "Unified AI proxy/router with 177+ providers, Auto-Combo 9-factor routing, RTK+Caveman compression, MCP/A2A protocols, and auto-fallback to FREE models.", "url": "https://github.com/diegosouzapw/OmniRoute", - "softwareVersion": "3.7.8", - "author": { - "@type": "Person", - "name": "Diego Souza" - }, + "softwareVersion": "3.8.0", + "author": { "@type": "Person", "name": "Diego Souza" }, "license": "https://github.com/diegosouzapw/OmniRoute/blob/main/LICENSE", "downloadUrl": "https://www.npmjs.com/package/omniroute", - "installUrl": "https://www.npmjs.com/package/omniroute", "features": [ - "160+ AI providers", - "Prompt compression (RTK + Caveman)", - "Auto-Combo 7-factor routing", - "Auto-fallback chain", - "MCP Server (37 tools)", - "A2A Protocol", - "40+ languages" - ] -} - - - -
-# ๐Ÿš€ OmniRoute โ€” The Free AI Gateway +# ๐Ÿš€ OmniRoute -### Never stop coding. Save 15-95% eligible tokens with RTK+Caveman compression + **Auto-Combo** intelligent routing to **FREE & low-cost AI models**. +### **The Free AI Gateway โ€” one endpoint, 177+ providers, zero downtime.** -_Auto-Combo uses 7-factor scoring (health, quota, cost, latency, capability, stability, tier) to automatically pick the best model for each request with self-healing fallback._ +**Auto-fallback to free models. Stop coding interruptions. Cut tokens 15-95%.** -_The most complete open-source AI proxy โ€” **one endpoint**, **160+ providers**, **13 routing strategies**, zero downtime. Multi-platform: **Web**, **Desktop (Electron)**, **Mobile (PWA + Termux)**. Fully extensible via **MCP Server (37 tools)**, **A2A Protocol**, and **Memory/Skills** systems. Available in **40+ languages**._ +[![npm](https://img.shields.io/npm/v/omniroute?logo=npm&style=flat-square)](https://www.npmjs.com/package/omniroute) +[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE) +[![Node](https://img.shields.io/badge/node-%E2%89%A520.20.2-brightgreen?style=flat-square)](package.json) +[![Stars](https://img.shields.io/github/stars/diegosouzapw/OmniRoute?style=social)](https://github.com/diegosouzapw/OmniRoute) +[![Trendshift](https://trendshift.io/api/badge/repositories/23589)](https://trendshift.io/repositories/23589) -**Chat Completions โ€ข Responses API โ€ข Embeddings โ€ข Image Generation โ€ข Video โ€ข Music โ€ข Audio Speech/Transcription โ€ข Reranking โ€ข Moderations โ€ข Web Search โ€ข MCP Server โ€ข A2A Protocol โ€ข 4,600+ Tests โ€ข 100% TypeScript** +[**Website**](https://omniroute.online) ยท [**Quick Start**](#-quick-start) ยท [**Docs**](#-documentation) ยท [**Discord/WhatsApp**](#-community) -
- - - Get $100 Free AI Credits - - -๐Ÿ”ฅ Limited offer: Sign up at AgentRouter and get $100 in free AI credits
Access GPT-5, Claude, Gemini, DeepSeek & 100+ models. No credit card required. Claim your credits โ†’
- -
- -diegosouzapw%2FOmniRoute | Trendshift - -[๐Ÿš€ Quick Start](#-quick-start) โ€ข [๐Ÿ’ก Features](#-key-features) โ€ข [๐Ÿ—œ๏ธ Compression](#%EF%B8%8F-prompt-compression--save-15-95-eligible-tokens-automatically) โ€ข [๐Ÿ’ฐ Pricing](#-pricing-at-a-glance) โ€ข [๐ŸŽฏ Use Cases](#-use-cases--ready-made-combo-playbooks) โ€ข [๐ŸŒ Proxy](#-bypass-geographic-blocks--use-ai-from-any-country) โ€ข [โ“ FAQ](#-frequently-asked-questions) โ€ข [๐Ÿ“– Docs](#-documentation) โ€ข [๐Ÿ’ฌ WhatsApp](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) +v3.8.0 ยท MIT ยท Production-ready ยท Self-hosted
--- -
+## โšก The Pitch (60 seconds) -๐ŸŒ **Available in:** ๐Ÿ‡บ๐Ÿ‡ธ [English](README.md) | ๐Ÿ‡ง๐Ÿ‡ท [Portuguรชs (Brasil)](docs/i18n/pt-BR/README.md) | ๐Ÿ‡ช๐Ÿ‡ธ [Espaรฑol](docs/i18n/es/README.md) | ๐Ÿ‡ซ๐Ÿ‡ท [Franรงais](docs/i18n/fr/README.md) | ๐Ÿ‡ฎ๐Ÿ‡น [Italiano](docs/i18n/it/README.md) | ๐Ÿ‡ท๐Ÿ‡บ [ะ ัƒััะบะธะน](docs/i18n/ru/README.md) | ๐Ÿ‡จ๐Ÿ‡ณ [ไธญๆ–‡ (็ฎ€ไฝ“)](docs/i18n/zh-CN/README.md) | ๐Ÿ‡ฉ๐Ÿ‡ช [Deutsch](docs/i18n/de/README.md) | ๐Ÿ‡ฎ๐Ÿ‡ณ [เคนเคฟเคจเฅเคฆเฅ€](docs/i18n/in/README.md) | ๐Ÿ‡น๐Ÿ‡ญ [เน„เธ—เธข](docs/i18n/th/README.md) | ๐Ÿ‡บ๐Ÿ‡ฆ [ะฃะบั€ะฐั—ะฝััŒะบะฐ](docs/i18n/uk-UA/README.md) | ๐Ÿ‡ธ๐Ÿ‡ฆ [ุงู„ุนุฑุจูŠุฉ](docs/i18n/ar/README.md) | ๐Ÿ‡ฏ๐Ÿ‡ต [ๆ—ฅๆœฌ่ชž](docs/i18n/ja/README.md) | ๐Ÿ‡ป๐Ÿ‡ณ [Tiแบฟng Viแป‡t](docs/i18n/vi/README.md) | ๐Ÿ‡ง๐Ÿ‡ฌ [ะ‘ัŠะปะณะฐั€ัะบะธ](docs/i18n/bg/README.md) | ๐Ÿ‡ฉ๐Ÿ‡ฐ [Dansk](docs/i18n/da/README.md) | ๐Ÿ‡ซ๐Ÿ‡ฎ [Suomi](docs/i18n/fi/README.md) | ๐Ÿ‡ฎ๐Ÿ‡ฑ [ืขื‘ืจื™ืช](docs/i18n/he/README.md) | ๐Ÿ‡ญ๐Ÿ‡บ [Magyar](docs/i18n/hu/README.md) | ๐Ÿ‡ฎ๐Ÿ‡ฉ [Bahasa Indonesia](docs/i18n/id/README.md) | ๐Ÿ‡ฐ๐Ÿ‡ท [ํ•œ๊ตญ์–ด](docs/i18n/ko/README.md) | ๐Ÿ‡ฒ๐Ÿ‡พ [Bahasa Melayu](docs/i18n/ms/README.md) | ๐Ÿ‡ณ๐Ÿ‡ฑ [Nederlands](docs/i18n/nl/README.md) | ๐Ÿ‡ณ๐Ÿ‡ด [Norsk](docs/i18n/no/README.md) | ๐Ÿ‡ต๐Ÿ‡น [Portuguรชs (Portugal)](docs/i18n/pt/README.md) | ๐Ÿ‡ท๐Ÿ‡ด [Romรขnฤƒ](docs/i18n/ro/README.md) | ๐Ÿ‡ต๐Ÿ‡ฑ [Polski](docs/i18n/pl/README.md) | ๐Ÿ‡ธ๐Ÿ‡ฐ [Slovenฤina](docs/i18n/sk/README.md) | ๐Ÿ‡ธ๐Ÿ‡ช [Svenska](docs/i18n/sv/README.md) | ๐Ÿ‡ต๐Ÿ‡ญ [Filipino](docs/i18n/phi/README.md) | ๐Ÿ‡จ๐Ÿ‡ฟ [ฤŒeลกtina](docs/i18n/cs/README.md) +| ๐ŸŽฏ The problem | โœ… How OmniRoute solves it | +| ------------------------------ | -------------------------------------------------------------------------------------------------- | +| Hit rate limits on Claude/GPT? | **Auto-fallback** across 177 providers โ€” never see a 429 again | +| Bored of switching API keys? | **One endpoint** (`localhost:20128`) speaks OpenAI, Anthropic, Gemini, Claude Code, Cursor formats | +| Paying $200/mo for AI? | **11 free providers** + intelligent routing โ†’ most users pay $0 | +| Tokens too expensive? | **RTK + Caveman compression** saves 15-95% on eligible payloads | +| Blocked region? | **4-level proxy** (account/provider/combo/global) + **1proxy free marketplace** | +| Want CLI agents free? | Plug **Cursor, Cline, Codex, Claude Code, Aider, 15+ CLIs** at OmniRoute | -
- -
- -[![npm version](https://img.shields.io/npm/v/omniroute?color=cb3837&logo=npm)](https://www.npmjs.com/package/omniroute) -![NPM Weekly](https://img.shields.io/npm/dw/omniroute?label=npm/week&color=cb3837&logo=npm) -![NPM Monthly](https://img.shields.io/npm/dm/omniroute?label=npm/month&color=cb3837&logo=npm) -![NPM Yearly](https://img.shields.io/npm/d18m/omniroute?label=npm/year&color=cb3837&logo=npm) - -[![Docker Hub](https://img.shields.io/docker/v/diegosouzapw/omniroute?label=Docker%20Hub&logo=docker&color=2496ED)](https://hub.docker.com/r/diegosouzapw/omniroute) -![Docker Pulls](https://img.shields.io/docker/pulls/diegosouzapw/omniroute?label=docker%20pulls&logo=docker&color=2496ED) -![Electron Downloads](https://img.shields.io/github/downloads/diegosouzapw/omniroute/total?style=flat&label=electron%20downloads&logo=electron&color=47848F) -[![license](https://custom-icon-badges.demolab.com/github/license/diegosouzapw/OmniRoute?logo=law)](https://github.com/diegosouzapw/OmniRoute/blob/main/LICENSE) - - - -[![total contributions](https://custom-icon-badges.demolab.com/badge/dynamic/json?logo=graph&logoColor=fff&color=blue&label=total%20contributions&query=%24.totalContributions&url=https%3A%2F%2Fstreak-stats.demolab.com%2F%3Fuser%3Ddiegosouzapw%26type%3Djson)](https://github.com/diegosouzapw) -[![github streak](https://custom-icon-badges.demolab.com/badge/dynamic/json?logo=fire&logoColor=fff&color=orange&label=github%20streak&query=%24.currentStreak.length&suffix=%20days&url=https%3A%2F%2Fstreak-stats.demolab.com%2F%3Fuser%3Ddiegosouzapw%26type%3Djson)](https://github.com/diegosouzapw) -[![Website](https://img.shields.io/badge/Website-omniroute.online-blue?logo=google-chrome&logoColor=white)](https://omniroute.online) -[![WhatsApp](https://img.shields.io/badge/WhatsApp-Community-25D366?logo=whatsapp&logoColor=white)](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) - -
+๐Ÿ“บ **Watch in action:** [Video demo](https://www.youtube.com/@diegosouza-pw) --- -## ๐Ÿ–ผ๏ธ Main Dashboard +## ๐Ÿ–ผ๏ธ Dashboard Preview -
- OmniRoute Dashboard -
- ---- - -## ๐Ÿ“ธ Dashboard Preview +![OmniRoute Main Dashboard](docs/screenshots/MainOmniRoute.png)
-Click to see dashboard screenshots +Click for more screenshots (Providers, Combos, Memory, MCP, Auditโ€ฆ) -| Page | Screenshot | -| -------------- | ------------------------------------------------- | -| **Providers** | ![Providers](docs/screenshots/01-providers.png) | -| **Combos** | ![Combos](docs/screenshots/02-combos.png) | -| **Analytics** | ![Analytics](docs/screenshots/03-analytics.png) | -| **Health** | ![Health](docs/screenshots/04-health.png) | -| **Translator** | ![Translator](docs/screenshots/05-translator.png) | -| **Settings** | ![Settings](docs/screenshots/06-settings.png) | -| **CLI Tools** | ![CLI Tools](docs/screenshots/07-cli-tools.png) | -| **Usage Logs** | ![Usage](docs/screenshots/08-usage.png) | -| **Endpoints** | ![Endpoints](docs/screenshots/09-endpoint.png) | +| | | +| --------------------------------------------- | ---------------------------------------------------- | +| ![Providers](docs/screenshots/Provedores.png) | ![Combos](docs/screenshots/Combos.png) | +| ![Routing](docs/screenshots/Roteamento.png) | ![Compression](docs/screenshots/Compress%C3%A3o.png) | +| ![Audit](docs/screenshots/Auditoria.png) | ![Analytics](docs/screenshots/Analytics.png) |
--- -### ๐Ÿค– Free AI Provider for your favorite coding agents - -_Connect any AI-powered IDE or CLI tool through OmniRoute โ€” free API gateway for unlimited coding._ - - - - - - - - - - - - - - - - -
- - OpenClaw
- OpenClaw -

- โญ 205K -
- - NanoBot
- NanoBot -

- โญ 20.9K -
- - PicoClaw
- PicoClaw -

- โญ 14.6K -
- - ZeroClaw
- ZeroClaw -

- โญ 9.9K -
- - IronClaw
- IronClaw -

- โญ 2.1K -
- - OpenCode
- OpenCode -

- โญ 106K -
- - Codex CLI
- Codex CLI -

- โญ 60.8K -
- - Claude Code
- Claude Code -

- โญ 67.3K -
- - Gemini CLI
- Gemini CLI -

- โญ 94.7K -
- - Kilo Code
- Kilo Code -

- โญ 15.5K -
- -๐Ÿ“ก All agents connect via http://localhost:20128/v1 or http://cloud.omniroute.online/v1 โ€” one config, unlimited models and quota - ---- - -## ๐Ÿ“บ OmniRoute in Action โ€” Video Guides - -
- - - - - - - -
- - OmniRoute โ€” Guia em Portuguรชs -
- ๐Ÿ‡ง๐Ÿ‡ท Portuguรชs
- Guia completo do OmniRoute -
- - OmniRoute โ€” English Guide -
- ๐Ÿ‡บ๐Ÿ‡ธ English
- Complete OmniRoute walkthrough -
- - OmniRoute โ€” ะ ัƒะบะพะฒะพะดัั‚ะฒะพ ะฝะฐ ั€ัƒััะบะพะผ -
- ๐Ÿ‡ท๐Ÿ‡บ ะ ัƒััะบะธะน
- ะŸะพะปะฝะพะต ั€ัƒะบะพะฒะพะดัั‚ะฒะพ ะฟะพ OmniRoute -
- -
- -> ๐ŸŽฌ **Made a video about OmniRoute?** We'd love to feature it here! Open an [issue](https://github.com/diegosouzapw/OmniRoute/issues/new) or [discussion](https://github.com/diegosouzapw/OmniRoute/discussions) with the link and we'll add it to this showcase. - ---- - -## ๐Ÿค” Why OmniRoute? - -**Stop wasting money, tokens and hitting limits:** - -โŒ Subscription quota expires unused every month -โŒ Rate limits stop you mid-coding -โŒ Tool outputs (`git diff`, `grep`, `ls`...) burn tokens fast -โŒ Expensive APIs ($20-50/month per provider) -โŒ Manual switching between providers -โŒ Each provider has a different API format -โŒ AI providers blocked in your country - -**OmniRoute solves all of this:** - -โœ… **Prompt Compression** โ€” auto-compress prompts & tool outputs, save 15-95% eligible tokens per request with RTK+Caveman stacked mode -โœ… **Maximize subscriptions** โ€” track quota, use every bit before reset -โœ… **Auto fallback** โ€” Subscription โ†’ API Key โ†’ Cheap โ†’ Free, zero downtime -โœ… **Multi-account** โ€” round-robin between accounts per provider -โœ… **Format translation** โ€” OpenAI โ†” Claude โ†” Gemini โ†” Responses API, any tool works -โœ… **3-level proxy** โ€” bypass geo-blocks with global, per-provider, and per-key proxies -โœ… **10 multi-modal APIs** โ€” chat, images, video, music, audio, search in one endpoint -โœ… **MCP + A2A** โ€” 29 MCP tools + agent-to-agent protocol, production-ready -โœ… **Universal** โ€” works with Claude Code, Codex, Gemini CLI, Cursor, Cline, OpenClaw, any CLI tool - -### Why OmniRoute Wins - -| Capability | OmniRoute | LiteLLM | Bifrost | Routerly | -| ------------------ | :----------: | :-----: | :-----: | :------: | -| **Providers** | **160+** | 100+ | 15+ | 8 | -| **Free Providers** | **11** | 1 | 0 | 0 | -| **Token Savings** | **15-95%** | โŒ | โŒ | โŒ | -| **Auto-Combo** | **6-factor** | โŒ | โŒ | Basic | -| **MCP Tools** | **37** | 0 | 1 | 0 | -| **A2A Protocol** | **โœ…** | โŒ | โŒ | โŒ | -| **Desktop App** | **โœ…** | โŒ | โŒ | โŒ | -| **i18n** | **40+** | โŒ | โŒ | โŒ | -| **Open Source** | **โœ…** | โœ… | โŒ | โŒ | -| **Self-Hosted** | **Free** | $50/mo | $19/mo | $9.99/mo | - ---- - -## ๐ŸŽฏ Auto-Combo โ€” Intelligent 6-Factor Routing - -> **Auto-Combo uses AI-powered scoring to automatically pick the best model for every request.** No manual combo configuration needed โ€” it learns from your usage patterns and self-heals when providers fail. - -### How It Works - -Auto-Combo evaluates each request against **7 scoring factors**: - -| Factor | Description | Weight | -| ---------------------- | ------------------------------ | ------ | -| **Health** | Provider uptime and error rate | 25% | -| **Quota Availability** | Remaining rate limits | 20% | -| **Cost Efficiency** | Price per 1M tokens | 20% | -| **Latency** | Historical response time | 15% | -| **Capability Match** | Model strengths for task | 10% | -| **Stability** | Avoid same-provider clustering | 5% | -| **Tier Priority** | Subscription tier preference | 5% | - -### Mode Packs - -Auto-Combo includes pre-configured **mode packs** optimized for different use cases: - -- **๐Ÿ–ฅ๏ธ Coding Pack** โ€” prioritizes reasoning + code generation models (Claude, DeepSeek, Qwen) -- **๐Ÿ‘๏ธ Vision Pack** โ€” multimodal models with image understanding (GPT-5, Claude, Gemini) -- **๐Ÿ“Š Analysis Pack** โ€” complex reasoning + math capabilities (o3, DeepSeek-R1, GLM) -- **๐Ÿ’ฌ Chat Pack** โ€” balanced speed + quality for conversation (Claude Haiku, Gemini Flash) - -### Self-Healing - -When a provider fails (rate limit, outage, error), Auto-Combo automatically: - -1. Detects the failure within seconds -2. Scores all remaining options -3. Reroutes to the next best provider -4. Updates its scoring model for future requests - -### Auto-Combo vs Manual Combos - -| Feature | Auto-Combo | Manual Combo | -| -------------------- | ---------------------------------------------------- | --------------------- | -| **Configuration** | Zero config โ€” works out of the box | Requires manual setup | -| **Adaptability** | Real-time scoring based on conditions | Static priority list | -| **Self-Healing** | Automatic rerouting on failure | Manual fallback only | -| **6-Factor Scoring** | โœ… Task, cost, latency, quota, capability, diversity | โŒ | -| **Mode Packs** | โœ… Coding, Vision, Analysis, Chat | โŒ | -| **Learning** | Improves over time based on usage | Static | - -### Competitor Comparison - -| Feature | OmniRoute Auto-Combo | Routerly | -| ---------------------- | ---------------------- | --------------------- | -| **6-Factor Scoring** | โœ… | โŒ (simple LLM-based) | -| **Self-Healing** | โœ… Automatic rerouting | โš ๏ธ Basic fallback | -| **Mode Packs** | โœ… 4 optimized packs | โŒ | -| **Free Providers** | โœ… 11 unlimited | โš ๏ธ 3 limited | -| **Prompt Compression** | โœ… RTK+Caveman 15-95% | โŒ | -| **Price** | **Free** (open-source) | $9.99/mo | - -๐Ÿ“– **Full Auto-Combo documentation:** [docs/AUTO-COMBO.md](docs/AUTO-COMBO.md) - ---- - -## ๐Ÿ“ง Support - -> ๐Ÿ’ฌ **Join our community!** [WhatsApp Group](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) โ€” Get help, share tips, and stay updated. - -- **Website**: [omniroute.online](https://omniroute.online) -- **GitHub**: [github.com/diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) -- **Issues**: [github.com/diegosouzapw/OmniRoute/issues](https://github.com/diegosouzapw/OmniRoute/issues) -- **WhatsApp**: [Community Group](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) -- ๐Ÿ‡ง๐Ÿ‡ท **Official Brazilian WhatsApp Group**: [Community Group](https://chat.whatsapp.com/CeGCxdFzqBe5Uki288wOvf) -- **Contributing**: See [CONTRIBUTING.md](CONTRIBUTING.md), open a PR, or pick a `good first issue` -- **Original Project**: [9router by decolua](https://github.com/decolua/9router) - -### ๐Ÿ› Reporting a Bug? - -When opening an issue, please run the system-info command and attach the generated file: - -```bash -npm run system-info -``` - -This generates a `system-info.txt` with your Node.js version, OmniRoute version, OS details, installed CLI tools (qoder, gemini, claude, codex, antigravity, droid, etc.), Docker/PM2 status, and system packages โ€” everything we need to reproduce your issue quickly. Attach the file directly to your GitHub issue. - ---- - -## ๐Ÿ› ๏ธ Supported CLI Tools - -OmniRoute works seamlessly with **16+ AI coding tools** โ€” one config, all tools: - - - - - - - - - - - - - - - - - - - - - - - - - - -
Claude Code
Anthropic
Codex CLI
OpenAI
Gemini CLI
Google
Cursor
IDE
OpenClaw
CLI
Antigravity
VS Code
Cline
Extension
Continue
Extension
Kilo Code
Extension
Kiro
AWS IDE
OpenCode
CLI
Droid
CLI
AMP
CLI
Copilot
GitHub
Windsurf
IDE
Hermes
CLI
Qwen CLI
Alibaba
Custom
Any tool
- -๐Ÿ“– Full setup for each tool: [`docs/CLI-TOOLS.md`](docs/CLI-TOOLS.md) - ---- - -## ๐ŸŒ Supported Providers โ€” 160+ - -### ๐Ÿ” OAuth Providers - - - - - - - - - - - - - - - -
Claude Code
Anthropic OAuth
Antigravity
Google OAuth
Codex
OpenAI OAuth
GitHub Copilot
GitHub OAuth
Cursor
Cursor OAuth
Kimi Coding
Moonshot OAuth
Kilo Code
Kilo OAuth
Cline
Cline OAuth
- -### ๐Ÿ†“ Free Providers (No Cost) - - - - - - - - - - - - - - -
๐ŸŸข Kiro AI
Claude Sonnet/Haiku
Unlimited FREE
๐ŸŸข Qoder AI
Kimi-K2, DeepSeek-R1
Unlimited FREE
๐ŸŸข Pollinations
GPT-5, Claude, Llama 4
No API key needed
๐ŸŸข Qwen Code
Qwen3 Coder Plus
Unlimited FREE
๐ŸŸข LongCat AI
Flash-Lite
50M tokens/day
๐ŸŸข Cloudflare AI
50+ models
10K neurons/day
๐ŸŸข Puter AI
GPT-4.1, Claude
Rate-limited free
๐ŸŸข NVIDIA NIM
Llama, Mistral
1K req/day free
- -### ๐Ÿ”‘ API Key Providers (120+) - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
OpenAIAnthropicGeminiDeepSeekGroqxAI (Grok)
MistralOpenRouterGLMKimiMiniMaxFireworks
Together AICerebrasCohereNVIDIAPerplexitySiliconFlow
NebiusHuggingFaceDeepInfraSambaNovaVertex AIAzure OpenAI
AWS BedrockSnowflakeDatabricksVenice.aiAI21 LabsMeta Llama
- -
-...and 90+ more providers - -Alibaba ยท Amazon Q ยท AssemblyAI ยท Baidu Qianfan ยท Baseten ยท Black Forest Labs ยท Blackbox ยท Brave Search ยท Bytez ยท CablyAI ยท Cartesia ยท ChatGPT Web ยท Chutes.ai ยท Clarifai ยท Codestral ยท CrofAI ยท DataRobot ยท Deepgram ยท ElevenLabs ยท Empower ยท Exa Search ยท Fal.ai ยท Featherless AI ยท FenayAI ยท FriendliAI ยท Galadriel ยท GigaChat ยท GitLab Duo ยท GLHF Chat ยท GoAPI ยท Heroku AI ยท Hyperbolic ยท IBM watsonx ยท Inference.net ยท Inworld ยท Jina AI ยท Kilo Gateway ยท Lambda AI ยท LaoZhang ยท Linkup Search ยท LlamaGate ยท Maritalk ยท Modal ยท Moonshot AI ยท Morph ยท Muse Spark ยท NanoBanana ยท NanoGPT ยท NLP Cloud ยท Nous Research ยท Novita AI ยท nScale ยท OCI ยท Ollama Cloud ยท OVHcloud ยท PiAPI ยท PlayHT ยท Poe ยท Predibase ยท PublicAI ยท Qwen Code ยท Recraft ยท Reka ยท Runway ยท SAP ยท Scaleway ยท SearchAPI ยท SearXNG ยท Serper ยท Stability AI ยท Synthetic ยท Tavily ยท TheB.AI ยท Topaz ยท Upstage ยท v0 (Vercel) ยท Vercel AI Gateway ยท Volcengine ยท Voyage AI ยท W&B Inference ยท Xiaomi MiMo ยท You.com ยท Z.AI ยท + OpenAI/Anthropic-compatible custom endpoints - -
- -### ๐Ÿ  Self-Hosted - - - - - - - - - - - - - - - - -
LM StudioOllamavLLMLlamafileDocker Model Runner
NVIDIA TritonXInferenceoobaboogaComfyUISD WebUI
- ---- - -## ๐Ÿ”„ How It Works - -``` -โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” -โ”‚ Your CLI โ”‚ (Claude Code, Codex, Gemini CLI, OpenClaw, Cursor, Cline...) -โ”‚ Tool โ”‚ -โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”˜ - โ”‚ http://localhost:20128/v1 - โ†“ -โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” -โ”‚ OmniRoute (Smart Router) โ”‚ -โ”‚ โ€ข ๐Ÿ—œ๏ธ Prompt Compression (save 15-95% eligible) โ”‚ -โ”‚ โ€ข Format translation (OpenAI โ†” Claude โ†” Gemini) โ”‚ -โ”‚ โ€ข Quota tracking + Embeddings + Images โ”‚ -โ”‚ โ€ข Auto token refresh + Rate limit management โ”‚ -โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ - โ”‚ - โ”œโ”€โ†’ [Tier 1: SUBSCRIPTION] Claude Code, Codex, Gemini CLI - โ”‚ โ†“ quota exhausted - โ”œโ”€โ†’ [Tier 2: API KEY] DeepSeek, Groq, xAI, Mistral, NVIDIA NIM, etc. - โ”‚ โ†“ budget limit - โ”œโ”€โ†’ [Tier 3: CHEAP] GLM ($0.6/1M), MiniMax ($0.2/1M) - โ”‚ โ†“ budget limit - โ””โ”€โ†’ [Tier 4: FREE] Qoder, Qwen, Kiro (unlimited) - -Result: Never stop coding, minimal cost + 15-95% eligible token savings -``` - ---- - -## ๐Ÿ—œ๏ธ Prompt Compression โ€” Save 15-95% Eligible Tokens Automatically - -> **Why use many token when few token do trick?** OmniRoute's built-in compression pipeline reduces token usage before requests reach the provider. It combines ideas from [RTK - Rust Token Killer](https://github.com/rtk-ai/rtk) and [Caveman](https://github.com/JuliusBrussee/caveman) (โญ 51K+). - -### How It Works - -Every request passes through the compression pipeline **transparently** โ€” no client changes needed: - -``` -โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” -โ”‚ Client sends โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚ OmniRoute Compression โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚ Provider โ”‚ -โ”‚ full prompt โ”‚ โ”‚ Pipeline (7 options) โ”‚ โ”‚ receives โ”‚ -โ”‚ (10,000 tok) โ”‚ โ”‚ โ”‚ โ”‚ compressed โ”‚ -โ”‚ โ”‚ โ”‚ ๐Ÿชถ Lite ........... ~15% โ”‚ โ”‚ (~1,080 tok)โ”‚ -โ”‚ โ”‚ โ”‚ ๐Ÿชจ Standard ....... ~30% โ”‚ โ”‚ โ”‚ -โ”‚ โ”‚ โ”‚ โšก Aggressive ..... ~50% โ”‚ โ”‚ ๐Ÿ’ฐ up to 95%โ”‚ -โ”‚ โ”‚ โ”‚ ๐Ÿ”ฅ Ultra .......... ~75% โ”‚ โ”‚ โ”‚ -โ”‚ โ”‚ โ”‚ ๐Ÿงฐ RTK ............ 60-90% โ”‚ โ”‚ โ”‚ -โ”‚ โ”‚ โ”‚ ๐Ÿ”— Stacked ........ 78-95% โ”‚ โ”‚ โ”‚ -โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ -``` - -### 7 Compression Options - -| Mode | Savings | Technique | Best For | -| ------------------------- | ------- | ----------------------------------------------------------------------------------------------- | -------------------------------------- | -| **Off** | 0% | No compression | When you need exact prompts | -| **๐Ÿชถ Lite** | ~15% | Whitespace collapse, dedup system prompts, image URL shortening | Always-on safe default | -| **๐Ÿชจ Standard (Caveman)** | ~30% | 30+ regex rules: filler removal, context condensation, structural compression, multi-turn dedup | Daily coding with Claude/Codex | -| **โšก Aggressive** | ~50% | All standard + progressive message aging + tool result summarization + LLM-based compression | Long sessions with many tool calls | -| **๐Ÿ”ฅ Ultra** | ~75% | All aggressive + heuristic token pruning + stopword removal + score-based filtering | Maximum savings when tokens are scarce | -| **๐Ÿงฐ RTK** | 60-90% | 49 command-aware filters, RTK-style JSON DSL, verify gate, trust-gated custom filters | Shell/test/build/git output in agents | -| **๐Ÿ”— Stacked** | 78-95% | RTK first, then Caveman input condensation; ~89% with upstream average math | Mixed prompts with tool logs + prose | - -### RTK + Caveman Savings Math - -These numbers are based on the upstream project READMEs under `_references/_outros`: - -| Source | Upstream claim used by OmniRoute docs | -| ------- | ------------------------------------------------------------------------------------------------------------------- | -| Caveman | `~75%` fewer output tokens; benchmark average `65%` output savings, range `22-87%`; `~46%` input compression tool | -| RTK | `60-90%` command-output token savings; sample session `~118,000 -> ~23,900` tokens, which is `79.7%` saved (`~80%`) | - -For the default stacked compression combo, OmniRoute runs: - -```txt -RTK -> Caveman -``` - -When both engines can act on the same tool/context payload, the savings compound: - -```txt -combined = 1 - (1 - RTK savings) * (1 - Caveman input savings) -average = 1 - (1 - 0.80) * (1 - 0.46) = 89.2% -range = 1 - (1 - 0.60..0.90) * (1 - 0.46) = 78.4-94.6% -``` - -Caveman output mode is separate from prompt compression. When enabled for responses, use Caveman's -own upstream output numbers: `65%` average, `~75%` headline, `22-87%` observed range. Total bill -savings depend on the prompt/output mix, but coding-agent sessions are often tool-context heavy, so -the `RTK -> Caveman` combo is the best default for maximum context savings. - -### Before & After (Standard/Caveman Mode) - -**๐Ÿ—ฃ๏ธ Before compression (69 tokens):** - -> "The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I would recommend using useMemo to memoize the object." - -**๐Ÿชจ After compression (19 tokens):** - -> "New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo." - -**Same answer. 72% less tokens. Zero accuracy loss.** - -### Architecture - -``` -Request Body - โ”‚ - โ”œโ”€ strategySelector.ts โ”€โ”€โ”€ Picks mode (config / combo override / auto-trigger) - โ”‚ - โ”œโ”€ lite.ts โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whitespace, dedup, image URLs, redundant content - โ”œโ”€ caveman.ts โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ 30+ regex rules via cavemanRules.ts - โ”‚ โ””โ”€ preservation.ts โ”€โ”€โ”€ Protects code blocks, URLs, JSON from compression - โ”œโ”€ engines/rtk/ โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Command detection + JSON DSL filters + raw-output recovery - โ”œโ”€ engines/registry.ts โ”€โ”€โ”€ Shared engine registry for caveman, RTK, and stacked - โ”œโ”€ aggressive.ts โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Summarizer + tool result compressor + progressive aging - โ”‚ โ”œโ”€ summarizer.ts โ”€โ”€โ”€โ”€โ”€ Rule-based message summarization - โ”‚ โ”œโ”€ toolResultCompressor.ts โ”€โ”€ file/grep/shell/JSON/error compression - โ”‚ โ””โ”€ progressiveAging.ts โ”€โ”€โ”€โ”€ Older messages โ†’ shorter summaries - โ””โ”€ ultra.ts โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Heuristic token scoring + pruning - โ””โ”€ ultraHeuristic.ts โ”€ Stopword detection, score thresholds, force-preserve -``` - -### Configuration - -``` -Dashboard โ†’ Context & Cache โ†’ Caveman / RTK / Compression Combos -``` - -Or per-combo override: - -```json -{ - "comboOverrides": { - "my-coding-combo": "standard", - "my-cheap-combo": "ultra" - } -} -``` - -Auto-trigger: set `autoTriggerTokens` to automatically enable compression when a request exceeds a token threshold. - -Compression combos can also assign a named compression pipeline to routing combos, so a coding combo can use RTK + Caveman while a paid subscription combo stays on lite mode. - -> ๐Ÿชจ **Fun fact:** The standard/caveman mode is inspired by [Caveman](https://github.com/JuliusBrussee/caveman) โ€” the viral project that reports 65% average output-token savings while keeping technical accuracy. OmniRoute takes this further with a **7-option pipeline** and a default `RTK -> Caveman` combo that can reach ~89% average savings on eligible tool/context payloads. - -๐Ÿ“– **Full compression documentation:** [`docs/COMPRESSION_GUIDE.md`](docs/COMPRESSION_GUIDE.md) โ€ข [`docs/RTK_COMPRESSION.md`](docs/RTK_COMPRESSION.md) โ€ข [`docs/COMPRESSION_ENGINES.md`](docs/COMPRESSION_ENGINES.md) โ€ข [`docs/COMPRESSION_RULES_FORMAT.md`](docs/COMPRESSION_RULES_FORMAT.md) โ€ข [`docs/COMPRESSION_LANGUAGE_PACKS.md`](docs/COMPRESSION_LANGUAGE_PACKS.md) - ---- - -## ๐ŸŽฏ What OmniRoute Solves - -> **Every developer using AI tools faces these problems daily.** OmniRoute solves them all. - -| # | Problem | OmniRoute Solution | -| --- | ---------------------------------------- | ----------------------------------------------------------------------------------------------- | -| ๐Ÿ’ธ | Subscription quota expires mid-coding | **Smart 4-Tier Fallback** โ€” auto-routes Subscription โ†’ API Key โ†’ Cheap โ†’ Free | -| ๐Ÿ”Œ | Each provider has a different API format | **Format Translation** โ€” unified endpoint translates OpenAI โ†” Claude โ†” Gemini โ†” Responses | -| ๐ŸŒ | AI providers block my country/region | **3-Level Proxy** โ€” global, per-provider, and per-key proxy with TLS fingerprint spoofing | -| ๐Ÿ†“ | Can't afford AI subscriptions | **11 Free Providers** โ€” Kiro, Qoder, Pollinations, LongCat, Cloudflare AI, NVIDIA NIM... | -| ๐Ÿ”’ | Gateway is exposed without protection | **API Key Management** โ€” scoping, rotation, IP filtering, rate limiting, prompt injection guard | -| ๐Ÿ›‘ | Provider went down, lost coding flow | **Circuit Breakers** โ€” auto-failover with cooldown, retry, anti-thundering herd | -| ๐Ÿ”ง | Configuring each CLI tool is tedious | **CLI Tools Dashboard** โ€” one-click setup for Claude Code, Codex, Cursor, OpenClaw, Kilo | -| ๐Ÿ”‘ | Managing OAuth tokens is hell | **Auto Token Refresh** โ€” OAuth PKCE for 8 providers, multi-account, LAN/remote fix | -| ๐Ÿ“Š | Don't know how much I'm spending | **Cost Analytics** โ€” per-token tracking, budget limits, usage stats per API key | -| ๐Ÿ› | Can't diagnose errors in AI calls | **Unified Logs** โ€” 4-tab dashboard (request, proxy, audit, console) + p50/p95/p99 telemetry | - -
-๐Ÿ“– See all 31 problems OmniRoute solves - -| # | Problem | Solution | -| --- | --------------------------------------------- | -------------------------------------------------------------------------------------------------- | -| 11 | Deploying/maintaining is complex | npm global, Docker multi-arch, Electron, Termux โ€” deploy anywhere | -| 12 | Interface is English-only | 40+ languages with RTL support | -| 13 | Need more than chat (images, audio, video) | 10 multi-modal APIs: embeddings, images, video, music, TTS, STT, moderation, rerank, search, batch | -| 14 | No way to test/compare models | LLM Evals, Translator Playground, Chat Tester, Live Monitor | -| 15 | Need to scale without losing performance | Semantic cache, request dedup, rate limit detection, queue & pacing | -| 16 | Want to control model behavior globally | System prompt injection, thinking budget, wildcard routing | -| 17 | Need MCP tools as first-class features | 29 MCP tools, 3 transports (stdio/SSE/HTTP), 10 scopes, audit trail | -| 18 | Need A2A orchestration | JSON-RPC 2.0 + SSE streaming, task lifecycle, sync + stream paths | -| 19 | Need real MCP process health | Runtime heartbeat, PID tracking, UI status cards | -| 20 | Need auditable MCP execution | SQLite-backed audit with filters, pagination, stats | -| 21 | Need scoped MCP permissions | 10 granular scopes per integration | -| 22 | Need operational controls without redeploying | Combo switches, resilience tuning, breaker resets from dashboard | -| 23 | Need A2A task lifecycle visibility | Task listing/filtering, drill-down, cancellation | -| 24 | Need active stream metrics | Active stream counters, per-state counts, A2A dashboard cards | -| 25 | Need standard agent discovery | Agent Card at `/.well-known/agent.json` | -| 26 | Need protocol discoverability | Consolidated Endpoints page with Proxy, MCP, A2A, API tabs | -| 27 | Need E2E protocol validation | Real MCP SDK + A2A client flows in `test:protocols:e2e` | -| 28 | Need unified observability | Health + audit + telemetry across OpenAI, MCP, and A2A layers | -| 29 | Need one runtime for proxy + tools + agents | OpenAI proxy + MCP + A2A in one stack with shared auth/resilience | -| 30 | Need agentic workflows without glue-code | Unified endpoint, protocol UIs, production-ready foundations | -| 31 | Long sessions crash with context limits | Proactive context compression, structural integrity guards, multi-layer dropping | - -
- -๐Ÿ“– **Deep dives:** [Resilience Guide](docs/RESILIENCE_GUIDE.md) โ€ข [Proxy Guide](docs/PROXY_GUIDE.md) โ€ข [Setup Guide](docs/SETUP_GUIDE.md) โ€ข [Compression Guide](docs/COMPRESSION_GUIDE.md) - ---- - -## ๐Ÿ†“ Start Free โ€” Zero Configuration Cost - -> Setup AI coding in minutes at **$0/month**. Connect these free accounts and use the built-in **Free Stack** combo. - -| Step | Action | Providers Unlocked | -| ---- | -------------------------------------------------- | ------------------------------------------------------------------ | -| 1 | Connect **Kiro** (AWS Builder ID OAuth) | Claude Sonnet 4.5, Haiku 4.5 โ€” **unlimited** | -| 2 | Connect **Qoder** (Google OAuth) | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1... โ€” **unlimited** | -| 3 | Connect **Qwen** (Device Code) | qwen3-coder-plus, qwen3-coder-flash... โ€” **unlimited** | -| 4 | Connect **Gemini CLI** (Google OAuth) | gemini-3-flash, gemini-2.5-pro โ€” **180K/mo free** | -| 5 | `/dashboard/combos` โ†’ **Free Stack ($0)** template | Round-robin all free providers automatically | - -**Point any IDE/CLI to:** `http://localhost:20128/v1` ยท API Key: `any-string` ยท Done. - -> **Optional extra coverage (also free):** Groq API key (30 RPM free), NVIDIA NIM (40 RPM free, 70+ models), Cerebras (1M tok/day), LongCat API key (50M tokens/day!), Cloudflare Workers AI (10K Neurons/day, 50+ models). - ## โšก Quick Start -### 1) Install and run - ```bash -npm install -g omniroute -omniroute -``` - -Dashboard opens at `http://localhost:20128` ยท API at `http://localhost:20128/v1`. - -### 2) Connect providers - -1. Dashboard โ†’ **Providers** โ†’ connect at least one provider (OAuth or API key) -2. Dashboard โ†’ **Endpoints** โ†’ create an API key -3. Dashboard โ†’ **Combos** โ†’ set your fallback chain (optional) - -### 3) Point your coding tool - -```txt -Base URL: http://localhost:20128/v1 -API Key: [copy from Endpoint page] -Model: if/kimi-k2-thinking (or any provider/model) -``` - -Works with Claude Code, Codex CLI, Gemini CLI, Cursor, Cline, OpenClaw, OpenCode, and any OpenAI-compatible tool. - -
-๐Ÿ“ฆ More install methods (Docker, source, Arch, Void, pnpm) - -**Docker:** - -```bash -docker run -d --name omniroute --restart unless-stopped -p 20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest -``` - -**From source:** - -```bash -cp .env.example .env && npm install -PORT=20128 DASHBOARD_PORT=20129 NEXT_PUBLIC_BASE_URL=http://localhost:20129 npm run dev -``` - -**pnpm:** `pnpm install -g omniroute && pnpm approve-builds -g && omniroute` - -**Arch Linux (AUR):** `yay -S omniroute-bin && systemctl --user enable --now omniroute.service` - -**MCP:** `omniroute --mcp` (stdio transport) - -**CLI options:** `omniroute setup`, `omniroute doctor`, `omniroute providers available`, `omniroute providers list`, `omniroute --port 3000`, `omniroute --no-open`, `omniroute --help` - -**Split-port mode:** `PORT=20128 DASHBOARD_PORT=20129 omniroute` - -**Uninstall:** `npm run uninstall` (keeps data) or `npm run uninstall:full` (removes everything) - -๐Ÿ“– Full details: [Setup Guide](#-setup-guide) ยท [Docker](#-docker) ยท [Void Linux template](#-quick-start) - -
- ---- - -## ๐Ÿณ Docker - -OmniRoute is available as a public Docker image on [Docker Hub](https://hub.docker.com/r/diegosouzapw/omniroute). - -**Quick run:** - -```bash -docker run -d \ - --name omniroute \ - --restart unless-stopped \ - --stop-timeout 40 \ - -p 20128:20128 \ - -v omniroute-data:/app/data \ - diegosouzapw/omniroute:latest -``` - -**With environment file:** - -```bash -# Copy and edit .env first -cp .env.example .env - -docker run -d \ - --name omniroute \ - --restart unless-stopped \ - --stop-timeout 40 \ - --env-file .env \ - -p 20128:20128 \ - -v omniroute-data:/app/data \ - diegosouzapw/omniroute:latest -``` - -**Using Docker Compose:** - -```bash -# Base profile (no CLI tools) -docker compose --profile base up -d - -# CLI profile (Claude Code, Codex, OpenClaw built-in) -docker compose --profile cli up -d -``` - -Dashboard support for Docker deployments now includes a one-click **Cloudflare Quick Tunnel** on `Dashboard โ†’ Endpoints`. The first enable downloads `cloudflared` only when needed, starts a temporary tunnel to your current `/v1` endpoint, and shows the generated `https://*.trycloudflare.com/v1` URL directly below your normal public URL. Endpoint tunnel panels, including Cloudflare, Tailscale, and ngrok, can be shown or hidden from `Settings โ†’ Appearance` without changing active tunnel state. - -Notes: - -- Quick Tunnel URLs are temporary and change after every restart. -- Quick Tunnels are not auto-restored after an OmniRoute or container restart. Re-enable them from the dashboard when needed. -- Managed install currently supports Linux, macOS, and Windows on `x64` / `arm64`. -- Managed Quick Tunnels default to HTTP/2 transport to avoid noisy QUIC UDP buffer warnings in constrained container environments. Set `CLOUDFLARED_PROTOCOL=quic` or `auto` if you want a different transport. -- Docker images bundle system CA roots and pass them to managed `cloudflared`, which avoids TLS trust failures when the tunnel bootstraps inside the container. -- SQLite runs in WAL mode. `docker stop` should be allowed to finish so OmniRoute can checkpoint the latest changes back into `storage.sqlite`. -- The bundled Compose files already set a 40s stop grace period. If you run the image directly, keep `--stop-timeout 40` (or similar) so manual stops do not cut off shutdown cleanup. -- Set `CLOUDFLARED_BIN=/absolute/path/to/cloudflared` if you want OmniRoute to use an existing binary instead of downloading one. - -**Using Docker Compose with Caddy (HTTPS Auto-TLS):** - -OmniRoute can be securely exposed using Caddy's automatic SSL provisioning. Ensure your domain's DNS A record points to your server's IP. - -```yaml -services: - omniroute: - image: diegosouzapw/omniroute:latest - container_name: omniroute - restart: unless-stopped - volumes: - - omniroute-data:/app/data - environment: - - PORT=20128 - - NEXT_PUBLIC_BASE_URL=https://your-domain.com - - caddy: - image: caddy:latest - container_name: caddy - restart: unless-stopped - ports: - - "80:80" - - "443:443" - command: caddy reverse-proxy --from https://your-domain.com --to http://omniroute:20128 - -volumes: - omniroute-data: -``` - -| Image | Tag | Size | Description | -| ------------------------ | -------- | ------ | --------------------- | -| `diegosouzapw/omniroute` | `latest` | ~250MB | Latest stable release | -| `diegosouzapw/omniroute` | `3.7.8` | ~250MB | Current version | - -๐Ÿ“– **Full Docker documentation:** [`docs/DOCKER_GUIDE.md`](docs/DOCKER_GUIDE.md) โ€” Compose profiles, Caddy HTTPS, Cloudflare tunnels, and more. - ---- - -## ๐Ÿ“ฑ Multi-Platform โ€” Run Anywhere - -> OmniRoute runs on **Web**, **Desktop (Electron)**, **Android (Termux)**, and as a **Progressive Web App (PWA)**. - -| Platform | Install | Highlights | -| -------------- | -------------------------------------------- | -------------------------------------------------------------------------- | -| ๐Ÿ–ฅ๏ธ **Desktop** | `npm run electron:build` | Native window, system tray, auto-start, offline mode โ€” Windows/macOS/Linux | -| ๐Ÿ“ฑ **Android** | `pkg install nodejs-lts && npx -y omniroute` | ARM native, no root, 24/7 via Termux:Boot โ€” your phone is an AI server | -| ๐Ÿ“ฒ **PWA** | "Add to Home Screen" in browser | Fullscreen, offline page, service worker caching โ€” Android/iOS/Desktop | - -
-๐Ÿ–ฅ๏ธ Desktop App details - -- Native Electron app with system tray, auto-start, native notifications -- One-click install: NSIS (Windows), DMG (macOS), AppImage (Linux) -- Dev: `npm run electron:dev` ยท Build: `npm run electron:build` -- ๐Ÿ“– Full docs: [`electron/README.md`](electron/README.md) - -
- -
-๐Ÿ“ฑ Android (Termux) details - -```bash -pkg update && pkg install nodejs-lts python build-essential git +# Run instantly (npx โ€” no install needed) npx -y omniroute@latest + +# Or install globally +npm install -g omniroute && omniroute + +# Or via Docker +docker run -d -p 20128:20128 diegosouzapw/omniroute:3.8.0 ``` -Access from any device on the same network: `http://PHONE_IP:20128/v1` +โ†’ Open **http://localhost:20128** โ†’ login with `admin` / `CHANGEME` โ†’ connect your first provider via OAuth or API key. -- ๐Ÿ“– Full guide: [`docs/TERMUX_GUIDE.md`](docs/TERMUX_GUIDE.md) - -
- -
-๐Ÿ“ฒ PWA details - -- **Android (Chrome):** โ‹ฎ โ†’ "Add to Home screen" -- **iOS (Safari):** Share โ†’ "Add to Home Screen" -- **Desktop (Chrome/Edge):** Install icon in address bar -- ๐Ÿ“– Full docs: [`docs/PWA_GUIDE.md`](docs/PWA_GUIDE.md) - -
- ---- - -## ๐ŸŒ Bypass Geographic Blocks โ€” Use AI From Any Country - -> ๐Ÿ‡ท๐Ÿ‡บ ๐Ÿ‡จ๐Ÿ‡ณ ๐Ÿ‡ฎ๐Ÿ‡ท ๐Ÿ‡จ๐Ÿ‡บ ๐Ÿ‡น๐Ÿ‡ท **In Russia, China, Iran, or any blocked region?** OmniRoute's 3-level proxy system solves this completely. - -| Level | Badge | Configure In | Use Case | -| ------------------ | ----- | ------------------ | ------------------------------- | -| **Global** | ๐ŸŸข | Settings โ†’ Proxy | All traffic through one proxy | -| **Per-Provider** | ๐ŸŸก | Provider โ†’ Proxy | Only specific providers proxied | -| **Per-Connection** | ๐Ÿ”ต | Connection โ†’ Proxy | Each API key uses its own proxy | - -**What gets proxied:** API requests โœ… โ€ข OAuth flows โœ… โ€ข Connection tests โœ… โ€ข Token refresh โœ… โ€ข Model sync โœ… - -**Protocols:** HTTP/HTTPS, SOCKS5 (`ENABLE_SOCKS5_PROXY=true`), Authenticated proxies - -### ๐Ÿ†“ 1proxy โ€” Free Proxy Marketplace - -> Contributed by [@oyi77](https://github.com/oyi77) โ€” [#1847](https://github.com/diegosouzapw/OmniRoute/pull/1847) - -No proxy? Use the built-in **1proxy** integration for **hundreds of free, validated proxies** worldwide: - -- One-click sync (up to 500 proxies) โ€ข Quality scores (0-100) โ€ข Country filter โ€ข Auto-rotation (quality/random/sequential) โ€ข Auto-degradation โ€ข Circuit breaker - -### Anti-Detection - -- ๐Ÿ”’ **TLS Fingerprint Spoofing** โ€” browser-like TLS via `wreq-js` -- ๐Ÿ” **CLI Fingerprint Matching** โ€” matches native CLI binary signatures -- ๐Ÿ  **Proxy IP Preservation** โ€” stealth + IP masking simultaneously - -๐Ÿ“– **Full proxy documentation:** [`docs/PROXY_GUIDE.md`](docs/PROXY_GUIDE.md) - ---- - ---- - -## ๐Ÿ’ฐ Pricing at a Glance - -| Tier | Provider | Cost | Quota Reset | Best For | -| ------------------- | --------------------------- | ------------------------- | ---------------- | --------------------------------- | -| **๐Ÿ’ณ SUBSCRIPTION** | Claude Code (Pro) | $20/mo | 5h + weekly | Already subscribed | -| | Codex (Plus/Pro) | $20-200/mo | 5h + weekly | OpenAI users | -| | Gemini CLI | **FREE** | 180K/mo + 1K/day | Everyone! | -| | GitHub Copilot | $10-19/mo | Monthly | GitHub users | -| **๐Ÿ”‘ API KEY** | NVIDIA NIM | **FREE** (dev forever) | ~40 RPM | 70+ open models | -| | Cerebras | **FREE** (1M tok/day) | 60K TPM / 30 RPM | World's fastest | -| | Groq | **FREE** (30 RPM) | 14.4K RPD | Ultra-fast Llama/Gemma | -| | DeepSeek V3.2 | $0.27/$1.10 per 1M | None | Best price/quality reasoning | -| | xAI Grok-4 Fast | **$0.20/$0.50 per 1M** ๐Ÿ†• | None | Fastest + tool calling, ultralow | -| | xAI Grok-4 (standard) | $0.20/$1.50 per 1M ๐Ÿ†• | None | Reasoning flagship from xAI | -| | Mistral | Free trial + paid | Rate limited | European AI | -| | OpenRouter | Pay-per-use | None | 100+ models aggr. | -| | AgentRouter ๐Ÿ†• | Pay-per-use | None | $200 free credits at signup | -| **๐Ÿ’ฐ CHEAP** | GLM-5 (via Z.AI) ๐Ÿ†• | $0.5/1M | Daily 10AM | 128K output, newest flagship | -| | GLM-4.7 | $0.6/1M | Daily 10AM | Budget backup | -| | MiniMax M2.5 ๐Ÿ†• | $0.3/1M input | 5-hour rolling | Reasoning + agentic tasks | -| | MiniMax M2.1 | $0.2/1M | 5-hour rolling | Cheapest option | -| | Kimi K2.5 (Moonshot API) ๐Ÿ†• | Pay-per-use | None | Direct Moonshot API access | -| | Kimi K2 | $9/mo flat | 10M tokens/mo | Predictable cost | -| **๐Ÿ†“ FREE** | Qoder | **$0** | Unlimited | 5 models unlimited | -| | Qwen | **$0** | Unlimited | 4 models unlimited | -| | Kiro | **$0** | Unlimited | Claude Sonnet/Haiku (AWS Builder) | -| | LongCat Flash-Lite ๐Ÿ†• | **$0** (50M tok/day ๐Ÿ”ฅ) | 1 RPS | Largest free quota on Earth | -| | Pollinations AI ๐Ÿ†• | **$0** (no key needed) | 1 req/15s | GPT-5, Claude, DeepSeek, Llama 4 | -| | Cloudflare Workers AI ๐Ÿ†• | **$0** (10K Neurons/day) | ~150 resp/day | 50+ models, global edge | -| | Scaleway AI ๐Ÿ†• | **$0** (1M tokens total) | Rate limited | EU/GDPR, Qwen3 235B, Llama 70B | - -> ๐Ÿ†• **New models added (Mar 2026):** Grok-4 Fast family at $0.20/$0.50/M (benchmarked at 1143ms โ€” 30% faster than Gemini 2.5 Flash), GLM-5 via Z.AI with 128K output, MiniMax M2.5 reasoning, DeepSeek V3.2 updated pricing, Kimi K2.5 via Moonshot direct API. - -**๐Ÿ’ก See the full [$0 Free Stack (11 providers)](#-free-models--11-providers-0-forever) below.** - -> ๐Ÿ’ก **Understanding Dashboard Costs:** -> -> The "cost" displayed in the Usage Analytics page is **for tracking and comparison purposes only**. -> OmniRoute itself **never charges you anything** โ€” it's free, open-source software running on your machine. -> If your dashboard shows "$290 total cost" while using free models, that's how much you **saved** compared to paid API pricing. -> Think of it as a **savings tracker**, not a bill. - ---- - -## ๐Ÿ†“ Free Models โ€” 11 Providers, $0 Forever - -> Combine all free providers into one unbreakable combo โ€” OmniRoute auto-routes between them when quota runs out. - -| Provider | Prefix | Free Models | Quota | -| ----------------- | ----------- | ------------------------------------------------------------- | -------------------- | -| **Kiro** | `kr/` | Claude Sonnet 4.5, Haiku 4.5, Opus 4.6 | 50 CREDITS per month | -| **Qoder** | `if/` | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1, minimax-m2.1 | โ™พ๏ธ Unlimited | -| **Qwen** | `qw/` | qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next | โ™พ๏ธ Unlimited | -| **Pollinations** | `pol/` | GPT-5, Claude, Gemini, DeepSeek, Llama 4, Mistral | No key needed | -| **LongCat** | `lc/` | LongCat-Flash-Lite | 50M tokens/day ๐Ÿ”ฅ | -| **Gemini CLI** | `gc/` | gemini-3-flash, gemini-2.5-pro | 180K tok/mo | -| **Cloudflare AI** | `cf/` | 50+ models (Llama, Gemma, Mistral, Whisper) | 10K Neurons/day | -| **Groq** | `groq/` | Llama 3.3 70B, Qwen3 32B, Kimi K2 | 14.4K RPD | -| **NVIDIA NIM** | `nvidia/` | 129 models (DeepSeek, Llama, GLM, Kimi) | ~40 RPM | -| **Cerebras** | `cerebras/` | Qwen3 235B, GPT-OSS 120B, Llama 3.1 | 1M tok/day | -| **Scaleway** | `scw/` | Qwen3 235B, Llama 70B, DeepSeek V3 | 1M tokens (EU) | - -
-๐Ÿ“– 25+ more free providers โ€” Groq, Cerebras, Mistral, GitHub Models, OpenRouter, and more - -**Also free (API Key required):** -Mistral (1B tok/month) ยท OpenRouter (35+ `:free` models) ยท GitHub Models (GPT-5, 45+ models) ยท -Cohere (1K calls/month) ยท Z.AI/GLM (permanent free Flash models) ยท SiliconFlow (1K RPM, 50K TPM) ยท -Kilo Code (~200 req/hr auto-router) ยท HuggingFace ($0.10/mo credits) ยท Ollama Cloud (400+ models) ยท -LLM7.io (30+ models) ยท Kluster AI ยท IBM watsonx (300K tok/month) ยท OpenCode Zen ยท Vercel AI Gateway ($5/mo) - -**Trial credits (one-time):** -Baseten ($30) ยท NLP Cloud ($15) ยท AI21 ($10) ยท Upstage ($10) ยท SambaNova ($5) ยท Modal ($5/mo) ยท -Fireworks ($1) ยท Nebius ($1) ยท Inference.net ($1 + $25 survey) ยท Hyperbolic ($1) ยท Novita ($0.50) - -**China-based (free tiers):** -ModelScope ยท Tencent Hunyuan ยท Volcengine ยท ChatAnywhere ยท InternAI ยท Bigmodel - -**Combined capacity: ~31,000+ RPD ยท ~32B+ tokens/month ยท 500+ models ยท $0** - -
- -๐Ÿ“– **Complete free provider directory:** [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md) โ€” 25+ providers, quotas, base URLs, model tables, and OmniRoute combo setup. - ---- - -## ๐ŸŽ™๏ธ Free Transcription Combo - -> Transcribe any audio/video for **$0** โ€” Deepgram leads with $200 free, AssemblyAI $50 fallback, Groq Whisper as unlimited emergency backup. - -| Provider | Free Credits | Best Model | Rate Limit | -| ----------------- | ---------------------- | -------------------------------------------- | ---------------------------- | -| ๐ŸŸข **Deepgram** | **$200 free** (signup) | `nova-3` โ€” best accuracy, 30+ languages | No RPM limit on free credits | -| ๐Ÿ”ต **AssemblyAI** | **$50 free** (signup) | `universal-3-pro` โ€” chapters, sentiment, PII | No RPM limit on free credits | -| ๐Ÿ”ด **Groq** | **Free forever** | `whisper-large-v3` โ€” OpenAI Whisper | 30 RPM (rate limited) | - ---- - -**Suggested combo in `/dashboard/combos`:** - -``` -Name: free-transcription -Strategy: Priority -Nodes: - [1] deepgram/nova-3 โ†’ uses $200 free first - [2] assemblyai/universal-3-pro โ†’ fallback when Deepgram credits run out - [3] groq/whisper-large-v3 โ†’ free forever, emergency fallback -``` - -Then in `/dashboard/media` โ†’ **Transcription** tab: upload any audio or video file โ†’ select your combo endpoint โ†’ get transcription in supported formats. - -## ๐Ÿ’ก Key Features - -> **4,690+ automated tests** across 517 test files. Not just a relay โ€” a full operational platform. - -| Feature | Why It Matters | -| ---------------------------------------------------------------------------------------------------- | -------------------------------- | -| ๐Ÿง  **Smart 4-Tier Fallback** โ€” Subscription โ†’ API โ†’ Cheap โ†’ Free | Never stop coding, zero downtime | -| ๐Ÿ”„ **Format Translation** โ€” OpenAI โ†” Claude โ†” Gemini โ†” Responses API | Works with ANY CLI tool | -| ๐Ÿ—œ๏ธ **Prompt Compression** โ€” 7 options including Caveman, RTK, and stacked pipelines | Save 15-95% eligible tokens | -| ๐Ÿค– **MCP Server** โ€” 37 tools, 3 transports (stdio/SSE/HTTP), 10 scopes | IDE/agent tool integration | -| ๐Ÿ›ก๏ธ **Resilience Engine** โ€” circuit breakers, cooldowns, TLS spoofing, anti-thundering herd | Auto-recovery from any failure | -| ๐ŸŽต **10 Multi-Modal APIs** โ€” chat, embed, images, video, music, TTS, STT, moderation, rerank, search | One endpoint for everything | -| ๐ŸŒ **3-Level Proxy** โ€” global, per-provider, per-key + 1proxy free marketplace | Access AI from any country | -| ๐Ÿ“Š **Full Observability** โ€” unified logs, p50/p95/p99 telemetry, cost tracking, budget controls | Know exactly what's happening | - -
-๐Ÿ“‹ Complete feature list โ€” 30+ capabilities - -**Routing & Intelligence** - -- 13 balancing strategies (priority, weighted, round-robin, P2C, cost-optimized, context-relay...) -- Task-aware smart routing (coding/vision/analysis) ยท Context relay session handoffs -- Thinking budget controls (passthrough/auto/custom) ยท Wildcard routing ยท System prompt injection - -**Translation & Compatibility** - -- Auto token refresh (OAuth PKCE for 8 providers) ยท Multi-account round-robin -- Responses API โ€” full `/v1/responses` for Codex ยท Batch API with Files API -- OpenAPI 3.0 live spec + Try-It UI - -**Protocols** - -- A2A Server โ€” JSON-RPC 2.0, SSE streaming, task lifecycle, skills -- ACP โ€” CLI agent discovery (14 agents + custom) - -**Platform** - -- Desktop (Electron) ยท Android (Termux) ยท PWA ยท Docker (AMD64 + ARM64) -- Cloudflare / Tailscale / ngrok tunnels ยท 40+ languages with RTL -- Semantic + signature cache (two-tier) ยท Request idempotency + deduplication - -**Observability** - -- Health dashboard โ€” uptime, breakers, cache, lockouts -- Evaluation framework โ€” golden set testing ยท Webhooks ยท Compliance audit - -**v3.6+ Highlights:** -V1 WebSocket Bridge ยท Sync Tokens & Config Bundle ยท GLM Thinking (glmt) ยท Hybrid Token Counting ยท -Safe Outbound Fetch ยท Wait For Cooldown ยท Runtime Env Validation ยท Vision Bridge ยท -Grok-4 Fast ยท GLM-5 via Z.AI ยท MiniMax M2.5 ยท toolCalling flag ยท -Multilingual Intent Detection ยท Benchmark-Driven Fallbacks ยท Request Deduplication - -**Architecture Examples:** - -```txt -Combo: "my-coding-stack" Format Translation: - 1. cc/claude-opus-4-7 CLI โ†’ OpenAI format - 2. nvidia/llama-3.3-70b OmniRoute โ†’ translates - 3. glm/glm-4.7 Provider โ†’ native format - 4. if/kimi-k2-thinking -``` - -๐Ÿ“– [MCP Server README](open-sse/mcp-server/README.md) ยท [A2A Server README](src/lib/a2a/README.md) ยท [Resilience Guide](docs/RESILIENCE_GUIDE.md) ยท [Features Gallery](docs/FEATURES.md) - -
- ---- - -## ๐ŸŽฏ Use Cases โ€” Ready-Made Combo Playbooks - -### Case 0: "I want zero-config, auto-routing NOW" - -**Problem:** Don't want to create combos manually. Just want AI routing to work immediately. +**Point any OpenAI-compatible client at OmniRoute:** ```bash -# No combo creation needed! Use auto/ prefix directly: -model: "auto" # Default LKGP routing across all connected providers -model: "auto/coding" # Quality-first weights for code generation -model: "auto/fast" # Low-latency routing (fastest provider first) -model: "auto/cheap" # Cost-optimized (cheapest per token) -model: "auto/offline" # High availability (most quota available) -model: "auto/smart" # Best discovery (10% exploration rate) +export OPENAI_BASE_URL=http://localhost:20128/v1 +export OPENAI_API_KEY=or_ ``` -**How it works:** - -1. Add providers in Dashboard โ†’ Providers (OAuth or API key) -2. Use `auto/` prefix in any AI tool โ€” **no combo creation needed** -3. OmniRoute dynamically builds a virtual combo from your active connections -4. Routes using LKGP (Last Known Good Provider) + 6-factor scoring -5. Session stickiness ensures consistent provider selection - -**Dashboard indicator:** A blue banner at the top shows "Auto-Routing Active" with a link to `/dashboard/combos` for configuration. - -**Monthly cost:** $0 (uses your existing free providers) or whatever your connected providers cost +That's it. Cursor, Cline, Codex, Continue, Aider, and any SDK now work. โ†’ Detailed setup: [`docs/SETUP_GUIDE.md`](docs/SETUP_GUIDE.md) --- -### Case 1: "I have a Claude Pro subscription" +## ๐ŸŒŸ What's new in v3.8.0 -**Problem:** Quota expires unused, rate limits during heavy coding sessions. +- ๐Ÿค– **Auto-Combo zero-config routing** โ€” just use `auto/coding`, `auto/cheap`, `auto/fast`, `auto/offline`, `auto/smart`, `auto/lkgp` as model IDs +- ๐ŸŽฏ **Manifest-aware tier routing W1-W4** โ€” automatic tier prioritization +- ๐Ÿ†• **Command Code provider** + **Z.AI quota labels** + **KIE video expansion** +- ๐Ÿ” **Windsurf + Devin CLI + GitLab Duo OAuth** flows +- ๐Ÿ†“ **9 new free providers**: LLM7, Lepton, Kluster, UncloseAI, BazaarLink, Completions, Enally, FreeTheAi, AgentRouter ($200 credits) +- ๐Ÿฉบ **Model Cooldowns dashboard** with manual re-enable +- ๐ŸŽจ **Cursor full OpenAI parity** (tools, streaming, sessions) +- ๐Ÿ“Œ **Per-session sticky routing** for Codex +- ๐Ÿ”Š **Inworld TTS** enhancements +- ๐Ÿง  **Reasoning Replay Cache** โ€” fixes 400s on DeepSeek V4, Kimi K2, Qwen-Thinking, GLM +- ๐Ÿ”„ **Reset-aware routing** strategy (14th strategy) +- ๐Ÿ› ๏ธ **20+ new CLI commands** (`omniroute setup/doctor/providers/combos`) -``` -Combo: "maximize-claude" - 1. cc/claude-opus-4-7 (use subscription fully) - 2. glm/glm-5.1 (cheap backup when quota out โ€” $0.5/1M) - 3. kr/claude-sonnet-4.5 (free emergency fallback via Kiro) - -Compression: standard (caveman) โ€” saves 30% tokens = stretch quota further -Monthly cost: $20 (subscription) + ~$3 (backup) = $23 total -vs. $20 + hitting limits + lost productivity = frustration -``` - -### Case 2: "I want $0 forever" - -**Problem:** Can't afford subscriptions, need reliable AI for coding. - -``` -Combo: "free-forever" - 1. kr/claude-sonnet-4.5 (Claude 4.5 free unlimited via Kiro) - 2. if/kimi-k2-thinking (reasoning model free via Qoder) - 3. pol/gpt-5 (GPT-5 free via Pollinations โ€” no key) - 4. lc/longcat-flash-lite (50M tokens/day free backup) - -Compression: aggressive โ€” saves 50% tokens = double your free quota -Monthly cost: $0 -Quality: Production-ready models + 50% token savings -``` - -### Case 3: "I need 24/7 coding, no interruptions" - -**Problem:** Deadlines, can't afford any downtime. - -``` -Combo: "always-on" - 1. cc/claude-opus-4-7 (best quality โ€” subscription) - 2. cx/gpt-5.5 (second subscription โ€” OpenAI) - 3. glm/glm-5.1 (cheap, resets daily โ€” $0.5/1M) - 4. minimax/MiniMax-M2.5 (cheapest paid โ€” $0.3/1M) - 5. kr/claude-sonnet-4.5 (free unlimited โ€” never fails) - -Compression: lite โ€” saves 15% tokens passively, zero risk -Result: 5 layers of fallback = zero downtime -Monthly cost: $20-200 (subscriptions) + $5-10 (backup) -``` - -### Case 4: "I'm in a blocked region (Russia, China, Iran...)" - -**Problem:** AI providers block my country, VPNs are slow. - -``` -Combo: "unblocked-ai" - 1. kr/claude-sonnet-4.5 (free via Kiro + proxy) - 2. pol/deepseek-r1 (Pollinations โ€” no geo-block) - 3. groq/llama-3.3-70b (Groq + proxy) - -Proxy: Global proxy set in Settings โ†’ or per-provider proxy override -Result: Access ALL providers from ANY country -Monthly cost: $0 (free providers) + $0 (1proxy free marketplace) -``` - -### Case 5: "I want maximum token savings" - -**Problem:** Token costs are eating my budget, need to squeeze every token. - -``` -Combo: "ultra-saver" - 1. cc/claude-opus-4-7 (subscription โ€” best quality) - 2. glm/glm-5.1 (cheap backup) - -Compression: ultra โ€” saves 75% tokens -Result: 10K token prompt โ†’ 2.5K tokens sent -Montly savings: ~$150-300/month in token costs for heavy users -``` - -## ๐Ÿงช Evaluations (Evals) - -OmniRoute includes a built-in evaluation framework to test LLM response quality against a golden set. Access it via **Analytics โ†’ Evals** in the dashboard. - -### Built-in Golden Set - -The pre-loaded "OmniRoute Golden Set" contains test cases for: - -- Greetings, math, geography, code generation -- JSON format compliance, translation, markdown generation -- Safety refusal (harmful content), counting, boolean logic - -### Evaluation Strategies - -| Strategy | Description | Example | -| ---------- | ------------------------------------------------ | -------------------------------- | -| `exact` | Output must match exactly | `"4"` | -| `contains` | Output must contain substring (case-insensitive) | `"Paris"` | -| `regex` | Output must match regex pattern | `"1.*2.*3"` | -| `custom` | Custom JS function returns true/false | `(output) => output.length > 10` | +โ†’ Full changelog: [`CHANGELOG.md`](CHANGELOG.md) --- -## ๐Ÿ“– Setup Guide +## ๐ŸŽฏ Why OmniRoute Wins -### Connect Your Coding Tool +| | OmniRoute v3.8 | LiteLLM | OpenRouter | +| ------------------ | ------------------- | ---------------- | ------------- | +| Providers | **177+** | ~50 | ~50 | +| Free providers | **11** | 0 | 0 | +| OAuth providers | **14** | 0 | 1 | +| Routing strategies | **14** | 3 | 1 | +| Auto routing | โœ… 9-factor scoring | โŒ | โŒ | +| Prompt compression | โœ… RTK + Caveman | โŒ | โŒ | +| MCP server | โœ… 37 tools | โŒ | โŒ | +| A2A protocol | โœ… v0.3 + 5 skills | โŒ | โŒ | +| Desktop app | โœ… Electron 41 | โŒ | โŒ | +| PWA | โœ… | โŒ | โŒ | +| Self-hosted | โœ… MIT | Limited | โŒ (cloud) | +| Pricing | **$0 forever** | OSS / Cloud paid | 10% fee + API | -Point any OpenAI-compatible tool to OmniRoute: +โ†’ Detailed comparison: [`docs/FEATURES.md`](docs/FEATURES.md) -```txt -Base URL: http://localhost:20128/v1 -API Key: [from Dashboard โ†’ Endpoints] -``` +--- -| Tool | Config Location | -| --------------- | ----------------------------------------------------------------------------------------- | -| **Claude Code** | `claude mcp add-server omniroute --type http --url http://localhost:20128/api/mcp/stream` | -| **Codex CLI** | `OPENAI_BASE_URL=http://localhost:20128/v1 OPENAI_API_KEY=your-key codex` | -| **Cursor** | Settings โ†’ Models โ†’ Add Model โ†’ Override Base URL | -| **Cline** | Extension settings โ†’ Custom API Base URL | -| **OpenClaw** | `OPENAI_BASE_URL=http://localhost:20128/v1 openclaw` | -| **Gemini CLI** | Uses native OAuth via OmniRoute โ€” connect in Providers | +## ๐Ÿ› ๏ธ Compatible CLI Tools (17+) -### Protocols (MCP + A2A) +All work out-of-the-box once you point `OPENAI_BASE_URL` at OmniRoute: + +**Claude family:** Claude Code ยท Cline ยท Continue ยท Kilo Code ยท Kimi Coding +**OpenAI family:** Codex CLI ยท Cursor ยท Aider ยท OpenClaw ยท Droid ยท AMP +**Google family:** Gemini CLI ยท Antigravity ยท Jules +**Others:** Windsurf ยท GitLab Duo ยท Devin CLI ยท Hermes ยท Amazon Q ยท Kiro ยท Qoder ยท Custom + +โ†’ Full setup: [`docs/CLI-TOOLS.md`](docs/CLI-TOOLS.md) + +--- + +## ๐ŸŒ Providers (177+) + +### ๐Ÿ†“ Free providers (11 โ€” no API key or unlimited tier) + +| Provider | Highlight | +| --------------------------------------------------- | --------------------------------------- | +| **Kiro AI** | 50 credits/month (Claude Sonnet/Haiku) | +| **Qoder AI** | Unlimited (Kimi-K2, Qwen3, DeepSeek-R1) | +| **Gemini CLI** | 180K tokens/month | +| **Amazon Q** | AWS Builder ID OAuth | +| **LongCat** | 50M tokens/day | +| **Pollinations** | No API key, GPT-5 + Claude | +| **AgentRouter** | $200 free credits | +| **LLM7** ยท **Lepton** ยท **Kluster** ยท **UncloseAI** | New v3.8 free tiers | + +โš ๏ธ Qwen Code OAuth was **discontinued on 2026-04-15** (use API key with `alicode` provider instead). + +โ†’ Curated guide: [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md) ยท Full catalog: [`docs/PROVIDER_REFERENCE.md`](docs/PROVIDER_REFERENCE.md) (auto-generated) + +### ๐Ÿ” OAuth providers (14) + +Claude Code ยท Codex ยท GitHub Copilot ยท Cursor ยท Antigravity ยท Gemini ยท Kimi Coding ยท Kilo Code ยท Cline ยท Qwen ยท Kiro ยท Qoder ยท Windsurf ยท GitLab Duo + +### ๐Ÿ”‘ API key providers (~123) + +OpenAI ยท Anthropic ยท Google ยท Mistral ยท Cohere ยท DeepSeek ยท Groq ยท Together ยท Fireworks ยท Cerebras ยท SambaNova ยท NVIDIA NIM ยท Bedrock ยท Vertex ยท Azure ยท Cloudflare AI ยท 100+ more. + +### ๐Ÿ  Self-hosted (10) + +Ollama ยท LM Studio ยท vLLM ยท Llamafile ยท Lemonade ยท Petals ยท Triton ยท Docker Model Runner ยท Xinference ยท Oobabooga + +--- + +## ๐Ÿค– Auto-Combo โ€” Zero-Config Routing + +Just use `auto/` as model ID. No combo setup needed. ```bash -# MCP (stdio transport) -omniroute --mcp - -# A2A (JSON-RPC 2.0) -curl http://localhost:20128/.well-known/agent.json +# 6 variants + plain `auto`: +auto/coding # โ†’ optimized for coding tasks +auto/cheap # โ†’ minimize cost +auto/fast # โ†’ minimize latency +auto/offline # โ†’ prefer local providers +auto/smart # โ†’ prefer top-tier models +auto/lkgp # โ†’ Last-Known-Good-Path (sticky) +auto # โ†’ balanced default ``` -### Key Environment Variables +**How it picks:** 9-factor scoring (health ยท quota ยท cost ยท latency ยท taskFit ยท stability ยท tierPriority ยท tierAffinity ยท specificityMatch) over a virtual candidate pool built from all enabled providers. -| Variable | Default | Purpose | -| -------------------- | -------------- | ----------------------------------------- | -| `PORT` | `20128` | API and dashboard port | -| `DASHBOARD_PORT` | โ€” | Separate dashboard port (split-port mode) | -| `REQUIRE_API_KEY` | `false` | Require API key for all requests | -| `DATA_DIR` | `~/.omniroute` | Database and config storage | -| `REQUEST_TIMEOUT_MS` | `600000` | Upstream response timeout | - -
-๐Ÿ“– Full Setup Guide โ€” All CLI tools, protocols, and environment variables - -๐Ÿ“– **Complete documentation:** - -- [User Guide](docs/USER_GUIDE.md) โ€” Providers, combos, CLI integration -- [API Reference](docs/API_REFERENCE.md) โ€” All endpoints with examples -- [MCP Server](open-sse/mcp-server/README.md) โ€” 37 tools, IDE configs -- [A2A Server](src/lib/a2a/README.md) โ€” JSON-RPC, skills, streaming -- [Environment Config](docs/ENVIRONMENT.md) โ€” Complete `.env` reference -- [VM Deployment](docs/VM_DEPLOYMENT_GUIDE.md) โ€” VM + nginx + Cloudflare - -
+โ†’ Full guide: [`docs/AUTO-COMBO.md`](docs/AUTO-COMBO.md) --- -## โ“ Frequently Asked Questions +## ๐Ÿ—œ๏ธ Prompt Compression โ€” Save 15-95% Tokens -
-๐Ÿ“Š Why does my dashboard show high costs if I'm using free models? +Two engines, stackable: -The dashboard tracks your token usage and displays **estimated costs** as if you were using paid APIs directly. This is **not actual billing** โ€” it's a reference to show how much you're saving. +- **Caveman** โ€” natural-language condensation (filler removal, hedging, repeated context). 30+ regex rules per language pack (en, es, pt-BR, de, fr, ja). +- **RTK** โ€” terminal/shell/git/test output. 49 declarative filters. -**Example:** +**Modes:** `off` ยท `lite` ยท `standard` ยท `aggressive` ยท `ultra` ยท `rtk` ยท `stacked` (RTKโ†’Caveman, max savings). -- **Dashboard shows:** "$290 total cost" -- **Reality:** You're using Kiro + Qoder (FREE unlimited) -- **Your actual cost:** **$0.00** -- **What $290 means:** Amount you **saved** by using free models instead of paid APIs! - -The cost display is a "savings tracker" to help you understand your usage patterns and optimization opportunities. - -
- -
-๐Ÿ’ณ Will I be charged by OmniRoute? - -**No.** OmniRoute is free, open-source software that runs on your own computer. It never charges you anything. - -**You only pay:** - -- โœ… **Subscription providers** (Claude Code $20/mo, Codex $20-200/mo) โ†’ Pay them directly on their websites -- โœ… **API key providers** (DeepSeek, xAI, etc.) โ†’ Pay them directly, OmniRoute just routes your requests -- โŒ **OmniRoute itself** โ†’ **Never charges anything, ever** - -OmniRoute is a local proxy/router. It doesn't have your credit card, can't send invoices, and has no billing system. It's completely free software. - -
- -
-๐Ÿ†“ Are FREE providers really unlimited? - -**Yes!** The current FREE providers are genuinely free with **no hidden charges**: - -- **Kiro AI**: Free unlimited Claude Sonnet/Haiku via AWS Builder ID / Google / GitHub OAuth -- **Qoder**: Free unlimited kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 via PAT token -- **Pollinations AI**: No API key needed โ€” GPT-5, Claude, DeepSeek, Llama 4 -- **LongCat Flash-Lite**: 50M tokens/day โ€” largest free quota available -- **Cloudflare Workers AI**: 10K Neurons/day โ€” 50+ models at the edge - -OmniRoute just routes your requests to them โ€” there's no "catch" or future billing. - -
- -
-๐Ÿ’ฐ How do I minimize my actual AI costs? - -**Free-First Strategy:** - -1. **Start with 100% free combo:** - - ``` - 1. kr/claude-sonnet-4.5 (Kiro โ€” unlimited free) - 2. if/kimi-k2-thinking (Qoder โ€” unlimited free) - 3. pol/gpt-5 (Pollinations โ€” no key needed) - ``` - - **Cost: $0/month** - -2. **Enable Prompt Compression** โ€” even `lite` mode saves ~15% passively - -3. **Add cheap backup** only if you need it: - - ``` - 4. glm/glm-5.1 ($0.5/1M tokens) - ``` - - **Additional cost: Only pay for what you actually use** - -4. **Use subscription providers last** โ€” only if you already have them. OmniRoute helps maximize their value through quota tracking. - -**Result:** Most users can operate at **$0/month** using only free tiers! - -
- -
-๐Ÿ—œ๏ธ Will compression affect response quality? - -**No.** Compression only affects the **input** (your prompt), not the model's response. Each mode has been designed to preserve technical accuracy: - -- **Lite** (~15%): Only whitespace/formatting โ€” zero semantic change -- **Standard** (~30%): Removes filler words ("please", "I think", "basically") โ€” same meaning -- **Aggressive** (~50%): Summarizes old messages + compresses tool outputs โ€” core context preserved -- **Ultra** (~75%): Heuristic pruning โ€” use only when token budget is critical - -Code blocks, URLs, JSON, and structured data are **always protected** from compression via the preservation engine. - -
- -
-๐ŸŒ Does OmniRoute work in countries where AI is blocked? - -**Yes!** OmniRoute has a 3-level proxy system: - -1. **Global proxy** โ€” all requests go through your proxy -2. **Per-provider proxy** โ€” different proxy per provider -3. **Per-API-key proxy** โ€” different proxy per key - -Plus the **1proxy free marketplace** for community-shared proxies. Users in Russia, China, Iran, and other restricted regions can access all 160+ providers through OmniRoute's proxy infrastructure. - -See the [Proxy Guide](docs/PROXY_GUIDE.md) for setup instructions. - -
+โ†’ [`docs/COMPRESSION_GUIDE.md`](docs/COMPRESSION_GUIDE.md) ยท [`docs/RTK_COMPRESSION.md`](docs/RTK_COMPRESSION.md) ยท [`docs/COMPRESSION_LANGUAGE_PACKS.md`](docs/COMPRESSION_LANGUAGE_PACKS.md) --- -## ๐Ÿ› Troubleshooting +## ๐ŸŒ Bypass Geographic Blocks -| Problem | Quick Fix | -| --------------------------------------------- | ------------------------------------------------------------------------------------ | -| **"Language model did not provide messages"** | Provider quota exhausted โ†’ check quota tracker, use combo fallback | -| **Rate limiting (429)** | Add fallback combo: `cc/claude โ†’ glm/glm-4.7 โ†’ if/kimi-k2-thinking` | -| **OAuth token expired** | Auto-refreshed by OmniRoute. If stuck: delete + re-auth in Providers | -| **`unsupported_country_region_territory`** | Configure proxy in Settings โ†’ Proxy (see [Proxy Guide](docs/PROXY_GUIDE.md)) | -| **Docker SQLite locks** | Use `--stop-timeout 40` for clean WAL checkpoint on shutdown | -| **Node.js runtime errors** | Use Node.js `>=20.20.2 <21`, `>=22.22.2 <23`, or `>=24.0.0 <25` (24 LTS recommended) | -| **`system-info` for bug reports** | Run `npm run system-info` and attach `system-info.txt` to your issue | +For users in **Russia, China, Iran, Cuba, Turkey** and other regions: -๐Ÿ“– **Full troubleshooting guide:** [`docs/TROUBLESHOOTING.md`](docs/TROUBLESHOOTING.md) +- **4-level outbound proxy** โ€” account / provider / combo / global scopes +- **1proxy free marketplace** โ€” auto-syncs working HTTP/SOCKS5 proxies +- **Anti-detection** โ€” TLS fingerprinting (JA3/JA4), CCH headshakes, header sanitization +- **Public tunnels** โ€” Cloudflare (Quick or Named), ngrok, Tailscale Funnel for OAuth callbacks + +โ†’ [`docs/PROXY_GUIDE.md`](docs/PROXY_GUIDE.md) ยท [`docs/TUNNELS_GUIDE.md`](docs/TUNNELS_GUIDE.md) ยท [`docs/STEALTH_GUIDE.md`](docs/STEALTH_GUIDE.md) --- -## ๐Ÿ“Š Performance Benchmarks +## ๐Ÿ“ฑ Multi-Platform -> OmniRoute is optimized for speed with minimal overhead. Run your own benchmarks to validate performance on your hardware. +| Platform | Install | Doc | +| --------------------------- | -------------------------------------------------------- | --------------------------------------------------------------- | +| **CLI / Server** | `npm install -g omniroute` | [`SETUP_GUIDE.md`](docs/SETUP_GUIDE.md) | +| **Desktop (Win/Mac/Linux)** | Electron installer from GitHub Releases | [`ELECTRON_GUIDE.md`](docs/ELECTRON_GUIDE.md) | +| **PWA** | Install from any modern browser | [`PWA_GUIDE.md`](docs/PWA_GUIDE.md) | +| **Android (Termux)** | `pkg install nodejs-lts && npm i -g omniroute` | [`TERMUX_GUIDE.md`](docs/TERMUX_GUIDE.md) | +| **Docker** | `docker compose up` (base/cli/host/cliproxyapi profiles) | [`DOCKER_GUIDE.md`](docs/DOCKER_GUIDE.md) | +| **VM / VPS** | Generic Ubuntu/Debian + nginx + systemd | [`VM_DEPLOYMENT_GUIDE.md`](docs/VM_DEPLOYMENT_GUIDE.md) | +| **Fly.io** | `fly deploy` | [`FLY_IO_DEPLOYMENT_GUIDE.md`](docs/FLY_IO_DEPLOYMENT_GUIDE.md) | -### Benchmark Your Instance +--- -```bash -# Install load testing tool -npm install -g autocannon +## ๐Ÿงฉ Extensibility -# Test local OmniRoute instance -autocannon -c 10 -d 30 -m POST \ - -H "Authorization: Bearer your-api-key" \ - -H "Content-Type: application/json" \ - -b '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hi"}]}' \ - http://localhost:20128/v1/chat/completions +| System | What it does | Docs | +| ----------------------- | ----------------------------------------------------------------------- | -------------------------------------------- | +| ๐Ÿง  **Skills** | Built-in skills + marketplace + sandboxed custom skills (Docker) | [`docs/SKILLS.md`](docs/SKILLS.md) | +| ๐Ÿ’พ **Memory** | Persistent conversational memory (SQLite FTS5 + Qdrant vector) | [`docs/MEMORY.md`](docs/MEMORY.md) | +| โ˜๏ธ **Cloud Agents** | Submit long tasks to Codex Cloud / Devin / Jules | [`docs/CLOUD_AGENT.md`](docs/CLOUD_AGENT.md) | +| ๐Ÿช **Webhooks** | HMAC-signed event delivery (request.completed, quota.exceeded, etc.) | [`docs/WEBHOOKS.md`](docs/WEBHOOKS.md) | +| ๐Ÿ›ก๏ธ **Guardrails** | PII masker, prompt injection guard, vision bridge โ€” hot-reload | [`docs/GUARDRAILS.md`](docs/GUARDRAILS.md) | +| ๐Ÿงช **Evals** | Suite-based regression testing (combos/models/cases/rubrics) | [`docs/EVALS.md`](docs/EVALS.md) | +| ๐Ÿ” **Compliance/Audit** | `audit_log` table, retention, `noLog` opt-out, SSRF logging | [`docs/COMPLIANCE.md`](docs/COMPLIANCE.md) | +| ๐Ÿ›ก๏ธ **MCP Server** | 37 tools, 3 transports (stdio/SSE/Streamable HTTP), ~13 scopes | [`docs/MCP-SERVER.md`](docs/MCP-SERVER.md) | +| ๐Ÿค **A2A Protocol** | v0.3 JSON-RPC, 5 skills (smart-routing, quota, discovery, cost, health) | [`docs/A2A-SERVER.md`](docs/A2A-SERVER.md) | + +--- + +## ๐Ÿ“š Documentation + +Everything you need, organized by area. + +### ๐Ÿš€ Start here + +- [`SETUP_GUIDE.md`](docs/SETUP_GUIDE.md) โ€” install + connect first provider +- [`USER_GUIDE.md`](docs/USER_GUIDE.md) โ€” end-user manual (modes, combos, CLIs, audio, ~1200 lines) +- [`FREE_TIERS.md`](docs/FREE_TIERS.md) โ€” start free, no card +- [`TROUBLESHOOTING.md`](docs/TROUBLESHOOTING.md) โ€” common issues + v3.8 known issues + +### ๐Ÿ›๏ธ Architecture + +- [`ARCHITECTURE.md`](docs/ARCHITECTURE.md) โ€” high-level architecture +- [`CODEBASE_DOCUMENTATION.md`](docs/CODEBASE_DOCUMENTATION.md) โ€” engineering reference +- [`REPOSITORY_MAP.md`](docs/REPOSITORY_MAP.md) โ€” every directory and root file +- [`FEATURES.md`](docs/FEATURES.md) โ€” full feature matrix + +### ๐Ÿ”Œ API & contracts + +- [`API_REFERENCE.md`](docs/API_REFERENCE.md) โ€” endpoint reference +- [`openapi.yaml`](docs/openapi.yaml) โ€” OpenAPI 3.0 spec +- [`PROVIDER_REFERENCE.md`](docs/PROVIDER_REFERENCE.md) โ€” full catalog (auto-generated) +- [`CLI-TOOLS.md`](docs/CLI-TOOLS.md) โ€” CLI integrations + internal CLI +- [`ENVIRONMENT.md`](docs/ENVIRONMENT.md) โ€” all env vars + +### ๐ŸŽฏ Routing & resilience + +- [`AUTO-COMBO.md`](docs/AUTO-COMBO.md) โ€” Auto-Combo (9-factor scoring, 14 strategies) +- [`RESILIENCE_GUIDE.md`](docs/RESILIENCE_GUIDE.md) โ€” circuit breaker + cooldown + lockout +- [`REASONING_REPLAY.md`](docs/REASONING_REPLAY.md) โ€” reasoning cache for DeepSeek/Kimi/Qwen +- [`STEALTH_GUIDE.md`](docs/STEALTH_GUIDE.md) โ€” TLS fingerprinting + obfuscation + +### ๐Ÿค– Agent protocols + +- [`AGENT_PROTOCOLS_GUIDE.md`](docs/AGENT_PROTOCOLS_GUIDE.md) โ€” A2A vs ACP vs Cloud Agents +- [`MCP-SERVER.md`](docs/MCP-SERVER.md) โ€” Model Context Protocol server +- [`A2A-SERVER.md`](docs/A2A-SERVER.md) โ€” Agent-to-Agent protocol +- [`CLOUD_AGENT.md`](docs/CLOUD_AGENT.md) โ€” Codex Cloud / Devin / Jules + +### ๐Ÿง  Extensions + +- [`SKILLS.md`](docs/SKILLS.md) โ€” Skills framework +- [`MEMORY.md`](docs/MEMORY.md) โ€” Memory system +- [`EVALS.md`](docs/EVALS.md) โ€” Eval framework +- [`GUARDRAILS.md`](docs/GUARDRAILS.md) โ€” PII / injection / vision +- [`WEBHOOKS.md`](docs/WEBHOOKS.md) โ€” Webhook delivery +- [`COMPLIANCE.md`](docs/COMPLIANCE.md) โ€” Audit + retention +- [`AUTHZ_GUIDE.md`](docs/AUTHZ_GUIDE.md) โ€” Authorization pipeline + +### ๐Ÿ—œ๏ธ Compression + +- [`COMPRESSION_GUIDE.md`](docs/COMPRESSION_GUIDE.md) +- [`COMPRESSION_ENGINES.md`](docs/COMPRESSION_ENGINES.md) +- [`COMPRESSION_RULES_FORMAT.md`](docs/COMPRESSION_RULES_FORMAT.md) +- [`COMPRESSION_LANGUAGE_PACKS.md`](docs/COMPRESSION_LANGUAGE_PACKS.md) +- [`RTK_COMPRESSION.md`](docs/RTK_COMPRESSION.md) + +### ๐Ÿš€ Deployment + +- [`DOCKER_GUIDE.md`](docs/DOCKER_GUIDE.md) +- [`VM_DEPLOYMENT_GUIDE.md`](docs/VM_DEPLOYMENT_GUIDE.md) +- [`FLY_IO_DEPLOYMENT_GUIDE.md`](docs/FLY_IO_DEPLOYMENT_GUIDE.md) +- [`ELECTRON_GUIDE.md`](docs/ELECTRON_GUIDE.md) +- [`PWA_GUIDE.md`](docs/PWA_GUIDE.md) +- [`TERMUX_GUIDE.md`](docs/TERMUX_GUIDE.md) +- [`TUNNELS_GUIDE.md`](docs/TUNNELS_GUIDE.md) +- [`PROXY_GUIDE.md`](docs/PROXY_GUIDE.md) + +### ๐Ÿ“‹ Operations + +- [`RELEASE_CHECKLIST.md`](docs/RELEASE_CHECKLIST.md) โ€” release flow with Claude Code skills +- [`COVERAGE_PLAN.md`](docs/COVERAGE_PLAN.md) โ€” test coverage state (current: 82.58%/82.58%/84.23%/75.22%) +- [`I18N.md`](docs/I18N.md) โ€” 30 supported locales +- [`UNINSTALL.md`](docs/UNINSTALL.md) + +### ๐Ÿค Contributing & policy + +- [`CONTRIBUTING.md`](CONTRIBUTING.md) โ€” contributor guide +- [`SECURITY.md`](SECURITY.md) โ€” security policy +- [`CODE_OF_CONDUCT.md`](CODE_OF_CONDUCT.md) +- [`CLAUDE.md`](CLAUDE.md) โ€” rules for Claude Code agents +- [`AGENTS.md`](AGENTS.md) โ€” rules for non-Claude agents +- [`GEMINI.md`](GEMINI.md) โ€” rules for Gemini agents + +--- + +## ๐Ÿ’ก Use Cases + +| Scenario | Solution | +| --------------------------------- | ------------------------------------------------------------- | +| "Claude Pro user, hit rate limit" | Combo: Claude โ†’ GLM โ†’ DeepSeek (auto-fallback) | +| "Want $0 forever" | `auto/cheap` โ†’ Kiro/Qoder/Pollinations fallback chain | +| "24/7 coding, no interruptions" | `auto/lkgp` (sticky to last-good) + Resilience | +| "Blocked region" | 1proxy free marketplace + Cloudflare Quick Tunnel | +| "Max token savings" | Stacked compression: `rtk โ†’ caveman` (78-95% on logs) | +| "Multi-agent system" | Expose OmniRoute as A2A node, route via `smart-routing` skill | +| "Long-running coding task" | Cloud Agents โ†’ Devin/Jules with management auth | + +โ†’ Detailed playbooks: [`USER_GUIDE.md`](docs/USER_GUIDE.md) ยท [`AUTO-COMBO.md`](docs/AUTO-COMBO.md) + +--- + +## ๐Ÿ“ก Protocols supported + +OmniRoute speaks all major AI protocols โ€” clients don't need to change: + +- **OpenAI** (Chat Completions, Responses, Embeddings, Images, Audio, Files, Batches, Rerank, Moderations) +- **Anthropic Messages** (Claude format, with thinking blocks + reasoning replay) +- **Google Gemini** (generateContent + Vertex) +- **Claude Code** (CLI-specific format with CCH + fingerprinting) +- **Cursor** (proprietary format with tool calls) +- **Kiro** (AWS Builder ID OAuth) +- **MCP** (Model Context Protocol โ€” 37 tools, stdio/SSE/Streamable HTTP) +- **A2A** (Agent-to-Agent v0.3 JSON-RPC โ€” agent card at `/.well-known/agent.json`) + +--- + +## ๐Ÿ—๏ธ Architecture (10-second tour) + +``` +Client โ†’ /v1/chat/completions โ†’ [CORS โ†’ Zod โ†’ Auth โ†’ Authz โ†’ Guardrails] + โ†’ handleChatCore() โ†’ [Cache โ†’ Rate limit โ†’ Combo routing] + โ†’ translateRequest โ†’ getExecutor โ†’ fetch upstream (with retry) + โ†’ response translation โ†’ SSE stream or JSON + โ†’ [Compliance audit] โ†’ response ``` -### Expected Performance +**Major pieces:** -| Metric | OmniRoute | LiteLLM | Bifrost (claimed) | -| -------------------- | --------- | ------- | ----------------- | -| **RPS (sustained)** | ~5,000+ | ~3,000 | ~13,925 | -| **Latency overhead** | <5ms | <10ms | <100ยตs | -| **Memory (idle)** | ~150MB | ~200MB | 32MB | -| **Cold start** | <2s | <5s | <1s | +- **`src/app/`** โ€” Next.js 16 App Router (60+ API routes + 30 dashboard pages) +- **`src/lib/`** โ€” 50+ domain modules (db, a2a, memory, skills, guardrails, evals, โ€ฆ) +- **`open-sse/`** โ€” Streaming engine workspace (31 executors, 9+8+9 translators, 80+ services, 37-tool MCP server) +- **`src/domain/`** โ€” Pure business logic (policies, fallback, cost rules) +- **`src/server/`** โ€” Server-only (authz pipeline, cors) -### What Impacts Performance - -- **Prompt compression** โ€” RTK/Caveman adds ~1-3ms latency but saves 15-95% tokens -- **Format translation** โ€” OpenAI โ†” Claude adds ~2-5ms for format conversion -- **Proxy overhead** โ€” Global proxy adds ~10-50ms depending on proxy location -- **Caching** โ€” Semantic cache hit avoids upstream entirely (0ms added) - -### Optimization Tips - -1. **Enable compression** โ€” Lite/Standard mode adds minimal latency for significant savings -2. **Use local providers** โ€” Same-region providers have lowest latency -3. **Configure circuit breakers** โ€” Prevent cascading failures during high load -4. **Enable semantic cache** โ€” Repeated queries hit cache without upstream call - -๐Ÿ“– **Full performance guide:** [`docs/PERFORMANCE.md`](docs/PERFORMANCE.md) +โ†’ Deep dive: [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md) ยท [`docs/CODEBASE_DOCUMENTATION.md`](docs/CODEBASE_DOCUMENTATION.md) --- -## ๐Ÿ† Competitive Comparison +## ๐ŸŒ i18n -> OmniRoute vs other AI Gateways โ€” See why developers choose OmniRoute +UI translated to **30 languages** with full RTL support for Arabic and Hebrew. -### Feature Comparison +๐ŸŒ [English](README.md) ยท [Portuguรชs](docs/i18n/pt-BR/README.md) ยท [Espaรฑol](docs/i18n/es/README.md) ยท [Franรงais](docs/i18n/fr/README.md) ยท [Deutsch](docs/i18n/de/README.md) ยท [ไธญๆ–‡](docs/i18n/zh-CN/README.md) ยท [ๆ—ฅๆœฌ่ชž](docs/i18n/ja/README.md) ยท [ํ•œ๊ตญ์–ด](docs/i18n/ko/README.md) ยท [ุงู„ุนุฑุจูŠุฉ](docs/i18n/ar/README.md) ยท [เคนเคฟเคจเฅเคฆเฅ€](docs/i18n/hi/README.md) ยท [ะ ัƒััะบะธะน](docs/i18n/ru/README.md) ยท [+ 19 more](docs/i18n/) -| Feature | OmniRoute | LiteLLM | Bifrost | Routerly | -| ------------------ | :--------------: | :-----: | :-----: | :------: | -| **Providers** | **160+** | 100+ | 15+ | 8 | -| **Token Savings** | **15-95%** | โŒ | โŒ | โŒ | -| **Free Tier** | **11 providers** | 1 | 0 | 0 | -| **Auto-Combo** | **6-factor** | โŒ | โŒ | basic | -| **MCP Server** | **37 tools** | โŒ | โœ… | โŒ | -| **A2A Protocol** | **โœ…** | โŒ | โŒ | โŒ | -| **Desktop App** | **โœ…** | โŒ | โŒ | โŒ | -| **Memory/Skills** | **โœ…** | โŒ | โŒ | โŒ | -| **i18n Languages** | **40+** | โŒ | โŒ | โŒ | -| **Price** | **Free** | $0 | $19/mo | $9.99/mo | - -### What Makes OmniRoute Unique - -#### ๐Ÿ—œ๏ธ **RTK+Caveman Compression** (Exclusive) - -No other gateway offers prompt compression. OmniRoute saves 15-95% tokens on eligible payloads using RTK (command-output) and Caveman (context condensation) engines. - -#### ๐Ÿ†“ **11 Free Providers** (Exclusive) - -- Kiro AI โ€” Claude Sonnet/Haiku (50 credits/month) -- Qoder AI โ€” Kimi-K2, Qwen3 unlimited -- Pollinations โ€” GPT-5, no API key needed -- LongCat โ€” 50M tokens/day -- 7 more - -#### ๐Ÿค– **Auto-Combo 6-Factor Routing** (Superior) - -OmniRoute scores each request on 7 factors (health, quota, cost, latency, capability, stability, tier) โ€” smarter than Routerly's basic LLM-based routing. - -#### ๐Ÿ”ง **MCP Server + A2A Protocol** (Exclusive) - -37 MCP tools for IDE integration and agent-to-agent communication. No other gateway offers both. - -#### ๐Ÿ–ฅ๏ธ **Desktop App** (Exclusive) - -Electron desktop app for Windows/macOS/Linux with system tray, auto-start, and offline mode. +โ†’ Adding a language: [`docs/I18N.md`](docs/I18N.md) --- -## ๐Ÿ†“ Free Forever โ€” No Credit Card Required +## ๐Ÿค Community -> Use OmniRoute's free providers for $0 lifetime access to AI. No API key needed for some, others just sign up free. - -### Free Provider Directory - -| Provider | Models | Quota | Auth Required | -| ----------------- | --------------------------------- | --------------- | -------------- | -| **Kiro (AWS)** | Claude Sonnet/Haiku | Unlimited | OAuth | -| **Qoder** | kimi-k2, qwen3-coder, deepseek-r1 | Unlimited | PAT token | -| **Pollinations** | GPT-5, Claude, Llama 4 | No key needed | โŒ | -| **Qwen Code** | qwen3-coder-plus | Unlimited | Device code | -| **LongCat** | Flash-Lite | 50M tokens/day | API key (free) | -| **Gemini CLI** | gemini-2.5-flash | 180K/mo | OAuth | -| **Cloudflare AI** | 50+ models | 10K neurons/day | โŒ | -| **Groq** | Llama 3.3 70B | 30 RPM | API key (free) | -| **NVIDIA NIM** | 129 models | ~40 RPM | API key (free) | -| **Cerebras** | Qwen3 235B | 1M tokens/day | API key (free) | -| **Scaleway** | Qwen3 235B | 1M tokens | API key (free) | - -### Total Free Capacity - -- **~31,000+ requests/day** combined -- **~32B+ tokens/month** -- **500+ models** -- **$0 forever** - -### Quick Start with Free Stack - -```bash -# Point any tool to OmniRoute's free endpoint -Base URL: http://localhost:20128/v1 -API Key: any-string -Model: auto # Auto-Combo picks best free model -``` - -Or use the built-in **Free Stack** combo in `/dashboard/combos` โ€” round-robins all free providers automatically. - -๐Ÿ“– **Full free provider guide:** [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md) +- ๐ŸŒ **Website:** [omniroute.online](https://omniroute.online) +- ๐Ÿ“ฆ **npm:** [omniroute](https://www.npmjs.com/package/omniroute) +- ๐Ÿณ **Docker Hub:** [diegosouzapw/omniroute](https://hub.docker.com/r/diegosouzapw/omniroute) +- ๐Ÿ’ฌ **WhatsApp (BR):** Brazilian community group โ€” see README link +- ๐Ÿ› **Issues:** [GitHub Issues](https://github.com/diegosouzapw/OmniRoute/issues) +- ๐Ÿ’ก **Discussions:** [GitHub Discussions](https://github.com/diegosouzapw/OmniRoute/discussions) --- -## ๐Ÿ› ๏ธ Tech Stack +## โค๏ธ Contributing -
-Click to expand tech stack details +We welcome PRs! Start with: -- **Runtime**: Node.js 20.20.2+, 22.22.2+, or 24.x LTS (24 LTS recommended) -- **Language**: TypeScript 5.9 โ€” **100% TypeScript** across `src/` and `open-sse/` (zero `any` in core modules since v2.0) -- **Framework**: Next.js 16 + React 19 + Tailwind CSS 4 -- **Database**: better-sqlite3 (SQLite) + LowDB (JSON legacy) โ€” domain state, proxy logs, MCP audit, routing decisions, memory, skills -- **Schemas**: Zod (MCP tool I/O validation, API contracts) -- **Protocols**: MCP (stdio/HTTP) + A2A v0.3 (JSON-RPC 2.0 + SSE) -- **Streaming**: Server-Sent Events (SSE) + WebSocket bridge (`/v1/ws`) -- **Auth**: OAuth 2.0 (PKCE) + JWT + API Keys + MCP Scoped Authorization -- **Testing**: Node.js test runner + Vitest (**4,690+ test cases** across 517 files โ€” unit, integration, E2E, security, ecosystem) -- **Platforms**: Desktop (Electron), Android (Termux), PWA (any browser) -- **CI/CD**: GitHub Actions (auto npm publish + Docker Hub on release) -- **Website**: [omniroute.online](https://omniroute.online) -- **Package**: [npmjs.com/package/omniroute](https://www.npmjs.com/package/omniroute) -- **Docker**: [hub.docker.com/r/diegosouzapw/omniroute](https://hub.docker.com/r/diegosouzapw/omniroute) -- **Resilience**: Circuit breaker, exponential backoff, anti-thundering herd, TLS spoofing, auto-combo self-healing +1. Read [`CONTRIBUTING.md`](CONTRIBUTING.md) โ€” setup, conventional commits, testing +2. Pick an issue labeled [`good first issue`](https://github.com/diegosouzapw/OmniRoute/labels/good%20first%20issue) +3. Branch from `main` (`feat/*`, `fix/*`, `docs/*`, `refactor/*`, `test/*`, `chore/*`) +4. Hooks will run lint + test on commit/push -
+**Adding a provider?** [`docs/ARCHITECTURE.md ยง Adding a New Provider`](docs/ARCHITECTURE.md) +**Adding an MCP tool?** [`docs/MCP-SERVER.md`](docs/MCP-SERVER.md) +**Adding an A2A skill?** [`docs/A2A-SERVER.md ยง Adding a New Skill`](docs/A2A-SERVER.md) --- -## ๐Ÿ“– Documentation +## ๐Ÿ”’ Security -### ๐Ÿ“˜ Getting Started - -| Document | Description | -| ------------------------------------- | ----------------------------------------------------------------------------- | -| [User Guide](docs/USER_GUIDE.md) | Providers, combos, CLI integration, deployment | -| [Setup Guide](docs/SETUP_GUIDE.md) | Full install methods, CLI tool configs, protocol setup, timeout tuning | -| [CLI Tools Guide](docs/CLI-TOOLS.md) | Per-tool setup for Claude Code, Codex, Cursor, Cline, OpenClaw, Kilo, Copilot | -| [Quick Start](README.md#-quick-start) | 3-step install โ†’ connect โ†’ configure | - -### ๐Ÿ”ง Operations & Deployment - -| Document | Description | -| ---------------------------------------------------- | -------------------------------------------------------------- | -| [Docker Guide](docs/DOCKER_GUIDE.md) | Docker run, Compose profiles, Caddy HTTPS, tunnels, image tags | -| [VM Deployment](docs/VM_DEPLOYMENT_GUIDE.md) | Complete guide: VM + nginx + Cloudflare setup | -| [Fly.io Deployment](docs/FLY_IO_DEPLOYMENT_GUIDE.md) | Deploy to Fly.io with persistent storage | -| [Termux Guide](docs/TERMUX_GUIDE.md) | Run OmniRoute on Android via Termux | -| [PWA Guide](docs/PWA_GUIDE.md) | Progressive Web App install, caching, architecture | -| [Uninstall Guide](docs/UNINSTALL.md) | Clean removal for all install methods | -| [Environment Config](docs/ENVIRONMENT.md) | Complete `.env` variables and references | - -### ๐Ÿง  Features & Architecture - -| Document | Description | -| ---------------------------------------------------------------- | ----------------------------------------------------------------------------- | -| [Architecture](docs/ARCHITECTURE.md) | System architecture, data flow, and internals | -| [Compression Guide](docs/COMPRESSION_GUIDE.md) | 7-option pipeline: off / lite / standard / aggressive / ultra / RTK / stacked | -| [RTK Compression](docs/RTK_COMPRESSION.md) | Command-output compression, filters, trust, verify, raw-output recovery | -| [Compression Engines](docs/COMPRESSION_ENGINES.md) | Caveman, RTK, stacked pipelines, dashboard/API/MCP surfaces | -| [Compression Rules Format](docs/COMPRESSION_RULES_FORMAT.md) | JSON rule-pack schemas for Caveman and RTK filters | -| [Compression Language Packs](docs/COMPRESSION_LANGUAGE_PACKS.md) | Language detection and Caveman rule-pack authoring | -| [Resilience Guide](docs/RESILIENCE_GUIDE.md) | Circuit breakers, cooldowns, queue, anti-thundering herd, TLS spoofing | -| [Auto-Combo Engine](docs/AUTO-COMBO.md) | 6-factor scoring, mode packs, self-healing | -| [Proxy Guide](docs/PROXY_GUIDE.md) | 3-level proxy system, 1proxy marketplace, registry CRUD | -| [Free Tiers](docs/FREE_TIERS.md) | 25+ free API providers consolidated directory | -| [Features Gallery](docs/FEATURES.md) | Visual dashboard tour with screenshots | -| [Codebase Documentation](docs/CODEBASE_DOCUMENTATION.md) | Beginner-friendly codebase walkthrough | - -### ๐Ÿค– Protocols & APIs - -| Document | Description | -| ------------------------------------------- | --------------------------------------------------- | -| [API Reference](docs/API_REFERENCE.md) | All endpoints with examples | -| [OpenAPI Spec](docs/openapi.yaml) | OpenAPI 3.0 specification | -| [MCP Server](open-sse/mcp-server/README.md) | 29 MCP tools, IDE configs, Python/TS/Go clients | -| [MCP Server Guide](docs/MCP-SERVER.md) | MCP installation, transports, and tool reference | -| [A2A Server](src/lib/a2a/README.md) | JSON-RPC 2.0 protocol, skills, streaming, task mgmt | -| [A2A Server Guide](docs/A2A-SERVER.md) | A2A agent card, tasks, skills, and streaming | - -### ๐Ÿ“‹ Project & Quality - -| Document | Description | -| ---------------------------------------------- | ----------------------------------------------- | -| [Contributing](CONTRIBUTING.md) | Development setup and guidelines | -| [Security Policy](SECURITY.md) | Vulnerability reporting and security practices | -| [i18n Guide](docs/I18N.md) | 40+ language support, translation workflow, RTL | -| [Release Checklist](docs/RELEASE_CHECKLIST.md) | Pre-release validation steps | -| [Coverage Plan](docs/COVERAGE_PLAN.md) | Test coverage strategy and 4,690+ test suite | - ---- - -## โญ Top Contributors - -> OmniRoute is shaped by a passionate open-source community. These individuals have made exceptional contributions that directly impact the quality, stability, and reach of the project. **Thank you.** - - - - - - - - - -
- - oyi77
- oyi77 -

- ๐Ÿฅ‡ 190 commits โ€ข +72K lines
- Analytics engine, SQL aggregations,
proxy marketplace, test coverage
-
- - Chris Staley
- Chris Staley -

- ๐Ÿฅˆ 72 commits โ€ข +5.7K lines
- SSE stream hardening, Responses API,
Gemini pagination, test regression fixes
-
- - zenobit
- zenobit -

- ๐Ÿฅ‰ 62 commits โ€ข +24K lines
- CI/CD pipeline, i18n for 33 languages,
Void Linux package, platform fixes
-
- - R.D. & Randi
- R.D. & Randi -

- ๐Ÿ… 107 commits โ€ข +28K lines
- Endpoints page, tunnel integrations,
Docker workflows, A2A status, compression UI
-
- - benzntech
- benzntech -

- ๐Ÿ… 20 commits โ€ข +7.5K lines
- Electron desktop app, auto-updater,
release build workflows, cross-platform CI
-
- -> ๐Ÿ™ These contributors' features, bug fixes, and infrastructure improvements are a **core part** of what makes OmniRoute reliable and feature-rich. Every pull request, every test case, and every i18n translation file matters. Open source is built by people like them. - ---- - -## ๐Ÿ‘ฅ Contributors - -[![Contributors](https://contrib.rocks/image?repo=diegosouzapw/OmniRoute&max=100&columns=20&anon=1)](https://github.com/diegosouzapw/OmniRoute/graphs/contributors) - -### How to Contribute - -1. Fork the repository -2. Create your feature branch (`git checkout -b feature/amazing-feature`) -3. Commit your changes (`git commit -m 'Add amazing feature'`) -4. Push to the branch (`git push origin feature/amazing-feature`) -5. Open a Pull Request - -See [CONTRIBUTING.md](CONTRIBUTING.md) for detailed guidelines. - -### Releasing a New Version - -```bash -# Create a release โ€” npm publish happens automatically -gh release create v2.0.0 --title "v2.0.0" --generate-notes -``` - ---- - -## ๐Ÿ“Š Star History - - - - - - Star History Chart - - - -## ๐ŸŒ StarMapper - - - - - - StarMapper - - - -## ๐Ÿ™ Acknowledgments - -Special thanks to **[9router](https://github.com/decolua/9router)** by **[decolua](https://github.com/decolua)** โ€” the original project that inspired this fork. OmniRoute builds upon that incredible foundation with additional features, multi-modal APIs, and a full TypeScript rewrite. - -Special thanks to **[CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI)** by **[router-for-me](https://github.com/router-for-me)** โ€” the original Go implementation that inspired this JavaScript port. - -Special thanks to **[Caveman](https://github.com/JuliusBrussee/caveman)** by **[JuliusBrussee](https://github.com/JuliusBrussee)** (โญ 51K+) โ€” the viral "why use many token when few token do trick" project whose caveman-speak compression philosophy inspired OmniRoute's standard compression mode and 30+ filler/condensation regex rules. - -Special thanks to **[RTK - Rust Token Killer](https://github.com/rtk-ai/rtk)** by **[RTK AI](https://github.com/rtk-ai)** โ€” the high-performance command-output compression project whose terminal, build, test, git, and tool-output filtering model inspired OmniRoute's RTK engine, JSON filter DSL, raw-output recovery, and stacked RTK โ†’ Caveman compression pipeline. +- **Reporting:** see [`SECURITY.md`](SECURITY.md) for disclosure policy +- **Supported versions:** 3.8.x (Active), 3.7.x (Security only) +- **Secrets:** never commit. Use `.env` (auto-generated from `.env.example` on first install) or vaults +- **Encryption:** credentials at rest with AES-256-GCM +- **Authz:** route-aware classification (`src/server/authz/`) โ€” see [`docs/AUTHZ_GUIDE.md`](docs/AUTHZ_GUIDE.md) +- **Guardrails:** PII masking, prompt injection detection โ€” hot-reloadable --- ## ๐Ÿ“„ License -MIT License - see [LICENSE](LICENSE) for details. +[MIT](LICENSE) ยฉ 2025-2026 [Diego Souza](https://github.com/diegosouzapw) + +Free forever. Self-hosted. No tracking. No cloud lock-in. ---
- Built with โค๏ธ for developers who code 24/7 -
- omniroute.online + +**[โฌ† Back to top](#-omniroute)** ยท Built with โค๏ธ for the open-source AI community. + +OmniRoute v3.8.0 ยท Node โ‰ฅ20.20.2 ยท MIT License +
- diff --git a/docs/REPOSITORY_MAP.md b/docs/REPOSITORY_MAP.md new file mode 100644 index 0000000000..e3c709ec0a --- /dev/null +++ b/docs/REPOSITORY_MAP.md @@ -0,0 +1,506 @@ +# Repository Map + +> **One-line description for every directory and root file.** +> Last updated: 2026-05-13 โ€” OmniRoute v3.8.0 +> +> Use this map to navigate the codebase quickly. For deep dives, follow links to dedicated docs. + +## Top-level tree + +``` +OmniRoute/ +โ”œโ”€โ”€ src/ # Next.js 16 application (UI + API routes + libs + domain + server) +โ”œโ”€โ”€ open-sse/ # Streaming engine workspace (handlers, executors, translator, MCP server) +โ”œโ”€โ”€ electron/ # Desktop wrapper (Electron 41 + electron-builder 26.10) +โ”œโ”€โ”€ bin/ # CLI entry point and command handlers +โ”œโ”€โ”€ scripts/ # Build, check, sync, and one-off scripts +โ”œโ”€โ”€ docs/ # Public documentation (you are here) +โ”œโ”€โ”€ tests/ # All test suites (unit, integration, e2e, protocols-e2e) +โ”œโ”€โ”€ public/ # Next.js static assets, PWA manifest, service worker, icons +โ”œโ”€โ”€ config/ # Static config files +โ”œโ”€โ”€ images/ # Marketing / README image assets +โ”œโ”€โ”€ .github/ # GitHub Actions workflows + issue templates + PR template +โ”œโ”€โ”€ .husky/ # Git hooks (pre-commit, pre-push) +โ”œโ”€โ”€ .claude/ # Claude Code slash commands (project-scoped) +โ”œโ”€โ”€ .agents/ # Codex / generic agent workflows + skills (mirror of .claude/) +โ”œโ”€โ”€ .vscode/ # VS Code workspace settings +โ”œโ”€โ”€ _ideia/ # Planning notes (informal; not shipped) +โ”œโ”€โ”€ _mono_repo/ # Historic subprojects (cloud, site, vscode-extension) +โ”œโ”€โ”€ _references/ # Read-only reference clones from related OSS projects +โ”œโ”€โ”€ _tasks/ # Per-release task tracking files (informal) +โ”œโ”€โ”€ .issues/ # Local issue cache (gitignored) +โ”œโ”€โ”€ .playwright-mcp/ # Playwright MCP test artifacts +โ”œโ”€โ”€ coverage/ # c8 coverage output (gitignored) +โ”œโ”€โ”€ logs/ # Runtime logs (gitignored) +โ”œโ”€โ”€ node_modules/ # Dependencies (gitignored) +โ”œโ”€โ”€ package/ # npm pack staging area (build artifact) +โ”œโ”€โ”€ .next/ # Next.js build output (gitignored) +โ””โ”€โ”€ (root files โ€” see below) +``` + +--- + +## Root files + +| File | Purpose | +| ------------------------------------------- | --------------------------------------------------------------------------------- | +| **README.md** | Marketing landing page + quick start + feature matrix (see also `llm.txt`) | +| **CHANGELOG.md** | Per-release changelog (auto-generated by `/version-bump-cc` skill) | +| **LICENSE** | MIT license text | +| **CLAUDE.md** | Project rules for Claude Code agents (hard rules, conventions, scenarios) | +| **AGENTS.md** | Same as CLAUDE.md but for non-Claude AI agents (Codex, Cursor, etc.) | +| **GEMINI.md** | Concise rules for Gemini-based agents (subset of CLAUDE.md) | +| **CONTRIBUTING.md** | Contributor guide: setup, conventional commits, testing, PR flow | +| **SECURITY.md** | Vulnerability reporting policy, supported versions, threat model | +| **CODE_OF_CONDUCT.md** | Contributor Covenant โ€” community behavior expectations | +| **llm.txt** | Plain-text landing optimized for LLM crawlers (SEO for AI assistants) | +| **Tuto_Qdrant.MD** | Standalone tutorial for enabling Qdrant vector memory (see also `docs/MEMORY.md`) | +| **package.json** | npm manifest, scripts, dependencies, engines, c8 coverage gate | +| **package-lock.json** | Locked dependency tree | +| **tsconfig.json** | Root TypeScript config | +| **tsconfig.typecheck-core.json** | Typecheck config for `src/` core | +| **tsconfig.typecheck-noimplicit-core.json** | Strict (`noImplicitAny`) typecheck | +| **tsconfig.tsbuildinfo** | TS incremental build cache (gitignored) | +| **next.config.mjs** | Next.js 16 build configuration (standalone output) | +| **next-env.d.ts** | Next.js auto-generated env types | +| **eslint.config.mjs** | ESLint flat config (rules per project area) | +| **prettier.config.mjs** | Prettier formatting rules | +| **postcss.config.mjs** | PostCSS config for Tailwind/CSS pipeline | +| **playwright.config.ts** | Playwright E2E test config | +| **vitest.config.ts** | Vitest config (default suite) | +| **vitest.mcp.config.ts** | Vitest config for MCP server / autoCombo / cache suites | +| **sonar-project.properties** | SonarQube/SonarCloud config (code quality) | +| **Dockerfile** | Multi-stage Docker build (builder โ†’ runner-base โ†’ runner-cli) | +| **docker-compose.yml** | Dev compose with 4 profiles (base, cli, host, cliproxyapi) + redis sidecar | +| **docker-compose.prod.yml** | Production compose (port 20130, redis, named volumes) | +| **.dockerignore** | Files excluded from Docker context | +| **fly.toml** | Fly.io deployment config (region `sin`, port 20128, /data volume) | +| **.env.example** | Template env file (815 lines, auto-copied to `.env` on first install) | +| **.gitignore** | Git ignore patterns | +| **.npmignore** | npm publish exclusion list | +| **.npmrc** | npm config (registry, lockfile policy) | +| **.node-version** | Node version pin (used by nvm-compatible tools) | +| **.nvmrc** | Node version pin for nvm | + +--- + +## `src/` โ€” Next.js application + +``` +src/ +โ”œโ”€โ”€ app/ # App Router (pages + API routes + status pages + landing) +โ”œโ”€โ”€ lib/ # Core libraries / domain modules (~50 subdirs + ~30 top-level files) +โ”œโ”€โ”€ domain/ # Pure domain logic (policy engine, fallback, cost, lockout, comboResolver, assessment) +โ”œโ”€โ”€ server/ # Server-only modules (authz pipeline, cors, auth middleware) โ€” cannot import from client +โ”œโ”€โ”€ shared/ # Shared between server and client where safe (constants, types, validation, contracts, utils) +โ”œโ”€โ”€ i18n/ # next-intl config + per-locale message JSON (30+ locales) +โ”œโ”€โ”€ middleware/ # Next.js middleware (request enrichment, locale detection) +โ”œโ”€โ”€ mitm/ # MITM proxy helpers (Linux cert install, antigravity stealth) +โ”œโ”€โ”€ models/ # Model adapter glue (legacy shim) +โ”œโ”€โ”€ scripts/ # In-tree maintenance scripts (e.g., backfillAggregation) +โ”œโ”€โ”€ sse/ # Legacy SSE handlers/services (chat.ts, chatHelpers.ts, services/auth.ts) +โ”œโ”€โ”€ store/ # Legacy in-memory store (being phased out for src/lib/db) +โ”œโ”€โ”€ types/ # Shared TS type files +โ”œโ”€โ”€ instrumentation.ts # Next.js telemetry hook (browser + edge) +โ”œโ”€โ”€ instrumentation-node.ts # Node-only instrumentation +โ”œโ”€โ”€ server-init.ts # Server bootstrap (DB migrations, jobs, cleanup) +โ””โ”€โ”€ proxy.ts # HTTP-proxy entry shim +``` + +### `src/app/` โ€” App Router (Next.js 16) + +| Path | Purpose | +| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `app/api/v1/` | Public OpenAI-compat API (~25 sub-routes: chat, completions, embeddings, files, batches, audio, images, videos, music, rerank, moderations, search, ws, agents, accounts, providers, etc.) | +| `app/api/v1beta/` | Gemini-style API endpoints | +| `app/api/` (non-v1) | Management/admin routes (~60 directories: providers, combos, settings, mcp, a2a, evals, memory, skills, webhooks, compliance, resilience, monitoring, tunnels, cli-tools, etc.) | +| `app/a2a/` | A2A JSON-RPC 2.0 entry point (`POST /a2a`) | +| `app/.well-known/agent.json/` | A2A Agent Card (discovery) | +| `app/(dashboard)/dashboard/` | Dashboard UI pages (~30 pages: providers, combos, settings, memory, skills, webhooks, evals, audit, batch, cache, costs, health, system, etc.) | +| `app/docs/` | Embedded documentation viewer (renders `docs/*.md`) | +| `app/landing/` | Marketing landing page | +| `app/login/`, `forgot-password/`, `forbidden/` | Auth-related pages | +| `app/{400,401,403,408,429,500,502,503}/` | HTTP error pages | +| `app/maintenance/`, `offline/`, `status/`, `privacy/`, `terms/`, `callback/` | Static/status pages | +| `app/layout.tsx`, `page.tsx`, `manifest.ts`, `globals.css` | Root layout, home, PWA manifest, global CSS | +| `app/error.tsx`, `global-error.tsx`, `not-found.tsx`, `loading.tsx` | Error boundaries | + +### `src/lib/` โ€” Core libraries (~50 modules) + +| Module | Purpose | +| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `a2a/` | A2A protocol task manager, skills (5), streaming | +| `acp/` | CLI Agent Registry (local CLI discovery โ€” see `docs/AGENT_PROTOCOLS_GUIDE.md`) | +| `api/` | Shared API helpers (`requireManagementAuth`, validation) | +| `auth/` | Session, password hashing, token validation | +| `batches/` | OpenAI Batches API handlers | +| `catalog/` | Provider catalog Zod validation + capability resolution | +| `cloudAgent/` | Cloud Agents (Codex Cloud, Devin, Jules) โ€” see `docs/CLOUD_AGENT.md` | +| `combos/` | Combo resolution + reorder helpers | +| `compliance/` | Audit log + provider audit โ€” see `docs/COMPLIANCE.md` | +| `compression/` | Compression engine glue (engines live in `open-sse/services/compression/`) | +| `config/` | Runtime config helpers | +| `db/` | 45+ domain DB modules + 55 migrations (always go through here for SQLite) | +| `display/` | UI formatting helpers (cost, latency, etc.) | +| `embeddings/` | Embeddings service helpers | +| `env/` | Env variable parsing + validation | +| `evals/` | Eval framework (suites, runner, runtime) โ€” see `docs/EVALS.md` | +| `guardrails/` | PII masker, prompt injection, vision bridge โ€” see `docs/GUARDRAILS.md` | +| `jobs/` | Background jobs (cron-like) | +| `memory/` | Conversational memory (SQLite FTS5 + Qdrant) โ€” see `docs/MEMORY.md` | +| `monitoring/` | Health checks, metrics emission | +| `oauth/` | OAuth flows for 14 providers (claude, codex, antigravity, cursor, github, gemini, kimi-coding, kilocode, cline, qwen, kiro, qoder, gitlab-duo, windsurf) | +| `plugins/` | Plugin registry | +| `promptCache/` | Anthropic-style prompt cache breakpoints | +| `skills/` | Skills framework (built-in + marketplace + SkillsSH) โ€” see `docs/SKILLS.md` | +| `webhookDispatcher.ts` | HMAC webhook delivery โ€” see `docs/WEBHOOKS.md` | +| `cloudflaredTunnel.ts`, `ngrokTunnel.ts` | Tunnel managers โ€” see `docs/TUNNELS_GUIDE.md` | +| `oneproxySync.ts`, `oneproxyRotator.ts` | 1proxy free proxy marketplace โ€” see `docs/PROXY_GUIDE.md` | +| `cloudSync.ts`, `initCloudSync.ts` | Optional cloud sync of state | +| `localDb.ts` | Re-export barrel for db modules (no logic โ€” re-exports only) | +| `cacheLayer.ts`, `idempotencyLayer.ts` | Request caching + idempotency | +| (~30 more top-level files) | Specialized helpers (logEnv, modelsDevSync, piiSanitizer, etc.) | + +### `src/db/` โ€” Database (45+ modules + 55 migrations) + +| Subdir | Purpose | +| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `db/core.ts` | `getDbInstance()` singleton with WAL journaling | +| `db/migrations/` | 55 versioned SQL files (idempotent, transactional, numbered `001`..`055`) | +| `db/.ts` | One module per domain: providers, combos, apiKeys, users, sessions, usage, audit*log, webhooks, skills, memory_entries, cloud_agent_tasks, evals*\*, reasoning_cache, etc. | + +### `src/domain/` + +| Module | Purpose | +| ---------------------- | ----------------------------------------------------------------------- | +| `policy.ts` | Policy engine | +| `fallbackPolicy.ts` | Fallback decision tree | +| `costRules.ts` | Cost calculation rules | +| `lockoutPolicy.ts` | Model/connection lockout policy | +| `tagRouter.ts` | Tag-based routing | +| `comboResolver.ts` | Combo resolution (used by combo engine) | +| `modelAvailability.ts` | Per-model availability check | +| `assessment/` | Model assessment (Phase 1 of RFC-AUTO-ASSESSMENT โ€” see `docs/archive/`) | + +### `src/server/` + +| Module | Purpose | +| -------- | --------------------------------------------------------------------------------------- | +| `authz/` | Authorization pipeline: `classify` โ†’ `policies` โ†’ `enforce` โ€” see `docs/AUTHZ_GUIDE.md` | +| `cors/` | CORS configuration | +| `auth/` | Session middleware | + +### `src/shared/` + +| Module | Purpose | +| -------------------------------- | ------------------------------------------------------------- | +| `constants/providers.ts` | **177 providers** with Zod validation (source of truth) | +| `constants/cliTools.ts` | External CLI tool registry | +| `constants/routingStrategies.ts` | **14 routing strategies** with priorities | +| `constants/publicApiRoutes.ts` | Routes that require Bearer (vs management) auth | +| `constants/upstreamHeaders.ts` | Header denylist for upstream requests | +| `validation/schemas.ts` | ~80 Zod schemas (single source of truth for API contracts) | +| `validation/helpers.ts` | Zod validation helpers (`validateBody`, etc.) | +| `types/` | Shared TS types | +| `contracts/` | Public API contracts (consumed by `files:` in `package.json`) | +| `utils/circuitBreaker.ts` | Provider circuit breaker (see `docs/RESILIENCE_GUIDE.md`) | +| `utils/apiAuth.ts` | API key validation, scope checking | +| `utils/fetchTimeout.ts` | Timeout/abort wrappers for upstream fetch | + +--- + +## `open-sse/` โ€” Streaming Engine Workspace + +Separate npm workspace (`@omniroute/open-sse`). Handles request processing + provider execution. + +``` +open-sse/ +โ”œโ”€โ”€ handlers/ # 15 files (11 handlers + 4 helpers): chatCore, responsesHandler, embeddings, audio, image, video, music, rerank, moderations, search, etc. +โ”œโ”€โ”€ executors/ # 31 provider-specific executors (extend BaseExecutor) +โ”œโ”€โ”€ translator/ # Format converters (9 request, 8 response, 9 helpers) +โ”œโ”€โ”€ transformer/ # Responses API โ†” Chat Completions (TransformStream) +โ”œโ”€โ”€ services/ # ~80+ service modules (combo, accountFallback, autoCombo, reasoningCache, claude code/chatgpt stealth, modelDeprecation, taskAwareRouter, workflowFSM, etc.) +โ”œโ”€โ”€ mcp-server/ # MCP server (37 tools, 3 transports, ~13 scopes) +โ”œโ”€โ”€ config/ # Provider/model registries, header config, model aliases +โ”œโ”€โ”€ utils/ # TLS client, proxy fetch/dispatcher, network helpers +โ”œโ”€โ”€ index.ts # Workspace entry +โ”œโ”€โ”€ package.json # Workspace manifest +โ”œโ”€โ”€ tsconfig.json # Workspace TS config +โ””โ”€โ”€ types.d.ts # Workspace type declarations +``` + +### `open-sse/mcp-server/` + +| Path | Purpose | +| --------------------------- | -------------------------------------------------------------------- | +| `server.ts` | MCP server lifecycle (stdio + HTTP transports) | +| `httpTransport.ts` | HTTP Streamable + SSE transports (`/api/mcp/sse`, `/api/mcp/stream`) | +| `audit.ts` | Audit logging to `mcp_tool_audit` table | +| `scopeEnforcement.ts` | Per-tool scope validation | +| `runtimeHeartbeat.ts` | Health heartbeat to `DATA_DIR/runtime/mcp-heartbeat.json` | +| `descriptionCompressor.ts` | Compress tool description metadata to save context | +| `schemas/tools.ts` | 30 base tool definitions + scopes | +| `tools/advancedTools.ts` | Advanced tool implementations | +| `tools/memoryTools.ts` | 3 memory tools (search/add/clear) | +| `tools/skillTools.ts` | 4 skill tools (list/enable/execute/executions) | +| `tools/compressionTools.ts` | 5 compression tools | +| `README.md` | Internal MCP server README (cross-linked from `docs/MCP-SERVER.md`) | + +--- + +## `electron/` โ€” Desktop Wrapper + +| File | Purpose | +| ---------------- | --------------------------------------------------------------------------------- | +| `main.js` | Electron main process (BrowserWindow, embedded Next.js server, tray, auto-update) | +| `preload.js` | IPC bridge (contextBridge โ†’ `window.omniroute`) | +| `package.json` | electron-builder config + Electron 41 + electron-builder 26.10 deps | +| `assets/` | App icons (Windows .ico, macOS .icns, Linux .png) | +| `dist-electron/` | Build output (gitignored) | +| `types.d.ts` | Type declarations for renderer bridge | +| `README.md` | Internal Electron README (see also `docs/ELECTRON_GUIDE.md`) | + +--- + +## `bin/` โ€” CLI + +| File | Purpose | +| ----------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | +| `omniroute.mjs` | Main CLI entry โ€” `omniroute serve`, `omniroute setup`, `omniroute doctor`, `omniroute providers`, `omniroute combos`, etc. | +| `reset-password.mjs` | Standalone password reset CLI | +| `cli/commands/setup.mjs` | Interactive + non-interactive setup wizard | +| `cli/commands/doctor.mjs` | System health diagnostics (8+ checks) | +| `cli/commands/providers.mjs` | Provider list/test/validate | +| `cli/{args,data-dir,encryption,io,provider-catalog,provider-store,provider-test,settings-store,sqlite}.mjs` | CLI helper modules | +| `nodeRuntimeSupport.mjs` | Validate supported Node.js version on install | + +--- + +## `scripts/` โ€” Build & Check Scripts + +| Script | Purpose | +| ----------------------------------- | -------------------------------------------------------------------------- | +| `run-next.mjs` | Dev/start runner with env hydration | +| `build-next-isolated.mjs` | Standalone build (Next.js 16 standalone) | +| `prepublish.ts` | Package preparation before `npm pack` | +| `postinstall.mjs` | Auto-create `.env` from `.env.example` on first install | +| `sync-env.mjs` | Re-sync `.env` keys with `.env.example` | +| `check-cycles.mjs` | Detect circular dependencies | +| `check-route-validation.mjs` | Validate all API routes have Zod validation | +| `check-t11-any-budget.mjs` | Enforce explicit `any` budget per file | +| `check-docs-sync.mjs` | Validate docs version sync (existing pre-commit) | +| **`check-env-doc-sync.mjs`** | NEW: cross-check env vars in code vs `.env.example` vs `ENVIRONMENT.md` | +| **`check-docs-counts-sync.mjs`** | NEW: validate counts (executors, strategies, OAuth, A2A skills) match docs | +| **`check-deprecated-versions.mjs`** | NEW: flag stale versions/dates in docs | +| `check-supported-node-runtime.ts` | Validate current Node version is supported | +| `check-pr-test-policy.mjs` | Enforce "tests required" rule on production code changes | +| **`gen-provider-reference.ts`** | NEW: auto-generate `docs/PROVIDER_REFERENCE.md` from catalog | +| `generate-docs-index.mjs` | Build `src/app/docs/lib/docs-auto-generated.ts` from `docs/*.md` | +| `i18n/generate-multilang.mjs` | Translate UI strings + docs via Google Translate | +| `i18n_autotranslate.py` | LLM-based doc translation pipeline | +| `validate_translation.py` | Per-locale translation validation | +| `check_translations.py` | Code-side i18n key check | +| `run-playwright-tests.mjs` | Playwright E2E runner | +| `run-protocol-clients-tests.mjs` | MCP/A2A E2E runner | +| `run-ecosystem-tests.mjs` | Ecosystem (provider integration) tests | +| `test-report-summary.mjs` | Generate coverage summary markdown | +| `smoke-electron-packaged.mjs` | Smoke-test packaged Electron build | +| `native-binary-compat.mjs` | Validate native deps (`better-sqlite3`) match Electron's Node | +| `validate-pack-artifact.ts` | Validate npm pack output | +| `responses-ws-proxy.mjs` | WebSocket bridge for Codex Responses API | +| `v1-ws-bridge.mjs` | WebSocket bridge for `/api/v1/ws` endpoint | +| `standalone-server-ws.mjs` | Standalone WS server runner | +| `system-info.mjs` | Print system/runtime info for support | +| `healthcheck.mjs` | One-shot health check (used by Docker HEALTHCHECK) | +| `uninstall.mjs` | Clean uninstall script | + +--- + +## `docs/` โ€” Public Documentation (44 files + 4 subdirs) + +### Top-level guides + +| Doc | Purpose | +| --------------------------- | ------------------------------------------------------------------------------------- | +| `ARCHITECTURE.md` | High-level architecture, subsystem map, dashboard surface | +| `CODEBASE_DOCUMENTATION.md` | Engineering reference: directories, modules, conventions | +| `FEATURES.md` | Feature matrix with v3.8 highlights | +| `USER_GUIDE.md` | End-user manual (setup, models, combos, CLIs, audio, etc.) | +| `API_REFERENCE.md` | API endpoint reference with auth model | +| `openapi.yaml` | OpenAPI 3.0 spec (121 paths) | +| `SETUP_GUIDE.md` | Install methods (npm, npx, Docker, Electron, Termux, source) | +| `ENVIRONMENT.md` | All env vars (~219 used in code, ~810 lines `.env.example`) | +| `TROUBLESHOOTING.md` | Common errors + v3.8.0 known issues | +| `RELEASE_CHECKLIST.md` | Full release flow (skills, husky, conventional commits, deploy) | +| `COVERAGE_PLAN.md` | Coverage goals and current state | +| `FREE_TIERS.md` | Curated free-tier providers (48+ free + 11 OAuth) | +| `CLI-TOOLS.md` | External CLI integrations + Internal OmniRoute CLI | +| `I18N.md` | i18n architecture, adding a language, 30 locales | +| `UNINSTALL.md` | Clean uninstall steps | +| `PROVIDER_REFERENCE.md` | **Auto-generated** catalog of 177 providers (regen: `npm run gen:provider-reference`) | + +### Subsystem deep-dives + +| Doc | Purpose | +| -------------------------- | ------------------------------------------------------------------- | +| `MCP-SERVER.md` | MCP server: 37 tools, 3 transports, ~13 scopes, REST endpoints | +| `A2A-SERVER.md` | A2A v0.3: JSON-RPC, 5 skills, REST helpers, agent card | +| `AGENT_PROTOCOLS_GUIDE.md` | Unified guide: A2A vs ACP vs Cloud Agents | +| `CLOUD_AGENT.md` | Codex Cloud / Devin / Jules orchestration | +| `SKILLS.md` | Skills framework (built-in + marketplace + SkillsSH + sandbox) | +| `MEMORY.md` | Memory system (SQLite FTS5 + Qdrant) | +| `EVALS.md` | Eval framework (suites, runs, rubrics) | +| `GUARDRAILS.md` | PII masker, prompt injection, vision bridge | +| `COMPLIANCE.md` | Audit log, retention, noLog opt-out | +| `WEBHOOKS.md` | HMAC-signed webhook delivery | +| `REASONING_REPLAY.md` | Hybrid memory/SQLite cache for `reasoning_content` | +| `AUTHZ_GUIDE.md` | Authorization pipeline (`classify` โ†’ `policies` โ†’ `enforce`) | +| `RESILIENCE_GUIDE.md` | Circuit breaker + cooldown + model lockout | +| `STEALTH_GUIDE.md` | TLS fingerprinting (JA3/JA4), Claude Code CCH, MITM cert | +| `AUTO-COMBO.md` | Auto Combo engine (9-factor scoring, 4 mode packs, virtual factory) | + +### Compression + +| Doc | Purpose | +| ------------------------------- | ---------------------------------------- | +| `COMPRESSION_GUIDE.md` | Overview of compression modes + roadmap | +| `COMPRESSION_ENGINES.md` | Caveman + RTK engines, registry contract | +| `COMPRESSION_RULES_FORMAT.md` | Caveman rule pack JSON schema | +| `COMPRESSION_LANGUAGE_PACKS.md` | Per-language rule pack inventory | +| `RTK_COMPRESSION.md` | RTK declarative pipeline (49 filters) | + +### Deployment + +| Doc | Purpose | +| ---------------------------- | ----------------------------------------------------------------- | +| `DOCKER_GUIDE.md` | Docker build, profiles (base/cli/host/cliproxyapi), Redis sidecar | +| `VM_DEPLOYMENT_GUIDE.md` | Generic VM/VPS deployment (Ubuntu/Debian + nginx + systemd) | +| `FLY_IO_DEPLOYMENT_GUIDE.md` | Fly.io deployment (currently Chinese-only) | +| `TERMUX_GUIDE.md` | Android headless via Termux | +| `PWA_GUIDE.md` | Progressive Web App install + service worker | +| `ELECTRON_GUIDE.md` | Desktop app build + sign + distribute | +| `TUNNELS_GUIDE.md` | Cloudflared + ngrok + Tailscale Funnel | +| `PROXY_GUIDE.md` | 4-level outbound proxy + 1proxy marketplace | + +### Subdirectories + +| Subdir | Purpose | +| ------------------------- | ------------------------------------------------------------------------------------- | +| `docs/archive/` | Archived/historical docs (e.g., `RFC-AUTO-ASSESSMENT-DRAFT.md` โ€” superseded by EVALS) | +| `docs/i18n/` | Localized doc translations (~40 locales) | +| `docs/screenshots/` | Image assets for guides | +| `docs/superpowers/plans/` | Implementation plans (generated by `superpowers:writing-plans` skill) | + +--- + +## `tests/` โ€” Test Suites + +| Subdir | Type | Runner | +| ---------------------- | --------------------------------------- | --------------------------------------- | +| `tests/unit/` | Unit tests (~500 files, fastest) | Node native test runner | +| `tests/integration/` | Multi-module + DB integration tests | Node native test runner (concurrency 1) | +| `tests/e2e/` | UI + workflow E2E | Playwright | +| `tests/protocols-e2e/` | MCP + A2A real-client E2E | Custom protocol clients | +| `tests/ecosystem/` | Provider integration (network-touching) | Node native test runner | + +--- + +## `public/` โ€” Static Assets + +| Path | Purpose | +| ------------------- | ---------------------------------------------------------------- | +| `public/` (root) | Favicons, robots.txt, manifest, service worker, marketing images | +| `public/providers/` | Provider logo PNG/SVG (used in dashboard) | + +--- + +## `config/` โ€” Static Configs + +Shipped configuration templates and sample files (referenced by setup wizard). + +--- + +## `.github/` โ€” GitHub Integration + +| Path | Purpose | +| ---------------------------------- | -------------------------------------------------------------- | +| `.github/workflows/` | GitHub Actions CI/CD workflows (lint, test, coverage, release) | +| `.github/ISSUE_TEMPLATE/` | Bug/feature issue templates | +| `.github/PULL_REQUEST_TEMPLATE.md` | PR template | +| `.github/dependabot.yml` | Dependency update config | + +--- + +## `.husky/` โ€” Git Hooks + +| File | Purpose | +| ------------ | ----------------------------------------------------------------- | +| `pre-commit` | Runs `lint-staged + check-docs-sync + check:any-budget:t11` | +| `pre-push` | Currently disabled (commented). Run `npm run test:unit` manually. | +| `_/` | Husky internals | + +--- + +## `.claude/` โ€” Claude Code Slash Commands + +| File | Purpose | +| ----------------------------------------------------------------- | -------------------------------------------------- | +| `commands/version-bump-cc.md` | `/version-bump-cc` โ€” bump version + auto-changelog | +| `commands/generate-release-cc.md` | `/generate-release-cc` โ€” full release workflow | +| `commands/deploy-vps-{local,akamai,both}-cc.md` | Deploy to VPS | +| `commands/capture-release-evidences-cc.md` | Browser-record new features as WebP | +| `commands/review-{prs,discussions}-cc.md` | Triage GitHub PRs/discussions | +| `commands/{issue-triage,resolve-issues,implement-features}-cc.md` | Issue workflows | +| `settings.local.json` | Per-project Claude Code settings | + +--- + +## `.agents/` โ€” Generic Agent Workflows (Codex / Cursor / etc.) + +| Path | Purpose | +| ------------------------ | ------------------------------------------------------- | +| `workflows/*-ag.md` | 11 workflow definitions (mirror of `.claude/commands/`) | +| `skills//SKILL.md` | 9 skill definitions with Codex Execution Notes | + +> **Note:** Workflows and commands are currently identical byte-by-byte. If `.agents/` is meant to target a different agent runtime (Codex), the variants need to diverge meaningfully. + +--- + +## `_ideia/`, `_mono_repo/`, `_references/`, `_tasks/` โ€” Out-of-tree + +These underscore-prefixed directories hold non-shipping content: + +- **`_ideia/`** โ€” design notes (defer / notfit / viable categories) +- **`_mono_repo/`** โ€” historic subprojects (omnirouteCloud, omnirouteSite, vscode-extension) +- **`_references/`** โ€” read-only clones of related OSS projects (LiteLLM, 9router, ClawRouter, CLIProxyAPI, modelrelay, new-api, etc.) for cross-reference during development +- **`_tasks/`** โ€” per-release task tracking files (informal) + +Not included in `npm pack` output. See `.npmignore`. + +--- + +## Generated / Gitignored + +| Path | Purpose | +| ---------------------- | ----------------------------- | +| `node_modules/` | npm dependencies | +| `.next/` | Next.js build output | +| `coverage/` | c8 coverage reports | +| `logs/` | Runtime logs | +| `package/` | npm pack staging | +| `.playwright-mcp/` | Playwright MCP test artifacts | +| `.issues/` | Local issue cache | +| `tsconfig.tsbuildinfo` | TS incremental cache | + +--- + +## Navigation tips + +- **New contributor?** Read `CONTRIBUTING.md` โ†’ `CLAUDE.md` โ†’ `docs/ARCHITECTURE.md` โ†’ `docs/CODEBASE_DOCUMENTATION.md`. +- **Adding a provider?** Follow `docs/ARCHITECTURE.md ยง Adding a New Provider` + cross-check `docs/PROVIDER_REFERENCE.md`. +- **Adding a route?** `docs/ARCHITECTURE.md ยง Adding a New API Route` + `src/shared/validation/schemas.ts`. +- **Adding an MCP tool?** `docs/MCP-SERVER.md ยง Adding a Tool`. +- **Adding an A2A skill?** `docs/A2A-SERVER.md ยง Adding a New Skill`. +- **Running locally?** `docs/SETUP_GUIDE.md`. +- **Deploying?** `docs/DOCKER_GUIDE.md` / `docs/VM_DEPLOYMENT_GUIDE.md` / `docs/FLY_IO_DEPLOYMENT_GUIDE.md`. +- **Releasing?** `docs/RELEASE_CHECKLIST.md` (and `/generate-release-cc` Claude Code skill).