diff --git a/AGENTS.md b/AGENTS.md index 22343a8663..328feeb2a5 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -3,12 +3,12 @@ ## Project Unified AI proxy/router — route any LLM through one endpoint. Multi-provider support -with **265 provider entries** (OpenAI, Anthropic, Gemini, DeepSeek, Groq, xAI, Mistral, Fireworks, +with **268 provider entries** (OpenAI, Anthropic, Gemini, DeepSeek, Groq, xAI, Mistral, Fireworks, Cohere, NVIDIA, Cerebras, Pollinations, Puter, Cloudflare AI, HuggingFace, DeepInfra, SambaNova, Meta Llama API, Moonshot AI, AI21 Labs, Databricks, Snowflake, and many more) with **MCP Server** (94 tools), **A2A v0.3 Protocol**, and **Electron desktop app**. -> **Live counts (v3.8.49)**: providers 265 · MCP tools 94 · MCP scopes 30 · A2A skills 6 · +> **Live counts (v3.8.49)**: providers 268 · MCP tools 104 · MCP scopes 30 · A2A skills 6 · > open-sse services 134 · routing strategies 17 · auto-combo scoring factors 12 · > DB modules 95 · DB migrations 110 · base tables 17 · search providers 11 · > i18n locales 42. **Refresh with `npm run check:docs-all`.** diff --git a/CLAUDE.md b/CLAUDE.md index edf6d63bc7..f8baadb4b6 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -35,7 +35,7 @@ For full test matrix, see `CONTRIBUTING.md` → "Running Tests". For deep archit ## Project at a Glance -**OmniRoute** — unified AI proxy/router. One endpoint, 265 LLM providers, auto-fallback. +**OmniRoute** — unified AI proxy/router. One endpoint, 268 LLM providers, auto-fallback. | Layer | Location | Purpose | | ------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | diff --git a/README.md b/README.md index d1c4d7e834..557487be3f 100644 --- a/README.md +++ b/README.md @@ -6,18 +6,23 @@ # 🚀 OmniRoute — The Free AI Gateway -### Never stop coding. Connect every AI tool to **265 providers** — **90+ free** — through one endpoint. +OmniRoute — Never stop coding. Every AI tool → 268 providers — 90+ free — through one endpoint. Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini with auto-fallback. RTK + Caveman stacked compression saves 15–95% tokens (~89% avg) — never hit limits. 268 AI providers · 90+ free tiers · ~1.6B free tokens/mo · 18 routing strategies · $0 to start. -**Plug Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini. Auto-fallback.** -
+ -**RTK + Caveman compression saves 15–95% tokens. Never hit limits.** +
-
+# 💰 ~1.6B Free Tokens / Month -**~1.6B documented free tokens/month** — up to **~2.1B in your first month** with signup credits — aggregated across the free tiers, plus a long tail of permanently-free, no-cap providers, and the compression above stretches every one further. ([how we count →](docs/reference/FREE_TIERS.md#tldr--how-much-free-inference-does-omniroute-actually-aggregate)) +
-
+> Stacking free tiers by hand is painful — dozens of SDKs, dozens of rate limits, and no idea how much you actually have. OmniRoute aggregates the **documented** free tiers of **40+ provider pools / 500+ models** into one honest number and shows it live on the dashboard (`/dashboard/free-tiers`). + +OmniRoute free-tier budget card: ~1.6B free tokens per month steady, up to ~2.1B in the first month with signup credits, from the documented free tiers of 40+ provider pools / 500+ models behind one endpoint. Honest pool-deduped math — each shared pool counted once (counting every rate limit 24/7 would read ~10B; not published), 15 providers ToS-flagged so you decide. Budget bar of the 21 countable free pools with per-model grid (Mistral Large 3 1B, GPT-4o mini 150M, LongCat 150M, Gemini 2.5 Flash 60M … Auto 25K), ~616M one-time first-month signup credits (vertex 300M, agentrouter 200M, predibase 25M, together 25M, glm-cn 20M, doubao 15M, ai21 10M, deepseek 5M, hyperbolic 5M), plus permanently-free no-token-cap providers (SiliconFlow, Z.AI GLM-Flash, Kilo, OpenCode Zen, baidu …) and a $10 OpenRouter top-up unlocking +24M/mo — surfaced separately so they never inflate the headline. Live used/remaining on /dashboard/free-tiers. + +> Animated summary of the live `/dashboard/free-tiers` page. Full methodology (pool dedupe, credit tiers, provider terms): **[docs/reference/FREE_TIERS.md](docs/reference/FREE_TIERS.md)**. + +

@@ -29,15 +34,6 @@ diegosouzapw%2FOmniRoute | Trendshift [![Star History Rank](https://api.star-history.com/badge?repo=diegosouzapw/OmniRoute&theme=dark)](https://www.star-history.com/diegosouzapw/omniroute) -
- -[![251 AI Providers](https://img.shields.io/badge/251-AI_Providers-6C5CE7?style=for-the-badge)](#-251-ai-providers--90-free) -[![90+ Free](https://img.shields.io/badge/90%2B-Free_Tiers-00B894?style=for-the-badge)](#-251-ai-providers--90-free) -[![1.6B Free Tokens/mo](https://img.shields.io/badge/1.6B-Free_Tokens%2Fmo-00B894?style=for-the-badge)](docs/reference/FREE_TIERS.md) -[![Token Savings](https://img.shields.io/badge/up_to_95%25-Token_Savings-E17055?style=for-the-badge)](#%EF%B8%8F-save-1595-tokens--automatically) -[![18 Strategies](https://img.shields.io/badge/18-Routing_Strategies-0984E3?style=for-the-badge)](#-combos--the-flagship) -[![$0 to start](https://img.shields.io/badge/%240-To_Start-FDCB6E?style=for-the-badge&logoColor=black)](#-quick-start) -
### 💬 Join the community @@ -61,14 +57,14 @@ ![Docker Pulls](https://img.shields.io/docker/pulls/diegosouzapw/omniroute?label=docker%20pulls&logo=docker&color=2496ED) ![Electron Downloads](https://img.shields.io/github/downloads/diegosouzapw/omniroute/total?style=flat&label=electron%20downloads&logo=electron&color=47848F) -[**🚀 Quick Start**](#-quick-start) • [**🎯 Combos**](#-combos--the-flagship) • [**🌐 Providers**](#-251-ai-providers--90-free) • [**🔌 CLI & MCP**](#-full-cli--a2a--mcp) • [**🗜️ Compression**](#%EF%B8%8F-save-1595-tokens--automatically) • [**🌍 Website**](https://omniroute.online) +[**🚀 Quick Start**](#-quick-start) • [**🎯 Combos**](#-combos--the-flagship) • [**🌐 Providers**](#-268-ai-providers--90-free) • [**🔌 CLI & MCP**](#-full-cli--a2a--mcp) • [**🗜️ Compression**](#%EF%B8%8F-save-1595-tokens--automatically) • [**🌍 Website**](https://omniroute.online) -[💥 The Promise](#-the-promise) • [🤔 Why](#-why-omniroute) • [🏆 What Sets Apart](#-what-sets-omniroute-apart) • [🤖 Compatible CLIs](#-compatible-clis--coding-agents) • [🖥️ Where It Runs](#%EF%B8%8F-where-omniroute-runs--anywhere) • [🔒 Private](#-private--local-first) • [🎬 In Action](#-omniroute-in-action) • [📚 Explore More](#-explore-more) • [📧 Support](#-support--community) +[💥 The Promise](#-the-promise) • [🤔 Why](#-why-omniroute) • [🏆 What Sets Apart](#-what-sets-omniroute-apart) • [🤖 Compatible CLIs](#-compatible-clis--coding-agents) • [🖥️ Where It Runs](#%EF%B8%8F-where-omniroute-runs--anywhere) • [🔒 Private](#-private--local-first) • [🎬 In Action](#-omniroute-in-action) • [📸 Screenshots](#-dashboard-screenshots) • [📧 Support](#-support--community)

- 🌐 In 42+ languages + 🌐 In 43 languages @@ -122,47 +118,13 @@
🇺🇸
-
- -
- -# 💰 ~1.6B Free Tokens / Month - -
- -> Stacking free tiers by hand is painful — dozens of SDKs, dozens of rate limits, and no idea how much you actually have. OmniRoute aggregates the **documented** free tiers of **40+ provider pools / 500+ models** into one honest number and shows it live on the dashboard (`/dashboard/free-tiers`). - -- **~1.6B free tokens / month** (steady) — and **up to ~2.1B in your first month** with signup credits. -- **Pool-deduped, honest** — we count each shared free pool **once**, so the headline isn't inflated by rate-limit ceilings the way multi-billion competitor claims are. (Counting every rate limit 24/7 would read ~10B; we don't publish that.) -- **Plus the un-countable** — permanently-free, no-token-cap providers (SiliconFlow, Z.AI GLM-Flash, Kilo, OpenCode Zen…) and a **$10 OpenRouter top-up** that unlocks **+24M/mo**, both surfaced separately so they never inflate the headline. -- **Per-model breakdown**, **live used / remaining** for the current month, and a transparent **terms flag** per provider. - -OmniRoute free-tier budget card: ~1.6B free tokens per month steady, up to ~2.1B in the first month with signup credits, from the documented free tiers of 40+ provider pools / 500+ models behind one endpoint. Honest pool-deduped math — each shared pool counted once (counting every rate limit 24/7 would read ~10B; not published), 15 providers ToS-flagged so you decide. Budget bar of the 21 countable free pools with per-model grid (Mistral Large 3 1B, GPT-4o mini 150M, LongCat 150M, Gemini 2.5 Flash 60M … Auto 25K), ~616M one-time first-month signup credits (vertex 300M, agentrouter 200M, predibase 25M, together 25M, glm-cn 20M, doubao 15M, ai21 10M, deepseek 5M, hyperbolic 5M), plus permanently-free no-token-cap providers (SiliconFlow, Z.AI GLM-Flash, Kilo, OpenCode Zen, baidu …) and a $10 OpenRouter top-up unlocking +24M/mo — surfaced separately so they never inflate the headline. Live used/remaining on /dashboard/free-tiers. - -> Animated summary of the live `/dashboard/free-tiers` page. Full methodology (pool dedupe, credit tiers, provider terms): **[docs/reference/FREE_TIERS.md](docs/reference/FREE_TIERS.md)**. - -
-
# 💥 The Promise
-> One endpoint. **265 providers.** Never stop building — and let OmniRoute pick the cheapest one that works. - - - - - - - - - - - - -
🚫 Never hit limits
Auto-fallback across 265 providers in milliseconds. Quota out? Next provider takes over — zero downtime.
💸 Save up to 95% tokens
RTK + Caveman stacked compression cuts 15–95% of eligible tokens (~89% avg on tool-heavy sessions).
🆓 $0 to start
90+ providers with a free tier, 11 free forever (Kiro, Qoder, Pollinations, LongCat…). No card needed.
🔌 Every tool works
24+ coding agents — Claude Code, Codex, Cursor, Cline, Copilot, Antigravity — through one config.
🧩 One endpoint
OpenAI ↔ Claude ↔ Gemini ↔ Responses API translation. Point any tool at /v1 and it just works.
🛡️ Production-grade
Circuit breakers, TLS stealth, MCP (94 tools), A2A, memory, guardrails, evals. 21,000+ tests.
+The Promise — One endpoint. 268 providers. Never stop building — OmniRoute picks the cheapest one that works. Six pillars: Never hit limits (auto-fallback across 268 providers in milliseconds, zero downtime) · Save up to 95% tokens (RTK + Caveman stacked compression cuts 15–95%, ~89% avg on tool-heavy sessions) · $0 to start (90+ free tiers, 40+ free forever — no card needed) · Every tool works (26 coding agents through one config) · One endpoint (OpenAI ↔ Claude ↔ Gemini ↔ Responses API at /v1) · Production-grade (circuit breakers, TLS stealth, MCP 104 tools, A2A, memory, guardrails, evals — 25,000+ tests).

@@ -173,16 +135,7 @@ -> Stop juggling 10 dashboards, dead API keys, and surprise bills. - -| ❌ The daily pain | ✅ How OmniRoute fixes it | -| ------------------------------------------------------ | ----------------------------------------------------------------------------- | -| 📉 Subscription quota expires unused every month | **Maximize subscriptions** — track quota, use every token before reset | -| 🛑 Rate limits stop you mid-coding | **4-tier auto-fallback** — Subscription → API → Cheap → Free, in milliseconds | -| 🔥 Tool outputs (`git diff`, `grep`, logs) burn tokens | **RTK + Caveman compression** — save 15–95% eligible tokens per request | -| 💸 Expensive APIs ($20–50/mo per provider) | **Cost-optimized routing** — auto-route to the cheapest viable model | -| 🧰 Each AI tool wants its own setup | **One endpoint, every tool, one dashboard** | -| 🌍 AI blocked in your country | **3-level proxy** + TLS fingerprint stealth — use AI from anywhere | +Why OmniRoute — stop juggling 10 dashboards, dead API keys and surprise bills. Ten daily pains vs fixes: quota expiring unused → maximize subscriptions; rate limits mid-coding → 4-tier auto-fallback (Subscription → API → Cheap → Free); tool outputs burning tokens → RTK + Caveman compression (15–95%); expensive APIs → cost-optimized routing; every tool its own setup → one endpoint, one dashboard; AI blocked → 3-level proxy + TLS stealth; dead keys → 3-layer resilience (circuit breakers, key cooldown, model lockout); team sharing one subscription → key pools with fair-share quotas; prompts through someone's cloud → local-first with AES-256-GCM encrypted keys; no spend visibility → live analytics (usage, quota, savings, p95 latency).
@@ -204,14 +157,14 @@ No combo to create. Set your model to `auto` (or a variant) and OmniRoute builds a virtual combo from your connected providers, scored live: -| Model ID | What it optimizes for
| -| -------------- | ---------------------------------------------------------------------------------------------- | -| `auto` | 🎯 Balanced default (LKGP — sticks to your last good provider) | -| `auto/coding` | 🧑‍💻 Quality-first weights for code generation | -| `auto/fast` | ⚡ Lowest latency first | -| `auto/cheap` | 💰 Cheapest per token first | -| `auto/offline` | 🔋 Most quota / rate-limit headroom first | -| `auto/smart` | 🔭 Quality-first + 10% exploration to discover better models | +| Model ID | What it optimizes for                                                                                                                                                                  | +| -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `auto` | 🎯 Balanced default (LKGP — sticks to your last good provider) | +| `auto/coding` | 🧑‍💻 Quality-first weights for code generation | +| `auto/fast` | ⚡ Lowest latency first | +| `auto/cheap` | 💰 Cheapest per token first | +| `auto/offline` | 🔋 Most quota / rate-limit headroom first | +| `auto/smart` | 🔭 Quality-first + 10% exploration to discover better models | ## @@ -219,26 +172,28 @@ No combo to create. Set your model to `auto` (or a variant) and OmniRoute builds All **18** strategies — mix & match per combo step: -| # | Strategy | What it does
| -| --- | ------------------- | ------------------------------------------------------------------------------------- | -| 1 | `priority` | First-target ordered list — drain each before the next 🥇 | -| 2 | `fill-first` | Fill each target's quota fully before moving on | -| 3 | `weighted` | Weighted random by per-target weight | -| 4 | `round-robin` | Cycle through targets in order | -| 5 | `p2c` | Power-of-two-choices random load balancing | -| 6 | `least-used` | Pick the target with the lowest current load | -| 7 | `random` | Uniform random pick (deduplicated) | -| 8 | `strict-random` | Random without de-duplicating repeats 🎲 | -| 9 | `cost-optimized` | Minimize $ per request from live catalog pricing 💸 | -| 10 | `headroom` | Pick the target with the most remaining quota | -| 11 | `reset-window` | Prefer the target whose quota window resets soonest | -| 12 | `reset-aware` | Rank by quota reset time — short windows first 📊 | -| 13 | `context-relay` | Hand off context across targets for long conversations 🧠 | -| 14 | `context-optimized` | Pick the best fit for the current context size | -| 15 | `lkgp` | Last-Known-Good Path — sticky to the last successful target | -| 16 | `auto` | 12-factor live scoring across every connection 🤖 | -| 17 | `fusion` | Fan out to a panel of models + a judge synthesizes one answer 🧬 | -| 18 | `pipeline` | Chain steps — each target's output feeds the next one 🔗 | +| # | Strategy | What it does                                                                                                                                                             | +| --- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| 1 | `priority` | First-target ordered list — drain each before the next 🥇 | +| 2 | `fill-first` | Fill each target's quota fully before moving on | +| 3 | `weighted` | Weighted random by per-target weight | +| 4 | `round-robin` | Cycle through targets in order | +| 5 | `p2c` | Power-of-two-choices random load balancing | +| 6 | `least-used` | Pick the target with the lowest current load | +| 7 | `random` | Uniform random pick (deduplicated) | +| 8 | `strict-random` | Random without de-duplicating repeats 🎲 | +| 9 | `cost-optimized` | Minimize $ per request from live catalog pricing 💸 | +| 10 | `headroom` | Pick the target with the most remaining quota | +| 11 | `reset-window` | Prefer the target whose quota window resets soonest | +| 12 | `reset-aware` | Rank by quota reset time — short windows first 📊 | +| 13 | `context-relay` | Hand off context across targets for long conversations 🧠 | +| 14 | `context-optimized` | Pick the best fit for the current context size | +| 15 | `lkgp` | Last-Known-Good Path — sticky to the last successful target | +| 16 | `auto` | 12-factor live scoring across every connection 🤖 | +| 17 | `fusion` | Fan out to a panel of models + a judge synthesizes one answer 🧬 | +| 18 | `pipeline` | Chain steps — each target's output feeds the next one 🔗 | + +All 18 combo routing strategies animated, one tile per strategy showing the flow it executes: priority (drain the 1st, then the next), fill-first (fill a target's quota, then move on), weighted (weighted random), round-robin (cycle in order), p2c (pick 2, take the lighter), least-used (lowest load wins), random (uniform, deduped), strict-random (repeats allowed), cost-optimized (cheapest $ per request), headroom (most remaining quota), reset-window (resets soonest → use it), reset-aware (rank by reset, short first), context-relay (hand off long context), context-optimized (fit the context size), lkgp (sticky to last success), auto (live 12-factor scoring), fusion (panel + judge → one answer), pipeline (each output feeds the next). The Auto-Combo engine scores every candidate on **12 factors** (health, quota, cost, latency, success rate, freshness…) — see [`docs/routing/AUTO-COMBO.md`](docs/routing/AUTO-COMBO.md). @@ -248,12 +203,12 @@ All **18** strategies — mix & match per combo step: > Running several keys against the **same upstream account** (one Codex Pro plan, one Kimi key, one GLM Coding seat)? A burst on one key can burn the whole 5-hour / hourly quota and lock everyone else out. **Quota-Share** distributes a provider's time-based quota **fairly** across the keys in a pool — and it's _work-conserving_, so an idle member's slice is lent out instead of wasted. -| Knob | What it controls
| -| ------------------------ | ----------------------------------------------------------------------------------------- | -| ⚖️ **Allocation weight** | each key's slice of the pool — e.g. `50 / 30 / 20` | -| 📐 **Dimensions** | track `%` · requests · tokens · `$`, per **5h / 7d / per-model** window | -| 🚦 **Policy** | `hard` (block over share) · `soft` (deprioritize) · `burst` (use idle headroom) | -| 🧱 **Cap** | absolute ceiling per key, independent of mode | +| Knob | What it controls                                                                                                                                                              | +| ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| ⚖️ **Allocation weight** | each key's slice of the pool — e.g. `50 / 30 / 20` | +| 📐 **Dimensions** | track `%` · requests · tokens · `$`, per **5h / 7d / per-model** window | +| 🚦 **Policy** | `hard` (block over share) · `soft` (deprioritize) · `burst` (use idle headroom) | +| 🧱 **Cap** | absolute ceiling per key, independent of mode | OmniRoute key pool 'team-codex': one Codex Pro account shared by 3 keys over a 5-hour window. alice weight 50 (up to 50% of the shared 5h quota), bob weight 30, ci-bot weight 20. In generous mode (under 50% pool used) idle shares are lent out; once the pool crosses 50% strict mode holds each key to its fair share. @@ -263,13 +218,7 @@ All **18** strategies — mix & match per combo step: ### 🧱 Resilience is built in (3 independent layers) -| Layer | Scope | What it does | -| -------------------------- | ----------------- | -------------------------------------------------------------------------- | -| 🔌 **Circuit breaker** | whole provider | Stops hammering a provider that's failing upstream; auto-probes to recover | -| 💤 **Connection cooldown** | one account / key | Skips a rate-limited key while other keys keep serving | -| 🎯 **Model lockout** | provider + model | Quarantines just one quota-limited model, not the whole connection | - -OmniRoute combo 'always-on' with priority strategy: requests go to 1. cc/claude-opus-4-7 (subscription, use it fully); on failure fall through to 2. cx/gpt-5.5 (second subscription), then 3. glm/glm-5.1 (cheap backup at $0.5/1M), then 4. kr/claude-sonnet-4.5 (free, unlimited, never fails). Result: 4 layers of fallback = zero downtime. +OmniRoute resilience — 3 independent self-healing layers, the right layer for the right failure. Layer 1 provider circuit breaker (whole provider): trips only on 408/5xx, thresholds OAuth 3× / API-key 5× / local 2×, resets 60s/30s/15s into a HALF-OPEN probe, lazy recovery; while OPEN the combo reroutes to the next provider. Layer 2 connection cooldown (one key/account): base 5s OAuth / 3s API-key, exponential ×2 backoff with anti-thundering-herd guard, 429 honors Retry-After, success clears all error state; one cooling key is skipped while sibling keys keep serving. Layer 3 model lockout (one model): per-model 429, local 404 or mode denials lock just that model — never the whole connection. Terminal states (banned, expired, credits exhausted) are for the operator, not cooldowns. 📖 [Auto-Combo Engine](docs/routing/AUTO-COMBO.md) · [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) @@ -283,18 +232,18 @@ All **18** strategies — mix & match per combo step: | Feature | OmniRoute | Other routers | | -------------------------------------- | ------------------------------------------------------------------- | ------------- | -| 🌐 Providers | **251** | 20–100 | -| 🆓 Free providers | **90+ (11 free forever)** | 1–5 | +| 🌐 Providers | **268** | 20–100 | +| 🆓 Free providers | **90+ (40+ free forever)** | 1–5 | | 🔀 Routing strategies | **18** (priority, weighted, cost-optimized, context-relay, fusion…) | 1–3 | | 🗜️ Token compression | **RTK + Caveman stacked (15–95%)** | None / 20–40% | -| 🧰 Built-in MCP server | **94 tools, 3 transports, 30 scopes** | Rare | +| 🧰 Built-in MCP server | **104 tools, 3 transports, 30 scopes** | Rare | | 🤝 A2A agent protocol | **6 skills, JSON-RPC 2.0** | None | | 🧠 Memory (FTS5 + vector) | **Yes** | Rare | | 🛡️ Guardrails (PII, injection, vision) | **Yes** | Rare | | ☁️ Cloud agents | **Codex, Cursor, Devin, Jules** | None | | 🥷 TLS fingerprint stealth | **JA3/JA4 via wreq-js** | None | | 🖥️ Multi-platform | **Web · Desktop · Termux · PWA** | Web only | -| 🌍 i18n | **42 locales** | 0–4 | +| 🌍 i18n | **43 locales** | 0–4 | 📊 Detailed comparison vs LiteLLM, OpenRouter & Portkey → [`docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md`](docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md) @@ -306,23 +255,23 @@ All **18** strategies — mix & match per combo step:
-> Recent highlights from **v3.8.20 → v3.8.47**. Full history in [`CHANGELOG.md`](CHANGELOG.md). +> Recent highlights from **v3.8.20 → v3.8.49**. Full history in [`CHANGELOG.md`](CHANGELOG.md). -- **🗜️ Compression hardening** — a default-on **inflation guard** (discard the stacked result and send the verbatim original whenever compression would _grow_ the prompt), completed **Caveman rule packs** for German / French / Japanese (dedup + ultra) plus a new **Chinese (文言 / wényán) input pack** with zh-vs-ja auto-detection, and **RTK filters for Gradle & .NET (`dotnet`)** build output. → [Compression](docs/compression/COMPRESSION_ENGINES.md) -- **💸 Honest flat-rate cost** — subscription / coding-plan providers (ChatGPT Web, grok-web, the Minimax / Kimi / GLM / Alibaba Coding plans, Xiaomi MiMo…) now read **$0** in cost analytics instead of an inflated per-token estimate, while budget / quota / routing keep estimating unchanged. → [API Reference](docs/reference/API_REFERENCE.md) -- **⚖️ Quota-Share routing** — a dedicated combo strategy that spreads load across accounts by _available quota_: Deficit-Round-Robin scheduling, per-connection `max_concurrent` with cooldown-wait queueing, multi-window usage buckets (5h / 7d / per-model), per-(key, model) caps, session stickiness for prompt-cache integrity (now with a per-combo / global disable toggle), and proactive saturation from upstream token-usage headers. → [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) -- **🤖 One-command CLI/agent setup** — a dedicated `setup-*` command configures each coding tool to route through OmniRoute (Claude Code, Codex, Cline, Continue, Cursor, Roo Code, Kilo Code, Crush, Goose, Qwen Code, Aider, OpenCode); `omniroute launch` / `omniroute launch-codex` are zero-config launchers. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md) -- **🛰️ Remote mode** — drive a remote OmniRoute from any machine with scoped access tokens (`omniroute connect` / `omniroute contexts` / `omniroute tokens`), plus an `omniroute login antigravity` helper that runs Google "native/desktop" OAuth on your own machine and pastes a credential blob into a remote/VPS install (where the loopback redirect is unreachable). → [Remote Mode](docs/guides/REMOTE-MODE.md) -- **🧭 Smarter auto-routing** — OpenRouter-style `auto/:` combos (e.g. `auto/coding:fast`, `auto/reasoning:pro`), a **Fusion** strategy (fan out to a panel of models in parallel, then synthesize via a judge), **task-aware routing** (best-fit connection per task type), per-request `X-Route-Model` override, live Arena-ELO + models.dev model intelligence, per-step account allowlists, provider-wildcard combo steps, nested combo-ref execution, sticky weighted selection, `web_search`-aware routing (now with **per-model web-search/web-fetch interception rules**), native **xAI Grok `/v1/responses`** routing, and **per-request Auto-Combo controls** (`X-OmniRoute-Mode` mode-preset override + `X-OmniRoute-Budget` hard USD cost ceiling, scoped to a single request). Embeddings-only and rerank-only models (JinaAI, OpenRouter custom, reranker models…) no longer disappear from the combo builder's model picker. → [Auto-Combo](docs/routing/AUTO-COMBO.md) -- **🗜️ Pluggable compression** — an async pipeline of **10 composable engines** with Compression Studios, an LLMLingua-2 ONNX engine and a heuristic/SLM two-tier **Ultra**, RTK, delegated Anthropic Context Editing, **Output Styles** (output-axis steering: terse-prose / less-code / terse-CJK), an **adaptive context-budget dial** (escalate only as far as needed to fit the context window), per-request `x-omniroute-compression` control, an opt-in offline eval harness, one-click **Headroom** proxy lifecycle management from the dashboard (Docker sidecar supported), a synthetic **compression playground** (Play lanes + A/B Compare with USD-capped fidelity verdicts), an opt-in **per-step fidelity gate** that rejects a lossy engine before it degrades the prompt, a **best-of-N candidate encoder** (GCF vs TOON — keep whichever is shorter, with an A/B bytes/token table in the studio), the vendored **GCF codec updated to spec v3.2** (nested flattening — deeply-nested payloads go from ~3% to ~32% compression vs JSON), a new **omniglyph** engine (context-as-image, ~10× fewer tokens on the converted block), **CCR ranged/grep/stats retrieval** (pull an exact byte/line slice or summary of a stored block instead of re-expanding it), a unified panel with named profiles + an active-profile selector, an opt-in **per-engine pipeline circuit-breaker**, an opt-in **LLM-tier engine** (a model pass for higher-ratio semantic compression), a **read-lifecycle engine** that collapses superseded file reads, **usage-observed prefix freeze**, a graduated **CCR retrieval-feedback ramp**, a `preserveSystemPrompt` mode enum, and a **drag-reorder pipeline editor** in the studio. → [Compression](docs/compression/COMPRESSION_ENGINES.md) -- **🕵️ Transparent MITM decrypt (TPROXY)** — capture & translate traffic from CLIs that ignore proxy env vars, with a per-SNI certificate authority and a trust-store installer. → [MITM/TPROXY](docs/security/MITM-TPROXY-DECRYPT.md) -- **💸 Cost telemetry everywhere** — `X-OmniRoute-*` cost/usage headers on every endpoint (including media), a non-token cost engine, a cache-HIT `X-OmniRoute-Cost-Saved` header, and per-key USD spend quotas. → [API Reference](docs/reference/API_REFERENCE.md) -- **🧠 Memory you control** — opt-in int8 vector quantization (Qdrant + sqlite-vec), opt-in **typed memory decay** (aged low-value memories fade on a per-type schedule), memory off by default, and a per-request `x-omniroute-no-memory` header. → [Memory](docs/frameworks/MEMORY.md) -- **🛡️ Security** — a prompt-injection guard across every LLM route (backed by a red-team suite), plus a free DuckDuckGo last-resort web search. → [Guardrails](docs/security/GUARDRAILS.md) -- **🖼️ New endpoints** — `/v1/ocr` (Mistral OCR) and `/v1/audio/translations` (Whisper-style audio translation) round out the media API surface. → [API Reference](docs/reference/API_REFERENCE.md) -- **🌍 Deployment & ops** — reverse-proxy `basePath` deployment (`OMNIROUTE_BASE_PATH`, e.g. serving OmniRoute under `/omniroute/`), browser-language auto-detect on first visit, per-API-key device/connection tracking (IP+UA fingerprint, masked, in-memory only), root-less MITM cert trust for user-namespaced containers (`OMNIROUTE_NO_SUDO`), server-side configured-only / available-only filters on the Free Provider Rankings page, and **Traditional Chinese (zh-TW)** localization for the frontend + CLI. → [Environment](docs/reference/ENVIRONMENT.md) -- **🤝 More providers & agents** — Cursor Cloud Agent (a 4th cloud agent), CodeBuddy CN (`copilot.tencent.com`), a Google Flow video-generation provider, new gateways **DGrid** and **Pioneer AI** (Fastino Labs), inbound **xAI Grok** translators plus **Grok Build (xAI)** with an OAuth import-token flow, GPT-4 / GPT-4o-mini on the GitHub Copilot provider, multi-model **Factory Droid**, **ZenMux Free** (session-cookie free tier), **Alibaba DashScope** text-to-video (`wan2.7-t2v`), a refreshed 250-provider catalog (OrcaRouter, Wafer AI, OpenAdapter, dit.ai, TokenRouter, …), Vertex AI media generation (speech/transcription/music/video), a first-class **Ollama** local-provider card, the **SenseNova** free Token Plan (chat + text-to-image), one-click account import from CLIProxyAPI (`~/.cli-proxy-api/`), **Claude Sonnet 5** wired end-to-end, a new provider wave (**Kenari**, **SumoPod**, **X5Lab**, **Charm Hyper**, **Nube.sh**, **b.ai**, **Qiniu**, **ModelScope**, **Augment/Auggie CLI**, **ClinePass**, NVIDIA NIM image generation), Codex account import from a raw ChatGPT access token, the **Requesty** gateway (BYOK, ~200 free req/day), **Yuanbao (web)** as a cookie-session provider (DeepSeek V3/R1 + Hunyuan), the **Zed** hosted LLM aggregator (OAuth), **Claude 5 Sonnet** on the Claude Web provider, Kiro **adaptive-thinking reasoning** surfaced as `reasoning_content`, **bulk API-key add for Cloudflare Workers AI**, and **OpenVecta** (AI inference gateway). → [Providers](docs/reference/PROVIDER_REFERENCE.md) -- **⚡ Local performance & infra** — a one-click local Redis launcher (`omniroute redis up`, plus a dashboard Redis panel), one-click **Cloudflare Workers** and **Deno Deploy** relay deployers wired into the proxy pool, a relay-backend selector (`OMNIROUTE_RELAY_BACKEND=ts|bifrost|auto`) so `/v1/relay` stays the stable surface while choosing the fastest backend internally, **Bifrost** (Go AI-gateway) and **Mux** (agent-orchestration daemon) promoted to first-class embedded/supervised services alongside 9Router/CLIProxyAPI, **Webshare** added as a paid fourth source in the free-proxy provider framework, and **shorthand proxy formats + protocol header mode** for bulk proxy import. → [Embedded Services](docs/frameworks/EMBEDDED-SERVICES.md) +- **🗜️ Compression hardening** — default-on inflation guard, Caveman packs for DE / FR / JA + Chinese (wényán), RTK filters for Gradle & .NET. → [Compression](docs/compression/COMPRESSION_ENGINES.md) +- **💸 Honest flat-rate cost** — subscription / coding-plan providers read **$0** in cost analytics; budget, quota & routing keep estimating. → [API Reference](docs/reference/API_REFERENCE.md) +- **⚖️ Quota-Share routing** — split load across accounts by _available quota_: DRR scheduling, per-connection concurrency, multi-window buckets, session stickiness. → [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) +- **🤖 One-command CLI/agent setup** — `setup-*` configures 12+ coding tools; `omniroute launch` / `launch-codex` are zero-config. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md) +- **🛰️ Remote mode** — drive a remote OmniRoute with scoped tokens (`connect` / `contexts` / `tokens`) + an `antigravity` OAuth helper for VPS installs. → [Remote Mode](docs/guides/REMOTE-MODE.md) +- **🧭 Smarter auto-routing** — `auto/:` combos, **Fusion** (model panel + judge), task-aware routing, per-request model / mode / USD-budget overrides. → [Auto-Combo](docs/routing/AUTO-COMBO.md) +- **🗜️ Pluggable compression** — 10 composable engines + Compression Studios: LLMLingua-2, two-tier Ultra, omniglyph, per-step fidelity gate, GCF v3.2, drag-reorder editor. → [Compression](docs/compression/COMPRESSION_ENGINES.md) +- **🕵️ Transparent MITM decrypt (TPROXY)** — capture CLIs that ignore proxy env vars, with a per-SNI CA + trust-store installer. → [MITM/TPROXY](docs/security/MITM-TPROXY-DECRYPT.md) +- **💸 Cost telemetry everywhere** — `X-OmniRoute-*` cost/usage headers on every endpoint, cache-HIT savings header, per-key USD spend quotas. → [API Reference](docs/reference/API_REFERENCE.md) +- **🧠 Memory you control** — off by default, opt-in int8 vector quantization + typed decay, per-request `x-omniroute-no-memory`. → [Memory](docs/frameworks/MEMORY.md) +- **🛡️ Security** — prompt-injection guard on every LLM route (red-team suite) + free DuckDuckGo last-resort web search. → [Guardrails](docs/security/GUARDRAILS.md) +- **🖼️ New endpoints** — `/v1/ocr` (Mistral OCR) and `/v1/audio/translations` (Whisper-style) round out the media surface. → [API Reference](docs/reference/API_REFERENCE.md) +- **🌍 Deployment & ops** — reverse-proxy `basePath`, browser-language auto-detect, per-key device tracking, root-less MITM trust, zh-TW localization. → [Environment](docs/reference/ENVIRONMENT.md) +- **🤝 More providers & agents** — Cursor Cloud Agent, Grok Build (xAI), Ollama first-class card, Claude Sonnet 5, Zed, Requesty, SenseNova, Yuanbao… and a refreshed 250-provider catalog. → [Providers](docs/reference/PROVIDER_REFERENCE.md) +- **⚡ Local performance & infra** — one-click local Redis, Cloudflare Workers / Deno Deploy relay deployers, Bifrost & Mux as supervised embedded services. → [Embedded Services](docs/frameworks/EMBEDDED-SERVICES.md)
@@ -335,28 +284,44 @@ All **18** strategies — mix & match per combo step:
- - - - - + + + + + + + + + - - - - - - + + + + + + + + + + + + + + + + + +
Claude Code
Claude Code
Codex CLI
Codex CLI
Cursor
Cursor
Copilot
Copilot
Continue
Continue
Claude Code
Claude Code
                           
Codex CLI
Codex CLI
                           
Cline
Cline
                           
Kilo Code
Kilo Code
                           
Roo CodeRoo Code
Roo Code
                           
Continue
Continue
                           
Qwen Code
Qwen Code
                           
Aider
Aider
                           
ForgeCode
ForgeCode
                           
OpenCode
OpenCode
Kilo Code
Kilo Code
Droid
Droid
OpenClaw
OpenClaw
Kiro
Kiro
Command Code
Command
jcode
jcode
                           
DeepSeek TUI
DeepSeek TUI
                           
CodeWhale
CodeWhale
                           
OpenCode
OpenCode
                           
Factory Droid
Factory Droid
                           
GitHub Copilot CLI
Copilot CLI
                           
Cursor CLI
Cursor CLI
                           
Smelt
Smelt
                           
Pi (pi-coding-agent)
Pi
                           
Grok Build (xAI)
Grok Build
                           
Hermes Agent (Nous Research)
Hermes Agent
                           
OpenClaw
OpenClaw
                           
Goose
Goose
                           
Open Interpreter
Open Interpreter
                           
Warp AI
Warp AI
                           
Agent Deck
Agent Deck
                           
-+ also works with · Cline · Antigravity · Windsurf · AMP · Hermes · Qwen CLI · Roo · Continue · any OpenAI-compatible tool ++ also works with · Kiro · Command Code · Antigravity · Windsurf · AMP · any OpenAI-compatible tool
-📖 Per-tool setup for all 24+ tools → [`docs/reference/CLI-TOOLS.md`](docs/reference/CLI-TOOLS.md) · 🧩 OpenCode plugin → [`@omniroute/opencode-provider`](https://www.npmjs.com/package/@omniroute/opencode-provider) +📖 Per-tool setup for all 26 tools (20 CLI Code's + 6 CLI Agents) → [`docs/reference/CLI-TOOLS.md`](docs/reference/CLI-TOOLS.md) · 🧩 OpenCode plugin → [`@omniroute/opencode-provider`](https://www.npmjs.com/package/@omniroute/opencode-provider) @@ -364,11 +329,11 @@ All **18** strategies — mix & match per combo step:
-# 🌐 251 AI Providers — 90+ Free +# 🌐 268 AI Providers — 90+ Free
-> The most complete catalog of any open-source router: **265 providers**, **90+ with a free tier**, **11 free forever**. +> The most complete catalog of any open-source router: **268 providers**, **90+ with a free tier**, **40+ free forever**.
@@ -376,28 +341,26 @@ All **18** strategies — mix & match per combo step: - - - - - - + + + + + + + + + - - - - - - - - - - - - - - + + + + + + + + +
OpenAI
OpenAI
Anthropic
Anthropic
Gemini
Gemini
xAI Grok
xAI Grok
DeepSeek
DeepSeek
Mistral
Mistral
OpenAI
OpenAI
                           
Anthropic
Anthropic
                           
Gemini
Gemini
                           
xAI Grok
xAI Grok
                           
DeepSeek
DeepSeek
                           
Mistral
Mistral
                           
Qwen
Qwen
                           
Meta Llama
Meta Llama
                           
Groq
Groq
                           
Qwen
Qwen
Meta Llama
Meta Llama
Groq
Groq
NVIDIA
NVIDIA
MiniMax
MiniMax
Cohere
Cohere
Perplexity
Perplexity
Hugging Face
HuggingFace
Together
Together
Fireworks
Fireworks
Cloudflare
Cloudflare
Baidu
Baidu
NVIDIA
NVIDIA
                           
MiniMax
MiniMax
                           
Cohere
Cohere
                           
Perplexity
Perplexity
                           
Hugging Face
HuggingFace
                           
Together
Together
                           
Fireworks
Fireworks
                           
Cloudflare
Cloudflare
                           
Baidu
Baidu
                           
@@ -409,15 +372,13 @@ All **18** strategies — mix & match per combo step: - - - - - - - - - + + + + + + +
AgentRouter
AgentRouter
GPT-5, Claude, Gemini
$100 free credits
Qoder AI
Qoder AI
Kimi-K2, DeepSeek-R1
Unlimited FREE
Pollinations
Pollinations
GPT-5, Claude, Llama 4
No key needed
LongCat
LongCat
LongCat-2.0
10M tokens one-time (KYC) 🔑
Cloudflare AI
Cloudflare AI
50+ models
10K neurons/day
NVIDIA NIM
NVIDIA NIM
129 models
~40 RPM free
Cerebras
Cerebras
Qwen3 235B
1M tokens/day
AgentRouter
AgentRouter
GPT-5, Claude, Gemini
$100 free credits

                                     
Qoder AI
Qoder AI
Kimi-K2, DeepSeek-R1
Unlimited FREE

                                     
Pollinations
Pollinations
GPT-5, Claude, Llama 4
No key needed

                                     
LongCat
LongCat
LongCat-2.0
10M tokens one-time (KYC) 🔑

                                     
Cloudflare AI
Cloudflare AI
50+ models
10K neurons/day

                                     
NVIDIA NIM
NVIDIA NIM
129 models
~40 RPM free

                                     
Cerebras
Cerebras
Qwen3 235B
1M tokens/day

                                     
@@ -455,13 +416,7 @@ All **18** strategies — mix & match per combo step:
-> Your keys, your machine, your data. OmniRoute is a **local proxy** — it never phones home. - -- 🏠 **Runs 100% on your hardware** — npm, Docker, desktop, or your phone. No OmniRoute cloud sits in the request path. -- 🔐 **Credentials encrypted at rest** — API keys & OAuth tokens sealed with **AES-256-GCM**. -- 🚫 **Zero telemetry by default** — your prompts go only to the providers _you_ choose, nowhere else. -- 🛡️ **Hardened gateway** — API-key scoping, IP filtering, rate limits, prompt-injection guard, loopback-only process routes. -- 📜 **MIT licensed & fully open-source** — audit every line, self-host forever. +Private and local-first — your keys, your machine, your data; OmniRoute is a local proxy that never phones home. Eleven guarantees: runs 100% on your hardware (0 cloud hops), zero telemetry by default, credentials encrypted at rest (AES-256-GCM), no account or sign-up, hardened gateway (API-key scoping, IP filtering, rate limits, prompt-injection guard), loopback-only process routes, upstream header scrubbing, strictly opt-in PII redaction, sanitized errors that never leak internals, a local audit trail in your own SQLite, and MIT-licensed fully open-source code. 📖 [Authorization](docs/architecture/AUTHZ_GUIDE.md) · [Guardrails](docs/security/GUARDRAILS.md) · [Compliance](docs/security/COMPLIANCE.md) @@ -510,12 +465,12 @@ Tokens are scoped `read` / `write` / `admin`; process-spawning routes stay loopb Expose OmniRoute over **MCP** or **A2A** and any capable agent gets the keys to the whole gateway — routing, providers, combos, cache, compression, memory — autonomously. -| Protocol | Endpoint | Use it for | -| ------------------ | ----------------------------------------------- | ------------------------------------------------------ | -| 🧰 **MCP (stdio)** | `omniroute --mcp` | Plug into Claude Desktop, Cursor, any MCP client | -| 🌊 **MCP (HTTP)** | `http://localhost:20128/api/mcp/stream` | Remote MCP — **94 tools**, 30 scopes, full audit trail | -| 📡 **MCP (SSE)** | `http://localhost:20128/api/mcp/sse` | Streaming MCP transport | -| 🤝 **A2A** | `http://localhost:20128/.well-known/agent.json` | Agent-to-agent, **JSON-RPC 2.0** + SSE, 6 skills | +| Protocol | Endpoint | Use it for | +| ------------------ | ----------------------------------------------- | ------------------------------------------------------- | +| 🧰 **MCP (stdio)** | `omniroute --mcp` | Plug into Claude Desktop, Cursor, any MCP client | +| 🌊 **MCP (HTTP)** | `http://localhost:20128/api/mcp/stream` | Remote MCP — **104 tools**, 30 scopes, full audit trail | +| 📡 **MCP (SSE)** | `http://localhost:20128/api/mcp/sse` | Streaming MCP transport | +| 🤝 **A2A** | `http://localhost:20128/.well-known/agent.json` | Agent-to-agent, **JSON-RPC 2.0** + SSE, 6 skills | ```bash # Give Claude Code the full OmniRoute toolset over MCP: @@ -553,14 +508,14 @@ Engines run in pipeline order; each is independently toggleable and configurable Code blocks, URLs and structured data are **always preserved** byte-perfect. **One-click presets** combine the engines: -| Mode | Savings | Best for
| -| ------------------------------ | ---------- | --------------------------------------------------------------------------------- | -| 🪶 **Lite** | ~15% | Always-on safe default | -| 🪨 **Standard (Caveman)** | ~30% | Daily coding | -| ⚡ **Aggressive** | ~50% | Long tool-heavy sessions | -| 🔥 **Ultra** | ~75% | Maximum savings | -| 🧰 **RTK** | 60–90% | Shell/test/build/git output | -| 🔗 **Stacked (RTK → Caveman)** | **78–95%** | Mixed prompts + tool logs | +| Mode | Savings | Best for                                                                                                                                        | +| ------------------------------ | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | +| 🪶 **Lite** | ~15% | Always-on safe default | +| 🪨 **Standard (Caveman)** | ~30% | Daily coding | +| ⚡ **Aggressive** | ~50% | Long tool-heavy sessions | +| 🔥 **Ultra** | ~75% | Maximum savings | +| 🧰 **RTK** | 60–90% | Shell/test/build/git output | +| 🔗 **Stacked (RTK → Caveman)** | **78–95%** | Mixed prompts + tool logs | **Real example — Standard mode:** @@ -772,137 +727,10 @@ same process on one port, so there is no separate CLI-only package today.
-# 📚 Explore More +# 📸 Dashboard Screenshots
-
-💰 Pricing at a glance & the $0 Free Stack (11 providers) - -
- -| Tier | Example
| Cost | -| --------------------------- | -------------------------------------------------------------------------------- | ---------- | -| 💳 **Subscription** | Claude Code Pro / Codex / Copilot | $10–200/mo | -| 🔑 **API Key (free tiers)** | NVIDIA NIM, Cerebras, Groq | **FREE** | -| 💰 **Cheap** | GLM-5 $0.5/1M · MiniMax M2.5 $0.3/1M | pennies | -| 🆓 **Free Forever** | Kiro, Qoder, Qwen, Pollinations, LongCat | **$0** | - -**The $0 Free Stack — combine into one unbreakable combo:** - -| Provider | Prefix | Free models
| Quota | -| ----------------- | ----------- | ------------------------------------------------------------------------------------ | ------------------ | -| **Kiro** | `kr/` | Claude Sonnet 4.5, Haiku 4.5, Opus 4.6 | 50 credits/mo | -| **Qoder** | `if/` | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 | ♾️ Unlimited | -| **Qwen** | `qw/` | qwen3-coder-plus/flash/next | ♾️ Unlimited | -| **Pollinations** | `pol/` | GPT-5, Claude, Gemini, DeepSeek, Llama 4 | No key needed | -| **LongCat** | `lc/` | LongCat-2.0 | 10M one-time (KYC) | -| **Cloudflare AI** | `cf/` | 50+ models | 10K neurons/day | -| **NVIDIA NIM** | `nvidia/` | 129 models | ~40 RPM | -| **Cerebras** | `cerebras/` | Qwen3 235B, GPT-OSS 120B | 1M tok/day | - -> 💡 The dashboard "cost" is a **savings tracker**, not a bill — OmniRoute never charges you. A "$290 total cost" using free models means **$290 saved**. - -📖 Complete free directory → [`docs/reference/FREE_TIERS.md`](docs/reference/FREE_TIERS.md) — 25+ providers, quotas, base URLs. - -
- -
-🎯 Use Cases — ready-made combo playbooks - -
- -**$0 forever:** - -``` -1. kr/claude-sonnet-4.5 (Kiro — ~50 credits/mo per acct) -2. if/kimi-k2-thinking (Qoder — unlimited) -3. pol/gpt-5 (Pollinations — no key) -4. lc/LongCat-2.0 (10M one-time backup, KYC) -Compression: aggressive (~50%) → double your free quota · Cost: $0/mo -``` - -**24/7 no interruptions:** chain 2 subscriptions → cheap → free for 5 layers of fallback. -**Blocked region:** free providers + global/per-provider proxy → access AI from any country. -**Max savings:** subscription + cheap backup + `ultra` compression (~75%) → ~$150–300/mo saved for heavy users. - -
- -
-🌍 Bypass geo-blocks — 3-level proxy + stealth - -
- -🇷🇺 🇨🇳 🇮🇷 🇨🇺 🇹🇷 In a blocked region? OmniRoute's **3-level proxy** (Global / Per-Provider / Per-Connection) proxies API requests, OAuth flows, connection tests, token refresh & model sync. - -- **Protocols:** HTTP/HTTPS, SOCKS5, authenticated proxies -- **🆓 1proxy marketplace** — hundreds of free validated proxies, quality scores, auto-rotation -- **Anti-detection** — TLS fingerprint spoofing (`wreq-js`), CLI fingerprint matching, proxy IP preservation - -📖 [`docs/ops/PROXY_GUIDE.md`](docs/ops/PROXY_GUIDE.md) - -
- -
-✨ Full feature list — 30+ capabilities (memory, evals, observability) - -
- -**Routing:** 18 strategies · task-aware smart routing · thinking budget controls · wildcard routing · system prompt injection. -**Compatibility:** OpenAI ↔ Claude ↔ Gemini ↔ Responses API · auto OAuth refresh (PKCE, 8 providers) · multi-account round-robin · Batch + Files API · live OpenAPI 3.0. -**Protocols:** MCP (94 tools, 3 transports, 30 scopes) · A2A (JSON-RPC 2.0, SSE, 6 skills) · ACP · cloud agents (Codex, Cursor, Devin, Jules). -**Plugins:** custom plugin marketplace (system-configured registry URL with SSRF-guarded fetch) · install / enable / disable · Notion + Obsidian knowledge-base integrations (WebDAV file server, vault search, note CRUD). -**Embedded services:** one-click install & lifecycle management of local sidecar services (CLIProxy, NineRouter). -**Quality & Ops:** built-in **Evals** (golden-set: exact/contains/regex/custom) · guardrails (PII, injection, vision) · health dashboard · p50/p95/p99 telemetry · webhooks · compliance audit. -**AI Agent Skills:** drop-in markdown manifests — point any agent at a `skills/*/SKILL.md` manifest. 43 skills available. - -📖 [MCP Server](open-sse/mcp-server/README.md) · [A2A Server](src/lib/a2a/README.md) · [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) · [Features Gallery](docs/guides/FEATURES.md) - -
- -
-📖 Setup, env vars & FAQ - -
- -| Env var | Default | Purpose
| -| ----------------- | -------------- | -------------------------------------------------------------------------------- | -| `PORT` | `20128` | API + dashboard port | -| `REQUIRE_API_KEY` | `false` | Require API key for all requests | -| `DATA_DIR` | `~/.omniroute` | Database & config storage | - -**Will I be charged by OmniRoute?** No — it's free, open-source software on your machine. You only pay paid providers directly. OmniRoute has no billing system. -**Are FREE providers really unlimited?** Mostly — Qoder, Pollinations, LongCat, and Cloudflare are free with no per-account credit cap. Kiro is free too but capped at ~50 credits/month per account. Stack multiple free providers in a combo and auto-fallback keeps you serving for $0. -**Will compression hurt quality?** No — it only compresses the **input**; code, URLs, JSON are always protected. -**Does it work where AI is blocked?** Yes — 3-level proxy + 1proxy marketplace reach all 265 providers. - -📖 [User Guide](docs/guides/USER_GUIDE.md) · [API Reference](docs/reference/API_REFERENCE.md) · [Environment Config](docs/reference/ENVIRONMENT.md) - -
- -
-🐛 Troubleshooting - -
- -| Problem | Quick fix
| -| ----------------------------------------- | ---------------------------------------------------------------------------------- | -| "Language model did not provide messages" | Provider quota exhausted → use a combo fallback | -| Rate limiting (429) | Add fallback: `cc/claude → glm/glm-4.7 → if/kimi-k2-thinking` | -| OAuth token expired | Auto-refreshed; if stuck, delete + re-auth in Providers | -| `unsupported_country_region_territory` | Configure proxy in Settings → Proxy | -| Docker SQLite locks | Use `--stop-timeout 40` for clean WAL checkpoint | -| Node runtime errors | Use Node `>=22.22.2 <23` or `>=24.0.0 <27` | - -🐛 **Reporting a bug?** Run `npm run system-info` and attach `system-info.txt`. 📖 [`docs/guides/TROUBLESHOOTING.md`](docs/guides/TROUBLESHOOTING.md) - -
- -
-📸 Dashboard screenshots - -
- | Page | Screenshot | Page | Screenshot | | ---------- | ------------------------------------------------- | ---------- | --------------------------------------------- | | Providers | ![Providers](docs/screenshots/01-providers.png) | Combos | ![Combos](docs/screenshots/02-combos.png) | @@ -910,8 +738,6 @@ Compression: aggressive (~50%) → double your free quota · Cost: $0/mo | Translator | ![Translator](docs/screenshots/05-translator.png) | Settings | ![Settings](docs/screenshots/06-settings.png) | | CLI Tools | ![CLI Tools](docs/screenshots/07-cli-tools.png) | Usage Logs | ![Usage](docs/screenshots/08-usage.png) | -
-
@@ -944,7 +770,7 @@ Compression: aggressive (~50%) → double your free quota · Cost: $0/mo - **Protocols**: MCP (stdio/HTTP) + A2A v0.3 (JSON-RPC 2.0 + SSE) - **Streaming**: Server-Sent Events (SSE) + WebSocket bridge (`/v1/ws`) - **Auth**: OAuth 2.0 (PKCE) + JWT + API Keys + MCP Scoped Authorization -- **Testing**: Node.js test runner + Vitest (**21,000+ test cases** across 2,586 files — unit, integration, E2E, security, ecosystem) +- **Testing**: Node.js test runner + Vitest (**25,000+ test cases** across 3,300+ files — unit, integration, E2E, security, ecosystem) - **Platforms**: Desktop (Electron), Android (Termux), PWA (any browser) - **CI/CD**: GitHub Actions (auto npm publish + Docker Hub on release) - **Website**: [omniroute.online](https://omniroute.online) @@ -962,66 +788,66 @@ Compression: aggressive (~50%) → double your free quota · Cost: $0/mo ### 📘 Getting Started -| Document | Description
| -| -------------------------------------------------------------- | ------------------------------------------------------------------------------------ | -| [User Guide](docs/guides/USER_GUIDE.md) | Providers, combos, CLI integration, deployment | -| [Setup Guide](docs/guides/SETUP_GUIDE.md) | Full install methods, CLI tool configs, protocol setup, timeout tuning | -| [CLI Tools Guide](docs/reference/CLI-TOOLS.md) | Per-tool setup for Claude Code, Codex, Cursor, Cline, OpenClaw, Kilo, Copilot | -| [Remote Mode](docs/guides/REMOTE-MODE.md) | Drive a remote OmniRoute (VPS) from your laptop CLI via scoped access tokens | -| [Claude Code Config](docs/guides/CLAUDE-CODE-CONFIGURATION.md) | Point Claude Code at OmniRoute (local/remote) with `launch` + per-model profiles | -| [Quick Start](README.md#-quick-start) | 3-step install → connect → configure | +| Document | Description                                                                                                                                                                          | +| -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| [User Guide](docs/guides/USER_GUIDE.md) | Providers, combos, CLI integration, deployment | +| [Setup Guide](docs/guides/SETUP_GUIDE.md) | Full install methods, CLI tool configs, protocol setup, timeout tuning | +| [CLI Tools Guide](docs/reference/CLI-TOOLS.md) | Per-tool setup for Claude Code, Codex, Cursor, Cline, OpenClaw, Kilo, Copilot | +| [Remote Mode](docs/guides/REMOTE-MODE.md) | Drive a remote OmniRoute (VPS) from your laptop CLI via scoped access tokens | +| [Claude Code Config](docs/guides/CLAUDE-CODE-CONFIGURATION.md) | Point Claude Code at OmniRoute (local/remote) with `launch` + per-model profiles | +| [Quick Start](README.md#-quick-start) | 3-step install → connect → configure | ### 🔧 Operations & Deployment -| Document | Description
| -| -------------------------------------------------------- | ------------------------------------------------------------------------------------ | -| [Docker Guide](docs/guides/DOCKER_GUIDE.md) | Docker run, Compose profiles, Caddy HTTPS, tunnels, image tags | -| [Podman Guide](contrib/podman/README.md) | Quadlet systemd integration, podman-compose, SELinux | -| [VM Deployment](docs/ops/VM_DEPLOYMENT_GUIDE.md) | Complete guide: VM + nginx + Cloudflare setup | -| [Fly.io Deployment](docs/ops/FLY_IO_DEPLOYMENT_GUIDE.md) | Deploy to Fly.io with persistent storage | -| [Termux Guide](docs/guides/TERMUX_GUIDE.md) | Run OmniRoute on Android via Termux | -| [PWA Guide](docs/guides/PWA_GUIDE.md) | Progressive Web App install, caching, architecture | -| [Uninstall Guide](docs/guides/UNINSTALL.md) | Clean removal for all install methods | -| [Environment Config](docs/reference/ENVIRONMENT.md) | Complete `.env` variables and references | +| Document | Description                                                                                                                                                                         | +| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| [Docker Guide](docs/guides/DOCKER_GUIDE.md) | Docker run, Compose profiles, Caddy HTTPS, tunnels, image tags | +| [Podman Guide](contrib/podman/README.md) | Quadlet systemd integration, podman-compose, SELinux | +| [VM Deployment](docs/ops/VM_DEPLOYMENT_GUIDE.md) | Complete guide: VM + nginx + Cloudflare setup | +| [Fly.io Deployment](docs/ops/FLY_IO_DEPLOYMENT_GUIDE.md) | Deploy to Fly.io with persistent storage | +| [Termux Guide](docs/guides/TERMUX_GUIDE.md) | Run OmniRoute on Android via Termux | +| [PWA Guide](docs/guides/PWA_GUIDE.md) | Progressive Web App install, caching, architecture | +| [Uninstall Guide](docs/guides/UNINSTALL.md) | Clean removal for all install methods | +| [Environment Config](docs/reference/ENVIRONMENT.md) | Complete `.env` variables and references | ### 🧠 Features & Architecture -| Document | Description
| -| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | -| [Architecture](docs/architecture/ARCHITECTURE.md) | System architecture, data flow, and internals | -| [Compression Guide](docs/compression/COMPRESSION_GUIDE.md) | 7-option pipeline: off / lite / standard / aggressive / ultra / RTK / stacked | -| [RTK Compression](docs/compression/RTK_COMPRESSION.md) | Command-output compression, filters, trust, verify, raw-output recovery | -| [Compression Engines](docs/compression/COMPRESSION_ENGINES.md) | Caveman, RTK, stacked pipelines, dashboard/API/MCP surfaces | -| [Compression Rules Format](docs/compression/COMPRESSION_RULES_FORMAT.md) | JSON rule-pack schemas for Caveman and RTK filters | -| [Compression Language Packs](docs/compression/COMPRESSION_LANGUAGE_PACKS.md) | Language detection and Caveman rule-pack authoring | -| [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) | Circuit breakers, cooldowns, queue, anti-thundering herd, TLS spoofing | -| [Auto-Combo Engine](docs/routing/AUTO-COMBO.md) | 12-factor scoring, mode packs, self-healing | -| [Proxy Guide](docs/ops/PROXY_GUIDE.md) | 3-level proxy system, 1proxy marketplace, registry CRUD | -| [Free Tiers](docs/reference/FREE_TIERS.md) | 25+ free API providers consolidated directory | -| [Features Gallery](docs/guides/FEATURES.md) | Visual dashboard tour with screenshots | -| [Codebase Documentation](docs/architecture/CODEBASE_DOCUMENTATION.md) | Beginner-friendly codebase walkthrough | +| Document | Description                                                                                                                                                        | +| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| [Architecture](docs/architecture/ARCHITECTURE.md) | System architecture, data flow, and internals | +| [Compression Guide](docs/compression/COMPRESSION_GUIDE.md) | 7-option pipeline: off / lite / standard / aggressive / ultra / RTK / stacked | +| [RTK Compression](docs/compression/RTK_COMPRESSION.md) | Command-output compression, filters, trust, verify, raw-output recovery | +| [Compression Engines](docs/compression/COMPRESSION_ENGINES.md) | Caveman, RTK, stacked pipelines, dashboard/API/MCP surfaces | +| [Compression Rules Format](docs/compression/COMPRESSION_RULES_FORMAT.md) | JSON rule-pack schemas for Caveman and RTK filters | +| [Compression Language Packs](docs/compression/COMPRESSION_LANGUAGE_PACKS.md) | Language detection and Caveman rule-pack authoring | +| [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) | Circuit breakers, cooldowns, queue, anti-thundering herd, TLS spoofing | +| [Auto-Combo Engine](docs/routing/AUTO-COMBO.md) | 12-factor scoring, mode packs, self-healing | +| [Proxy Guide](docs/ops/PROXY_GUIDE.md) | 3-level proxy system, 1proxy marketplace, registry CRUD | +| [Free Tiers](docs/reference/FREE_TIERS.md) | 25+ free API providers consolidated directory | +| [Features Gallery](docs/guides/FEATURES.md) | Visual dashboard tour with screenshots | +| [Codebase Documentation](docs/architecture/CODEBASE_DOCUMENTATION.md) | Beginner-friendly codebase walkthrough | ### 🤖 Protocols & APIs -| Document | Description
| -| ------------------------------------------------- | ------------------------------------------------------------------------------------ | -| [API Reference](docs/reference/API_REFERENCE.md) | All endpoints with examples | -| [OpenAPI Spec](docs/openapi.yaml) | OpenAPI 3.0 specification | -| [MCP Server](open-sse/mcp-server/README.md) | 95 MCP tools, IDE configs, Python/TS/Go clients | -| [MCP Server Guide](docs/frameworks/MCP-SERVER.md) | MCP installation, transports, and tool reference | -| [A2A Server](src/lib/a2a/README.md) | JSON-RPC 2.0 protocol, skills, streaming, task mgmt | -| [A2A Server Guide](docs/frameworks/A2A-SERVER.md) | A2A agent card, tasks, skills, and streaming | +| Document | Description                                                                                                                                                                             | +| ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| [API Reference](docs/reference/API_REFERENCE.md) | All endpoints with examples | +| [OpenAPI Spec](docs/openapi.yaml) | OpenAPI 3.0 specification | +| [MCP Server](open-sse/mcp-server/README.md) | 104 MCP tools, IDE configs, Python/TS/Go clients | +| [MCP Server Guide](docs/frameworks/MCP-SERVER.md) | MCP installation, transports, and tool reference | +| [A2A Server](src/lib/a2a/README.md) | JSON-RPC 2.0 protocol, skills, streaming, task mgmt | +| [A2A Server Guide](docs/frameworks/A2A-SERVER.md) | A2A agent card, tasks, skills, and streaming | ### 📋 Project & Quality -| Document | Description
| -| -------------------------------------------------- | ------------------------------------------------------------------------------------ | -| [Contributing](CONTRIBUTING.md) | Development setup and guidelines | -| [Changelog](CHANGELOG.md) | Full per-version release history | -| [Security Policy](SECURITY.md) | Vulnerability reporting and security practices | -| [i18n Guide](docs/guides/I18N.md) | 40+ language support, translation workflow, RTL | -| [Release Checklist](docs/ops/RELEASE_CHECKLIST.md) | Pre-release validation steps | -| [Coverage Plan](docs/ops/COVERAGE_PLAN.md) | Test coverage strategy and 21,000+ test suite | +| Document | Description                                                                                                                                                                              | +| -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| [Contributing](CONTRIBUTING.md) | Development setup and guidelines | +| [Changelog](CHANGELOG.md) | Full per-version release history | +| [Security Policy](SECURITY.md) | Vulnerability reporting and security practices | +| [i18n Guide](docs/guides/I18N.md) | 40+ language support, translation workflow, RTL | +| [Release Checklist](docs/ops/RELEASE_CHECKLIST.md) | Pre-release validation steps | +| [Coverage Plan](docs/ops/COVERAGE_PLAN.md) | Test coverage strategy and 25,000+ test suite |
@@ -1251,7 +1077,7 @@ MIT License - see [LICENSE](LICENSE) for details. **[⬆ Back to top](#-omniroute)** · Built with ❤️ for the open-source AI community. -OmniRoute v3.8.43 · Node ≥22.22.2 · MIT License · omniroute.online +OmniRoute v3.8.49 · Node ≥22.22.2 · MIT License · omniroute.online
diff --git a/changelog.d/maintenance/7769-readme-animated-cards-overhaul.md b/changelog.d/maintenance/7769-readme-animated-cards-overhaul.md new file mode 100644 index 0000000000..56e757fe3b --- /dev/null +++ b/changelog.d/maintenance/7769-readme-animated-cards-overhaul.md @@ -0,0 +1 @@ +- README: unified animated card system — numbers audited against the v3.8.49 tree (268 providers, 104 MCP tools, 25k+ tests, 26 CLIs, 40+ free-forever, 43 locales, regenerated provider reference); one flat style contract across all cards; 5 new SMIL cards (hero fused with the budget card, "Why" 10-row pain-vs-fix ledger, 18-strategy flow grid, "Private & Local-First" 11-row guarantee ledger, 3-layer resilience card replacing the always-on combo card) plus a rebuilt compact half-height CLI terminal; every animation is pause-at-t0-safe — the first frame is always the finished composition (includes the budget-bar freeze fix on the shipped free-tier card) (#7769) diff --git a/docs/architecture/ARCHITECTURE.md b/docs/architecture/ARCHITECTURE.md index 059b8db9c2..6841590774 100644 --- a/docs/architecture/ARCHITECTURE.md +++ b/docs/architecture/ARCHITECTURE.md @@ -17,7 +17,7 @@ It provides a single OpenAI-compatible endpoint (`/v1/*`) and routes traffic acr Core capabilities: -- OpenAI-compatible API surface for CLI/tools (237 providers, 75 executors) +- OpenAI-compatible API surface for CLI/tools (268 providers, 84 executors) - Request/response translation across provider formats - Model combo fallback (multi-model sequence) - Structured combo steps (`provider + model + connection`) with runtime ordering by `compositeTiers` diff --git a/docs/architecture/CODEBASE_DOCUMENTATION.md b/docs/architecture/CODEBASE_DOCUMENTATION.md index 513b144ee0..fe0e187eb2 100644 --- a/docs/architecture/CODEBASE_DOCUMENTATION.md +++ b/docs/architecture/CODEBASE_DOCUMENTATION.md @@ -452,7 +452,7 @@ open-sse/ ├── types.d.ts ├── config/ Provider registries, header profiles, identity, … ├── handlers/ Request handlers (chat, embeddings, audio, image, …) -├── executors/ 75 provider-specific HTTP executors +├── executors/ 84 provider-specific HTTP executors ├── translator/ Format conversion (OpenAI ↔ Claude ↔ Gemini ↔ Cursor ↔ Kiro) ├── transformer/ Responses API ↔ Chat Completions stream transformer ├── services/ 80+ service modules (combos, fallback, quotas, identity, …) @@ -482,7 +482,7 @@ open-sse/ ### 4.2 `open-sse/executors/` -75 provider executors, each extending `BaseExecutor` (`base.ts`): +84 provider executors, each extending `BaseExecutor` (`base.ts`): `antigravity`, `azure-openai`, `blackbox-web`, `chatgpt-web`, `cliproxyapi`, `cloudflare-ai`, `codex`, `commandCode`, `cursor`, `default`, `devin-cli`, @@ -491,7 +491,7 @@ open-sse/ (shared identity helper) and `index.ts` (registry). > Note: providers not listed here are served by `default.ts` using the generic -> OpenAI-compatible executor. The full provider catalog (237 entries) lives in +> OpenAI-compatible executor. The full provider catalog (268 entries) lives in > `src/shared/constants/providers.ts`. ### 4.3 `open-sse/translator/` diff --git a/docs/diagrams/README.md b/docs/diagrams/README.md index 8f318a5e32..dee34594f5 100644 --- a/docs/diagrams/README.md +++ b/docs/diagrams/README.md @@ -27,14 +27,20 @@ Not every diagram comes from a `.mmd` source. Hand-authored SVGs live at this directory's root and animate with SMIL only (no JS, no external fonts), so they play inside GitHub's `` sandbox: -| File | Used in | Notes | -| ---------------------------------------- | ---------------- | ---------------------------------------------------------------------------- | -| [tier-cascade.svg](./tier-cascade.svg) | README.md (root) | Animated 4-tier auto-fallback cascade (16s loop, 4 acts). Edit the SVG directly — there is no `.mmd` source. | -| [pool-fair-share.svg](./pool-fair-share.svg) | README.md (root) | Animated key-pool fair-share quota (generous → strict, 16s loop). Edit the SVG directly — there is no `.mmd` source. | -| [combo-always-on.svg](./combo-always-on.svg) | README.md (root) | Animated priority-combo fallback (4 layers, 16s loop). Edit the SVG directly — there is no `.mmd` source. | -| [cli-terminal.svg](./cli-terminal.svg) | README.md (root) | Animated terminal cycling 3 CLI commands (providers/combo/health) + subcommand ticker (18s loop). Edit the SVG directly — there is no `.mmd` source. | -| [compression-pipeline.svg](./compression-pipeline.svg) | README.md (root) | Animated 10-engine compression funnel (8s loop). Edit the SVG directly — there is no `.mmd` source. | -| [free-tier-budget.svg](./free-tier-budget.svg) | README.md (root) | Animated free-tier budget card (~1.6B/mo headline, 21-pool budget bar, per-model grid, signup credits, 10s loop). Edit the SVG directly — there is no `.mmd` source. | +| File | Used in | Notes | +| ------------------------------------------------------ | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| [tier-cascade.svg](./tier-cascade.svg) | README.md (root) | Animated 4-tier auto-fallback cascade (16s loop, 4 acts). Edit the SVG directly — there is no `.mmd` source. | +| [pool-fair-share.svg](./pool-fair-share.svg) | README.md (root) | Animated key-pool fair-share quota (generous → strict, 16s loop). Edit the SVG directly — there is no `.mmd` source. | +| [combo-always-on.svg](./combo-always-on.svg) | style reference | Animated priority-combo fallback (4 layers, 16s loop). Edit the SVG directly — there is no `.mmd` source. | +| [cli-terminal.svg](./cli-terminal.svg) | README.md (root) | Compact half-height animated terminal (1200×350): 3 real CLI commands cycling with typewriter + scrolling subcommand ticker; first frame = completed providers screen. Edit the SVG directly — there is no `.mmd` source. | +| [compression-pipeline.svg](./compression-pipeline.svg) | README.md (root) | Animated 10-engine compression funnel (8s loop). Edit the SVG directly — there is no `.mmd` source. | +| [free-tier-budget.svg](./free-tier-budget.svg) | README.md (root) | Animated free-tier budget card (~1.6B/mo headline, 21-pool budget bar, per-model grid, signup credits, 10s loop). Edit the SVG directly — there is no `.mmd` source. | +| [readme-hero.svg](./readme-hero.svg) | README.md (root) | Animated hero card (tagline, 268-provider/90+ free headline, full-width compression bar demo, 6 stat chips). Edit the SVG directly — there is no `.mmd` source. | +| [promise-pillars.svg](./promise-pillars.svg) | README.md (root) | Animated "The Promise" 6-pillar card (12s border-highlight sweep). Edit the SVG directly — there is no `.mmd` source. | +| [why-pain-fix.svg](./why-pain-fix.svg) | README.md (root) | Animated "Why OmniRoute" 10-row pain-vs-fix ledger (15s green row sweep). Edit the SVG directly — there is no `.mmd` source. | +| [strategies-grid.svg](./strategies-grid.svg) | README.md (root) | Animated 6×3 grid of all 18 routing-strategy flows (one micro-stage per strategy, staggered dot loops). Edit the SVG directly — there is no `.mmd` source. | +| [privacy-local.svg](./privacy-local.svg) | README.md (root) | Animated "Private & Local-First" 11-row guarantee ledger with receipt chips (16s green row sweep). Edit the SVG directly — there is no `.mmd` source. | +| [resilience-layers.svg](./resilience-layers.svg) | README.md (root) | Animated 3-layer resilience card (breaker states CLOSED→OPEN→HALF-OPEN, key cooldown with ×2 backoff, model lockout — 18s loops). Edit the SVG directly — there is no `.mmd` source. | ## How to update diff --git a/docs/diagrams/cli-terminal.svg b/docs/diagrams/cli-terminal.svg index 4a853aca10..832ac12fd8 100644 --- a/docs/diagrams/cli-terminal.svg +++ b/docs/diagrams/cli-terminal.svg @@ -1,120 +1,42 @@ - - - - - - - - - - - - - - - - - - - - - - omniroute — 80+ commands - - - - - $ - - omniroute providers list - - - - - - - - OmniRoute Providers - - 1f3a9c2e  anthropic   Claude Max 20x    active - - 8c2d5b1a  codex       Codex Pro (team)  active - - f4e0a97b  glm         GLM Coding Plan   active - - 03bd6e5f  kimi        Kimi K2 free      active - - 9d4e1a08  gemini-cli  Gemini 2.5 Pro    active - - b7c3f26d  copilot     GPT-5 (Copilot)   active - - … 253 more connections - - - - - - - $ - - omniroute combo list - - - - - - - - OmniRoute Combos - -    always-on     [priority      ] enabled - -    cost-saver    [cost-optimized] enabled - -    fusion-panel  [fusion        ] enabled - -    context-relay [context-relay ] enabled - - - - - - - $ - - omniroute health - - - - - - - - OmniRoute Health - -   Status: healthy - -   Uptime: 4d 12h 33m - -   Requests (24h): 18,412 - -   Circuit Breakers - -     ● closed     24 - -     ○ half-open  1 - -     ○ open       0 - - - - - - - - - providers · oauth · keys · combo · nodes · models · cache · compression · cost · usage · quota · health · resilience · telemetry · logs · audit · mcp · a2a · cloud · memory · skills · eval · tunnel · backup · sync · webhooks · policy · pricing · translator · simulate … - providers · oauth · keys · combo · nodes · models · cache · compression · cost · usage · quota · health · resilience · telemetry · logs · audit · mcp · a2a · cloud · memory · skills · eval · tunnel · backup · sync · webhooks · policy · pricing · translator · simulate … - - - + +Compact animated terminal cycling three real OmniRoute CLI commands with a typewriter effect and a scrolling subcommand ticker; the first frame shows the completed providers-list screen. + + + + + +omniroute — 80+ commands +omniroute providers listOmniRoute Providers1f3a9c2e  anthropic   Claude Max 20x    active8c2d5b1a  codex       Codex Pro (team)  activef4e0a97b  glm         GLM Coding Plan   active03bd6e5f  kimi        Kimi K2 free      active… 264 more providers + +$ +omniroute providers list + + + + +OmniRoute Providers1f3a9c2e  anthropic   Claude Max 20x    active8c2d5b1a  codex       Codex Pro (team)  activef4e0a97b  glm         GLM Coding Plan   active03bd6e5f  kimi        Kimi K2 free      active… 264 more providers + + +$ +omniroute combo list + + + + +OmniRoute Combos   always-on     [priority      ] enabled   cost-saver    [cost-optimized] enabled   fusion-panel  [fusion        ] enabled   context-relay [context-relay ] enabled… run: omniroute combo create + + +$ +omniroute health + + + + +OmniRoute Health  Status: healthy   Uptime: 4d 12h 33m  Requests (24h): 18,412   p95: 412ms  Breakers: ● 24 closed  ◒ 1 half-open  ○ 0 open  Providers: 268 registered   90+ free tiers… live: /dashboard · omniroute status + + + + +providers · oauth · keys · combo · nodes · models · cache · compression · cost · usage · quota · health · resilience · telemetry · logs · audit · mcp · a2a · cloud · memory · skills · eval · doctor · repl · tunnel · backup · sync · webhooks · policy · pricing · translator · simulate …providers · oauth · keys · combo · nodes · models · cache · compression · cost · usage · quota · health · resilience · telemetry · logs · audit · mcp · a2a · cloud · memory · skills · eval · doctor · repl · tunnel · backup · sync · webhooks · policy · pricing · translator · simulate … + + \ No newline at end of file diff --git a/docs/diagrams/free-tier-budget.svg b/docs/diagrams/free-tier-budget.svg index 3236d4fb08..89b7d23d1c 100644 --- a/docs/diagrams/free-tier-budget.svg +++ b/docs/diagrams/free-tier-budget.svg @@ -22,7 +22,7 @@ - + @@ -70,7 +70,7 @@ THE HONEST MATH ~10B - + every rate limit · 24/7 we don't publish that @@ -108,8 +108,8 @@ - - + + each segment = one free pool · widths floored so every provider shows · honest numbers below diff --git a/docs/diagrams/privacy-local.svg b/docs/diagrams/privacy-local.svg new file mode 100644 index 0000000000..b571eb7936 --- /dev/null +++ b/docs/diagrams/privacy-local.svg @@ -0,0 +1,25 @@ + + Animated privacy ledger: eleven fully readable rows on the first frame; a soft green highlight sweeps down the rows in a continuous cycle. + + + + + + + + + + + + + + PRIVATE & LOCAL-FIRST + + Your keys, your machine, your data. OmniRoute is a local proxy — it never phones home. + + + + + + Runs 100% on your hardware — npm, Docker, desktop, or your phone — no OmniRoute cloud in the request path0 CLOUD HOPSZero telemetry by default — your prompts go only to the providers you choose, nowhere elseDEFAULTCredentials encrypted at rest — API keys & OAuth tokens sealed on your own diskAES-256-GCMNo account, no sign-up — a local password guards the dashboard — OmniRoute never asks who you areLOCAL AUTHHardened gateway — API-key scoping, IP filtering, rate limits, prompt-injection guardAUTHZ TIERSProcess routes are loopback-only — a token leaked through a tunnel can’t spawn processes127.0.0.1Upstream header scrubbing — deny-listed headers stripped before every provider callDENY-LISTPII redaction & response sanitization — built in, strictly opt-in — payloads are never mutated by defaultOPT-INSanitized errors — responses never leak stack traces, paths or internalsNO LEAKSLocal audit trail — MCP tool calls & admin actions logged in your SQLite, not oursYOUR DBMIT licensed & fully open-source — audit every line, self-host foreverMIT + diff --git a/docs/diagrams/promise-pillars.svg b/docs/diagrams/promise-pillars.svg new file mode 100644 index 0000000000..23d0e59143 --- /dev/null +++ b/docs/diagrams/promise-pillars.svg @@ -0,0 +1,139 @@ + + Animated promise card: six pillar tiles fade in in reading order, then a soft colored border highlight sweeps from tile to tile in a continuous cycle. + + + + + + + + + + + + + + + + + + THE PROMISE + + + + One endpoint. 268 providers. Never stop building — OmniRoute picks the cheapest one that works. + + + + + + + + + + + + + + + + Never hit limits + Auto-fallback across 268 providers in + milliseconds. Quota out? The next provider + takes over — zero downtime. + + + + + + + + + + + + + + + Save up to 95% tokens + RTK + Caveman stacked compression cuts + 15–95% of eligible tokens — ~89% average + on tool-heavy sessions. + + + + + + + + + + + + + + $0 to start + 90+ providers with a free tier, 40+ free + forever — Qoder, Pollinations, Cloudflare, + SiliconFlow… No card needed. + + + + + + + + + + + + + + + Every tool works + 26 coding agents — Claude Code, Codex, + Cursor, Cline, Copilot, Antigravity — + through one config. + + + + + + + + + + + + + + One endpoint + OpenAI ↔ Claude ↔ Gemini ↔ Responses API + translation. Point any tool at /v1 — + it just works. + + + + + + + + + + + + + + Production-grade + Circuit breakers, TLS stealth, MCP (104 + tools), A2A, memory, guardrails, evals — + 25,000+ tests. + + + + + + $ npm i -g omniroute  ·  point your tool at http://localhost:20128/v1  ·  $0 + MIT · OPEN SOURCE + + diff --git a/docs/diagrams/readme-hero.svg b/docs/diagrams/readme-hero.svg new file mode 100644 index 0000000000..c058c8d41c --- /dev/null +++ b/docs/diagrams/readme-hero.svg @@ -0,0 +1,87 @@ + + Animated hero card: a pulse travels the divider line and a compression bar demo repeatedly shrinks a prompt by up to 95 percent; all headline content is static and readable on the first frame. + + + + + + + + + + + + + + + + + + + + + + OMNIROUTE — THE FREE AI GATEWAY + ONE ENDPOINT · /v1 + + + Never stop coding. + + + Every AI tool → 268 providers90+ free — through one endpoint. + + + Claude Code · Codex · Cursor · Cline · Copilot · Antigravity  →  FREE Claude / GPT / Gemini · auto-fallback + + + + + + + + + + + + + + RTK + CAVEMAN · STACKED COMPRESSION + Save 15–95% tokens + + + + + + your prompt + + + −89% avg on tool-heavy sessions + + + never hit limits + auto-fallback keeps you coding + $ npm i -g omniroute + + + + + + 268 + AI PROVIDERS + + 90+ + FREE TIERS + + ~1.6B + FREE TOKENS / MO + + 15–95% + TOKEN SAVINGS + + 18 + ROUTING STRATEGIES + + $0 + TO START + + diff --git a/docs/diagrams/resilience-layers.svg b/docs/diagrams/resilience-layers.svg new file mode 100644 index 0000000000..b4d49b7f31 --- /dev/null +++ b/docs/diagrams/resilience-layers.svg @@ -0,0 +1,22 @@ + + Animated resilience card: three stacked layer panels, each replaying its healing loop — breaker states cycling CLOSED, OPEN, HALF-OPEN; a cooling key with backoff while other keys serve; a locked model while sibling models keep serving. First frame is fully readable. + + + + + + + + + + + + + + RESILIENCE · 3 SELF-HEALING LAYERS + + The right layer for the right failure — never kill more than what actually broke. + PROVIDERCONNECTION / KEYMODEL + LAYER 1 · SCOPE: WHOLE PROVIDERProvider circuit breakerisolate a provider failing upstream —reroute now, auto-probe to recovertrips only on 408 · 500 · 502 · 503 · 504threshold — oauth 3× · api-key 5× · local 2×reset — 60s · 30s · 15s → HALF-OPEN probelazy recovery — reads refresh expired staterouterprovider Afails ×5provider B ← nextCLOSEDOPENHALF-OPENLAYER 2 · SCOPE: ONE KEY / ACCOUNTConnection cooldownskip one rate-limited key while theother keys keep serving the providerbase cooldown — oauth 5s · api-key 3srepeat fails — backoff ×2 (anti-herd guard)429 honors Retry-After / reset headerssuccess → clearAccountError() resets allprovider · 3 keyskey-1429key-2key-3cooling ×2ⁿLAYER 3 · SCOPE: ONE MODELModel lockoutquarantine a single model — never killthe whole connection for one 429scope — provider + connection + modelper-model 429 · local 404 · mode denialslocked model ≠ dead keyother models keep serving instantlykey-1model-amodel-bmodel-c + which failure trips what → 5xx / 408 : breaker · key 429 / 401 : cooldown · one-model 429 / 404 : lockout · banned / expired / credits : terminal (operator) + \ No newline at end of file diff --git a/docs/diagrams/strategies-grid.svg b/docs/diagrams/strategies-grid.svg new file mode 100644 index 0000000000..6ef2ae7f7c --- /dev/null +++ b/docs/diagrams/strategies-grid.svg @@ -0,0 +1,110 @@ + + Animated grid of 18 routing-strategy tiles; static tracks show each flow on the first frame while a small dot repeatedly travels the path each strategy takes. + + + + + + + + + + + + + + COMBO ROUTING · ALL 18 STRATEGY FLOWS + + ● request → ▮ targets — each tile animates the path its strategy takes · green outline = the pick + + +priority + +drain the 1st, then the next + + +fill-first + +fill T1’s quota, then move on + + +weighted +60%30%10% +weighted random pick + + +round-robin + +cycle in order 1→2→3→4 + + +p2c +20%65% +pick 2, take the lighter + + +least-used +75%20%55%40% +lowest current load wins + + +random + +uniform random (deduped) + + +strict-random +×2! +pure random — repeats ok + + +cost-optimized +$9$3$0.5FREE +cheapest $ per request + + +headroom +10%80%40%60% +most remaining quota + + +reset-window +58m2m!31m12h +resets soonest → use it + + +reset-aware +3rd1st2nd4th +rank by reset — short first + + +context-relay + +hand off long context + + +context-optimized +8K200K32K128K +fit the context size + + +lkgp + +sticky to last success + + +auto +72916455 +live 12-factor scoring + + +fusion + +panel + judge → one answer + + +pipeline +123 +each output feeds the next + + diff --git a/docs/diagrams/why-pain-fix.svg b/docs/diagrams/why-pain-fix.svg new file mode 100644 index 0000000000..6b3051098c --- /dev/null +++ b/docs/diagrams/why-pain-fix.svg @@ -0,0 +1,107 @@ + + Animated comparison ledger: ten pain-versus-fix rows are fully readable on the first frame; a soft green highlight sweeps down the rows in a continuous cycle. + + + + + + + + + + + + + + + + + WHY OMNIROUTE + + Stop juggling 10 dashboards, dead API keys, and surprise bills. + + + + + + THE DAILY PAIN + + + + HOW OMNIROUTE FIXES IT + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + Subscription quota expires unused every month + Rate limits stop you mid-coding + Tool outputs (git diff, grep, logs) burn tokens + Expensive APIs — $20–50/mo per provider + Each AI tool wants its own setup + AI blocked in your country + Dead keys and banned accounts kill your flow + One subscription, a whole team fighting over it + Your prompts routed through someone else's cloud + No idea where tokens and money go + + + + + Maximize subscriptions — track quota, use every token before reset + 4-tier auto-fallback — Subscription → API → Cheap → Free, in ms + RTK + Caveman compression — save 15–95% eligible tokens + Cost-optimized routing — auto-route to the cheapest viable model + One endpoint — every tool, one config, one dashboard + 3-level proxy + TLS stealth — use AI from anywhere + 3-layer resilience — circuit breakers, key cooldown, model lockout + Key pools — fair-share quotas per member, hot-path enforced + Local-first — your machine, keys encrypted (AES-256-GCM) + Live analytics — usage, quota, savings & p95 latency per provider + + + diff --git a/docs/reference/PROVIDER_REFERENCE.md b/docs/reference/PROVIDER_REFERENCE.md index 89d0a3f625..370f7d7bea 100644 --- a/docs/reference/PROVIDER_REFERENCE.md +++ b/docs/reference/PROVIDER_REFERENCE.md @@ -1,16 +1,16 @@ --- title: "Provider Reference" version: 3.8.49 -lastUpdated: 2026-07-18 +lastUpdated: 2026-07-19 --- # Provider Reference > **Auto-generated** from `src/shared/constants/providers.ts` — do not edit by hand. > Regenerate with: `npm run gen:provider-reference` -> **Last generated:** 2026-07-18 +> **Last generated:** 2026-07-19 -Total providers: **265**. See category breakdown below. +Total providers: **268**. See category breakdown below. ## Categories @@ -31,7 +31,7 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each --- -## OAuth Providers (22) +## OAuth Providers (23) | ID | Alias | Name | Tags | Website | Notes | |----|-------|------|------|---------|-------| @@ -49,12 +49,13 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each | `gitlab-duo` | `gitlab-duo` | GitLab Duo | OAuth | [link](https://docs.gitlab.com/user/duo_agent_platform/code_suggestions/) | OAuth application with ai_features + read_user scopes. Configure GITLAB_DUO_OAUTH_CLIENT_ID and optionally GITLAB_DUO_OAUTH_CLIENT_SECRET on this OmniRoute instance. | | `grok-cli` | `gc` | Grok Build | OAuth | — | Paste your ~/.grok/auth.json (or the JWT access token) from the Grok Build CLI; refresh_token is rotated automatically. | | `kilocode` | `kc` | Kilo Code | OAuth | — | — | -| `kimi-coding` | `kmc` | Kimi Coding | OAuth | — | — | +| `kimi-coding` | `kmc` | Kimi Code CLI | OAuth | — | Sign in with the same Kimi account used by Kimi Code CLI. OmniRoute uses the CLI OAuth flow and Kimi Coding Plan endpoints. | | `kiro` | `kr` | Kiro AI | OAuth | — | Free tier: 50 credits/month (~25K–100K tokens). ⚠️ Kiro ToS prohibits third-party proxy/harness use. | | `qoder` | `if` | Qoder | OAuth | — | — | | `qwen` | `qw` | Qwen Code | OAuth | — | ⚠️ **DEPRECATED.** Qwen OAuth free tier was discontinued on 2026-04-15. Use 'bailian-coding-plan', 'alibaba', 'alibaba-cn', or 'openrouter' provider with API key instead. | | `trae` | `tr` | Trae | OAuth | [link](https://trae.ai) | Trae is an AI-native IDE by ByteDance (SOLO remote agent). Authorize via trae.ai in the popup, or sign in at solo.trae.ai and paste the Cloud-IDE-JWT (sent as 'Authorization: Cloud-IDE-JWT ', ~14-day lifetime) as the access token; web_id/biz_user_id/user_unique_id/scope/tenant/region propagate via providerSpecificData. No headless refresh for pasted tokens — re-paste on expiry. | | `windsurf` | `ws` | Windsurf (Devin CLI) | OAuth | [link](https://windsurf.com) | In the Windsurf / VS Code IDE, open the command palette and run `Windsurf: Provide Auth Token` (or click the Jupyter "Get Windsurf Authentication Token" button), then copy the shown token and paste it here. Note: opening windsurf.com/show-auth-token directly only renders a "Redirecting" page — the IDE must initiate the flow (it adds a `?state=...` param) for the token to appear. | +| `xai-oauth` | `xao` | xAI OAuth (Grok) | OAuth | [link](https://x.ai) | Sign in with xAI to use api.x.ai models such as Grok 4.5. This is separate from Grok Build JWT sessions, which use cli-chat-proxy.grok.com and grok-build model aliases. | | `zed` | `zd` | Zed IDE | OAuth | [link](https://zed.dev) | Zed stores LLM provider credentials (OpenAI, Anthropic, Google, Mistral, xAI) in the OS keychain. Use the Import button below to discover and import them automatically. | | `zed-hosted` | — | Zed Hosted Models | OAuth | [link](https://zed.dev) | Sign in with your Zed account (native-app sign-in). OmniRoute generates a one-time RSA keypair and opens zed.dev to authorize it — on a remote/headless install, copy the resulting 127.0.0.1 callback URL from your browser's address bar and paste it back here. Distinct from the 'Zed IDE' credential-import entry above: this proxies chat completions through Zed's own hosted model aggregator (cloud.zed.dev), fronting Anthropic/OpenAI/Google/xAI models under your Zed plan. | @@ -75,7 +76,7 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each | `grok-web` | `gw` | Grok Web (Subscription) | Web cookie | [link](https://grok.com) | Paste the full grok.com cookie line from DevTools → Application → Cookies. Include both `sso` and `sso-rw` (e.g. `sso=...; sso-rw=...`) — Grok's anti-bot rejects `sso` on its own. | | `huggingchat` | `huggingchat` | HuggingChat (Free) | Web cookie | [link](https://huggingface.co/chat) | Paste the full Cookie header from huggingface.co/chat (DevTools → Network → /chat/conversation → Request Headers → Cookie). It should include hf-chat and may also include token / aws-waf-token. | | `inner-ai` | `in-ai` | Inner.ai (Subscription) | Web cookie | [link](https://app.innerai.com) | Paste your token cookie and email separated by a space: open DevTools → Application → Cookies → .innerai.com, copy the token value, then append a space and your Inner.ai login email. Example: eyJhbG... user@example.com | -| `kimi-web` | `kimi-web` | Kimi Web (Moonshot AI) | Web cookie | [link](https://www.kimi.com) | Paste `access_token` from www.kimi.com DevTools → Application → Local Storage. A legacy `kimi-auth` cookie is also accepted. | +| `kimi-web` | `kimi-web` | Kimi Web | Web cookie | [link](https://www.kimi.com) | Paste access_token from www.kimi.com DevTools → Application → Local Storage. A legacy kimi-auth cookie is also accepted. | | `lmarena` | `lma` | Arena (Free) | Web cookie | [link](https://arena.ai) | Paste the full Cookie header from arena.ai (DevTools → Network → request → Cookie). Include arena-auth-prod-v1.0/.1… and cf_clearance/__cf_bm when present. OmniRoute uses Chrome TLS impersonation; if Arena still 403s, set providerSpecificData.recaptchaV3Token from a live browser session. | | `microsoft-designer-web` | `msdesigner` | Microsoft Designer (Image Generation) | Web cookie | [link](https://designer.microsoft.com) | Sign in at designer.microsoft.com, then open DevTools → Network, generate an image, and find the request to DallE.ashx?action=GetDallEImagesCogSci. Copy the value of its Authorization: Bearer header (the access_token — no 'Bearer ' prefix). The token is short-lived; this is an unofficial, reverse-engineered integration. | | `muse-spark-web` | `ms-web` | Muse Spark Web (Meta AI) | Web cookie | [link](https://www.meta.ai) | Paste your ecto_1_sess value or full cookie header from meta.ai | @@ -90,12 +91,13 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each | `zai-web` | `zw` | Z.ai Web (Free) | Web cookie | [link](https://chat.z.ai) | Paste the full Cookie header from chat.z.ai (must include the token= cookie) | | `zenmux-free` | `zmf` | ZenMux Free (Web) | Web cookie | [link](https://zenmux.ai) | Login at zenmux.ai, then export all cookies using EditThisCookie or Cookie-Editor and paste the full Cookie header string here. Refresh every ~30 days. | -## API Key Providers (paid / paid-with-free-credits) (177) +## API Key Providers (paid / paid-with-free-credits) (179) | ID | Alias | Name | Tags | Website | Notes | |----|-------|------|------|---------|-------| | `360ai` | `360ai` | 360 AI | API key | [link](https://ai.360.cn) | Get API key at ai.360.cn | | `agentrouter` | `agentrouter` | AgentRouter | API key, aggregator | [link](https://agentrouter.org) | $200 free credits on signup - multi-model routing gateway | +| `agnes` | `agnes` | Agnes AI | API key | [link](https://agnes-ai.com) | Get API key at agnes-ai.com | | `ai21` | `ai21` | AI21 Labs | API key | [link](https://www.ai21.com) | $10 trial credits on signup (valid 3 months), no credit card required | | `aimlapi` | `aiml` | AI/ML API | API key, aggregator | [link](https://aimlapi.com) | Free tier paused (2026) — AI/ML API is now pay-as-you-go only (min $20 top-up); no recurring free credits. | | `alibaba` | `ali` | Alibaba | API key | [link](https://bailian.console.alibabacloud.com/) | — | @@ -128,6 +130,7 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each | `command-code` | `cmd` | Command Code | API key | [link](https://commandcode.ai/) | Use a Command Code API key. Requests are sent to Command Code's /alpha/generate endpoint. | | `coze` | `coze` | Coze | API key | [link](https://coze.com) | Get API key at coze.com/open/api | | `crof` | `crof` | CrofAI | API key | [link](https://crof.ai) | — | +| `dahl` | `dahl` | Dahl | API key | [link](https://inference.dahl.global) | Click 'Add Account' to auto-generate a token. | | `databricks` | `databricks` | Databricks | API key, enterprise | [link](https://www.databricks.com) | — | | `datarobot` | `datarobot` | DataRobot | API key, enterprise | [link](https://docs.datarobot.com) | Use your DataRobot API token. Optional Base URL can be the account root (for LLM Gateway) or a deployment URL under /api/v2/deployments/. | | `deepinfra` | `deepinfra` | DeepInfra | API key | [link](https://deepinfra.com) | Free signup credits for API testing and model exploration | @@ -180,8 +183,8 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each | `kenari` | `kenari` | Kenari | API key | [link](https://kenari.id) | Use your Kenari API key (kn-...) in Authorization: Bearer . Fully OpenAI-compatible. API base URL: https://kenari.id/v1. | | `kie` | `kie` | KIE.AI | API key | [link](https://kie.ai) | — | | `kilo-gateway` | `kg` | Kilo Gateway | API key, aggregator | [link](https://kilo.ai) | — | -| `kimi` | `kimi` | Kimi | API key | [link](https://platform.moonshot.ai) | — | -| `kimi-coding-apikey` | `kmca` | Kimi Coding (API Key) | API key | [link](https://www.kimi.com/code) | — | +| `kimi` | `kimi` | Kimi (Legacy Moonshot API) | API key | [link](https://platform.moonshot.ai) | — | +| `kimi-coding-apikey` | `kmca` | Kimi Code API Key | API key | [link](https://www.kimi.com/code) | — | | `lambda-ai` | `lambda` | Lambda AI | API key | [link](https://lambda.ai) | — | | `laozhang` | `lz` | LaoZhang AI | API key, aggregator | [link](https://api.laozhang.ai) | — | | `leonardo` | `leo` | Leonardo AI | API key, video | [link](https://leonardo.ai) | Get API key at leonardo.ai/developer | diff --git a/public/providers/cli-generic.svg b/public/providers/cli-generic.svg new file mode 100644 index 0000000000..3427d8ce70 --- /dev/null +++ b/public/providers/cli-generic.svg @@ -0,0 +1 @@ +