diff --git a/README.md b/README.md index 373c1ef558..9a8e877f16 100644 --- a/README.md +++ b/README.md @@ -17,9 +17,9 @@ -> Stacking free tiers by hand is painful — dozens of SDKs, dozens of rate limits, and no idea how much you actually have. OmniRoute aggregates the **documented** free tiers of **42 provider pools / 495 models** into one honest number and shows it live on the dashboard (`/dashboard/free-tiers`). +> Stacking free tiers by hand is painful — dozens of SDKs, dozens of rate limits, and no idea how much you actually have. OmniRoute catalogs **455 free-tier entries across 40 recurring pool keys** and computes the token headline from the **20 pools with a published positive monthly budget**, deduplicated by shared pool. The result stays visible on the dashboard (`/dashboard/free-tiers`). -OmniRoute free-tier budget card: ~1.51B free tokens per month steady, up to ~2.13B in the first month with signup credits, from the documented free tiers of 42 provider pools / 495 models behind one endpoint. Honest pool-deduped math — each shared pool counted once (counting every rate limit 24/7 would read ~10B; not published), 15 providers ToS-flagged so you decide. Budget bar of the countable free pools with per-model grid (Mistral Large 3 1B, GPT-4o mini 150M, Gemini 2.5 Flash 60M … Claude Sonnet 4.5 25K), one-time first-month signup credits (vertex 300M, agentrouter 200M, predibase 25M, together 25M, glm-cn 20M, doubao 15M, ai21 10M, longcat 10M, deepseek 5M, hyperbolic 5M, nscale 5M), plus permanently-free no-token-cap providers (SiliconFlow, Z.AI GLM-Flash, Kilo, OpenCode Zen, baidu …) and a $10 OpenRouter top-up unlocking +24M/mo — surfaced separately so they never inflate the headline. Live used/remaining on /dashboard/free-tiers. +OmniRoute free-tier budget card: ~1.51B free tokens per month steady, up to ~2.13B in the first month with signup credits, from 40 documented recurring pool keys covering 455 cataloged free-tier entries behind one endpoint. Honest pool-deduped math — each shared pool counted once, including 20 recurring pools with a published positive monthly token budget; 15 providers are marked avoid in the terms-risk catalog so you decide. Budget bar includes Mistral 1B, LLM7 150M, Nara 150M, Gemini 60M and smaller pools, plus first-month signup credits and permanently-free no-token-cap providers surfaced separately so they never inflate the headline. Live used/remaining on /dashboard/free-tiers. > Animated summary of the live `/dashboard/free-tiers` page. Full methodology (pool dedupe, credit tiers, provider terms): **[docs/reference/FREE_TIERS.md](docs/reference/FREE_TIERS.md)**. > @@ -61,14 +61,14 @@
-| | v3.8.49 | **v3.8.50** | `v3.8.51+` | -| ------------------------- | :-----: | :---------: | :---------: | -| 🌐 Providers | 290 | **342** | more queued | -| 🧠 Documented models | 1185 | **1202** | — | -| 🖼️ Modality Bridge | — | 🆕 vision | video | -| 📡 Radar free catalog | — | 🆕 opt-in | — | -| ⚖️ Quota-aware scheduling | — | — | 🔭 next | -| 📊 Quota telemetry | — | — | 🔭 next | +| | v3.8.49 | **v3.8.50** | `v3.8.51+` | +| ------------------------- | :-----: | :-----------------------: | :---------: | +| 🌐 Providers | 290 | **350** | more queued | +| 🧠 Unique chat model IDs | 1185 | **1312** | — | +| 🖼️ Modality Bridge | — | 🆕 vision + audio + video | — | +| 📡 Radar free catalog | — | 🆕 opt-in | — | +| ⚖️ Quota-aware scheduling | — | 🆕 Quota-Share | — | +| 📊 Quota telemetry | — | 🆕 live | — | **→ [Roadmap](ROADMAP.md) — riding the rail to `v3.9.0 LTS`** @@ -101,7 +101,7 @@ ⚙️ Features 🎯 Combos - 🌐 Providers + 🌐 Providers 🔌 CLI & MCP @@ -126,7 +126,7 @@ 📦 Project 🛠️ Tech Stack 📖 Docs - 👥 Contributors + 👥 Contributors @@ -210,7 +210,7 @@ curl http://localhost:20128/v1/chat/completions \
-The Promise — One endpoint. 350 providers. Never stop building — OmniRoute picks the cheapest one that works. Six pillars: Never hit limits (auto-fallback across 350 providers in milliseconds, zero downtime) · Save up to 95% tokens (RTK + Caveman stacked compression cuts 15–95%, ~89% avg on tool-heavy sessions) · $0 to start (90+ free tiers, 56 free forever — no card needed) · Every tool works (33 coding agents through one config) · One endpoint (OpenAI ↔ Claude ↔ Gemini ↔ Responses API at /v1) · Production-grade (circuit breakers, TLS stealth, MCP 110 tools, A2A, memory, guardrails, evals — 25,000+ tests). +The Promise — One endpoint and 350 providers. Automatic fallback keeps routing while another healthy target is available. Six pillars: resilient fallback across 350 providers · up to 95% token savings on eligible workloads · $0 to start with 90+ free tiers and 56 recurring/keyless free-forever providers · 35 CLI/agent integrations through one config · OpenAI, Claude, Gemini and Responses API compatibility at /v1 · production controls including circuit breakers, TLS stealth, MCP 110 tools, A2A, memory, guardrails, evals and 39,000+ static test declarations across 5,100+ tracked test files.

@@ -225,7 +225,7 @@ curl http://localhost:20128/v1/chat/completions \
-OmniRoute request flow: your IDE or CLI (Claude Code, Cursor, Cline…) calls one local endpoint (http://localhost:20128/v1); the OmniRoute Smart Router (RTK + Caveman compression, 19 routing strategies, circuit breakers, TLS stealth, MCP, A2A, guardrails) auto-falls back across 4 provider tiers — Tier 1 Subscription (Claude Code, Codex, Copilot), quota out? Tier 2 API Key (DeepSeek, Groq, xAI), budget hit? Tier 3 Cheap (GLM $0.5, MiniMax $0.2), budget hit? Tier 4 Free (Kiro, Qoder, Pollinations) — always on. +OmniRoute request flow: your IDE or CLI (Claude Code, Cursor, Cline…) calls one local endpoint (http://localhost:20128/v1); the OmniRoute Smart Router (RTK + Caveman compression, 19 routing strategies, circuit breakers, TLS stealth, MCP, A2A, guardrails) can fall back across 4 provider tiers while an eligible healthy target remains — Tier 1 Subscription, Tier 2 API Key, Tier 3 Cheap and Tier 4 Free.
@@ -318,7 +318,7 @@ curl http://localhost:20128/v1/chat/completions \ All 19 combo routing strategies animated — one tile per strategy: priority, fill-first, weighted, round-robin, p2c, least-used, random, strict-random, cost-optimized, headroom, reset-window, reset-aware, context-relay, context-optimized, cache-optimized, lkgp, auto, fusion, pipeline. See the table above for what each one does. -> A **combo** is a chain of models OmniRoute routes across **automatically**. Quota runs out, a provider fails, or costs spike — the combo silently slides to the next model. **This is what makes OmniRoute unbreakable.** 🛡️ +> A **combo** is a chain of models OmniRoute routes across **automatically**. If quota runs out, a provider fails, or costs spike, the combo can move to the next eligible healthy model. 🛡️ ### ⚡ Zero-config — just use `auto` @@ -429,7 +429,7 @@ All **19** strategies — mix & match per combo step: 17 auto - 14-factor live scoring across every connection 🤖 + 15-factor live scoring across every connection 🤖 18 @@ -443,7 +443,7 @@ All **19** strategies — mix & match per combo step: -The Auto-Combo engine scores every candidate on **14 factors** (health, quota, cost, latency, success rate, freshness…) — see [`docs/routing/AUTO-COMBO.md`](docs/routing/AUTO-COMBO.md). +The Auto-Combo engine scores every candidate on **15 factors** (health, quota, cost, latency, task fit, quality, session availability…) — see [`docs/routing/AUTO-COMBO.md`](docs/routing/AUTO-COMBO.md). ## @@ -461,7 +461,7 @@ All **19** strategies — mix & match per combo step: -What sets OmniRoute apart — comparison table vs 9router, OpenRouter, CLIProxyAPI and LiteLLM across 13 capabilities. OmniRoute: 350 providers, 90+ free providers built-in, 19 routing strategies, 12-engine token compression, built-in MCP server with 110 tools, A2A agent protocol, persistent memory, guardrails, cloud agents, TLS fingerprint stealth, Desktop/Termux/PWA, 43 i18n UI locales, 100% MIT self-hosted. OmniRoute is the only one with the full set; competitors show a mix of checks, partials and crosses. Verified from each project's docs. +What sets OmniRoute apart — a dated feature snapshot vs 9router, OpenRouter, CLIProxyAPI and LiteLLM across 13 capabilities. OmniRoute: 350 providers, 90+ free tiers built in, 19 routing strategies, 12-engine token compression, built-in MCP server with 110 tools, A2A agent protocol, persistent memory, guardrails, cloud agents, TLS fingerprint stealth, Desktop/Termux/PWA and 43 i18n UI locales. OmniRoute is MIT-licensed and self-hostable. Competitor capabilities and counts may change; see the linked methodology. 📊 Full methodology & per-feature detail vs 9router, OpenRouter, CLIProxyAPI & LiteLLM → [`docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md`](docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md) @@ -517,9 +517,9 @@ Pix copia-e-cola: ## 📡 OmniRoute Radar -The main free-tier headline remains **~1.53B tokens/month** from the documented, +The main free-tier headline remains **~1.51B tokens/month** from the documented, pool-deduplicated catalog above. Temporary provider signup credits can separately lift the first -month to **~2.15B**. Radar is an optional, signed catalog overlay for people who want fresher +month to **~2.13B**. Radar is an optional, signed catalog overlay for people who want fresher free-model availability between OmniRoute releases; the community catalog and every existing free feature remain free. @@ -548,7 +548,7 @@ the current catalog at **[radar.omniroute.online/planos](https://radar.omniroute - **🗜️ Compression hardening** — default-on inflation guard, Caveman packs for DE / FR / JA + Chinese (wényán), RTK filters for Gradle & .NET. → [Compression](docs/compression/COMPRESSION_ENGINES.md) - **💸 Honest flat-rate cost** — subscription / coding-plan providers read **$0** in cost analytics; budget, quota & routing keep estimating. → [API Reference](docs/reference/API_REFERENCE.md) - **⚖️ Quota-Share routing** — split a shared account's quota fairly across pooled keys, work-conserving so idle slices are lent out. → [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md) -- **🤖 One-command CLI/agent setup** — `setup-*` configures 12+ coding tools; `omniroute run` launches 7 CLIs (Claude Code, Codex, Aider, Goose, OpenCode, Qwen Code, Gemini CLI) with zero config written; `omniroute configure` is an interactive provider+model picker with per-context favorites. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md) +- **🤖 One-command CLI/agent setup** — 12 registered `setup-*` commands; `omniroute run` launches 7 CLIs (Claude Code, Codex, Aider, Goose, OpenCode, Qwen Code, Gemini CLI); `omniroute configure` supports 9 targets with an interactive provider+model picker and per-context favorites. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md) - **🛰️ Remote mode** — drive a remote OmniRoute with scoped tokens (`connect` / `contexts` / `tokens`) + an `antigravity` OAuth helper for VPS installs. → [Remote Mode](docs/guides/REMOTE-MODE.md) - **🧭 Smarter auto-routing** — `auto/:` combos, **Fusion** (model panel + judge), task-aware routing, per-request model / mode / USD-budget overrides. → [Auto-Combo](docs/routing/AUTO-COMBO.md) - **🗜️ Pluggable compression** — 12 composable engines + Compression Studios: LLMLingua-2, two-tier Ultra, omniglyph, per-step fidelity gate, GCF v3.2, drag-reorder editor. → [Compression](docs/compression/COMPRESSION_ENGINES.md) @@ -642,11 +642,11 @@ of your shell history. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md)
-## 🌐 349 AI Providers — 90+ Free +## 🌐 350 AI Providers — 154 Catalog-Marked Free
-> The most complete catalog of any open-source router: **350 providers**, **90+ with a free tier**, **56 free forever**. +> **350 registered providers** across the canonical chat, media, search, local, cloud-agent and system collections, including **154 carrying `hasFree: true` discovery metadata**. The chat model registry covers **268 providers / 2,566 distinct provider-model pairs / 1,312 raw model IDs**; the separate free-budget catalog has **455 per-model rows**, **40 recurring pools** and **56 recurring/keyless free-forever providers**. These are different denominators by design; definitions and pool-deduped calculations live in the [Provider Reference](docs/reference/PROVIDER_REFERENCE.md) and [Free Tiers](docs/reference/FREE_TIERS.md).
@@ -679,7 +679,7 @@ of your shell history. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md) -…and 220+ more — every icon resolves live from the dashboard's provider catalog. 📖 [Provider Reference](docs/reference/PROVIDER_REFERENCE.md) +…and 330+ more — every icon resolves live from the dashboard's provider catalog. 📖 [Provider Reference](docs/reference/PROVIDER_REFERENCE.md)
@@ -769,7 +769,7 @@ From inside the editor: open the **Extensions** view, search **"OmniRoute"**, cl
-Private and local-first — your keys, your machine, your data; OmniRoute is a local proxy that never phones home. Eleven guarantees: runs 100% on your hardware (0 cloud hops), zero telemetry by default, credentials encrypted at rest (AES-256-GCM), no account or sign-up, hardened gateway (API-key scoping, IP filtering, rate limits, prompt-injection guard), loopback-only process routes, upstream header scrubbing, strictly opt-in PII redaction, sanitized errors that never leak internals, a local audit trail in your own SQLite, and MIT-licensed fully open-source code. +Private and local-first — OmniRoute's gateway and control plane run on your machine. Prompts are sent to the upstream provider selected for each request; OmniRoute adds no hosted prompt-processing hop and telemetry is disabled by default. Credentials are encrypted at rest with AES-256-GCM; controls include API-key scoping, IP filtering, rate limits, prompt-injection guards, upstream-header scrubbing, opt-in PII redaction, sanitized errors and a local SQLite audit trail. OmniRoute is MIT-licensed and self-hostable. 📖 [Authorization](docs/architecture/AUTHZ_GUIDE.md) · [Guardrails](docs/security/GUARDRAILS.md) · [Compliance](docs/security/COMPLIANCE.md) @@ -810,7 +810,7 @@ Tokens are scoped `read` / `write` / `admin`; process-spawning routes stay loopb
-Animated terminal demoing the OmniRoute CLI — omniroute providers list, omniroute combo list, omniroute health — cycling over the 80+ command surface: providers · oauth · keys · combo · nodes · models · cache · compression · cost · usage · quota · health · resilience · telemetry · logs · audit · mcp · a2a · cloud · memory · skills · eval · tunnel · backup · sync · webhooks · policy · pricing · translator · simulate … +Animated terminal demoing the OmniRoute CLI — omniroute providers list, omniroute combo list and omniroute health — cycling over the 85-command top-level surface: providers · oauth · keys · combo · nodes · models · cache · compression · cost · usage · quota · health · resilience · telemetry · logs · audit · mcp · a2a · cloud · memory · skills · eval · tunnel · backup · sync · webhooks · policy · pricing · translator · simulate …
@@ -846,7 +846,7 @@ claude mcp add-server omniroute --type http --url http://localhost:20128/api/mcp ### 📖 How it works — pipeline, architecture & savings math -OmniRoute compression pipeline: a client request of 10,000 tokens passes through 12 stacked engines — Session-Dedup, CCR, Lite, RTK, Responses Tool Output, Headroom, Relevance, Caveman, Aggressive, LLMLingua-2, Ultra, OmniGlyph — and reaches the provider at about 1,080 tokens, up to 95% saved. Code, URLs and JSON are always preserved byte-perfect. +OmniRoute compression pipeline: an illustrative 10,000-token client request passes through 12 composable engines — Session-Dedup, CCR, Lite, RTK, Responses Tool Output, Headroom, Relevance, Caveman, Aggressive, LLMLingua-2, Ultra and OmniGlyph — and can reach the provider at about 1,080 tokens in the documented stacked example. Structured content is protected by preservation guards and per-step fidelity gates; explicit lossy or experimental modes may transform eligible content. Default stacked combo runs `RTK → Caveman`. When both act on the same tool/context payload, savings compound: @@ -1013,6 +1013,7 @@ Full table: [Docker Guide — runtime RAM](docs/guides/DOCKER_GUIDE.md#runtime-r **🥟 Bun** Standard `bun install` and global installation (`bun install -g omniroute`) are supported via Bun runtime detection: + - **Built-in `bun:sqlite`**: OmniRoute uses Bun's built-in `bun:sqlite` driver when running under Bun, falling back to `better-sqlite3` on Node.js or `sql.js`. - **Automatic Webpack bundler selection**: Development (`bun run dev`) and production builds (`bun run build`) automatically detect Bun and disable Turbopack in favor of Webpack to prevent native V8 binding incompatibilities. - **Dedicated Bun Dockerfile**: Multi-stage `Dockerfile.bun` for native Bun production deployments (`docker build -f Dockerfile.bun -t omniroute:bun .`). @@ -1105,7 +1106,7 @@ same process on one port, so there is no separate CLI-only package today.
-Dados de cobertura social em 2026-08-17 · YT: 741 | TT: 137 | IG: 124 · Frescor (dias): YT 0 · TT 14 · IG 15 +Snapshot do painel em 2026-08-24 · Catálogo bruto: YT 809 | TT 137 | IG 124 · Frescor (dias): YT 1 | TT 21 | IG 22 @@ -1114,52 +1115,52 @@ same process on one port, so there is no separate CLI-only package today. Instagram Reel
🎬 #1 — Instagram
- nick_saraev — 1,628,910 views + nick_saraev — 3,042,474 views + + + - -
+ + Instagram Reel — theopenstack +
+ 🎬 #2 — Instagram
+ theopenstack — 692,419 views +
+ + TikTok — milesreevesai +
+ 🎬 #3 — TikTok
+ milesreevesai — 620,400 views
YouTube — Vaibhav Sisinty
- 🎬 #2 — YouTube
- Vaibhav Sisinty — 373,084 views + 🎬 #4 — YouTube
+ Vaibhav Sisinty — 391,109 views
- - YouTube Shorts + + Instagram Reel — buildwithai.club
- 🎬 #3 — YouTube Shorts
- Nick Automates — 207,714 views -
- - TikTok Thumbnail -
- 🎬 #4 — TikTok
- milesreevesai — 620,400 views -
- - Valency Labs -
- 🎬 #5 — YouTube
- Valency Labs — 135,974 views + 🎬 #5 — Instagram
+ buildwithai.club — 347,652 views
-**Ranking completo (`v > 0`, maior alcance):** +**Ranking completo (URLs canônicas deduplicadas, `v > 0`, maior alcance):** -| #1 | #2 | #3 | #4 | #5 | -| -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | -| [nick_saraev — Instagram](https://www.instagram.com/reel/Da8ZthUPK98/) — **1,628,910** | [milesreevesai — TikTok](https://www.tiktok.com/@milesreevesai/video/7667980059189366019) — **620,400** | [Vaibhav Sisinty — YouTube](https://www.youtube.com/watch?v=QucgvbO5gsM) — **373,084** | [Nick Automates — YouTube Shorts](https://www.youtube.com/shorts/fZIBK_4fKq8) — **207,714** | [midudev — TikTok](https://www.tiktok.com/@midudev/video/7664636453544152342) — **177,800** | +| #1 | #2 | #3 | #4 | #5 | +| -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | +| [nick_saraev — Instagram](https://www.instagram.com/reel/Da8ZthUPK98/) — **3,042,474** | [theopenstack — Instagram](https://www.instagram.com/reel/DaSs65mMrHk/) — **692,419** | [milesreevesai — TikTok](https://www.tiktok.com/@milesreevesai/video/7667980059189366019) — **620,400** | [Vaibhav Sisinty — YouTube](https://www.youtube.com/watch?v=QucgvbO5gsM) — **391,109** | [buildwithai.club — Instagram](https://www.instagram.com/reel/DbIt9AjK7-U/) — **347,652** | -| #6 | #7 | #8 | #9 | #10 | -| ------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | -| [theopenstack — Instagram](https://www.instagram.com/reel/DaSs65mMrHk/) — **155,453** | [t.ghoush.ai — TikTok](https://www.tiktok.com/@t.ghoush.ai/video/7669497680527248656) — **152,800** | [Valency Labs — YouTube](https://www.youtube.com/watch?v=LkP6ocAoQkk) — **135,974** | [Asati — YouTube](https://www.youtube.com/watch?v=JjPtJcqwhqg) — **126,130** | [Vaibhav Sisinty — YouTube](https://www.youtube.com/watch?v=NuNDpeZYQ28) — **122,672** | +| #6 | #7 | #8 | #9 | #10 | +| ----------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | +| [nivedan.ai — Instagram](https://www.instagram.com/reel/DbIrCksJiqq/) — **331,973** | [vaibhavsisinty — Instagram](https://www.instagram.com/reel/Dae05TSAK1l/) — **263,744** | [Nick Automates — YouTube Shorts](https://www.youtube.com/shorts/fZIBK_4fKq8) — **218,174** | [theroshankrishna — Instagram](https://www.instagram.com/reel/Dapjs58z0P0/) — **186,786** | [midudev — TikTok](https://www.tiktok.com/@midudev/video/7664636453544152342) — **177,800** | -Métricas de validação: 1002 vídeos rastreados · 7,069,190 visualizações conhecidas · 595 perfis/canais · 13+ idiomas · 13+ criadores. +Métricas canônicas em 2026-08-24: **1.029 vídeos únicos** · **11.132.922 visualizações conhecidas** (`v > 0`) · **639 canais/perfis por rede**. O painel bruto contém 1.070 linhas; 41 duplicatas do Instagram foram normalizadas pela URL canônica, mantendo a maior contagem por vídeo. > 🎬 **Made a video about OmniRoute?** Open an [issue](https://github.com/diegosouzapw/OmniRoute/issues/new) or [discussion](https://github.com/diegosouzapw/OmniRoute/discussions) with the link — we'll feature it here. @@ -1211,7 +1212,7 @@ Métricas de validação: 1002 vídeos rastreados · 7,069,190 visualizações c Stealthwreq-js — JA3 / JA4 TLS fingerprint impersonation, 3-level proxy ResilienceCircuit breaker, exponential backoff, anti-thundering-herd, auto-combo self-healing Loggingpino — structured JSON logs with request context - TestingNode.js test runner + Vitest — 25,000+ test cases across 3,300+ files (unit, integration, E2E, security, ecosystem) + TestingNode.js test runner + Vitest — 39,000+ static test declarations across 5,100+ tracked test files (unit, integration, E2E, security, ecosystem) PlatformsDesktop (Electron) · Android (Termux) · PWA (any browser) CI/CDGitHub Actions — auto npm publish + Docker Hub on release LinksWebsite · npm · Docker Hub @@ -1262,9 +1263,9 @@ Métricas de validação: 1002 vídeos rastreados · 7,069,190 visualizações c Compression Rules FormatJSON rule-pack schemas for Caveman and RTK filters Compression Language PacksLanguage detection and Caveman rule-pack authoring Resilience GuideCircuit breakers, cooldowns, queue, anti-thundering herd, TLS spoofing - Auto-Combo Engine14-factor scoring, mode packs, self-healing + Auto-Combo Engine15-factor scoring, mode packs, self-healing Proxy Guide3-level proxy system, 1proxy marketplace, registry CRUD - Free Tiers90+ free providers consolidated directory (42 documented token pools / 495 models) + Free TiersConsolidated directory: 40 documented recurring pools / 455 cataloged free-tier entries Features GalleryVisual dashboard tour with screenshots Codebase DocumentationBeginner-friendly codebase walkthrough @@ -1275,7 +1276,7 @@ Métricas de validação: 1002 vídeos rastreados · 7,069,190 visualizações c DocumentDescription API ReferenceAll endpoints with examples OpenAPI SpecOpenAPI 3.0 specification - MCP Server109 MCP tools, IDE configs, Python/TS/Go clients + MCP Server110 MCP tools, IDE configs, Python/TS/Go clients MCP Server GuideMCP installation, transports, and tool reference A2A ServerJSON-RPC 2.0 protocol, skills, streaming, task mgmt A2A Server GuideA2A agent card, tasks, skills, and streaming @@ -1291,7 +1292,7 @@ Métricas de validação: 1002 vídeos rastreados · 7,069,190 visualizações c Security PolicyVulnerability reporting and security practices i18n Guide43-language support, translation workflow, RTL Release ChecklistPre-release validation steps - Coverage PlanTest coverage strategy and 25,000+ test suite + Coverage PlanTest coverage strategy for 39,000+ static test declarations across 5,100+ tracked test files
@@ -1302,93 +1303,123 @@ Métricas de validação: 1002 vídeos rastreados · 7,069,190 visualizações c > OmniRoute is shaped by a passionate open-source community. These individuals have made exceptional contributions that directly impact the quality, stability, and reach of the project. **Thank you.** +### External contributors by merged pull requests + + + + + + + + + + + + + + + + + + + + + + + + +
RankContributorMerged PRs~Changed lines
1backryun190227,977
2oyi77180407,678
3rdself14580,663
4JxnLexn128387,049
5KooshaPari101125,747
6herjarsa88230,872
7RaviTharuma7955,106
8maxmad64bis69394,715
9artickc5933,260
10HouMinXi5147,334
10chirag127515,153
12xz-dev50245,976
13hartmark4752,185
14rqzbeh39143,181
15dhaern3419,559
16Dingding-leo331,986
17NomenAK3213,854
18MumuTW3016,953
19benzntech2911,641
20pacocartones249,331
20Prudhvivuda246,312
+ +Frozen at live release/v3.8.50 tip dafb4ae808, with merges through 2026-08-24 05:26:03 UTC. The paginated GitHub GraphQL census contains 5,911 merged PRs: 2,707 by the repository owner, 179 by Dependabot, and 3,025 external PRs from 535 distinct contributors. “Changed lines” is GitHub additions + deletions and includes generated files, lockfiles, catalogs, translations and documentation; it is churn, not authored LOC. Ties at the cutoff are retained. + +### GitHub-attributed commits + - - - - - - - + + + + + + + +
- - oyi77
- oyi77 -

- 🥇 213 commits • +114K lines
- Analytics engine, SQL aggregations,
proxy marketplace, test coverage
-
- - R.D. & Randi
- R.D. & Randi -

- 🥈 108 commits • +38K lines
- Endpoints page, tunnel integrations,
Docker workflows, A2A status, compression UI
-
- - Chris Staley
- Chris Staley -

- 🥉 70 commits • +1.8K lines
- SSE stream hardening, Responses API,
Gemini pagination, test regression fixes
-
- - zenobit
- zenobit -

- 🏅 62 commits • +22K lines
- CI/CD pipeline, i18n for 33 languages,
Void Linux package, platform fixes
-
- - Jan Leon
- Jan Leon -

- 🏅 58 commits • +22K lines
- Reasoning-effort routing, proxy controls,
quota visibility, Live Zone compression
-
backryun
backryun

- 🏅 53 commits • +70K lines
- Provider catalog curation — Perplexity, Kimi,
Cerebras, Copilot, LMArena refreshes
+ 🥇 220 GitHub-attributed commits
- - Chirag Singhal
- Chirag Singhal +
+ Paijo
+ Paijo

- 🏅 46 commits • +4.8K lines
- Error sanitization, MITM prefill fix,
fusion judge, breaker/429 correctness
+ 🥈 219 GitHub-attributed commits
- - kfiramar
- kfiramar +
+ Randi
+ Randi

- 🏅 38 commits • +1.7K lines
- Codex websocket + passthrough, auth/onboarding,
Electron hardening, DB migrations
+ 🥉 108 GitHub-attributed commits
- - Benson K B
- Benson K B +
+ Ravi Tharuma
+ Ravi Tharuma

- 🏅 28 commits • +9.2K lines
- Electron desktop app, auto-updater,
release build workflows, cross-platform CI
+ 🏅 81 GitHub-attributed commits
- - Hernan J. Ardila
- Hernan J. Ardila +
+ Chris
+ Chris

- 🏅 25 commits • +174K lines
- Zero-latency combos, vision-bridge auto-routing,
catalog context-length, resilience 429 hints
+ 🏅 70 GitHub-attributed commits +
+ + Markus Hartung
+ Markus Hartung +

+ 🏅 69 GitHub-attributed commits · tied #6 +
+ + Dizzle
+ Dizzle +

+ 🏅 69 GitHub-attributed commits · tied #6 +
+ + Jan Leon
+ Jan Leon +

+ 🏅 64 GitHub-attributed commits +
+ + zenobit
+ zenobit +

+ 🏅 62 GitHub-attributed commits +
+ + Bob.Hou
+ Bob.Hou +

+ 🏅 51 GitHub-attributed commits · tied #10 +
+ + Xiangzhe
+ Xiangzhe +

+ 🏅 51 GitHub-attributed commits · tied #10
+Rechecked at 2026-08-24 06:14:31 UTC: GitHub-attributed commits reported by the repository Contributors API for the release/v3.8.50 default branch. The API returned 525 identities (415 users, 2 bots, 108 anonymous); this table excludes the maintainer, bots and anonymous identities and retains competition ties. It is distinct from both the merged-PR ranking above and the 639-person Git-metadata census below. + > 🙏 These contributors' features, bug fixes, and infrastructure improvements are a **core part** of what makes OmniRoute reliable and feature-rich. Every pull request, every test case, and every i18n translation file matters. Open source is built by people like them. @@ -1405,25 +1436,48 @@ A heartfelt thank-you to the people who fund OmniRoute out of their own pocket + + +
+ + Andrew
+ Andrew +

+ 💛 Active monthly sponsor +
+ + Vlad I
+ Vlad I +

+ 💛 Active monthly sponsor +
+ + Paco Cartones
+ Paco Cartones +

+ 💛 Active one-time sponsor +
Professor Igor Morais Vasconcelos
Prof. Igor Morais

- 💛 Sponsor + 💛 Past one-time supporter
longtao
longtao

- 💛 Sponsor + 💛 Past one-time supporter
… and others who prefer to stay private 💛 +Public GitHub Sponsors revalidated on 2026-08-24. GitHub's activeOnly status determines the active labels above; previously disclosed public one-time supporters remain thanked, and private sponsors remain anonymous. + 💖 Become a sponsor → — every dollar keeps OmniRoute free and independent. @@ -1432,11 +1486,13 @@ A heartfelt thank-you to the people who fund OmniRoute out of their own pocket
-## 👥 320+ Contributors +## 👥 600+ Contributors
-[![Contributors](https://contrib.rocks/image?repo=diegosouzapw/OmniRoute&max=400&columns=20&anon=1)](https://github.com/diegosouzapw/OmniRoute/graphs/contributors) +[![Contributors](https://contrib.rocks/image?repo=diegosouzapw/OmniRoute&max=639&columns=20&anon=1)](https://github.com/diegosouzapw/OmniRoute/graphs/contributors) + +Audited on 2026-08-24 at frozen base ac02c5b42f and rechecked at live release/v3.8.50 tip dafb4ae808: 639 normalized human Git identities — 407 appear as commit authors (including the maintainer) and 232 only in explicit Co-authored-by trailers. The census normalizes GitHub noreply handles, excludes 26 bot/agent/service/placeholder identities, and does not merge ordinary email addresses merely because their display names match. ### How to Contribute @@ -1453,7 +1509,8 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) for detailed guidelines. ```bash # Create a release — npm publish happens automatically -gh release create v3.8.2 --title "v3.8.2" --generate-notes +VERSION=x.y.z +gh release create "v${VERSION}" --title "v${VERSION}" --generate-notes ```
@@ -1495,88 +1552,108 @@ gh release create v3.8.2 --title "v3.8.2" --generate-notes OmniRoute stands on the shoulders of giants. It started as a fork of **[9router](https://github.com/decolua/9router)** and a TypeScript port of the Go project **[CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI)** — and from there, every subsystem below was inspired by an open-source project that got there first. Each one shaped a concrete piece of OmniRoute. This is our thank-you to all of them. 🙏 -> ⭐ star counts as of July 2026 — go give these projects a star. +> ⭐ star counts verified from GitHub's REST API on August 24, 2026 — go give these projects a star. Counts are an exact dated snapshot and will naturally change. ### 🧬 Lineage & gateway - - - + + + + + + + + + + + + + + +
ProjectHow it inspired OmniRoute
9router22.7kThe original project this fork is built on — extended here with multi-modal APIs and a full TypeScript rewrite.
CLIProxyAPI43.6kThe Go implementation that inspired this JavaScript / TypeScript port.
LiteLLM54.0kThe AI gateway whose public pricing dataset feeds our cost-tracking sync and whose provider-normalization model informed our routing.
9router26,161The original project this fork is built on — extended here with multi-modal APIs and a full TypeScript rewrite.
CLIProxyAPI48,497The Go implementation that inspired this JavaScript / TypeScript port.
LiteLLM57,100The AI gateway whose public pricing dataset feeds our cost-tracking sync and whose provider-normalization model informed our routing.
codex-chatgpt-web1,410MIT source adapted into the vendored ChatGPT Web → Codex Responses bridge, including browser-session, response-framing, usage and web-search adapters.
free-claude-code48,112Patterns ported into stream recovery, no-thinking aliases, fallback web search, sliding-window limits, log redaction and hardened launcher flows.
composer-api322Cursor Composer tool-choice, output-constraint and tool-commit patterns adapted into the native Cursor executor.
codex-multi-auth457Fresh-login and refresh-token rotation patterns ported into Codex OAuth reauthentication.
opencode-anthropic-auth510Claude Code-compatible transform defaults and billing-header behavior generalized into OmniRoute's config-driven bridge.
grok2api-merged2Its Grok model mappings, fake-TypeError Statsig generator, request and device defaults, and NDJSON response processor were materially adapted into OmniRoute's Grok Web executor.
TQZHR/grok2api705The principal transitive code source behind grok2api-merged; its model, header, payload, Statsig and processor implementations are preserved in the Grok Web lineage.
chenyme/grok2api7,520The underlying MIT source for Grok payload and device defaults, the Statsig generator, and the result.response processor carried through TQZHR and grok2api-merged.
grok2api-pro27A transitive source credited by grok2api-merged for its proxy-pool layer; OmniRoute preserves that lineage notice but does not claim a proxy-pool port in its bounded Grok Web executor.
GrokProxy50Its cookie-authenticated Grok proxy and result.response.token streaming pattern informed OmniRoute's Grok Web transport.
GrokBridge5The original Grok Web implementation consulted its HTTP/browser upstream design; its direct HTTP path derives from GrokProxy, so no independent code port is claimed.
grok-web-api14Its Rust ChatOptions and response-envelope schemas informed OmniRoute's TypeScript Grok request and streaming-response types.
### 🗜️ Context & token compression — engines - - - - - - - + + + + + + + +
ProjectHow it inspired OmniRoute
Caveman90.8kThe viral "why use many token when few token do trick" project — its caveman-speak philosophy powers our standard compression mode and 30+ filler/condensation rules.
RTK – Rust Token Killer71.8kHigh-performance command-output compression — inspired our RTK engine, JSON filter DSL, raw-output recovery and the stacked RTK → Caveman pipeline.
headroom60.1kReversible context-compression (SmartCrusher) — inspired our headroom engine and the ccr retrieve-marker pattern.
LLMLingua6.5kPrompt-compression research (LLMLingua / LLMLingua-2) — inspired our async, code-safe, fail-open llmlingua engine.
llmlingua-2-js30The JS/ONNX port (MobileBERT / XLM-RoBERTa) used as the worker-thread backend for our LLMLingua engine.
Troglodita26PT-BR token compression — powers our pt-BR language pack: pleonasm reduction and filler removal tuned for Brazilian-Portuguese grammar.
ponytail86.0kThe viral "lazy senior dev" YAGNI-coder skill — inspired our less-code Output Style: smallest-working-change steering that cuts _generated_ code (the output-axis sibling to Caveman's terse prose).
Caveman100,538The viral "why use many token when few token do trick" project — its caveman-speak philosophy powers our standard compression mode and 30+ filler/condensation rules.
RTK – Rust Token Killer77,185High-performance command-output compression — inspired our RTK engine, JSON filter DSL, raw-output recovery and the stacked RTK → Caveman pipeline.
headroom67,310Reversible context-compression (SmartCrusher) — inspired our headroom engine and the ccr retrieve-marker pattern.
LLMLingua6,598Prompt-compression research (LLMLingua / LLMLingua-2) — inspired our async, code-safe, fail-open llmlingua engine.
llmlingua-2-js31The JS/ONNX port (MobileBERT / XLM-RoBERTa) used as the worker-thread backend for our LLMLingua engine.
Troglodita40PT-BR token compression — powers our pt-BR language pack: pleonasm reduction and filler removal tuned for Brazilian-Portuguese grammar.
ponytail108,957The viral "lazy senior dev" YAGNI-coder skill — inspired our less-code Output Style: smallest-working-change steering that cuts _generated_ code (the output-axis sibling to Caveman's terse prose).
i-have-adhd23,526Its action-first, ADHD-friendly response style was adapted into OmniRoute's concise output style across five languages.
### 🧩 Compact formats, token research & code-aware tooling - - - - - - - + + + + + + + + - - + + - +
ProjectHow it inspired OmniRoute
TOON24.9kToken-Oriented Object Notation — its columnar, header-plus-rows model shaped our tabular compaction stage.
GCF – Graph Compact Format22First inspired our tabular compaction stage; now its zero-dependency, lossless generic-profile encoder is vendored directly as the Headroom codec (MIT, SPDX-marked), with later numeric-domain and count-mismatch correctness fixes.
token-optimizer-mcp444Brotli/SQLite cache + per-session context-delta — inspired our session-dedup engine.
token-savior1.1kBash-output compaction + MCP profiles — inspired our compression bail-out discipline and MCP tool-manifest reduction.
token-saver117Content-aware, per-file-type output compression with failure-aware bail-out — validated our per-type dispatch and minimum-gain skip.
token-optimizer1.7k"Find the ghost tokens" — its offload + recoverable-handle pattern informed our CCR offload thinking.
TokenMizer16A session-graph + cross-turn line-dedup blueprint that informed our session-dedup design.
TOON25,233Token-Oriented Object Notation — its columnar, header-plus-rows model shaped our tabular compaction stage.
GCF – Graph Compact Format41Its compact graph format and generic-profile design informed OmniRoute's tabular compaction and Headroom codec format.
gcf-typescript4The MIT TypeScript implementation directly vendored and extended as the Headroom generic-profile codec.
token-optimizer-mcp494Brotli/SQLite cache + per-session context-delta — inspired our session-dedup engine.
token-savior1,122Bash-output compaction + MCP profiles — inspired our compression bail-out discipline and MCP tool-manifest reduction.
token-saver138Content-aware, per-file-type output compression with failure-aware bail-out — validated our per-type dispatch and minimum-gain skip.
token-optimizer1,951"Find the ghost tokens" — its offload + recoverable-handle pattern informed our CCR offload thinking.
TokenMizer28A session-graph + cross-turn line-dedup blueprint that informed our session-dedup design.
OmniCompress3Rust columnar-JSON + content-addressed retrieve + cross-message dedup — validated our headroom/ccr/session-dedup engine design and the cache-stable "compressed form is position-independent" invariant.
mcp-compressor98MCP tool-schema/description compression — informed our MCP tool-manifest cardinality reduction.
RepoMapper187Aider-style repo-map ranking — informed our repo-map / retrieval-ranking exploration.
mcp-compressor113MCP tool-schema/description compression — informed our MCP tool-manifest cardinality reduction.
RepoMapper197Aider-style repo-map ranking — informed our repo-map / retrieval-ranking exploration.
quiet-shell-mcp4Declarative shell-output reduction over MCP — validated our declarative bash-output compaction.
ts-morph6.1kTypeScript Compiler API toolkit — inspired our parser-based comment removal that preserves string, template and regex literals.
ts-morph6,162TypeScript Compiler API toolkit — inspired our parser-based comment removal that preserves string, template and regex literals.
### 🧠 Memory & RAG - - - + + +
ProjectHow it inspired OmniRoute
Mem061.2kUniversal memory layer — its proxy-as-write/read-boundary model shaped our memory architecture.
Letta (MemGPT)23.9kStateful agents with tiered memory — inspired our Context Control & Recovery (CCR) tiered model.
WFGY1.8kThe ProblemMap taxonomy of 16 recurring RAG/LLM failure modes — the shared vocabulary in our troubleshooting guide.
Mem063,902Universal memory layer — its proxy-as-write/read-boundary model shaped our memory architecture.
Letta (MemGPT)24,382Stateful agents with tiered memory — inspired our Context Control & Recovery (CCR) tiered model.
WFGY1,781The ProblemMap taxonomy of 16 recurring RAG/LLM failure modes — the shared vocabulary in our troubleshooting guide.
### 🛰️ Traffic inspection, MITM & transparent proxy - - + +
ProjectHow it inspired OmniRoute
llm-interceptor49MITM interception/analysis of coding-assistant ↔ LLM traffic — our Traffic Inspector ports its SSE merge, conversation normalization, host passthrough and secret masking (MIT).
ProxyBridge5.5kTransparent per-process proxy routing — inspired our crash-safe MITM teardown, socket idle-timeouts, /proc process attribution and TPROXY capture.
llm-interceptor66MITM interception/analysis of coding-assistant ↔ LLM traffic — our Traffic Inspector ports its SSE merge, conversation normalization, host passthrough and secret masking. The upstream's complete license text is still under provenance review.
ProxyBridge5,995Transparent per-process proxy routing — inspired our crash-safe MITM teardown, socket idle-timeouts, /proc process attribution and TPROXY capture.
### 📚 Model data, observability & UI - - - - - - + + + + + + +
ProjectHow it inspired OmniRoute
models.dev6.0kOpen database of AI model specs, pricing and capabilities — synced natively into our model catalog.
React Flow / xyflow37.7kThe node-based graph library powering our real-time Compression Studio and Combo/Routing Studio.
LangGraph37.6kLangGraph Studio's live workflow-graph visualization inspired our Studios' real-time cascade view.
Langfuse31.4kIts trace → span → generation observability model shaped our Compression Studio waterfall.
Kiali3.6kIstio service-mesh observability — inspired our circuit-breaker badges and error-edge visuals in the Routing/Combo Studio.
lobe-icons2.2kAI/LLM brand logos that render the provider icons across our dashboard.
models.dev6,555Open database of AI model specs, pricing and capabilities — synced natively into our model catalog.
React Flow / xyflow38,108The node-based graph library powering our real-time Compression Studio and Combo/Routing Studio.
LangGraph40,314LangGraph Studio's live workflow-graph visualization inspired our Studios' real-time cascade view.
Langfuse33,592Its trace → span → generation observability model shaped our Compression Studio waterfall.
Kiali3,631Istio service-mesh observability — inspired our circuit-breaker badges and error-edge visuals in the Routing/Combo Studio.
lobe-icons2,428AI/LLM brand logos that render the provider icons across our dashboard.
flag-icons12,354Provides the MIT-licensed SVG flags used by the README language selector.
### 🛡️ Security - +
ProjectHow it inspired OmniRoute
awesome-secure-defaults710A curated list of secure-by-default libraries that guides our security choices (Helmet.js, DOMPurify, ssrf-req-filter, safe-regex, Google Tink).
awesome-secure-defaults721A curated list of secure-by-default libraries that guides our security choices (Helmet.js, DOMPurify, ssrf-req-filter, safe-regex, Google Tink).
### 🧭 Complementary tools + + + + +
ProjectHow it inspired OmniRoute
ClawRouter6,564Inspired request deduplication, emergency zero-cost fallback, pluggable Auto-Combo strategies and multilingual intent classification.
Antigravity-Manager30,652Its account-aware model remapping, executable-path validation and plan-label behavior informed OmniRoute's Antigravity runtime.
vscode-antigravity-cockpit4,817Its compact quota-reset countdown format inspired the corresponding provider-limit display in OmniRoute.
AionUi32,230Its ACP integrations inspired OmniRoute's automatic detection of installed CLI agents.
CodexBar20,507Identified the Grok Build quota surface; OmniRoute then verified and corrected the live wire format independently.
## 📄 License @@ -1589,7 +1666,7 @@ MIT License - see [LICENSE](LICENSE) for details. **[⬆ Back to top](#-omniroute)** · Built with ❤️ for the open-source AI community. -OmniRoute v3.8.49 · Node ≥22.22.2 · MIT License · omniroute.online +OmniRoute v3.8.50 · Node ≥22.22.2 · MIT License · omniroute.online diff --git a/changelog.d/maintenance/11356-readme-live-metrics.md b/changelog.d/maintenance/11356-readme-live-metrics.md new file mode 100644 index 0000000000..3d06965062 --- /dev/null +++ b/changelog.d/maintenance/11356-readme-live-metrics.md @@ -0,0 +1,5 @@ +- **docs(readme):** reconcile live v3.8.50 provider, free-tier, CLI, routing, test, + community, sponsor, acknowledgment, and SVG metrics with their audited source + denominators, including a deduplicated OmniRoute-in-Action snapshot and distinct + contributor rankings for merged pull requests, GitHub-attributed commits, and Git history + ([#11356](https://github.com/diegosouzapw/OmniRoute/pull/11356)). diff --git a/docs/diagrams/README.md b/docs/diagrams/README.md index 7b6fd18f9a..ccd553555c 100644 --- a/docs/diagrams/README.md +++ b/docs/diagrams/README.md @@ -16,7 +16,7 @@ Mermaid sources (`.mmd`) and exported SVGs for OmniRoute v3.8.0 architecture flo | [auto-combo-12factor.mmd](./auto-combo-12factor.mmd) | [SVG](./exported/auto-combo-12factor.svg) | docs/routing/AUTO-COMBO.md | | [resilience-3layers.mmd](./resilience-3layers.mmd) | [SVG](./exported/resilience-3layers.svg) | docs/architecture/RESILIENCE_GUIDE.md, CLAUDE.md | | [i18n-flow.mmd](./i18n-flow.mmd) | [SVG](./exported/i18n-flow.svg) | docs/guides/I18N.md | -| [mcp-tools-107.mmd](./mcp-tools-107.mmd) | [SVG](./exported/mcp-tools-107.svg) | docs/frameworks/MCP-SERVER.md | +| [mcp-tools-107.mmd](./mcp-tools-107.mmd) | [SVG](./exported/mcp-tools-107.svg) | docs/frameworks/MCP-SERVER.md | | [cloud-agent-flow.mmd](./cloud-agent-flow.mmd) | [SVG](./exported/cloud-agent-flow.svg) | docs/frameworks/CLOUD_AGENT.md | | [authz-pipeline.mmd](./authz-pipeline.mmd) | [SVG](./exported/authz-pipeline.svg) | docs/architecture/AUTHZ_GUIDE.md | | [db-schema-overview.mmd](./db-schema-overview.mmd) | [SVG](./exported/db-schema-overview.svg) | docs/architecture/CODEBASE_DOCUMENTATION.md | @@ -34,11 +34,11 @@ inside GitHub's `` sandbox: | [combo-always-on.svg](./combo-always-on.svg) | style reference | Animated priority-combo fallback (4 layers, 16s loop). Edit the SVG directly — there is no `.mmd` source. | | [cli-terminal.svg](./cli-terminal.svg) | README.md (root) | Compact half-height animated terminal (1200×350): 3 real CLI commands cycling with typewriter + scrolling subcommand ticker; first frame = completed providers screen. Edit the SVG directly — there is no `.mmd` source. | | [compression-pipeline.svg](./compression-pipeline.svg) | README.md (root) | Animated 10-engine compression funnel (8s loop). Edit the SVG directly — there is no `.mmd` source. | -| [free-tier-budget.svg](./free-tier-budget.svg) | README.md (root) | Animated free-tier budget card (~1.53B/mo quantified headline, 19-pool budget bar, per-model grid, signup credits, 10s loop). Edit the SVG directly — there is no `.mmd` source. | -| [readme-hero.svg](./readme-hero.svg) | README.md (root) | Animated hero card (tagline, live provider/free-access headline, full-width compression bar demo, 6 stat chips). Edit the SVG directly — there is no `.mmd` source. | +| [free-tier-budget.svg](./free-tier-budget.svg) | README.md (root) | Animated free-tier budget card (~1.51B/mo quantified headline, 20-pool budget bar, per-pool grid, signup credits, 10s loop). Edit the SVG directly — there is no `.mmd` source. | +| [readme-hero.svg](./readme-hero.svg) | README.md (root) | Animated hero card (tagline, live provider/free-access headline, full-width compression bar demo, 6 stat chips). Edit the SVG directly — there is no `.mmd` source. | | [promise-pillars.svg](./promise-pillars.svg) | README.md (root) | Animated "The Promise" 6-pillar card (12s border-highlight sweep). Edit the SVG directly — there is no `.mmd` source. | | [why-pain-fix.svg](./why-pain-fix.svg) | README.md (root) | Animated "Why OmniRoute" 10-row pain-vs-fix ledger (15s green row sweep). Edit the SVG directly — there is no `.mmd` source. | -| [strategies-grid.svg](./strategies-grid.svg) | README.md (root) | Animated grid illustrating 18 of the 19 routing strategies; `cache-optimized` remains documented in the adjacent table. Edit the SVG directly — there is no `.mmd` source. | +| [strategies-grid.svg](./strategies-grid.svg) | README.md (root) | Animated grid illustrating 18 of the 19 routing strategies; `cache-optimized` remains documented in the adjacent table. Edit the SVG directly — there is no `.mmd` source. | | [privacy-local.svg](./privacy-local.svg) | README.md (root) | Animated "Private & Local-First" 11-row guarantee ledger with receipt chips (16s green row sweep). Edit the SVG directly — there is no `.mmd` source. | | [resilience-layers.svg](./resilience-layers.svg) | README.md (root) | Animated 3-layer resilience card (breaker states CLOSED→OPEN→HALF-OPEN, key cooldown with ×2 backoff, model lockout — 18s loops). Edit the SVG directly — there is no `.mmd` source. | diff --git a/docs/diagrams/auto-combo-12factor.mmd b/docs/diagrams/auto-combo-12factor.mmd index 3c7f967534..a5e54711ee 100644 --- a/docs/diagrams/auto-combo-12factor.mmd +++ b/docs/diagrams/auto-combo-12factor.mmd @@ -1,24 +1,28 @@ -%% Auto-Combo 13-factor scoring +%% Auto-Combo 15-factor scoring %% Reflects: open-sse/services/autoCombo/scoring.ts (DEFAULT_WEIGHTS, sum = 1.0) -%% v3.8.49 +%% v3.8.50 +%% svg-title: OmniRoute Auto-Combo 15-factor scoring +%% svg-description: Flow from an incoming request through eligible candidates, the 15 weighted scoring factors, descending score sort, top-N selection, and sequential dispatch. flowchart TB Request["Incoming request"] --> Candidates["Eligible candidates
(provider × model × account)"] Candidates --> Score["Compute composite score
per candidate"] - subgraph Factors["13-factor scoring weights (sum = 1.0)"] - f1["health (0.20)"] - f2["quota (0.15)"] - f3["costInv (0.15)"] - f4["latencyInv (0.12)"] - f5["taskFit (0.08)"] - f6["stability (0.05)"] - f7["tierPriority (0.05)"] - f8["tierAffinity (0.05)"] - f9["specificityMatch (0.05)"] - f10["contextAffinity (0.05)"] - f11["connectionDensity (0.05)"] - f12["cacheAffinity (0.00)"] - f13["resetWindowAffinity (0.00)"] + subgraph Factors["15-factor scoring weights (sum = 1.0)"] + f1["quota (0.1429)"] + f2["health (0.1605)"] + f3["costInv (0.1429)"] + f4["latencyInv (0.1143)"] + f5["taskFit (0.0762)"] + f6["stability (0.0476)"] + f7["tierPriority (0.0476)"] + f8["tierAffinity (0.0476)"] + f9["specificityMatch (0.0476)"] + f10["contextAffinity (0.0476)"] + f11["cacheAffinity (0.0000)"] + f12["sessionAvailability (0.0476)"] + f13["resetWindowAffinity (0.0000)"] + f14["connectionDensity (0.0476)"] + f15["quality (0.0300)"] end Score --> Factors diff --git a/docs/diagrams/cli-terminal.svg b/docs/diagrams/cli-terminal.svg index 3a8d056e5c..1a2f50bd9e 100644 --- a/docs/diagrams/cli-terminal.svg +++ b/docs/diagrams/cli-terminal.svg @@ -1,12 +1,12 @@ - + Compact animated terminal cycling three real OmniRoute CLI commands with a typewriter effect and a scrolling subcommand ticker; the first frame shows the completed providers-list screen. -omniroute — 80+ commands -omniroute providers listOmniRoute Providers1f3a9c2e  anthropic   Claude Max 20x    active8c2d5b1a  codex       Codex Pro (team)  activef4e0a97b  glm         GLM Coding Plan   active03bd6e5f  kimi        Kimi K2 free      active… 334 more providers +omniroute — 85 top-level commands +omniroute providers listOmniRoute Providers1f3a9c2e  anthropic   Claude Max 20x    active8c2d5b1a  codex       Codex Pro (team)  activef4e0a97b  glm         GLM Coding Plan   active03bd6e5f  kimi        Kimi K2 free      active… 346 more providers $ omniroute providers list @@ -14,7 +14,7 @@ -OmniRoute Providers1f3a9c2e  anthropic   Claude Max 20x    active8c2d5b1a  codex       Codex Pro (team)  activef4e0a97b  glm         GLM Coding Plan   active03bd6e5f  kimi        Kimi K2 free      active… 334 more providers +OmniRoute Providers1f3a9c2e  anthropic   Claude Max 20x    active8c2d5b1a  codex       Codex Pro (team)  activef4e0a97b  glm         GLM Coding Plan   active03bd6e5f  kimi        Kimi K2 free      active… 346 more providers $ @@ -32,11 +32,11 @@ -OmniRoute Health  Status: healthy   Uptime: 4d 12h 33m  Requests (24h): 18,412   p95: 412ms  Breakers: ● 24 closed  ◒ 1 half-open  ○ 0 open  Providers: 338 registered   90+ free tiers… live: /dashboard · omniroute status +OmniRoute Health  Status: healthy   Uptime: 4d 12h 33m  Requests (24h): 18,412   p95: 412ms  Breakers: ● 24 closed  ◒ 1 half-open  ○ 0 open  Providers: 350 registered   90+ free tiers… live: /dashboard · omniroute status providers · oauth · keys · combo · nodes · models · cache · compression · cost · usage · quota · health · resilience · telemetry · logs · audit · mcp · a2a · cloud · memory · skills · eval · doctor · repl · tunnel · backup · sync · webhooks · policy · pricing · translator · simulate …providers · oauth · keys · combo · nodes · models · cache · compression · cost · usage · quota · health · resilience · telemetry · logs · audit · mcp · a2a · cloud · memory · skills · eval · doctor · repl · tunnel · backup · sync · webhooks · policy · pricing · translator · simulate … - \ No newline at end of file + diff --git a/docs/diagrams/comparison-table.svg b/docs/diagrams/comparison-table.svg index 24018c7fed..fbbef865f5 100644 --- a/docs/diagrams/comparison-table.svg +++ b/docs/diagrams/comparison-table.svg @@ -23,7 +23,7 @@ Providers - 338 + 350 40+ 400+* ~5 @@ -57,7 +57,7 @@ Built-in MCP server (own tools) - 109 + 110 diff --git a/docs/diagrams/exported/auto-combo-12factor.svg b/docs/diagrams/exported/auto-combo-12factor.svg index d0f2e00d7c..7f165626c3 100644 --- a/docs/diagrams/exported/auto-combo-12factor.svg +++ b/docs/diagrams/exported/auto-combo-12factor.svg @@ -1 +1 @@ -

13-factor scoring weights (sum = 1.0)

health (0.20)

quota (0.15)

costInv (0.15)

latencyInv (0.12)

taskFit (0.08)

stability (0.05)

tierPriority (0.05)

tierAffinity (0.05)

specificityMatch (0.05)

contextAffinity (0.05)

connectionDensity (0.05)

cacheAffinity (0.00)

resetWindowAffinity (0.00)

Incoming request

Eligible candidates
(provider × model × account)

Compute composite score
per candidate

Sort by score
(desc)

Pick top-N targets

Dispatch sequentially
(short-circuit on success)

\ No newline at end of file +OmniRoute Auto-Combo 15-factor scoringFlow from an incoming request through eligible candidates, the 15 weighted scoring factors, descending score sort, top-N selection, and sequential dispatch.

15-factor scoring weights (sum = 1.0)

quota (0.1429)

health (0.1605)

costInv (0.1429)

latencyInv (0.1143)

taskFit (0.0762)

stability (0.0476)

tierPriority (0.0476)

tierAffinity (0.0476)

specificityMatch (0.0476)

contextAffinity (0.0476)

cacheAffinity (0.0000)

sessionAvailability (0.0476)

resetWindowAffinity (0.0000)

connectionDensity (0.0476)

quality (0.0300)

Incoming request

Eligible candidates
(provider × model × account)

Compute composite score
per candidate

Sort by score
(desc)

Pick top-N targets

Dispatch sequentially
(short-circuit on success)

\ No newline at end of file diff --git a/docs/diagrams/free-tier-budget.svg b/docs/diagrams/free-tier-budget.svg index 393ee00594..b96da3272d 100644 --- a/docs/diagrams/free-tier-budget.svg +++ b/docs/diagrams/free-tier-budget.svg @@ -1,4 +1,5 @@ - + + Pool-deduplicated chart of the 20 recurring free-token pools with positive published budgets, plus signup credits and uncapped providers shown separately. @@ -63,7 +64,7 @@ ~1.51B FREE TOKENS / MONTH · STEADY up to ~2.13B in your first month — signup credits - documented free tiers · 40 provider pools · 495 models · one endpoint + documented free tiers · 40 recurring pools · 455 catalog entries · one endpoint @@ -79,59 +80,61 @@ counted once ✓ 15 providers ToS-flagged — we flag it · you decide - - WHERE IT COMES FROM · 19 COUNTABLE FREE POOLS + + WHERE IT COMES FROM · 20 QUANTIFIED RECURRING POOLS - - - - - - - - - - - - - - - - - - - + + + + + + + + + + + + + + + + + + + + - each segment = one free pool · widths floored so every provider shows · honest numbers below + each segment = one recurring pool · widths floored so every pool shows · audited pool budgets below - + - Mistral Large 3 1.00B - GPT-4o mini 150M - Gemini 2.5 Flash 60M - GLM 4.7 30M - Llama 3.3 70B 30M - Grok-3 24M - DeepSeek V4 Pro 20M - GPT-4.1 18M - Llama 4 Scout 15M - GPT-4o 7M - MiniMax-M2.7 6M - Arcee Trinity Large Prev 5M - Auto Free 4M - Auto 1M - Command A Reasoning 800K - ERNIE 4.5 VL 424B 500K - morph-v3-large 400K - Llama 3.1 8B 200K - Claude Sonnet 4.5 25K + Mistral 1.00B + LLM7 150M + Nara 150M + Gemini 60M + Cerebras 30M + Cloudflare AI 30M + API Airforce 24M + Ollama Cloud 20M + Groq 15M + Bluesminds 7.2M + SambaNova 6M + Arcee 4.8M + Navy 4.5M + BazaarLink 3.6M + OpenRouter 1.2M + Cohere 800K + HuggingChat 500K + Morph 400K + Hugging Face 200K + Kiro 25K diff --git a/docs/diagrams/promise-pillars.svg b/docs/diagrams/promise-pillars.svg index f0d30f74a3..232065d928 100644 --- a/docs/diagrams/promise-pillars.svg +++ b/docs/diagrams/promise-pillars.svg @@ -1,4 +1,4 @@ - + Animated promise card: six pillar tiles fade in in reading order, then a soft colored border highlight sweeps from tile to tile in a continuous cycle. @@ -40,7 +40,7 @@ Never hit limits Auto-fallback across 350 providers in milliseconds. Quota out? The next provider - takes over — zero downtime. + takes over while a healthy target remains.
@@ -91,7 +91,7 @@
Every tool works - 33 coding agents — Claude Code, Codex, + 35 CLI/agent integrations — Claude Code, Codex, Cursor, Cline, Copilot, Antigravity — through one config. @@ -127,7 +127,7 @@ Production-grade Circuit breakers, TLS stealth, MCP (110 tools), A2A, memory, guardrails, evals — - 25,000+ tests. + 39,000+ static test declarations. diff --git a/docs/diagrams/readme-hero.svg b/docs/diagrams/readme-hero.svg index e3faa34758..52e4d168db 100644 --- a/docs/diagrams/readme-hero.svg +++ b/docs/diagrams/readme-hero.svg @@ -66,7 +66,7 @@ - 338 + 350 AI PROVIDERS 90+ diff --git a/docs/diagrams/resilience-layers.svg b/docs/diagrams/resilience-layers.svg index 022e35f365..369277e0fb 100644 --- a/docs/diagrams/resilience-layers.svg +++ b/docs/diagrams/resilience-layers.svg @@ -17,6 +17,6 @@ The right layer for the right failure — never kill more than what actually broke. PROVIDERCONNECTION / KEYMODEL - LAYER 1 · SCOPE: WHOLE PROVIDERProvider circuit breakerisolate a provider failing upstream —reroute now, auto-probe to recovertrips only on 408 · 500 · 502 · 503 · 504threshold — oauth 3× · api-key 5× · local 2×reset — 60s · 30s · 15s → HALF-OPEN probelazy recovery — reads refresh expired staterouterprovider Afails ×15provider B ← nextCLOSEDOPENHALF-OPENLAYER 2 · SCOPE: ONE KEY / ACCOUNTConnection cooldownskip one rate-limited key while theother keys keep serving the providerbase cooldown — oauth 5s · api-key 3srepeat fails — backoff ×2 (anti-herd guard)429 honors Retry-After / reset headerssuccess → clearAccountError() resets allprovider · 3 keyskey-1429key-2key-3cooling ×2ⁿLAYER 3 · SCOPE: ONE MODELModel lockoutquarantine a single model — never killthe whole connection for one 429scope — provider + connection + modelper-model 429 · local 404 · mode denialslocked model ≠ dead keyother models keep serving instantlykey-1model-amodel-bmodel-c + LAYER 1 · SCOPE: WHOLE PROVIDERProvider circuit breakerisolate a provider failing upstream —reroute now, auto-probe to recovertrips only on 408 · 500 · 502 · 503 · 504threshold — oauth 10× · api-key 15× · local 2×reset — 60s · 30s · 15s → HALF-OPEN probelazy recovery — reads refresh expired staterouterprovider Afails ×15provider B ← nextCLOSEDOPENHALF-OPENLAYER 2 · SCOPE: ONE KEY / ACCOUNTConnection cooldownskip one rate-limited key while theother keys keep serving the providerbase cooldown — oauth 5s · api-key 3srepeat fails — backoff ×2 (anti-herd guard)429 honors Retry-After / reset headerssuccess → clearAccountError() resets allprovider · 3 keyskey-1429key-2key-3cooling ×2ⁿLAYER 3 · SCOPE: ONE MODELModel lockoutquarantine a single model — never killthe whole connection for one 429scope — provider + connection + modelper-model 429 · local 404 · mode denialslocked model ≠ dead keyother models keep serving instantlykey-1model-amodel-bmodel-c which failure trips what → 5xx / 408 : breaker · key 429 / 401 : cooldown · one-model 429 / 404 : lockout · banned / expired / credits : terminal (operator) \ No newline at end of file diff --git a/docs/diagrams/strategies-grid.svg b/docs/diagrams/strategies-grid.svg index d518f85d06..1706f36e7e 100644 --- a/docs/diagrams/strategies-grid.svg +++ b/docs/diagrams/strategies-grid.svg @@ -95,7 +95,7 @@ auto 72916455 -live 13-factor scoring +live 15-factor scoring fusion diff --git a/docs/getting-started/FREE-TIERS-GUIDE.md b/docs/getting-started/FREE-TIERS-GUIDE.md index 6fd9dcc35b..400343affe 100644 --- a/docs/getting-started/FREE-TIERS-GUIDE.md +++ b/docs/getting-started/FREE-TIERS-GUIDE.md @@ -1,6 +1,6 @@ # Free Tiers Guide: Understand and Combine Free AI Access -> **TL;DR**: OmniRoute registers 329 providers, with **155 catalog entries marked free/no-auth**. The stricter audited budget currently covers **43 recurring pools / 522 model budget entries**. Connect several suitable providers for broader fallback capacity; every quota, approval rule, privacy policy, and paid-overage condition still applies. +> **TL;DR**: OmniRoute registers 350 provider IDs, with **154 provider-catalog entries marked `hasFree`**. The stricter audited free-model catalog covers **40 recurring pool keys / 455 entries** (448 active + 7 discontinued). Connect several suitable providers for broader fallback capacity; every quota, approval rule, privacy policy, and paid-overage condition still applies. --- @@ -21,38 +21,38 @@ OmniRoute **aggregates** these free tiers into one endpoint. Instead of signing These providers have a recurring, keyless, or uncapped free-access path in the audited catalog. “Uncapped” means no published token cap; rate, concurrency, account, regional, and policy limits can still apply: -| Provider | Models | Quota | How to Connect | -|----------|--------|-------|----------------| -| **Kiro AI** | Claude Sonnet 4.5, Haiku 4.5, DeepSeek V3.2, and others | Audited catalog estimates a 25K-token shared monthly pool | OAuth/account flow; ToS flagged `avoid` in the catalog | -| **OpenCode Free** | Current `*-free` model set in the provider registry | Keyless; no published token cap | No provider credential; ToS flagged `avoid` | -| **Pollinations** | Current keyless model set; some former models are discontinued or key-required | Keyless; no published token cap | No provider credential for the keyless models | -| **Logfare** | kimi-k3, deepseek-v4-pro, glm-5.2, gpt-5.6-luna, minimax-m3, and more | Free API key (no rate limits, no card); **every request is logged** for research (opt out at logfare.ai/consent) | Instant key at logfare.ai/register; ToS/privacy at logfare.ai/tos and logfare.ai/privacy | -| **Cloudflare AI** | Workers AI catalog | Audited pool estimates ~30M tokens/month from published usage units | Cloudflare account and API credentials | -| **Gemini** | Gemini Flash family | Audited pool estimates ~60M tokens/month | Google AI Studio API key; rate limits apply | -| **Groq** | Llama, GPT-OSS, and Qwen models | Audited pool estimates ~15M tokens/month | Groq API key; rate limits apply | -| **Cerebras** | GLM 4.7 and GPT-OSS 120B | Audited pool estimates ~30M tokens/month | Cerebras API key; rate limits apply | +| Provider | Models | Quota | How to Connect | +| ----------------- | ------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | +| **Kiro AI** | Claude Sonnet 4.5, Haiku 4.5, DeepSeek V3.2, and others | Audited catalog estimates a 25K-token shared monthly pool | OAuth/account flow; ToS flagged `avoid` in the catalog | +| **OpenCode Free** | Current `*-free` model set in the provider registry | Keyless; no published token cap | No provider credential; ToS flagged `avoid` | +| **Pollinations** | Current keyless model set; some former models are discontinued or key-required | Keyless; no published token cap | No provider credential for the keyless models | +| **Logfare** | kimi-k3, deepseek-v4-pro, glm-5.2, gpt-5.6-luna, minimax-m3, and more | Free API key (no rate limits, no card); **every request is logged** for research (opt out at logfare.ai/consent) | Instant key at logfare.ai/register; ToS/privacy at logfare.ai/tos and logfare.ai/privacy | +| **Cloudflare AI** | Workers AI catalog | Audited pool estimates ~30M tokens/month from published usage units | Cloudflare account and API credentials | +| **Gemini** | Gemini Flash family | Audited pool estimates ~60M tokens/month | Google AI Studio API key; rate limits apply | +| **Groq** | Llama, GPT-OSS, and Qwen models | Audited pool estimates ~15M tokens/month | Groq API key; rate limits apply | +| **Cerebras** | GLM 4.7 and GPT-OSS 120B | Audited pool estimates ~30M tokens/month | Cerebras API key; rate limits apply | ### Signup Grants and Provider-Specific Credits These providers give you **free credits** when you sign up: -| Provider | Free Credits | Models | How to Get | -|----------|-------------|--------|------------| -| **DeepSeek** | 5M free tokens | DeepSeek V4 | Sign up at platform.deepseek.com | -| **LongCat** | 10M-token one-time grant | LongCat 2.0 | API key + KYC; pay-as-you-go after the grant | -| **Together** | $25 signup credit represented as ~25M tokens in the budget model | Provider catalog | Sign up and verify current terms | +| Provider | Free Credits | Models | How to Get | +| ------------- | ------------------------------------------------------------------ | ------------------------- | --------------------------------------------------------- | +| **DeepSeek** | 5M free tokens | DeepSeek V4 | Sign up at platform.deepseek.com | +| **LongCat** | 10M-token one-time grant | LongCat 2.0 | API key + KYC; pay-as-you-go after the grant | +| **Together** | $25 signup credit represented as ~25M tokens in the budget model | Provider catalog | Sign up and verify current terms | | **Vertex AI** | $300 signup credit represented as ~300M tokens in the budget model | Gemini and partner models | Google Cloud account; billing and eligibility rules apply | ### Other Limited Access These providers have **free tiers** with specific limits: -| Provider | Free Limit | Models | Best For | -|----------|-----------|--------|----------| -| **GitHub Models** | Audited shared pool estimates ~18M tokens/month | Broad model evaluation | -| **Hugging Face** | Small recurring monthly pool | Experiments and model variety | -| **OpenRouter free models** | Shared request-limited pool; optional one-time top-up increases the recurring allowance | Broad fallback catalog | -| **AI Horde** | Keyless community capacity; availability varies | Opportunistic distributed inference | +| Provider | Free Limit | Models | Best For | +| -------------------------- | --------------------------------------------------------------------------------------- | ----------------------------------- | -------- | +| **GitHub Models** | Audited shared pool estimates ~18M tokens/month | Broad model evaluation | +| **Hugging Face** | Small recurring monthly pool | Experiments and model variety | +| **OpenRouter free models** | Shared request-limited pool; optional one-time top-up increases the recurring allowance | Broad fallback catalog | +| **AI Horde** | Keyless community capacity; availability varies | Opportunistic distributed inference | --- @@ -70,6 +70,7 @@ Connect several providers to reduce dependence on any single quota: 4. **LongCat** — one-time signup grant (requires KYC) Then use `model: "auto"` and OmniRoute will: + - Try the highest-ranked eligible connection first - If its quota or health check fails → try the next configured provider - If the keyless provider is unavailable → continue through the remaining targets @@ -135,6 +136,7 @@ If one free provider is busy or down, OmniRoute automatically tries the next one ### 2. Smart Routing OmniRoute picks the **best free provider** for each request based on: + - Speed — Which provider is fastest right now? - Quality — Which provider is best for this task? - Capacity — Which provider has quota remaining? @@ -157,13 +159,13 @@ provider's quota or access policy. The live, pool-deduplicated catalog currently reports: -| Metric | Current audited value | Interpretation | -| --- | ---: | --- | -| Recurring quantified grant | **~1.53B tokens/month** | Shared pools counted once; excludes uncapped providers from the sum | -| First month with signup grants | **~2.15B tokens** | Recurring total plus one-time and recurring credits | -| Quantified inventory | **43 pools / 522 model budget entries** | Budget-model coverage, not the full 329-provider catalog | -| Recurring/keyless/uncapped providers represented | **58** | Provider presence in recurring forms of the audited budget catalog | -| Free/no-auth discovery entries | **155** | Broader provider metadata; not all have a quantifiable recurring quota | +| Metric | Current audited value | Interpretation | +| ---------------------------------------------------- | -----------------------------------------------: | ----------------------------------------------------------------------------------------- | +| Recurring quantified grant | **~1.51B tokens/month** | Shared pools counted once; excludes uncapped providers from the sum | +| First month with signup grants | **~2.13B tokens** | Recurring total plus one-time and recurring credits | +| Audited free-model inventory | **40 recurring pool keys / 455 catalog entries** | 448 active + 7 discontinued; distinct from the 350-provider catalog | +| Recurring/keyless free-forever providers represented | **56** | Unique providers across recurring daily/monthly/credit/uncapped and keyless catalog types | +| Provider catalog entries marked `hasFree` | **154 / 350** | Broader provider metadata; not all have a quantifiable recurring quota | These values are computed from `open-sse/config/freeModelCatalog.ts`; see the [Free Tiers Reference](../reference/FREE_TIERS.md) for pool deduplication, ToS flags, diff --git a/docs/routing/AUTO-COMBO.md b/docs/routing/AUTO-COMBO.md index bad9aa43dd..eacd1a700b 100644 --- a/docs/routing/AUTO-COMBO.md +++ b/docs/routing/AUTO-COMBO.md @@ -183,30 +183,31 @@ See [#7992](https://github.com/diegosouzapw/OmniRoute/issues/7992) and [#7111](h ## How It Works (Persisted Auto-Combos) -The Auto-Combo Engine dynamically selects the best provider/model for each request using a **14-factor scoring function** (defined in `open-sse/services/autoCombo/scoring.ts` → `DEFAULT_WEIGHTS`). Weights form a normalized distribution (custom weights are renormalized by `normalizeScoringWeights()`). +The Auto-Combo Engine dynamically selects the best provider/model for each request using a **15-factor scoring function** (defined in `open-sse/services/autoCombo/scoring.ts` → `DEFAULT_WEIGHTS`). The default weights sum to `1.0`; custom weights are renormalized by `normalizeScoringWeights()`. -![Auto-Combo 14-factor scoring](../diagrams/exported/auto-combo-12factor.svg) +![Auto-Combo 15-factor scoring](../diagrams/exported/auto-combo-12factor.svg) -> Source: [diagrams/auto-combo-12factor.mmd](../diagrams/auto-combo-12factor.mmd) (regenerate via `npm run docs:render-diagrams`). The filename predates the current factor set; the diagram shows 13 of the 14 factors (missing `sessionAvailability`). +> Source: [diagrams/auto-combo-12factor.mmd](../diagrams/auto-combo-12factor.mmd) (regenerate via `npm run docs:render-diagrams`). The filename is historical; the source and rendered diagram show all 15 factors declared in `DEFAULT_WEIGHTS`. | Factor | Default Weight | Description | | :-------------------- | :------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| `health` | 0.20 | Health score from circuit breaker (CLOSED=1.0, HALF_OPEN=0.5, OPEN=0.0) | -| `quota` | 0.15 | Remaining quota / rate-limit headroom [0..1] | -| `costInv` | 0.15 | Inverse **blended** cost (60% input + 40% output token price, normalized) — cheaper = higher score | -| `latencyInv` | 0.12 | Inverse p95 latency normalized to pool — faster = higher score | -| `taskFit` | 0.08 | Task-type fitness (coding, review, planning, analysis, debugging, docs) | -| `stability` | 0.05 | Variance-based stability (low latency stdDev / error rate) | -| `tierPriority` | 0.05 | Account-tier priority — Ultra=1.0, Pro=0.67, Standard=0.33, Free=0.0 | -| `tierAffinity` | 0.05 | Affinity between the candidate's tier and the manifest-recommended tier | -| `specificityMatch` | 0.05 | Match between request specificity (manifest hint) and model tier | -| `contextAffinity` | 0.05 | Affinity between the request's context-window need and the model's context window | -| `sessionAvailability` | 0.05 | OAuth session availability of the candidate connection for this session (`getOAuthSessionAvailability()`; non-OAuth connections score 1.0) | -| `connectionDensity` | 0.05 | Spreads load across connections of the same provider (anti-concentration) | +| `quota` | 0.1429 | Remaining quota / rate-limit headroom [0..1] | +| `health` | 0.1605 | Health score from circuit breaker (CLOSED=1.0, HALF_OPEN=0.5, OPEN=0.0) | +| `costInv` | 0.1429 | Inverse **blended** cost (60% input + 40% output token price, normalized) — cheaper = higher score | +| `latencyInv` | 0.1143 | Inverse p95 latency normalized to pool — faster = higher score | +| `taskFit` | 0.0762 | Task-type fitness (coding, review, planning, analysis, debugging, docs) | +| `stability` | 0.0476 | Variance-based stability (low latency stdDev / error rate) | +| `tierPriority` | 0.0476 | Account-tier priority — Ultra=1.0, Pro=0.67, Standard=0.33, Free=0.0 | +| `tierAffinity` | 0.0476 | Affinity between the candidate's tier and the manifest-recommended tier | +| `specificityMatch` | 0.0476 | Match between request specificity (manifest hint) and model tier | +| `contextAffinity` | 0.0476 | Affinity between the request's context-window need and the model's context window | +| `sessionAvailability` | 0.0476 | OAuth session availability of the candidate connection for this session (`getOAuthSessionAvailability()`; non-OAuth connections score 1.0) | +| `connectionDensity` | 0.0476 | Spreads load across connections of the same provider (anti-concentration) | | `cacheAffinity` | 0.00 | Rendezvous-hash affinity toward the connection likeliest to already hold this request's prompt-cache prefix (`open-sse/services/combo/promptCacheAffinity.ts`); disabled by default (#8008) | | `resetWindowAffinity` | 0.00 | Bias toward connections whose quota reset window is favorable (disabled by default) | +| `quality` | 0.03 | Feedback-driven output-quality signal from the routing-event quality tracker; candidates without observations receive a neutral 0.5 | -**Sum:** `0.20 + 0.15 + 0.15 + 0.12 + 0.08 + 0.05 + 0.05 + 0.05 + 0.05 + 0.05 + 0.05 + 0.05 + 0.00 + 0.00 = 1.05` as literally declared in `DEFAULT_WEIGHTS`; user-configured weights are renormalized into a distribution by `normalizeScoringWeights()` before scoring. +**Sum:** `0.1429 + 0.1605 + 0.1429 + 0.1143 + 0.0762 + (7 × 0.0476) + 0.00 + 0.00 + 0.03 = 1.0` as declared in `DEFAULT_WEIGHTS`; user-configured weights are renormalized into a distribution by `normalizeScoringWeights()` before scoring. ## Mode Packs @@ -677,8 +678,8 @@ Including the bare `auto` (default) plus the 6 `AutoVariant` values declared in ## How tiers fit Auto-Combo -The 14-factor scoring function (`open-sse/services/autoCombo/scoring.ts`) treats tier -membership as two signals: `tierPriority` (0.05) and `tierAffinity` (0.05). See the +The 15-factor scoring function (`open-sse/services/autoCombo/scoring.ts`) treats tier +membership as two signals: `tierPriority` (0.0476) and `tierAffinity` (0.0476). See the canonical [scoring factor table](#how-it-works-persisted-auto-combos) above for the full `DEFAULT_WEIGHTS` set — the per-pack overrides (ship-fast/cost-saver/quality-first/ offline-friendly) are listed in the "Weight profiles per pack" table. diff --git a/docs/screenshots/free-tier-budget-card.svg b/docs/screenshots/free-tier-budget-card.svg index 4a861867d6..42bf89d20c 100644 --- a/docs/screenshots/free-tier-budget-card.svg +++ b/docs/screenshots/free-tier-budget-card.svg @@ -1,77 +1,80 @@ - + Static dashboard preview of recurring token pools, first-month signup grants, and uncapped but rate-limited free-access providers. OmniRoute · /dashboard/free-tiers · preview mockup Monthly free-token budget -43 provider pools · 522 model entries · one endpoint +40 recurring pools · 455 catalog entries · one endpoint Steady / month -~1.53B +~1.51B First month (+ signup credits) -~2.15B +~2.13B ToS-flagged (you decide) 15 providers - - - - - - - - - - - - - - - - - - - + + + + + + + + + + + + + + + + + + + + -Each segment = one of 19 quantified recurring pools · 43 total pools / 522 entries in the audited catalog. +Each segment = one of 20 quantified recurring pools · 40 pools / 455 entries in the audited catalog. -Mistral Large 3 1.00B +Mistral 1.00B -GPT-4o mini 150M +LLM7 150M -Gemini 2.5 Flash 60M +Nara 150M -GLM 4.7 30M +Gemini 60M -Llama 3.3 70B 30M +Cerebras 30M -Grok-3 24M +Cloudflare AI 30M -DeepSeek V4 Pro 20M +API Airforce 24M -GPT-4.1 18M +Ollama Cloud 20M -Llama 4 Scout 15M +Groq 15M -GPT-4o 7M +Bluesminds 7.2M -MiniMax-M2.7 6M +SambaNova 6M -Arcee Trinity Large Prev 5M +Arcee 4.8M -Auto Free 4M +Navy 4.5M -Auto 1M +BazaarLink 3.6M -Command A Reasoning 800K +OpenRouter 1.2M -ERNIE 4.5 VL 424B 500K +Cohere 800K -morph-v3-large 400K +HuggingChat 500K -Llama 3.1 8B 200K +Morph 400K -Claude Sonnet 4.5 25K +Hugging Face 200K + +Kiro 25K + First month: one-time signup credits (~626M) @@ -98,5 +101,5 @@ nscale 5M Pool-deduped, honest counting — no inflated rate-limit ceilings. Some terms suggest personal-use only; we flag them so you decide. -+ 13 recurring uncapped* providers (rate/concurrency-limited) · OpenRouter $10 → +24M/mo. ++ 14 recurring uncapped* providers (rate/concurrency-limited) · OpenRouter $10 → +24M/mo. diff --git a/scripts/docs/render-diagrams.mjs b/scripts/docs/render-diagrams.mjs index 01cd7c7377..148013adb8 100644 --- a/scripts/docs/render-diagrams.mjs +++ b/scripts/docs/render-diagrams.mjs @@ -18,11 +18,13 @@ * gate on it. */ import { spawnSync } from "node:child_process"; -import { existsSync, mkdirSync, readdirSync, writeFileSync } from "node:fs"; +import { existsSync, mkdirSync, readFileSync, readdirSync, writeFileSync } from "node:fs"; import { dirname, join, resolve } from "node:path"; import { fileURLToPath } from "node:url"; import { tmpdir } from "node:os"; +import { ensureSvgAccessibility, validateSvgFile } from "./validate-svg.mjs"; + const __dirname = dirname(fileURLToPath(import.meta.url)); const repoRoot = resolve(__dirname, "..", ".."); const srcDir = resolve(repoRoot, "docs", "diagrams"); @@ -75,6 +77,31 @@ for (const src of sources) { if (result.status !== 0) { console.error(` [FAIL] ${src} (exit ${result.status})`); failures += 1; + continue; + } + + const source = readFileSync(input, "utf8"); + const title = source.match(/^%%\s*svg-title:\s*(.+)$/im)?.[1]?.trim(); + const description = source.match(/^%%\s*svg-description:\s*(.+)$/im)?.[1]?.trim(); + if (title && description) { + const svg = readFileSync(output, "utf8"); + writeFileSync( + output, + ensureSvgAccessibility(svg, { + title, + description, + idBase: src.replace(/\.mmd$/, ""), + }) + ); + } else if (title || description) { + console.warn(` [WARN] ${src}: svg-title and svg-description must be provided together`); + } + + const validation = validateSvgFile(output); + for (const warning of validation.warnings) console.warn(` [WARN] ${src}: ${warning}`); + if (validation.errors.length > 0) { + for (const error of validation.errors) console.error(` [FAIL] ${src}: ${error}`); + failures += 1; } } diff --git a/scripts/docs/validate-svg.mjs b/scripts/docs/validate-svg.mjs new file mode 100644 index 0000000000..b3c65b07f6 --- /dev/null +++ b/scripts/docs/validate-svg.mjs @@ -0,0 +1,167 @@ +#!/usr/bin/env node + +import { readFileSync, writeFileSync } from "node:fs"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +import { XMLParser, XMLValidator } from "fast-xml-parser"; + +const parser = new XMLParser({ + ignoreAttributes: false, + attributeNamePrefix: "@_", + preserveOrder: true, +}); + +function collectIds(value, ids) { + if (Array.isArray(value)) { + for (const entry of value) collectIds(entry, ids); + return; + } + if (!value || typeof value !== "object") return; + + const attributes = value[":@"]; + if (attributes && typeof attributes === "object" && typeof attributes["@_id"] === "string") { + ids.push(attributes["@_id"]); + } + for (const entry of Object.values(value)) collectIds(entry, ids); +} + +function escapeXml(value) { + return value + .replaceAll("&", "&") + .replaceAll("<", "<") + .replaceAll(">", ">") + .replaceAll('"', """) + .replaceAll("'", "'"); +} + +function replaceRootAttribute(openingTag, name, value) { + const attribute = new RegExp(`\\s${name}=(?:"[^"]*"|'[^']*')`, "i"); + const withoutExisting = openingTag.replace(attribute, ""); + return withoutExisting.replace(/>$/, ` ${name}="${escapeXml(value)}">`); +} + +function escapeRegExp(value) { + return value.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"); +} + +export function ensureSvgAccessibility(svg, { title, description, idBase }) { + const xmlResult = XMLValidator.validate(svg); + if (xmlResult !== true) throw new Error(`invalid XML: ${xmlResult.err.msg}`); + + const titleId = `${idBase}-title`; + const descriptionId = `${idBase}-desc`; + const priorTitle = new RegExp( + `]*\\bid=["']${escapeRegExp(titleId)}["'][^>]*>[\\s\\S]*?<\\/title>`, + "i" + ); + const priorDescription = new RegExp( + `]*\\bid=["']${escapeRegExp(descriptionId)}["'][^>]*>[\\s\\S]*?<\\/desc>`, + "i" + ); + const withoutPriorAccessibleName = svg.replace(priorTitle, "").replace(priorDescription, ""); + const match = withoutPriorAccessibleName.match(/]*>/i); + if (!match) throw new Error("document root is not an SVG element"); + + let openingTag = replaceRootAttribute(match[0], "role", "img"); + openingTag = replaceRootAttribute(openingTag, "aria-labelledby", `${titleId} ${descriptionId}`); + const accessibleName = + `${escapeXml(title)}` + + `${escapeXml(description)}`; + + return withoutPriorAccessibleName.replace(match[0], `${openingTag}${accessibleName}`); +} + +export function validateSvgText(svg) { + const xmlResult = XMLValidator.validate(svg); + if (xmlResult !== true) { + return { errors: [`invalid XML: ${xmlResult.err.msg}`], warnings: [] }; + } + + const document = parser.parse(svg); + const ids = []; + collectIds(document, ids); + const duplicates = [...new Set(ids.filter((id, index) => ids.indexOf(id) !== index))].sort(); + + const openingTag = svg.match(/]*>/i)?.[0] ?? ""; + const warnings = []; + if (!/\srole=["']img["']/i.test(openingTag)) warnings.push('root role is not "img"'); + const hasAccessibleName = + /\saria-(?:label|labelledby)=["'][^"']+["']/i.test(openingTag) || + /]*>[^<]+<\/title>/i.test(svg); + if (!hasAccessibleName) { + warnings.push("missing accessible name (title, aria-label, or aria-labelledby)"); + } + if (!/]*>[^<]+<\/desc>/i.test(svg)) warnings.push("missing desc element"); + if (/ 0 ? [`duplicate IDs: ${duplicates.join(", ")}`] : [], + warnings, + }; +} + +export function validateSvgFile(file) { + return validateSvgText(readFileSync(file, "utf8")); +} + +function isDirectExecution() { + if (!process.argv[1]) return false; + return fileURLToPath(import.meta.url) === path.resolve(process.argv[1]); +} + +if (isDirectExecution()) { + const args = process.argv.slice(2); + let fixAccessibility = false; + let title; + let description; + const files = []; + for (let index = 0; index < args.length; index += 1) { + const arg = args[index]; + if (arg === "--fix-a11y") { + fixAccessibility = true; + } else if (arg === "--title") { + title = args[++index]; + } else if (arg === "--description") { + description = args[++index]; + } else { + files.push(arg); + } + } + if (files.length === 0) { + console.error( + "Usage: node scripts/docs/validate-svg.mjs [--fix-a11y --title TEXT --description TEXT] [...]" + ); + process.exit(2); + } + + if (fixAccessibility && (!title || !description)) { + console.error("--fix-a11y requires both --title and --description"); + process.exit(2); + } + + let failures = 0; + for (const file of files) { + if (fixAccessibility) { + const idBase = path.basename(file, path.extname(file)); + const updated = ensureSvgAccessibility(readFileSync(file, "utf8"), { + title, + description, + idBase, + }); + writeFileSync(file, updated); + } + const result = validateSvgFile(file); + for (const warning of result.warnings) console.warn(`WARN ${file}: ${warning}`); + if (result.errors.length === 0) { + console.log(`PASS ${file}`); + continue; + } + failures += 1; + for (const error of result.errors) console.error(`FAIL ${file}: ${error}`); + } + if (failures > 0) process.exit(1); +} diff --git a/tests/unit/docs-validate-svg.test.ts b/tests/unit/docs-validate-svg.test.ts new file mode 100644 index 0000000000..a53ff75afd --- /dev/null +++ b/tests/unit/docs-validate-svg.test.ts @@ -0,0 +1,106 @@ +import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; +import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import test from "node:test"; +import { fileURLToPath } from "node:url"; + +const here = path.dirname(fileURLToPath(import.meta.url)); +const validator = path.resolve(here, "../../scripts/docs/validate-svg.mjs"); + +test("SVG validator ignores Mermaid data-id attributes when checking duplicate IDs", () => { + const fixtureDir = mkdtempSync(path.join(tmpdir(), "omniroute-svg-validator-")); + const fixture = path.join(fixtureDir, "mermaid.svg"); + writeFileSync( + fixture, + '' + + "Fixture diagram." + + '' + + '' + + "" + ); + + try { + const result = spawnSync(process.execPath, [validator, fixture], { encoding: "utf8" }); + assert.equal(result.status, 0, `${result.stdout}${result.stderr}`); + assert.match(result.stdout, /PASS/); + assert.doesNotMatch(`${result.stdout}${result.stderr}`, /WARN/); + } finally { + rmSync(fixtureDir, { recursive: true, force: true }); + } +}); + +test("SVG validator rejects duplicate XML id attributes", () => { + const fixtureDir = mkdtempSync(path.join(tmpdir(), "omniroute-svg-validator-")); + const fixture = path.join(fixtureDir, "duplicate.svg"); + writeFileSync( + fixture, + '' + + '' + + "" + ); + + try { + const result = spawnSync(process.execPath, [validator, fixture], { encoding: "utf8" }); + assert.equal(result.status, 1, `${result.stdout}${result.stderr}`); + assert.match(result.stderr, /duplicate IDs: edge-a/); + } finally { + rmSync(fixtureDir, { recursive: true, force: true }); + } +}); + +test("SVG validator adds explicit accessible naming when requested for a generated diagram", () => { + const fixtureDir = mkdtempSync(path.join(tmpdir(), "omniroute-svg-validator-")); + const fixture = path.join(fixtureDir, "auto-combo.svg"); + writeFileSync( + fixture, + '' + ); + + try { + const result = spawnSync( + process.execPath, + [ + validator, + "--fix-a11y", + "--title", + "Auto-Combo scoring", + "--description", + "How OmniRoute scores eligible routing targets with 15 factors.", + fixture, + ], + { encoding: "utf8" } + ); + assert.equal(result.status, 0, `${result.stdout}${result.stderr}`); + + const repeated = spawnSync( + process.execPath, + [ + validator, + "--fix-a11y", + "--title", + "Auto-Combo scoring", + "--description", + "How OmniRoute scores eligible routing targets with 15 factors.", + fixture, + ], + { encoding: "utf8" } + ); + assert.equal(repeated.status, 0, `${repeated.stdout}${repeated.stderr}`); + + const updated = readFileSync(fixture, "utf8"); + assert.match(updated, /role="img"/); + assert.match(updated, /aria-labelledby="auto-combo-title auto-combo-desc"/); + assert.match(updated, /Auto-Combo scoring<\/title>/); + assert.match( + updated, + /How OmniRoute scores eligible routing targets with 15 factors\.<\/desc>/ + ); + assert.equal([...updated.matchAll(/id="auto-combo-title"/g)].length, 1); + assert.equal([...updated.matchAll(/id="auto-combo-desc"/g)].length, 1); + } finally { + rmSync(fixtureDir, { recursive: true, force: true }); + } +});