+
+> Combinare manualmente i tier gratuiti è scomodo: decine di SDK, decine di rate limit e nessuna idea chiara di quanta capacità sia davvero disponibile. OmniRoute aggrega i tier gratuiti **documentati** di **42 pool di provider / 495 modelli** in un unico numero trasparente e lo mostra in tempo reale nella dashboard (`/dashboard/free-tiers`).
+
+
+
+> Riepilogo animato della pagina live `/dashboard/free-tiers`. Metodologia completa (deduplicazione dei pool, tier di credito, condizioni dei provider): **[docs/reference/FREE_TIERS.md](../../reference/FREE_TIERS.md)**.
+>
+> Questi valori vengono ricontrollati ogni due settimane rispetto al catalogo live e **possono sia salire sia scendere**: se un provider termina un tier gratuito, il numero diminuisce; se ne arriva uno nuovo, aumenta. Pubblichiamo ciò che il catalogo calcola realmente, mai una stima ottimistica arrotondata verso l'alto.
+
+
+
+
+
+
+
+⭐ Metti una stella alla repo se OMNIROUTE ti ha aiutato a risparmiare e a lavorare meglio.
+
+
+
+[](https://github.com/diegosouzapw/OmniRoute)
+
+[](https://www.star-history.com/diegosouzapw/omniroute)
+[](https://olud.ai/project/diegosouzapw-omniroute.html)
+
+### 💬 Unisciti alla community
+
+**👋 Segui il maintainer — scopri per primo nuovi provider, release e suggerimenti:**
+
+[](https://www.linkedin.com/in/diegosouzapw/)
+[](https://github.com/diegosouzapw)
+
+[](https://discord.gg/U47eFqAXCn)
+[](https://t.me/omnirouteOficial)
+[](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t)
+[](https://chat.whatsapp.com/LTSpdFhXTxjH4R6CCNiKWz)
+[](https://omniroute.online)
+
+**Domande, suggerimenti sui provider, roadmap e supporto → [Discord](https://discord.gg/U47eFqAXCn) · [Telegram](https://t.me/omnirouteOficial) · WhatsApp [🌍 Global](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) / [🇧🇷 Brasil](https://chat.whatsapp.com/LTSpdFhXTxjH4R6CCNiKWz)**
+
+
+
+## 📈 Il gateway continua a crescere
+
+
-
-
----
-
-### 🤖 Free AI Provider for your favorite coding agents
-
-_Connect any AI-powered IDE or CLI tool through OmniRoute — free-access AI gateway; provider limits and terms apply._
-
-
-
-📡 All agents connect via http://localhost:20128/v1 or http://cloud.omniroute.online/v1 — one config; model access and quotas depend on providers
-
----
-
-## 🤔 Why OmniRoute?
-
-**Stop wasting money and hitting limits:**
-
-- Subscription quota expires unused every month
-- Rate limits stop you mid-coding
-- Expensive APIs ($20-50/month per provider)
-- Manual switching between providers
-
-**OmniRoute solves this:**
-
-- ✅ **Maximize subscriptions** - Track quota, use every bit before reset
-- ✅ **Auto fallback** - Subscription → API Key → Cheap → Free; availability depends on eligible upstream routes
-- ✅ **Multi-account** - Round-robin between accounts per provider
-
----
-
-## 📧 Support
-
-> 💬 **Join our community!** [WhatsApp Group](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) — Get help, share tips, and stay updated.
-
-- **Website**: [omniroute.online](https://omniroute.online)
-- **GitHub**: [github.com/diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute)
-- **Issues**: [github.com/diegosouzapw/OmniRoute/issues](https://github.com/diegosouzapw/OmniRoute/issues)
-- **WhatsApp**: [Community Group](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t)
-- **Contributing**: See [CONTRIBUTING.md](CONTRIBUTING.md), open a PR, or pick a `good first issue`
-
-### 🐛 Reporting a Bug?
-
-When opening an issue, please run the system-info command and attach the generated file:
+
```bash
-npm run system-info
+# Fresh install, zero credentials — `auto` already works:
+curl http://localhost:20128/v1/chat/completions \
+ -H "Content-Type: application/json" \
+ -d '{"model":"auto","messages":[{"role":"user","content":"Hello!"}]}'
```
-This generates a `system-info.txt` with your Node.js version, OmniRoute version, OS details, installed CLI tools (qoder, gemini, claude, codex, antigravity, droid, etc.), Docker/PM2 status, and system packages — everything we need to reproduce your issue quickly. Attach the file directly to your GitHub issue.
+Preferisci uno specifico backend gratuito? Chiamalo direttamente, ad esempio `oc/…` (OpenCode Free) o `felo/…` (Felo). Poi passa a `auto` e lascia che sia OmniRoute a scegliere.
----
+📦 Script di avvio rapido pronti da copiare per **Python, Node.js, PHP e cURL** → [`examples/quickstart/`](../../../examples/quickstart/)
-## 🔄 How It Works
+
+
+
+
+> **Vuoi diventare un Open Source Friend?** Queste sono le aziende che sostengono l'open source e aiutano OmniRoute a continuare a crescere — e dichiariamo pubblicamente dove viene usato ogni token che ci forniscono. Contatto: [diegosouza.pw@outlook.com](mailto:diegosouza.pw@outlook.com)
+
+
+ Grazie a Kimi (Moonshot AI), il nostro Open Source Friend fondatore, per il sostegno al progetto! Kimi è il laboratorio AI dietro le famiglie di modelli open-weight K2 e K3 — Kimi K3 offre una finestra di contesto da 1M token, vision nativa e capacità di coding di frontiera a una frazione del prezzo dei modelli chiusi, e funziona subito con Claude Code, Codex e ogni strumento di coding supportato da OmniRoute.
+
+ Cosa rende possibile il supporto di Kimi: i crediti API di Kimi alimentano la pipeline di release validata dall'AI di OmniRoute — la fase merge validation powered by Kimi K3 che esamina ogni pull request prima del rilascio — oltre allo sviluppo quotidiano delle funzionalità. Il supporto Kimi di prima classe è disponibile su entrambi i canali: la Kimi API diretta (kimi-k3) e il piano di coding Kimi Code (OAuth e API key). OmniRoute è anche il primo progetto open source brasiliano nel programma di supporto di Kimi. Ottieni una Kimi API key con il 15% di crediti extra →
+
+ Grazie a Cheaper Inference, un Open Source Friend di OmniRoute, per il sostegno al progetto! Cheaper Inference è un gateway ordinato per costo che rivende 42 modelli di frontiera — Claude, GPT-5.x, Gemini, Kimi K3, GLM, DeepSeek, Grok e MiniMax — dietro un unico endpoint compatibile con OpenAI, instradando ogni richiesta verso il provider idoneo più economico senza mai addebitare più del prezzo di listino del produttore del modello.
+
+ Supporto di prima classe in OmniRoute: Chat Completions, endpoint nativo /v1/responses, vision, tool calling e 3 modelli immagine (grok-imagine, nano-banana-pro, nano-banana-2, raggiungibili come cheaperinference/<model>). Ottieni una API key →
+
+
+
+
+I link contrassegnati con aff=omniroute sono link partner. Finanziano il progetto senza costi aggiuntivi per te.
+
+
+
+
+🎟️ Promo affiliati — coupon gratuiti di registrazione da provider che non sponsorizziamo (clicca per espandere)
+
+Questa sezione contiene soltanto codici referral/coupon. Le partnership sponsorizzate sono riportate sopra in 🤝 Supportato dai nostri amici dell'Open Source. OmniRoute non ha sponsorizzazioni o partnership con i provider elencati qui: sono coupon pubblici utilizzabili da chiunque.
+
+
+ AgentRouter — registrazione affiliata · $100 di crediti gratuiti alla registrazione (server gratuito, aspettati una latenza maggiore — ideale per test, non per produzione). Supporto di prima classe in OmniRoute dalla v3.8.50: Chat Completions, formato wire compatibile con Anthropic e percorso compatibile con OpenAI. I modelli disponibili includono claude-opus-4-8, claude-opus-5, gpt-5.6-sol e altri. Ottieni i tuoi $100 →
+
+ ⚠️ Link affiliato — OmniRoute non ha sponsorizzazioni o partnership con questo provider.
+
+
+
+
+Conosci un altro provider con un generoso coupon gratuito di registrazione utile agli utenti OmniRoute? Apri una issue e lo aggiungeremo qui.
+
+
+
+
+
+
+
+
+## 🎯 Combo — La funzionalità di punta
+
+
+
+
+
+> Una **combo** è una catena di modelli tra cui OmniRoute instrada le richieste **automaticamente**. La quota finisce, un provider fallisce o i costi aumentano: la combo passa silenziosamente al modello successivo. **È questo che rende OmniRoute resistente ai guasti.** 🛡️
+
+### ⚡ Zero-config — usa semplicemente `auto`
+
+Non devi creare nessuna combo. Imposta il modello su `auto` (o una sua variante) e OmniRoute costruisce una combo virtuale a partire dai provider collegati, assegnando i punteggi in tempo reale:
+
+
🧑💻 Pesi orientati prima alla qualità per la generazione di codice
+
auto/fast
⚡ Prima la latenza più bassa
+
auto/cheap
💰 Prima il costo per token più basso
+
auto/offline
🔋 Prima il maggiore margine di quota / rate limit
+
auto/smart
🔭 Prima la qualità + 10% di esplorazione per scoprire modelli migliori
+
+
+##
+
+### 🔀 Oppure creane una tua — 19 strategie di routing
+
+Tutte e **19** le strategie — combinabili liberamente per ogni passaggio della combo:
+
+
+
+
#
+
Strategia
+
Cosa fa
+
+
1
priority
Lista ordinata con priorità al primo target — esaurisce ciascuno prima di passare al successivo 🥇
+
2
fill-first
Usa completamente la quota di ogni target prima di passare oltre
+
3
weighted
Scelta casuale pesata in base al peso assegnato a ogni target
+
4
round-robin
Scorre ciclicamente i target in ordine
+
5
p2c
Bilanciamento casuale del carico Power-of-Two-Choices
+
6
least-used
Sceglie il target con il carico corrente più basso
+
7
random
Scelta casuale uniforme (con deduplicazione)
+
8
strict-random
Casuale senza deduplicare le ripetizioni 🎲
+
9
cost-optimized
Riduce al minimo il costo per richiesta usando i prezzi live del catalogo 💸
+
10
headroom
Sceglie il target con la maggiore quota residua
+
11
reset-window
Preferisce il target la cui finestra di quota si resetta prima
+
12
reset-aware
Ordina in base al reset della quota — prima le finestre più brevi 📊
+
13
context-relay
Passa il contesto tra i target nelle conversazioni lunghe 🧠
+
14
context-optimized
Sceglie il target più adatto alla dimensione corrente del contesto
+
15
cache-optimized
Fissa ogni prefisso di prompt riutilizzabile allo stesso account — massimizza gli hit della prompt cache 🎯
+
16
lkgp
Last-Known-Good Path — resta sull'ultimo target che ha risposto correttamente
+
17
auto
Punteggio live su 14 fattori per ogni connessione 🤖
+
18
fusion
Invia la richiesta a un gruppo di modelli + un giudice sintetizza una sola risposta 🧬
+
19
pipeline
Concatena i passaggi — l'output di ogni target alimenta il successivo 🔗
+
+
+Il motore Auto-Combo valuta ogni candidato su **14 fattori** (salute, quota, costo, latenza, tasso di successo, freschezza…) — consulta [`docs/routing/AUTO-COMBO.md`](../../routing/AUTO-COMBO.md).
+
+##
+
+### 🧱 Resilienza integrata (3 livelli indipendenti)
+
+
+
+📖 [Motore Auto-Combo](../../routing/AUTO-COMBO.md) · [Guida alla resilienza](../../architecture/RESILIENCE_GUIDE.md)
+
+
+
+
+
+
+## 🏆 Cosa distingue OmniRoute
+
+
+
+
+
+📊 Metodologia completa e dettaglio per funzionalità rispetto a 9router, OpenRouter, CLIProxyAPI e LiteLLM → [`docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md`](../../comparison/OMNIROUTE_VS_ALTERNATIVES.md)
+
+
+
+
+## 💚 Supporta OmniRoute
+
+OmniRoute è distribuito con licenza MIT e mantenuto apertamente. Se ti fa risparmiare tempo o denaro, ecco come aiutarlo a restare indipendente — scegli ciò che preferisci. Le sponsorizzazioni non influenzano mai la priorità del routing: acquistano visibilità, non posizionamento.
+
+
+
+**🇧🇷 PIX** — istantaneo, senza commissioni (Brasile)
+
+
+
+Chiave (casuale): `5d865059-bc44-483a-962d-43ceb80126eb`
+
+Pix copia-e-cola:
```
-┌─────────────┐
-│ Your CLI │ (Claude Code, Codex, OpenClaw, Cursor, Cline...)
-│ Tool │
-└──────┬──────┘
- │ http://localhost:20128/v1
- ↓
-┌─────────────────────────────────────────┐
-│ OmniRoute (Smart Router) │
-│ • Format translation (OpenAI ↔ Claude) │
-│ • Quota tracking + Embeddings + Images │
-│ • Auto token refresh │
-└──────┬──────────────────────────────────┘
- │
- ├─→ [Tier 1: SUBSCRIPTION] Claude Code, Codex
- │ ↓ quota exhausted
- ├─→ [Tier 2: API KEY] DeepSeek, Groq, xAI, Mistral, NVIDIA NIM, etc.
- │ ↓ budget limit
- ├─→ [Tier 3: CHEAP] GLM ($0.6/1M), MiniMax ($0.2/1M)
- │ ↓ budget limit
- └─→ [Tier 4: FREE] Qoder, Qwen, Kiro (provider limits apply)
-
-Result: broader fallback coverage and cost control; availability is not guaranteed
+00020101021126580014br.gov.bcb.pix01365d865059-bc44-483a-962d-43ceb80126eb5204000053039865802BR5922OMNIROUTE CONTRIBUICAO6006BRASIL62070503***630475DD
```
----
-
-## 🎯 What OmniRoute Solves — 30 Real Pain Points & Use Cases
-
-> **Every developer using AI tools faces these problems daily.** OmniRoute was built to solve them all — from cost overruns to regional blocks, from broken OAuth flows to protocol operations and enterprise observability.
+
-💸 1. "I pay for an expensive subscription but still get interrupted by limits"
+₿ Crypto — BTC · ETH · USDT-TRC20 · USDC-Solana (clicca per espandere)
-Developers pay $20–200/month for Claude Pro, Codex Pro, or GitHub Copilot. Even paying, quota has a ceiling — 5h of usage, weekly limits, or per-minute rate limits. Mid-coding session, the provider stops responding and the developer loses flow and productivity.
+
+
₿ BTC
Bitcoin (SegWit)
bc1qh00smz004sy85wyl28v77tenkt3ckl6eaep7fd
+
Ξ ETH
Ethereum (ERC20)
0x64Cf6B68A6Ff34288e89172950a2d00102337a84
+
₮ USDT
Tron (TRC20)
TKAF41JpuQrHbKTnsQa9svJE2T192Hvsc2
+
$ USDC
Solana
2emNNZzVVWQc3FQ2wk9M6qXUQmW8AKdjjL174fXR28Tu
+
-**How OmniRoute solves it:**
-
-- **Smart 4-Tier Fallback** — If subscription quota runs out, automatically redirects to API Key → Cheap → Free with zero manual intervention
-- **Provider Limits Tracking** — Cached quota snapshots refresh on a server-side schedule (default `PROVIDER_LIMITS_SYNC_INTERVAL_MINUTES=70`) with manual refresh available in the UI
-- **Multi-Account Support** — Multiple accounts per provider with auto round-robin — when one runs out, switches to the next
-- **Custom Combos** — Customizable fallback chains with 13 balancing strategies (priority, weighted, fill-first, round-robin, P2C, random, least-used, cost-optimized, strict-random, auto, lkgp, context-optimized, **context-relay**)
-- **Structured Combo Builder** — Build combos step-by-step with explicit provider + model + account selection, including repeated providers and fixed-account targets
-- **Quota-Aware P2C** — Power-of-two account selection now factors quota headroom, backoff, recent errors, and consecutive use
-- **Codex Business Quotas** — Business/Team workspace quota monitoring directly in the dashboard
+⚠️ Invia ogni moneta esclusivamente sulla rete indicata: inviarla sulla rete sbagliata può causare la perdita dei fondi.
-
-🔌 2. "I need to use multiple providers but each has a different API"
+🐛 Hai trovato un bug o vuoi lasciare un feedback? Apri una [Discussion](https://github.com/diegosouzapw/OmniRoute/discussions).
+
+
+
+
Note per gli sviluppatori: il progetto può generare un file locale .env durante npm install/postinstall per comodità nello sviluppo. Questo file viene intenzionalmente ignorato tramite .gitignore (vedi .gitignore) e non deve mai essere incluso nei commit; se viene committato accidentalmente, ruota ogni secret esposto e rimuovi il file dalla cronologia. Consulta docs/DEVELOPER-ENVIRONMENT.md per le indicazioni sulla gestione dei file di ambiente locali e dei secret.
+
+## 📡 OmniRoute Radar
+
+Il valore principale dei tier gratuiti resta **~1,53 miliardi di token/mese**, calcolato sul catalogo documentato con deduplicazione dei pool riportato sopra. I crediti temporanei di registrazione dei provider possono separatamente portare il primo mese a **~2,15 miliardi**. Radar è un overlay opzionale e firmato del catalogo, pensato per chi vuole informazioni più aggiornate sulla disponibilità dei modelli gratuiti tra una release di OmniRoute e la successiva; il catalogo della community e tutte le funzionalità gratuite esistenti restano gratuiti.
+
+I sostenitori possono ricevere il catalogo live e ulteriori opportunità offerte dai provider. Il relativo tetto separato e variabile è di **circa 3 miliardi di token/mese al massimo**, a seconda della disponibilità dei provider. Questo limite non è una garanzia: i provider possono modificare quote, requisiti, modelli o regioni in qualsiasi momento.
+
+Radar è opt-in e usa soltanto richieste GET. Il client OmniRoute non carica prompt, traffico, configurazione dei provider, telemetria d'uso o lo stato locale di chiusura degli annunci. Dettagli sui requisiti e sul catalogo corrente su **[radar.omniroute.online/planos](https://radar.omniroute.online/planos)**.
+
+
+
+
+
+
+## ✨ Novità
+
+
+
+> Novità principali da **v3.8.20 → v3.8.50**. Cronologia completa in [`CHANGELOG.md`](../../../CHANGELOG.md).
+
+- **🎛️ OmniConductor** — delega A2A in ingresso alla tua flotta di agenti, skill Conductor nell'Agent Card e un pannello dashboard con chat vocale push-to-talk Faro. → [A2A Server](../../frameworks/A2A-SERVER.md)
+- **🛂 Admission adattiva e protezione dal sovraccarico** — le richieste chat pesanti vengono messe in coda invece di ricevere 503, con lease RPM rolling atomici per connessione. → [Guida alla resilienza](../../architecture/RESILIENCE_GUIDE.md)
+- **🗂️ Ordinamento canonico di `/v1/models`** — un blocco contiguo raggruppato per provider per ciascun provider (combo sempre in testa), stabile tra tutte le fonti del catalogo. → [Riferimento API](../../reference/API_REFERENCE.md)
+- **🗜️ Rafforzamento della compressione** — protezione dall'inflazione attiva per impostazione predefinita, pack Caveman per DE / FR / JA + cinese (wényán), filtri RTK per Gradle e .NET. → [Compressione](../../compression/COMPRESSION_ENGINES.md)
+- **💸 Costo flat-rate trasparente** — i provider in abbonamento / coding plan risultano a **$0** nelle analytics dei costi; budget, quote e routing continuano a fare stime. → [Riferimento API](../../reference/API_REFERENCE.md)
+- **⚖️ Routing Quota-Share** — divide equamente la quota di un account condiviso tra chiavi in pool, in modo work-conserving così le porzioni inattive vengono prestate. → [Guida alla resilienza](../../architecture/RESILIENCE_GUIDE.md)
+- **🤖 Configurazione CLI/agente con un comando** — `setup-*` configura oltre 12 strumenti di coding; `omniroute run` avvia 7 CLI (Claude Code, Codex, Aider, Goose, OpenCode, Qwen Code, Gemini CLI) senza scrivere configurazioni; `omniroute configure` è un selettore interattivo provider+modello con preferiti per contesto. → [Integrazioni CLI](../../guides/CLI-INTEGRATIONS.md)
+- **🛰️ Modalità remota** — controlla un OmniRoute remoto con token scoped (`connect` / `contexts` / `tokens`) + helper OAuth `antigravity` per installazioni VPS. → [Modalità remota](../../guides/REMOTE-MODE.md)
+- **🧭 Auto-routing più intelligente** — combo `auto/:`, **Fusion** (gruppo di modelli + giudice), routing task-aware, override per-request di modello / modalità / budget USD. → [Auto-Combo](../../routing/AUTO-COMBO.md)
+- **🗜️ Compressione pluggable** — 12 motori componibili + Compression Studios: LLMLingua-2, Ultra a due livelli, omniglyph, fidelity gate per passaggio, GCF v3.2, editor drag-reorder. → [Compressione](../../compression/COMPRESSION_ENGINES.md)
+- **🕵️ Decrittazione MITM trasparente (TPROXY)** — cattura le CLI che ignorano le variabili d'ambiente del proxy, con CA per-SNI + installer del trust store. → [MITM/TPROXY](../../security/MITM-TPROXY-DECRYPT.md)
+- **💸 Telemetria dei costi ovunque** — header di costo/utilizzo `X-OmniRoute-*` su ogni endpoint, header del risparmio su cache HIT, quote di spesa USD per chiave. → [Riferimento API](../../reference/API_REFERENCE.md)
+- **🧠 Memoria sotto il tuo controllo** — disattivata per impostazione predefinita, quantizzazione vettoriale int8 opt-in + decadimento tipizzato, `x-omniroute-no-memory` per-request. → [Memoria](../../frameworks/MEMORY.md)
+- **🛡️ Sicurezza** — guard contro la prompt injection su ogni route LLM (suite red-team), guardrail opzionale per il masking delle credenziali (oscura API key/secret trapelati in entrambe le direzioni), web search DuckDuckGo gratuita come ultima risorsa e gate di login OIDC opzionale per la dashboard (il login con password resta sempre disponibile). → [Guardrail](../../security/GUARDRAILS.md)
+- **🖼️ Nuovi endpoint** — `/v1/ocr` (Mistral OCR) e `/v1/audio/translations` (stile Whisper) completano la superficie media. → [Riferimento API](../../reference/API_REFERENCE.md)
+- **🎨 Generazione immagini / video / audio** — una sola API per i media: xAI Grok Imagine e Novita AI video, ComfyUI, Freepik, Adobe Firefly, Microsoft Designer, Segmind, EdgeTTS. → [Riferimento API](../../reference/API_REFERENCE.md)
+- **🌍 Deployment e operazioni** — `basePath` del reverse proxy, rilevamento automatico della lingua del browser, tracking dei dispositivi per chiave, trust MITM senza root, localizzazione zh-TW. → [Ambiente](../../reference/ENVIRONMENT.md)
+- **🤝 Più provider e agenti** — Cursor Cloud Agent, Grok Build (xAI) con login browser + OAuth, scheda Ollama di prima classe, Claude Opus 5 e Sonnet 5, partnership ufficiale Kimi (Code/Web/Moonshot), Zed, Requesty, SenseNova, Yuanbao, Agnes AI… e un catalogo aggiornato di **350 provider**. → [Provider](../../reference/PROVIDER_REFERENCE.md)
+- **📡 Trasparenza del routing** — ogni risposta include un header `X-OmniRoute-Decision` con strategia/provider/latenza che l'ha servita; una nuova strategia combo `cache-optimized` + il fattore `cacheAffinity` di Auto-Combo riportano le richieste ripetute alla connessione che possiede il prefisso in cache; un endpoint read-only `/v1/auto-combo/{channel}/candidates` espone il pool di candidati live di un canale `auto/*`. → [Auto-Combo](../../routing/AUTO-COMBO.md)
+- **⚡ Prestazioni e infrastruttura locali** — Redis locale con un clic, deployer relay Cloudflare Workers / Deno Deploy, Bifrost e Mux come servizi embedded supervisionati. → [Servizi embedded](../../frameworks/EMBEDDED-SERVICES.md)
+
+
+
+
+
+
+## 🤖 CLI e agenti di coding compatibili
+
+> Una sola configurazione — `http://localhost:20128/v1` — e **qualsiasi** IDE o CLI AI può usare modelli gratuiti e a basso costo.
+
+
+
+
+
+**Avvia qualsiasi CLI supportata tramite OmniRoute con un solo comando** — senza scrivere file di configurazione,
+con le credenziali iniettate per singolo processo e una home temporanea isolata per Qwen/Gemini:
+
+```bash
+omniroute run claude --model openai/gpt-5.4 # Claude Code
+omniroute run codex --model glm/glm-5.2 # OpenAI Codex CLI
+omniroute run aider --model glm/glm-5.2 -- --message "reply OK"
+omniroute run goose --model glm/glm-5.2
+omniroute run opencode --model glm/glm-5.2 -- run "reply OK"
+omniroute run qwen --model glm/glm-5.2 -- -p "reply OK"
+omniroute run gemini --model glm/glm-5.2 -- --skip-trust -p "reply OK"
+
+# Or pick provider+model interactively and write the tool's own config:
+omniroute configure codex # also: claude opencode qwen aider goose cline continue kilo
+```
+
+Ogni comando rispetta il contesto remoto attivo (`omniroute connect `); `--dry-run`
+mostra in anteprima env/argomenti esatti senza eseguire nulla, mentre `--api-key-env NAME` evita che i segreti
+finiscano nella cronologia della shell. → [Integrazioni CLI](../../guides/CLI-INTEGRATIONS.md)
+
+
+
+
+
+
+## 🌐 349 provider AI — oltre 90 gratuiti
+
+
+
+> Il catalogo più completo tra i router open source: **349 provider**, **oltre 90 con un piano gratuito**, **56 gratuiti per sempre**.
+
+
+
+### 🏢 Tutti i principali laboratori — tramite un solo endpoint
+
+
+
+
OpenAI
+
Anthropic
+
Gemini
+
xAI Grok
+
DeepSeek
+
Mistral
+
+
+
Qwen
+
Meta Llama
+
Groq
+
NVIDIA
+
MiniMax
+
Cohere
+
+
+
Perplexity
+
HuggingFace
+
Together
+
Fireworks
+
Cloudflare
+
Baidu
+
+
+
+…e oltre 220 altri — ogni icona viene risolta in tempo reale dal catalogo provider della dashboard. 📖 [Riferimento provider](../../reference/PROVIDER_REFERENCE.md)
+
+
+
+### 🆓 Gratuiti per sempre — $0, nessuna carta
+
+
+
+
OpenCode Zen DeepSeek V4, Nemotron 3 Nessun limite di token
+
Kilo Code Auto-router, Tencent Hy3 Gratuito per sempre
+
Requesty GPT-OSS 120B, Nemotron Gratuito per sempre
+
SiliconFlow DeepSeek V3.2 / R1 Piano gratuito
+
Z.AI GLM GLM-4.7 / 4.5-Flash Gratuito per sempre
+
Baidu ERNIE ERNIE 4.0 Gratuito per sempre
+
+
+
Qoder AI Qwen3-Max, Kimi-K2 GRATUITO senza limiti
+
Pollinations GPT, Llama, Claude Nessuna chiave necessaria
+
Cloudflare AI 50+ modelli 10K neuroni/giorno
+
NVIDIA NIM GLM, MiniMax ~40 RPM gratuiti
+
Cerebras GLM 4.7, GPT-OSS 1M token/giorno
+
OpenRouter modelli :free +$10 → RPM più elevati
+
+
-OpenAI uses one format, Claude (Anthropic) uses another, Gemini yet another. If a dev wants to test models from different providers or fallback between them, they need to reconfigure SDKs, change endpoints, deal with incompatible formats. Custom providers (FriendLI, NIM) have non-standard model endpoints.
+📖 Catalogo completo leggibile dalle macchine → [`docs/reference/PROVIDER_REFERENCE.md`](../../reference/PROVIDER_REFERENCE.md)
-**How OmniRoute solves it:**
+
+
-- **Unified Endpoint** — A single `http://localhost:20128/v1` serves as proxy for all 329 provider catalog entries
-- **Format Translation** — Automatic and transparent: OpenAI ↔ Claude ↔ Gemini ↔ Responses API
-- **Response Sanitization** — Strips non-standard fields (`x_groq`, `usage_breakdown`, `service_tier`) that break OpenAI SDK v1.83+
-- **Role Normalization** — Converts `developer` → `system` for non-OpenAI providers; `system` → `user` for GLM/ERNIE
-- **Think Tag Extraction** — Extracts `` blocks from models like DeepSeek R1 into standardized `reasoning_content`
-- **Structured Output for Gemini** — `json_schema` → `responseMimeType`/`responseSchema` automatic conversion
-- **`stream` defaults to `false`** — Aligns with OpenAI spec, avoiding unexpected SSE in Python/Rust/Go SDKs
+
+
-
+## 🖥️ Dove gira OmniRoute — ovunque
-
-🌐 3. "My AI provider blocks my region/country"
+
-Providers like OpenAI/Codex block access from certain geographic regions. Users get errors like `unsupported_country_region_territory` during OAuth and API connections. This is especially frustrating for developers from developing countries.
+> La stessa app, sulla tua macchina, secondo le tue regole. Da un'installazione npm globale fino al **tuo telefono** tramite Termux.
-**How OmniRoute solves it:**
+
+
Piattaforma
Installazione
Punti di forza
+
📦 npm (globale)
npm install -g omniroute
Un comando, qualsiasi OS
+
🐳 Docker
docker run … diegosouzapw/omniroute
Multi-arch AMD64 + ARM64
+
🖥️ Desktop (Electron)
npm run electron:build
Finestra nativa + system tray — Windows / macOS / Linux
+
💪 ARM
nativo arm64
Raspberry Pi, server ARM, Apple Silicon
+
📱 Android (Termux)
pkg install nodejs && npx -y omniroute
Gira sul tuo telefono, 24/7, senza root
+
📲 PWA
"Aggiungi alla schermata Home"
Schermo intero, offline, installabile dal browser
+
🧩 Plugin OpenCode
@omniroute/opencode-provider
Integrazione nativa con OpenCode
+
🤖 VS Code Copilot Chat
installa l'estensione OmniCopilot
Tutti i modelli OmniRoute nel selettore nativo di Copilot Chat — Stable e Insiders
+
🛠️ Da sorgente
npm install && npm run dev
Modificalo e contribuisci
+
-- **3-Level Proxy Config** — Configurable proxy at 3 levels: global (all traffic), per-provider (one provider only), and per-connection/key
-- **Color-Coded Proxy Badges** — Visual indicators: 🟢 global proxy, 🟡 provider proxy, 🔵 connection proxy, always showing the IP
-- **OAuth Token Exchange Through Proxy** — OAuth flow also goes through the proxy, solving `unsupported_country_region_territory`
-- **Connection Tests via Proxy** — Connection tests use the configured proxy (no more direct bypass)
-- **SOCKS5 Support** — Full SOCKS5 proxy support for outbound routing
-- **TLS Fingerprint Spoofing** — Browser-like TLS fingerprint via `wreq-js` to bypass bot detection
-- **🔏 CLI Fingerprint Matching** — Reorders headers and body fields to match native CLI binary signatures, drastically reducing account flagging risk. The proxy IP is preserved — you get both stealth **and** IP masking simultaneously
+📖 [Guida Docker](../../guides/DOCKER_GUIDE.md) · [Desktop](../../../electron/README.md) · [Termux](../../guides/TERMUX_GUIDE.md) · [PWA](../../guides/PWA_GUIDE.md) · [OpenCode](../../frameworks/OPENCODE.md)
-
+
-
-🆓 4. "I want to use AI for coding but I have no money"
+
-Not everyone can pay $20–200/month for AI subscriptions. Students, devs from emerging countries, hobbyists, and freelancers need access to quality models at zero cost.
+### 🧩 Novità: OmniRoute dentro il Copilot Chat nativo di VS Code
-**How OmniRoute solves it:**
+
-- **Ollama Cloud** — Cloud-hosted Ollama models at `api.ollama.com` with free "Light usage" tier; use `ollamacloud/` prefix
-- **Free-Only Combos** — Chain `if/kimi-k2-thinking → qw/qwen3-coder-plus` can use currently listed $0 access; limits and availability apply
-- **NVIDIA NIM Free Access** — ~40 RPM free access as currently listed; provider terms and model availability apply at build.nvidia.com (transitioning from credits to pure rate limits)
-- **Cost Optimized Strategy** — Routing strategy that automatically chooses the cheapest available provider
+> Nessuna nuova barra laterale, nessuna nuova UI di chat — ogni modello servito da OmniRoute compare direttamente nel
+> **selettore modelli di Copilot Chat che usi già**. Da VS Code 1.122, i modelli dei provider funzionano
+> senza accesso GitHub né abbonamento Copilot — modalità agent, tool calling e vision, gratuitamente.
-
+Installa l'estensione **[OmniCopilot](https://github.com/diegosouzapw/OmniCopilot)**, collegala
+al tuo server OmniRoute (predefinito `localhost:20128`), poi apri Copilot Chat → selettore modelli
+→ **Manage Models…** → **OmniRoute**.
-
-🔒 5. "I need to protect my AI gateway from unauthorized access"
+
-When exposing an AI gateway to the network (LAN, VPS, Docker), anyone with the address can consume the developer's tokens/quota. Without protection, APIs are vulnerable to misuse, prompt injection, and abuse.
+Dall'editor: apri la vista **Extensions**, cerca **"OmniRoute"**, fai clic su **Install**
+— funziona allo stesso modo su entrambi gli store. Sorgenti, issue e runbook di pubblicazione sono su
+[diegosouzapw/OmniCopilot](https://github.com/diegosouzapw/OmniCopilot).
-**How OmniRoute solves it:**
+📖 [Guida VS Code Copilot Chat](../../guides/VSCODE-COPILOT.md) — configurazione, contenuto del selettore, dashboard in una scheda, risoluzione dei problemi
-- **API Key Management** — Generation, rotation, and scoping per provider with a dedicated `/dashboard/api-manager` page
-- **Model-Level Permissions** — Restrict API keys to specific models (`openai/*`, wildcard patterns), with Allow All/Restrict toggle
-- **API Endpoint Protection** — Require a key for `/v1/models` and block specific providers from the listing
-- **Auth Guard + CSRF Protection** — All dashboard routes protected with `withAuth` middleware + CSRF tokens
-- **Rate Limiter** — Per-IP rate limiting with configurable windows
-- **IP Filtering** — Allowlist/blocklist for access control
-- **Prompt Injection Guard** — Sanitization against malicious prompt patterns
-- **AES-256-GCM Encryption** — Credentials encrypted at rest
+
-
+
+
-
-🛑 6. "My provider went down and I lost my coding flow"
+## 🔒 Privato e local-first
-AI providers can become unstable, return 5xx errors, or hit temporary rate limits. If a dev depends on a single provider, they're interrupted. Without circuit breakers, repeated retries can crash the application.
+
-**How OmniRoute solves it:**
+
-- **Request Queue & Pacing** — Per-connection request buckets smooth bursts before they hit upstream rate caps
-- **Connection Cooldown** — A single connection cools down after retryable failures with optional upstream `Retry-After` hints and exponential backoff
-- **Provider Circuit Breaker** — The provider only trips after fallback is exhausted and the provider request still fails with provider-wide transient errors; connection-scoped `429` rate limits stay in Connection Cooldown
-- **Wait For Cooldown** — The server can wait for the earliest connection cooldown to expire and retry the same client request automatically
-- **Anti-Thundering Herd** — Mutex + semaphore protection against concurrent retry storms
-- **Combo Fallback Chains** — If the primary provider fails, automatically falls through the chain with no intervention
-- **Health Dashboard** — Uptime monitoring, provider circuit breaker states, cooldowns, cache stats, p50/p95/p99 latency
+📖 [Autorizzazione](../../architecture/AUTHZ_GUIDE.md) · [Guardrail](../../security/GUARDRAILS.md) · [Conformità](../../security/COMPLIANCE.md)
-
+
-
-🔧 7. "Configuring each AI tool is tedious and repetitive"
+
+
-**How OmniRoute solves it:**
+## 🔌 CLI completa + A2A e MCP
-- **CLI Tools Dashboard** — Dedicated page with one-click setup for Claude Code, Codex CLI, OpenClaw, Kilo Code, Antigravity, Cline
-- **GitHub Copilot Config Generator** — Generates `chatLanguageModels.json` for VS Code with bulk model selection
-- **Onboarding Wizard** — Guided 4-step setup for first-time users
-- **One endpoint, all models** — Configure `http://localhost:20128/v1` once, access 329 provider catalog entries
+
-
+> Oltre al server, OmniRoute è una **console completa da riga di comando** con **oltre 80 comandi**, più protocolli agent aperti che permettono a un agent AI di gestirlo **autonomamente**.
-
-🔑 8. "Managing OAuth tokens from multiple providers is hell"
+### ⌨️ Una vera CLI (non solo `start`)
-Claude Code, Codex, Copilot — all use OAuth 2.0 with expiring tokens. Developers need to re-authenticate constantly, deal with `client_secret is missing`, `redirect_uri_mismatch`, and failures on remote servers. OAuth on LAN/VPS is particularly problematic.
+```bash
+omniroute # serve gateway + dashboard (port 20128)
+omniroute chat # interactive TUI chat client (slash: /model /combo /skill /memory)
+omniroute setup # guided first-run wizard
+omniroute doctor # diagnose providers, ports, native deps
+```
-**How OmniRoute solves it:**
+### 🛰️ Modalità remota — esegui qui la CLI, OmniRoute su un VPS
-- **Auto Token Refresh** — OAuth tokens refresh in background before expiration
-- **OAuth 2.0 (PKCE) Built-in** — Automatic flow for Claude Code, Codex, Copilot, Kiro, Qwen, Qoder
-- **Multi-Account OAuth** — Multiple accounts per provider via JWT/ID token extraction
-- **OAuth LAN/Remote Fix** — Private IP detection for `redirect_uri` + manual URL mode for remote servers
-- **OAuth Behind Nginx** — Uses `window.location.origin` for reverse proxy compatibility
-- **Remote OAuth Guide** — Step-by-step guide for Google Cloud credentials on VPS/Docker
+OmniRoute gira su un server? Gestiscilo dal laptop con la **stessa CLI**. Accedi una volta
+con un token di accesso con scope; da quel momento ogni comando punta all'istanza remota.
-
+```bash
+omniroute connect 192.168.0.15 # password → scoped token, saved as a context
+omniroute models list # ← runs against the REMOTE server
+omniroute configure codex # ← picks a remote model, writes a local Codex profile
+omniroute tokens create --name ci --scope read # mint narrower tokens for other machines
+omniroute contexts use default # ← switch back to the local server
+```
-
-📊 9. "I don't know how much I'm spending or where"
+I token hanno scope `read` / `write` / `admin`; le route che avviano processi restano limitate al loopback.
+📖 [Modalità remota](../../guides/REMOTE-MODE.md)
-Developers use multiple paid providers but have no unified view of spending. Each provider has its own billing dashboard, but there's no consolidated view. Unexpected costs can pile up.
+
-**How OmniRoute solves it:**
+
-- **Cost Analytics Dashboard** — Per-token cost tracking and budget management per provider
-- **Budget Limits per Tier** — Spending ceiling per tier that triggers automatic fallback
-- **Per-Model Pricing Configuration** — Configurable prices per model
-- **Usage Statistics Per API Key** — Request count and last-used timestamp per key
-- **Analytics Dashboard** — Stat cards, model usage chart, provider table with success rates and latency
+
-
+### 🤝 Collega un agent — e controllerà OmniRoute stesso
-
-🐛 10. "I can't diagnose errors and problems in AI calls"
+Esponi OmniRoute tramite **MCP**, **A2A**, una **REST API**, **webhook** o una **CLI remota** — qualsiasi agent compatibile (o il tuo codice) ottiene accesso al gateway: routing, provider, combo, cache, compressione, memoria — in autonomia. Gli endpoint HTTP qui sotto sono serviti su `http://localhost:20128`.
-When a call fails, the dev doesn't know if it was a rate limit, expired token, wrong format, or provider error. Fragmented logs across different terminals. Without observability, debugging is trial-and-error.
+
+
Interfaccia
Endpoint / comando
A cosa serve
+
🧰 MCP (stdio)
omniroute --mcp
Collegamento a Claude Desktop, Cursor e qualsiasi client MCP
Compatibile con OpenAI — chat, embedding, immagini, audio, OCR
+
🔔 Webhook
/api/webhooks
Invia eventi (utilizzo, quota, errori, routing) al tuo URL
+
🛰️ CLI remota
omniroute connect
Gestisci un'istanza remota con token di accesso con scope
+
-**How OmniRoute solves it:**
+```bash
+# Give Claude Code the full OmniRoute toolset over MCP:
+claude mcp add-server omniroute --type http --url http://localhost:20128/api/mcp/stream
+```
-- **Unified Logs Dashboard** — 4 tabs: Request Logs, Proxy Logs, Audit Logs, Console
-- **Console Log Viewer** — Real-time terminal-style viewer with color-coded levels, auto-scroll, search, filter
-- **SQLite Summary Logs** — Request and proxy log indexes stay queryable across restarts without loading large payload blobs into SQLite
-- **Translator Playground** — 4 debugging modes: Playground (format translation), Chat Tester (round-trip), Test Bench (batch), Live Monitor (real-time)
-- **Request Telemetry** — p50/p95/p99 latency + X-Request-Id tracing
-- **File-Based Detail Artifacts** — App logs rotate by size, retention days, and archive count; detailed request/response payloads live in `DATA_DIR/call_logs/` and rotate independently of SQLite summaries
-- **System Info Report** — `npm run system-info` generates `system-info.txt` with your full environment (Node version, OmniRoute version, OS, CLI tools, Docker/PM2 status). Attach it when reporting issues for instant triage.
+📖 [MCP Server](../../frameworks/MCP-SERVER.md) · [A2A Server](../../frameworks/A2A-SERVER.md) · [Protocolli agent](../../frameworks/AGENT_PROTOCOLS_GUIDE.md)
-
+
-
-🏗️ 11. "Deploying and maintaining the gateway is complex"
+
+
-Installing, configuring, and maintaining an AI proxy across different environments (local, VPS, Docker, cloud) is labor-intensive. Problems like hardcoded paths, `EACCES` on directories, port conflicts, and cross-platform builds add friction.
+## 🗜️ Risparmia il 15–95% dei token — automaticamente
-**How OmniRoute solves it:**
+
-- **npm global install** — `npm install -g omniroute && omniroute` — done
-- **Docker Multi-Platform** — AMD64 + ARM64 native (Apple Silicon, AWS Graviton, Raspberry Pi)
-- **Docker Compose Profiles** — `base` (no CLI tools) and `cli` (with Claude Code, Codex, OpenClaw)
-- **Electron Desktop App** — Native app for Windows/macOS/Linux with system tray, auto-start, offline mode
-- **Split-Port Mode** — API and Dashboard on separate ports for advanced scenarios (reverse proxy, container networking)
-- **Cloud Sync** — Config synchronization across devices via Cloudflare Workers
-- **DB Backups** — Automatic backup, restore, export and import of all settings, with `DISABLE_SQLITE_AUTO_BACKUP` for externally managed backups
+### 📖 Come funziona — pipeline, architettura e calcolo del risparmio
-
+
-
-🌍 12. "The interface is English-only and my team doesn't speak English"
-
-Teams in non-English-speaking countries, especially in Latin America, Asia, and Europe, struggle with English-only interfaces. Language barriers reduce adoption and increase configuration errors.
-
-**How OmniRoute solves it:**
-
-- **Dashboard i18n — 30 Languages** — All 500+ keys translated including Arabic, Bulgarian, Danish, German, Spanish, Finnish, French, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Malay, Dutch, Norwegian, Polish, Portuguese (PT/BR), Romanian, Russian, Slovak, Swedish, Thai, Ukrainian, Vietnamese, Chinese, Filipino, English
-- **RTL Support** — Right-to-left support for Arabic and Hebrew
-- **Multi-Language READMEs** — 30 complete documentation translations
-- **Language Selector** — Globe icon in header for real-time switching
-
-
-
-
-🔄 13. "I need more than chat — I need embeddings, images, audio"
-
-AI isn't just chat completion. Devs need to generate images, transcribe audio, create embeddings for RAG, rerank documents, and moderate content. Each API has a different endpoint and format.
-
-**How OmniRoute solves it:**
-
-- **Embeddings** — `/v1/embeddings` with 6 providers and 9+ models
-- **Image Generation** — `/v1/images/generations` with 10 providers and 20+ models (OpenAI, xAI, Together, Fireworks, Nebius, Hyperbolic, NanoBanana, Antigravity, SD WebUI, ComfyUI)
-- **Text-to-Video** — `/v1/videos/generations` — ComfyUI (AnimateDiff, SVD) and SD WebUI
-- **Text-to-Music** — `/v1/music/generations` — ComfyUI (Stable Audio Open, MusicGen)
-- **Audio Transcription** — `/v1/audio/transcriptions` — Whisper + Nvidia NIM, HuggingFace, Qwen3
-- **Text-to-Speech** — `/v1/audio/speech` — ElevenLabs, Nvidia NIM, HuggingFace, Coqui, Tortoise, Qwen3, **Inworld**, **Cartesia**, **PlayHT**, + existing providers
-- **Moderations** — `/v1/moderations` — Content safety checks
-- **Reranking** — `/v1/rerank` — Document relevance reranking
-- **Responses API** — Full `/v1/responses` support for Codex
-
-
-
-
-🧪 14. "I have no way to test and compare quality across models"
-
-Developers want to know which model is best for their use case — code, translation, reasoning — but comparing manually is slow. No integrated eval tools exist.
-
-**How OmniRoute solves it:**
-
-- **LLM Evaluations** — Golden set testing with 10 pre-loaded cases covering greetings, math, geography, code generation, JSON compliance, translation, markdown, safety refusal
-- **4 Match Strategies** — `exact`, `contains`, `regex`, `custom` (JS function)
-- **Translator Playground Test Bench** — Batch testing with multiple inputs and expected outputs, cross-provider comparison
-- **Chat Tester** — Full round-trip with visual response rendering
-- **Live Monitor** — Real-time stream of all requests flowing through the proxy
-
-
-
-
-📈 15. "I need to scale without losing performance"
-
-As request volume grows, without caching the same questions generate duplicate costs. Without idempotency, duplicate requests waste processing. Per-provider rate limits must be respected.
-
-**How OmniRoute solves it:**
-
-- **Semantic Cache** — Two-tier cache (signature + semantic) reduces cost and latency
-- **Request Idempotency** — 5s deduplication window for identical requests
-- **Rate Limit Detection** — Per-provider RPM, min gap, and max concurrent tracking
-- **Request Queue & Pacing** — Configurable queue, pacing, and concurrency defaults in Settings → Resilience
-- **API Key Validation Cache** — 3-tier cache for production performance
-- **Health Dashboard with Telemetry** — p50/p95/p99 latency, cache stats, uptime
-
-
-
-
-🤖 16. "I want to control model behavior globally"
-
-Developers who want all responses in a specific language, with a specific tone, or want to limit reasoning tokens. Configuring this in every tool/request is impractical.
-
-**How OmniRoute solves it:**
-
-- **System Prompt Injection** — Global prompt applied to all requests
-- **Thinking Budget Validation** — Reasoning token allocation control per request (passthrough, auto, custom, adaptive)
-- **9 Routing Strategies** — Global strategies that determine how requests are distributed
-- **Wildcard Router** — `provider/*` patterns route dynamically to any provider
-- **Combo Enable/Disable Toggle** — Toggle combos directly from the dashboard
-- **Manual Combo Ordering** — Drag combo cards by handle and persist the order in SQLite
-- **Provider Toggle** — Enable/disable all connections for a provider with one click
-- **Blocked Providers** — Exclude specific providers from `/v1/models` listing
-
-
-
-
-🧰 17. "I need MCP tools as first-class product capabilities"
-
-Many AI gateways expose MCP only as a hidden implementation detail. Teams need a visible, manageable operation layer.
-
-**How OmniRoute solves it:**
-
-- MCP appears in the dashboard navigation and endpoint protocol tab
-- Dedicated MCP management page with process, tools, scopes, and audit
-- Built-in quick-start for `omniroute --mcp` and client onboarding
-
-
-
-
-🧠 18. "I need A2A orchestration with sync + stream task paths"
-
-Agent workflows need both direct replies and long-running streamed execution with lifecycle control.
-
-**How OmniRoute solves it:**
-
-- A2A JSON-RPC endpoint (`POST /a2a`) with `message/send` and `message/stream`
-- SSE streaming with terminal state propagation
-- Task lifecycle APIs for `tasks/get` and `tasks/cancel`
-
-
-
-
-🛰️ 19. "I need real MCP process health, not guessed status"
-
-Operational teams need to know if MCP is actually alive, not just whether an API is reachable.
-
-**How OmniRoute solves it:**
-
-- Runtime heartbeat file with PID, timestamps, transport, tool count, and scope mode
-- MCP status API combining heartbeat + recent activity
-- UI status cards for process/uptime/heartbeat freshness
-
-
-
-
-📋 20. "I need auditable MCP tool execution"
-
-When tools mutate config or trigger ops actions, teams need forensic traceability.
-
-**How OmniRoute solves it:**
-
-- SQLite-backed audit logging for MCP tool calls
-- Filters by tool, success/failure, API key, and pagination
-- Dashboard audit table + stats endpoints for automation
-
-
-
-
-🔐 21. "I need scoped MCP permissions per integration"
-
-Different clients should have least-privilege access to tool categories.
-
-**How OmniRoute solves it:**
-
-- 32 granular MCP scopes for controlled tool access
-- Scope enforcement and visibility in MCP management UI
-- Safe default posture for operational tooling
-
-
-
-
-⚙️ 22. "I need operational controls without redeploying"
-
-Teams need quick runtime changes during incidents or cost events.
-
-**How OmniRoute solves it:**
-
-- Switch combo activation directly from MCP dashboard
-- Tune queue, cooldown, breaker, and wait settings from the dedicated Resilience page
-- Review live provider breaker state from the Health dashboard
-
-
-
-
-🔄 23. "I need live A2A task lifecycle visibility and cancellation"
-
-Without lifecycle visibility, task incidents become hard to triage.
-
-**How OmniRoute solves it:**
-
-- Task listing/filtering by state/skill with pagination
-- Drill-down on task metadata, events, and artifacts
-- Task cancellation endpoint and UI action with confirmation
-
-
-
-
-🌊 24. "I need active stream metrics for A2A load"
-
-Streaming workflows require operational insight into concurrency and live connections.
-
-**How OmniRoute solves it:**
-
-- Active stream counters integrated into A2A status
-- Last task timestamp and per-state counts
-- A2A dashboard cards for real-time ops monitoring
-
-
-
-
-🪪 25. "I need standard agent discovery for clients"
-
-External clients and orchestrators need machine-readable metadata for onboarding.
-
-**How OmniRoute solves it:**
-
-- Agent Card exposed at `/.well-known/agent.json`
-- Capabilities and skills shown in management UI
-- A2A status API includes discovery metadata for automation
-
-
-
-
-🧭 26. "I need protocol discoverability in the product UX"
-
-If users cannot discover protocol surfaces, adoption and support quality drop.
-
-**How OmniRoute solves it:**
-
-- Consolidated **Endpoints** page with tabs for Proxy, MCP, A2A, and API Endpoints
-- Inline service status toggles (Online/Offline) for MCP and A2A
-- Links from overview to dedicated management tabs
-
-
-
-
-🧪 27. "I need end-to-end protocol validation with real clients"
-
-Mock tests are not enough to validate protocol compatibility before release.
-
-**How OmniRoute solves it:**
-
-- E2E suite that boots app and uses real MCP SDK client transport
-- A2A client tests for discovery, send, stream, get, and cancel flows
-- Cross-check assertions against MCP audit and A2A tasks APIs
-
-
-
-
-📡 28. "I need unified observability across all interfaces"
-
-Splitting observability by protocol creates blind spots and longer MTTR.
-
-**How OmniRoute solves it:**
-
-- Unified dashboards/logs/analytics in one product
-- Health + audit + request telemetry across OpenAI, MCP, and A2A layers
-- Operational APIs for status and automation
-
-
-
-
-💼 29. "I need one runtime for proxy + tools + agent orchestration"
-
-Running many separate services increases operational cost and failure modes.
-
-**How OmniRoute solves it:**
-
-- OpenAI-compatible proxy, MCP server, and A2A server in one stack
-- Shared auth, resilience, data store, and observability
-- Consistent policy model across all interaction surfaces
-
-
-
-
-🚀 30. "I need to ship agentic workflows without glue-code sprawl"
-
-Teams lose velocity when stitching multiple ad-hoc services and scripts.
-
-**How OmniRoute solves it:**
-
-- Unified endpoint strategy for clients and agents
-- Built-in protocol management UIs and smoke validation paths
-- Production-ready foundations (security, logging, resilience, backup)
-
-
-
-
-📚 31. "My long sessions crash with 'context_length_exceeded' limits"
-
-During deep debugging, long histories with tool results quickly exceed provider token windows, causing failed requests and orphaned context.
-
-**How OmniRoute solves it:**
-
-- **Proactive Context Compression** — Evaluates token budgets before the request hits upstream and proactively prunes old conversation history with a smart binary-search mechanism.
-- **Structural Integrity Guards** — Automatically tracks explicit `tool_use` definitions and ensures that if a tool input is truncated, its corresponding `tool_result` is also safely removed, preventing API validation errors.
-- **Multi-Layer Dropping** — Progressively drops system messages, regular messages, and finally enforces strict length limits without breaking conversational logic.
-
-
-
-### Example Playbooks (Integrated Use Cases)
-
-**Playbook A: Maximize paid subscription + cheap backup**
+La combinazione in cascata predefinita esegue `RTK → Caveman`. Quando entrambi intervengono sullo stesso payload di tool/contesto, i risparmi si compongono:
```txt
-Combo: "maximize-claude"
- 1. cc/claude-opus-4-7
- 2. glm/glm-4.7
- 3. if/kimi-k2-thinking
-
-Monthly cost: $20 + small backup spend
-Outcome: higher quality, near-zero interruption
+combined = 1 − (1 − RTK) × (1 − Caveman_input)
+average = 1 − (1 − 0.80) × (1 − 0.46) = 89.2%
+range = 78.4 – 94.6%
```
-**Playbook B: Zero-cost coding stack**
+Blocchi di codice, URL, JSON e dati strutturati sono **sempre protetti** dal motore di preservazione.
-```txt
-Combo: "free-access"
- 1. if/kimi-k2-thinking (no published token cap; limits apply)
- 2. qw/qwen3-coder-plus (no published token cap; limits apply)
+> **Perché usare molti token quando ne bastano pochi?** Ogni richiesta attraversa la pipeline di compressione di OmniRoute **in modo trasparente** — senza modifiche al client. Ora è una **stack di 12 motori componibili** eseguiti in ordine e combinabili per ciascun routing combo — basati anche su idee di [RTK](https://github.com/rtk-ai/rtk), [Caveman](https://github.com/JuliusBrussee/caveman) (⭐ 90K+), [LLMLingua-2](https://github.com/microsoft/LLMLingua) e [Troglodita](https://github.com/leninejunior/troglodita) (PT-BR).
-Monthly cost: $0
-Outcome: broader free-access fallback; upstream availability is not guaranteed
-```
+### 🧱 La stack di 12 motori
-**Playbook C: 24/7 always-on fallback chain**
+I motori vengono eseguiti nell'ordine della pipeline; ciascuno può essere attivato/disattivato e configurato indipendentemente per combo:
-```txt
-Combo: "multi-layer-fallback"
- 1. cc/claude-opus-4-7
- 2. cx/gpt-5.2-codex
- 3. glm/glm-4.7
- 4. minimax/MiniMax-M2.1
- 5. if/kimi-k2-thinking
+
+
#
Motore
Cosa fa
+
1
Session-Dedup
Elimina contenuti ripetuti tra i turni (content-addressed, cross-turn)
+
2
CCR
Archivia blocchi grandi dietro marker di recupero, caricati su richiesta
+
3
Lite
Riduzione di spazi e URL immagine (baseline a bassa latenza)
+
4
RTK
Filtro intelligente dei risultati dei tool, deduplica e troncamento (consapevole del comando)
Compattazione tabellare lossless di array JSON (~30%) tramite codec GCF incluso nel progetto
+
7
Relevance
Valutazione estrattiva delle frasi rispetto all'ultima richiesta dell'utente
+
8
Caveman
Compressione della prosa basata su regole (~65–75% sull'output)
+
9
Aggressive
Riepilogo + invecchiamento progressivo dei turni precedenti
+
10
LLMLingua-2
Pruning semantico ML tramite MobileBERT ONNX — code-safe, asincrono
+
11
Ultra
Pruning euristico dei token con livello opzionale basato su piccolo modello (SLM)
+
12
OmniGlyph
Codifica sperimentale del contesto come immagine per Claude Fable 5 misurato sul protocollo Anthropic diretto; i transformer GPT 5.6 restano fail-closed in attesa di ricevute del provider. Quattro profili di compressione (aggressive predefinito, balanced, coding-safe, passthrough) (il più aggressivo; opt-in)
+
-Outcome: deep fallback depth for deadline-critical workloads
-```
+Blocchi di codice, URL e dati strutturati sono **sempre preservati** byte per byte. I **preset con un clic** combinano i motori:
-**Playbook D: Agent ops with MCP + A2A**
+
+
Modalità
Risparmio
Ideale per
+
🪶 Lite
~15%
Impostazione predefinita sicura sempre attiva
+
🪨 Standard (Caveman)
~30%
Coding quotidiano
+
⚡ Aggressive
~50%
Sessioni lunghe con molti tool
+
🔥 Ultra
~75%
Massimo risparmio
+
🧰 RTK
60–90%
Output di shell/test/build/git
+
🔗 Stacked (RTK → Caveman)
78–95%
Prompt misti + log dei tool
+
-```txt
-1) Start MCP transport (`omniroute --mcp`) for tool-driven operations
-2) Run A2A tasks via `message/send` and `message/stream`
-3) Observe via /dashboard/endpoint (MCP and A2A tabs)
-4) Toggle services via inline status controls
-```
+**Esempio reale — modalità Standard:**
----
+> **Prima (69 token):** _"The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I would recommend using useMemo to memoize the object."_
+>
+> **Dopo (19 token):** _"New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo."_
+>
+> **Stessa risposta. 72% di token in meno. Nessuna perdita di accuratezza.** ✅
-## 🆓 Start Free — Zero Configuration Cost
+**Esempio PT-BR — modalità [Troglodita](https://github.com/leninejunior/troglodita):**
-> Setup AI coding in minutes at **$0/month**. Connect these free accounts and use the built-in **Free Stack** combo.
+> **Antes (42 tokens):** _"O problema é que o componente está re-renderizando porque uma nova referência de objeto está sendo criada em cada ciclo de renderização. Eu recomendaria usar useMemo."_
+>
+> **Depois (12 tokens):** _"Re-render: ref nova cada ciclo (objeto inline recriado). Usar `useMemo`."_
+>
+> **Mesma resposta. ~70% menos tokens. Precisão técnica intacta.** ✅
-| Step | Action | Providers Unlocked |
-| ---- | -------------------------------------------------- | ------------------------------------------------------------------ |
-| 1 | Connect **Kiro** (AWS Builder ID OAuth) | Claude Sonnet 4.5, Haiku 4.5 — provider/account limits apply |
-| 2 | Connect **Qoder** (Google OAuth) | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1... — provider/account limits apply |
-| 3 | Connect **Qwen** (Device Code) | qwen3-coder-plus, qwen3-coder-flash... — provider/account limits apply |
-| 4 | `/dashboard/combos` → **Free Stack ($0)** template | Round-robin all free providers automatically |
+
-**Point any IDE/CLI to:** `http://localhost:20128/v1` · API Key: `any-string` · Done.
+### 🎚️ Oltre i motori — output style, regolazione adattiva e controllo per richiesta
-> **Optional extra coverage (current terms apply):** Groq, NVIDIA NIM, Cerebras, LongCat and Cloudflare Workers AI can provide free access or signup credits where currently listed. Quotas, models, accounts, regions and provider terms can change; see [`FREE_TIERS.md`](../../reference/FREE_TIERS.md).
+I 12 motori sopra riducono ciò che entra **in input**. Altri tre livelli definiscono **come**, **quando** e cosa esce **in output**:
-## Avvio Rapido
+- **🪄 Output Styles** _(controllo dell'output)_ — iniettano istruzioni deterministiche e cache-safe per modellare la risposta; sono combinabili, ciascuno con intensità `lite` / `full` / `ultra`. Aggiungere uno style richiede una sola voce nel registry:
+ - **Terse prose** — elimina riempitivi / articoli / esitazioni; mantiene esatto il contenuto tecnico.
+ - **Less code** — YAGNI da "senior dev pigro": modifica minima funzionante, nessuna infrastruttura non richiesta.
+ - **Terse CJK (文言)** — stile cinese classico ultra-conciso (limitato alla locale `zh`).
+- **🎯 Adaptive context-budget** _(la regolazione)_ — invece di una singola soglia token on/off, aumenta gradualmente l'uso dei motori più economici e lossless solo quanto necessario per **rientrare nella context window del modello**. Policy: `reserve-output` (predefinita, model-aware) · `percentage` · `absolute`. Modalità: `floor` (garantisce il fit) · `replace-autotrigger` (vince la tua scelta esplicita) · `off` (soglia legacy).
+- **🎛️ Dove viene decisa la compressione** _(precedenza, alta → bassa)_ — header per richiesta `x-omniroute-compression` › override del routing combo › profilo nominato attivo › adaptive / auto-trigger › impostazione predefinita del pannello › off. Il piano applicato viene restituito nell'header di risposta `X-OmniRoute-Compression: ; source=`.
-### 1) Install and run
+Puoi attivare l'auto-trigger tramite soglia token, abilitare la regolazione adattiva, fissare un profilo nominato, impostare una scelta una tantum per richiesta oppure assegnare una pipeline a ciascun routing combo — scegli ciò che si adatta al carico di lavoro. Un **eval harness** offline opt-in (`npm run eval:compression`) misura fedeltà e risparmio su un corpus fissato prima di promuovere una modifica.
+
+📖 [`COMPRESSION_GUIDE.md`](../../compression/COMPRESSION_GUIDE.md) · [`RTK_COMPRESSION.md`](../../compression/RTK_COMPRESSION.md) · [`COMPRESSION_ENGINES.md`](../../compression/COMPRESSION_ENGINES.md)
+
+
+
+
+
+
+# ⚡ Avvio rapido
+
+
+
+**1) Installa e avvia**
```bash
npm install -g omniroute
omniroute
```
-> **pnpm users:** Pass `--allow-build` at install time to enable native build scripts required by `better-sqlite3` and `@swc/core` (the `approve-builds -g` command is not supported for global installs on pnpm v11):
->
-> ```bash
-> pnpm add -g omniroute@latest --allow-build=better-sqlite3 --allow-build=@swc/core
-> omniroute
-> ```
+> 💡 Vedi `npm warn ERESOLVE` o avvisi sulle peer dependency? [Sono innocui](../../guides/TROUBLESHOOTING.md#npm-install-warnings-eresolve--peer--deprecated).
-Dashboard opens at `http://localhost:20128` and API base URL is `http://localhost:20128/v1`.
+Dashboard su `http://localhost:20128` · API su `http://localhost:20128/v1`.
-#### Arch Linux (AUR)
+**2) Collega un provider GRATUITO (senza registrazione)**
-Arch Linux users can install the [AUR package](https://aur.archlinux.org/packages/omniroute-bin), which installs OmniRoute and provides a systemd user service:
+Dashboard → **Providers** → collega **Kiro AI** (Claude gratuito, ~50 crediti/mese per account) oppure **OpenCode Free** (nessuna autenticazione) → fatto.
-```bash
-yay -S omniroute-bin
-systemctl --user enable --now omniroute.service
-```
-
-| Command | Description |
-| ----------------------- | ----------------------------------------------------------- |
-| `omniroute` | Start server (`PORT=20128`, API and dashboard on same port) |
-| `omniroute --port 3000` | Set canonical/API port to 3000 |
-| `omniroute --mcp` | Start MCP server (stdio transport) |
-| `omniroute --no-open` | Don't auto-open browser |
-| `omniroute --help` | Show help |
-
-Optional split-port mode:
-
-```bash
-PORT=20128 DASHBOARD_PORT=20129 omniroute
-# API: http://localhost:20128/v1
-# Dashboard: http://localhost:20129
-```
-
-### 2) Uninstalling
-
-When you no longer need OmniRoute, we provide two quick scripts for a clean removal:
-
-| Command | Action |
-| ------------------------ | ----------------------------------------------------------------------------------- |
-| `npm run uninstall` | Removes the system app but **keeps your DB and configurations** in `~/.omniroute`. |
-| `npm run uninstall:full` | Removes the app AND permanently **erases all configurations, keys, and databases**. |
-
-> Note: To run these commands, navigate to the OmniRoute project folder (if you cloned it) and run them. Alternatively, if globally installed, you can simply run `npm uninstall -g omniroute`.
-
-### Long-Running Streaming Timeouts
-
-For most deployments, you only need:
-
-| Variable | Default | Purpose |
-| ------------------------ | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
-| `REQUEST_TIMEOUT_MS` | `600000` | Shared baseline for upstream response-start timeout, hidden Undici timeouts, TLS fingerprint requests, and API bridge request/proxy timeouts |
-| `STREAM_IDLE_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` | Maximum gap between streaming chunks before OmniRoute aborts the SSE stream |
-
-Backward compatibility is preserved: existing `FETCH_TIMEOUT_MS`, `API_BRIDGE_PROXY_TIMEOUT_MS`, and other per-layer timeout vars still work and override the shared baseline.
-
-For Claude Code-compatible upstreams (`anthropic-compatible-cc-*`), OmniRoute also derives the outbound `X-Stainless-Timeout` header from the resolved fetch timeout so provider-side read timeouts stay aligned with your env configuration.
-
-For third-party Claude Code-compatible reverse proxies, OmniRoute keeps the default
-`anthropic-beta` set conservative and, when `Client Cache Control` is left on `Auto`,
-only forwards client-provided `cache_control` markers. If the request does not include
-`cache_control`, OmniRoute does not inject bridge-owned markers.
-
-Advanced overrides are available if you need finer control:
-
-| Variable | Default | Purpose |
-| ---------------------------------------- | ------------------------------------------ | -------------------------------------------------------------------- |
-| `FETCH_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` | Upstream response-start timeout used until response headers arrive |
-| `FETCH_HEADERS_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Undici time limit for receiving upstream response headers |
-| `FETCH_BODY_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Undici time limit between upstream body chunks (`0` disables it) |
-| `FETCH_CONNECT_TIMEOUT_MS` | `30000` | Undici TCP connect timeout |
-| `FETCH_KEEPALIVE_TIMEOUT_MS` | `4000` | Undici idle keep-alive socket timeout |
-| `TLS_CLIENT_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Timeout for TLS fingerprint requests made through `wreq-js` |
-| `API_BRIDGE_PROXY_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` or `600000` | Timeout for `/v1` proxy forwarding from API port to dashboard port |
-| `API_BRIDGE_SERVER_REQUEST_TIMEOUT_MS` | `max(API_BRIDGE_PROXY_TIMEOUT_MS, 300000)` | Incoming request timeout on the API bridge server |
-| `API_BRIDGE_SERVER_HEADERS_TIMEOUT_MS` | `60000` | Incoming header timeout on the API bridge server |
-| `API_BRIDGE_SERVER_KEEPALIVE_TIMEOUT_MS` | `5000` | Keep-alive timeout on the API bridge server |
-| `API_BRIDGE_SERVER_SOCKET_TIMEOUT_MS` | `0` | Socket inactivity timeout on the API bridge server (`0` disables it) |
-
-For streaming requests, `FETCH_TIMEOUT_MS` only covers connection setup / waiting for the first upstream response. Once the stream is active, OmniRoute will only abort on an actual stall (`STREAM_IDLE_TIMEOUT_MS`) or Undici body inactivity (`FETCH_BODY_TIMEOUT_MS`).
-
-If you run OmniRoute behind Nginx, Caddy, Cloudflare, or another reverse proxy, make sure the proxy
-timeouts are also higher than your OmniRoute stream/fetch timeouts.
-
-### 2) Connect providers and create your API key
-
-1. Open Dashboard → `Providers` and connect at least one provider (OAuth or API key).
-2. Open Dashboard → `Endpoints` and create an API key.
-3. (Optional) Open Dashboard → `Combos` and set your fallback chain.
-
-### 3) Point your coding tool to OmniRoute
+**3) Configura il tuo strumento di coding**
```txt
Base URL: http://localhost:20128/v1
-API Key: [copy from Endpoint page]
-Model: if/kimi-k2-thinking (or any provider/model prefix)
+API Key: [copy from Dashboard → Endpoints]
+Model: auto (zero-config smart routing — or any provider/model)
```
-### 4) Enable and validate protocols (v2.0)
-
-**MCP (for tool-driven operations):**
+**4) Verifica che funzioni**
```bash
-omniroute --mcp
+curl http://localhost:20128/v1/models -H "Authorization: Bearer YOUR_KEY"
```
-Then connect your MCP client over `stdio` and test tools like:
+Dovresti vedere elencati i modelli collegati. 🎉 Tutto qui — inizia a programmare: OmniRoute instrada automaticamente le richieste ed esegue il fallback quando serve.
-- `omniroute_get_health`
-- `omniroute_list_combos`
-
-**A2A (for agent-to-agent workflows):**
-
-```bash
-curl http://localhost:20128/.well-known/agent.json
-```
-
-```bash
-curl -X POST http://localhost:20128/a2a \
- -H 'content-type: application/json' \
- -d '{"jsonrpc":"2.0","id":"quickstart","method":"message/send","params":{"skill":"quota-management","messages":[{"role":"user","content":"Give me a short quota summary."}]}}'
-```
-
-### 5) Validate everything end-to-end (recommended)
-
-```bash
-npm run test:protocols:e2e
-```
-
-This suite validates real MCP and A2A client flows against a running app.
-
-### Alternative: run from source
-
-```bash
-cp .env.example .env
-npm install
-PORT=20128 DASHBOARD_PORT=20129 NEXT_PUBLIC_BASE_URL=http://localhost:20129 npm run dev
-```
-
-
-Void Linux (`xbps-src` template)
-
-For Void Linux users, you can build a native package using `xbps-src`. Save this block as `srcpkgs/omniroute/template`:
-
-```bash
-# Template file for 'omniroute'
-pkgname=omniroute
-version=3.4.1
-revision=1
-hostmakedepends="nodejs python3 make"
-depends="openssl"
-short_desc="Universal AI gateway with smart routing for multiple LLM providers"
-maintainer="zenobit "
-license="MIT"
-homepage="https://github.com/diegosouzapw/OmniRoute"
-distfiles="https://github.com/diegosouzapw/OmniRoute/archive/refs/tags/v${version}.tar.gz"
-checksum=009400afee90a9f32599d8fe734145cfd84098140b7287990183dde45ae2245b
-system_accounts="_omniroute"
-omniroute_homedir="/var/lib/omniroute"
-export NODE_ENV=production
-export npm_config_engine_strict=false
-export npm_config_loglevel=error
-export npm_config_fund=false
-export npm_config_audit=false
-
-do_build() {
- # Determine target CPU arch for node-gyp
- local _gyp_arch
- case "$XBPS_TARGET_MACHINE" in
- aarch64*) _gyp_arch=arm64 ;;
- armv7*|armv6*) _gyp_arch=arm ;;
- i686*) _gyp_arch=ia32 ;;
- *) _gyp_arch=x64 ;;
- esac
-
- # 1) Install all deps – skip scripts (no network in do_build, native modules
- # compiled separately below; better-sqlite3 is serverExternalPackage so
- # Next.js does not execute it during next build)
- NODE_ENV=development npm ci --ignore-scripts
-
- # 2) Build the Next.js standalone bundle
- npm run build
-
- # 3) Copy static assets into standalone
- cp -r .next/static .next/standalone/.next/static
- [ -d public ] && cp -r public .next/standalone/public || true
-
- # 4) Compile better-sqlite3 native binding for the target architecture.
- # Use node-gyp directly so CC/CXX from xbps-src cross-toolchain are used
- # without npm altering them.
- local _node_gyp=/usr/lib/node_modules/npm/node_modules/node-gyp/bin/node-gyp.js
- (cd node_modules/better-sqlite3 && node "$_node_gyp" rebuild --arch="$_gyp_arch")
-
- # 5) Place the compiled binding into the standalone bundle
- local _bs3_release=.next/standalone/node_modules/better-sqlite3/build/Release
- mkdir -p "$_bs3_release"
- cp node_modules/better-sqlite3/build/Release/better_sqlite3.node "$_bs3_release/"
-
- # 6) Remove arch-specific sharp bundles – upstream sets images.unoptimized=true
- # so sharp is not used at runtime; x64 .so files would break aarch64 strip
- rm -rf .next/standalone/node_modules/@img
-
- # 7) Copy pino runtime deps omitted by Next.js static analysis:
- # pino-abstract-transport – required by pino's worker thread
- # split2 – dep of pino-abstract-transport
- # process-warning – dep of pino itself
- for _mod in pino-abstract-transport split2 process-warning; do
- cp -r "node_modules/$_mod" .next/standalone/node_modules/
- done
-}
-
-do_check() {
- npm run test:unit
-}
-
-do_install() {
- vmkdir usr/lib/omniroute/.next
-
- vcopy .next/standalone/. usr/lib/omniroute/.next/standalone
-
- # Prevent removal of empty Next.js app router dirs by the post-install hook
- for _d in \
- .next/standalone/.next/server/app/dashboard \
- .next/standalone/.next/server/app/dashboard/settings \
- .next/standalone/.next/server/app/dashboard/providers; do
- touch "${DESTDIR}/usr/lib/omniroute/${_d}/.keep"
- done
-
- cat > "${WRKDIR}/omniroute" <<'EOF'
-#!/bin/sh
-export PORT="${PORT:-20128}"
-export DATA_DIR="${DATA_DIR:-${XDG_DATA_HOME:-${HOME}/.local/share}/omniroute}"
-export APP_LOG_TO_FILE="${APP_LOG_TO_FILE:-false}"
-mkdir -p "${DATA_DIR}"
-exec node /usr/lib/omniroute/.next/standalone/server.js "$@"
-EOF
- vbin "${WRKDIR}/omniroute"
-}
-
-post_install() {
- vlicense LICENSE
-}
-```
-
-
-
----
-
-## 🐳 Docker
-
-OmniRoute is available as a public Docker image on [Docker Hub](https://hub.docker.com/r/diegosouzapw/omniroute).
-
-**Quick run:**
-
-```bash
-docker run -d \
- --name omniroute \
- --restart unless-stopped \
- --stop-timeout 40 \
- -p 20128:20128 \
- -v omniroute-data:/app/data \
- diegosouzapw/omniroute:latest
-```
-
-**With environment file:**
-
-```bash
-# Copy and edit .env first
-cp .env.example .env
-
-docker run -d \
- --name omniroute \
- --restart unless-stopped \
- --stop-timeout 40 \
- --env-file .env \
- -p 20128:20128 \
- -v omniroute-data:/app/data \
- diegosouzapw/omniroute:latest
-```
-
-**Using Docker Compose:**
-
-```bash
-# Base profile (no CLI tools)
-docker compose --profile base up -d
-
-# CLI profile (Claude Code, Codex, OpenClaw built-in)
-docker compose --profile cli up -d
-```
-
-Dashboard support for Docker deployments now includes a one-click **Cloudflare Quick Tunnel** on `Dashboard → Endpoints`. The first enable downloads `cloudflared` only when needed, starts a temporary tunnel to your current `/v1` endpoint, and shows the generated `https://*.trycloudflare.com/v1` URL directly below your normal public URL.
-
-Notes:
-
-- Quick Tunnel URLs are temporary and change after every restart.
-- Quick Tunnels are not auto-restored after an OmniRoute or container restart. Re-enable them from the dashboard when needed.
-- Managed install currently supports Linux, macOS, and Windows on `x64` / `arm64`.
-- Managed Quick Tunnels default to HTTP/2 transport to avoid noisy QUIC UDP buffer warnings in constrained container environments. Set `CLOUDFLARED_PROTOCOL=quic` or `auto` if you want a different transport.
-- Docker images bundle system CA roots and pass them to managed `cloudflared`, which avoids TLS trust failures when the tunnel bootstraps inside the container.
-- SQLite runs in WAL mode. `docker stop` should be allowed to finish so OmniRoute can checkpoint the latest changes back into `storage.sqlite`.
-- The bundled Compose files already set a 40s stop grace period. If you run the image directly, keep `--stop-timeout 40` (or similar) so manual stops do not cut off shutdown cleanup.
-- Set `CLOUDFLARED_BIN=/absolute/path/to/cloudflared` if you want OmniRoute to use an existing binary instead of downloading one.
-
-**Using Docker Compose with Caddy (HTTPS Auto-TLS):**
-
-OmniRoute can be securely exposed using Caddy's automatic SSL provisioning. Ensure your domain's DNS A record points to your server's IP.
-
-```yaml
-services:
- omniroute:
- image: diegosouzapw/omniroute:latest
- container_name: omniroute
- restart: unless-stopped
- volumes:
- - omniroute-data:/app/data
- environment:
- - PORT=20128
- - NEXT_PUBLIC_BASE_URL=https://your-domain.com
-
- caddy:
- image: caddy:latest
- container_name: caddy
- restart: unless-stopped
- ports:
- - "80:80"
- - "443:443"
- command: caddy reverse-proxy --from https://your-domain.com --to http://omniroute:20128
-
-volumes:
- omniroute-data:
-```
-
-| Image | Tag | Size | Description |
-| ------------------------ | -------- | ------ | --------------------- |
-| `diegosouzapw/omniroute` | `latest` | ~250MB | Latest stable release |
-| `diegosouzapw/omniroute` | `3.6.2` | ~250MB | Current version |
-
----
-
-## 🖥️ Desktop App — Offline & Always-On
-
-> 🆕 **NEW!** OmniRoute is now available as a **native desktop application** for Windows, macOS, and Linux.
-
-Run OmniRoute as a standalone desktop app — no terminal, no browser, no internet required for local models. The Electron-based app includes:
-
-- 🖥️ **Native Window** — Dedicated app window with system tray integration
-- 🔄 **Auto-Start** — Launch OmniRoute on system login
-- 🔔 **Native Notifications** — Get alerts for quota exhaustion or provider issues
-- ⚡ **One-Click Install** — NSIS (Windows), DMG (macOS), AppImage (Linux)
-- 🌐 **Offline Mode** — Works fully offline with bundled server
-
-### Avvio Rapido
-
-```bash
-# Development mode
-npm run electron:dev
-
-# Build for your platform
-npm run electron:build # Current platform
-npm run electron:build:win # Windows (.exe)
-npm run electron:build:mac # macOS (.dmg) — x64 & arm64
-npm run electron:build:linux # Linux (.AppImage)
-```
-
-### System Tray
-
-When minimized, OmniRoute lives in your system tray with quick actions:
-
-- Open dashboard
-- Change server port
-- Quit application
-
-📖 Full documentation: [`electron/README.md`](electron/README.md)
-
----
-
-## 💰 Pricing at a Glance
-
-| Tier | Provider | Cost | Quota Reset | Best For |
-| ------------------- | --------------------------- | ------------------------------------- | --------------------- | ---------------------------------- |
-| **💳 SUBSCRIPTION** | Claude Code (Pro) | $20/mo | 5h + weekly | Already subscribed |
-| | Codex (Plus/Pro) | $20-200/mo | 5h + weekly | OpenAI users |
-| | GitHub Copilot | $10-19/mo | Monthly | GitHub users |
-| **🔑 API KEY** | NVIDIA NIM | **FREE ACCESS** (current terms apply) | ~40 RPM | 70+ open models |
-| | Cerebras | **FREE** (1M tok/day) | 60K TPM / 30 RPM | World's fastest |
-| | Groq | **FREE** (30 RPM) | 14.4K RPD | Ultra-fast Llama/Gemma |
-| | DeepSeek V3.2 | $0.27/$1.10 per 1M | None | Best price/quality reasoning |
-| | xAI Grok-4 Fast | **$0.20/$0.50 per 1M** 🆕 | None | Fastest + tool calling, ultralow |
-| | xAI Grok-4 (standard) | $0.20/$1.50 per 1M 🆕 | None | Reasoning flagship from xAI |
-| | Mistral | Free trial + paid | Rate limited | European AI |
-| | OpenRouter | Pay-per-use | None | 100+ models aggr. |
-| **💰 CHEAP** | GLM-5 (via Z.AI) 🆕 | $0.5/1M | Daily 10AM | 128K output, newest flagship |
-| | GLM-4.7 | $0.6/1M | Daily 10AM | Budget backup |
-| | MiniMax M2.5 🆕 | $0.3/1M input | 5-hour rolling | Reasoning + agentic tasks |
-| | MiniMax M2.1 | $0.2/1M | 5-hour rolling | Cheapest option |
-| | Kimi K2.5 (Moonshot API) 🆕 | Pay-per-use | None | Direct Moonshot API access |
-| | Kimi K2 | $9/mo flat | 10M tokens/mo | Predictable cost |
-| **🆓 FREE ACCESS** | Qoder | **$0** | Limits apply | Selected models; terms apply |
-| | Qwen | **$0** | Limits apply | Selected models; terms apply |
-| | Kiro | **$0** | Credit/account limits | Claude access; current terms apply |
-| | LongCat signup credit | **$0** (10M one-time; KYC) | One-time | Signup grant; not recurring |
-| | Pollinations AI 🆕 | **$0** (no key needed) | 1 req/15s | GPT-5, Claude, DeepSeek, Llama 4 |
-| | Cloudflare Workers AI 🆕 | **$0** (10K Neurons/day) | ~150 resp/day | 50+ models, global edge |
-| | Scaleway AI 🆕 | **$0** (1M tokens total) | Rate limited | EU/GDPR, Qwen3 235B, Llama 70B |
-
-> 🆕 **New models added (Mar 2026):** Grok-4 Fast family at $0.20/$0.50/M (benchmarked at 1143ms — 30% faster than Gemini 2.5 Flash), GLM-5 via Z.AI with 128K output, MiniMax M2.5 reasoning, DeepSeek V3.2 updated pricing, Kimi K2.5 via Moonshot direct API.
-
-**💡 $0 Combo Stack — The Complete Free Setup:**
-
-```
-# 🆓 Free-access examples — provider limits and terms apply
-Kiro (kr/) → Claude access — account/credit limits apply
-Qoder (if/) → selected models — no published token cap; rate/account limits apply
-LongCat (lc/) → LongCat-2.0 — 10M one-time signup credit; KYC required
-Pollinations (pol/) → GPT-5, Claude, DeepSeek, Llama 4 — no key needed
-Qwen (qw/) → selected models — no published token cap; rate/account limits apply
-Gemini (gemini/) → selected free-tier models — current API quotas apply
-Cloudflare AI (cf/) → Llama 70B, Gemma 3, Mistral — 10K Neurons/day
-Scaleway (scw/) → Qwen3 235B, Llama 70B — 1M free tokens (EU)
-Groq (groq/) → selected models — current per-model rate limits apply
-NVIDIA NIM (nvidia/) → selected models — current rate limits apply
-Cerebras (cerebras/) → Llama/Qwen world-fastest — 1M tok/day
-```
-
-**Current $0 access where listed; availability is not guaranteed.** A combo can try the next eligible route when a quota or upstream fails.
-
----
-
----
-
-## 🆓 Free Models — What You Actually Get
-
-> The entries below summarize access that was listed as free when audited. Provider quotas, card/account/KYC requirements, models, regions and terms can change. A combo broadens fallback coverage but does not guarantee uninterrupted $0 access.
-
-### 🔵 CLAUDE MODELS (via Kiro — AWS Builder ID)
-
-| Model | Prefix | Limit | Rate Limit |
-| ------------------- | ------ | ------------- | --------------------- |
-| `claude-sonnet-4.5` | `kr/` | No published token cap | Provider/account limits may apply |
-| `claude-haiku-4.5` | `kr/` | No published token cap | Provider/account limits may apply |
-| `claude-opus-4.6` | `kr/` | No published token cap | Latest Opus; provider/account limits apply |
-
-### 🟢 QODER MODELS (Free PAT via qodercli)
-
-| Model | Prefix | Limit | Rate Limit |
-| ------------------ | ------ | ------------- | --------------- |
-| `kimi-k2-thinking` | `if/` | No published token cap | Provider/account limits may apply |
-| `qwen3-coder-plus` | `if/` | No published token cap | Provider/account limits may apply |
-| `deepseek-r1` | `if/` | No published token cap | Provider/account limits may apply |
-| `minimax-m2.1` | `if/` | No published token cap | Provider/account limits may apply |
-| `kimi-k2` | `if/` | No published token cap | Provider/account limits may apply |
-
-> Recommended connection method: **Personal Access Token + `qodercli`**. Browser OAuth is
-> experimental and disabled by default unless `QODER_OAUTH_*` environment variables are configured.
-
-### 🟡 QWEN MODELS (Device Code Auth)
-
-| Model | Prefix | Limit | Rate Limit |
-| ------------------- | ------ | ------------- | ------------------- |
-| `qwen3-coder-plus` | `qw/` | No published token cap | Provider/account limits may apply |
-| `qwen3-coder-flash` | `qw/` | No published token cap | Provider/account limits may apply |
-| `qwen3-coder-next` | `qw/` | No published token cap | Provider/account limits may apply |
-| `vision-model` | `qw/` | No published token cap | Multimodal; provider/account limits may apply |
-
-### ⚫ NVIDIA NIM (Free API Key — build.nvidia.com)
-
-| Tier | Daily Limit | Rate Limit | Notes |
-| ---------- | ------------ | ----------- | ------------------------------------------------------ |
-| Free (Dev) | No token cap | **~40 RPM** | 70+ models; transitioning to pure rate limits mid-2025 |
-
-Popular free models: `moonshotai/kimi-k2.5` (Kimi K2.5), `z-ai/glm4.7` (GLM 4.7), `deepseek-ai/deepseek-v3.2` (DeepSeek V3.2), `nvidia/llama-3.3-70b-instruct`, `deepseek/deepseek-r1`
-
-### ⚪ CEREBRAS (Free API Key — inference.cerebras.ai)
-
-| Tier | Daily Limit | Rate Limit | Notes |
-| ---- | ----------------- | ---------------- | ------------------------------------------- |
-| Free | **1M tokens/day** | 60K TPM / 30 RPM | World's fastest LLM inference; resets daily |
-
-Available free: `llama-3.3-70b`, `llama-3.1-8b`, `deepseek-r1-distill-llama-70b`
-
-### 🔴 GROQ (Free API Key — console.groq.com)
-
-| Tier | Daily Limit | Rate Limit | Notes |
-| ---- | ------------- | ---------------- | ----------------------------------------- |
-| Free | **14.4K RPD** | 30 RPM per model | No credit card; 429 on limit, not charged |
-
-Available free: `llama-3.3-70b-versatile`, `gemma2-9b-it`, `mixtral-8x7b`, `whisper-large-v3`
-
-### 🔴 LONGCAT AI (Signup credit — KYC required)
-
-| Model | Prefix | Current catalog grant | Notes |
-| ------------- | ------ | ----------------------- | --------------------------------------------------- |
-| `LongCat-2.0` | `lc/` | **10M tokens one-time** | Signup grant; not a recurring monthly or daily pool |
-
-> Provider terms, eligibility and model availability can change. See [`FREE_TIERS.md`](../../reference/FREE_TIERS.md) for the audited catalog entry.
-
-### 🟢 POLLINATIONS AI (No API Key Required) 🆕
-
-| Model | Prefix | Rate Limit | Provider Behind |
-| ---------- | ------ | ---------- | ------------------ |
-| `openai` | `pol/` | 1 req/15s | GPT-5 |
-| `claude` | `pol/` | 1 req/15s | Anthropic Claude |
-| `gemini` | `pol/` | 1 req/15s | Google Gemini |
-| `deepseek` | `pol/` | 1 req/15s | DeepSeek V3 |
-| `llama` | `pol/` | 1 req/15s | Meta Llama 4 Scout |
-| `mistral` | `pol/` | 1 req/15s | Mistral AI |
-
-> ✨ **Zero friction:** No signup, no API key. Add the Pollinations provider with an empty key field and it works immediately.
-
-### 🟠 CLOUDFLARE WORKERS AI (Free API Key — cloudflare.com) 🆕
-
-| Tier | Daily Neurons | Equivalent Usage | Notes |
-| ---- | ------------- | --------------------------------------- | ----------------------- |
-| Free | **10,000** | ~150 LLM resp / 500s audio / 15K embeds | Global edge, 50+ models |
-
-Popular free models: `@cf/meta/llama-3.3-70b-instruct`, `@cf/google/gemma-3-12b-it`, `@cf/openai/whisper-large-v3-turbo` (free audio!), `@cf/qwen/qwen2.5-coder-15b-instruct`
-
-> Requires API Token + Account ID from [dash.cloudflare.com](https://dash.cloudflare.com). Store Account ID in provider settings.
-
-### 🟣 SCALEWAY AI (1M Free Tokens — scaleway.com) 🆕
-
-| Tier | Free Quota | Location | Notes |
-| ---- | ------------- | ------------ | ----------------------------------- |
-| Free | **1M tokens** | 🇫🇷 Paris, EU | No credit card needed within limits |
-
-Available free: `qwen3-235b-a22b-instruct-2507` (Qwen3 235B!), `llama-3.1-70b-instruct`, `mistral-small-3.2-24b-instruct-2506`, `deepseek-v3-0324`
-
-> EU/GDPR compliant. Get API key at [console.scaleway.com](https://console.scaleway.com).
-
-> **💡 Free-access examples (provider limits and terms apply):**
->
-> ```
-> Kiro (kr/) → Claude access — account/credit limits apply
-> Qoder (if/) → selected models — no published token cap; limits apply
-> LongCat (lc/) → LongCat-2.0 — 10M one-time signup credit; KYC required
-> Pollinations (pol/) → GPT-5, Claude, DeepSeek, Llama 4 — no key needed
-> Qwen (qw/) → selected models — no published token cap; limits apply
-> Gemini (gemini/) → selected free-tier models — current quotas apply
-> Cloudflare AI (cf/) → 50+ models — 10K Neurons/day
-> Scaleway (scw/) → Qwen3 235B, Llama 70B — 1M free tokens (EU)
-> Groq (groq/) → selected models — current per-model rate limits apply
-> NVIDIA NIM (nvidia/) → selected models — current rate limits apply
-> Cerebras (cerebras/) → Llama/Qwen world-fastest — 1M tok/day
-> ```
-
-## 🎙️ Free Transcription Combo
-
-> Transcription access depends on each upstream allowance — Deepgram and AssemblyAI signup credits can lead, with Groq Whisper as a rate-limited fallback.
-
-| Provider | Free Credits | Best Model | Rate Limit |
-| ----------------- | --------------------------- | -------------------------------------------- | ---------------------------------------- |
-| 🟢 **Deepgram** | **$200 free** (signup) | `nova-3` — best accuracy, 30+ languages | No RPM limit on free credits |
-| 🔵 **AssemblyAI** | **$50 free** (signup) | `universal-3-pro` — chapters, sentiment, PII | No RPM limit on free credits |
-| 🔴 **Groq** | **Free tier; limits apply** | `whisper-large-v3` — OpenAI Whisper | Current model-specific rate limits apply |
-
-**Suggested combo in `/dashboard/combos`:**
-
-```
-Name: free-transcription
-Strategy: Priority
-Nodes:
- [1] deepgram/nova-3 → uses $200 free first
- [2] assemblyai/universal-3-pro → fallback when Deepgram credits run out
- [3] groq/whisper-large-v3 → free access; rate limits apply
-```
-
-Then in `/dashboard/media` → **Transcription** tab: upload any audio or video file → select your combo endpoint → get transcription in supported formats.
-
-## 💡 Key Features
-
-OmniRoute v3.6 is built as an operational platform, not just a relay proxy.
-
-### 🆕 New — v3.6.x Highlights (Apr 2026)
-
-| Feature | What It Does |
-| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
-| 🌐 **V1 WebSocket Bridge** | OpenAI-compatible WebSocket traffic upgraded and proxied via `/v1/ws` — full streaming over WS with session auth (API key or session cookie) |
-| 🔑 **Sync Tokens & Config Bundle** | Issue/revoke sync tokens for config sync endpoints. Config bundles versioned with ETag for bandwidth-efficient polling |
-| 🧠 **GLM Thinking (glmt) Preset** | GLM Thinking registered first-class: 65 536 max tokens, 24 576 thinking budget, 900s timeout, usage sync & pricing — Claude-compatible API |
-| 🔢 **Hybrid Token Counting** | Uses provider-side `/messages/count_tokens` when available; falls back to estimation — accurate usage tracking without guessing |
-| 🌱 **Model Alias Auto-Seed** | 30+ cross-proxy dialect aliases normalised at startup — no more routing mismatches |
-| 🛡️ **Safe Outbound Fetch** | All provider validation and model discovery go through a guarded fetch layer blocking private/local URLs with retry, timeout, and SSRF protection |
-| ⏳ **Wait For Cooldown** | Server-side chat retries when every candidate connection is cooling down; configurable `enabled`, `maxRetries`, and `maxRetryWaitSec` |
-| 🔍 **Runtime Env Validation** | Startup validates all env vars with Zod schemas — clear errors for missing secrets, invalid URLs, or wrong types |
-| 📋 **Compliance Audit Expansion** | Structured audit logs with pagination, request context, auth events, provider CRUD events, and SSRF-blocked validation logging |
-| 🔐 **TPS Log Metric** | Log details modal shows Tokens Per Second (TPS) — quick performance at-a-glance for every request |
-| 🗑️ **Uninstall / Full Uninstall** | `npm run uninstall` keeps data, `npm run uninstall:full` removes everything — clean removal for all install methods |
-| 🔧 **OAuth Env Repair** | One-click "Repair env" action for OAuth providers restores missing env vars and fixes broken auth state |
-| 🔒 **Graceful Electron Shutdown** | Electron `before-quit` shuts down Next.js gracefully, preventing SQLite WAL database locks on desktop close |
-| 👁️ **Model Visibility Toggle** | Per-model visibility toggle (👁 icon) with search filter and active-count badge (`N/M active`) on provider pages |
-| 📧 **Email Privacy Masking** | OAuth account emails masked (`di*****@g****.com`), full address visible on hover |
-| 🔗 **Context Relay Strategy** | Combo strategy preserving session continuity via structured handoff summaries when accounts rotate mid-conversation |
-| 🛡️ **Proxy Hardening** | Token health check, API key validation, and undici dispatcher all honor proxy config |
-| ⚠️ **Node.js 24 Login Warning** | Login page proactively detects incompatible Node.js versions and shows a clear warning banner |
-| 📎 **Gemini PDF Attachments** | PDF attachments correctly routed to Gemini via `inline_data` and generic base64 detection |
-| 🔒 **CodeQL Security Hardening** | Resolved SSRF, insecure randomness, polynomial ReDoS, and incomplete URL sanitization alerts |
-
-### 🆕 New — ClawRouter-Inspired Improvements (Mar 2026)
-
-| Feature | What It Does |
-| ------------------------------------ | ------------------------------------------------------------------------------------------- |
-| ⚡ **Grok-4 Fast Family** | xAI models at $0.20/$0.50/M — benchmarked 1143ms (30% faster than Gemini 2.5 Flash) |
-| 🧠 **GLM-5 via Z.AI** | 128K output context, $0.5/1M — newest flagship from the GLM family |
-| 🔮 **MiniMax M2.5** | Reasoning + agentic tasks at $0.30/1M — significant upgrade from M2.1 |
-| 🎯 **toolCalling Flag per Model** | Per-model `toolCalling: true/false` in registry — AutoCombo skips non-tool-capable models |
-| 🌍 **Multilingual Intent Detection** | PT/ZH/ES/AR keywords in AutoCombo scoring — better model selection for non-English content |
-| 📊 **Benchmark-Driven Fallbacks** | Real p95 latency from live requests feeds combo scoring — AutoCombo learns from actual data |
-| 🔁 **Request Deduplication** | Content-hash based dedup window — multi-agent safe, prevents duplicate charges |
-| 🔌 **Pluggable RouterStrategy** | Extensible `RouterStrategy` interface — add custom routing logic as plugins |
-
-### 🚀 Previous v2.0.9+ — Playground, CLI Fingerprints & ACP
-
-| Feature | What It Does |
-| --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
-| 🎮 **Model Playground** | Dashboard page to test any model directly — provider/model/endpoint selectors, Monaco Editor, streaming, abort, timing |
-| 🔏 **CLI Fingerprint Matching** | Per-provider header/body ordering to match native CLI signatures — toggle per provider in Settings > Security. **Your proxy IP is preserved** |
-| 🤖 **ACP Agents Dashboard** | Debug › Agents page — grid of 14 agents with install status, version, custom agent form for any CLI tool. **OpenCode** users get a "Download opencode.json" button that auto-generates a ready-to-use config with all available models. |
-| 🔧 **Custom Model `apiFormat` Routing** | Custom models with `apiFormat: "responses"` now correctly route to the Responses API translator |
-| 🏢 **Codex Workspace Isolation** | Multiple Codex workspaces per email — OAuth correctly separates connections by workspace ID |
-| 🔄 **Electron Auto-Update** | Desktop app checks for updates + auto-install on restart |
-
-### 🤖 Agent & Protocol Operations (v2.0)
-
-| Feature | What It Does |
-| ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
-| 🔧 **MCP Server (107 tools)** | IDE/agent tools via 3 transports: stdio, SSE (`/api/mcp/sse`), Streamable HTTP (`/api/mcp/stream`). 107 unique tools across the registered tool families; enabled skills may add dynamic tools at runtime |
-| 🤝 **A2A Server (JSON-RPC + SSE)** | Agent-to-agent task execution with sync and streaming flows |
-| 🧭 **Consolidated Endpoints Page** | Tabbed management page with Endpoint Proxy, MCP, A2A, and API Endpoints tabs |
-| 🎚️ **Service Enable/Disable Toggles** | ON/OFF switches for MCP and A2A with settings persistence (default: OFF) |
-| 🛰️ **MCP Runtime Heartbeat** | Real process status (pid, uptime, heartbeat age, transport, scope mode) |
-| 📋 **MCP Audit Trail** | Filterable audit logs with success/failure and key attribution |
-| 🔐 **MCP Scope Enforcement** | 32 granular scope permissions for controlled tool access |
-| 📡 **A2A Task Lifecycle Management** | List/filter tasks, inspect events/artifacts, cancel running tasks |
-| 📋 **Agent Card Discovery** | `/.well-known/agent.json` for client auto-discovery |
-| 🧪 **Protocol E2E Test Harness** | Real MCP SDK + A2A client flows in `test:protocols:e2e` |
-| ⚙️ **Operational Controls** | Switch combos, tune resilience settings, and review breaker state from dedicated Health and Settings surfaces |
-
-### 🧠 Routing & Intelligence
-
-| Feature | What It Does |
-| ---------------------------------- | ------------------------------------------------------------------------ |
-| 🎯 **Smart 4-Tier Fallback** | Auto-route: Subscription → API Key → Cheap → Free |
-| 📊 **Real-Time Quota Tracking** | Live token count + reset countdown per provider |
-| 🔄 **Format Translation** | OpenAI ↔ Claude ↔ Gemini ↔ Responses with schema-safe conversions |
-| 👥 **Multi-Account Support** | Multiple accounts per provider with intelligent selection |
-| 🔄 **Auto Token Refresh** | OAuth tokens refresh automatically with retry |
-| 🎨 **Custom Combos** | 13 balancing strategies + fallback chain control |
-| 🔗 **Context Relay** | Session continuity handoffs when account rotation happens mid-session |
-| 🌐 **Wildcard Router** | `provider/*` dynamic routing |
-| 🧠 **Thinking Budget Controls** | Passthrough, auto, custom, and adaptive reasoning limits |
-| 🔀 **Model Aliases** | Built-in + custom model aliasing and migration safety |
-| ⚡ **Background Degradation** | Route low-priority background tasks to cheaper models |
-| 🧪 **Task-Aware Smart Routing** | Auto-select model by content type (coding/vision/analysis/summarization) |
-| 🔄 **A2A Agent Workflows** | Deterministic FSM orchestrator for stateful multi-step agent executions |
-| 🔀 **Adaptive Routing** | Dynamic strategy override based on token volume and prompt complexity |
-| 🎲 **Provider Diversity** | Shannon entropy scoring balancing auto-combo traffic distribution |
-| 💬 **System Prompt Injection** | Global behavior controls applied consistently |
-| 📄 **Responses API Compatibility** | Full `/v1/responses` support for Codex and advanced agentic workflows |
-
-### 🎵 Multi-Modal APIs
-
-| Feature | What It Does |
-| -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
-| 🖼️ **Image Generation** | `/v1/images/generations` with cloud and local backends |
-| 📐 **Embeddings** | `/v1/embeddings` for search and RAG pipelines |
-| 🎤 **Audio Transcription** | `/v1/audio/transcriptions` — 7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
-| 🔊 **Text-to-Speech** | `/v1/audio/speech` — 10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) with correct error messages |
-| 🎬 **Video Generation** | `/v1/videos/generations` (ComfyUI + SD WebUI workflows) |
-| 🎵 **Music Generation** | `/v1/music/generations` (ComfyUI workflows) |
-| 🛡️ **Moderations** | `/v1/moderations` safety checks |
-| 🔀 **Reranking** | `/v1/rerank` for relevance scoring |
-| 🔍 **Web Search** 🆕 | `/v1/search` — 5 providers (Serper, Brave, Perplexity, Exa, Tavily), 6,500+ free/month, auto-failover, cache |
-
-### 🛡️ Resilience, Security & Governance
-
-| Feature | What It Does |
-| ----------------------------------- | ------------------------------------------------------------------------------------------------------- |
-| 🔌 **Provider Circuit Breakers** | Provider-wide trip/recover after fallback exhaustion with configurable thresholds |
-| 🔒 **Daily Quota Lock** 🆕 | Detects exhaustion signals and locks routing for the specific model until midnight |
-| 🎯 **Endpoint-Aware Models** | Custom models declare supported endpoints + API format |
-| 🛡️ **Anti-Thundering Herd** | Mutex + semaphore protections on retry/rate events |
-| 🧠 **Semantic + Signature Cache** | Cost/latency reduction with two cache layers |
-| ⚡ **Request Idempotency** | Duplicate protection window |
-| 🔒 **TLS Fingerprint Spoofing** | Browser-like TLS fingerprint — **reduces bot detection and account flagging** |
-| 🔏 **CLI Fingerprint Matching** | Matches native CLI request signatures — **reduces ban risk while preserving proxy IP** |
-| 🌐 **IP Filtering** | Allowlist/blocklist control for exposed deployments |
-| 🚦 **Request Queue & Pacing** | Configurable per-connection request buckets for RPM, spacing, concurrency, and max wait |
-| 📉 **Graceful Degradation** | Multi-layer capability fallbacks protecting core gateway operations |
-| 📜 **Config Audit Trail** | Diff-based change tracking preventing operational drift with simple rollbacks |
-| ⏳ **Provider Health Sync** | Proactive token expiration monitoring triggering alerts before authorization failures |
-| ❄️ **Connection Cooldown** | Retryable 408/429/5xx failures cool down a single connection with optional upstream hints |
-| 🚪 **Auto-Disable Banned Accounts** | Permanently blocked token accounts can be disabled automatically |
-| 🔑 **API Key Management + Scoping** | Secure key issuance/rotation and model/provider controls |
-| 👁️ **Scoped API Key Reveal** 🆕 | Opt-in recovery of API keys via `ALLOW_API_KEY_REVEAL` |
-| 🛡️ **Protected `/models`** | Optional auth gating and provider hiding for model catalog |
-| 🛡️ **Safe Outbound Fetch** 🆕 | Guarded fetch for provider calls — blocks private/local URLs, retries, SSRF protection |
-| ⏳ **Wait For Cooldown** 🆕 | Auto-retry chat after connection cooldowns; configurable `enabled`, `maxRetries`, and `maxRetryWaitSec` |
-| 🔍 **Runtime Env Validation** 🆕 | Zod-based env schema validation at startup with actionable error messages |
-| 📋 **Compliance Audit v2** 🆕 | Pagination, request context, auth events, provider CRUD, and SSRF-blocked logging |
-
-### 📊 Observability & Analytics
-
-| Feature | What It Does |
-| -------------------------------- | ----------------------------------------------------- |
-| 📝 **Request + Proxy Logging** | Full request/response and proxy logging |
-| 📉 **Streamed Detailed Logs** | Reconstructs SSE payload streams cleanly into the UI |
-| 🏷️ **Real-Time Model Badges** 🆕 | Live model status and daily quota countdown timers |
-| 📋 **Unified Logs Dashboard** | Request, proxy, audit, and console views in one page |
-| 🔍 **Request Telemetry** | p50/p95/p99 latency and request tracing |
-| 🏥 **Health Dashboard** | Uptime, breaker states, lockouts, cache stats |
-| 💰 **Cost Tracking** | Budget controls and per-model pricing visibility |
-| 📈 **Analytics Visualizations** | Model/provider usage insights and trend views |
-| 🧪 **Evaluation Framework** | Golden set testing with configurable match strategies |
-| 📡 **Live Diagnostics** 🆕 | Semantic cache bypass for accurate combo live testing |
-| 🔐 **TPS Log Metric** 🆕 | Tokens Per Second badge in log details modal |
-
-### ☁️ Deployment & Platform
-
-| Feature | What It Does |
-| ------------------------------ | --------------------------------------------------------------------- |
-| 🌐 **Deploy Anywhere** | Localhost, VPS, Docker, Cloud environments |
-| 🚇 **Cloudflare Tunnel** 🆕 | One-click Quick Tunnel integration from the dashboard |
-| 🔑 **API Key Model Filtering** | Native /v1/models response filtered via assigned Bearer context roles |
-| ⚡ **Smart Cache Bypass** | Configurable TTL heuristics and forced refetch controls |
-| 🔄 **Backup/Restore** | Export/import and disaster recovery flows |
-| 🧙 **Onboarding Wizard** | First-run guided setup |
-| 🔧 **CLI Tools Dashboard** | One-click setup for popular coding tools |
-| 🎮 **Model Playground** | Test any provider/model/endpoint from the dashboard |
-| 🔏 **CLI Fingerprint Toggle** | Per-provider fingerprint matching in Settings > Security |
-| 🌐 **i18n (30 languages)** | Full dashboard + docs language support with RTL coverage |
-| 🧹 **Clear All Models** | One-click model list clearing in provider details |
-| 👁️ **Sidebar Controls** 🆕 | Hide components and integrations from Appearance Settings |
-| 📋 **Issue Templates** | Standardized GitHub templates for bugs and features |
-| 📂 **Custom Data Directory** | `DATA_DIR` override for storage location |
-| 🌐 **V1 WebSocket Bridge** 🆕 | OpenAI-compatible WebSocket traffic proxied via `/v1/ws` |
-| 🔑 **Sync Tokens & Bundle** 🆕 | Config sync tokens + versioned bundle endpoint with ETag support |
-
-### Feature Deep Dive
-
-#### Smart fallback with practical cost control
+Se il tuo client non può inviare header personalizzati, OmniRoute espone anche alias di compatibilità con token incorporato:
```txt
-Combo: "my-coding-stack"
- 1. cc/claude-opus-4-7
- 2. nvidia/llama-3.3-70b
- 3. glm/glm-4.7
- 4. if/kimi-k2-thinking
+OpenAI catalog: http://localhost:20128/vscode/YOUR_KEY/
+OpenAI models: http://localhost:20128/vscode/YOUR_KEY/models
+OpenAI chat: http://localhost:20128/vscode/YOUR_KEY/chat/completions
+OpenAI responses: http://localhost:20128/vscode/YOUR_KEY/responses
+Ollama chat: http://localhost:20128/vscode/YOUR_KEY/api/chat
+Ollama tags: http://localhost:20128/vscode/YOUR_KEY/api/tags
```
-When quota, rate, or health fails, OmniRoute automatically moves to the next candidate without manual switching.
+Usali solo con client che non possono aggiungere `Authorization: Bearer ...`. L'autenticazione tramite header resta la modalità consigliata.
-#### Protocol management that is visible and operable
+
-- MCP + A2A are discoverable in UI and docs (not hidden)
-- Protocol status APIs expose live operational data (`/api/mcp/*`, `/api/a2a/*`)
-- Dashboards include actions for day-2 ops (combo toggles, breaker resets, task cancellation)
+
+## 📦 Altri metodi di installazione — Docker, sorgente, pnpm, Arch
-#### Translator + validation workflow
-
-The Translator area includes:
-
-- **Playground**: request transformation checks
-- **Chat Tester**: full request/response round-trip
-- **Test Bench**: multiple cases in one run
-- **Live Monitor**: real-time traffic view
-
-Plus protocol validation with real clients via `npm run test:protocols:e2e`.
-
-> 📖 **[MCP Server README](open-sse/mcp-server/README.md)** — Tool reference, IDE configs, and client examples
->
-> 📖 **[A2A Server README](src/lib/a2a/README.md)** — Skills, JSON-RPC methods, streaming, and task lifecycle
-
-## 🧪 Evaluations (Evals)
-
-OmniRoute includes a built-in evaluation framework to test LLM response quality against a golden set. Access it via **Analytics → Evals** in the dashboard.
-
-### Built-in Golden Set
-
-The pre-loaded "OmniRoute Golden Set" contains test cases for:
-
-- Greetings, math, geography, code generation
-- JSON format compliance, translation, markdown generation
-- Safety refusal (harmful content), counting, boolean logic
-
-### Evaluation Strategies
-
-| Strategy | Description | Example |
-| ---------- | ------------------------------------------------ | -------------------------------- |
-| `exact` | Output must match exactly | `"4"` |
-| `contains` | Output must contain substring (case-insensitive) | `"Paris"` |
-| `regex` | Output must match regex pattern | `"1.*2.*3"` |
-| `custom` | Custom JS function returns true/false | `(output) => output.length > 10` |
-
----
-
-## 📖 Setup Guide
-
-### Protocol Setup (MCP + A2A)
-
-
-🧩 MCP Setup (Model Context Protocol)
-
-Start MCP transport in stdio mode:
+**🐳 Docker**
```bash
-omniroute --mcp
+docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \
+ -p 127.0.0.1:20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest
```
-Recommended validation flow:
+`:latest` segue la versione SemVer stabile **pubblicata** più alta. Non segue il branch git `main`. Per GitOps, fissa `:X.Y.Z`. Vedi [Canali di release Docker](../../guides/DOCKER_GUIDE.md#release-channels). L'immagine imposta **`OMNIROUTE_MEMORY_MB=1024`**. È sufficiente per la dashboard e una chat leggera. I **coding agent** (`POST /v1/responses` da Claude Code, Codex, Grok, …) richiedono un heap V8 molto più grande, altrimenti il processo va in `FATAL ERROR` a ~12 GiB con due contesti lunghi sovrapposti. Dimensiona il container oltre l'heap (i buffer nativi si trovano fuori da V8):
-1. Connect your MCP client over stdio.
-2. Run `omniroute_get_health`.
-3. Run `omniroute_list_combos`.
-4. Open `/dashboard/mcp` to confirm heartbeat, activity, and audit.
-
-Useful APIs for automation:
-
-- `GET /api/mcp/status`
-- `GET /api/mcp/tools`
-- `GET /api/mcp/audit`
-- `GET /api/mcp/audit/stats`
-
-
-
-
-🤝 A2A Setup (Agent2Agent)
-
-Discover the agent:
+| Carico di lavoro | Heap (`-e OMNIROUTE_MEMORY_MB`) | Container (`--memory`) |
+| ----------------------------------- | ------------------------------- | ---------------------- |
+| Dashboard / chat leggera | `1024` (predefinito immagine) | ≥2 g |
+| Un coding agent | `8192` | ≥10 g |
+| Due `/v1/responses` lunghe simultanee | `10240`–`12288` | ≥12–16 g |
```bash
-curl http://localhost:20128/.well-known/agent.json
+docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \
+ -e OMNIROUTE_MEMORY_MB=8192 --memory=10g \
+ -p 127.0.0.1:20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest
```
-Send a task:
+Tabella completa: [Guida Docker — RAM di runtime](../../guides/DOCKER_GUIDE.md#runtime-ram-for-coding-agents).
+
+> **Canale Docker pre-release:** `diegosouzapw/omniroute:next` e
+> `diegosouzapw/omniroute:next-web` seguono l'attuale branch `release/v*` predefinito.
+> Questi tag mutabili sono destinati esclusivamente al test di fix non ancora rilasciati e
+> **non sono supportati in produzione**. Vedi
+> [Canali di release Docker](../../guides/DOCKER_GUIDE.md#release-channels).
+
+**🥟 Bun**
+
+Sono supportati `bun install` standard e l'installazione globale (`bun install -g omniroute`) tramite rilevamento del runtime Bun:
+- **`bun:sqlite` integrato**: OmniRoute usa il driver integrato `bun:sqlite` quando gira con Bun, con fallback a `better-sqlite3` su Node.js o a `sql.js`.
+- **Selezione automatica del bundler Webpack**: sviluppo (`bun run dev`) e build di produzione (`bun run build`) rilevano automaticamente Bun e disabilitano Turbopack a favore di Webpack per evitare incompatibilità dei binding V8 nativi.
+- **Dockerfile Bun dedicato**: `Dockerfile.bun` multi-stage per deployment di produzione nativi Bun (`docker build -f Dockerfile.bun -t omniroute:bun .`).
```bash
-curl -X POST http://localhost:20128/a2a \
- -H 'content-type: application/json' \
- -d '{"jsonrpc":"2.0","id":"setup-a2a","method":"message/send","params":{"skill":"quota-management","messages":[{"role":"user","content":"Summarize quota status."}]}}'
+# Install and run with Bun
+bun install
+bun run dev
```
-Manage lifecycle:
-
-- `GET /api/a2a/status`
-- `GET /api/a2a/tasks`
-- `GET /api/a2a/tasks/:id`
-- `POST /api/a2a/tasks/:id/cancel`
-
-Operational UI:
-
-- `/dashboard/a2a` for task/state/stream observability and smoke actions
-
-
-
-
-🧪 End-to-end protocol validation
-
-Validate both protocols with real clients:
+**🛠️ Da sorgente**
```bash
-npm run test:protocols:e2e
+cp .env.example .env && npm install
+PORT=20128 npm run dev
```
-This verifies:
-
-- MCP SDK client connect/list/call
-- A2A discovery/send/stream/get/cancel
-- Cross-check data in MCP audit and A2A task management APIs
-
-
-
-
-💳 Subscription Providers
-
-### Claude Code (Pro/Max)
+**📦 pnpm**
```bash
-Dashboard → Providers → Connect Claude Code
-→ OAuth login → Auto token refresh
-→ 5-hour + weekly quota tracking
-
-Models:
- cc/claude-opus-4-7
- cc/claude-sonnet-4-5-20250929
- cc/claude-haiku-4-5-20251001
+pnpm add -g omniroute@latest --allow-build=better-sqlite3 --allow-build=@swc/core && omniroute
```
-**Pro Tip:** Use Opus for complex tasks, Sonnet for speed. OmniRoute tracks quota per model!
-
-### OpenAI Codex (Plus/Pro)
+**🐧 Arch Linux (AUR)**
```bash
-Dashboard → Providers → Connect Codex
-→ OAuth login (port 1455)
-→ 5-hour + weekly reset
-
-Models:
- cx/gpt-5.2-codex
- cx/gpt-5.1-codex-max
+yay -S omniroute-bin && systemctl --user enable --now omniroute.service
```
-#### Codex Account Limit Management (5h + Weekly)
-
-Each Codex account now has policy toggles in `Dashboard -> Providers`:
-
-- `5h` (ON/OFF): enforce the 5-hour window threshold policy.
-- `Weekly` (ON/OFF): enforce the weekly window threshold policy.
-- Threshold behavior: when an enabled window reaches >=90% usage, that account is skipped.
-- Rotation behavior: OmniRoute routes to the next eligible Codex account automatically.
-- Reset behavior: when the provider `resetAt` time passes, the account becomes eligible again automatically.
-
-Scenarios:
-
-- `5h ON` + `Weekly ON`: account is skipped when either window reaches threshold.
-- `5h OFF` + `Weekly ON`: only weekly usage can block the account.
-- `5h ON` + `Weekly OFF`: only 5-hour usage can block the account.
-- `resetAt` passed: account re-enters rotation automatically (no manual re-enable).
-
-### GitHub Copilot
+**🔧 Nix (Flake)**
```bash
-Dashboard → Providers → Connect GitHub
-→ OAuth via GitHub
-→ Monthly reset (1st of month)
-
-Models:
- gh/gpt-5
- gh/claude-4.5-sonnet
- gh/gemini-3.1-pro-preview
-```
-
-
-
-
-🔑 API Key Providers
-
-### NVIDIA NIM (FREE developer access — 70+ models)
-
-1. Sign up: [build.nvidia.com](https://build.nvidia.com)
-2. Get free API key (1000 inference credits included)
-3. Dashboard → Add Provider → NVIDIA NIM:
- - API Key: `nvapi-your-key`
-
-**Models:** `nvidia/llama-3.3-70b-instruct`, `nvidia/mistral-7b-instruct`, and 50+ more
-
-**Pro Tip:** OpenAI-compatible API — works seamlessly with OmniRoute's format translation!
-
-### DeepSeek
-
-1. Sign up: [platform.deepseek.com](https://platform.deepseek.com)
-2. Get API key
-3. Dashboard → Add Provider → DeepSeek
-
-**Models:** `deepseek/deepseek-chat`, `deepseek/deepseek-coder`
-
-### Groq (Free Tier Available!)
-
-1. Sign up: [console.groq.com](https://console.groq.com)
-2. Get API key (free tier included)
-3. Dashboard → Add Provider → Groq
-
-**Models:** `groq/llama-3.3-70b`, `groq/mixtral-8x7b`
-
-**Pro Tip:** Ultra-fast inference — best for real-time coding!
-
-### OpenRouter (100+ Models)
-
-1. Sign up: [openrouter.ai](https://openrouter.ai)
-2. Get API key
-3. Dashboard → Add Provider → OpenRouter
-
-**Models:** Access 100+ models from all major providers through a single API key.
-
-**Dashboard behavior:** OpenRouter models are managed from **Available Models**. Manual add, import, and auto-sync all update the same list.
-
-
-
-
-💰 Cheap Providers (Backup)
-
-### GLM-4.7 (Daily reset, $0.6/1M)
-
-1. Sign up: [Zhipu AI](https://open.bigmodel.cn/)
-2. Get API key from Coding Plan
-3. Dashboard → Add API Key:
- - Provider: `glm`
- - API Key: `your-key`
-
-**Use:** `glm/glm-4.7`
-
-**Pro Tip:** Coding Plan offers 3× quota at 1/7 cost! Reset daily 10:00 AM.
-
-### MiniMax M2.1 (5h reset, $0.20/1M)
-
-1. Sign up: [MiniMax](https://www.minimax.io/)
-2. Get API key
-3. Dashboard → Add API Key
-
-**Use:** `minimax/MiniMax-M2.1`
-
-**Pro Tip:** Cheapest option for long context (1M tokens)!
-
-### Kimi K2 ($9/month flat)
-
-1. Subscribe: [Moonshot AI](https://platform.moonshot.ai/)
-2. Get API key
-3. Dashboard → Add API Key
-
-**Use:** `kimi/kimi-latest`
-
-**Pro Tip:** Fixed $9/month for 10M tokens = $0.90/1M effective cost!
-
-
-
-
-🆓 FREE Providers (Emergency Backup)
-
-### Qoder (5 FREE models via OAuth)
-
-```bash
-Dashboard → Connect Qoder
-→ Qoder OAuth login
-→ Access is subject to current provider limits
-
-Models:
- if/kimi-k2-thinking
- if/qwen3-coder-plus
- if/glm-4.7
- if/minimax-m2
- if/deepseek-r1
-```
-
-### Qwen (4 FREE models via Device Code)
-
-```bash
-Dashboard → Connect Qwen
-→ Device code authorization
-→ Access is subject to current provider limits
-
-Models:
- qw/qwen3-coder-plus
- qw/qwen3-coder-flash
-```
-
-### Kiro (Claude FREE)
-
-```bash
-Dashboard → Connect Kiro
-→ AWS Builder ID or Google/GitHub
-→ Access is subject to current provider limits
-
-Models:
- kr/claude-sonnet-4.5
- kr/claude-haiku-4.5
-```
-
-
-
-
-🎨 Create Combos
-
-### Example 1: Maximize Subscription → Cheap Backup
-
-```
-Dashboard → Combos → Create New
-
-Name: premium-coding
-Models:
- 1. cc/claude-opus-4-7 (Subscription primary)
- 2. glm/glm-4.7 (Cheap backup, $0.6/1M)
- 3. minimax/MiniMax-M2.1 (Cheapest fallback, $0.20/1M)
-
-Use in CLI: premium-coding
-```
-
-### Example 2: Free-Only (Zero Cost)
-
-```
-Name: free-combo
-Models:
- 1. if/kimi-k2-thinking (no published token cap; provider limits may apply)
- 2. qw/qwen3-coder-plus (no published token cap; provider limits may apply)
-
-Cost: currently listed as $0; terms and availability may change
-```
-
-
-
-
-🔧 CLI Integration
-
-### Cursor IDE
-
-```
-Settings → Models → Advanced:
- OpenAI API Base URL: http://localhost:20128/v1
- OpenAI API Key: [from OmniRoute dashboard]
- Model: cc/claude-opus-4-7
-```
-
-### Claude Code
-
-Use the **CLI Tools** page in the dashboard for one-click configuration, or edit `~/.claude/settings.json` manually.
-
-### Codex CLI
-
-```bash
-export OPENAI_BASE_URL="http://localhost:20128"
-export OPENAI_API_KEY="your-omniroute-api-key"
-
-codex "your prompt"
-```
-
-### OpenClaw
-
-**Option 1 — Dashboard (recommended):**
-
-```
-Dashboard → CLI Tools → OpenClaw → Select Model → Apply
-```
-
-**Option 2 — Manual:** Edit `~/.openclaw/openclaw.json`:
-
-```json
-{
- "models": {
- "providers": {
- "omniroute": {
- "baseUrl": "http://127.0.0.1:20128/v1",
- "apiKey": "sk_omniroute",
- "api": "openai-completions"
- }
- }
- }
-}
-```
-
-> **Note:** OpenClaw only works with local OmniRoute. Use `127.0.0.1` instead of `localhost` to avoid IPv6 resolution issues.
-
-### Cline / Continue / RooCode
-
-```
-Settings → API Configuration:
- Provider: OpenAI Compatible
- Base URL: http://localhost:20128/v1
- API Key: [from OmniRoute dashboard]
- Model: if/kimi-k2-thinking
-```
-
-### OpenCode
-
-**Step 1:** Add OmniRoute as a custom provider:
-
-```bash
-opencode
-/connect
-# Select "Other" → Enter ID: "omniroute" → Enter your OmniRoute API key
-```
-
-**Step 2:** Create/edit `opencode.json` in your project root:
-
-```json
-{
- "$schema": "https://opencode.ai/config.json",
- "provider": {
- "omniroute": {
- "npm": "@ai-sdk/openai-compatible",
- "name": "OmniRoute",
- "options": {
- "baseURL": "http://localhost:20128/v1"
- },
- "models": {
- "cc/claude-sonnet-4-20250514": { "name": "Claude Sonnet 4" },
- "gg/gemini-2.5-pro": { "name": "Gemini 2.5 Pro" },
- "if/kimi-k2-thinking": { "name": "Kimi K2 (Free)" }
- }
- }
- }
-}
-```
-
-**Step 3:** Select the model in OpenCode:
-
-```bash
-/models
-# Select any OmniRoute model from the list
-```
-
-> **Tip:** Add any model available in your OmniRoute `/v1/models` endpoint to the `models` section. Use the format `provider/model-id` from your OmniRoute dashboard.
-
-
-
----
-
-## Risoluzione dei Problemi
-
-
-Click to expand troubleshooting guide
-
-**"Language model did not provide messages"**
-
-- Provider quota exhausted → Check dashboard quota tracker
-- Solution: Use combo fallback or switch to cheaper tier
-
-**Rate limiting**
-
-- Subscription quota out → Fallback to GLM/MiniMax
-- Add combo: `cc/claude-opus-4-7 → glm/glm-4.7 → if/kimi-k2-thinking`
-
-**OAuth token expired**
-
-- Auto-refreshed by OmniRoute
-- If issues persist: Dashboard → Provider → Reconnect
-
-**High costs**
-
-- Check usage stats in Dashboard → Costs
-- Switch primary model to GLM/MiniMax
-
-**Dashboard/API ports are wrong**
-
-- `PORT` is the canonical base port (and API port by default)
-- `API_PORT` overrides only OpenAI-compatible API listener
-- `DASHBOARD_PORT` overrides only dashboard/Next.js listener
-- Set `NEXT_PUBLIC_BASE_URL` to your dashboard/public URL (for OAuth callbacks)
-
-**Cloud sync errors**
-
-- Verify `BASE_URL` points to your running instance
-- Verify `CLOUD_URL` points to your expected cloud endpoint
-- Keep `NEXT_PUBLIC_*` values aligned with server-side values
-
-**First login not working**
-
-- Check `INITIAL_PASSWORD` in `.env`
-- If unset, fallback password is `123456`
-
-**No request logs**
-
-- `call_logs` in SQLite stores summary metadata for the Request Logs table and analytics views
-- Detailed request/response payloads are written to `DATA_DIR/call_logs/` as one JSON artifact per request
-- Enable pipeline capture from Dashboard → Logs → Request Logs if you need detailed per-stage payloads
-- `Export Logs` reads the artifact files on demand, while `Export All` includes the `call_logs/` directory alongside `storage.sqlite`
-- Set `APP_LOG_TO_FILE=true` if you also want application console logs in `logs/application/app.log`
-- Adjust `APP_LOG_MAX_FILE_SIZE`, `APP_LOG_RETENTION_DAYS`, `APP_LOG_MAX_FILES`, and `CALL_LOG_MAX_ENTRIES` as needed
-
-**Connection test shows "Invalid" for OpenAI-compatible providers**
-
-- Many providers don't expose a `/models` endpoint
-- OmniRoute v1.0.6+ includes fallback validation via chat completions
-- Ensure base URL includes `/v1` suffix
-
-### 🔐 OAuth on a Remote Server
-
-
-
-
-> **⚠️ Important for users running OmniRoute on a VPS, Docker, or any remote server**
-
-The OAuth credentials bundled in OmniRoute are registered **for `localhost` only**. When you access OmniRoute on a remote server (e.g. `https://omniroute.myserver.com`), Google rejects the authentication with:
-
-```
-Error 400: redirect_uri_mismatch
-```
-
-#### Solution: Configure your own OAuth credentials
-
-You need to create an **OAuth 2.0 Client ID** in Google Cloud Console with your server's URI.
-
-#### Step-by-step
-
-**1. Open Google Cloud Console**
-
-Go to: [https://console.cloud.google.com/apis/credentials](https://console.cloud.google.com/apis/credentials)
-
-**2. Create a new OAuth 2.0 Client ID**
-
-- Click **"+ Create Credentials"** → **"OAuth client ID"**
-- Application type: **"Web application"**
-- Name: anything you like (e.g. `OmniRoute Remote`)
-
-**3. Add Authorized Redirect URIs**
-
-In the **"Authorized redirect URIs"** field, add:
-
-```
-https://your-server.com/callback
-```
-
-> Replace `your-server.com` with your server's domain or IP (include the port if needed, e.g. `http://45.33.32.156:20128/callback`).
-
-**4. Save and copy the credentials**
-
-After creating, Google will show the **Client ID** and **Client Secret**.
-
-**5. Set environment variables**
-
-In your `.env` (or Docker environment variables):
-
-```bash
-# For Antigravity:
-ANTIGRAVITY_OAUTH_CLIENT_ID=your-client-id.apps.googleusercontent.com
-ANTIGRAVITY_OAUTH_CLIENT_SECRET=GOCSPX-your-secret
-
-GEMINI_OAUTH_CLIENT_ID=your-client-id.apps.googleusercontent.com
-GEMINI_OAUTH_CLIENT_SECRET=GOCSPX-your-secret
-```
-
-**6. Restart OmniRoute**
-
-```bash
-# npm:
+# Using Nix flakes
+nix develop
npm run dev
-# Docker:
-docker restart omniroute
+# Or using devbox
+devbox run npm run dev
```
-**7. Try connecting again**
+📖 [Guida Docker](../../guides/DOCKER_GUIDE.md) — profili Compose, Caddy HTTPS, tunnel Cloudflare.
-Google will now redirect correctly to `https://your-server.com/callback`.
+**🦭 Podman**
+
+```bash
+# 1. Prepare the bind-mounted data directory
+mkdir -p data
+
+# 2. Linux + local rootless Podman only (never a remote Podman Machine client):
+podman unshare chown 1000:1000 ./data
+
+# 3. Set the runtime hint, build the local Compose image, and start
+echo "CONTAINER_HOST=podman" >> .env
+podman compose --profile base up -d --build
+```
+
+Su macOS o Windows, Podman usa una Podman Machine remota: salta `podman unshare` e
+segui le [indicazioni sui permessi della directory dati specifiche per topologia](../../../contrib/podman/README.md#data-directory-permissions-by-topology).
+
+📖 [Guida Podman](../../../contrib/podman/README.md) — build Compose, Podman Machine e
+configurazione Quadlet Linux/systemd.
+
+**⚡ Installazione più rapida / leggera (salta la build nativa)**
+
+Il motore SQLite nativo (`better-sqlite3`) è una dipendenza **opzionale**, quindi un'installazione
+globale non si blocca mai per compilare da sorgente: usa un binario precompilato quando disponibile
+per la tua piattaforma/Node e altrimenti passa in modo trasparente a un motore pure-JS
+(`node:sqlite` su Node 22+, altrimenti `sql.js` WASM incluso) — senza richiedere strumenti di build.
+
+Per saltare completamente il warm-up nativo post-installazione (CI, sistemi headless o macchine lente):
+
+```bash
+OMNIROUTE_SKIP_POSTINSTALL=1 npm install -g omniroute # CI=1 also skips it
+```
+
+Per installazioni più rapide preferisci **pnpm** (store content-addressed + hard link — vedi sopra).
+Per un runtime headless senza dashboard usa il profilo Docker `base` (sopra) oppure la
+[guida Termux](../../guides/TERMUX_GUIDE.md). CLI e dashboard web sono servite dallo
+stesso processo su una sola porta, quindi oggi non esiste un pacchetto separato solo CLI.
+
+
+
+
+
+
+# 🎬 OmniRoute in azione
+
+
+
+## 📹 Guide video
+
+
+
+Dati di copertura social al 2026-08-17 · YT: 741 | TT: 137 | IG: 124 · Aggiornamento (giorni): YT 0 · TT 14 · IG 15
+
+
-If you don't want to set up your own credentials right now, you can still use the **manual URL flow**:
+
+## 🛠️ Stack tecnologico
-1. OmniRoute opens the Google authorization URL
-2. After authorizing, Google tries to redirect to `localhost` (which fails on the remote server)
-3. **Copy the full URL** from your browser's address bar (even if the page doesn't load)
-4. Paste that URL into the field shown in the OmniRoute connection modal
-5. Click **"Connect"**
+
-> This works because the authorization code in the URL is valid regardless of whether the redirect page loaded.
+
Strategia di copertura dei test e suite da oltre 25.000 test
+
+
+
+
+
+
+# ⭐ Principali contributor
+
+> OmniRoute è plasmato da una community open source appassionata. Queste persone hanno apportato contributi eccezionali che incidono direttamente su qualità, stabilità e diffusione del progetto. **Grazie.**
+
+
+
+> 🙏 Funzionalità, bug fix e miglioramenti infrastrutturali di questi contributor sono una **parte fondamentale** di ciò che rende OmniRoute affidabile e ricco di funzionalità. Ogni pull request, ogni caso di test e ogni file di traduzione i18n conta. L'open source è costruito da persone come loro.
+
+
-
+Un grazie di cuore alle persone che finanziano OmniRoute di tasca propria — ogni contributo aiuta a mantenere il progetto gratuito, indipendente e in evoluzione.
----
+
-- 🔗 **OpenCode Integration** — Native provider support for the OpenCode AI coding IDE
-- 🔗 **TRAE Integration** — Full support for the TRAE AI development framework
-- 📦 **Batch API** — Asynchronous batch processing for bulk requests
-- 🎯 **Tag-Based Routing** — Route requests based on custom tags and metadata
-- 💰 **Lowest-Cost Strategy** — Automatically select the cheapest available provider
+[](https://github.com/diegosouzapw/OmniRoute/graphs/contributors)
-> 📝 Full feature specifications available in [`docs/new-features/`](docs/new-features/) (217 detailed specs)
+### Come contribuire
----
+1. Fai un fork del repository
+2. Crea il branch dalla punta della `release/vX.Y.Z` **attiva** (non da `main`) — vedi [Modello di branching e release](../../ops/BRANCHING_MODEL.md)
+3. Crea il tuo feature branch (`git checkout -b feat/amazing-feature`)
+4. Esegui il commit delle modifiche (`git commit -m 'feat: add amazing feature'`)
+5. Esegui il push del branch (`git push origin feat/amazing-feature`)
+6. Apri una Pull Request con **base = quel branch `release/vX.Y.Z`**
-## 👥 Contributors
+Vedi [CONTRIBUTING.md](../../../CONTRIBUTING.md) per le linee guida complete.
-[](https://github.com/diegosouzapw/OmniRoute/graphs/contributors)
-
-### How to Contribute
-
-1. Fork the repository
-2. Create your feature branch (`git checkout -b feature/amazing-feature`)
-3. Commit your changes (`git commit -m 'Add amazing feature'`)
-4. Push to the branch (`git push origin feature/amazing-feature`)
-5. Open a Pull Request
-
-See [CONTRIBUTING.md](CONTRIBUTING.md) for detailed guidelines.
-
-### Releasing a New Version
+### Pubblicare una nuova versione
```bash
# Create a release — npm publish happens automatically
-gh release create v2.0.0 --title "v2.0.0" --generate-notes
+gh release create v3.8.2 --title "v3.8.2" --generate-notes
```
----
+
-## 📊 Star History
+
-## 🙏 Acknowledgments
+
-Special thanks to **[CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI)** — the original Go implementation that inspired this JavaScript port.
+
----
+## 🙏 Ringraziamenti
-## Licenza
+
-MIT License - see [LICENSE](LICENSE) for details.
+OmniRoute è costruito sulle spalle di giganti. È nato come fork di **[9router](https://github.com/decolua/9router)** e come port TypeScript del progetto Go **[CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI)** — da lì, ogni sottosistema qui sotto è stato ispirato da un progetto open source arrivato prima. Ognuno ha influenzato una parte concreta di OmniRoute. Questo è il nostro ringraziamento a tutti loro. 🙏
+
+> ⭐ conteggio stelle a luglio 2026 — vai a lasciare una stella a questi progetti.
+
+### 🧬 Origini e gateway
+
+
Il gateway AI il cui dataset pubblico dei prezzi alimenta la sincronizzazione del cost tracking e il cui modello di normalizzazione dei provider ha influenzato il nostro routing.
+
+
+### 🗜️ Compressione di contesto e token — motori
+
+
Il progetto virale "why use many token when few token do trick" — la sua filosofia caveman-speak alimenta la nostra modalità di compressione standard e oltre 30 regole di rimozione riempitivi/condensazione.
Compressione ad alte prestazioni dell'output dei comandi — ha ispirato il nostro motore RTK, la DSL per filtri JSON, il recupero dell'output grezzo e la pipeline stacked RTK → Caveman.
Compressione token PT-BR — alimenta il nostro language pack pt-BR: riduzione dei pleonasmi e rimozione dei riempitivi ottimizzate per la grammatica portoghese brasiliana.
La skill virale da "lazy senior dev" basata su YAGNI — ha ispirato il nostro Output Style less-code: orientamento alla modifica minima funzionante che riduce il codice _generato_ (l'equivalente sull'asse output della prosa concisa di Caveman).
+
+
+### 🧩 Formati compatti, ricerca sui token e tooling code-aware
+
+
Ha inizialmente ispirato la nostra fase di compattazione tabellare; ora il suo encoder generic-profile lossless e senza dipendenze è incluso direttamente come codec Headroom (MIT, con marcatura SPDX), insieme ai successivi fix di correttezza per dominio numerico e discrepanze nei conteggi.
Compattazione dell'output Bash + profili MCP — ha ispirato la nostra disciplina di bail-out nella compressione e la riduzione del manifest dei tool MCP.
Compressione dell'output consapevole del contenuto e del tipo di file, con bail-out in caso di errore — ha validato il nostro dispatch per tipo e lo skip basato sul guadagno minimo.
JSON colonnare in Rust + retrieve content-addressed + deduplica cross-message — ha validato il design dei nostri motori headroom/ccr/session-dedup e l'invariante cache-stable "la forma compressa è indipendente dalla posizione".
Toolkit per la TypeScript Compiler API — ha ispirato la nostra rimozione dei commenti basata su parser, che preserva stringhe, template e literal regex.
Intercettazione/analisi MITM del traffico coding-assistant ↔ LLM — il nostro Traffic Inspector adatta il suo merge SSE, la normalizzazione delle conversazioni, il passthrough degli host e il masking dei segreti (MIT).
Routing proxy trasparente per processo — ha ispirato il teardown MITM crash-safe, gli idle timeout dei socket, l'attribuzione dei processi tramite /proc e la cattura TPROXY.
Una raccolta curata di librerie secure-by-default che guida le nostre scelte di sicurezza (Helmet.js, DOMPurify, ssrf-req-filter, safe-regex, Google Tink).
+
+
+### 🧭 Strumenti complementari
+
+
+
Progetto
⭐
Come ha ispirato OmniRoute
+
+
+## 📄 Licenza
+
+Licenza MIT - vedi [LICENSE](../../../LICENSE) per i dettagli.
---
- Built with ❤️ for developers who code 24/7
-
- omniroute.online
+
+**[⬆ Torna all'inizio](#-omniroute)** · Realizzato con ❤️ per la community AI open source.
+
+OmniRoute v3.8.49 · Node ≥22.22.2 · Licenza MIT · omniroute.online
+