From 03d05e0490742e1c86b9de27acafd61ea2aea5eb Mon Sep 17 00:00:00 2001 From: diegosouzapw Date: Mon, 16 Feb 2026 14:25:52 -0300 Subject: [PATCH] docs: track USER_GUIDE, API_REFERENCE, and TROUBLESHOOTING in git These files were referenced in README but ignored by .gitignore. Added them to the exception list so they appear in the repository. --- .gitignore | 3 + docs/API_REFERENCE.md | 428 +++++++++++++++++++++++++++ docs/TROUBLESHOOTING.md | 211 +++++++++++++ docs/USER_GUIDE.md | 639 ++++++++++++++++++++++++++++++++++++++++ 4 files changed, 1281 insertions(+) create mode 100644 docs/API_REFERENCE.md create mode 100644 docs/TROUBLESHOOTING.md create mode 100644 docs/USER_GUIDE.md diff --git a/.gitignore b/.gitignore index f083d9eb20..f2a0ff389c 100644 --- a/.gitignore +++ b/.gitignore @@ -56,6 +56,9 @@ docs/* !docs/ARCHITECTURE.md !docs/CODEBASE_DOCUMENTATION.md !docs/CONTRIBUTING.md +!docs/USER_GUIDE.md +!docs/API_REFERENCE.md +!docs/TROUBLESHOOTING.md !docs/EXECUTION_CONTEXT_PROVIDER_SYNC.md !docs/TASK_NEBIUS_BACKEND_ENABLEMENT.md !docs/frontend-backend-provider-gap-report.md diff --git a/docs/API_REFERENCE.md b/docs/API_REFERENCE.md new file mode 100644 index 0000000000..392fdd48b3 --- /dev/null +++ b/docs/API_REFERENCE.md @@ -0,0 +1,428 @@ +# API Reference + +Complete reference for all OmniRoute API endpoints. + +--- + +## Table of Contents + +- [Chat Completions](#chat-completions) +- [Embeddings](#embeddings) +- [Image Generation](#image-generation) +- [List Models](#list-models) +- [Compatibility Endpoints](#compatibility-endpoints) +- [Semantic Cache](#semantic-cache) +- [Dashboard & Management](#dashboard--management) +- [Request Processing](#request-processing) +- [Authentication](#authentication) + +--- + +## Chat Completions + +```bash +POST /v1/chat/completions +Authorization: Bearer your-api-key +Content-Type: application/json + +{ + "model": "cc/claude-opus-4-6", + "messages": [ + {"role": "user", "content": "Write a function to..."} + ], + "stream": true +} +``` + +### Custom Headers + +| Header | Direction | Description | +| ------------------------ | --------- | --------------------------------- | +| `X-OmniRoute-No-Cache` | Request | Set to `true` to bypass cache | +| `X-OmniRoute-Progress` | Request | Set to `true` for progress events | +| `Idempotency-Key` | Request | Dedup key (5s window) | +| `X-Request-Id` | Request | Alternative dedup key | +| `X-OmniRoute-Cache` | Response | `HIT` or `MISS` (non-streaming) | +| `X-OmniRoute-Idempotent` | Response | `true` if deduplicated | +| `X-OmniRoute-Progress` | Response | `enabled` if progress tracking on | + +--- + +## Embeddings + +```bash +POST /v1/embeddings +Authorization: Bearer your-api-key +Content-Type: application/json + +{ + "model": "nebius/Qwen/Qwen3-Embedding-8B", + "input": "The food was delicious" +} +``` + +Available providers: Nebius, OpenAI, Mistral, Together AI, Fireworks, NVIDIA. + +```bash +# List all embedding models +GET /v1/embeddings +``` + +--- + +## Image Generation + +```bash +POST /v1/images/generations +Authorization: Bearer your-api-key +Content-Type: application/json + +{ + "model": "openai/dall-e-3", + "prompt": "A beautiful sunset over mountains", + "size": "1024x1024" +} +``` + +Available providers: OpenAI (DALL-E), xAI (Grok Image), Together AI (FLUX), Fireworks AI. + +```bash +# List all image models +GET /v1/images/generations +``` + +--- + +## List Models + +```bash +GET /v1/models +Authorization: Bearer your-api-key + +→ Returns all chat, embedding, and image models + combos in OpenAI format +``` + +--- + +## Compatibility Endpoints + +| Method | Path | Format | +| ------ | --------------------------- | ---------------------- | +| POST | `/v1/chat/completions` | OpenAI | +| POST | `/v1/messages` | Anthropic | +| POST | `/v1/responses` | OpenAI Responses | +| POST | `/v1/embeddings` | OpenAI | +| POST | `/v1/images/generations` | OpenAI | +| GET | `/v1/models` | OpenAI | +| POST | `/v1/messages/count_tokens` | Anthropic | +| GET | `/v1beta/models` | Gemini | +| POST | `/v1beta/models/{...path}` | Gemini generateContent | +| POST | `/v1/api/chat` | Ollama | + +### Dedicated Provider Routes + +```bash +POST /v1/providers/{provider}/chat/completions +POST /v1/providers/{provider}/embeddings +POST /v1/providers/{provider}/images/generations +``` + +The provider prefix is auto-added if missing. Mismatched models return `400`. + +--- + +## Semantic Cache + +```bash +# Get cache stats +GET /api/cache + +# Clear all caches +DELETE /api/cache +``` + +Response example: + +```json +{ + "semanticCache": { + "memorySize": 42, + "memoryMaxSize": 500, + "dbSize": 128, + "hitRate": 0.65 + }, + "idempotency": { + "activeKeys": 3, + "windowMs": 5000 + } +} +``` + +--- + +## Dashboard & Management + +### Authentication + +| Endpoint | Method | Description | +| ----------------------------- | ------- | --------------------- | +| `/api/auth/login` | POST | Login | +| `/api/auth/logout` | POST | Logout | +| `/api/settings/require-login` | GET/PUT | Toggle login required | + +### Provider Management + +| Endpoint | Method | Description | +| ---------------------------- | --------------- | ------------------------ | +| `/api/providers` | GET/POST | List / create providers | +| `/api/providers/[id]` | GET/PUT/DELETE | Manage a provider | +| `/api/providers/[id]/test` | POST | Test provider connection | +| `/api/providers/[id]/models` | GET | List provider models | +| `/api/providers/validate` | POST | Validate provider config | +| `/api/provider-nodes*` | Various | Provider node management | +| `/api/provider-models` | GET/POST/DELETE | Custom models | + +### OAuth Flows + +| Endpoint | Method | Description | +| -------------------------------- | ------- | ----------------------- | +| `/api/oauth/[provider]/[action]` | Various | Provider-specific OAuth | + +### Routing & Config + +| Endpoint | Method | Description | +| --------------------- | -------- | ----------------------------- | +| `/api/models/alias` | GET/POST | Model aliases | +| `/api/models/catalog` | GET | All models by provider + type | +| `/api/combos*` | Various | Combo management | +| `/api/keys*` | Various | API key management | +| `/api/pricing` | GET | Model pricing | + +### Usage & Analytics + +| Endpoint | Method | Description | +| --------------------------- | ------ | -------------------- | +| `/api/usage/history` | GET | Usage history | +| `/api/usage/logs` | GET | Usage logs | +| `/api/usage/request-logs` | GET | Request-level logs | +| `/api/usage/[connectionId]` | GET | Per-connection usage | + +### Settings + +| Endpoint | Method | Description | +| ------------------------------- | ------- | ---------------------- | +| `/api/settings` | GET/PUT | General settings | +| `/api/settings/proxy` | GET/PUT | Network proxy config | +| `/api/settings/proxy/test` | POST | Test proxy connection | +| `/api/settings/ip-filter` | GET/PUT | IP allowlist/blocklist | +| `/api/settings/thinking-budget` | GET/PUT | Reasoning token budget | +| `/api/settings/system-prompt` | GET/PUT | Global system prompt | + +### Monitoring + +| Endpoint | Method | Description | +| ------------------------ | ---------- | ----------------------- | +| `/api/sessions` | GET | Active session tracking | +| `/api/rate-limits` | GET | Per-account rate limits | +| `/api/monitoring/health` | GET | Health check | +| `/api/cache` | GET/DELETE | Cache stats / clear | + +### Cloud Sync + +| Endpoint | Method | Description | +| ---------------------- | ------- | --------------------- | +| `/api/sync/cloud` | Various | Cloud sync operations | +| `/api/sync/initialize` | POST | Initialize sync | +| `/api/cloud/*` | Various | Cloud management | + +### CLI Tools + +| Endpoint | Method | Description | +| ---------------------------------- | ------ | ------------------- | +| `/api/cli-tools/claude-settings` | GET | Claude CLI status | +| `/api/cli-tools/codex-settings` | GET | Codex CLI status | +| `/api/cli-tools/droid-settings` | GET | Droid CLI status | +| `/api/cli-tools/openclaw-settings` | GET | OpenClaw CLI status | +| `/api/cli-tools/runtime/[toolId]` | GET | Generic CLI runtime | + +CLI responses include: `installed`, `runnable`, `command`, `commandPath`, `runtimeMode`, `reason`. + +### Resilience & Rate Limits + +| Endpoint | Method | Description | +| ----------------------- | ------- | ------------------------------- | +| `/api/resilience` | GET/PUT | Get/update resilience profiles | +| `/api/resilience/reset` | POST | Reset circuit breakers | +| `/api/rate-limits` | GET | Per-account rate limit status | +| `/api/rate-limit` | GET | Global rate limit configuration | + +### Evals + +| Endpoint | Method | Description | +| ------------ | -------- | --------------------------------- | +| `/api/evals` | GET/POST | List eval suites / run evaluation | + +### Policies + +| Endpoint | Method | Description | +| --------------- | --------------- | ----------------------- | +| `/api/policies` | GET/POST/DELETE | Manage routing policies | + +### Compliance + +| Endpoint | Method | Description | +| --------------------------- | ------ | ----------------------------- | +| `/api/compliance/audit-log` | GET | Compliance audit log (last N) | + +### v1beta (Gemini-Compatible) + +| Endpoint | Method | Description | +| -------------------------- | ------ | --------------------------------- | +| `/v1beta/models` | GET | List models in Gemini format | +| `/v1beta/models/{...path}` | POST | Gemini `generateContent` endpoint | + +These endpoints mirror Gemini's API format for clients that expect native Gemini SDK compatibility. + +### Internal / System APIs + +| Endpoint | Method | Description | +| --------------- | ------ | ---------------------------------------------------- | +| `/api/init` | GET | Application initialization check (used on first run) | +| `/api/tags` | GET | Ollama-compatible model tags (for Ollama clients) | +| `/api/restart` | POST | Trigger graceful server restart | +| `/api/shutdown` | POST | Trigger graceful server shutdown | + +> **Note:** These endpoints are used internally by the system or for Ollama client compatibility. They are not typically called by end users. + +--- + +## Audio Transcription + +```bash +POST /v1/audio/transcriptions +Authorization: Bearer your-api-key +Content-Type: multipart/form-data +``` + +Transcribe audio files using Deepgram or AssemblyAI. + +**Request:** + +```bash +curl -X POST http://localhost:20128/v1/audio/transcriptions \ + -H "Authorization: Bearer your-api-key" \ + -F "file=@recording.mp3" \ + -F "model=deepgram/nova-3" +``` + +**Response:** + +```json +{ + "text": "Hello, this is the transcribed audio content.", + "task": "transcribe", + "language": "en", + "duration": 12.5 +} +``` + +**Supported providers:** `deepgram/nova-3`, `assemblyai/best`. + +**Supported formats:** `mp3`, `wav`, `m4a`, `flac`, `ogg`, `webm`. + +--- + +## Ollama Compatibility + +For clients that use Ollama's API format: + +```bash +# Chat endpoint (Ollama format) +POST /v1/api/chat + +# Model listing (Ollama format) +GET /api/tags +``` + +Requests are automatically translated between Ollama and internal formats. + +--- + +## Telemetry + +```bash +# Get latency telemetry summary (p50/p95/p99 per provider) +GET /api/telemetry/summary +``` + +**Response:** + +```json +{ + "providers": { + "claudeCode": { "p50": 245, "p95": 890, "p99": 1200, "count": 150 }, + "github": { "p50": 180, "p95": 620, "p99": 950, "count": 320 } + } +} +``` + +--- + +## Budget + +```bash +# Get budget status for all API keys +GET /api/usage/budget + +# Set or update a budget +POST /api/usage/budget +Content-Type: application/json + +{ + "keyId": "key-123", + "limit": 50.00, + "period": "monthly" +} +``` + +--- + +## Model Availability + +```bash +# Get real-time model availability across all providers +GET /api/models/availability + +# Check availability for a specific model +POST /api/models/availability +Content-Type: application/json + +{ + "model": "claude-sonnet-4-5-20250929" +} +``` + +--- + +## Request Processing + +1. Client sends request to `/v1/*` +2. Route handler calls `handleChat`, `handleEmbedding`, `handleAudioTranscription`, or `handleImageGeneration` +3. Model is resolved (direct provider/model or alias/combo) +4. Credentials selected from local DB with account availability filtering +5. For chat: `handleChatCore` — format detection, translation, cache check, idempotency check +6. Provider executor sends upstream request +7. Response translated back to client format (chat) or returned as-is (embeddings/images/audio) +8. Usage/logging recorded +9. Fallback applies on errors according to combo rules + +Full architecture reference: [`ARCHITECTURE.md`](ARCHITECTURE.md) + +--- + +## Authentication + +- Dashboard routes (`/dashboard/*`) use `auth_token` cookie +- Login uses saved password hash; fallback to `INITIAL_PASSWORD` +- `requireLogin` toggleable via `/api/settings/require-login` +- `/v1/*` routes optionally require Bearer API key when `REQUIRE_API_KEY=true` diff --git a/docs/TROUBLESHOOTING.md b/docs/TROUBLESHOOTING.md new file mode 100644 index 0000000000..365d991df4 --- /dev/null +++ b/docs/TROUBLESHOOTING.md @@ -0,0 +1,211 @@ +# Troubleshooting + +Common problems and solutions for OmniRoute. + +--- + +## Quick Fixes + +| Problem | Solution | +| ----------------------------- | ------------------------------------------------------------------ | +| First login not working | Check `INITIAL_PASSWORD` in `.env` (default: `123456`) | +| Dashboard opens on wrong port | Set `PORT=20128` and `NEXT_PUBLIC_BASE_URL=http://localhost:20128` | +| No request logs under `logs/` | Set `ENABLE_REQUEST_LOGS=true` | + +--- + +## Provider Issues + +### "Language model did not provide messages" + +**Cause:** Provider quota exhausted. + +**Fix:** + +1. Check dashboard quota tracker +2. Use a combo with fallback tiers +3. Switch to cheaper/free tier + +### Rate Limiting + +**Cause:** Subscription quota exhausted. + +**Fix:** + +- Add fallback: `cc/claude-opus-4-6 → glm/glm-4.7 → if/kimi-k2-thinking` +- Use GLM/MiniMax as cheap backup + +### OAuth Token Expired + +OmniRoute auto-refreshes tokens. If issues persist: + +1. Dashboard → Provider → Reconnect +2. Delete and re-add the provider connection + +--- + +## Cloud Issues + +### Cloud Sync Errors + +1. Verify `BASE_URL` points to your running instance (e.g., `http://localhost:20128`) +2. Verify `CLOUD_URL` points to your cloud endpoint (e.g., `https://omniroute.dev`) +3. Keep `NEXT_PUBLIC_*` values aligned with server-side values + +### Cloud `stream=false` Returns 500 + +**Symptom:** `Unexpected token 'd'...` on cloud endpoint for non-streaming calls. + +**Cause:** Upstream returns SSE payload while client expects JSON. + +**Workaround:** Use `stream=true` for cloud direct calls. Local runtime includes SSE→JSON fallback. + +### Cloud Says Connected but "Invalid API key" + +1. Create a fresh key from local dashboard (`/api/keys`) +2. Run cloud sync: Enable Cloud → Sync Now +3. Old/non-synced keys can still return `401` on cloud + +--- + +## Docker Issues + +### CLI Tool Shows Not Installed + +1. Check runtime fields: `curl http://localhost:20128/api/cli-tools/runtime/codex | jq` +2. For portable mode: use image target `runner-cli` (bundled CLIs) +3. For host mount mode: set `CLI_EXTRA_PATHS` and mount host bin directory as read-only +4. If `installed=true` and `runnable=false`: binary was found but failed healthcheck + +### Quick Runtime Validation + +```bash +curl -s http://localhost:20128/api/cli-tools/codex-settings | jq '{installed,runnable,commandPath,runtimeMode,reason}' +curl -s http://localhost:20128/api/cli-tools/claude-settings | jq '{installed,runnable,commandPath,runtimeMode,reason}' +curl -s http://localhost:20128/api/cli-tools/openclaw-settings | jq '{installed,runnable,commandPath,runtimeMode,reason}' +``` + +--- + +## Cost Issues + +### High Costs + +1. Check usage stats in Dashboard → Usage +2. Switch primary model to GLM/MiniMax +3. Use free tier (Gemini CLI, iFlow) for non-critical tasks +4. Set cost budgets per API key: Dashboard → API Keys → Budget + +--- + +## Debugging + +### Enable Request Logs + +Set `ENABLE_REQUEST_LOGS=true` in your `.env` file. Logs appear under `logs/` directory. + +### Check Provider Health + +```bash +# Health dashboard +http://localhost:20128/dashboard/health + +# API health check +curl http://localhost:20128/api/monitoring/health +``` + +### Runtime Storage + +- Main state: `${DATA_DIR}/db.json` (providers, combos, aliases, keys, settings) +- Usage: `${DATA_DIR}/usage.json`, `${DATA_DIR}/log.txt`, `${DATA_DIR}/call_logs/` +- Request logs: `/logs/...` (when `ENABLE_REQUEST_LOGS=true`) + +--- + +## Circuit Breaker Issues + +### Provider stuck in OPEN state + +When a provider's circuit breaker is OPEN, requests are blocked until the cooldown expires. + +**Fix:** + +1. Go to **Dashboard → Settings → Resilience** +2. Check the circuit breaker card for the affected provider +3. Click **Reset All** to clear all breakers, or wait for the cooldown to expire +4. Verify the provider is actually available before resetting + +### Provider keeps tripping the circuit breaker + +If a provider repeatedly enters OPEN state: + +1. Check **Dashboard → Health → Provider Health** for the failure pattern +2. Go to **Settings → Resilience → Provider Profiles** and increase the failure threshold +3. Check if the provider has changed API limits or requires re-authentication +4. Review latency telemetry — high latency may cause timeout-based failures + +--- + +## Audio Transcription Issues + +### "Unsupported model" error + +- Ensure you're using the correct prefix: `deepgram/nova-3` or `assemblyai/best` +- Verify the provider is connected in **Dashboard → Providers** + +### Transcription returns empty or fails + +- Check supported audio formats: `mp3`, `wav`, `m4a`, `flac`, `ogg`, `webm` +- Verify file size is within provider limits (typically < 25MB) +- Check provider API key validity in the provider card + +--- + +## Translator Debugging + +Use **Dashboard → Translator** to debug format translation issues: + +| Mode | When to Use | +| ---------------- | -------------------------------------------------------------------------------------------- | +| **Playground** | Compare input/output formats side by side — paste a failing request to see how it translates | +| **Chat Tester** | Send live messages and inspect the full request/response payload including headers | +| **Test Bench** | Run batch tests across format combinations to find which translations are broken | +| **Live Monitor** | Watch real-time request flow to catch intermittent translation issues | + +### Common format issues + +- **Thinking tags not appearing** — Check if the target provider supports thinking and the thinking budget setting +- **Tool calls dropping** — Some format translations may strip unsupported fields; verify in Playground mode +- **System prompt missing** — Claude and Gemini handle system prompts differently; check translation output + +--- + +## Resilience Settings + +### Auto rate-limit not triggering + +- Auto rate-limit only applies to API key providers (not OAuth/subscription) +- Verify **Settings → Resilience → Provider Profiles** has auto-rate-limit enabled +- Check if the provider returns `429` status codes or `Retry-After` headers + +### Tuning exponential backoff + +Provider profiles support these settings: + +- **Base delay** — Initial wait time after first failure (default: 1s) +- **Max delay** — Maximum wait time cap (default: 30s) +- **Multiplier** — How much to increase delay per consecutive failure (default: 2x) + +### Anti-thundering herd + +When many concurrent requests hit a rate-limited provider, OmniRoute uses mutex + auto rate-limiting to serialize requests and prevent cascading failures. This is automatic for API key providers. + +--- + +## Still Stuck? + +- **GitHub Issues**: [github.com/diegosouzapw/OmniRoute/issues](https://github.com/diegosouzapw/OmniRoute/issues) +- **Architecture**: See [`docs/ARCHITECTURE.md`](ARCHITECTURE.md) for internal details +- **API Reference**: See [`docs/API_REFERENCE.md`](API_REFERENCE.md) for all endpoints +- **Health Dashboard**: Check **Dashboard → Health** for real-time system status +- **Translator**: Use **Dashboard → Translator** to debug format issues diff --git a/docs/USER_GUIDE.md b/docs/USER_GUIDE.md new file mode 100644 index 0000000000..492e644f67 --- /dev/null +++ b/docs/USER_GUIDE.md @@ -0,0 +1,639 @@ +# User Guide + +Complete guide for configuring providers, creating combos, integrating CLI tools, and deploying OmniRoute. + +--- + +## Table of Contents + +- [Pricing at a Glance](#-pricing-at-a-glance) +- [Use Cases](#-use-cases) +- [Provider Setup](#-provider-setup) +- [CLI Integration](#-cli-integration) +- [Deployment](#-deployment) +- [Available Models](#-available-models) +- [Advanced Features](#-advanced-features) + +--- + +## 💰 Pricing at a Glance + +| Tier | Provider | Cost | Quota Reset | Best For | +| ------------------- | ----------------- | ----------- | ---------------- | -------------------- | +| **💳 SUBSCRIPTION** | Claude Code (Pro) | $20/mo | 5h + weekly | Already subscribed | +| | Codex (Plus/Pro) | $20-200/mo | 5h + weekly | OpenAI users | +| | Gemini CLI | **FREE** | 180K/mo + 1K/day | Everyone! | +| | GitHub Copilot | $10-19/mo | Monthly | GitHub users | +| **🔑 API KEY** | DeepSeek | Pay per use | None | Cheap reasoning | +| | Groq | Pay per use | None | Ultra-fast inference | +| | xAI (Grok) | Pay per use | None | Grok 4 reasoning | +| | Mistral | Pay per use | None | EU-hosted models | +| | Perplexity | Pay per use | None | Search-augmented | +| | Together AI | Pay per use | None | Open-source models | +| | Fireworks AI | Pay per use | None | Fast FLUX images | +| | Cerebras | Pay per use | None | Wafer-scale speed | +| | Cohere | Pay per use | None | Command R+ RAG | +| | NVIDIA NIM | Pay per use | None | Enterprise models | +| **💰 CHEAP** | GLM-4.7 | $0.6/1M | Daily 10AM | Budget backup | +| | MiniMax M2.1 | $0.2/1M | 5-hour rolling | Cheapest option | +| | Kimi K2 | $9/mo flat | 10M tokens/mo | Predictable cost | +| **🆓 FREE** | iFlow | $0 | Unlimited | 8 models free | +| | Qwen | $0 | Unlimited | 3 models free | +| | Kiro | $0 | Unlimited | Claude free | + +**💡 Pro Tip:** Start with Gemini CLI (180K free/month) + iFlow (unlimited free) combo = $0 cost! + +--- + +## 🎯 Use Cases + +### Case 1: "I have Claude Pro subscription" + +**Problem:** Quota expires unused, rate limits during heavy coding + +``` +Combo: "maximize-claude" + 1. cc/claude-opus-4-6 (use subscription fully) + 2. glm/glm-4.7 (cheap backup when quota out) + 3. if/kimi-k2-thinking (free emergency fallback) + +Monthly cost: $20 (subscription) + ~$5 (backup) = $25 total +vs. $20 + hitting limits = frustration +``` + +### Case 2: "I want zero cost" + +**Problem:** Can't afford subscriptions, need reliable AI coding + +``` +Combo: "free-forever" + 1. gc/gemini-3-flash (180K free/month) + 2. if/kimi-k2-thinking (unlimited free) + 3. qw/qwen3-coder-plus (unlimited free) + +Monthly cost: $0 +Quality: Production-ready models +``` + +### Case 3: "I need 24/7 coding, no interruptions" + +**Problem:** Deadlines, can't afford downtime + +``` +Combo: "always-on" + 1. cc/claude-opus-4-6 (best quality) + 2. cx/gpt-5.2-codex (second subscription) + 3. glm/glm-4.7 (cheap, resets daily) + 4. minimax/MiniMax-M2.1 (cheapest, 5h reset) + 5. if/kimi-k2-thinking (free unlimited) + +Result: 5 layers of fallback = zero downtime +Monthly cost: $20-200 (subscriptions) + $10-20 (backup) +``` + +### Case 4: "I want FREE AI in OpenClaw" + +**Problem:** Need AI assistant in messaging apps, completely free + +``` +Combo: "openclaw-free" + 1. if/glm-4.7 (unlimited free) + 2. if/minimax-m2.1 (unlimited free) + 3. if/kimi-k2-thinking (unlimited free) + +Monthly cost: $0 +Access via: WhatsApp, Telegram, Slack, Discord, iMessage, Signal... +``` + +--- + +## 📖 Provider Setup + +### 🔐 Subscription Providers + +#### Claude Code (Pro/Max) + +```bash +Dashboard → Providers → Connect Claude Code +→ OAuth login → Auto token refresh +→ 5-hour + weekly quota tracking + +Models: + cc/claude-opus-4-6 + cc/claude-sonnet-4-5-20250929 + cc/claude-haiku-4-5-20251001 +``` + +**Pro Tip:** Use Opus for complex tasks, Sonnet for speed. OmniRoute tracks quota per model! + +#### OpenAI Codex (Plus/Pro) + +```bash +Dashboard → Providers → Connect Codex +→ OAuth login (port 1455) +→ 5-hour + weekly reset + +Models: + cx/gpt-5.2-codex + cx/gpt-5.1-codex-max +``` + +#### Gemini CLI (FREE 180K/month!) + +```bash +Dashboard → Providers → Connect Gemini CLI +→ Google OAuth +→ 180K completions/month + 1K/day + +Models: + gc/gemini-3-flash-preview + gc/gemini-2.5-pro +``` + +**Best Value:** Huge free tier! Use this before paid tiers. + +#### GitHub Copilot + +```bash +Dashboard → Providers → Connect GitHub +→ OAuth via GitHub +→ Monthly reset (1st of month) + +Models: + gh/gpt-5 + gh/claude-4.5-sonnet + gh/gemini-3-pro +``` + +### 💰 Cheap Providers + +#### GLM-4.7 (Daily reset, $0.6/1M) + +1. Sign up: [Zhipu AI](https://open.bigmodel.cn/) +2. Get API key from Coding Plan +3. Dashboard → Add API Key: Provider: `glm`, API Key: `your-key` + +**Use:** `glm/glm-4.7` — **Pro Tip:** Coding Plan offers 3× quota at 1/7 cost! Reset daily 10:00 AM. + +#### MiniMax M2.1 (5h reset, $0.20/1M) + +1. Sign up: [MiniMax](https://www.minimax.io/) +2. Get API key → Dashboard → Add API Key + +**Use:** `minimax/MiniMax-M2.1` — **Pro Tip:** Cheapest option for long context (1M tokens)! + +#### Kimi K2 ($9/month flat) + +1. Subscribe: [Moonshot AI](https://platform.moonshot.ai/) +2. Get API key → Dashboard → Add API Key + +**Use:** `kimi/kimi-latest` — **Pro Tip:** Fixed $9/month for 10M tokens = $0.90/1M effective cost! + +### 🆓 FREE Providers + +#### iFlow (8 FREE models) + +```bash +Dashboard → Connect iFlow → OAuth login → Unlimited usage + +Models: if/kimi-k2-thinking, if/qwen3-coder-plus, if/glm-4.7, if/minimax-m2, if/deepseek-r1 +``` + +#### Qwen (3 FREE models) + +```bash +Dashboard → Connect Qwen → Device code auth → Unlimited usage + +Models: qw/qwen3-coder-plus, qw/qwen3-coder-flash +``` + +#### Kiro (Claude FREE) + +```bash +Dashboard → Connect Kiro → AWS Builder ID or Google/GitHub → Unlimited + +Models: kr/claude-sonnet-4.5, kr/claude-haiku-4.5 +``` + +--- + +## 🎨 Combos + +### Example 1: Maximize Subscription → Cheap Backup + +``` +Dashboard → Combos → Create New + +Name: premium-coding +Models: + 1. cc/claude-opus-4-6 (Subscription primary) + 2. glm/glm-4.7 (Cheap backup, $0.6/1M) + 3. minimax/MiniMax-M2.1 (Cheapest fallback, $0.20/1M) + +Use in CLI: premium-coding +``` + +### Example 2: Free-Only (Zero Cost) + +``` +Name: free-combo +Models: + 1. gc/gemini-3-flash-preview (180K free/month) + 2. if/kimi-k2-thinking (unlimited) + 3. qw/qwen3-coder-plus (unlimited) + +Cost: $0 forever! +``` + +--- + +## 🔧 CLI Integration + +### Cursor IDE + +``` +Settings → Models → Advanced: + OpenAI API Base URL: http://localhost:20128/v1 + OpenAI API Key: [from omniroute dashboard] + Model: cc/claude-opus-4-6 +``` + +### Claude Code + +Edit `~/.claude/config.json`: + +```json +{ + "anthropic_api_base": "http://localhost:20128/v1", + "anthropic_api_key": "your-omniroute-api-key" +} +``` + +### Codex CLI + +```bash +export OPENAI_BASE_URL="http://localhost:20128" +export OPENAI_API_KEY="your-omniroute-api-key" +codex "your prompt" +``` + +### OpenClaw + +Edit `~/.openclaw/openclaw.json`: + +```json +{ + "agents": { + "defaults": { + "model": { "primary": "omniroute/if/glm-4.7" } + } + }, + "models": { + "providers": { + "omniroute": { + "baseUrl": "http://localhost:20128/v1", + "apiKey": "your-omniroute-api-key", + "api": "openai-completions", + "models": [{ "id": "if/glm-4.7", "name": "glm-4.7" }] + } + } + } +} +``` + +**Or use Dashboard:** CLI Tools → OpenClaw → Auto-config + +### Cline / Continue / RooCode + +``` +Provider: OpenAI Compatible +Base URL: http://localhost:20128/v1 +API Key: [from dashboard] +Model: cc/claude-opus-4-6 +``` + +--- + +## 🚀 Deployment + +### VPS Deployment + +```bash +git clone https://github.com/diegosouzapw/OmniRoute.git +cd OmniRoute && npm install && npm run build + +export JWT_SECRET="your-secure-secret-change-this" +export INITIAL_PASSWORD="your-password" +export DATA_DIR="/var/lib/omniroute" +export PORT="20128" +export HOSTNAME="0.0.0.0" +export NODE_ENV="production" +export NEXT_PUBLIC_BASE_URL="http://localhost:20128" +export API_KEY_SECRET="endpoint-proxy-api-key-secret" + +npm run start +# Or: pm2 start npm --name omniroute -- start +``` + +### Docker + +```bash +# Build image (default = runner-cli with codex/claude/droid preinstalled) +docker build -t omniroute:cli . + +# Portable mode (recommended) +docker run -d --name omniroute -p 20128:20128 --env-file ./.env -v omniroute-data:/app/data omniroute:cli +``` + +For host-integrated mode with CLI binaries, see the Docker section in the main docs. + +### Environment Variables + +| Variable | Default | Description | +| --------------------- | ------------------------------------ | ------------------------------------------------------- | +| `JWT_SECRET` | `omniroute-default-secret-change-me` | JWT signing secret (**change in production**) | +| `INITIAL_PASSWORD` | `123456` | First login password | +| `DATA_DIR` | `~/.omniroute` | Data directory (db, usage, logs) | +| `PORT` | framework default | Service port (`20128` in examples) | +| `HOSTNAME` | framework default | Bind host (Docker defaults to `0.0.0.0`) | +| `NODE_ENV` | runtime default | Set `production` for deploy | +| `BASE_URL` | `http://localhost:20128` | Server-side internal base URL | +| `CLOUD_URL` | `https://omniroute.dev` | Cloud sync endpoint base URL | +| `API_KEY_SECRET` | `endpoint-proxy-api-key-secret` | HMAC secret for generated API keys | +| `REQUIRE_API_KEY` | `false` | Enforce Bearer API key on `/v1/*` | +| `ENABLE_REQUEST_LOGS` | `false` | Enables request/response logs | +| `AUTH_COOKIE_SECURE` | `false` | Force `Secure` auth cookie (behind HTTPS reverse proxy) | + +For the full environment variable reference, see the [README](../README.md). + +--- + +## 📊 Available Models + +
+View all available models + +**Claude Code (`cc/`)** — Pro/Max: `cc/claude-opus-4-6`, `cc/claude-sonnet-4-5-20250929`, `cc/claude-haiku-4-5-20251001` + +**Codex (`cx/`)** — Plus/Pro: `cx/gpt-5.2-codex`, `cx/gpt-5.1-codex-max` + +**Gemini CLI (`gc/`)** — FREE: `gc/gemini-3-flash-preview`, `gc/gemini-2.5-pro` + +**GitHub Copilot (`gh/`)**: `gh/gpt-5`, `gh/claude-4.5-sonnet` + +**GLM (`glm/`)** — $0.6/1M: `glm/glm-4.7` + +**MiniMax (`minimax/`)** — $0.2/1M: `minimax/MiniMax-M2.1` + +**iFlow (`if/`)** — FREE: `if/kimi-k2-thinking`, `if/qwen3-coder-plus`, `if/deepseek-r1` + +**Qwen (`qw/`)** — FREE: `qw/qwen3-coder-plus`, `qw/qwen3-coder-flash` + +**Kiro (`kr/`)** — FREE: `kr/claude-sonnet-4.5`, `kr/claude-haiku-4.5` + +**DeepSeek (`ds/`)**: `ds/deepseek-chat`, `ds/deepseek-reasoner` + +**Groq (`groq/`)**: `groq/llama-3.3-70b-versatile`, `groq/llama-4-maverick-17b-128e-instruct` + +**xAI (`xai/`)**: `xai/grok-4`, `xai/grok-4-0709-fast-reasoning`, `xai/grok-code-mini` + +**Mistral (`mistral/`)**: `mistral/mistral-large-2501`, `mistral/codestral-2501` + +**Perplexity (`pplx/`)**: `pplx/sonar-pro`, `pplx/sonar` + +**Together AI (`together/`)**: `together/meta-llama/Llama-3.3-70B-Instruct-Turbo` + +**Fireworks AI (`fireworks/`)**: `fireworks/accounts/fireworks/models/deepseek-v3p1` + +**Cerebras (`cerebras/`)**: `cerebras/llama-3.3-70b` + +**Cohere (`cohere/`)**: `cohere/command-r-plus-08-2024` + +**NVIDIA NIM (`nvidia/`)**: `nvidia/nvidia/llama-3.3-70b-instruct` + +
+ +--- + +## 🧩 Advanced Features + +### Custom Models + +Add any model ID to any provider without waiting for an app update: + +```bash +# Via API +curl -X POST http://localhost:20128/api/provider-models \ + -H "Content-Type: application/json" \ + -d '{"provider": "openai", "modelId": "gpt-4.5-preview", "modelName": "GPT-4.5 Preview"}' + +# List: curl http://localhost:20128/api/provider-models?provider=openai +# Remove: curl -X DELETE "http://localhost:20128/api/provider-models?provider=openai&model=gpt-4.5-preview" +``` + +Or use Dashboard: **Providers → [Provider] → Custom Models**. + +### Dedicated Provider Routes + +Route requests directly to a specific provider with model validation: + +```bash +POST http://localhost:20128/v1/providers/openai/chat/completions +POST http://localhost:20128/v1/providers/openai/embeddings +POST http://localhost:20128/v1/providers/fireworks/images/generations +``` + +The provider prefix is auto-added if missing. Mismatched models return `400`. + +### Network Proxy Configuration + +```bash +# Set global proxy +curl -X PUT http://localhost:20128/api/settings/proxy \ + -d '{"global": {"type":"http","host":"proxy.example.com","port":"8080"}}' + +# Per-provider proxy +curl -X PUT http://localhost:20128/api/settings/proxy \ + -d '{"providers": {"openai": {"type":"socks5","host":"proxy.example.com","port":"1080"}}}' + +# Test proxy +curl -X POST http://localhost:20128/api/settings/proxy/test \ + -d '{"proxy":{"type":"socks5","host":"proxy.example.com","port":"1080"}}' +``` + +**Precedence:** Key-specific → Combo-specific → Provider-specific → Global → Environment. + +### Model Catalog API + +```bash +curl http://localhost:20128/api/models/catalog +``` + +Returns models grouped by provider with types (`chat`, `embedding`, `image`). + +### Cloud Sync + +- Sync providers, combos, and settings across devices +- Automatic background sync with timeout + fail-fast +- Prefer server-side `BASE_URL`/`CLOUD_URL` in production + +### LLM Gateway Intelligence (Phase 9) + +- **Semantic Cache** — Auto-caches non-streaming, temperature=0 responses (bypass with `X-OmniRoute-No-Cache: true`) +- **Request Idempotency** — Deduplicates requests within 5s via `Idempotency-Key` or `X-Request-Id` header +- **Progress Tracking** — Opt-in SSE `event: progress` events via `X-OmniRoute-Progress: true` header + +--- + +### Translator Playground + +Access via **Dashboard → Translator**. Debug and visualize how OmniRoute translates API requests between providers. + +| Mode | Purpose | +| ---------------- | -------------------------------------------------------------------------------------- | +| **Playground** | Select source/target formats, paste a request, and see the translated output instantly | +| **Chat Tester** | Send live chat messages through the proxy and inspect the full request/response cycle | +| **Test Bench** | Run batch tests across multiple format combinations to verify translation correctness | +| **Live Monitor** | Watch real-time translations as requests flow through the proxy | + +**Use cases:** + +- Debug why a specific client/provider combination fails +- Verify that thinking tags, tool calls, and system prompts translate correctly +- Compare format differences between OpenAI, Claude, Gemini, and Responses API formats + +--- + +### Routing Strategies + +Configure via **Dashboard → Settings → Routing**. + +| Strategy | Description | +| ------------------------------ | ------------------------------------------------------------------------------------------------ | +| **Fill First** | Uses accounts in priority order — primary account handles all requests until unavailable | +| **Round Robin** | Cycles through all accounts with a configurable sticky limit (default: 3 calls per account) | +| **P2C (Power of Two Choices)** | Picks 2 random accounts and routes to the healthier one — balances load with awareness of health | + +#### Wildcard Model Aliases + +Create wildcard patterns to remap model names: + +``` +Pattern: claude-sonnet-* → Target: cc/claude-sonnet-4-5-20250929 +Pattern: gpt-* → Target: gh/gpt-5.1-codex +``` + +Wildcards support `*` (any characters) and `?` (single character). + +#### Fallback Chains + +Define global fallback chains that apply across all requests: + +``` +Chain: production-fallback + 1. cc/claude-opus-4-6 + 2. gh/gpt-5.1-codex + 3. glm/glm-4.7 +``` + +--- + +### Resilience & Circuit Breakers + +Configure via **Dashboard → Settings → Resilience**. + +OmniRoute implements provider-level resilience with three components: + +1. **Circuit Breaker** — Tracks failures per provider and automatically opens the circuit when a threshold is reached: + - **CLOSED** (Healthy) — Requests flow normally + - **OPEN** — Provider is temporarily blocked after repeated failures + - **HALF_OPEN** — Testing if provider has recovered + +2. **Provider Profiles** — Per-provider configuration for: + - Failure threshold (how many failures before opening) + - Cooldown duration + - Rate limit detection sensitivity + - Exponential backoff parameters + +3. **Rate Limit Auto-Detection** — Monitors `429` and `Retry-After` headers to proactively avoid hitting provider rate limits. + +**Pro Tip:** Use **Reset All** button to clear all circuit breakers and cooldowns when a provider recovers from an outage. + +--- + +### Costs & Budget Management + +Access via **Dashboard → Costs**. + +| Tab | Purpose | +| ----------- | ---------------------------------------------------------------------------------------- | +| **Budget** | Set spending limits per API key with daily/weekly/monthly budgets and real-time tracking | +| **Pricing** | View and edit model pricing entries — cost per 1K input/output tokens per provider | + +```bash +# API: Set a budget +curl -X POST http://localhost:20128/api/usage/budget \ + -H "Content-Type: application/json" \ + -d '{"keyId": "key-123", "limit": 50.00, "period": "monthly"}' + +# API: Get current budget status +curl http://localhost:20128/api/usage/budget +``` + +**Cost Tracking:** Every request logs token usage and calculates cost using the pricing table. View breakdowns in **Dashboard → Usage** by provider, model, and API key. + +--- + +### Audio Transcription + +OmniRoute supports audio transcription via the OpenAI-compatible endpoint: + +```bash +POST /v1/audio/transcriptions +Authorization: Bearer your-api-key +Content-Type: multipart/form-data + +# Example with curl +curl -X POST http://localhost:20128/v1/audio/transcriptions \ + -H "Authorization: Bearer your-api-key" \ + -F "file=@audio.mp3" \ + -F "model=deepgram/nova-3" +``` + +Available providers: **Deepgram** (`deepgram/`), **AssemblyAI** (`assemblyai/`). + +Supported audio formats: `mp3`, `wav`, `m4a`, `flac`, `ogg`, `webm`. + +--- + +### Combo Balancing Strategies + +Configure per-combo balancing in **Dashboard → Combos → Create/Edit → Strategy**. + +| Strategy | Description | +| ------------------ | ------------------------------------------------------------------------ | +| **Round-Robin** | Rotates through models sequentially | +| **Priority** | Always tries the first model; falls back only on error | +| **Random** | Picks a random model from the combo for each request | +| **Weighted** | Routes proportionally based on assigned weights per model | +| **Least-Used** | Routes to the model with the fewest recent requests (uses combo metrics) | +| **Cost-Optimized** | Routes to the cheapest available model (uses pricing table) | + +Global combo defaults can be set in **Dashboard → Settings → Routing → Combo Defaults**. + +--- + +### Health Dashboard + +Access via **Dashboard → Health**. Real-time system health overview with 6 cards: + +| Card | What It Shows | +| --------------------- | ----------------------------------------------------------- | +| **System Status** | Uptime, version, memory usage, data directory | +| **Provider Health** | Per-provider circuit breaker state (Closed/Open/Half-Open) | +| **Rate Limits** | Active rate limit cooldowns per account with remaining time | +| **Active Lockouts** | Providers temporarily blocked by the lockout policy | +| **Signature Cache** | Deduplication cache stats (active keys, hit rate) | +| **Latency Telemetry** | p50/p95/p99 latency aggregation per provider | + +**Pro Tip:** The Health page auto-refreshes every 10 seconds. Use the circuit breaker card to identify which providers are experiencing issues.