--- title: "Troubleshooting" version: 3.8.49 lastUpdated: 2026-07-15 --- # Troubleshooting > **For Users**: Looking for quick fixes? See the [Quick Reference](#quick-reference) below. ๐ŸŒ **Languages:** ๐Ÿ‡บ๐Ÿ‡ธ [English](./TROUBLESHOOTING.md) | ๐Ÿ‡ง๐Ÿ‡ท [Portuguรชs (Brasil)](../i18n/pt-BR/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ช๐Ÿ‡ธ [Espaรฑol](../i18n/es/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ซ๐Ÿ‡ท [Franรงais](../i18n/fr/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฎ๐Ÿ‡น [Italiano](../i18n/it/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ท๐Ÿ‡บ [ะ ัƒััะบะธะน](../i18n/ru/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡จ๐Ÿ‡ณ [ไธญๆ–‡ (็ฎ€ไฝ“)](../i18n/zh-CN/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฉ๐Ÿ‡ช [Deutsch](../i18n/de/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฎ๐Ÿ‡ณ [เคนเคฟเคจเฅเคฆเฅ€](../i18n/in/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡น๐Ÿ‡ญ [เน„เธ—เธข](../i18n/th/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡บ๐Ÿ‡ฆ [ะฃะบั€ะฐั—ะฝััŒะบะฐ](../i18n/uk-UA/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ธ๐Ÿ‡ฆ [ุงู„ุนุฑุจูŠุฉ](../i18n/ar/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฏ๐Ÿ‡ต [ๆ—ฅๆœฌ่ชž](../i18n/ja/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ป๐Ÿ‡ณ [Tiแบฟng Viแป‡t](../i18n/vi/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ง๐Ÿ‡ฌ [ะ‘ัŠะปะณะฐั€ัะบะธ](../i18n/bg/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฉ๐Ÿ‡ฐ [Dansk](../i18n/da/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ซ๐Ÿ‡ฎ [Suomi](../i18n/fi/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฎ๐Ÿ‡ฑ [ืขื‘ืจื™ืช](../i18n/he/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ญ๐Ÿ‡บ [Magyar](../i18n/hu/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฎ๐Ÿ‡ฉ [Bahasa Indonesia](../i18n/id/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฐ๐Ÿ‡ท [ํ•œ๊ตญ์–ด](../i18n/ko/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ฒ๐Ÿ‡พ [Bahasa Melayu](../i18n/ms/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ณ๐Ÿ‡ฑ [Nederlands](../i18n/nl/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ณ๐Ÿ‡ด [Norsk](../i18n/no/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ต๐Ÿ‡น [Portuguรชs (Portugal)](../i18n/pt/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ท๐Ÿ‡ด [Romรขnฤƒ](../i18n/ro/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ต๐Ÿ‡ฑ [Polski](../i18n/pl/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ธ๐Ÿ‡ฐ [Slovenฤina](../i18n/sk/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ธ๐Ÿ‡ช [Svenska](../i18n/sv/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡ต๐Ÿ‡ญ [Filipino](../i18n/phi/docs/guides/TROUBLESHOOTING.md) | ๐Ÿ‡จ๐Ÿ‡ฟ [ฤŒeลกtina](../i18n/cs/docs/guides/TROUBLESHOOTING.md) Common problems and solutions for OmniRoute. --- ## Quick Reference **New to OmniRoute?** Start here โ€” these solve 90% of problems: | I see this | What it means | What to do | | ----------------------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- | | "Can't connect" | OmniRoute isn't running | Run `omniroute` or `docker restart omniroute` | | "Invalid API key" | Your key is wrong or expired | Re-copy the key from the provider's website | | "Rate limit exceeded" | You're sending too many requests | Wait 1 minute, or use `model: "auto"` for automatic fallback | | "Quota exceeded" | You've used up your free/paid quota | Connect more providers, or use free providers (Kiro, Pollinations) | | "Slow responses" | Provider is busy or far away | Use `model: "auto/fast"` or connect a faster provider (Groq, Cerebras) | | "Wrong provider used" | `auto` picked a different provider | That's normal! `auto` picks the best one. Force a specific provider with `model: "openai/gpt-4o"` | | "502 Bad Gateway" | Provider is down | Wait and retry, or use `model: "auto"` to switch providers | | "401 Unauthorized" | Your credentials are wrong | Check your API key or re-authenticate with OAuth | | "429 Too Many Requests" | Rate limited | Wait 1 minute, or connect more providers | **Still stuck?** See the [detailed troubleshooting](#detailed-troubleshooting) below, or ask on [Discord](https://discord.gg/U47eFqAXCn). --- ## Detailed Troubleshooting --- ## Quick Fixes | Problem | Solution | | --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | First login not working | Set `INITIAL_PASSWORD` in `.env` (no hardcoded default) | | Dashboard opens on wrong port | Set `PORT=20128` and `NEXT_PUBLIC_BASE_URL=http://localhost:20128` | | No logs written to disk | Set `APP_LOG_TO_FILE=true` and verify call log capture is enabled | | EACCES: permission denied | Set `DATA_DIR=/path/to/writable/dir` to override `~/.omniroute` | | Routing strategy not saving | Update to the latest v3.x release (Zod schema fix for settings persistence shipped in earlier versions) | | Login crash / blank page | Check Node.js version โ€” see [Node.js Compatibility](#nodejs-compatibility) below | | `dlopen` / `slice is not valid mach-o file` (macOS) | Run `cd $(npm root -g)/omniroute/app && npm rebuild better-sqlite3 && omniroute` โ€” see [macOS native module rebuild](#macos-native-module-rebuild) below | | Proxy "fetch failed" | Ensure proxy config is set at the correct level โ€” see [Proxy Issues](#proxy-issues) below | | Docker `curl: (56) Recv failure: Connection reset by peer` | Your Docker port bind may be landing on IPv6. Use `-p 127.0.0.1:20128:20128` to force IPv4, or test with `curl -4`. See [Docker IPv6](#docker-ipv6) below | | Antivirus quarantines `README.md` | False positive โ€” see [Antivirus false positives](#antivirus-false-positives) below | | Kaspersky flags the Desktop app as a Trojan | Behavioral false positive on the unsigned installer โ€” see [Antivirus false positives](#antivirus-false-positives) below | --- ## Antivirus False Positives ### Avast/AVG quarantine `README.md` with `MD:HttpRequest-inf[Susp]` **This is a false positive. Nothing is infected, and no action is required.** Avast and AVG run a heuristic that flags plain-text/Markdown files containing many HTTP-request-looking links. OmniRoute's `README.md` ships inside the npm package (it is listed in `package.json` โ†’ `files`), so it lands at `node_modules/omniroute/README.md` on a global install โ€” and it contains ~15 `http://localhost:20128/...` examples (the MCP HTTP/SSE endpoints, the A2A `.well-known` URL, and `curl` snippets). That link density is enough to trip the heuristic. If this started only recently: the file did not change in kind. The README grew its endpoints table (MCP HTTP + SSE + A2A were added) and more `curl` examples, which pushed it past the threshold. The file is inert documentation with zero executable content. You can safely restore it from quarantine. **What to do:** 1. **Stop the notifications** โ€” exclude the install directory in your antivirus (Avast: Settings โ†’ Exceptions), adding your global `node_modules` path and/or the OmniRoute data dir (`~/.omniroute/`). 2. **Report the false positive** โ€” , attaching the quarantined `README.md`. This is the fix that helps everyone, since it is the vendor's heuristic overreacting to a text file. **Why we do not "fix" this on our side:** the examples are all `http://localhost`, and localhost cannot be `https` without self-signed-certificate friction. Mangling the docs to dodge one vendor's heuristic would hurt every reader to satisfy a scanner bug. ### Kaspersky flags the Desktop app as `PDM:Trojan.Win32.Generic` **This is a false positive from a behavioral heuristic. Nothing is infected.** Kaspersky's `PDM:` prefix means the verdict comes from its Proactive Defense Module (System Watcher), which judges what the installer *does* rather than matching it against known malware. When it fires, Kaspersky "rolls back" the whole installation โ€” deleting files it had already written โ€” so the app ends up broken or missing. The files it flags are stock parts of declared, open-source dependencies bundled with the desktop app, for example: - `resources/app/.build/next/node_modules/playwright-/lib/โ€ฆ/agentParser.js` and `workerProcessEntry.js` โ€” [Playwright](https://playwright.dev), the browser-automation library used for in-app provider login and browser-backed chat. - `resources/app/.build/next/node_modules/tls-client-node-/bin/tls-client-windows-64-.dll` โ€” the native binary from `tls-client-node`, used for Cloudflare-tolerant HTTP on some web providers. **Why it fires:** the Windows installer is **not yet code-signed**, so an unsigned NSIS installer has zero reputation and behavioral heuristics run at maximum aggression. Combined with a bundled native DLL and hundreds of `.js` files written under `%LOCALAPPDATA%\Programs\OmniRoute` (including hash-suffixed package directories from the Next.js standalone build), that is enough to trip the heuristic. Code signing is planned; until it lands, new releases can repeat this. **What to do:** 1. **Verify your download first** (rules out a tampered file). Every release publishes `latest.yml`, whose `sha512` field (base64) covers the `OmniRoute.Setup..exe` installer. In PowerShell, from the folder containing the installer: ```powershell $b = [System.Security.Cryptography.SHA512]::Create().ComputeHash( [System.IO.File]::ReadAllBytes("$PWD\OmniRoute.Setup..exe")) [Convert]::ToBase64String($b) ``` The output must match `latest.yml` โ†’ `sha512`. If it does not, delete the file and re-download only from the [GitHub releases page](https://github.com/diegosouzapw/OmniRoute/releases). 2. **Restore + exclude** โ€” restore the rolled-back items from quarantine and add an exclusion for `%LOCALAPPDATA%\Programs\OmniRoute` (Kaspersky โ†’ Settings โ†’ Threats and Exclusions), then reinstall. 3. **Report the false positive** โ€” . User-submitted FP reports genuinely speed up allowlisting. --- ## Node.js Compatibility ### Login page crashes or shows "Module self-registration" error **Cause:** You are running a Node.js version outside OmniRoute's approved secure runtime floor. The most common case is running an older Node 22 or 24 patch level that falls below the patched security floor OmniRoute requires. **Symptoms:** - Login page shows a blank screen or a server error - Console shows `Error: Module did not self-register` or similar native binding errors - The login page shows an **orange warning banner** with your Node version if the runtime is outside the supported secure policy **Fix:** 1. Install a supported Node.js LTS release (recommended: Node.js 24.x): ```bash nvm install 24 nvm use 24 ``` 2. Verify your version: `node --version` should show `v24.0.0` or newer on the 24.x LTS line 3. Reinstall OmniRoute: `npm install -g omniroute` 4. Restart: `omniroute` > **Supported secure versions:** `>=22.22.2 <23` or `>=24.0.0 <27`. Node.js 24.x LTS (Krypton) and Node.js 26 are fully supported. ### macOS: `dlopen` / "slice is not valid mach-o file" **Cause:** After a global `npm install -g omniroute`, the `better-sqlite3` native binary inside the package may have been compiled for a different architecture or Node.js ABI than what is running locally. This is common on macOS (both Apple Silicon and Intel) when the pre-built binary does not match your environment. **Symptoms:** - Server fails immediately on startup with a `dlopen` error - Error contains `slice is not valid mach-o file` - Full example: ``` dlopen(/Users//.nvm/versions/node/v24.14.1/lib/node_modules/omniroute/app/node_modules/better-sqlite3/build/Release/better_sqlite3.node, 0x0001): tried: '...' (slice is not valid mach-o file) ``` **Fix โ€” rebuild for your local environment (no Node.js downgrade required):** ```bash cd $(npm root -g)/omniroute/app npm rebuild better-sqlite3 omniroute ``` > **Note:** This recompiles the native binding against your local Node.js version and CPU architecture, resolving the binary mismatch. The officially supported runtime range is **`>=22.22.2 <23` or `>=24.0.0 <27`** (`SUPPORTED_NODE_RANGE` in `src/shared/utils/nodeRuntimeSupport.ts`, aligned with the `package.json` `engines` field). Node.js 24.x LTS (Krypton) and Node.js 26 are fully supported with `better-sqlite3` v12.x. --- ## Proxy Issues ### Provider validation shows "fetch failed" **Cause:** The API key validation endpoint (`POST /api/providers/validate`) was previously bypassing proxy configuration, causing failures in environments that require proxy routing. **Fix (v3.5.5+):** This is now fixed. Provider validation routes through `runWithProxyContext`, honoring provider-level and global proxy settings automatically. ### Token health check fails with "fetch failed" **Cause:** Background OAuth token refresh was not resolving proxy configuration per connection. **Fix (v3.5.5+):** The token health check scheduler now resolves proxy config per connection before attempting refresh. Update to v3.5.5+. ### SOCKS5 proxy returns "invalid onRequestStart method" **Cause:** On Node.js 22, the undici@8 dispatcher is incompatible with Node's built-in `fetch()` implementation. **Fix (v3.5.5+):** OmniRoute now uses undici's own `fetch()` function when a proxy dispatcher is active, ensuring consistent behavior. Update to v3.5.5+. ### MITM proxy under WSL: desktop apps on the Windows host are not intercepted **Cause:** The MITM proxy and its CA certificate install into the environment where OmniRoute runs. Under WSL that environment is the Linux guest, while the AI desktop apps (Kiro, Trae, Copilot, Zed, โ€ฆ) run on the Windows host. The host apps do not trust the guest's certificate store and do not route through the guest's system proxy, so desktop interception does not engage there. **Recommendation:** Run OmniRoute natively on the same OS as the desktop apps you want to intercept (Windows for Windows apps; macOS/Linux likewise). Keeping OmniRoute inside WSL while targeting host apps requires manually trusting the generated CA certificate on the Windows host and pointing each host app's network/proxy settings at the WSL proxy endpoint โ€” an unsupported, fragile setup. --- ## Provider Issues ### "Language model did not provide messages" **Cause:** Provider quota exhausted. **Fix:** 1. Check dashboard quota tracker 2. Use a combo with fallback tiers 3. Switch to cheaper/free tier ### Rate Limiting **Cause:** Subscription quota exhausted. **Fix:** - Add fallback: `cc/claude-opus-4-6 โ†’ glm/glm-4.7 โ†’ if/qwen3.8-max-preview` - Use GLM/MiniMax as cheap backup ### OAuth Token Expired OmniRoute auto-refreshes tokens. If issues persist: 1. Dashboard โ†’ Provider โ†’ Reconnect 2. Delete and re-add the provider connection ### Kiro multi-account: second account invalidates the first **Cause:** Kiro's backend enforces a single active session per OIDC client registration. When two accounts share the same registered client (connections imported before v3.8.0), refreshing one account's token invalidates the other's refresh token. **Fix (v3.8.0+):** Re-import affected connections. Starting with v3.8.0, every new Kiro connection created via **Import Token**, **Google/GitHub social login**, or **Auto-Import** automatically registers its own dedicated OIDC client. The connection is therefore fully isolated and refreshing one account has no effect on any other account. Connections that were imported _before_ v3.8.0 do not carry a per-connection client registration. Those connections continue to use the shared social-auth refresh endpoint. To gain isolation, delete the old connection from Dashboard โ†’ Providers and re-add it via any of the three import flows. For full details and step-by-step instructions for adding two Kiro accounts side by side, see [`docs/guides/KIRO_SETUP.md`](./KIRO_SETUP.md). --- ## Cloud Issues ### Cloud Sync Errors 1. Verify `BASE_URL` points to your running instance (e.g., `http://localhost:20128`) 2. Verify `CLOUD_URL` points to your cloud endpoint (e.g., `https://omniroute.dev`) 3. Keep `NEXT_PUBLIC_*` values aligned with server-side values ### Cloud `stream=false` Returns 500 **Symptom:** `Unexpected token 'd'...` on cloud endpoint for non-streaming calls. **Cause:** Upstream returns SSE payload while client expects JSON. **Workaround:** Use `stream=true` for cloud direct calls. Local runtime includes SSEโ†’JSON fallback. ### Cloud Says Connected but "Invalid API key" 1. Create a fresh key from local dashboard (`/api/keys`) 2. Run cloud sync: Enable Cloud โ†’ Sync Now 3. Old/non-synced keys can still return `401` on cloud --- ## Docker Issues ### Docker IPv6 / Connection Reset **Symptoms:** `curl http://localhost:20128/v1/models` returns `curl: (56) Recv failure: Connection reset by peer`. Dashboard and unauthenticated endpoints work, but authenticated endpoints fail โ€” it looks like an auth problem but isn't. **Cause:** `docker run -p 20128:20128` publishes on both `0.0.0.0` (IPv4) and `::` (IPv6), but the process inside the container listens on IPv4 only. On hosts where `localhost` resolves to `::1` first, the connection lands on the IPv6 published port with no listener behind it โ†’ connection reset. **Fix:** 1. **Quick diagnostic:** Run `curl -4 http://localhost:20128/v1/models`. If it works with `-4` but fails without, you have an IPv6 bind mismatch. 2. **Permanent fix:** Bind to IPv4 explicitly by using `-p 127.0.0.1:20128:20128` in your `docker run` command: ```bash docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \ -p 127.0.0.1:20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest ``` This forces the IPv4 bind and also avoids exposing the proxy on all host interfaces. --- ### CLI Tool Shows Not Installed 1. Check runtime fields: `curl http://localhost:20128/api/cli-tools/runtime/codex | jq` 2. For portable mode: use image target `runner-cli` (bundled CLIs) 3. For host mount mode: set `CLI_EXTRA_PATHS` and mount host bin directory as read-only 4. If `installed=true` and `runnable=false`: binary was found but failed healthcheck ### Quick Runtime Validation ```bash curl -s http://localhost:20128/api/cli-tools/codex-settings | jq '{installed,runnable,commandPath,runtimeMode,reason}' curl -s http://localhost:20128/api/cli-tools/claude-settings | jq '{installed,runnable,commandPath,runtimeMode,reason}' curl -s http://localhost:20128/api/cli-tools/openclaw-settings | jq '{installed,runnable,commandPath,runtimeMode,reason}' ``` --- ## Cost Issues ### High Costs 1. Check usage stats in Dashboard โ†’ Usage 2. Switch primary model to GLM/MiniMax 3. Use free tier (Qoder, Kiro) for non-critical tasks 4. Set cost budgets per API key: Dashboard โ†’ API Keys โ†’ Budget --- ## Debugging ### Enable Log Files Set `APP_LOG_TO_FILE=true` in your `.env` file. Application logs are written under `logs/`. Request artifacts are stored under `${DATA_DIR}/call_logs/` when the call log pipeline is enabled in settings. When pipeline capture is enabled, set `CALL_LOG_PIPELINE_CAPTURE_STREAM_CHUNKS=false` to omit stream chunk payloads, or tune `CALL_LOG_PIPELINE_MAX_SIZE_KB` to change the artifact cap in KB. ### Check Provider Health ```bash # Health dashboard http://localhost:20128/dashboard/health # API health check curl http://localhost:20128/api/monitoring/health ``` ### Runtime Storage - Main state: `${DATA_DIR}/storage.sqlite` (providers, combos, aliases, keys, settings) - Usage: SQLite tables in `storage.sqlite` (`usage_history`, `call_logs`, `proxy_logs`) + optional `${DATA_DIR}/call_logs/` - Application logs: `/logs/...` (when `APP_LOG_TO_FILE=true`) - Call log artifacts: `${DATA_DIR}/call_logs/YYYY-MM-DD/...` when the call log pipeline is enabled The Request Logs page's **Clean history** action clears `call_logs`, legacy `request_detail_logs`, and the local `${DATA_DIR}/call_logs/` artifact directory. --- ## Circuit Breaker Issues ### Provider stuck in OPEN state When a provider's circuit breaker is OPEN, requests are blocked until the cooldown expires. **Fix:** 1. Go to **Dashboard โ†’ Settings โ†’ Resilience** 2. Check the circuit breaker card for the affected provider 3. Click **Reset All** to clear all breakers, or wait for the cooldown to expire 4. Verify the provider is actually available before resetting ### Provider keeps tripping the circuit breaker If a provider repeatedly enters OPEN state: 1. Check **Dashboard โ†’ Health โ†’ Provider Health** for the failure pattern 2. Go to **Settings โ†’ Resilience โ†’ Provider Profiles** and increase the failure threshold 3. Check if the provider has changed API limits or requires re-authentication 4. Review latency telemetry โ€” high latency may cause timeout-based failures --- ## Audio Transcription Issues ### "Unsupported model" error - Ensure you're using the correct prefix: `deepgram/nova-3` or `assemblyai/best` - Verify the provider is connected in **Dashboard โ†’ Providers** ### Transcription returns empty or fails - Check supported audio formats: `mp3`, `wav`, `m4a`, `flac`, `ogg`, `webm` - Verify file size is within provider limits (typically < 25MB) - Check provider API key validity in the provider card --- ## Translator Debugging Use **Dashboard โ†’ Translator** to debug format translation issues: | Mode | When to Use | | ---------------- | -------------------------------------------------------------------------------------------- | | **Playground** | Compare input/output formats side by side โ€” paste a failing request to see how it translates | | **Chat Tester** | Send live messages and inspect the full request/response payload including headers | | **Test Bench** | Run batch tests across format combinations to find which translations are broken | | **Live Monitor** | Watch real-time request flow to catch intermittent translation issues | ### Common format issues - **Thinking tags not appearing** โ€” Check if the target provider supports thinking and the thinking budget setting - **Tool calls dropping** โ€” Some format translations may strip unsupported fields; verify in Playground mode - **System prompt missing** โ€” Claude and Gemini handle system prompts differently; check translation output - **SDK returns raw string instead of object** โ€” Resolved in v1.x; response sanitizer strips non-standard fields (`x_groq`, `usage_breakdown`, etc.) that cause OpenAI SDK Pydantic validation failures. If you still see this on v3.x+, please file an issue. - **GLM/ERNIE rejects `system` role** โ€” Resolved in v1.x; role normalizer automatically merges system messages into user messages for incompatible models. If you still see this on v3.x+, please file an issue. - **`developer` role not recognized** โ€” Resolved in v1.x; automatically converted to `system` for non-OpenAI providers. If you still see this on v3.x+, please file an issue. - **`json_schema` not working with Gemini** โ€” Resolved in v1.x; `response_format` is now converted to Gemini's `responseMimeType` + `responseSchema`. If you still see this on v3.x+, please file an issue. --- ## Resilience Settings ### Auto rate-limit not triggering - Auto rate-limit only applies to API key providers (not OAuth/subscription) - Verify **Settings โ†’ Resilience โ†’ Provider Profiles** has auto-rate-limit enabled - Check if the provider returns `429` status codes or `Retry-After` headers ### Tuning exponential backoff Provider profiles support these settings: - **Base delay** โ€” Initial wait time after first failure (default: 1s) - **Max delay** โ€” Maximum wait time cap (default: 30s) - **Multiplier** โ€” How much to increase delay per consecutive failure (default: 2x) ### Anti-thundering herd When many concurrent requests hit a rate-limited provider, OmniRoute uses mutex + auto rate-limiting to serialize requests and prevent cascading failures. This is automatic for API key providers. --- ## Optional RAG / LLM failure taxonomy (16 problems) Some OmniRoute users place the gateway in front of RAG or agent stacks. In those setups it is common to see a strange pattern: OmniRoute looks healthy (providers up, routing profiles ok, no rate limit alerts) but the final answer is still wrong. In practice these incidents usually come from the downstream RAG pipeline, not from the gateway itself. If you want a shared vocabulary to describe those failures you can use the WFGY ProblemMap, an external MIT license text resource that defines sixteen recurring RAG / LLM failure patterns. At a high level it covers: - retrieval drift and broken context boundaries - empty or stale indexes and vector stores - embedding versus semantic mismatch - prompt assembly and context window issues - logic collapse and overconfident answers - long chain and agent coordination failures - multi agent memory and role drift - deployment and bootstrap ordering problems The idea is simple: 1. When you investigate a bad response, capture: - user task and request - route or provider combo in OmniRoute - any RAG context used downstream (retrieved documents, tool calls, etc) 2. Map the incident to one or two WFGY ProblemMap numbers (`No.1` โ€ฆ `No.16`). 3. Store the number in your own dashboard, runbook, or incident tracker next to the OmniRoute logs. 4. Use the corresponding WFGY page to decide whether you need to change your RAG stack, retriever, or routing strategy. Full text and concrete recipes live here (MIT license, text only): [WFGY ProblemMap README](https://github.com/onestardao/WFGY/blob/main/ProblemMap/README.md) You can ignore this section if you do not run RAG or agent pipelines behind OmniRoute. --- ## v3.8.0 Known Issues Issues specific to the v3.8.0 release and their current workarounds. If a fix lands in a later patch, the entry will be updated or removed. ### Windsurf OAuth flow fails with 401 **Symptoms:** - "401 unauthorized" while completing the Windsurf OAuth flow from the dashboard - Windsurf provider card stays in "needs reconnection" state after the callback **Causes:** - `WINDSURF_FIREBASE_API_KEY` env var missing or empty - `WINDSURF_API_KEY` misconfigured or pointing at a stale token - Local firewall/proxy blocking the OAuth callback **Fix:** 1. Verify both `WINDSURF_FIREBASE_API_KEY` and `WINDSURF_API_KEY` are set in `.env` 2. Restart OmniRoute so the new env values are picked up 3. Re-run the OAuth flow from **Dashboard โ†’ Providers โ†’ Windsurf โ†’ Reconnect** ### Devin CLI auth failures **Symptoms:** - "Devin CLI not found" or "auth failed" when invoking Devin-backed tools - CLI runtime check reports `installed=false` **Causes:** - `CLI_DEVIN_BIN` points to a path that does not exist - Devin CLI is not installed on the host **Fix:** 1. Install the Devin CLI for your platform 2. Set `CLI_DEVIN_BIN=/usr/local/bin/devin` (or the real path) in `.env` 3. Restart OmniRoute and re-test from **Dashboard โ†’ CLI Tools** ### Model cooldown stuck (manual reset) **Symptoms:** - A model stays listed in cooldown even after the expiration time has passed - Requests still skip the model in combo routing despite the timestamp being in the past **Manual reset:** - **Dashboard:** **Settings โ†’ Model Cooldowns** โ†’ click **Re-enable** on the affected card - **API:** `DELETE /api/resilience/model-cooldowns` with management auth headers ### Command Code provider connection fails with 403 **Symptoms:** - 403 when testing the Command Code provider connection - The provider card shows "unauthorized" after a fresh add **Cause:** The OAuth flow did not complete (callback not received or token not persisted). **Fix:** - Run `omniroute providers` from the CLI to re-trigger the OAuth flow, or - Re-run OAuth from **Dashboard โ†’ Providers โ†’ Command Code โ†’ Reconnect** ### ModelScope returns aggressive 429 cooldowns **Symptoms:** - Very short or immediate cooldowns on ModelScope after a small burst of requests - Combo routing skips ModelScope earlier than expected **Cause:** ModelScope emits provider-specific `Retry-After` headers. v3.8.0 ships dedicated handling for those headers, so older versions misread them as generic rate-limit hints. **Fix:** - Ensure you are on v3.8.0 or later - Verify the `useUpstream429BreakerHints` toggle is enabled under **Settings โ†’ Resilience** ### OMNIROUTE_WS_BRIDGE_SECRET missing in production **Symptoms:** - 401 on every Codex/Responses WebSocket bridge request when running on a remote production host - WebSocket bridge handshake closes immediately after connect **Cause:** The `OMNIROUTE_WS_BRIDGE_SECRET` env var is missing from the production environment. **Fix:** 1. Generate a random secret: `openssl rand -hex 32` 2. Set `OMNIROUTE_WS_BRIDGE_SECRET=` in the production server env (and any client that talks to the bridge) 3. Restart OmniRoute ### Responses API: background mode degraded to synchronous **Symptoms:** - Warning logged: `background mode degraded to synchronous` - A `background: true` request returns a normal synchronous response instead of a background job handle **Cause:** v3.8.0 intentionally degrades `background: true` on the Responses API to synchronous execution while emitting a warning. Full async background execution is a future deliverable. **Fix:** - Adjust the client to call without `background`, or - Wait for a later release that ships full async background mode (track the changelog) --- ## Still Stuck? - **GitHub Issues**: [github.com/diegosouzapw/OmniRoute/issues](https://github.com/diegosouzapw/OmniRoute/issues) - **Architecture**: See [`docs/architecture/ARCHITECTURE.md`](../architecture/ARCHITECTURE.md) for internal details - **API Reference**: See [`docs/reference/API_REFERENCE.md`](../reference/API_REFERENCE.md) for all endpoints - **Health Dashboard**: Check **Dashboard โ†’ Health** for real-time system status - **Translator**: Use **Dashboard โ†’ Translator** to debug format issues