23 KiB
title, version, lastUpdated
| title | version | lastUpdated |
|---|---|---|
| Resilience Guide | 3.8.40 | 2026-06-28 |
Resilience Guide
OmniRoute has three distinct but related resilience mechanisms. Each has a different scope and purpose. Keep them separate when debugging routing behavior.
Source: diagrams/resilience-3layers.mmd
1. Provider Circuit Breaker
Scope: entire provider (e.g., glm, openai, anthropic).
Purpose: stop sending traffic to a provider that is repeatedly failing at the upstream/service level.
Implementation:
- Core class:
src/shared/utils/circuitBreaker.ts - Wiring:
src/sse/handlers/chatHelpers.ts,src/sse/handlers/chat.ts - Status API:
GET /api/monitoring/health - Reset API:
POST /api/resilience/reset - Wrappers:
open-sse/services/accountFallback.ts - DB table:
domain_circuit_breakers
States:
CLOSED— normal traffic allowedDEGRADED— traffic still allowed, but elevated provider failures are being trackedOPEN— provider temporarily blocked; combo routing skips itHALF_OPEN— reset timeout elapsed; probe request allowed
Configurable defaults (open-sse/config/constants.ts, exposed in Dashboard → Settings → Resilience):
| Class | Degraded at | Opens at | Reset timeout |
|---|---|---|---|
| OAuth | 5 failures | 8 failures | 60s |
| API-key | 7 failures | 12 failures | 30s |
| Local | derived | 2 failures | 15s |
degradationThreshold controls when a provider enters DEGRADED; failureThreshold controls when it opens and is skipped. Local provider profiles are not exposed on the Resilience settings page yet.
Trip codes: only provider-level statuses [408, 500, 502, 503, 504]. Do NOT trip for account-level errors (most 401/403/429 — those belong to cooldown or lockout).
Lazy recovery: when OPEN expires, getStatus(), canExecute(), getRetryAfterMs() refresh state to HALF_OPEN. No background timer needed.
2. Connection Cooldown
Scope: single provider connection/account/key.
Purpose: skip one bad key while other connections for the same provider keep serving.
Implementation:
- Mark unavailable:
src/sse/services/auth.ts::markAccountUnavailable() - Selection:
getProviderCredentials*in same file - Cooldown calc:
open-sse/services/accountFallback.ts::checkFallbackError() - Settings:
src/lib/resilience/settings.ts
Fields per connection:
rateLimitedUntil— timestamp until cooldown expirestestStatus: "unavailable"lastError,lastErrorType,errorCodebackoffLevel— exponential backoff counter
Default cooldowns:
- OAuth base: 5s
- API-key base: 3s
- API-key 429: prefers upstream
Retry-After/reset headers/parseable reset text - Backoff:
baseCooldownMs * 2 ** failureIndex
Anti-thundering-herd guard: prevents concurrent failures from over-extending cooldown or double-incrementing backoffLevel.
Terminal states (NOT cooldowns):
banned— set by banned-keyword / account-ban detection (see BAN_DETECTION)expiredcredits_exhausted
These persist until credentials change or an operator resets them. Do not overwrite terminal states with transient cooldown state.
Lazy recovery: when rateLimitedUntil is past, connection becomes eligible again. On successful use, clearAccountError() clears all error fields.
Session affinity (#7274)
Scope: one client session (X-Session-Id / x-codex-session-id / x-omniroute-session header) pinned to one connection, for any provider.
Purpose: keep a multi-turn agent (Claude Code, aider, custom agents) on the same account across requests, reducing cross-account context loss and repeated cold-start 429s on providers with per-account session state.
Implementation:
- TTL resolution:
src/sse/services/sessionAffinityPin.ts::resolveSessionAffinityTtlMs() - Pin selection/creation:
src/sse/services/sessionAffinityPin.ts::selectSessionAffinityConnection() - Header extraction (generic, any provider):
src/sse/services/auth.ts::extractSessionAffinityKey() - Persisted pin table:
sessionAccountAffinity(src/lib/db/sessionAccountAffinity.ts) - Setting:
sessionAffinityTtlMs(global TTL in ms,0disables) —src/lib/db/settings.ts. Renamed from the Codex-onlycodexSessionAffinityTtlMsby migration124_generic_session_affinity_ttl.sql, which carries over any previously-configured Codex TTL as the new default.
Before #7274, resolveSessionAffinityTtlMs() hard-bailed to 0 for every provider except codex, so the TTL setting (and the session headers) had no effect anywhere else even though the pinning mechanism and header extraction were already provider-agnostic. The fix removed that early-return; the TTL now applies uniformly to every provider once set globally above 0.
The three session-affinity headers are never forwarded upstream — executors build their own upstream headers from scratch rather than passing client headers through, so this stays an internal correlation id only.
3. Model Lockout
Scope: provider + connection + model triple.
Purpose: avoid disabling a whole connection when only one model is unavailable or quota-limited.
Examples:
- Per-model quota providers returning 429
- Local providers returning 404 for one missing model
- Provider-specific mode/model permission failures (e.g., Grok modes)
Implementation: open-sse/services/accountFallback.ts — lockModel(), clearModelLock(), getAllModelLockouts().
Model Cooldowns Dashboard (v3.8.0)
UI: Settings → Model Cooldowns (src/app/(dashboard)/dashboard/settings/components/ModelCooldownsCard.tsx)
Lists active lockouts with: provider, connection, model, reason, expiresAt. Operators can manually re-enable a model from the card.
REST API:
GET /api/resilience/model-cooldowns— list active lockoutsDELETE /api/resilience/model-cooldowns— manual re-enable. Body:{provider, connection, model}. Auth: management.
Lockout settings UI + success-decay recovery (v3.8.23)
Model lockout went from always-on hardcoded behavior to a fully configurable, opt-in feature with its own settings card and a self-healing recovery path.
Settings card: Settings → Model Lockout
(src/app/(dashboard)/dashboard/settings/components/ModelLockoutCard.tsx).
This is distinct from the read-only ModelCooldownsCard above (which only
lists active lockouts) — the new card configures the parameters. Defaults
live in DEFAULT_MODEL_LOCKOUT_SETTINGS
(src/lib/resilience/modelLockoutSettings.ts):
| Setting | Default | Meaning |
|---|---|---|
enabled |
false |
Master toggle — model lockout is off by default. |
errorCodes |
[403, 404, 429, 502, 503, 504] |
Upstream statuses that count as a model-scoped failure. |
baseCooldownMs |
120_000 (120 s) |
Initial lockout duration for the first failure. |
maxCooldownMs |
1_800_000 (30 min) |
Cap on the escalated cooldown. |
maxBackoffSteps |
10 |
Max exponential-backoff escalation steps. |
useExponentialBackoff |
true |
Whether repeated failures escalate the cooldown exponentially. |
Settings persist through the normal settings store and validate via the
resilience settings schema; the card clamps baseCooldownMs/maxCooldownMs
(with maxCooldownMs ≥ baseCooldownMs) and maxBackoffSteps.
Success-decay recovery: recovery is not purely timer expiry. A healthy
response walks the model's failure count back down so a model that recovered
mid-window stops escalating (and clears) before its timer would. On a successful
combo target, open-sse/services/combo.ts calls decayModelFailureCount()
(open-sse/services/accountFallback.ts), which halves the stored
failureCount (Math.floor(failureCount / 2)); when it reaches 0 the lockout
entry is deleted entirely. The counterpart recordModelLockoutFailure()
increments the count (and escalates the cooldown) on failures within the
escalation window. This success-decay is in addition to plain timer expiry —
either path can re-enable a model.
State: lockouts are held in-memory (per-process Maps of
ModelLockoutEntry keyed by provider:connectionId:model), not persisted to
the DB — they are lost on restart. The settings are persisted; the active
lockout state is ephemeral.
4. Quota-Share Concurrency Control (v3.8.36)
Subscription accounts (GLM, MiniMax, etc.) often accept only ~1–3 concurrent
requests; exceeding that triggers 429s and cooldowns. This is acute under
quota-share (qtSd/…) combos, where several API keys share one upstream
account. Three layers keep a shared account from being flooded.
Per-connection concurrency cap (max_concurrent)
Each provider connection can declare a max_concurrent ceiling
(provider_connections.max_concurrent, set in the connection modal / API / DB).
Leave it empty for no limit. This is the single knob that drives the serialization
layer below — set it to the account's real concurrency (e.g. GLM ~1, MiniMax ~2).
Quota-share request serialization
When a quota-share dispatch targets a connection that declares a positive
max_concurrent, concurrent requests to that account are serialized through a
per-connection semaphore (key qsconn:<connectionId>): excess requests wait in
the queue instead of flooding the account. It is fail-open — a saturated
queue or timeout proceeds without a slot rather than ever rejecting a dispatchable
request. Toggle in Settings → Resilience → Quota-share per-connection
concurrency (resilienceSettings.quotaShareConcurrencyLimit.enabled, default
on). Without a max_concurrent cap the behavior is unchanged.
The quota-share routing gate (
selectQuotaShareTarget, DRR + P2C) is itself fail-open and only deprioritizes an at-cap connection — with a single-connection pool it cannot hard-limit, so this semaphore is what actually contains the flood.
Combo cooldown-aware retry
For every combo strategy (when enabled), a request that would crystallize a 429
for a SHORT transient cooldown waits it out and re-dispatches instead of
returning the 429 — this covers Gemini-class TPM/RPM windows (~60s retry-after)
on multi-model combos, e.g. both targets of a 2-model combo hitting a per-model
rate limit. Bounded by comboCooldownWait (enabled, maxWaitMs, maxAttempts,
budgetMs) in Settings → Resilience. It never waits on quota_exhausted
(locked until midnight) or auth/not-found reasons.
5. Request Queue Admission Control (v3.8.49 · issue #6593)
Scope: the local per-provider+connection rate-limit queue (open-sse/services/rateLimitManager.ts,
backed by Bottleneck), one layer below the three mechanisms above.
maxWaitMs is a legacy persisted name for execution expiration.
resilienceSettings.requestQueue.maxWaitMs is passed to Bottleneck as a job
expiration, whose timer starts only after dispatch. It therefore bounds
limiter-managed execution, not time spent in the local queue. Expiration is
surfaced as trusted local code: "RATE_LIMIT_EXECUTION_TIMEOUT" (HTTP 504);
the former queue-timeout code name is accepted only for trusted internal
backward compatibility. The default is 15000ms; override via
RATE_LIMIT_MAX_WAIT_MS (env) or the dashboard (Settings → Resilience,
1–30000ms UI ceiling). Queue residence has no time deadline; use
maxQueueDepth below to bound queued callers.
maxQueueDepth — opt-in admission cap (new). resilienceSettings.requestQueue.maxQueueDepth
bounds how many requests may sit queued (not yet dispatched) for one
provider+connection at once. When the queue already holds maxQueueDepth
requests, a new request is fast-rejected with a typed
code: "RATE_LIMIT_QUEUE_FULL" error before it ever reaches limiter.schedule()
— so the rejection is cheap and happens ahead of any downstream
prompt-compression / translation work for that request. Default 0 =
disabled, preserving the existing unbounded-queue behavior; bounded 0–100000.
Override via RATE_LIMIT_MAX_QUEUE_DEPTH (env) or
resilienceSettings.requestQueue.maxQueueDepth (dashboard/API patch).
The admission check itself is a pure function
(open-sse/services/rateLimitManager/admission.ts::checkQueueAdmission) so
it is unit-testable without a real Bottleneck limiter.
The RFC that opened #6593 also proposed a
bypassCompressionOnRateLimitflag. This repo'sopen-sse/services/compression/pipeline is prompt/context compression on the outbound LLM request (chatCore.ts, around theresolveCompressionSettings/selectCompressionStrategyblock), not HTTP response compression on synthesized 429 bodies — there is no matching code path for a literal bypass flag. That prompt-compression step also currently runs beforewithRateLimit()in the request pipeline, so reordering to skip it on a queue-full rejection is a separate, larger change than this issue's scope; it was intentionally not implemented here and is left as a follow-up if the CPU-saving win is worth the reordering risk.
6. Slow-stream throughput watchdog (#9709)
The optional resilienceSettings.streamRecovery.throughputWatchdog guard detects
an upstream that is still sending chunks but producing assistant output below the
configured useful-output rate. It is deliberately distinct from the idle timeout:
heartbeats and metadata reset neither timer and do not count as progress. It is also
distinct from the hard attempt deadline (#9153), which remains an absolute safety
ceiling regardless of output quality.
The watchdog requires a warm-up period followed by a complete rolling window before
it can abort. It counts text deltas from Chat Completions and Responses API output
events (a conservative UTF-8 byte proxy), ignores usage-only and empty events, and
suspends judgement while tool-call or reasoning events are in flight. It is disabled
by default and can be enabled with STREAM_THROUGHPUT_WATCHDOG_ENABLED=true; the
window, warm-up, minimum rate, and minimum measurable output are bounded by the
normal resilience-settings normalization layer.
When enabled, a watchdog abort is applied only to the active upstream attempt. Before any client-visible bytes, the existing same-account early-recovery path may reopen the attempt. After commit, the stream is never blindly replayed; only the existing safe mid-stream continuation contract can stitch a suffix. Finalization remains single-shot, so usage accounting and semaphore release are not duplicated.
7. Upstream Status Restatement (misstated quota errors)
Scope: one upstream gateway that reports temporary quota exhaustion with the wrong HTTP status.
Purpose: correct a misleading status BEFORE classification, so downstream consumers (fallback engine, combo aggregation, the client-facing response) see the true retryable nature of the failure.
Some gateways signal TEMPORARY quota exhaustion with a non-retryable HTTP
status. agentrouter.org returns 403 (sometimes 400) with a Chinese body
(用户额度不足 / 额度不足) instead of the standard 429. Clients like Claude
Code treat 403 as permanent and abort the session, and without correction
the fallback engine would classify it as AUTH_ERROR instead of a quota
event.
Implementation:
- Registry + matcher:
open-sse/config/upstreamStatusRestatement.ts— a per-provider list of rules ({id, fromStatuses, toStatus, textMarkers, excludeMarkers, defaultRetryAfterMs}), matched viaapplyStatusRestatement(). - Call site: the
providerFailure:block inopen-sse/handlers/chatCore.ts(around line 3654), right afterparseUpstreamError()parses an upstream response with an error HTTP status (!providerResponse.ok), and before any classification runs, so every downstream consumer sees the corrected status. Errors embedded inside a200SSE stream follow a separate, later stream-parsing path and are not covered by this hook today — a known limitation, not yet needed for agentrouter's misstatus (which surfaces as an error HTTP status). - Retry eligibility:
429is inRETRY_AFTER_ELIGIBLE_STATUSES(open-sse/services/combo/unavailableRetryGate.ts), so a restated error carries a real retry window instead of surfacing as a dead403.
Permanent errors (agentrouter's 无权访问模型 — no access to this model) are
NEVER restated: excludeMarkers vetoes the rule even when textMarkers hit,
so the error keeps its original status and nothing retries it forever. A
separate provider classification rule
(agentrouter-model-access-denied in open-sse/config/providerErrorRules.ts)
locks only the affected model instead (Model Lockout, §3), so the connection
keeps serving other models.
Two-stage design: status restatement, then classification
Status restatement (upstreamStatusRestatement.ts) and provider
classification rules (open-sse/config/providerErrorRules.ts,
providerRuleRegistry) are separate registries that both key on provider id
and text markers, but they run in different places and serve different
purposes: restatement rewrites the HTTP status early in chatCore.ts;
classification rules pick the fallback reason and lock scope
(model / provider / connection) inside checkFallbackError()
(open-sse/services/accountFallback.ts).
Classification rules only see full error text (needed to match body
markers like 额度不足) for providers listed in the FULL_TEXT_RULE_PROVIDERS
allowlist in providerErrorRules.ts — currently only "agentrouter". For
every other provider, checkFallbackError hands getProviderErrorRuleMatch
only the structured error ({code, type}), which is enough for
header/status/code-based rules but blind to body-text markers. The helper
resolveRuleMatchBody() performs this selection: full error text for
allowlisted providers, the structured error otherwise. Adding a provider to
FULL_TEXT_RULE_PROVIDERS is an explicit per-provider opt-in — it exists so
that the default path for every provider not on the list stays
byte-for-byte unchanged.
Adding a new quota-misstating gateway
- Register one rule array in
statusRestatementRegistry(open-sse/config/upstreamStatusRestatement.ts). KeeptextMarkersprovider-specific; never reuse generic English phrases that collide withCREDITS_EXHAUSTED_SIGNALS(open-sse/services/accountFallback.ts). - Optionally register classification rules in
open-sse/config/providerErrorRules.ts(providerRuleRegistry) to pick the right lock scope (connectionfor account-wide quota,modelfor per-model errors). This step only takes effect in production for providers whose rules need the full error text (body markers): add the provider id toFULL_TEXT_RULE_PROVIDERSin the same file — otherwisecheckFallbackErroronly ever hands the rule the structured{code, type}error and a body-text rule will never match live traffic. Rules that match purely onstatus/headers(like Opencode's or Minimax's) do not need this opt-in. - Add unit tests mirroring
tests/unit/upstream-status-restatement.test.tsandtests/unit/agentrouter-error-rules.test.ts(including the not-permanent / not-creditsExhausted guards, and — if the provider needs the allowlist — a test assertingresolveRuleMatchBody()returns the full text only for that provider).
No changes to chatCore.ts, classifyError, or combo are needed.
Other Resilience Features
- 19 routing strategies (priority, weighted, round-robin, context-relay, fill-first, p2c, random, least-used, cost-optimized, reset-aware, reset-window, headroom, strict-random, auto, lkgp, context-optimized, cache-optimized, fusion, pipeline) — see AUTO-COMBO.md.
- Reset-aware routing (v3.8.0) — prioritizes connections by quota reset time.
- Background mode degradation — Responses API
background: truedegraded to sync with warning. - Dynamic tool limit detection — backs off providers when tool count limits hit.
- Emergency fallback — controlled by
OMNIROUTE_EMERGENCY_FALLBACK; operators can override it from the Feature Flags page without a restart.
Debugging
- All keys for a provider skipped → check both circuit breaker state AND each connection's
rateLimitedUntil/testStatus. - Provider permanently excluded after reset window → code reading raw
stateinstead ofgetStatus()/canExecute(). - One key fails, others should work → prefer connection cooldown over circuit breaker.
- Only one model fails → prefer model lockout over connection cooldown.
- State should self-recover but doesn't → check for future timestamp + read path that refreshes expired state. Permanent statuses require manual changes.
TLS Fingerprinting & Stealth
Provider-specific stealth (JA3/JA4, CCH, obfuscation) is separately documented — see STEALTH_GUIDE.md.
Resilience testing (Phase 8 · Block C)
Beyond unit tests for resilience logic, three tests exercise the runtime under real stress/failure conditions (all integration/nightly — none block PRs):
| Test | What | Run |
|---|---|---|
| Chaos | Fake-upstream node injects real latency/reset/timeout/503; validates that the circuit breaker opens/recovers and checkFallbackError classifies 503 as recoverable fallback. |
RUN_CHAOS_INT=1 npm run test:chaos |
| Heap-growth | ~500 streams per createSSEStream under --expose-gc; fails if the heap grows beyond the ceiling (OOM guard #3069). |
npm run test:heap |
| k6 soak | Sustained load against /api/monitoring/health; p95/error thresholds. |
k6 run tests/load/k6-soak.js (nightly) |
Orchestrated by .github/workflows/nightly-resilience.yml (cron + dispatch). In the
default test:integration, chaos and heap self-skip (without RUN_CHAOS_INT/--expose-gc).
See Also
- Architecture Guide — System architecture and internals
- User Guide — Providers, combos, CLI integration
- Auto-Combo Engine — 13-factor scoring, mode packs