mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-07-26 09:52:11 +03:00
* fix(sse): Gemini TPM classification, combo-cooldown-wait for auto/quota-share, and target-timeout floor
Gemini TPM/RPM 429s were misclassified as QUOTA_EXHAUSTED because
sanitizeErrorMessage() truncates to the first line, hiding Google's
metric name and retry hint on lines 2-3. Added a rawMessage field
(internal-only, never reaches the client) and classifyGeminiQuotaMetricFromText()
to classify from the untruncated text, reordered ahead of the generic
credits/daily-quota checks.
Widened comboCooldownWaitEnabled (wait out a short transient cooldown
instead of crystallizing a 429/503) from quota-share-only to also cover
auto-strategy combos, and raised the wait ceiling to 65s/130s-budget/90s-cap
to match Gemini's ~60s TPM/RPM windows.
The per-target timeout (DEFAULT_COMBO_TARGET_TIMEOUT_MS, 120s) was shorter
than the new 130s cooldown-wait budget, so a target could get cut off
mid-wait with a synthetic 524 instead of completing the retry. Added
resolveComboTargetTimeoutMsForCombo()/isComboCooldownWaitEligible() in
comboConfig.ts to raise the per-target floor to budgetMs+buffer only for
wait-eligible strategies (auto/quota-share), verified live: a 12-request
concurrent burst against a TPM-exhausted combo went from 2/12 succeeding
(10 x 524) to 12/12 succeeding with zero 503/524.
Also: liveGeminiShared.ts's sendAndValidate now fails fast on a 503
instead of retrying past it, and the health dashboard + request logger
surface TPM stats alongside RPM/RPD.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): combo-exhausted rejection logs now capture request body + attempted models
recordRejectedRequestUsage() (the fast path for combo requests that never
reach handleChatCore, e.g. all targets locked by resilience cooldown)
hardcoded provider: "-" and never passed a request body to saveCallLog(),
so /dashboard/logs entries for these failures were nearly useless for
debugging: no way to see the client's request or which models were tried.
- recordRejectedRequestUsage() now accepts requestBody and persists it
through the existing saveCallLog() artifact mechanism (same path
handleChatCore's own logging uses).
- Added summarizeComboAttemptedModels(), which reads the combo's own model
list (always available, unlike the response's combo-diagnostics headers —
a model-level resilience-lockout skip never touches the
exhaustedProviders/exhaustedConnections sets those headers are built
from) to populate a real "provider" value instead of "-".
- Wired both into the call site in src/sse/handlers/chat.ts.
NOTE: unrelated to the Gemini TPM/combo-cooldown-wait fix on this branch —
landed here per operator request, to be split into its own branch/PR.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* feat(sse): synthetic streaming keep-alive event + 5-minute Gemini cooldown-wait ceiling
Many clients enforce a first-SSE-byte timeout, which made it unsafe to wait
out a longer upstream rate-limit cooldown on a streaming request — the
client would abandon the connection before any bytes arrived. This landed
in two parts:
1. Synthetic startup "thinking" event (OpenAI chat/completions format):
the already-existing withEarlyStreamKeepalive wrapper (open-sse/utils/
earlyStreamKeepalive.ts, wired into /v1/chat/completions, /v1/messages,
/v1/responses since #2544) opens the SSE stream immediately once a
request runs past its threshold, but only ever sent empty/no-op
keepalive frames. Added a `startupFrame` option (defaults to
`keepaliveFrame` — zero behavior change unless a route opts in) so the
very first frame can carry real content instead. Wired
OPENAI_STARTUP_THINKING_FRAME (a reasoning_content delta: "OmniRoute:
got request, sending to provider") into /v1/chat/completions only —
Claude Messages and Responses API formats both require a preceding
envelope event (message_start / response.created) that a synthetic
pre-dispatch frame can't safely fabricate without risking a duplicate
envelope once the real stream arrives, so those two routes keep their
existing (safe, proven) keepalive frames unchanged.
2. Raised the "wait out a known cooldown, then retry" ceiling to 5 minutes
for both retry mechanisms, now that a client-side first-byte timeout is
no longer a risk on the (opted-in) route:
- comboCooldownWait (auto/quota-share combos, open-sse/services/combo.ts):
maxWaitMs hard clamp raised 90s -> 300s (src/lib/resilience/settings/
normalize.ts); defaults raised to maxWaitMs:90s/maxAttempts:5/
budgetMs:300s. comboConfig.ts's resolveComboTargetTimeoutMsForCombo
already derives the per-target timeout floor from budgetMs, so it
tracks the new ceiling with no further changes.
- waitForCooldown (direct, non-combo model requests, src/sse/handlers/
chat.ts): this mechanism had NO cumulative cap before — only a
per-wait cap (maxRetryWaitMs) and a retry count (maxRetries), so
maxRetries x maxRetryWaitMs could exceed 5 minutes with no ceiling.
Added a budgetMs field (mirrors comboCooldownWait) to
WaitForCooldownSettings/CooldownAwareRetrySettings, threaded a
requestRetryBudgetLeftMs tracker through chat.ts's requestAttemptLoop
(mirrors combo.ts's comboCooldownBudgetLeftMs), and made
getCooldownAwareRetryDecision refuse to wait once the cumulative
budget is exhausted even if the single wait is under maxRetryWaitMs.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): extend the synthetic keep-alive thinking event to /v1/responses
Live incident (OpenClaw, log id 1784407081908-cbc24f): a /v1/responses
request to gemini/gemma-4-31b-it took 56s to produce a first byte and the
client disconnected (499 request_signal_aborted) — the same client-first-byte-
timeout problem the previous commit fixed for /v1/chat/completions, but
/v1/responses only had the generic bare-comment keepalive (no content), so it
wasn't covered.
Added RESPONSES_STARTUP_THINKING_FRAME: a self-contained synthetic reasoning
item (response.output_item.added -> reasoning_summary_part.added ->
reasoning_summary_text.delta -> reasoning_summary_part.done), opened AND
closed within this one frame rather than left dangling — it never carries a
response_id, so it can't collide with the real upstream response's own
independent response.created lifecycle that follows. Mirrors the abbreviated
delta+part.done close pattern open-sse/utils/stream.ts's own
emitSyntheticResponsesReasoningSummary already uses for real mid-stream
reasoning content.
Wired into src/app/api/v1/responses/route.ts via the startupFrame option
added in the previous commit.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): combo cooldown-wait vars reset every setTry, crystallizing a bogus 503 instead of waiting
Live incident (log id 1784416706646-51): a request to the "default" combo
(strategy=auto, maxSetRetries=3) hit a real Gemini TPM 429 on both gemma-4
targets, correctly classified as a short 40s rate_limit lockout — then
crystallized a 503 "all upstream accounts are inactive" in 6.9s instead of
ever reaching the cooldown-aware wait.
Root cause: `lastError`/`earliestRetryAfter`/`lastStatus` were declared with
`let` INSIDE the `for (setTry...)` loop body, so they reset to null at the
start of every set-try. When both targets lock out on setTry 0, every
subsequent setTry (1..maxSetRetries) pre-skips both targets via the
isModelLocked check with no real dispatch — so on the FINAL setTry (the only
one whose values the post-loop decision reads, since it's gated behind
`if (setTry < maxSetRetries) continue`), lastStatus was null, hitting the
"!lastStatus" branch (ALL_ACCOUNTS_INACTIVE 503) and completely bypassing the
comboCooldownWaitEnabled / earliestRetryAfter wait logic — even though a
real 429 with a known ~40s retry-after WAS observed on setTry 0.
This bug predates today's Gemini TPM work (any combo with maxSetRetries > 0
whose targets all lock out on the first pass was affected) but was masked in
existing tests: the "auto strategy (2 models...)" regression test uses
maxSetRetries: 0, so it only ever runs ONE setTry iteration and never
exercises the reset-on-retry path. It also explains why the dedicated
12-concurrent-request burst test passed cleanly — with concurrent requests,
timing variance meant some request's FINAL setTry iteration still had a live
target to dispatch to, giving lastStatus/earliestRetryAfter fresh data. A
single isolated request has no such luck.
Fix: hoist lastError/earliestRetryAfter/lastStatus to just inside
dispatchWithCooldownRetry, before the setTry loop, so they persist across
set-tries (still reset fresh on each recursive dispatchWithCooldownRetry()
call after a wait, which is correct). recordedAttempts/fallbackCount/
exhaustedProviders etc. are intentionally left per-iteration (unrelated to
this bug).
New regression test in tests/unit/combo-quota-share-cooldown-wait.test.ts
reproduces the exact live scenario (2 targets, both lock out on setTry 0,
maxSetRetries: 3) — confirmed red (503) against the pre-fix code, green
(200, waits and retries) against the fix.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* test(sse): extend live Gemini workload to Responses API + add large-context TPM test
Two additions to the live Gemini test suite, both live-verified against the
dev instance:
1. sendAndValidate() (tests/integration/liveGeminiShared.ts) now accepts an
apiFormat: "chat" | "responses" parameter, building the Responses-API
request shape (input array, max_output_tokens) and parsing its SSE events
(response.output_text.delta / response.reasoning_summary_text.delta /
response.completed) via the new readResponsesSSEStream(). Wired into two
new tests in live-gemini-workload.test.ts ([30]/[31]), mirroring the
existing Chat Completions streaming coverage. Verified live: 24/25 + 5/5
payloads succeeded end-to-end through the new code path (the one failure
was a ~300s test-client fetch timeout unrelated to the Responses API code
itself — a separate, not-yet-addressed test-harness limitation).
2. genHugeContextMessage() builds a single message large enough (~4
chars/token estimate) to approach or exceed Gemini's free-tier TPM ceiling
(16000 input tokens/min for gemma-4) by itself. Every other prompt
generator in this file tops out around 1-2k tokens — nowhere near that
ceiling — so none of the existing workload tests ever exercised a REAL TPM
429, only RPM-style rate limiting. tests/integration/gemini-large-context-tpm.test.ts
sends two ~12-13k-token requests back-to-back (comfortably exceeding
16000/min together) to exercise the full path against production Gemini:
TPM classification, the comboCooldownWait retry, and the synthetic
keep-alive frame on a genuinely slow request.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): abandoned combo target dispatch now observes its own per-target timeout, fixing a permanent "pending" dashboard leak
Live incident (dashboard log id 1784418258231-14961a, reported as "an ongoing
request even though there's already a 200"): a combo target dispatch
abandoned by comboTargetTimeoutMs (open-sse/services/combo/targetTimeoutRunner.ts)
left a permanent phantom "pending" entry in the dashboard, even after the
overall combo request had already succeeded via a different retry.
Root cause: chatCore.ts's createStreamController — and everything downstream
that depends on it (withRateLimit's Promise.race against Bottleneck,
acquireAccountSemaphore) — only ever watches clientRawRequest.signal, which
is the ORIGINAL client's request signal (set once via buildClientRawRequest
and reused unchanged across every target dispatch in a combo). It has no
connection to targetTimeoutRunner.ts's OWN AbortController
(target.modelAbortSignal), which is what actually fires when
comboTargetTimeoutMs (300s) elapses. src/sse/handlers/chat.ts's
handleSingleModel bridge between combo.ts and handleSingleModelChat received
`target.modelAbortSignal` but silently dropped it — never forwarded it
anywhere. So when a target got abandoned (e.g. stuck inside a wedged
Bottleneck rate-limiter queue, see the WEDGED force-reset log line from the
same incident), its per-target timeout fired and let the COMBO move on and
retry successfully elsewhere — but the abandoned dispatch's own promise
chain never learned it had been superseded, so it hung forever waiting on a
signal that was never going to fire, and trackPendingRequest(false) (the
finalize call) never ran.
Fix: thread target.modelAbortSignal through as a new modelAbortSignal
runtimeOption, and merge it into clientRawRequest.signal (via the existing
mergeAbortSignals helper from open-sse/executors/base.ts) right before
dispatch, so an abandoned target's own promise chain now observes its abort
and can reach its cleanup path — new resolveDispatchClientRawRequest() makes
this mechanically testable in isolation.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): combo cooldown-wait state recording, rate-limit wedge recovery, OpenAI-format SSE error frames
Five related fixes surfaced by live incidents (dashboard log ids 1784457764961-73,
1784465227489-a2cbc0, 1784504040241-6f8b9a) while validating the Gemini TPM/cooldown-wait
work on this branch against real OpenClaw traffic:
- combo.ts: the model-lockout bail-out branches in dispatchWithCooldownRetry never
recorded lastStatus, so once every target in a set hit an existing lockout the final
check crystallized a bogus ALL_ACCOUNTS_INACTIVE 503 instead of reaching the
cooldown-wait decision, even with a real 429 + short retry-after observed.
- combo.ts/combo/types.ts: the "all credentials cooling down" pre-dispatch rejection
(buildModelCooldownBody) nests its retry hint as error.retry_after/reset_seconds, not
the top-level retryAfter every other 429 shape uses — combo's extraction only read the
latter, so earliestRetryAfter stayed null for this shape even after lastStatus was fixed.
- rateLimitManager.ts: the wedge-recovery watchdog used disconnect(), which releases the
heartbeat timer but never rejects jobs already QUEUED on that instance — orphaned
dispatches hung until the outer ~300s per-target timeout, well past real clients'
patience. Switched to stop({ dropWaitingJobs: true }), safe because the wedge condition
already requires RUNNING===0 && EXECUTING===0.
- earlyStreamKeepalive.ts: the in-band error frame emitted after committing to a 200 SSE
stream was hardcoded to Anthropic's `event: error` convention for every route, including
the OpenAI-format ones (/v1/chat/completions, /v1/responses) where that framing is
either invisible or malformed to a plain data-line parser. Added per-route
OPENAI_CHAT_ERROR_FRAME / OPENAI_RESPONSES_ERROR_FRAME and wired them in.
- chatCore.ts: persisted a synthetic clientResponse error body even when the client had
already disconnected (AbortError) before that body was ever computed — misleading the
dashboard into showing "what the client received" for a response that was never sent.
Also: RequestLoggerDetail.tsx — Provider/Client Event Stream panes lost their collapse
toggle when StreamSection replaced the collapsible PayloadSection (692d6be80, unifying
active/finished request views) without carrying the toggle over.
Each fix has a TDD regression test with a confirmed red-before-green cycle.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* test(sse): free-tier model + gemma-4 TPM-ceiling benchmark harness
Adds a live benchmark comparing free models OmniRoute exposes across
configured providers plus previously-unexercised no-auth providers
(felo-web, aihorde, opencode, duckduckgo-web — none need a connection
row, they were just never tried). Reuses liveGeminiShared.ts's SSE
parsers and CASE_BUILDERS instead of duplicating them.
Also adds a targeted TPM-stress test firing back-to-back large-context
prompts at the gemma-4-31b model across its 3 free hosts (Gemini,
NVIDIA, AI Horde) to isolate whether the documented 16k-tokens/minute
free-tier ceiling is Gemini-specific enforcement or an inherent
model property.
FORCE_TOOL_CHOICE_REQUIRED is a test-only, default-off env flag added
to liveGeminiShared.ts and live-gemini-agentic-loop.test.ts for an
earlier live A/B comparison of tool_choice: required vs unset — kept
as a reusable knob for future runs.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* test(sse): benchmark for the 2026-07-22 newly-enabled provider batch
Adds NEWLY_ENABLED_MODELS to freeModelBenchmarkShared.ts (Mistral
Leanstral, OpenRouter's live "free"-tagged roster, OpenCode Zen's
current free models — refetched live from
https://opencode.ai/zen/v1/models since the static catalog had
drifted) and a dedicated workload benchmark test for them.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* test(sse): sync geminiRateLimitTracker tests with e74a1722b's corrected Gemma 4 limits
e74a1722b updated geminiRateLimits.json's gemma-4-* entries from the stale
15/1500/-1 (rpm/rpd/tpm) to the real published free-tier values
16000/14400/16000, but never updated the tests asserting the old numbers.
Surfaced by running the full test:unit suite as a post-rebase sanity check.
Co-Authored-By: Markus Hartung <markus.hartream@gmail.com>
* chore(quality): file-size baseline for own-growth (#8213)
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
311 lines
16 KiB
Markdown
311 lines
16 KiB
Markdown
---
|
||
title: "Resilience Guide"
|
||
version: 3.8.40
|
||
lastUpdated: 2026-06-28
|
||
---
|
||
|
||
# Resilience Guide
|
||
|
||
OmniRoute has three distinct but related resilience mechanisms. Each has a different scope and purpose. Keep them separate when debugging routing behavior.
|
||
|
||

|
||
|
||
> Source: [diagrams/resilience-3layers.mmd](../diagrams/resilience-3layers.mmd)
|
||
|
||
## 1. Provider Circuit Breaker
|
||
|
||
**Scope:** entire provider (e.g., `glm`, `openai`, `anthropic`).
|
||
|
||
**Purpose:** stop sending traffic to a provider that is repeatedly failing at the upstream/service level.
|
||
|
||
**Implementation:**
|
||
|
||
- Core class: `src/shared/utils/circuitBreaker.ts`
|
||
- Wiring: `src/sse/handlers/chatHelpers.ts`, `src/sse/handlers/chat.ts`
|
||
- Status API: `GET /api/monitoring/health`
|
||
- Reset API: `POST /api/resilience/reset`
|
||
- Wrappers: `open-sse/services/accountFallback.ts`
|
||
- DB table: `domain_circuit_breakers`
|
||
|
||
**States:**
|
||
|
||
- `CLOSED` — normal traffic allowed
|
||
- `DEGRADED` — traffic still allowed, but elevated provider failures are being tracked
|
||
- `OPEN` — provider temporarily blocked; combo routing skips it
|
||
- `HALF_OPEN` — reset timeout elapsed; probe request allowed
|
||
|
||
**Configurable defaults (`open-sse/config/constants.ts`, exposed in Dashboard → Settings → Resilience):**
|
||
|
||
| Class | Degraded at | Opens at | Reset timeout |
|
||
| ------- | ----------- | ----------- | ------------- |
|
||
| OAuth | 5 failures | 8 failures | 60s |
|
||
| API-key | 7 failures | 12 failures | 30s |
|
||
| Local | derived | 2 failures | 15s |
|
||
|
||
`degradationThreshold` controls when a provider enters `DEGRADED`; `failureThreshold` controls when it opens and is skipped. Local provider profiles are not exposed on the Resilience settings page yet.
|
||
|
||
**Trip codes:** only provider-level statuses `[408, 500, 502, 503, 504]`. Do NOT trip for account-level errors (most 401/403/429 — those belong to cooldown or lockout).
|
||
|
||
**Lazy recovery:** when `OPEN` expires, `getStatus()`, `canExecute()`, `getRetryAfterMs()` refresh state to `HALF_OPEN`. No background timer needed.
|
||
|
||
---
|
||
|
||
## 2. Connection Cooldown
|
||
|
||
**Scope:** single provider connection/account/key.
|
||
|
||
**Purpose:** skip one bad key while other connections for the same provider keep serving.
|
||
|
||
**Implementation:**
|
||
|
||
- Mark unavailable: `src/sse/services/auth.ts::markAccountUnavailable()`
|
||
- Selection: `getProviderCredentials*` in same file
|
||
- Cooldown calc: `open-sse/services/accountFallback.ts::checkFallbackError()`
|
||
- Settings: `src/lib/resilience/settings.ts`
|
||
|
||
**Fields per connection:**
|
||
|
||
- `rateLimitedUntil` — timestamp until cooldown expires
|
||
- `testStatus: "unavailable"`
|
||
- `lastError`, `lastErrorType`, `errorCode`
|
||
- `backoffLevel` — exponential backoff counter
|
||
|
||
**Default cooldowns:**
|
||
|
||
- OAuth base: 5s
|
||
- API-key base: 3s
|
||
- API-key 429: prefers upstream `Retry-After`/reset headers/parseable reset text
|
||
- Backoff: `baseCooldownMs * 2 ** failureIndex`
|
||
|
||
**Anti-thundering-herd guard:** prevents concurrent failures from over-extending cooldown or double-incrementing `backoffLevel`.
|
||
|
||
**Terminal states (NOT cooldowns):**
|
||
|
||
- `banned` — set by banned-keyword / account-ban detection (see [BAN_DETECTION](../security/BAN_DETECTION.md))
|
||
- `expired`
|
||
- `credits_exhausted`
|
||
|
||
These persist until credentials change or an operator resets them. Do not overwrite terminal states with transient cooldown state.
|
||
|
||
**Lazy recovery:** when `rateLimitedUntil` is past, connection becomes eligible again. On successful use, `clearAccountError()` clears all error fields.
|
||
|
||
### Session affinity (#7274)
|
||
|
||
**Scope:** one client session (`X-Session-Id` / `x-codex-session-id` / `x-omniroute-session` header) pinned to one connection, for **any** provider.
|
||
|
||
**Purpose:** keep a multi-turn agent (Claude Code, aider, custom agents) on the same account across requests, reducing cross-account context loss and repeated cold-start 429s on providers with per-account session state.
|
||
|
||
**Implementation:**
|
||
|
||
- TTL resolution: `src/sse/services/sessionAffinityPin.ts::resolveSessionAffinityTtlMs()`
|
||
- Pin selection/creation: `src/sse/services/sessionAffinityPin.ts::selectSessionAffinityConnection()`
|
||
- Header extraction (generic, any provider): `src/sse/services/auth.ts::extractSessionAffinityKey()`
|
||
- Persisted pin table: `sessionAccountAffinity` (`src/lib/db/sessionAccountAffinity.ts`)
|
||
- Setting: `sessionAffinityTtlMs` (global TTL in ms, `0` disables) — `src/lib/db/settings.ts`. Renamed from the Codex-only `codexSessionAffinityTtlMs` by migration `124_generic_session_affinity_ttl.sql`, which carries over any previously-configured Codex TTL as the new default.
|
||
|
||
Before #7274, `resolveSessionAffinityTtlMs()` hard-bailed to `0` for every provider except `codex`, so the TTL setting (and the session headers) had no effect anywhere else even though the pinning mechanism and header extraction were already provider-agnostic. The fix removed that early-return; the TTL now applies uniformly to every provider once set globally above `0`.
|
||
|
||
The three session-affinity headers are never forwarded upstream — executors build their own upstream headers from scratch rather than passing client headers through, so this stays an internal correlation id only.
|
||
|
||
---
|
||
|
||
## 3. Model Lockout
|
||
|
||
**Scope:** provider + connection + model triple.
|
||
|
||
**Purpose:** avoid disabling a whole connection when only one model is unavailable or quota-limited.
|
||
|
||
**Examples:**
|
||
|
||
- Per-model quota providers returning 429
|
||
- Local providers returning 404 for one missing model
|
||
- Provider-specific mode/model permission failures (e.g., Grok modes)
|
||
|
||
**Implementation:** `open-sse/services/accountFallback.ts` — `lockModel()`, `clearModelLock()`, `getAllModelLockouts()`.
|
||
|
||
### Model Cooldowns Dashboard (v3.8.0)
|
||
|
||
UI: Settings → Model Cooldowns (`src/app/(dashboard)/dashboard/settings/components/ModelCooldownsCard.tsx`)
|
||
|
||
Lists active lockouts with: provider, connection, model, reason, expiresAt. Operators can manually re-enable a model from the card.
|
||
|
||
**REST API:**
|
||
|
||
- `GET /api/resilience/model-cooldowns` — list active lockouts
|
||
- `DELETE /api/resilience/model-cooldowns` — manual re-enable. Body: `{provider, connection, model}`. Auth: management.
|
||
|
||
### Lockout settings UI + success-decay recovery (v3.8.23)
|
||
|
||
Model lockout went from always-on hardcoded behavior to a fully configurable,
|
||
opt-in feature with its own settings card and a self-healing recovery path.
|
||
|
||
**Settings card:** Settings → Model Lockout
|
||
(`src/app/(dashboard)/dashboard/settings/components/ModelLockoutCard.tsx`).
|
||
This is **distinct** from the read-only `ModelCooldownsCard` above (which only
|
||
_lists_ active lockouts) — the new card _configures the parameters_. Defaults
|
||
live in `DEFAULT_MODEL_LOCKOUT_SETTINGS`
|
||
(`src/lib/resilience/modelLockoutSettings.ts`):
|
||
|
||
| Setting | Default | Meaning |
|
||
| ----------------------- | -------------------------------- | -------------------------------------------------------------- |
|
||
| `enabled` | `false` | Master toggle — model lockout is **off by default**. |
|
||
| `errorCodes` | `[403, 404, 429, 502, 503, 504]` | Upstream statuses that count as a model-scoped failure. |
|
||
| `baseCooldownMs` | `120_000` (120 s) | Initial lockout duration for the first failure. |
|
||
| `maxCooldownMs` | `1_800_000` (30 min) | Cap on the escalated cooldown. |
|
||
| `maxBackoffSteps` | `10` | Max exponential-backoff escalation steps. |
|
||
| `useExponentialBackoff` | `true` | Whether repeated failures escalate the cooldown exponentially. |
|
||
|
||
Settings persist through the normal settings store and validate via the
|
||
resilience settings schema; the card clamps `baseCooldownMs`/`maxCooldownMs`
|
||
(with `maxCooldownMs ≥ baseCooldownMs`) and `maxBackoffSteps`.
|
||
|
||
**Success-decay recovery:** recovery is **not** purely timer expiry. A healthy
|
||
response walks the model's failure count back down so a model that recovered
|
||
mid-window stops escalating (and clears) before its timer would. On a successful
|
||
combo target, `open-sse/services/combo.ts` calls `decayModelFailureCount()`
|
||
(`open-sse/services/accountFallback.ts`), which **halves** the stored
|
||
`failureCount` (`Math.floor(failureCount / 2)`); when it reaches `0` the lockout
|
||
entry is deleted entirely. The counterpart `recordModelLockoutFailure()`
|
||
increments the count (and escalates the cooldown) on failures within the
|
||
escalation window. This success-decay is in addition to plain timer expiry —
|
||
either path can re-enable a model.
|
||
|
||
**State:** lockouts are held **in-memory** (per-process `Map`s of
|
||
`ModelLockoutEntry` keyed by `provider:connectionId:model`), not persisted to
|
||
the DB — they are lost on restart. The _settings_ are persisted; the active
|
||
lockout _state_ is ephemeral.
|
||
|
||
---
|
||
|
||
## 4. Quota-Share Concurrency Control (v3.8.36)
|
||
|
||
Subscription accounts (GLM, MiniMax, etc.) often accept only ~1–3 concurrent
|
||
requests; exceeding that triggers 429s and cooldowns. This is acute under
|
||
**quota-share** (`qtSd/…`) combos, where several API keys share one upstream
|
||
account. Three layers keep a shared account from being flooded.
|
||
|
||
### Per-connection concurrency cap (`max_concurrent`)
|
||
|
||
Each provider connection can declare a `max_concurrent` ceiling
|
||
(`provider_connections.max_concurrent`, set in the connection modal / API / DB).
|
||
Leave it empty for no limit. This is the single knob that drives the serialization
|
||
layer below — set it to the account's real concurrency (e.g. GLM ~1, MiniMax ~2).
|
||
|
||
### Quota-share request serialization
|
||
|
||
When a quota-share dispatch targets a connection that declares a positive
|
||
`max_concurrent`, concurrent requests to that **account** are serialized through a
|
||
per-connection semaphore (key `qsconn:<connectionId>`): excess requests **wait in
|
||
the queue** instead of flooding the account. It is **fail-open** — a saturated
|
||
queue or timeout proceeds without a slot rather than ever rejecting a dispatchable
|
||
request. Toggle in **Settings → Resilience → Quota-share per-connection
|
||
concurrency** (`resilienceSettings.quotaShareConcurrencyLimit.enabled`, default
|
||
on). Without a `max_concurrent` cap the behavior is unchanged.
|
||
|
||
> The quota-share routing gate (`selectQuotaShareTarget`, DRR + P2C) is itself
|
||
> fail-open and only _deprioritizes_ an at-cap connection — with a
|
||
> single-connection pool it cannot hard-limit, so this semaphore is what actually
|
||
> contains the flood.
|
||
|
||
### Combo cooldown-aware retry
|
||
|
||
For quota-share and `auto` combos, a request that would crystallize a 429 for a
|
||
SHORT transient cooldown waits it out and re-dispatches instead of returning
|
||
the 429 — this covers Gemini-class TPM/RPM windows (~60s retry-after) on a
|
||
multi-model `auto` combo, e.g. both targets of a 2-model combo hitting a
|
||
per-model rate limit. Bounded by `comboCooldownWait` (`enabled`, `maxWaitMs`
|
||
65s, `maxAttempts` 2, `budgetMs` 130s, hard ceiling 90s) in **Settings →
|
||
Resilience**. It never waits on `quota_exhausted` (locked until midnight) or
|
||
auth/not-found reasons.
|
||
|
||
---
|
||
|
||
## 5. Request Queue Admission Control (v3.8.49 · issue #6593)
|
||
|
||
**Scope**: the local per-provider+connection rate-limit queue (`open-sse/services/rateLimitManager.ts`,
|
||
backed by Bottleneck), one layer below the three mechanisms above.
|
||
|
||
**`maxWaitMs` default lowered 120s → 15s.** `resilienceSettings.requestQueue.maxWaitMs`
|
||
bounds how long a request may wait in the local queue before it is dropped
|
||
(`code: "RATE_LIMIT_QUEUE_TIMEOUT"`, #4165). The factory default fell from 120000ms to
|
||
15000ms so a saturated queue fails fast instead of holding a caller for two
|
||
minutes; override via `RATE_LIMIT_MAX_WAIT_MS` (env) or the dashboard
|
||
(**Settings → Resilience**, 1–30000ms UI ceiling).
|
||
|
||
**`maxQueueDepth` — opt-in admission cap (new).** `resilienceSettings.requestQueue.maxQueueDepth`
|
||
bounds how many requests may sit queued (not yet dispatched) for one
|
||
provider+connection at once. When the queue already holds `maxQueueDepth`
|
||
requests, a new request is fast-rejected with a typed
|
||
`code: "RATE_LIMIT_QUEUE_FULL"` error **before** it ever reaches `limiter.schedule()`
|
||
— so the rejection is cheap and happens ahead of any downstream
|
||
prompt-compression / translation work for that request. Default `0` =
|
||
disabled, preserving the existing unbounded-queue behavior; bounded 0–100000.
|
||
Override via `RATE_LIMIT_MAX_QUEUE_DEPTH` (env) or
|
||
`resilienceSettings.requestQueue.maxQueueDepth` (dashboard/API patch).
|
||
|
||
The admission check itself is a pure function
|
||
(`open-sse/services/rateLimitManager/admission.ts::checkQueueAdmission`) so
|
||
it is unit-testable without a real Bottleneck limiter.
|
||
|
||
> The RFC that opened #6593 also proposed a `bypassCompressionOnRateLimit`
|
||
> flag. This repo's `open-sse/services/compression/` pipeline is
|
||
> prompt/context compression on the outbound LLM request (`chatCore.ts`,
|
||
> around the `resolveCompressionSettings`/`selectCompressionStrategy` block),
|
||
> not HTTP response compression on synthesized 429 bodies — there is no
|
||
> matching code path for a literal bypass flag. That prompt-compression step
|
||
> also currently runs *before* `withRateLimit()` in the request pipeline, so
|
||
> reordering to skip it on a queue-full rejection is a separate, larger
|
||
> change than this issue's scope; it was intentionally **not** implemented
|
||
> here and is left as a follow-up if the CPU-saving win is worth the
|
||
> reordering risk.
|
||
|
||
---
|
||
|
||
## Other Resilience Features
|
||
|
||
- **18 routing strategies** (priority, weighted, round-robin, context-relay, fill-first, p2c, random, least-used, cost-optimized, reset-aware, reset-window, headroom, strict-random, auto, lkgp, context-optimized, fusion, pipeline) — see [AUTO-COMBO.md](../routing/AUTO-COMBO.md).
|
||
- **Reset-aware routing** (v3.8.0) — prioritizes connections by quota reset time.
|
||
- **Background mode degradation** — Responses API `background: true` degraded to sync with warning.
|
||
- **Dynamic tool limit detection** — backs off providers when tool count limits hit.
|
||
- **Emergency fallback** — controlled by `OMNIROUTE_EMERGENCY_FALLBACK`; operators can override it from the Feature Flags page without a restart.
|
||
|
||
---
|
||
|
||
## Debugging
|
||
|
||
- All keys for a provider skipped → check both circuit breaker state AND each connection's `rateLimitedUntil`/`testStatus`.
|
||
- Provider permanently excluded after reset window → code reading raw `state` instead of `getStatus()`/`canExecute()`.
|
||
- One key fails, others should work → prefer connection cooldown over circuit breaker.
|
||
- Only one model fails → prefer model lockout over connection cooldown.
|
||
- State should self-recover but doesn't → check for future timestamp + read path that refreshes expired state. Permanent statuses require manual changes.
|
||
|
||
---
|
||
|
||
## TLS Fingerprinting & Stealth
|
||
|
||
Provider-specific stealth (JA3/JA4, CCH, obfuscation) is separately documented — see [STEALTH_GUIDE.md](../security/STEALTH_GUIDE.md).
|
||
|
||
---
|
||
|
||
## Resilience testing (Phase 8 · Block C)
|
||
|
||
Beyond unit tests for resilience logic, three tests exercise the runtime under
|
||
real stress/failure conditions (all integration/nightly — none block PRs):
|
||
|
||
| Test | What | Run |
|
||
| ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- |
|
||
| Chaos | Fake-upstream node injects real latency/reset/timeout/503; validates that the circuit breaker opens/recovers and `checkFallbackError` classifies 503 as recoverable fallback. | `RUN_CHAOS_INT=1 npm run test:chaos` |
|
||
| Heap-growth | ~500 streams per `createSSEStream` under `--expose-gc`; fails if the heap grows beyond the ceiling (OOM guard #3069). | `npm run test:heap` |
|
||
| k6 soak | Sustained load against `/api/monitoring/health`; p95/error thresholds. | `k6 run tests/load/k6-soak.js` (nightly) |
|
||
|
||
Orchestrated by `.github/workflows/nightly-resilience.yml` (cron + dispatch). In the
|
||
default `test:integration`, chaos and heap self-skip (without `RUN_CHAOS_INT`/`--expose-gc`).
|
||
|
||
---
|
||
|
||
## See Also
|
||
|
||
- [Architecture Guide](./ARCHITECTURE.md) — System architecture and internals
|
||
- [User Guide](../guides/USER_GUIDE.md) — Providers, combos, CLI integration
|
||
- [Auto-Combo Engine](../routing/AUTO-COMBO.md) — 12-factor scoring, mode packs
|