4 Commits

Author SHA1 Message Date
Ahmet Çetinkaya
84b5caeeb9 fix(sse): bound Antigravity 429 retry loop and lock quota-exhausted accounts for full reset window (#3122)
* fix(sse): bound Antigravity short-retry 429 loop per endpoint

A persistent 429 on the short-retry branch (retryAfterMs ≤ 60s) looped
forever on the same endpoint because the branch did `urlIndex--; continue`
without checking the shared retry counter. Production log showed 77
consecutive 429s on one daily endpoint/account with zero fallback.

Gate the short-retry branch on `retryAttemptsByUrl[urlIndex] < MAX_AUTO_RETRIES`
(mirroring the already-bounded sibling), so a persistent 429 retries at most
3× per endpoint across all 3 base URLs then returns the 429 to the account-
fallback layer.

Regression test: 'bounds a persistent short-retry 429' in
tests/unit/executor-antigravity.test.ts — asserts 12 total attempts
(3 endpoints × 4) and a returned 429 with zero hang.

* fix(sse): lock Antigravity quota-exhausted account for full reset window

After the retry-loop bound, OmniRoute fell over to the next account but
re-selected the exhausted one first on every subsequent request (~60s wasted
per request). Root cause: the 429 body 'Individual quota reached. Contact
your administrator to enable overages. Resets in 164h27m24s.' was not
recognized as quota exhaustion, so the model was locked for only ~5s instead
of the real 6.8-day reset window.

Two detector fixes (mirrors Antigravity-Manager rate_limit.rs set_lockout_until):

1. classify429.ts — add QUOTA_PATTERNS: /individual quota reached/i,
   /quota reached/i, /enable overages/i so looksLikeQuotaExhausted() fires.
2. accountFallback.ts — same patterns in classifyErrorText(); extend
   parseRetryFromErrorText() to parse 'Resets? in XhYmZs' (reusing the
   existing computeDurationMs helper) so the exact reset duration reaches
   recordModelLockoutFailure as exactCooldownMs (uncapped, per user choice).

The lockout machinery already stores until = now + cooldownMs with no clamp,
bypasses getScaledCooldown when exactCooldownMs > 0, and keeps the longer of
existing/new, so the full 164h window flows through intact.

New patterns stay specific — plain 'too many requests'/'rate limit exceeded'
messages still classify as rate_limit.

Verified end-to-end against the real message:
- classify429 → quota_exhausted
- parseRetryFromErrorText → 592044000 ms (164h27m24s exactly)
- checkFallbackError → usedUpstreamRetryHint: true, cooldownMs: 592044000

* refactor(account-fallback): simplify error parsing and add cooldown safety

- Implement a 30-day cap on parsed retry durations to prevent indefinite account lockouts.
- Replace manual string matching with `looksLikeQuotaExhausted` for more robust quota detection.
- Streamline regex logic in `parseRetryFromErrorText` for better readability.
- Add unit tests for extreme cooldown values and free-tier exhaustion scenarios.

* fix(429): drop over-broad /quota reached/ pattern, keep specific matches

The bare /quota reached/ would also flag transient per-minute limits like
'request quota reached, retry in 60s' as quota_exhausted (multi-hour lock).
The Antigravity message is still caught by /individual quota reached/. Added a
regression assertion proving the transient case stays a rate_limit.

---------

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
2026-06-03 18:29:59 -03:00
Gioxa
fdb4c63244 fix(kiro): avoid treating high-traffic 429s as quota exhaustion (#2153)
Integrated into release/v3.8.0 — fixes transient Kiro 429s being incorrectly classified as quota exhaustion
2026-05-11 10:05:48 -03:00
diegosouzapw
87dfbcca62 Merge remote-tracking branch 'origin/release/v3.8.0' into feat/zero-config-auto-routing
# Conflicts:
#	CHANGELOG.md
#	Dockerfile
#	docs/i18n/ar/CHANGELOG.md
#	docs/i18n/bg/CHANGELOG.md
#	docs/i18n/bn/CHANGELOG.md
#	docs/i18n/cs/CHANGELOG.md
#	docs/i18n/da/CHANGELOG.md
#	docs/i18n/de/CHANGELOG.md
#	docs/i18n/es/CHANGELOG.md
#	docs/i18n/fa/CHANGELOG.md
#	docs/i18n/fi/CHANGELOG.md
#	docs/i18n/fr/CHANGELOG.md
#	docs/i18n/gu/CHANGELOG.md
#	docs/i18n/he/CHANGELOG.md
#	docs/i18n/hi/CHANGELOG.md
#	docs/i18n/hu/CHANGELOG.md
#	docs/i18n/id/CHANGELOG.md
#	docs/i18n/it/CHANGELOG.md
#	docs/i18n/ja/CHANGELOG.md
#	docs/i18n/ko/CHANGELOG.md
#	docs/i18n/mr/CHANGELOG.md
#	docs/i18n/ms/CHANGELOG.md
#	docs/i18n/nl/CHANGELOG.md
#	docs/i18n/no/CHANGELOG.md
#	docs/i18n/phi/CHANGELOG.md
#	docs/i18n/pl/CHANGELOG.md
#	docs/i18n/pt-BR/CHANGELOG.md
#	docs/i18n/pt/CHANGELOG.md
#	docs/i18n/ro/CHANGELOG.md
#	docs/i18n/ru/CHANGELOG.md
#	docs/i18n/sk/CHANGELOG.md
#	docs/i18n/sv/CHANGELOG.md
#	docs/i18n/sw/CHANGELOG.md
#	docs/i18n/ta/CHANGELOG.md
#	docs/i18n/te/CHANGELOG.md
#	docs/i18n/th/CHANGELOG.md
#	docs/i18n/tr/CHANGELOG.md
#	docs/i18n/uk-UA/CHANGELOG.md
#	docs/i18n/ur/CHANGELOG.md
#	docs/i18n/vi/CHANGELOG.md
#	docs/i18n/zh-CN/CHANGELOG.md
#	open-sse/config/providerRegistry.ts
#	open-sse/handlers/chatCore.ts
#	open-sse/services/usage.ts
#	open-sse/utils/streamReadiness.ts
#	scripts/check-docs-sync.mjs
#	src/app/(dashboard)/dashboard/cache/media/MediaPageClient.tsx
#	src/app/(dashboard)/dashboard/providers/[id]/page.tsx
#	src/app/(dashboard)/dashboard/settings/components/ProxyTab.tsx
#	src/app/(dashboard)/dashboard/settings/components/RoutingTab.tsx
#	src/app/(dashboard)/dashboard/usage/components/ProviderLimits/utils.tsx
#	src/app/api/usage/analytics/route.ts
#	src/i18n/messages/zh-CN.json
#	src/lib/embeddings/service.ts
#	src/lib/usage/providerLimits.ts
#	src/mitm/cert/install.ts
#	src/shared/constants/providers.ts
#	src/sse/handlers/chat.ts
#	tests/unit/compression/rtk-code-stripper.test.ts
#	tests/unit/usage-service-hardening.test.ts
2026-05-10 18:33:20 -03:00
eleata
e0928f6b37 feat(circuit-breaker): classify 429 errors and apply per-kind cooldowns (#2116)
Integrated into release/v3.8.0
2026-05-10 09:43:22 -03:00