Phase 1 of client-side quota tracking for NVIDIA NIM (no rate-limit
headers, no usage API):
- Register nvidia in PROVIDER_DEFAULT_RATE_LIMITS (40 RPM sliding
window, matching the documented free-tier note), operator-overridable
via a new ResilienceSettings.providerQuotaOverrides map.
- Per-connection concurrency cap (default 6) via a new
nvidiaConcurrencyGate leaf module wrapping rateLimitSemaphore,
wired into DefaultExecutor.execute().
- Per-model 429 lockout: confirmed already satisfied by #6773's
passthroughModels flag on the nvidia registry entry (no new code
needed) — added as a regression-guard test instead.
Phase 2 (AIMD adaptive ceiling learning) and Phase 3 (dashboard quota
card + combo-routing headroom preference) are explicitly deferred to
follow-up issues, per the plan's own scope note.
OmniRoute v3.8.29 — 115 commits since v3.8.28. Full CHANGELOG + 41 i18n mirrors. All content quality gates green (build, unit 8/8, vitest 188/188, PR test policy, quality gates extended, docs sync, quality ratchet). Remaining red CI checks are pre-existing release flakes (coverage-shard/integration/node-compat teardown), a new transitive undici advisory in electron devDeps, and a workflow-level CodeQL fail (0 open alerts). VPS-validated by the operator.