Files
OmniRoute/docs/routing/QUOTA_SHARE.md
Diego Rodrigues de Sa e Souza dc40911583 Release v3.8.39 (#5164)
* chore(release): open v3.8.39 development cycle

* docs(changelog): backfill 5 v3.8.38 bullets merged after release finalize

These PRs squash-merged into release/v3.8.38 between the CHANGELOG finalize
(ff57be32f) and the merge-to-main (ae6e2342d), so they shipped in the v3.8.38
tag but had no bullet:

- feat(compression): Ionizer engine (lossy JSON-array sampling + CCR) (#5148)
- fix(sse): preserve non-stream reasoning fields (#5155, @rdself)
- fix(i18n): add missing English UI labels (#5153, @rdself)
- test(combo): gated live smoke (#5151) + release-expectations refresh (#5150, @KooshaPari)

(#5129 exact-host Anthropic baseUrl is already covered by the #5130 bullet — same CodeQL #674.)
Synced 41 i18n CHANGELOG mirrors.

* feat(compression): TOON best-of-N candidate encoder + encoder A/B table (#5163)

Integrated into release/v3.8.39. TOON best-of-N candidate encoder (GCF default, fail-open). 17/17 unit tests pass on merge result; CI reds were base-stale + Quality Ratchet DRIFT.

* fix(zenmux): normalize vendor-prefixed GLM system roles (#5158)

Integrated into release/v3.8.39. ZenMux vendor-prefixed GLM system-role normalization; 12/12 role-normalizer tests pass on merge result. CI reds base-stale.

* [codex] fix xAI OAuth test and reasoning effort (#5157)

Integrated into release/v3.8.39. xAI reasoning-effort normalization (max/xhigh→high) + OAuth test config; 46/46 xai-translator tests pass on merge result. CI reds base-stale.

* docs(i18n): add Traditional Chinese (zh-TW) README and update zh-CN to latest (#5162)

Integrated into release/v3.8.39. Traditional Chinese (zh-TW) README + zh-CN refresh; docs-only.

* test(security): guard PII redaction stays opt-in (default off) + Hard Rule #20 (#5159)

Integrated into release/v3.8.39. PII opt-in regression guard + Hard Rule #20; rebased to strip base-drift (+81/-1). 5/5 guard tests pass; flip-proof verified.

* test(combo): deterministic context-relay universal-handoff coverage (closes phase-2 TODO) (#5168)

Integrated into release/v3.8.39. Deterministic context-relay universal-handoff coverage (3 tests); 3/3 pass on merge result.

* docs(i18n): full sync zh-TW and zh-CN README with canonical English v3.8.39 (#5171)

Integrated into release/v3.8.39. Full zh-TW docs tree + zh-CN sync with canonical English v3.8.39; docs-only.

* fix(serve): honour HOSTNAME from .env instead of hardcoding 0.0.0.0 (#5134) (#5170)

Integrated into release/v3.8.39. HOSTNAME env override in serve (#5134) + regression test (4/4, TDD flip-proof verified).

* fix(sse): resolve nameless deepseek-web tool blocks via parameter-schema match (#5154) (#5173)

Integrated into release/v3.8.39. Schema-based nameless deepseek-web tool-block resolution (#5154); 6/6 tests pass on merge result (incl. ambiguous/no-match negatives + named-tag no-regression).

* fix(sse): normalize array user content for Command Code to avoid upstream 400 (#5166) (#5174)

Integrated into release/v3.8.39. Normalize array user content for Command Code (#5166, user-array/400 symptom); 4/4 tests pass on merge result.

* fix(sse): defer </think> close so it never leaks before tool_calls (#5123) (#5175)

Integrated into release/v3.8.39. Defer </think> close so it never leaks before tool_calls (#5123); 4/4 tests pass (incl. #4633 no-regression). CHANGELOG synced to keep all 3 v3.8.39 fixes.

* fix(dashboard): use amber for home update-step warning icon (#5176)

Integrated into release/v3.8.39. Amber for home update-step warning icon; 1/1 UI test.

* fix(api): LAN/Tailscale dashboard — host-aware CSP + GET-exempt version route + combo field errors (#5083) (#5177)

Integrated into release/v3.8.39. Host-aware CSP (ReDoS/injection-safe host validation) + GET-exempt /api/system/version (POST/spawn stays LOCAL_ONLY, exact-match safe-methods-only) + COMBO_002 firstField. 44/44 tests + route-guard membership gate green. CHANGELOG synced to keep all 4 v3.8.39 fixes.

* fix(api): replace #5083 global middleware CSP with declarative ws: scheme (#5083)

Follow-up to PR #5177 (merged): that version implemented the LAN-CSP fix (Bug 1)
with a new global `src/middleware.ts` + `src/server/csp.ts`, which contradicts the
project's documented architecture — 'No global Next.js middleware — interception is
route-specific' (CLAUDE.md / AGENTS.md) — and was merged unverified (middleware vs
next.config header precedence was never confirmed in a real build).

This replaces that approach with the minimal, declarative equivalent:
  • next.config.mjs: connect-src now permits the bare `ws:` scheme (symmetric with the
    bare `wss:` already allowed) so the dashboard can reach its own Live WS server from
    a LAN/Tailscale host. No middleware.
  • Removes src/middleware.ts, src/server/csp.ts, and tests/unit/csp-host-aware.test.ts.
  • Adds tests/unit/csp-lan-ws-5083.test.ts (incl. a guard asserting src/middleware.ts
    does NOT exist, so the global-middleware approach cannot silently return).

Bugs 2 (GET-exempt /api/system/version) and 3 (COMBO_002 field surfacing) from #5177
are unaffected and remain in place.

Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>

* test(combo): end-to-end quota-share DRR routing-decision coverage (matrix parity) (#5179)

Integrated into release/v3.8.39. Quota-share DRR routing-decision coverage (matrix parity); 2/2 pass on merge result.

* feat(agent-bridge): graceful cert-install fallback with manual guide for containers (#4546) (#5178)

Integrated into release/v3.8.39. Agent-bridge graceful cert-install fallback + manual guide (#4546); 6/6 tests pass on merge result.

* fix(antigravity): family-scoped quota lockout (gemini/claude buckets) (#5180)

Integrated into release/v3.8.39 — family-scoped antigravity quota lockout. Rebased from v3.8.37 + validated (vitest 5/5, typecheck clean, full combo-matrix green, model-lockout 99/0). Same-model cross-account retry (chat.ts) deferred pending live antigravity VPS validation.

* fix(cli): force NODE_ENV to match dev/start run mode in custom Next server (#5189)

Integrated into release/v3.8.39. Force NODE_ENV to match dev/start run mode in custom Next server; 2/2 source-scan+ordering tests pass on merge result.

* feat(compression): CCR ranged/grep/stats retrieval (ReDoS-safe, backward-compat) (#5187)

Integrated into release/v3.8.39. CCR ranged/grep/stats retrieval (safe-regex ReDoS guard + length/match caps); 17/17 tests pass on merge result.

* docs(combo): sync all combo/routing-strategy docs to current state + document test coverage (#5185)

Integrated into release/v3.8.39. Combo/routing-strategy docs sync; docs-only.

* fix(mcp): return 404 (not 400) for unknown Streamable HTTP session id (#5169) (#5191)

* fix(api): respect blocked Auto (Zero-Config) provider in /v1/models catalog (#5192) (#5194)

* test(combo): deterministic context-relay codex quota-handoff coverage (closes last gap) (#5195)

* test(ci): wire antigravity-quota-family under test:vitest (fix test-discovery orphan) (#5196)

* fix(oauth): antigravity login no longer hangs — fire-and-forget onboarding + bounded post-exchange (#5193)

Antigravity OAuth hang fix (no-PKCE/no-openid + bounded post-exchange + exchange-500 fix). Includes #5200 (Koosha) revert + owner rebaseline to keep documented comments. Integrated into release/v3.8.39.

* feat(oauth): remote Antigravity login via local helper + paste-credentials (#5203)

Remote Antigravity login: local helper (omniroute login antigravity) + paste-credentials. Integrated into release/v3.8.39.

* fix(translator): accept Claude Messages shape in non-stream malformed-200 guard (#5156)

Integrated into release/v3.8.39

* fix(cli): default dev bundler to Turbopack (16.2.x panic no longer reproduces) (#5206)

Integrated into release/v3.8.39

* fix(cli): auto-calibrate server V8 heap from physical RAM (#5172) (#5213)

The server was spawned with a fixed --max-old-space-size=512 (omniroute serve)
or no heap flag at all (Electron), so RAM-rich boxes still OOM-crashed under
load (Ineffective mark-compacts near heap limit ~500MB) with many providers/
accounts and large model catalogs. New calibrateHeapFallbackMb(os.totalmem())
defaults the heap to ~35% of RAM clamped [512,4096], wired into serve.mjs and
electron/main.js. Explicit OMNIROUTE_MEMORY_MB still wins (#2939 unchanged).

Also addresses #5160 (same OOM root); #5152 (docker) benefits via the same knob.

Closes #5172

* fix(proxy): coalesce fast-fail health probes (#5208)

Integrated into release/v3.8.39

* fix(proxy): close dispatchers when clearing cache (#5202)

Integrated into release/v3.8.39

* fix(cli): raise dev server Node heap limit to 8GB to prevent OOM (#5198)

Integrated into release/v3.8.39

* fix(auth): allow synthetic no-auth fallback for mimocode (#5205)

Integrated into release/v3.8.39

* fix(oauth): preserve Antigravity refresh_token on empty/omitted upstream response (#3850) (#5214)

Google's OAuth refresh tokens are non-rotating: the refresh response usually
omits refresh_token and occasionally returns it as an empty string. The
Antigravity executor used `typeof tokens.refresh_token === "string" ? ... `
which accepts "" (typeof "" === "string") and overwrote the stored token with
empty, nulling it on first refresh. Now treats non-string OR empty as absent and
preserves credentials.refreshToken, matching refreshGoogleToken semantics.

Closes #3850

* fix(responses): normalize non-array input (#5204)

Integrated into release/v3.8.39

* fix(stream): normalize safety finish reasons via shared helper (#5197)

Integrated into release/v3.8.39

* fix(request-logger): never render negative '(-100%)' compression badge (#5201)

Integrated into release/v3.8.39

* fix(combo): reject empty responses api output (#5207)

Integrated into release/v3.8.39 — combo failover now rejects empty Responses API output (validateQuality). Baseline rebaseline dropped (main-measured drift; maintainer rebaselines at release).

* fix(pwa): prefer cached navigation before offline page (#5209)

Integrated into release/v3.8.39 — PWA service worker prefers cached navigation before offline page (#5165).

* chore(release): v3.8.39 — 2026-06-28

* chore(release): rebaseline openapi+i18n coverage ratchet drift for v3.8.39

---------

Co-authored-by: Arthur Bodera <abodera@gmail.com>
Co-authored-by: Nguyen Minh <lop123thcs@gmail.com>
Co-authored-by: lunkerchen <labanchen@gmail.com>
Co-authored-by: Ankit <177378174+anki1kr@users.noreply.github.com>
Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>
Co-authored-by: Ardem2025 <ardemb22@gmail.com>
Co-authored-by: backryun <bakryun0718@proton.me>
Co-authored-by: Anton <39598727+NomenAK@users.noreply.github.com>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: Wilson <pedbookmed@gmail.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
2026-06-28 06:58:29 -03:00

12 KiB
Raw Blame History

title, version, lastUpdated
title version lastUpdated
Quota Sharing Engine 3.8.6 2026-05-28

Quota Sharing Engine

Doc reference: docs/routing/QUOTA_SHARE.md Part of Group B (plans 16 + 22).


Overview

The Quota Sharing Engine distributes a provider's time-based quota (e.g. Codex 5-hour window, Kimi 1500 req/h) fairly across multiple API keys that share the same connection.

Problem it solves: OmniRoute proxies many API keys against the same upstream provider account. Without sharing logic, a burst from key A can exhaust the provider quota for the hour, leaving keys B and C blocked until the window resets. The engine prevents this by:

  1. Tracking each key's rolling consumption per dimension (%, requests, tokens, $).
  2. Applying a work-conserving fair-share algorithm: a key may borrow from idle shares while the global pool is not saturated.
  3. Enforcing the result in the hot path (chatCore.ts) before the request reaches the upstream executor.

Algorithm: Fair-Share Work-Conserving

Implemented in src/lib/quota/fairShare.ts.

Modes

Condition Mode Behaviour
globalUsedPercent < saturationThreshold Generous Key may borrow up to global limit minus consumed-total
globalUsedPercent >= saturationThreshold Strict Enforce individual fair share strictly

Default saturationThreshold = 0.5 (env QUOTA_SATURATION_THRESHOLD).

Per-dimension decision

For each active dimension in the pool, the engine computes:

fairShareAllowed = poolLimit × (allocationWeight / 100)
consumed        = current rolling value for this key (from QuotaStore.peek)
remaining       = fairShareAllowed - consumed

Then:

  • policy = hard: if consumed > fairShareAllowed and mode is strict → block.
  • policy = soft: if consumed > fairShareAllowed and mode is strict → penalize (deprioritize in combo; never hard-block).
  • policy = burst: allow while global headroom exists regardless of fair share.

Cap absoluto

capValue + capUnit on an allocation is a hard ceiling independent of mode or policy. Any dimension where consumed >= capValue always blocks the request.

Multi-dimension check

A request is blocked if any dimension in the pool would block it. Dimensions are independent — a 5h% exhaustion does not affect the weekly% dimension.

Borrowing

In generous mode, a key whose allocation is under-consumed can use surplus from other keys' unallocated shares. The formula is:

maxAllowed = globalLimit - consumedByOtherKeys

where consumedByOtherKeys = consumedTotal - consumedByThisKey. The teto global (pool limit for that dimension) is always the hard ceiling.


Sliding Window Counter

Implemented in src/lib/quota/sqliteQuotaStore.ts and redisQuotaStore.ts.

Two buckets per (apiKeyId, dimensionKey):

  • curr: current bucket (floor(nowMs / windowMs))
  • prev: previous bucket (curr - 1)

Effective rolling value:

effectiveBucketIndex = floor(nowMs / windowMs)
bucketStartMs        = effectiveBucketIndex × windowMs
elapsed              = nowMs - bucketStartMs
weight               = 1 - elapsed / windowMs

effective = prev × weight + curr

Precision: ~99% accurate. The error is at most 1% of the window size at the boundary between buckets (inherent to the 2-bucket approximation).

Concurrency

SQLite driver: in-memory mutex per (apiKeyId | dimensionKey) key prevents the read-modify-write race. Pattern mirrors src/sse/services/auth.ts anti-thundering-herd.

Redis driver: Lua EVAL script for atomic increment — runs as a single Redis command.


Drivers

SQLite (default, 0-install)

  • Table: quota_consumption (see migration 073_quota_pools.sql / 074_quota_consumption.sql).
  • Best for single-instance deployments.
  • All persistence is in the existing OmniRoute SQLite DB (DATA_DIR/storage.sqlite).

Redis (optional, multi-instance)

  • Requires ioredis npm package.
  • Counters stored in Redis; metadata (pools/allocations) still in SQLite.
  • Best for multi-replica deployments where counters must be shared.

Switching drivers

Via settings UI (/dashboard/settings → Quota Store), or via env vars:

QUOTA_STORE_DRIVER=redis
QUOTA_STORE_REDIS_URL=redis://localhost:6379

DB setting has precedence over env. If driver=redis but URL is absent or ioredis is not installed, the factory falls back to SQLite and logs a warning.

Driver selection order:

  1. DB setting quotaStore.driver
  2. Env QUOTA_STORE_DRIVER
  3. Default: sqlite

Multi-Dimension

A pool can have multiple dimensions. Each dimension is independent:

QuotaDimension {
  unit: "percent" | "requests" | "tokens" | "usd",
  window: "5h" | "hourly" | "daily" | "weekly" | "monthly",
  limit: number,  // global pool ceiling for this dimension
}

Example: Codex plan (5h% + weekly%):

[
  { "unit": "percent", "window": "5h",    "limit": 100 },
  { "unit": "percent", "window": "weekly","limit": 100 }
]

A request must satisfy all dimensions to be allowed.


Plan Resolver

Implemented in src/lib/quota/planResolver.ts.

Precedence (highest to lowest):

  1. Manual DB overrideprovider_plans table, per connectionId.
  2. Known catalogsrc/lib/quota/planRegistry.ts (data-only).
  3. Empty plan — no dimensions, manual configuration required.

Known catalog

Provider Dimensions
codex percent/5h/100, percent/weekly/100
glm tokens/5h (limit=0, unknown), tokens/weekly
minimax tokens/5h, tokens/weekly
bailian percent/5h/100, percent/weekly/100, percent/monthly/100
kimi requests/hourly/1500
alibaba requests/monthly/90000
openai, anthropic No default — manual configuration required

Pipeline Integration

PRE hook (open-sse/handlers/chatCore.ts)

Runs before the upstream executor, after auth and policy checks:

resolveComboTargets / handleSingleModel
  → enforceQuotaShare(apiKeyId, connectionId, provider, estimatedCost)
      → getQuotaStore().peek() per dimension
      → fairShare.decideFairShare()
      → if block → return 429 (buildErrorBody, Hard Rule #12)
      → if allow + deprioritize → set quotaSoftPenalty=true on candidate
  → executor.execute()

Fail-open: if enforceQuotaShare throws, the request is allowed through with a pino.warn log. This prevents a quota-engine bug from blocking all traffic.

POST hook (record consumption)

After a successful response:

executor returns success
  → spendRecorder.recordConsumption(apiKeyId, connectionId, provider, actualCost)
      → getQuotaStore().consume() per dimension
      → fail-open: errors logged as pino.warn, never propagated to client

Drift note: if consume fails post-response, the rolling counter under-counts. The saturation signal from the provider (e.g. anthropic-ratelimit-unified-5h-utilization) corrects the global estimate on the next request.

Combo soft penalty (open-sse/services/combo.ts)

When decision.deprioritize === true:

if (candidate.quotaSoftPenalty) {
  score *= QUOTA_SOFT_DEPRIORITIZE_FACTOR;  // default 0.7
}

The penalty is applied after all other scoring factors. It lowers the auto-combo probability of selecting a saturated key without hard-blocking it.


UI Walkthrough

/dashboard/costs/quota-share — Main pools page

Components (all in src/app/(dashboard)/dashboard/costs/quota-share/):

Component Purpose
QuotaConceptCard Introductory card explaining quota sharing to new users
CreatePoolModal Create a new quota pool (connection + name + initial allocations)
PoolCard Per-pool summary: name, connection, allocation count
DimensionBar Per-dimension stacked bar: each key's share + global usage
AllocationTable Table with consumed, fair share, deficit/surplus, borrowing flag
BurnRateChart EMA burn-rate line chart (lazy Recharts via dynamic())
EditAllocationsModal Edit allocation weights, caps, and policies for a pool

The page hooks:

  • usePools — fetches GET /api/quota/pools every 30s.
  • usePoolUsage — fetches GET /api/quota/pools/[id]/usage on demand.
  • useLocalStoragePoolMigration — runs once on mount to migrate legacy LS data.

/dashboard/costs/quota-share/plans — Provider plan config

  • ProviderPlanConfigClient.tsx: dropdown to select a provider, view resolved plan (auto from catalog or manual override), and edit dimensions.
  • Changes write to PUT /api/quota/plans/[connectionId].
  • Deletion reverts to catalog or empty plan.

Environment Variables

Variable Default Description
QUOTA_STORE_DRIVER sqlite Driver to use: sqlite or redis
QUOTA_STORE_REDIS_URL (empty) Redis URL, e.g. redis://localhost:6379
QUOTA_SATURATION_THRESHOLD 0.5 0..1; >= threshold activates strict mode
QUOTA_SOFT_DEPRIORITIZE_FACTOR 0.7 0..1; multiplier for soft-policy combo score
QUOTA_CONSUMPTION_RETENTION_DAYS 14 Days before GC removes old quota_consumption buckets

DB settings (quotaStore.*) override env vars.


Troubleshooting

Redis configured but not connecting

Check that ioredis is installed (npm ls ioredis) and QUOTA_STORE_REDIS_URL is reachable. On connection failure the factory falls back to SQLite (logged at warn).

peek returns stale / fail-open

If peek throws, enforceQuotaShare treats the result as "allow" (fail-open). Check pino logs for quota:enforce and quota:factory entries to identify the root cause.

Consumption counter drift

If the actual provider usage differs from the counters, it is expected — the 2-bucket sliding window has ~1% error at window boundaries, and consume is fire-and-forget post-response. The saturation signal (saturationSignals.ts) reads the real provider utilization with a 30s TTL and adjusts globalUsedPercent accordingly.

Pool shows "no data" for burn rate

computeBurnRate requires at least 2 historical samples. New pools without prior consume calls will show tokensPerSecond: 0 and timeToExhaustionMs: null.


Migration from localStorage

When /dashboard/costs/quota-share first loads, the hook useLocalStoragePoolMigration checks:

  1. localStorage.getItem("omniroute:quota-share:pools") is non-empty.
  2. GET /api/quota/pools returns [] (DB is empty).

If both are true, it posts each legacy pool to POST /api/quota/pools in batch, then removes the localStorage key. The migration is idempotent: condition 2 prevents re-migration.


Internal Strategy Classification

quota-share is an internal-only routing strategy (INTERNAL_ROUTING_STRATEGY_VALUES in src/shared/constants/routingStrategies.ts). It is used exclusively by system-minted qtSd/ pool combos and is deliberately excluded from ROUTING_STRATEGY_VALUES so it never appears as a user-selectable option in the UI or API.


Test Coverage

Two layers of automated coverage ship with the quota-share engine:

Suite Command What it covers
Unit (29 tests) node --import tsx/esm --test tests/unit/quota-share-strategy.test.ts DRR scheduler, saturation gating, concurrency caps, fairShare math, backlog queueing
Integration matrix npm run test:combo:matrix End-to-end routing decision through the real combo pipeline; DRR fairness + saturation deprioritization via live seams (registerQuotaFetcher, setLKGP, __setHeadroomSaturationFetcherForTests)

The integration matrix runs in CI alongside the other 17 public strategies. The unit suite can be run standalone.


DB Schema Summary

Three tables added by migrations 073075:

  • quota_pools + quota_allocations — pool definitions and per-key allocations.
  • quota_consumption — rolling 2-bucket counters per (apiKeyId, dimensionKey).
  • provider_plans — manual provider plan overrides (dimensions JSON per connectionId).

All tables added via idempotent CREATE TABLE IF NOT EXISTS migrations.