Compare commits

..

52 Commits

Author SHA1 Message Date
diegosouzapw
f52b60d846 fix(sse): re-align stream-readiness-policy tests with minimax's openai format
PR #9463 switched minimax/minimax-cn from claude to openai format so images
work. The stream-readiness bump for Claude-format replicas is keyed off the
registry's format field (single source of truth), so minimax legitimately
falls out of that group now. Swap the "Claude-format replica" test fixtures
to agentrouter (still format: "claude") and add explicit coverage that
minimax no longer gets the claude_format_heavy_reasoning bump.
2026-08-05 19:59:32 -03:00
Diego Rodrigues de Sa e Souza
aebd481607 Merge branch 'release/v3.8.50' into fix/minimax-openai-vision 2026-08-05 16:23:59 -03:00
diegosouzapw
7589c9f71c fix(docs): repair the #7786 squash contamination on release/v3.8.50
The #7786 squash accidentally committed its worktree copy
(.claude/worktrees/feat-7786/**, since untracked) and leaked probe tests
(repro-8522/probe-9033/repro-8956 — each now green via #9355/#9385/#9354)
plus a stray changelog.d/fixes/9159-fix.plan.md describing an UNMERGED fix
(would fabricate a changelog entry at release time — removed; #9159's own
PR ships its fragment).

This restores the PR's actual deliverable at the right paths: the
management-auth terminology guide (now with the required MDX frontmatter),
its docs test (3/3 green) and its changelog fragment.
2026-08-05 16:16:33 -03:00
Diego Rodrigues de Sa e Souza
9e3126828e fix(auto-update): skip synthetic Next.js standalone package.json without name field in resolveProjectRoot (#8956) (#9354)
A Next.js standalone build writes a synthetic .build/next/package.json
({"type":"commonjs"}) that lacks a "name" field. The resolveProjectRoot()
walk-up was stopping at this marker instead of continuing to the real repo
root, making PROJECT_ROOT point at .build/next where no .git exists, which
caused the source-mode validation to report "Not a git repository."

Fix: only accept a package.json as a project-root marker when its parsed
content has a non-empty "name" field. Keep .git as a hard marker.
Add isValidPackageMarker() helper for testability.

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-05 16:07:31 -03:00
Diego Rodrigues de Sa e Souza
5e344a3a99 fix(auth): IP blacklist now blocks on direct connections via trusted peer stamp and re-reads config without restart (#9033) (#9385)
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-05 16:07:21 -03:00
Diego Rodrigues de Sa e Souza
12c64667a9 Merge branch 'release/v3.8.50' into fix/minimax-openai-vision 2026-08-05 13:21:17 -03:00
diegosouzapw
9fcefcce9f fix(quality): tighten eslintWarnings baseline to the gate's real measurement (0)
The 2026-08-05 TS7 rebaseline wrote 5000 measured WITHOUT the suppressions
file, but the PR gate (quality.yml lint:json + quality:collect) measures WITH
suppressions applied and reads 0 - so require-tighten failed every code PR
with delta 5000 > slack. Measured 0 on the pure tip ed122b2caf after the
stale-suppression prune (#9509). TS7 debt remains tracked in
config/quality/eslint-suppressions.json; any NEW warning outside it is an
immediate red, which is the policy.
2026-08-05 13:19:56 -03:00
Bob.Hou
ed122b2caf fix(quality): prune a stale entry from the ESLint suppressions baseline (#9509)
release/v3.8.50 fails its own "No new ESLint warnings" gate right now,
independent of what any PR changes. Measured directly: a worktree
checked out at the current tip alone, no PR merged in, exits 2 with
"There are suppressions left that do not occur anymore." Cross-checked
against two unrelated open PRs (#9499, #9497) hitting the identical
failure, ruling out anything content-specific.

The mass-freeze commit that regenerated config/quality/eslint-suppressions.json
for the TypeScript 7 migration left one entry pointing at a violation
that no longer exists: src/lib/usage/providerLimits.ts no longer
triggers no-restricted-imports, but the suppression entry for it does.
ESLint's own suppression bookkeeping treats an unmatched entry as a
hard failure, separate from and in addition to real unsuppressed
errors.

--prune-suppressions removes exactly that one entry. It also drops the
informal "_comment" key documenting the freeze's origin, since ESLint's
suppression writer only round-trips file-keyed entries it manages
itself -- that context is not lost, it is still readable at the
mass-freeze commit (6b0e11e37) in git history.

This is one of two independent problems behind the same gate failure,
not the whole fix. Two files (tests/unit/issue-9407-gemini-web-validation-false-positive.test.ts,
tests/unit/v1-models-auth-leak-9320.test.ts) carry real, currently
unsuppressed no-explicit-any errors with no entry covering them at
all -- pruning cannot add what was never there. #9484 fixes those at
the source. Verified here that after this change alone, the gate
moves from exit 2 (stale suppressions) to the ordinary exit 1 those
two remaining errors cause -- both this and #9484 need to land before
the gate is green again.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
2026-08-05 12:53:15 -03:00
Diego Rodrigues de Sa e Souza
7d5e8235da fix(quality): add base-relative file-size check so inherited drift does not red innocent PRs (#8522) (#9355)
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-05 12:53:05 -03:00
Diego Rodrigues de Sa e Souza
a316c8db52 Merge branch 'release/v3.8.50' into fix/minimax-openai-vision 2026-08-05 12:00:19 -03:00
Diego Rodrigues de Sa e Souza
3022df548e fix(docs): add required MDX frontmatter to AGENTROUTER_WAF.md (#9503)
Missing title/version/lastUpdated frontmatter broke the production build
(fumadocs-mdx requires title on every docs/**/*.md file).

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-05 11:36:22 -03:00
Diego Rodrigues de Sa e Souza
ef3f554665 fix(tests): clear the two base-reds on release/v3.8.50 (#9488)
* fix(tests): clear the two base-reds on release/v3.8.50

Both sat on the release itself and turned every open PR red as soon as it
merged the release, independently of the PR's own content.

- tests/snapshots/provider/translate-path.json: #9064 (b0501642dd) added the
  code-execution-2025-08-25 and skills-2025-10-02 beta flags to the Anthropic
  header but did not regenerate the golden. provider-translate-path-golden
  failed on the bare release tip — 2 pass / 1 fail with zero PRs boarded.
  Regenerated; the diff is 24 lines, all the same header in 6 variants.

- tests/unit/v1-models-auth-leak-9320.test.ts:83: shipped a (k: any) in a file
  new enough that eslint-suppressions.json does not cover it. With
  @typescript-eslint/no-explicit-any as error under tests/, that one cast
  failed 'No new ESLint warnings' for every PR. The callback parameter infers
  correctly, so the cast was redundant.

Verified on the fix branch: eslint exit 0, typecheck:core exit 0, and both
test files green (5/5).

* fix(tests): drop a third unsuppressed any (gemini-web validation test)

A full-repo lint on this branch surfaced one more file in the same class:
tests/unit/issue-9407-gemini-web-validation-false-positive.test.ts:50 casts
(executor as any).testConnection. The file entered the release at f1ea77fd04
(23:28), newer than the frozen eslint-suppressions.json, so the cast is not
covered — and it alone kept 'No new ESLint warnings' red on this very PR.

testConnection is a declared public method on GeminiWebExecutor
(open-sse/executors/gemini-web.ts:359), so the cast was redundant rather than
load-bearing; removed outright.

eslint exit 0, typecheck:core exit 0, 15/15 across the three touched tests.
2026-08-05 11:36:19 -03:00
diegosouzapw
6b0e11e378 refactor: update quality baseline and test masking allowlist
- Updated the quality baseline to set eslintWarnings value to 5000, reflecting the migration to TypeScript 7 and the new warning thresholds.
- Modified the test masking allowlist to account for removed tests and sources, ensuring proper tracking of deprecated features.
- Enhanced ESLint configuration to ignore additional directories containing non-source files.
- Removed the .npmignore file as its contents are now managed in package.json.
- Adjusted KimiWeb model configuration to correctly map K3 to the K2D5 scenario, reflecting changes in the underlying logic.
- Updated artifact packing policy to prevent nested node_modules from being published, ensuring a leaner package size.
- Added tests to verify the exclusion of node_modules from published artifacts and to ensure the integrity of the package.json files array.
2026-08-05 08:46:22 -03:00
diegosouzapw
a549db7dee feat(infra): add systemd autostart unit for Linux (#8635) 2026-08-05 08:45:00 -03:00
diegosouzapw
f4e93f339d docs: add management authentication terminology guide (#7786) 2026-08-05 08:45:00 -03:00
Xiangzhe
2c966c28af test(mutation): include adaptive admission coverage 2026-08-05 08:45:00 -03:00
Xiangzhe
8ca40e7971 feat(api): wire shared admission across LLM routes
Acquire admission once after API-key policy, preserve lazy raw-request snapshots, and bind lease settlement to JSON, SSE, abort, deadline, and failure lifecycles. Expose a low-cardinality health summary and preserve non-SSE Ollama errors unchanged.
2026-08-05 08:44:59 -03:00
Xiangzhe
a61020153c feat(admission): add adaptive overload and pressure controls
Add bounded weighted admission with fair queuing, deadline and cancellation handling, exact lease accounting, and a default-shadow runtime. Keep asynchronous resource-pressure shedding as an independent safety fuse and bound request feature estimation.
2026-08-05 08:32:38 -03:00
Xiangzhe
ce764bc6f3 fix(combo): classify local target timeouts as gateway timeouts
Return a typed HTTP 504 for OmniRoute's per-target timer, keep fallback active, and classify the local timeout as request-scoped so it cannot degrade provider connection health.
2026-08-05 08:32:38 -03:00
Diego Rodrigues de Sa e Souza
2cb7567d66 fix(providers): treat claude-web 429 as unhealthy and forward Retry-After (#9406) 2026-08-04 23:28:41 -03:00
Diego Rodrigues de Sa e Souza
b840628de8 fix(providers): add tool_use handling to claude-web stream parser (#9408) 2026-08-04 23:28:37 -03:00
Diego Rodrigues de Sa e Souza
d969555417 fix(security): require explicit tool envelope to prevent bare JSON tool_calls (#9343) 2026-08-04 23:28:33 -03:00
Diego Rodrigues de Sa e Souza
f1ea77fd04 fix(providers): detect expired gemini-web sessions and add testConnection (#9407) 2026-08-04 23:28:30 -03:00
Diego Rodrigues de Sa e Souza
85f30d4da8 fix(api): fall back to slugified provider name when prefix is empty (#9416) 2026-08-04 23:28:26 -03:00
Diego Rodrigues de Sa e Souza
ab560cce7b fix(providers): map kimi-web/K3 to K2D5 scenario to fix resource_exhausted (#9338) 2026-08-04 23:28:21 -03:00
Diego Rodrigues de Sa e Souza
6b531fbacd fix(claude): remove unconditional always-mode return in claudeClassifierCompat (#9276) 2026-08-04 21:36:54 -03:00
Diego Rodrigues de Sa e Souza
7d6a64b054 fix(mcp): break circular import between googApiKeyAuth.ts and auth.ts (#9297) 2026-08-04 21:36:47 -03:00
Diego Rodrigues de Sa e Souza
b07182c72a fix(security): require auth for /v1/models when management auth is configured (#9320) 2026-08-04 21:36:41 -03:00
Diego Rodrigues de Sa e Souza
7e55abbc41 fix(vision-bridge): do not select unreachable describe-model when no vision provider is connected (#8430) 2026-08-04 21:36:34 -03:00
Diego Rodrigues de Sa e Souza
b0501642dd fix(providers): anthropic strips code-execution/skills beta flag, causing container rejection (#9064) 2026-08-04 21:36:18 -03:00
Diego Rodrigues de Sa e Souza
7d46d4039f fix(perplexity-web): update catalog to use 'copilot' mode and fix model IDs (#8989) 2026-08-04 21:36:13 -03:00
Diego Rodrigues de Sa e Souza
d502f144b9 fix(providers): copilot-m365-web enterprise turns send disconnectBehavior=continue (#8971) 2026-08-04 21:36:09 -03:00
Diego Rodrigues de Sa e Souza
eaea0347ac fix(executor): guard claude/anthropic buildHeaders against empty credentials and extend dual-Bearer parity for third-party baseUrls (#8653) 2026-08-04 21:36:04 -03:00
Diego Rodrigues de Sa e Souza
37edd74f2d fix(proxy-health): include credentials in proxy health check URLs (#8853) 2026-08-04 21:36:00 -03:00
Diego Rodrigues de Sa e Souza
0b70a14a3b fix(auth): setting first dashboard login password no longer fails with HTTP 400 PASSWORD_REQUIRED (#8950) 2026-08-04 21:35:29 -03:00
Diego Rodrigues de Sa e Souza
28a1f4d1b6 fix(deps): bump transitive deps for 20 Dependabot CVE alerts
Bumps ip-address, hono, fast-uri, socket.io-parser, undici (v6+v7), protobufjs, tar via targeted package.json overrides. Closes 20 Dependabot alerts (2026-08-04). npm audit → 0 vulnerabilities.
2026-08-04 19:08:34 -03:00
diegosouzapw
ed2c4dbab3 fix(deps): bump transitive deps for 20 Dependabot CVE alerts
Bumps ip-address, hono, fast-uri, socket.io-parser, undici (v6+v7),
protobufjs, and tar via targeted package.json overrides.

All patches are lockfile-only (no code change, range already covers).
Verified: npm audit → 0 vulnerabilities.
Note: brace-expansion NOT in overrides (separate major lines need
different patches; each resolved within its parent range).

Co-authored-by: wgordon17 <22222756+wgordon17@users.noreply.github.com>
2026-08-04 18:51:49 -03:00
diegosouzapw
25cf9d9065 fix(providers): switch minimax from claude to openai format so images work
The Anthropic-compatible /anthropic/v1/messages endpoint rejects image
input with 403. MiniMax's OpenAI-compatible /v1/chat/completions endpoint
supports image_url natively for MiniMax-M3.

- minimax + minimax-cn: format claude→openai, baseUrl→/v1/chat/completions
- Remove Anthropic-Version header + ?beta=true suffix (not needed for openai)
- Remove minimax/minimax-cn from ?beta=true executor case
- Update cache-control tests (openai format uses different caching path)
- Fix reasoning-split test names (no longer claude format)

TDD: 2 registry tests assert format=openai (red→green).
Refs: Hermes Agent #15715, MiniMax OpenAI-compatible API docs.
2026-08-04 18:25:58 -03:00
Diego Rodrigues de Sa e Souza
0965b041fa chore(ci): stop dependabot from grouping ioredis majors with routine bumps (#9425)
* chore(ci): stop dependabot from grouping ioredis majors with routine bumps

ioredis is loaded through a dynamic import in the distributed quota store, so a
breaking major passes build, typecheck and both test suites and only surfaces at
runtime for operators running Redis-backed quota. #9310 grouped ioredis 5.10.1 to
6.0.0 with 9 unrelated production bumps; majors get their own PR from now on.

* docs(changelog): add fragment for #9425

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-04 18:07:16 -03:00
Nick Sullivan
7b8055c7f8 fix(resilience): count STREAM_EARLY_EOF as a provider failure in combo routing (#9251)
* fix(resilience): count STREAM_EARLY_EOF as a provider failure in combo routing

A STREAM_EARLY_EOF is an upstream that accepted the request (HTTP 200), opened
the SSE stream, then closed it without emitting a single non-ping event. The
combo path classified it together with STREAM_READINESS_TIMEOUT through
isStreamReadinessFailureErrorBody(), and the readiness exemption in
shouldRecordProviderBreakerFailure meant the whole-provider circuit breaker
never saw it.

During a provider-wide outage that makes the breaker blind. Over a 7-day window
on our router we recorded 311 of these events, 302 of them on one model, 265
inside the upstream's published incident window — and the provider breaker sat
at CLOSED / failure_count=0 the entire time. Every request kept being dispatched
to the failing provider instead of shedding to the next combo target.

The two codes are different signals. The readiness probe is a pre-flight
liveness check on a connection we have not committed to, so failing it means
"this connection looks stale". An early EOF means the provider took the request
and then failed to serve it. The single-model path already treats it that way:
shouldTripProviderBreakerForResult has no readiness exemption, so a 502 early
EOF trips the breaker there. This makes the combo path consistent.

isStreamReadinessFailureErrorBody keeps matching both codes, because the
transient-retry and round-robin semaphore-cooldown paths in combo.ts do want
identical treatment for both. Only the breaker needs to tell them apart, so the
distinction is added as a narrow predicate and an optional argument rather than
by changing the shared classifier. Omitting the new argument reproduces the
previous behaviour exactly.

Follows the additive-override pattern established by the isProxyUnreachable
work, and leaves the existing exclusions for client aborts and plain 429s
untouched.

* test: register stream-early-eof-breaker in stryker tap.testFiles

The mutation test-coverage gate (check:mutation-test-coverage --strict)
detects unit tests that cover a mutated module but are missing from
stryker.conf.json tap.testFiles, so their mutant kills would not count.

comboPredicates.ts is one of the mutated modules, and the new
stream-early-eof-breaker.test.ts covers it, so the gate correctly flagged
the omission. 8376-econnrefused-breaker.test.ts -- the test this one is
modeled on -- is already registered; this just brings the new file in line.

No production code change.

---------

Co-authored-by: Nick Sullivan <nick@technick.ai>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-08-04 18:07:09 -03:00
Xiangzhe
3440c118e0 feat(usage): show Grok Build billing limits (#9205)
* feat(usage): show Grok Build billing limits

* test(usage): keep Grok quota reset fixture in the future

* fix(i18n): add Grok billing labels to pt-BR

* fix(i18n): add Grok billing labels to Vietnamese
2026-08-04 18:07:02 -03:00
nguyenha935
712910612b fix(db): bundle and verify the sql.js fallback (#9044)
Co-authored-by: nguyenha935 <208228297+nguyenha935@users.noreply.github.com>
2026-08-04 18:06:55 -03:00
Dizzle
0538eec05e test(dashboard): drop stale next-intl mock breaking ProviderDetailPageClient smoke (#9150)
The local vi.mock("next-intl") predates the #7935 global polyfill and returns a
useTranslations without .rich, crashing t.rich() in ProviderParamFilterSection:199.
The global polyfill (backed by the real createTranslator) now covers this file;
assertions only check DOM/fetch, never translated text.

Co-authored-by: Max <maxmad64@gmail.com>
2026-08-04 18:06:46 -03:00
Dizzle
0ca25d61f4 fix(dashboard): apply provider Auto Sync per connection and fan out the master toggle (#9149)
* fix(dashboard): add per-connection autoSync toggle handler

* fix(dashboard): render per-connection autoSync toggle in ConnectionRow

* fix(dashboard): wire canAutoSync into ConnectionsListPanel

* fix(dashboard): wire per-connection autoSync toggle into provider page

* fix(dashboard): make master autoSync toggle all-on with fan-out

* docs(dashboard): add changelog fragment for per-connection autoSync

* fix(dashboard): correct disable toast and assert fan-out classification

* test(dashboard): pin fan-out classification branches symmetrically

* docs(dashboard): fill changelog fragment with PR number

* fix(dashboard): port autoSync i18n keys to vi and pt-BR locales

* fix(dashboard): localize autoSync keys across all 43 locales

---------

Co-authored-by: Max <maxmad64@gmail.com>
2026-08-04 18:06:35 -03:00
Bob.Hou
8027c60726 test(sse): expect the trailing period in the no-credentials message (#9392)
#9275 started appending a candidate-alias hint to the zero-active-credentials
error and terminated the provider name with a period, so the two sentences read
as one message. The two vscode tokenized-route tests still assert the old
unterminated string and now fail on every pull request opened against this
branch.

The Quality Gates workflow only runs on pull_request to release/**, never on
push, so the branch itself never re-runs these shards and the drift stayed
invisible after the merge.

Assert what the handler actually produces. Keeping the comparison exact rather
than loosening it to a prefix match is deliberate -- the exact form is what
caught the drift.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
2026-08-04 17:08:14 -03:00
Diego Rodrigues de Sa e Souza
4a3dcf6b0b fix(routing): only let Codex-native bare ids preempt a provider when codex is active (#9447)
* fix(routing): only let Codex-native bare ids preempt a provider when codex is active

#9275 widened CODEX_NATIVE_UNPREFIXED_MODELS from a single id to gpt-5.5 plus the
gpt-5.6-sol/terra/luna tiers, so bare Codex CLI ids would reach the ChatGPT
subscription instead of fanning out to whichever provider won the inference race.
The early return it added never consulted the active-provider set, which made the
codex-only guard 30 lines below unreachable for every id in the set:

  if (CODEX_NATIVE_UNPREFIXED_MODELS.has(modelId)) return { provider: "codex", ... }

An OpenAI-only install therefore had bare gpt-5.5 routed to codex and failed with
'no active credentials for provider: codex' on a model OpenAI serves, and an install
whose codex connection was merely inactive failed identically. This also silently
reverted #5887's compatibility boundary.

The preference now only PREEMPTS another provider when a codex connection is active.
Ids that no other provider catalogs (codex-auto-review) still resolve to codex with no
connection at all — there is nothing to preempt and 'no codex credentials' is the
honest error. With codex active the preference still beats OpenAI, which is the point
of #9275, and an explicit openai/ prefix overrides it either way.

Tests: the three assertions that encode the intended #9275 change now expect codex
(plus a new one pinning the explicit-prefix override); the rest were already correct
and pass again untouched. Adds a regression test for the OpenAI-only case.

* docs(changelog): correct fragment id to #9447

* test(routing): seed an active codex connection in the bare-precedence guards

The two files #9275 added assert that bare gpt-5.5 / gpt-5.6-sol reach codex, but
they ran against an empty database — so they also pinned 'codex wins with no codex
connection at all', which is the regression #9447 removes. That put them in direct
contradiction with plan3-p0 / chat-helpers / codex-gpt55-routing-5887, which assert
openai for the very same input: no implementation could satisfy both, which is why
the release could not go green.

Seeding an active codex connection keeps the contract these files were written to
guard (codex beats openai for a Codex-native bare id) while dropping the accidental
'even with no codex configured' half. Cases that need no connection are left as they
were: the tier-only ids and codex-auto-review have no alternative provider to preempt,
and the explicit-prefix overrides are unaffected.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-04 17:08:08 -03:00
Diego Rodrigues de Sa e Souza
16ed707148 feat(providers): filter detail connections server-side (#9247)
* feat(providers): filter detail connections server-side

Filter provider detail requests at the database boundary while preserving
the full per-provider connection set needed by search, pagination, and bulk
actions. Alias-backed provider pages keep their existing aggregate behavior.

Co-authored-by: RobertsXML <RobertsXML@proton.me>
Inspired-by: https://github.com/decolua/9router/pull/2998

* chore(changelog): fragment for #9247

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: RobertsXML <RobertsXML@proton.me>
2026-08-04 14:34:22 -03:00
Dizzle
8b97ef99aa fix(db): persist account egress IP into proxy_logs (#9291)
* fix(db): persist account egress IP into proxy_logs

The account egress IP (outbound IP the upstream saw, resolved via proxyEgress.ts
echo-IP probe with 5-min cache) was computed and surfaced in the proxy_logs
console and ring buffer, but never persisted: proxy_logs.egress_ip did not
exist, so the value was lost on restart and real traffic could not be
attributed to the node/IP active at that instant.

- migration 134 adds proxy_logs.egress_ip (nullable, backward-compatible)
- schemaColumns.ensureProxyLogsColumns() idempotent reconciler
- proxyLogger self-heals the schema in loadFromDb(), persists egress_ip on
  INSERT, and matches it in search
Follows the session_tag (#8249) migration + schemaColumns reconciler pattern;
base SCHEMA_SQL untouched.

* docs(changelog): add 9291 fragment for proxy_logs egress_ip

---------

Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
2026-08-04 14:34:17 -03:00
Dizzle
e50f2329dc fix(lib): memoize catalog pricing/capability lookups to fix cold /v1/models freeze (#8697) (#8987)
Root cause: a cold GET /v1/models catalog rebuild froze the entire server 41-54s.
node --prof profiling found a systemic missing-memoization pattern — a per-model
function rescanning a static or synced data structure with Object.entries()/
Object.keys() (or hitting SQLite) on every call instead of once per rebuild. Fixed
6 instances of the same pattern, found by iteratively re-profiling the full catalog
sweep after each fix (plus a whitebox review pass) until no further hotspot of this
shape remained:

1. getModelsDevPricing() (modelsDevSync.ts) — re-ran a synchronous SQLite query and
   re-JSON.parse'd ~180 blobs on every call (up to ~6091x instead of once per
   request). Memoized via the existing modelCatalogCacheVersion invalidation signal
   (same pattern as getCachedRawProviderConnections/getCachedProviderNodes in
   db/readCache.ts). Dominant cost of the original 41-54s freeze.

2. findInsensitive() (modelMetadataRegistry.ts, resolveCatalogPricing) — rebuilt a
   full Object.entries() scan on every case-insensitive lookup miss, twice per
   model. Replaced with a lowercase-key index built once per distinct pricing
   object and cached by identity (WeakMap). Warns once at index-build time on a
   case-insensitive key collision instead of silently discarding the second value.

3. getSyncedCapability() (modelsDevSync.ts) — ran a per-model SQLite SELECT on cold
   cache instead of self-warming the whole-table cache; no caller in the
   /v1/models build path ever primed it, so a cold rebuild ran one SQLite
   round-trip per model per call site. Now self-warms via the existing bulk
   getSyncedCapabilities() on first miss. Measured as the dominant remaining cost
   after fixes 1-2 (~70% of a full catalog sweep).

4. getCanonicalModelSpecId() (shared/constants/modelSpecs.ts) — up to 3 separate
   linear scans over the static MODEL_SPECS table per call (exact ci, alias ci,
   prefix). Replaced with a lazy, lowercase-key index built once (MODEL_SPECS never
   changes at runtime); prefix-match iteration order preserved exactly so
   resolution outcomes are unchanged.

5. getStaticSpecCanonicalModelId() (modelCapabilities.ts) — duplicated the same
   exact+alias scan as (4) in a second, separate rescan. Now reuses the shared
   index via a new exported helper (findModelSpecIdByExactOrAlias) instead of
   maintaining a second cache over the same static table.
   reverseModelsDevProviders() (modelCapabilities.ts) — rescanned
   Object.entries(MODELS_DEV_PROVIDER_MAP) (also static) on every call; memoized
   by provider key. Result is frozen (readonly) since it is now shared across
   calls instead of freshly allocated each time.

6. resolveModelAlias() (shared/constants/modelSpecs.ts) — rescanned
   Object.entries(MODEL_SPECS) unconditionally once per model (verified 1:1 call
   ratio, no short-circuit). Case-sensitive exact match (Array.includes(), no
   .toLowerCase()) — uses a dedicated exact-match index, deliberately not the
   case-insensitive alias index from fix 4/5 (would silently broaden matches).

Measured on a 1940-pair real-catalog sample (static PROVIDER_MODELS registry):
cold sweep 828ms -> 356ms after fixes 3-5 on top of 1-2, extrapolating to roughly
1s on the real ~6091-model catalog, down from the original 41-54s freeze.

Complementary to the stale-serve fix in #8801 (upstream) — neither alone
eliminates the freeze.

Tests: call-count regression guards for every fix (DB prepare / Object.entries /
Object.keys call counts staying constant instead of scaling with iteration count),
plus correctness coverage for case-insensitive/case-sensitive resolution. All
pre-existing consumer suites re-verified passing (96 tests total across 19 files).

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-04 14:34:10 -03:00
Xiangzhe
455906c181 fix(reasoning): forward Ollama Cloud thinking (#9290) 2026-08-04 14:34:02 -03:00
Bob.Hou
b6bcc491bc fix(token-refresh): exempt transient errors from exponential backoff (#9242)
* fix(token-refresh): exempt transient errors from exponential backoff

A refresh that failed on a network timeout was treated exactly like one
that failed on a revoked token: the streak incremented and the circuit
backed off exponentially, up to four hours. A brief upstream blip could
therefore park a healthy account for the rest of the day.

Transient failures now take a flat two-minute retry window instead of
advancing the streak. Classification checks structured signals first
(err.name for AbortError/TimeoutError, then err.code and err.cause.code)
and only falls back to matching the message text, so it does not depend
on upstream wording. Everything else keeps the existing exponential path.

Two properties worth preserving on sight:

  - A transient failure never shortens a longer permanent backoff. The
    new window is only adopted when the existing one is not already
    further out.
  - testStatus is preserved on both paths, so a connection whose access
    token is still valid keeps serving requests while its refresh
    retries.

Only a successful refresh clears the circuit. A successful request does
not, because requests do not refresh tokens.

* chore(quality): rebaseline file-size for tokenHealthCheck.ts

src/lib/tokenHealthCheck.ts lands at 1021 lines, above the 1000 cap. The
file consolidates token-refresh health checking that was previously split
across auth.ts and tokenRefresh.ts, and the refresh circuit state machine
does not divide cleanly, so splitting it to satisfy the cap would cost
more than it buys.

Scoped to this file only. Baseline entries for files this branch does not
touch are left at their upstream values.
2026-08-04 14:33:54 -03:00
NOXX - Commiter
45d375aa0b fix(api): defer media body size limits to providers (#8843)
Image and video payloads vary by provider and base64 encoding adds substantial overhead. Exempt media routes from OmniRoute's global request-body cap so provider-specific validation determines whether a request is too large. Keep finite body limits for non-media routes and cover both header and streamed-body admission paths.
2026-08-04 14:33:48 -03:00
285 changed files with 18697 additions and 2730 deletions

View File

@@ -119,11 +119,10 @@ omnirouteSite/
# 4. Diretorios de dados / runtime locais (storage, env, secrets, scratch)
# ─────────────────────────────────────────────────────────────────────────────
data/
src/lib/env/
src/app/api/agent-skills/coverage/
src/app/api/cloud/
src/app/api/sync/cloud/
src/app/api/system/env/
# NOTA: src/lib/env/, src/app/api/{cloud,sync/cloud,system/env,agent-skills/coverage}/
# foram removidos daqui (2026-08-05). Os nomes sugerem dados/segredos locais, mas os
# 8 arquivos sao route handlers e modulos rastreados no git — escondia-los do grafo
# criava pontos cegos em buscas e em analise de impacto.
tests/golden-set/data/
# Logs e saida de teste
@@ -142,6 +141,10 @@ obsidian-plugin/node_modules/
# 6. Diretorios de documentacao interna / workflow
# ─────────────────────────────────────────────────────────────────────────────
docs/superpowers/
# Docs traduzidas: 1.215 arquivos / 94 MB (inclui 20+ copias do CHANGELOG).
# Sao traducoes do tree em ingles, ja indexado — no grafo so geram ruido em
# search_code e consomem o auto_index_limit.
docs/i18n/
# ─────────────────────────────────────────────────────────────────────────────
# 7. Arquivos especificos (nao diretorios inteiros)
@@ -188,8 +191,9 @@ audit-report.json
scripts/i18n/_audit.json
scripts/i18n/_pending-keys.json
# Cli binario local (scratch)
bin/omniroute.mjs
# NOTA: bin/omniroute.mjs foi removido daqui (2026-08-05). Estava marcado como
# "scratch", mas e o entrypoint real do CLI publicado (package.json -> bin.omniroute)
# e consta em PACK_ARTIFACT_REQUIRED_PATHS. Precisa estar no grafo.
# Deploy / docker backups
deploy.sh

View File

@@ -7,7 +7,13 @@
**/.vscode
# Dependencies and build output
# `node_modules` alone matches the ROOT only — Docker's matcher does not cross
# `/` like .gitignore does. Without the `**/` form, nested installs ship in the
# build context (e.g. @omniroute/opencode-provider/node_modules, ~79 MB of
# devDependencies). Both forms are kept: the bare one is the documented root
# rule, the `**/` one covers every nested package.
node_modules
**/node_modules
.next
.build
out
@@ -37,6 +43,17 @@ tests
test-results
playwright-report
blob-report
output
.playwright-cli
.playwright-mcp
.stryker-tmp
reports/mutation
# Local caches and quality-gate artifacts (all gitignored). `_*` does not match
# dot-prefixed names, so these need explicit entries.
.artifacts
.eslintcache
.eslintcache-complexity
# Documentation
# Issue #2348: The Dashboard Docs viewer reads markdown from `/app/docs` at
@@ -49,6 +66,10 @@ blob-report
# (English) sources at runtime, so translations are not required in the
# container image.
docs/i18n/**
# Internal planning artifacts (gitignored). `*.md` above only matches the root,
# so without this rule these land in /app/docs and become readable through the
# dashboard's Docs viewer at runtime.
docs/superpowers/**
docs/diagrams/**/*.png
docs/diagrams/**/*.jpg
docs/diagrams/**/*.jpeg

View File

@@ -39,6 +39,17 @@ updates:
# the duplication gate — migrate the gate intentionally, not via dependabot.
- dependency-name: "jscpd"
update-types: ["version-update:semver-major"]
# ioredis is a SOFT/optional dependency loaded through a dynamic import
# (src/lib/quota/redisQuotaStore.ts — "Redis driver requires ioredis package"),
# so a breaking major never fails at build or typecheck time: the only consumers
# are the distributed quota store (redisQuotaStore.ts, storeFactory.ts) and the
# `import type Redis` in src/shared/utils/rateLimiter.ts. Nothing in the unit or
# vitest suites exercises a live Redis connection, so a v5→v6 API break would ship
# green and only surface at runtime for operators running distributed quota — the
# exact users least able to absorb it. #9310 grouped that major with 9 harmless
# bumps; majors here need their own PR and a deliberate migration review.
- dependency-name: "ioredis"
update-types: ["version-update:semver-major"]
# @huggingface/transformers is HARD-PINNED at 3.5.2 (exact, no caret) — FROZEN.
# It is load-bearing for the LLMLingua ONNX compression engine (open-sse/services/
# compression/engines/llmlingua/ — worker.ts pins @huggingface/transformers@3.5.2)

View File

@@ -155,7 +155,18 @@ jobs:
- run: npm run check:fetch-targets
# docs-all / openapi-routes / docs-symbols live in docs-gates (path-filtered).
- run: npm run check:deps
- run: npm run check:file-size
# #8522: --base-ref mode for PR events — compare against max(frozen, base) so
# inherited drift (base already over frozen cap) doesn't red an innocent PR.
# workflow_dispatch (no PR base) falls back to absolute comparison.
- name: File-size ratchet (base-relative on PR)
env:
PR_BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: |
if [ -n "$PR_BASE_SHA" ]; then
npm run check:file-size -- --base-ref "$PR_BASE_SHA"
else
npm run check:file-size
fi
- run: npm run check:error-helper
- run: npm run check:migration-numbering
- run: npm run check:public-creds

9
.gitignore vendored
View File

@@ -235,7 +235,10 @@ omniroute.md
# mise configuration
mise.toml
_artifacts/ # release-green artifacts
# release-green artifacts (.gitignore has no inline comments — a trailing
# `# ...` becomes part of the pattern, so it must sit on its own line).
# Already covered by /_*/ above; kept explicit for discoverability.
_artifacts/
.claude-flow/
# ESLint file cache (npm run lint --cache / complexity ratchets)
@@ -253,3 +256,7 @@ tests/homolog/ui/.auth/
homolog-report/
docker-compose.yml.bak
.playwright-cli/
# Playwright screenshot/log output. Today every artifact happens to land inside
# output/**/.playwright-cli/ (covered above), but anything written directly to
# output/ would otherwise show up as untracked.
/output/

View File

@@ -4,11 +4,14 @@ data/
**/db.json
# VS Code extension test runtime (large binary, not needed in npm package)
app/vscode-extension/
**/data/
**/db.json
# Source code (pre-built app/ is published instead)
# Source code (pre-built dist/ is published instead)
#
# NOTA (2026-08-05): as entradas `app/*` foram removidas — o diretorio `app/`
# foi renomeado para `dist/` na Layer 1 e nao existe mais. Elas sugeriam um
# layout que ja nao e o do projeto.
#
# NOTE (#3578 / #3821-review): package.json "files" is the source of truth for what
# ships. It now allowlists the backend source closure the MCP server needs at runtime
@@ -49,8 +52,6 @@ scripts/
.vscode/
.agents/
.env*
app/.env
app/.env*
eslint.config.mjs
prettier.config.mjs
postcss.config.mjs
@@ -82,8 +83,6 @@ bun.lock
*.deb
*.rpm
electron/
app/electron/
app/vscode-extension/
# Subprojects
clipr/
@@ -93,10 +92,6 @@ vscode-extension/
# Root-level underscore-prefixed directories (private/draft — never publish)
/_*/
app/_*/
app/coverage/
app/logs/
app/tests/
# Consistent with .gitignore and .dockerignore
.DS_Store

View File

@@ -1,6 +1,11 @@
# Long reference tables are manually aligned; formatting the whole file causes noisy diffs.
docs/reference/ENVIRONMENT.md
# Generated by `npm run gen:provider-reference`; the generator aligns the tables and
# is their formatter of record. Without this, lint-staged reformats the file whenever
# it is staged and the next generator run reverts it — a diff ping-pong.
docs/reference/PROVIDER_REFERENCE.md
# Dense auto-generated free-tier budget rows (one object per line) — prettier multi-line expand blows past file-size cap 800.
open-sse/config/freeModelCatalog.data.ts

View File

@@ -0,0 +1 @@
- **docs:** add management authentication terminology guide ([#7786](https://github.com/diegosouzapw/OmniRoute/issues/7786))

View File

@@ -0,0 +1 @@
- **feat(providers):** filter provider detail connections server-side while preserving full-page search and pagination. (thanks @RobertsXML)

View File

@@ -0,0 +1,3 @@
- fix(vision-bridge): describe-model no longer returns unreachable "openai/gpt-4o-mini" when every vision-capable provider is unreachable on the instance — returns null instead and surfaces a clear error (#8430)
- fix(vision-bridge): validate fixedModel against usable credentials before short-circuiting in getBestVisionModel, so the default "openai/gpt-4o-mini" is not unconditionally selected when no OpenAI connection exists (#8430)
- fix(vision-bridge): in the combo describe path, replace raw images with an error text stub when all describe attempts fail, instead of forwarding images to a confirmed non-vision backend that would reject them with an opaque serde error (#8430)

View File

@@ -0,0 +1 @@
- fix(quality): add base-relative file-size check so inherited drift does not red innocent PRs (#8522)

View File

@@ -0,0 +1 @@
- fix(executor): guard claude/anthropic buildHeaders against empty credentials and extend dual-Bearer parity for third-party baseUrls (#8653)

View File

@@ -0,0 +1 @@
- **fix(api):** Let image and video providers enforce their own request-size limits instead of rejecting media payloads at OmniRoute's 10 MB global default ([#8843](https://github.com/diegosouzapw/OmniRoute/pull/8843)) — thanks @artickc

View File

@@ -0,0 +1 @@
- fix(proxy-health): include credentials in proxy health check URLs (#8853)

View File

@@ -0,0 +1 @@
- fix(auth): setting first dashboard login password no longer fails with HTTP 400 PASSWORD_REQUIRED (#8950)

View File

@@ -0,0 +1 @@
- fix(auto-update): skip synthetic Next.js standalone package.json without `name` field in resolveProjectRoot (#8956)

View File

@@ -0,0 +1 @@
- fix(providers): copilot-m365-web enterprise turns send disconnectBehavior=continue (#8971)

View File

@@ -0,0 +1 @@
- fix(auth): IP blacklist now blocks on direct connections via trusted peer stamp and re-reads config without restart (#9033)

View File

@@ -0,0 +1 @@
- fix(providers): anthropic strips code-execution/skills beta flag, causing container rejection (#9064)

View File

@@ -0,0 +1 @@
- **fix(dashboard):** the provider "Auto Sync" toggle now applies to every active connection and each connection gets its own Auto Sync toggle — previously only the lowest-priority connection was updated. ([#9149](https://github.com/diegosouzapw/OmniRoute/pull/9149))

View File

@@ -0,0 +1 @@
- fix(claude): remove unconditional "always" return in claudeClassifierCompat so normal chat requests are not swallowed (#9276)

View File

@@ -0,0 +1 @@
- **fix(db):** persist the account egress IP into `proxy_logs.egress_ip` (migration 134 + schema reconciler) so real traffic stays attributable to the actual node/IP even after restart — the egress IP was previously computed and logged but silently dropped from persistence ([#9291](https://github.com/diegosouzapw/OmniRoute/pull/9291)) — thanks @maxmad64bis

View File

@@ -0,0 +1 @@
- fix(mcp): break circular import between googApiKeyAuth.ts and auth.ts to fix esbuild SyntaxError in MCP server bundle (#9297)

View File

@@ -0,0 +1 @@
- fix(security): require auth for /v1/models when management auth is configured (#9320)

View File

@@ -0,0 +1 @@
- fix(providers): map kimi-web/K3 to K2D5 scenario instead of OK Computer premium mode to fix resource_exhausted on non-subscriber accounts (#9338)

View File

@@ -0,0 +1 @@
- fix(security): require explicit tool envelope to prevent bare JSON from being promoted to real tool_calls (#9343)

View File

@@ -0,0 +1,2 @@
- fix(providers): treat claude-web 429 as unhealthy and forward upstream Retry-After header (#9406)
- fix(providers): treat muse-spark-web 429 as unhealthy (#9406)

View File

@@ -0,0 +1 @@
- fix(providers): detect expired gemini-web sessions via ServiceLogin redirect and add testConnection override (#9407)

View File

@@ -0,0 +1 @@
- fix(providers): add tool_use block handling to claude-web stream parser for OpenAI tool_calls projection (#9408)

View File

@@ -0,0 +1 @@
- fix(api): fall back to slugified provider name when prefix is empty to prevent UUID leak in /v1/models (#9416)

View File

@@ -0,0 +1 @@
- **fix(routing):** a Codex-native bare model id (`gpt-5.5`, the `gpt-5.6-sol`/`terra`/`luna` tiers) no longer routes to `codex` when no codex connection is active — an OpenAI-only install was getting `no active credentials for provider: codex` for a model OpenAI serves, and an install whose codex connection was merely inactive failed the same way. With codex active the Codex preference still wins over OpenAI, and ids only codex catalogs (`codex-auto-review`) still resolve to codex with no connection at all ([#9447](https://github.com/diegosouzapw/OmniRoute/pull/9447))

View File

@@ -0,0 +1 @@
- **chore(ci):** stopped dependabot from grouping `ioredis` majors with routine production bumps — the package is resolved through a dynamic import in the distributed quota store, so a breaking major passes build, typecheck and both test suites and only surfaces at runtime for operators running Redis-backed quota ([#9425](https://github.com/diegosouzapw/OmniRoute/pull/9425))

View File

@@ -0,0 +1 @@
- **chore(tests):** cleared two base-reds sitting on `release/v3.8.50` itself, both of which turned every open PR red the moment it merged the release. `tests/snapshots/provider/translate-path.json` was stale: #9064 added the `code-execution-2025-08-25` and `skills-2025-10-02` beta flags to the Anthropic header without regenerating the golden, so `provider-translate-path-golden` failed on the bare release tip (2 pass / 1 fail with zero PRs boarded). And `tests/unit/v1-models-auth-leak-9320.test.ts:83` shipped a `(k: any)` in a file new enough that `config/quality/eslint-suppressions.json` does not cover it — with `@typescript-eslint/no-explicit-any` set to `error` under `tests/`, that single cast failed the `No new ESLint warnings` gate repo-wide. Same class in `tests/unit/issue-9407-gemini-web-validation-false-positive.test.ts:50` (`(executor as any).testConnection`), which entered at `f1ea77fd04``testConnection` is a declared public method on `GeminiWebExecutor`, so that cast was redundant too. Regenerating the snapshot and dropping both casts restores the gates.

View File

@@ -0,0 +1 @@
- fix(quality): tighten eslintWarnings baseline 5000->0 to match the gate's suppressions-applied measurement (unblocks require-tighten on every code PR)

File diff suppressed because it is too large Load Diff

View File

@@ -283,6 +283,7 @@
"_rebaseline_2026_07_27_3850_relax_filesize_cap_v2_20pct": "OWNER-APPROVED TEMPORARY relax for v3.8.50-3.8.54 PREPARE phase (docs/ROADMAP.md). v1 was cap 800->900 / testCap 800->900 on 2026-07-27; v2 = v1 +20% buffer = cap 900->1000 (+100), testCap 900->1000 (+100). Justification: same as complexity v2 — the v3.8.50 release cut coincides with high-merge activity; owner accepted enlarging the headroom to cover the entire PREPARE phase (5 minor cycles .50-.54) without per-PR rebaseline noise. Targets: decompose-existing-frozen unchanged (frozen still only-shrink — see frozen[] entries and the 105 files >900 that still need structural decomposition regardless of cap); this only relaxes the cap for NEW files in the decompose/extract-while-PREPARE phase (.51='executor registry in-place' and .52='combo.ts decomposition' create new leaf modules above 800). RE-TIGHTENING MANDATORY in v3.8.51: cap target 850 = 850 once decomposition wave stabilizes (gives 150 units of post-tighten headroom vs the new 1000 ceiling). Tracked via same roadmap issue as complexity v2. Window: v3.8.50 (release cut) → v3.8.54 close (RE-TIGHTEN at v3.8.51 prep merge per ROADMAP.md). Last entry unless measured regression. v1 entry retained below for audit trail.",
"_rebaseline_2026_07_27_3850_relax_filesize_cap": "OWNER-APPROVED TEMPORARY relax for v3.8.50-3.8.54 PREPARE phase (docs/ROADMAP.md). cap 800->900 (+100), testCap 800->900 (+100). Targets: decompose-existing-frozen unchanged (frozen still only-shrink); this only relaxes the cap for NEW files in the decompose/extract-while-PREPARE phase (.51='executor registry in-place' and .52='combo.ts decomposition' create new leaf modules above 800). RE-TIGHTENING MANDATORY in v3.8.51: cap target 850 = 850 once decomposition wave stabilizes. SUPERSEDED by _rebaseline_2026_07_27_3850_relax_filesize_cap_v2_20pct (v1 +20% buffer) — retained for audit. Tracked via same roadmap issue.",
"_rebaseline_2026_07_27_v3849_train1h": "Merge-train 1H (31 PRs) — owner-approved 2026-07-27. Two distinct causes, kept separate on purpose: (1) GENUINE irreducible growth at existing chokepoints — providerLimits/auth (#8632 Kimi quota-reset recovery), rateLimitManager (#8616 idle wedged limiters), models-catalog-route.test (#8610 OpenCode Go effort aliases); (2) COLLISION with #8585, which banked shrinks measured on the pre-train release tip while 30 sibling PRs in the SAME train grew those files again — chat/accountFallback (#8628), chatCore (#8613), videoGeneration (#8581), imageGeneration. The zero-headroom frozen entries cannot absorb either. Ceilings re-pinned to the post-merge tip; #8612 (also in this train) automates shrink-banking so this self-inflicted drift stops recurring. Detail: src/lib/usage/providerLimits.ts 1006->1013 (#8632); src/sse/services/auth.ts 2492->2508 (#8632); open-sse/services/rateLimitManager.ts 1014->1060 (#8616); src/sse/handlers/chat.ts 1842->1845 (#8628); open-sse/handlers/chatCore.ts 4939->4955 (#8613); open-sse/handlers/imageGeneration.ts 3100->3101 ((sem PR — teto do #8585)); open-sse/handlers/videoGeneration.ts 1038->1063 (#8581); open-sse/services/accountFallback.ts 1965->1966 (#8628); tests/unit/models-catalog-route.test.ts 1608->1636 (#8610)",
"_rebaseline_2026_08_02_9242_token_health_transient": "PR #9242 (fix/refresh-circuit-transient): src/lib/tokenHealthCheck.ts 1021 (new file, above cap 1000). The file consolidates token-refresh health checking logic that was previously scattered across auth.ts and tokenRefresh.ts. Cohesive single-responsibility module for refresh circuit state management; not extractable without splitting the refresh state machine. Covered by tests/unit/tokenHealthCheck-transient.test.ts.",
"frozen": {
"_rebaseline_2026_06_22_4644_deepseek_web_tools": "PR #4644 (BugsBag/robust deepseek-web tool-call parsing): open-sse/executors/deepseek-web.ts 1117->1125 (+8). The new agentic tool-call path emits surrounding text + reasoning before tool_calls and swaps to the dedicated deepseekWebTools.ts parser; the +8 lines are cohesive wiring at the existing transformSSE chokepoint (the parser itself lives in the new deepseekWebTools.ts file, already under cap). The PR's own fast-gate (PR->release) does not run check:file-size, so this surfaced only at release reconcile. Covered by tests/unit/deepseek-web-tools-variants.test.ts + deepseek-web-tools-execute.test.ts.",
"_rebaseline_2026_06_23_4712_deepseek_web_tool_results": "PR for #4712 (deepseek-web drops role:tool): open-sse/executors/deepseek-web.ts 1125->1148 (+23). messagesToPrompt() now folds role:\"tool\" results into the single-prompt transcript (recovering the tool name from the preceding assistant tool_calls by tool_call_id) instead of silently dropping them; the lines are cohesive wiring inside the existing function. Covered by tests/unit/deepseek-web-tool-result-prompt-4712.test.ts.",
@@ -385,10 +386,10 @@
"src/app/(dashboard)/dashboard/settings/components/SystemStorageTab.tsx": 1573,
"src/app/(dashboard)/dashboard/usage/components/BudgetTab.tsx": 1028,
"src/app/(dashboard)/dashboard/usage/components/EvalsTab.tsx": 2148,
"_rebaseline_2026_07_30_8916_quota_compact_layout": "PR #8916 (apoapostolov, feat/improve-provider-quota-layouts) own growth: src/app/(dashboard)/dashboard/usage/components/ProviderLimits/index.tsx 1109->1153 (+44) adds Full/Compact layout toggle — LS_LAYOUT_MODE constant, LayoutMode type, layoutMode state, toggleLayoutMode callback, toggle button with icon. At existing filter/settings chokepoint. Not extractable without splitting state + toolbar away from data-fetching. Covered by tests/unit/quota-card-grid-compact-layout-8916.test.ts.",
"src/app/(dashboard)/dashboard/usage/components/ProviderLimits/index.tsx": 1153,
"src/app/(dashboard)/dashboard/usage/components/ProviderLimits/index.tsx": 1109,
"src/app/api/providers/[id]/models/route.ts": 2250,
"src/app/api/v1/models/catalog.ts": 1549,
"src/lib/tokenHealthCheck.ts": 1021,
"src/lib/db/apiKeys.ts": 1529,
"src/lib/db/core.ts": 1637,
"src/lib/db/migrationRunner.ts": 1077,
@@ -414,6 +415,6 @@
"_rebaseline_2026_07_28_8860_tokenrefresh_projectid": "PR #8860 (fix/antigravity-projectid-centralized) own test growth: tests/unit/token-refresh-service.test.ts 1311->1378 (+67 = 4 cases covering projectId discovery on the tokenRefresh.ts path — the Dashboard/health-check refresh route, which #8842 did not reach since that fixed the executor path). Covered by the same file.",
"_rebaseline_2026_07_28_8861_xiaomi_token_plan": "PR #8861 (feat/xiaomi-token-plan-protocol-selector) own growth: EditConnectionModal.tsx 1283->1316 (+33 = the per-connection API-protocol selector field) and open-sse/executors/base.ts 1540->1562 (+22 = alternate-format resolution at the existing buildUrl/headers chokepoint). Both are irreducible wiring at existing call sites.",
"_rebaseline_2026_07_28_8863_firefly_detail_level": "PR #8863 (fix/adobe-firefly-gpt-detail-level-max) own growth: adobeFireflyClient.ts 2317->2322 (+5 = gpt-image detailLevel defaulting to maximal at the existing payload-build site). Covered by tests/unit/adobe-firefly.test.ts.",
"_rebaseline_2026_07_29_8281_home_quickstart_prefetch": "Release v3.8.49 base-red fix (no PR — captain sweep): src/app/(dashboard)/dashboard/HomePageClient.tsx 1377->1381 (+4). #8292 added prefetch={false} to the sidebar but left /home's five quick-start Links prefetching, so first paint still fired 12 speculative RSC requests — caught by navigation.spec.ts only after the e2e helper bug (APP_ROUTE_PATTERN missing /home) was repaired in the same cycle. Growth is the five prefetch attributes; it was offset first by extracting the repeated className literals (INLINE_LINK x4, DOCS_LINK x1), which collapsed five wrapped <Link> blocks back to one line each — a naive fix measured 1391. Guard: tests/unit/sidebar-prefetch-policy-8281.test.ts.",
"_rebaseline_2026_07_29_8281_home_quickstart_prefetch": "Release v3.8.49 base-red fix (no PR — captain sweep): src/app/(dashboard)/dashboard/HomePageClient.tsx 1377->1381 (+4). #8292 added prefetch={false} to the sidebar but left /home's five quick-start Links prefetching, so first paint still fired 12 speculative RSC requests — caught by navigation.spec.ts only after the e2e helper bug (APP_ROUTE_PATTERN missing /home) was repaired in the same cycle. Growth is the five prefetch attributes; it was offset first by extracting the repeated className literals (INLINE_LINK x4, DOCS_LINK x1), which collapsed five wrapped <Link> blocks back to one line each — a naive fix measured 1391. Guard: tests/unit/sidebar-prefetch-policy-8281.test.ts.",
"_rebaseline_2026_08_02_v3850_agentrouter_responses": "Release v3.8.50 AgentRouter/Codex compatibility reconciliation. open-sse/executors/base.ts 1562->1578: #9190 wires AgentRouter's selected Claude/OpenAI/Responses protocol through the existing executor URL, auth, identity-header and fingerprint chokepoints; the reusable alternate resolver remains outside base.ts. open-sse/utils/stream.ts 2887->2889: #9213 evaluates Responses ID and usage normalization independently so response.completed always receives finite usage.total_tokens instead of short-circuiting after an ID rewrite. tests/unit/chatcore-translation-paths.test.ts 2769->2776: #9191 updates the existing Claude-Code bridge assertions for the dynamic AgentRouter wire image. PR #9224 offsets its own chatCore growth by extracting the AgentRouter protocol decisions into chatCore/agentRouterProtocol.ts, leaving chatCore below its frozen ceiling. Covered by agentrouter executor/chatCore protocol tests, chatcore translation-path tests, and responses-commentary-passthrough tests."
}

View File

@@ -3,23 +3,8 @@
"metrics": {
"eslintWarnings": {
"value": 0,
"_rebaseline_2026_07_03_v3844_residual_release_green": "4270->4279 (+9). v3.8.44 residual drift on release tip 716041223 (moving target: eslint 4270->4279 as the branch advanced past the prior rebaseline). Inherited from parallel-session merges (Quality Ratchet not on PR->release fast-gates).",
"_rebaseline_2026_07_03_v3844_ipfilter_release_green": "4256->4270 (+14). v3.8.44 cycle drift measured on release tip 32e4c906e during the #6131/#5975 release-green rebaseline. Inherited from the merge burst (Quality Ratchet does not run on PR->release fast-gates). route-edge-coverage +7 is my #5975 test comment; the rest is parallel-session drift. Tighten via --update next cycle.",
"_rebaseline_2026_07_03_v3844_review_prs_fix_batch": "4199->4256 (+57). Inherited v3.8.44 cycle drift surfaced by the release-green pre-flight (the Quality Ratchet does NOT run on PR->release fast-gates, so warnings accrue unmeasured across the cycle). 4256 = measured by `node scripts/quality/collect-metrics.mjs` on the release tip 72ee80649 during the /review-prs fix-batch round. The round's own merges (#5958 SSE-accept, #5988 deepseek-web, #6013/#5974 retry-after-json, #5975 embeddings-proxy, #5973 non-json-guard) plus the parallel-session merge burst into release/v3.8.44 account for the delta; all `any`-warn-allowed in open-sse/ + tests/. Cyclomatic is already green (2012 < baseline 2015) and needs no bump. Tighten via --require-tighten next cycle.",
"_rebaseline_2026_07_02_v3843_release_close": "4158->4199 (+41). v3.8.43 release-close drift measured by the release-green pre-flight (the Quality Ratchet does NOT run on PR->release fast-gates, so warnings accrued unmeasured across the ~120 commits merged after the mid-cycle 4158 rebaseline — the compression T02/T05/T06/T07/T08/T10 engine families, memory typed decay, provider adds Ollama/SenseNova, ~55 SSE/translator/kiro/oauth/dashboard fixes, and the god-file decomposition wave). Trust-but-verify: measured 4199 via `npm run lint` on the release-finalize working tree INCLUDING my changes (CHANGELOG/i18n/README docs + kiro pricing data entry + the 3 base-red CODE fixes: opencode fabrication removal, resolveEffectiveKey type-widen, openai-to-claude claudeFinishEmitted flag + 4 test-alignment files + golden snapshot regen) — the code fixes NET-REMOVE lines and add no `any`/unused, and lint reported 4199 both before and after them, so all +41 is inherited cycle drift (`any` warn-allowed in open-sse/ + tests/). Tighten via --require-tighten next cycle.",
"direction": "down",
"_rebaseline_2026_07_01_v3843_release": "4121->4158 (+37). v3.8.43 cycle drift surfaced by the release-green pre-flight; the Quality Ratchet does NOT run on PR->release fast-gates, so warnings accrued unmeasured across this cycle. 4158 = the value measured by the CI Quality Ratchet on the release tip fce85136c (release PR #5609). Trust-but-verify: the fix/release-v3843-ci-reds branch touches only test files (rtk-mcp-tools de-flake, compression-studio e2e anchor, oauth-error-linkify hardening test) + src/shared/utils/linkify.ts (eslint-clean, 0 warnings) + stryker.conf.json + this baseline -> 0 new warnings, so all +37 is inherited cycle drift (any warn-allowed in open-sse/ + tests/). Tighten via --require-tighten next cycle.",
"_rebaseline_2026_06_30_v3842_release": "4116->4121 (+5). v3.8.42 cycle drift surfaced by the release-green pre-flight (the Quality Ratchet does NOT run on PR->release fast-gates, so warnings accrued unmeasured across this cycle's 90 commits — chatgpt-web PoW sha3-512 BoringSSL fix #5540, provider baseUrl/i18n umbrella #5511, proxy union proxyUrlMap+acct.proxy #5521, dead-code + duplication waves #5468-#5495, tls-options packaging #5503, release-freeze + .npmrc fetch-retries #5506, dast-smoke spawn-prefix client-safe extraction #5546, plus ~30 SSE/translator/combo/dashboard fixes). Trust-but-verify: measured 4121 via `npm run check:release-green` on the working tree INCLUDING my reconciliation (CHANGELOG/i18n/golden snapshot + file-size baseline) — those touch only config JSON + a provider snapshot (eslint-ignored) and contribute 0 warnings; all +5 is inherited cycle drift (`any` warn-allowed in open-sse/ + tests/). Tighten via --require-tighten next cycle.",
"_rebaseline_2026_06_29_v3841_release": "4103->4116 (+13). v3.8.41 cycle drift surfaced by the release-green collect (the Quality Ratchet does NOT run on PR->release fast-gates, so warnings accrued unmeasured across this cycle's 52 commits — relay backend #5315, gemini catalog #5337, services dashboard #5299, empty-Claude-messages guard #5342, thinking-budget/redacted-replay + marker opt-out #5312/#5352/#5367, opencode proxy-pool + observability #5217/#5370/#5351, cors + HTTPS-serve #5242/#5360/#5361, grok cf_clearance #5350/#5358, oauth/chatgpt-web/routing/cli/dashboard/rerank #5326/#5240/#5239/#5238/#5264/#5332, partially offset by the dead-code sweep #5321-#5371). Trust-but-verify: measured 4116 via `npm run quality:collect` on the working tree INCLUDING my reconciliation (CHANGELOG/i18n/README/env docs + baselines) AND the lint-fix in useServiceLogs.ts — that fix REMOVES a setState-in-effect ERROR (eslintErrors stays 0) and adds an `open` listener with no `any`/unused, contributing 0 warnings; all +13 is inherited cycle drift (`any` warn-allowed in open-sse/ + tests/). Tighten via --require-tighten next cycle.",
"_rebaseline_2026_06_29_v3840_release": "4090->4103 (+13). v3.8.40 cycle drift surfaced by the release-green pre-flight + the release PR Quality Ratchet (the ratchet does NOT run on PR->release fast-gates, so warnings accrued unmeasured across this cycle's ~57 commits — compression roadmap relevance/hard-budget/memoization/transparency/saliency/splitter/tool_search/RTK/QuantumLock #5289/#5288/#5286/#5284/#5285/#5283/#5269/#5268/#5260, ~20 SSE/translator/combo fixes #5248/#5250/#5254/#5261/#5255/#5273/#5258, M365 Copilot provider #5302, public-origin centralization #5278). Trust-but-verify: measured 4103 locally via `npm run quality:collect` on the release tip INCLUDING my reconciliation commits (CHANGELOG + main merge + the 2 regression test fixes 165c823f5) — the test fixes add 0 `any`/warnings (health-autopilot added a NextRequest import + asserts; chat-pipeline changed one Accept string + a comment), so all +13 is inherited cycle drift (`any` warn-allowed in open-sse/ + tests/). Tighten via --require-tighten next cycle.",
"_rebaseline_2026_06_28_v3839_release": "4002->4090 (+88). v3.8.39 cycle drift surfaced by the release-green pre-flight (the Quality Ratchet does NOT run on PR->release fast-gates, so warnings accrued unmeasured across this cycle's 40 commits — antigravity remote-login + quota-family #5203/#5180/#5193, compression CCR-retrieve + TOON encoder #5187/#5163, ~20 SSE/translator/responses fixes #5156/#5154/#5197/#5204/#5158/#5123/#5166, proxy/health hardening #5202/#5208/#5209/#5201 from @KooshaPari, combo quota-share/context-relay E2E tests #5179/#5168/#5195). Trust-but-verify: this release-finalize working tree touches ONLY CHANGELOG.md, docs/i18n/*/CHANGELOG.md mirrors, README.md and these baselines — 0 production-code change, so all +88 is inherited cycle drift (`any` warn-allowed in open-sse/ + tests/). Tighten via --require-tighten next cycle.",
"_rebaseline_2026_06_27_v3838_release": "3987->4002 (+15). v3.8.38 cycle drift surfaced by the release-green pre-flight (the Quality Ratchet does NOT run on PR->release fast-gates, so warnings accrued unmeasured across this cycle's ~78 commits — provider adds Factory/Grok-Build/ZenMux-Free/Alibaba-video, ~30 SSE/translator/diagnostics fixes, compression fidelity-gate + playground #5080/#5143, Fusion editor #5074, salvage batches #5138/#5141). Trust-but-verify: this release-finalize working tree touches ONLY CHANGELOG.md, docs/i18n/*/CHANGELOG.md mirrors, README.md and these baselines — 0 production-code change, so all +15 is inherited cycle drift (`any` warn-allowed in open-sse/ + tests/). Tighten via --require-tighten next cycle.",
"_rebaseline_2026_06_25_v3836_release": "v3.8.36 cycle drift surfaced by the post-merge fix PR #5029 (the Quality Ratchet was SKIPPED on the release PR #4854 itself, and does NOT run on the PR→release fast-gates, so warnings accrued unmeasured across this cycle's 137 commits — Quota-Share Fase 2/3 features, god-file decomposition #3501/#4811-#4956, 14 external contributor PRs). 3912→3970 (+58), the exact value measured by the CI Quality Ratchet on #5029. Trust-but-verify: this fix PR touches ONLY scripts/build/pack-artifact-policy.ts (a string-literal allowlist array, scripts/ is eslint-light) and tests/integration/resilience-http-e2e.test.ts (2 string keys, no `any`) — 0 new warnings, so all +58 is inherited cycle drift (`any` warn-allowed in open-sse/ + tests/). Same precedent as _rebaseline_2026_06_23_v3835_release. Tighten via --require-tighten next cycle.",
"_rebaseline_2026_06_23_v3835_release": "v3.8.35 cycle drift surfaced by the release-green pre-flight (the Quality Ratchet does NOT run on PR→release fast-gates, so warnings accrued across this cycle's parallel-session merges — Compression Phase 4 #4694/#4707/#4716/#4720, chatCore #3501 leaf extractions, contributor PRs #4726/#4753/#4774/#4781/#4783/#4793, etc.). 3907→3912 (+5). Verified my release-finalize working tree touches ONLY docs/*.md (THREAT_MODEL), CHANGELOG.md, baselines, and 1 string line in scripts/check/check-fabricated-docs.mjs — 0 production-code change, so all +5 is inherited contributor drift. No coverage/openapi/i18n regressions.",
"_rebaseline_2026_06_22_v3834_release": "v3.8.34 cycle drift surfaced by the release-green pre-flight (the Quality Ratchet does NOT run on PR→release fast-gates, so warnings accrued across this cycle's parallel-session merges — #4583-4586/#4588-4593/#4606-4621/#4644/#4647/#4696/etc.). 3900→3907 (+7). Verified my release-finalize working tree touches ONLY CHANGELOG.md (git status: 0 code changes), so all +7 is inherited contributor drift. No coverage/openapi/i18n regressions.",
"_rebaseline_2026_06_22_v3833_release": "Cumulative cycle drift surfaced by the release PR full CI. 3867→3900 (+33).",
"_rebaseline_2026_06_26_v3837_release": "3970->3987. v3.8.37 cycle drift surfaced by the release-green pre-flight (the Quality Ratchet does NOT run on PR->release fast-gates, so warnings/complexity accrued unmeasured across this cycle's 76 commits — provider adds DGrid/Pioneer/xAI, headroom proxy lifecycle #4649, ~50 SSE/translator fixes, Engine Combos #5062). Trust-but-verify: this release-finalize working tree touches ONLY CHANGELOG.md, docs/i18n/*/CHANGELOG.md mirrors, and these baselines — 0 production-code change, so all drift is inherited cycle drift (`any` warn-allowed in open-sse/ + tests/). Tighten via --require-tighten next cycle.",
"_rebaseline_2026_07_04_pacote4_no_new_warnings": "4279->0. Pacote 4 do plano mestre testes+CI: a divida pre-existente (4279 warnings + violacoes das 3 regras promovidas a error em src/**) foi CONGELADA em config/quality/eslint-suppressions.json (ESLint bulk suppressions nativo) e passa a ser bloqueada NO PR que a introduziria (job lint-guard no quality.yml + npm run lint + lint-staged, todos suppressions-aware; fork = report-only, Principio Zero). collect-metrics agora mede sob o baseline congelado -> a metrica vira 'divida liquida NOVA' (~0 em regime). O aperto do ESTOQUE congelado acontece via `npx eslint . --prune-suppressions --suppressions-location config/quality/eslint-suppressions.json` na reconciliacao da release. Fim das rebaselines-surpresa de +41/+88 por ciclo."
"_rebaseline_2026_08_05_post_prune": "Apertado 5000->0 em 2026-08-05: o gate mede via lint:json COM as suppressions aplicadas (config/quality/eslint-suppressions.json congela a divida da migracao TS7), entao a contagem real do gate e 0. O 5000 anterior foi medido SEM suppressions (4139 brutos) e fazia o require-tighten reprovar todo PR de codigo (delta 5000>slack). Divida TS7 continua rastreada nas suppressions; warning NOVO (fora delas) agora e red imediato, que e a politica."
},
"eslintErrors": {
"value": 0,

View File

@@ -64,6 +64,18 @@
"tests/unit/ui/provider-plan-config.test.tsx": {
"replacement": "tests/unit/quota-plans-route-retired.test.ts",
"reason": "v3.8.49 #7127: fix(tests) suíte vitest UI de volta ao verde — a rota Plans e o ProviderPlanConfigClient foram APOSENTADOS; o replacement inverte a asserção e guarda a aposentadoria (o arquivo da rota e o ProviderPlanConfigClient não existem mais, costs-quota-plans saiu do sidebarVisibility e da navegação)."
},
"tests/unit/plugin-sandbox-permissions.test.ts": {
"sourceRemoved": [
"src/lib/plugins/pluginWorker.ts",
"src/lib/plugins/sandbox.ts",
"src/lib/plugins/signing.ts"
],
"reason": "v3.8.50 #9126 (commit 8fac6bcd48): pluginWorker.ts, sandbox.ts e signing.ts foram removidos por completo (\"zero importers confirmed\") — o subsistema de sandbox de plugins com worker-thread nunca foi ligado a nenhum consumidor. O teste era source-scan sobre pluginWorker.ts (ver docstring do arquivo deletado); sem o arquivo-fonte não há mais o que testar. OMNIROUTE_PLUGINS_ALLOW_EXEC também foi removido de .env.example e da doc na mesma release. Sem substituto porque a feature foi extinta, não migrada."
},
"tests/unit/plugins-sandbox.test.ts": {
"sourceRemoved": ["src/lib/plugins/sandbox.ts"],
"reason": "v3.8.50 #9126 (commit 8fac6bcd48): sandbox.ts foi removido por completo junto com pluginWorker.ts e signing.ts (\"zero importers confirmed\", subsistema de sandbox de plugins nunca ligado a nenhum consumidor). O teste cobria SandboxLevel/getSandboxLabel exportados por sandbox.ts; sem o arquivo-fonte não há mais símbolo a testar. Mesma causa-raiz de tests/unit/plugin-sandbox-permissions.test.ts nesta entrada."
}
},
"tests/unit/catalog-updates-v3x.test.ts": "v3.8.45 #6248: fix(providers) remove deprecated MiMo V2 entries — os 5 asserts removidos pinavam specs de modelos mimo-v2-* que deixaram de existir no catálogo (54→49). Asserts seguem a remoção dos modelos, não enfraquecimento. Verificado legítimo. Prune após v3.8.45 mergear para main.",
@@ -96,5 +108,6 @@
"tests/unit/usage-providers.test.ts": "v3.8.49 #7866: o case \"qwen\" saiu de getUsageForProvider (não há mais case \"qwen\" no switch de open-sse/services/usage.ts); o teste cobria esse ramo extinto (net 20→19). Verificado legítimo. Prune após v3.8.49 mergear para main.",
"tests/unit/usage-service-hardening.test.ts": "v3.8.49 #7866/#8565/#8013: qwen removido (3 asserts); o Kimi/Kiro builder-id (uso profileless) passou a ter SUCESSO real em vez de erro de ARN — supportsProfilelessKiroUsage(\"builder-id\") retorna true —, trocando 1 assert de regex de erro por 3 asserts de valor; e os ids de bucket de quota do Antigravity foram atualizados para o catálogo atual. Rodado no HEAD: 23/23 passam. Net 210→209. Verificado legítimo. Prune após v3.8.49 mergear para main.",
"tests/unit/virtual-auto-combo.test.ts": "v3.8.49 #7928/#8183: o pooling de contas passou a agrupar conexões web-session do mesmo provider numa entrada lógica com allowedConnectionIds (campo confirmado em open-sse/services/autoCombo/virtualFactory.ts), e o pool no-auth virou uma allowlist fixa (AUTO_COMBO_NOAUTH_ALLOWLIST = opencode, felo-web) — os testes antigos esperavam duplicatas e a inclusão de duckduckgo-web/theoldllm/chipotle, que hoje são corretamente excluídos. Guard dedicado em noauth-autocombo-allowlist.test.ts. Rodado no HEAD: 10/10 passam. Net 39→31. Verificado legítimo. Prune após v3.8.49 mergear para main.",
"open-sse/services/__tests__/tierResolver.test.ts": "v3.8.49 #7866: refactor(qwen) remove o provider OAuth legado — o teste \"classifies Qwen as free\" e a entrada de qwen na lista do batch saíram junto com o provider, e os índices do batch desceram de 10 para 9 elementos (net 61→59). Superfície extinta, não enfraquecimento. Verificado legítimo. Prune após v3.8.49 mergear para main."
"open-sse/services/__tests__/tierResolver.test.ts": "v3.8.49 #7866: refactor(qwen) remove o provider OAuth legado — o teste \"classifies Qwen as free\" e a entrada de qwen na lista do batch saíram junto com o provider, e os índices do batch desceram de 10 para 9 elementos (net 61→59). Superfície extinta, não enfraquecimento. Verificado legítimo. Prune após v3.8.49 mergear para main.",
"tests/unit/plugins-welcome-banner-e2e.test.ts": "v3.8.50 #9126 (commit 8fac6bcd48): o teste único 'BUILTIN_EVENTS has all 14 events' (13 asserts .ok/.equal) foi reestruturado em 3 testes mais específicos — 'contains only emitted/public events' (assert.deepEqual da lista completa), 'does not advertise dead events' (7 asserts .equal(false) para eventos sem emissor real: onModelSelect/onComboResolve/onRateLimit/onQuotaExhaust/onProviderError/onStreamStart/onStreamEnd) e 'lifecycle events remain represented' (4 asserts .ok). Contrato mais forte (agora também nega presença dos eventos mortos), não mais fraco — a contagem líquida cai (73→61) porque o assert.deepEqual único substitui múltiplos assert.ok redundantes com a mesma cobertura. Asserts restruturados, não removidos sem substituição. Verificado legítimo."
}

View File

@@ -0,0 +1,47 @@
---
title: "Management Authentication"
version: 3.8.50
lastUpdated: 2026-08-05
---
# Management Authentication
OmniRoute uses four distinct credential families for management access. This guide
distinguishes them by purpose, scope, and locality.
| Credential | Scope | Locality | Use Case |
|-------------------------|--------------------|---------------|-----------------------------------|
| Dashboard JWT session | Full management | Localhost | Web dashboard login |
| CLI machine-id token | Full management | Per-machine | `omniroute` CLI commands |
| Scoped `oma_` token | Configurable scope | External | Automation / CI / API access |
| Manage-scope API key | `manage` scope | External | Management API calls |
## Dashboard JWT Session
Generated on dashboard login (`/api/auth/login`). Stored in HTTP-only cookie.
Valid for the session duration. Cannot be used from external hosts.
## CLI Machine-ID Token
Created by `omniroute auth login` on first use. Stored in `~/.omniroute/auth.json`.
Used by the CLI for all management operations. Tied to the machine identity.
## Scoped `oma_` Access Token
Created via dashboard or CLI with configurable scopes (e.g., `manage`, `read`).
Format: `oma_<random-hex>`. Used for programmatic access from external systems.
## Manage-Scope API Key
Standard API key with the `manage` scope enabled. Created in dashboard API Keys page.
Used for management API calls from external hosts.
## Header Examples
```
Authorization: Bearer oma_abc123def456
Authorization: Bearer <standard-api-key-with-manage-scope>
Cookie: omniroute_session=<jwt-token>
```
See `docs/reference/API_REFERENCE.md` for endpoint-specific auth requirements.

View File

@@ -1,3 +1,9 @@
---
title: "AgentRouter WAF"
version: 3.8.50
lastUpdated: 2026-08-03
---
# agentrouter.org WAF (Web Application Firewall)
The `agentrouter` upstream gateway runs a keyword-based content filter on

View File

@@ -22,8 +22,7 @@ const LOCAL_DB_IMPORT_RESTRICTION = {
const EXECUTOR_IMPORT_RESTRICTION = {
regex: "^(?:@omniroute/)?open-sse/executors(?:/|$)",
message:
"Executor implementations must stay behind an open-sse handler or service boundary.",
message: "Executor implementations must stay behind an open-sse handler or service boundary.",
};
const PROP_TYPES_RESTRICTION = {
@@ -165,6 +164,14 @@ const eslintConfig = [
// their files move mid-scan, so never lint them from the main checkout.
".claude/**",
".omnivscodeagent/**",
// _tasks/ — planning/handoff/research artifacts (gitignored, external code)
"_tasks/**",
// .agents/ — skill definitions + their helper scripts (gitignored; the
// canonical copy lives here and is symlinked into .claude/).
".agents/**",
// .source/ — fumadocs codegen output (@ts-nocheck + bundler-only import
// query params like `?collection=docs`, which are not valid TS on their own).
".source/**",
// VS Code extension and its large test fixtures
"vscode-extension/**",
"_references/**",

View File

@@ -1,8 +0,0 @@
node_modules/
*.log
.DS_Store
test/
*.test.js
.env
.env.*

View File

@@ -24,6 +24,8 @@ const ANTHROPIC_BETA_BASE = Object.freeze([
"advisor-tool-2026-03-01",
"extended-cache-ttl-2025-04-11",
"cache-diagnosis-2026-04-07",
"code-execution-2025-08-25",
"skills-2025-10-02",
]);
const CLAUDE_OAUTH_EXTRA_BETAS = Object.freeze(["fine-grained-tool-streaming-2025-05-14"]);
@@ -53,6 +55,8 @@ export const ANTHROPIC_BETA_CLAUDE_OAUTH = [
export const FORWARDABLE_CLIENT_BETAS = Object.freeze([
"tool-search-tool-2025-10-19",
"context-1m-2025-08-07",
"code-execution-2025-08-25",
"skills-2025-10-02",
]);
/**

View File

@@ -12,16 +12,10 @@ export interface KimiWebModelConfig {
const STATIC_MODEL_CONFIGS: Record<string, KimiWebModelConfig> = {
k3: {
scenario: "SCENARIO_OK_COMPUTER",
kimiPlusId: "ok-computer",
supportedReasoningEfforts: [
"REASONING_EFFORT_LOW",
"REASONING_EFFORT_HIGH",
"REASONING_EFFORT_MAX",
],
defaultReasoningEffort: "REASONING_EFFORT_MAX",
supportedContextLengths: ["CONTEXT_LENGTH_L", "CONTEXT_LENGTH_XL"],
defaultContextLength: "CONTEXT_LENGTH_L",
scenario: "SCENARIO_K2D5",
supportedReasoningEfforts: ["REASONING_EFFORT_NONE", "REASONING_EFFORT_LOW"],
defaultReasoningEffort: "REASONING_EFFORT_NONE",
supportedContextLengths: [],
},
k2d6: {
scenario: "SCENARIO_K2D5",

View File

@@ -1,17 +1,14 @@
import type { RegistryEntry } from "../../../shared.ts";
import { getAnthropicCompatHeaders, ANTHROPIC_VERSION_HEADER } from "../../../shared.ts";
export const minimax_cnProvider: RegistryEntry = {
id: "minimax-cn",
alias: "minimax-cn", // unique alias (was colliding with minimax)
format: "claude",
format: "openai",
executor: "default",
baseUrl: "https://api.minimaxi.com/anthropic/v1/messages",
baseUrl: "https://api.minimaxi.com/v1/chat/completions",
modelsUrl: "https://api.minimaxi.com/v1/models",
urlSuffix: "?beta=true",
authType: "apikey",
authHeader: "bearer",
headers: getAnthropicCompatHeaders(),
models: [
// Keep parity with minimax to ensure model discovery works for minimax-cn connections.
// #3110: MiniMax M3 — frontier coding model with 1M context

View File

@@ -1,17 +1,14 @@
import type { RegistryEntry } from "../../shared.ts";
import { getAnthropicCompatHeaders, ANTHROPIC_VERSION_HEADER } from "../../shared.ts";
export const minimaxProvider: RegistryEntry = {
id: "minimax",
alias: "minimax",
format: "claude",
format: "openai",
executor: "default",
baseUrl: "https://api.minimax.io/anthropic/v1/messages",
baseUrl: "https://api.minimax.io/v1/chat/completions",
modelsUrl: "https://api.minimax.io/v1/models",
urlSuffix: "?beta=true",
authType: "apikey",
authHeader: "bearer",
headers: getAnthropicCompatHeaders(),
models: [
// T12/T28: MiniMax default upgraded from M2.5 to M2.7
// #3110: MiniMax M3 — frontier coding model with 1M context

View File

@@ -12,6 +12,18 @@ export const ollama_cloudProvider: RegistryEntry = {
// Note: rate limits vary by plan (free = "Light usage", Pro = more, Max = 5x Pro).
// Users can generate API keys at https://ollama.com/settings/keys
models: [
{
id: "gpt-oss:20b",
name: "GPT-OSS 20B",
supportsReasoning: true,
supportedThinkingEfforts: ["low", "medium", "high"],
},
{
id: "gpt-oss:120b",
name: "GPT-OSS 120B",
supportsReasoning: true,
supportedThinkingEfforts: ["low", "medium", "high"],
},
{ id: "deepseek-v4-pro", name: "DeepSeek V4 Pro", supportsReasoning: true },
{ id: "deepseek-v4-flash", name: "DeepSeek V4 Flash", supportsReasoning: true },
{ id: "kimi-k2.6", name: "Kimi K2.6" },

View File

@@ -15,7 +15,7 @@ export const perplexity_webProvider: RegistryEntry = {
{ id: "pplx-gpt-5.6-sol", name: "GPT-5.6 Sol (via Perplexity)", toolCalling: false },
{ id: "pplx-gemini", name: "Gemini 3.1 Pro (via Perplexity)", toolCalling: false },
{ id: "pplx-sonnet", name: "Claude Sonnet 5.0 (via Perplexity)", toolCalling: false },
{ id: "pplx-opus", name: "Claude Opus 4.8 (via Perplexity)", toolCalling: false },
{ id: "pplx-opus", name: "Claude Opus 5.0 (via Perplexity)", toolCalling: false },
{ id: "pplx-glm", name: "GLM-5.2 (via Perplexity)", toolCalling: false },
{ id: "pplx-kimi", name: "Kimi K2.6 (via Perplexity)", toolCalling: false },
{ id: "pplx-grok-4.5", name: "Grok 4.5 (via Perplexity)", toolCalling: false },

View File

@@ -48,6 +48,7 @@ export interface RegistryModel {
aliases?: readonly string[];
toolCalling?: boolean;
supportsReasoning?: boolean;
supportedThinkingEfforts?: readonly string[];
supportsVision?: boolean;
supportsXHighEffort?: boolean;
maxOutputTokens?: number;

View File

@@ -213,14 +213,21 @@ function makeErrorResponse(
details?: unknown;
type?: string;
code?: string;
extraHeaders?: Record<string, string>;
}
): Response {
const body = buildErrorBody(status, message, options?.details);
if (options?.type) body.error.type = options.type;
if (options?.code) body.error.code = options.code;
const headers: Record<string, string> = { "Content-Type": "application/json" };
if (options?.extraHeaders) {
for (const [key, value] of Object.entries(options.extraHeaders)) {
headers[key] = value;
}
}
return new Response(JSON.stringify(body), {
status,
headers: { "Content-Type": "application/json" },
headers,
});
}
@@ -302,7 +309,12 @@ async function errorResponseForTransport(
return makeErrorResponse(401, "Session expired or invalid");
}
if (result.status === 429) {
return makeErrorResponse(429, "Rate limited by Claude Web API");
const extraHeaders: Record<string, string> = {};
const upstreamRetryAfter = result.headers.get("retry-after");
if (upstreamRetryAfter) {
extraHeaders["Retry-After"] = upstreamRetryAfter;
}
return makeErrorResponse(429, "Rate limited by Claude Web API", { extraHeaders });
}
if (isClaudeWebChallenge({ ...result, bodyText })) {
return makeErrorResponse(403, "Claude Web returned a Cloudflare browser challenge", {

View File

@@ -217,6 +217,20 @@ function messageText(content: unknown): string {
return content.map(contentPartText).filter(Boolean).join("\n");
}
function buildPromptFromMessages(messages: unknown[]): string {
const parts: string[] = [];
for (const candidate of messages) {
if (!isRecord(candidate)) continue;
const role = candidate.role;
const text = messageText(candidate.content);
if (!text) continue;
if (role === "user" || role === "tool") {
parts.push(text);
}
}
return parts.join("\n\n");
}
function latestUserPrompt(messages: unknown[]): string {
let prompt = "";
for (const candidate of messages) {
@@ -308,7 +322,9 @@ export function transformToClaude(
const messages = Array.isArray(body.messages) ? body.messages : [];
const reasoningEffort = resolveClaudeWebReasoningEffort(body);
const resolvedModel = model || DEFAULT_CLAUDE_MODEL;
const resolvedTurn = turn ?? defaultTurn(latestUserPrompt(messages));
const prompt =
turn?.prompt ?? (buildPromptFromMessages(messages) || latestUserPrompt(messages));
const resolvedTurn = turn ?? defaultTurn(prompt);
if (resolvedTurn.operation === "completion" && !resolvedTurn.prompt.trim()) {
throw new Error("No user message found in request");

View File

@@ -13,14 +13,22 @@ export interface ClaudeWebStreamOptions {
}
type StreamPhase = "awaiting_message" | "in_message" | "stopped" | "failed";
type BlockKind = "thinking" | "text" | "other";
type BlockKind = "thinking" | "text" | "tool_use" | "other";
const MAX_CLAUDE_WEB_SSE_PENDING_CHARS = 1024 * 1024;
type SemanticEvent =
| { kind: "content"; text: string }
| { kind: "reasoning"; text: string }
| { kind: "tool_call"; index: number; id: string; name: string; input: string }
| { kind: "metadata"; eventType: string; data: Record<string, unknown> }
| { kind: "finish"; stopReason: string };
interface ToolBlockInfo {
id: string;
name: string;
inputParts: string[];
initialInput: string;
}
const KNOWN_METADATA_EVENTS = new Set([
"ping",
"completion",
@@ -193,6 +201,7 @@ function thinkingSummaryText(delta: Record<string, unknown>): string {
interface ProtocolState {
phase: StreamPhase;
openBlocks: Map<number, BlockKind>;
toolBlocks: Map<number, ToolBlockInfo>;
stopReason: string;
}
@@ -241,6 +250,7 @@ function handleMessageStart(state: ProtocolState): null {
function blockKind(block: Record<string, unknown>): BlockKind {
if (block.type === "thinking") return "thinking";
if (block.type === "text") return "text";
if (block.type === "tool_use") return "tool_use";
return "other";
}
@@ -252,17 +262,35 @@ function handleContentBlockStart(
const index = requireBlockIndex(event);
if (state.openBlocks.has(index)) protocolFailure(state, "Content block was opened twice");
const kind = blockKind(requireRecord(event.content_block, "content_block"));
const contentBlock = requireRecord(event.content_block, "content_block");
const kind = blockKind(contentBlock);
state.openBlocks.set(index, kind);
if (kind === "tool_use") {
const id = typeof contentBlock.id === "string" ? contentBlock.id : "";
const name = typeof contentBlock.name === "string" ? contentBlock.name : "";
let initialInput = "";
if (contentBlock.input !== undefined) {
try {
initialInput = JSON.stringify(contentBlock.input);
} catch {
initialInput = "";
}
}
state.toolBlocks.set(index, { id, name, inputParts: [], initialInput });
return null;
}
return kind === "thinking" ? { kind: "reasoning", text: "" } : null;
}
function handleContentBlockDelta(
event: Record<string, unknown>,
state: ProtocolState
): SemanticEvent {
): SemanticEvent | null {
assertInMessage(state, "content_block_delta");
const block = state.openBlocks.get(requireBlockIndex(event));
const index = requireBlockIndex(event);
const block = state.openBlocks.get(index);
if (!block) protocolFailure(state, "Content delta has no open block");
const delta = requireRecord(event.delta, "delta");
@@ -275,14 +303,42 @@ function handleContentBlockDelta(
if (delta.type === "thinking_summary_delta" && block === "thinking") {
return { kind: "reasoning", text: thinkingSummaryText(delta) };
}
if (delta.type === "input_json_delta" && block === "tool_use") {
const toolBlock = state.toolBlocks.get(index);
if (!toolBlock) protocolFailure(state, "input_json_delta has no tool block state");
if (typeof delta.partial_json === "string") {
toolBlock.inputParts.push(delta.partial_json);
}
return null;
}
return protocolFailure(state, "Content delta type does not match its block");
}
function handleContentBlockStop(event: Record<string, unknown>, state: ProtocolState): null {
function handleContentBlockStop(
event: Record<string, unknown>,
state: ProtocolState
): SemanticEvent | null {
assertInMessage(state, "content_block_stop");
if (!state.openBlocks.delete(requireBlockIndex(event))) {
protocolFailure(state, "Content block stop has no open block");
const index = requireBlockIndex(event);
const kind = state.openBlocks.get(index);
if (!kind) protocolFailure(state, "Content block stop has no open block");
state.openBlocks.delete(index);
if (kind === "tool_use") {
const toolBlock = state.toolBlocks.get(index);
state.toolBlocks.delete(index);
if (!toolBlock) protocolFailure(state, "Tool block stop has no tool state");
let inputStr = "";
if (toolBlock.inputParts.length > 0) {
inputStr = toolBlock.inputParts.join("");
} else if (toolBlock.initialInput) {
inputStr = toolBlock.initialInput;
}
return { kind: "tool_call", index, id: toolBlock.id, name: toolBlock.name, input: inputStr };
}
return null;
}
@@ -336,6 +392,7 @@ async function* parseClaudeWebEvents(
const state: ProtocolState = {
phase: "awaiting_message",
openBlocks: new Map(),
toolBlocks: new Map(),
stopReason: "end_turn",
};
@@ -447,6 +504,7 @@ async function createBufferedResponse(
let assistantText = "";
let reasoningText = "";
let stopReason = "end_turn";
const toolCalls: Array<{ id: string; name: string; input: string }> = [];
const metadataEvents: Array<{ type: string; data: Record<string, unknown> }> = [];
const control: StreamControl = { reader: null, cancelled: false };
@@ -454,12 +512,30 @@ async function createBufferedResponse(
for await (const event of parseClaudeWebEvents(source, control)) {
if (event.kind === "content") assistantText += event.text;
if (event.kind === "reasoning") reasoningText += event.text;
if (event.kind === "tool_call") {
toolCalls.push({ id: event.id, name: event.name, input: event.input });
}
if (event.kind === "metadata") {
metadataEvents.push({ type: event.eventType, data: event.data });
}
if (event.kind === "finish") stopReason = event.stopReason;
}
notifyComplete(options, { assistantText, stopReason });
const message: Record<string, unknown> = {
role: "assistant",
content: assistantText || null,
...(reasoningText ? { reasoning_content: reasoningText } : {}),
};
if (toolCalls.length > 0) {
message.tool_calls = toolCalls.map((tc) => ({
id: tc.id,
type: "function",
function: { name: tc.name, arguments: tc.input },
}));
}
return new Response(
JSON.stringify({
id,
@@ -469,11 +545,7 @@ async function createBufferedResponse(
choices: [
{
index: 0,
message: {
role: "assistant",
content: assistantText,
...(reasoningText ? { reasoning_content: reasoningText } : {}),
},
message,
finish_reason: openAiFinishReason(stopReason),
logprobs: null,
},
@@ -569,6 +641,31 @@ async function queueSemanticEvent(
);
return;
}
if (event.kind === "tool_call") {
state.pendingChunks.push(
encodeStreamEvent(
state,
makeChunk(
state.id,
state.created,
options,
{
tool_calls: [
{
index: event.index,
id: event.id,
type: "function",
function: { name: event.name, arguments: event.input },
},
],
},
null
)
)
);
return;
}
if (event.kind === "metadata") {
state.pendingChunks.push(
encodeStreamEvent(

View File

@@ -166,6 +166,13 @@ export interface ChatInvocationOptions {
tone?: string;
/** Tier-specific allowed message types; defaults to {@link ALLOWED_MESSAGE_TYPES}. */
allowedMessageTypes?: readonly string[];
/**
* Tier-specific disconnect behavior sent in every type:4 chat invocation. The work
* Surface rejects any value other than exactly "continue" (#8971). Defaults to ""
* for individual/consumer/EDU tiers; {@link resolveChatInvocationOverrides} returns
* "continue" for the enterprise tier.
*/
disconnectBehavior?: string;
}
/**
@@ -178,18 +185,21 @@ export function resolveChatInvocationOverrides(tier: string | undefined): {
optionsSets: string[];
tone: string;
allowedMessageTypes: readonly string[];
disconnectBehavior: string;
} {
if (tier === "enterprise") {
return {
optionsSets: [...M365_ENTERPRISE_OPTION_SETS],
tone: "Magic",
allowedMessageTypes: [...ALLOWED_MESSAGE_TYPES, ...M365_ENTERPRISE_EXTRA_MESSAGE_TYPES],
disconnectBehavior: "continue",
};
}
return {
optionsSets: [...M365_DEFAULT_OPTION_SETS],
tone: "",
allowedMessageTypes: ALLOWED_MESSAGE_TYPES,
disconnectBehavior: "",
};
}
@@ -253,7 +263,7 @@ export function buildChatInvocation(opts: ChatInvocationOptions): Record<string,
isSbsSupported: false,
tone: opts.tone ?? "",
renderReferencesBehindEOS: true,
disconnectBehavior: "",
disconnectBehavior: opts.disconnectBehavior ?? "",
},
],
};

View File

@@ -289,8 +289,6 @@ export class DefaultExecutor extends BaseExecutor {
case "glm":
case "glmt":
case "kimi-coding":
case "minimax":
case "minimax-cn":
return `${this.config.baseUrl}?beta=true`;
case "agentrouter":
return this.usesClaudeCodeProtocol(credentials)
@@ -395,9 +393,26 @@ export class DefaultExecutor extends BaseExecutor {
}
case "claude":
case "anthropic":
effectiveKey
? (headers["x-api-key"] = effectiveKey)
: (headers["Authorization"] = `Bearer ${credentials.accessToken}`);
if (effectiveKey) {
headers["x-api-key"] = effectiveKey;
// Port of decolua/9router commit b977bf74:
// Third-party Anthropic-compatible gateways frequently require
// Authorization: Bearer ALONGSIDE x-api-key — without it they
// return 401 missing_api_key on every forward. Only emit the
// Bearer fallback for non-official upstreams; api.anthropic.com
// (and the empty/default baseUrl that targets it) must keep the
// x-api-key-only behavior to avoid regressing the official path.
const baseUrl = credentials?.providerSpecificData?.baseUrl || "";
const isOfficial = isOfficialAnthropicBaseUrl(baseUrl);
if (!isOfficial) {
headers["Authorization"] = `Bearer ${effectiveKey}`;
}
} else if (credentials.accessToken) {
headers["Authorization"] = `Bearer ${credentials.accessToken}`;
}
// If neither effectiveKey nor accessToken is available, emit no
// auth header — the handler will produce a clean "no credentials"
// 4xx instead of forwarding garbage auth headers to the upstream.
break;
case "glm":
case "glmt":

View File

@@ -348,6 +348,30 @@ export class GeminiWebExecutor extends BaseExecutor {
super("gemini-web", { id: "gemini-web", baseUrl: GEMINI_URL });
}
/**
* testConnection — validates the cookie format without making a network call
* or launching Playwright. Returns true when the cookie is non-empty and
* contains at least one name=value pair with a non-empty value. This is a
* lightweight pre-check before the browser automation path; full session
* validation is done by validateGeminiWebProvider in the connection test
* flow (#9407).
*/
async testConnection(
credentials: Record<string, unknown>,
_signal?: AbortSignal
): Promise<boolean> {
try {
const cookie = resolveGeminiWebCookie(
credentials as unknown as ExecuteInput["credentials"]
);
if (!cookie) return false;
const pairs = parseCookies(cookie);
return pairs.some((p) => p.value.length > 0);
} catch {
return false;
}
}
/**
* Read the live Playwright cookie jar back after a successful run and, if
* Google rotated any of the __Secure-1PSID* cookies, forward the merged
@@ -593,6 +617,30 @@ export class GeminiWebExecutor extends BaseExecutor {
transformedBody: body,
};
}
// #9407: Playwright selector/click timeout errors are terminal — they indicate
// the page DOM does not match expectations (e.g. Gemini changed their UI or
// the session is so expired it lands on a different page). Return 400 so the
// account-fallback system does NOT retry this request as a transient 5xx.
if (
error instanceof Error &&
(error.name === "TimeoutError" ||
rawMessage.includes("waitForSelector") ||
rawMessage.includes("Timeout") ||
rawMessage.includes("actionability") ||
rawMessage.includes("interception"))
) {
return {
response: new Response(
JSON.stringify({
error: sanitizeErrorMessage(rawMessage),
}),
{ status: 400, headers: { "Content-Type": "application/json" } }
),
url: GEMINI_URL,
headers: {},
transformedBody: body,
};
}
return {
response: new Response(
JSON.stringify({

View File

@@ -388,7 +388,10 @@ export class PerplexityWebExecutor extends BaseExecutor {
let pplxMode: string;
let modelPref: string;
if (thinking && THINKING_MAP[model]) {
pplxMode = "search";
// "copilot", not "search": the backend downgrades "search" to CONCISE and drops
// model_preference, so the thinking variant would fail the same way the catalog
// models do (see the note above MODEL_MAP).
pplxMode = "copilot";
modelPref = THINKING_MAP[model];
log?.info?.("PPLX-WEB", `Thinking mode → ${model} using ${modelPref}`);
} else if (MODEL_MAP[model]) {

View File

@@ -51,31 +51,40 @@ export const PPLX_STREAM_EOF_SYMBOL = "event: end_of_stream";
export const PPLX_USER_AGENT =
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:148.0) Gecko/20100101 Firefox/148.0";
// mode / model_preference pairs. Live www.perplexity.ai still posts mode:"copilot"
// for the default turbo path; search mode is used for the curated catalog models.
// mode / model_preference pairs — every entry posts mode:"copilot", like the live
// www.perplexity.ai client does when a model is picked from the catalog.
//
// mode:"search" must NOT be used here. The backend now downgrades it to CONCISE and
// drops model_preference entirely, answering with status:"FAILED" and the text
// "Error in processing query." Verified against a paid `subscription_tier: "max"`
// account: mode:"search" + claude50sonnet → {"mode":"CONCISE","status":"FAILED"},
// while mode:"copilot" + the same preference → {"mode":"COPILOT",
// "display_model":"claude50sonnet"} and a normal stream. Same for every other
// catalog model, so "search" breaks the whole catalog, not just one entry.
export const MODEL_MAP: Record<string, [string, string]> = {
// pplx-auto/pplx-sonar use "copilot" mode (was "search", which for pplx-sonar
// maps to "experimental" — that model no longer streams answer-text blocks
// for many sessions → empty content, issue #6955). The live web client uses
// mode:"copilot" + model_preference:"turbo" for the default turbo path.
// pplx-auto/pplx-sonar were already on "copilot" (with "search", pplx-sonar maps to
// "experimental" — that model no longer streams answer-text blocks for many
// sessions → empty content, issue #6955).
"pplx-auto": ["copilot", "pplx_pro"],
"pplx-sonar": ["copilot", "turbo"],
"pplx-gpt-5.6-terra": ["search", "gpt56_terra"],
"pplx-gpt-5.6-sol": ["search", "gpt56_sol"],
"pplx-gemini": ["search", "gemini31pro_high"],
"pplx-sonnet": ["search", "claude50sonnet"],
"pplx-opus": ["search", "claude48opus"],
"pplx-glm": ["search", "glm_5_2"],
"pplx-kimi": ["search", "kimik26instant"],
"pplx-grok-4.5": ["search", "grok45low"],
"pplx-nemotron": ["search", "nv_nemotron_3_ultra"],
"pplx-gpt-5.6-terra": ["copilot", "gpt56_terra"],
"pplx-gpt-5.6-sol": ["copilot", "gpt56_sol"],
"pplx-gemini": ["copilot", "gemini31pro_high"],
"pplx-sonnet": ["copilot", "claude50sonnet"],
// Perplexity's catalog moved Opus to 5.0; claude48opus is still accepted but
// answers from the older model.
"pplx-opus": ["copilot", "claude50opus"],
"pplx-glm": ["copilot", "glm_5_2"],
"pplx-kimi": ["copilot", "kimik26instant"],
"pplx-grok-4.5": ["copilot", "grok45low"],
"pplx-nemotron": ["copilot", "nv_nemotron_3_ultra"],
};
export const THINKING_MAP: Record<string, string> = {
"pplx-gpt-5.6-terra": "gpt56_terra_thinking",
"pplx-gpt-5.6-sol": "gpt56_sol_thinking",
"pplx-sonnet": "claude50sonnetthinking",
"pplx-opus": "claude48opusthinking",
"pplx-opus": "claude50opusthinking",
"pplx-kimi": "kimik26thinking",
"pplx-grok-4.5": "grok45medium",
};

View File

@@ -67,7 +67,7 @@ import {
resolveMemoryOwnerId,
} from "./chatCore/memoryExtraction.ts";
import { CORS_HEADERS } from "../utils/cors.ts";
import { checkHeapPressureGuard } from "../utils/heapPressure.ts";
import { checkResourcePressureGuard } from "../utils/resourcePressure.ts";
import { normalizeHeaders } from "../utils/headers.ts";
import { resolveChatCoreRequestFormat } from "./chatCore/requestFormat.ts";
import { resolveChatCoreTargetFormat } from "./chatCore/targetFormat.ts";
@@ -359,13 +359,6 @@ import {
isRpmExhausted,
} from "../services/geminiRateLimitTracker.ts";
// ── Global memory pressure guard ────────────────────────────────────────
// Prevents OOM by rejecting new requests when V8 heap exceeds threshold.
// Self-healing: no counters to leak, no cleanup needed. The threshold
// auto-calibrates to 85% of the actual V8 heap ceiling (see heapPressure.ts) so
// it tracks --max-old-space-size across 1GB/2GB/large VPS instead of a fixed
// 200MB that sat below the app's own ~260MB baseline and rejected every request.
import { isSmallEnoughForSemanticCache } from "../utils/estimateSize.ts";
/**
@@ -415,17 +408,16 @@ export async function handleChatCore({
createPiiTransform = null,
correlationId = null,
modelPinned = false,
skipResourcePressureGuard = false,
}) {
let { provider, model, extendedContext } = modelInfo;
// ── Memory pressure guard ────────────────────────────────────────────
// Reject early if V8 heap is already near the 256MB limit. Prevents
// cascading OOM when many large-context requests arrive concurrently.
try {
const heapUsedMB = process.memoryUsage().heapUsed / (1024 * 1024);
const heapGuard = checkHeapPressureGuard(heapUsedMB);
if (heapGuard) return heapGuard;
} catch {
/* memoryUsage() never throws */
if (!skipResourcePressureGuard) {
try {
const pressureGuard = checkResourcePressureGuard();
if (pressureGuard) return pressureGuard;
} catch {
/* fail open */
}
}
// Per-request model-routing metadata (first extracted slice of the request-setup phase).

View File

@@ -41,8 +41,8 @@ function extractSystemTexts(body: Record<string, unknown> | null | undefined): s
* True when the inbound request should be default-allowed without calling upstream.
*
* - `mode === "off"` (default): never short-circuits.
* - `mode === "always"`: short-circuits every Claude-format request (operator has
* decided every `/v1/messages` call through this route is the classifier).
* - `mode === "always"`: short-circuits only when the request carries the classifier's
* system-prompt marker (same body-awareness as "auto").
* - `mode === "auto"`: only short-circuits when the request carries the classifier's
* system-prompt marker. `</block>` in `stop_sequences` is corroborating evidence but
* is never sufficient alone — the marker is the strong, classifier-unique signal;
@@ -56,7 +56,6 @@ export function shouldDefaultAllowClassifier(
): boolean {
if (mode !== "auto" && mode !== "always") return false;
if (sourceFormat !== FORMATS.CLAUDE) return false;
if (mode === "always") return true;
return extractSystemTexts(body).some((text) => text.includes(SECURITY_MONITOR_MARKER));
}

View File

@@ -0,0 +1,168 @@
import type { AdmissionPressure, AdmissionReleaseOutcome } from "./types.ts";
export interface AdaptationParams {
minLimit: number;
maxLimit: number;
windowMs: number;
shortLatencyAlpha: number;
longLatencyAlpha: number;
increaseStep: number;
decreaseFactor: number;
criticalDecreaseFactor: number;
highUtilizationThreshold: number;
lowUtilizationThreshold: number;
latencyGradientThreshold: number;
maxIncreasePerWindow: number;
}
export interface AdaptationState {
currentLimit: number;
shortLatencyEwma: number;
longLatencyEwma: number;
pressure: AdmissionPressure;
/** Sum of admitted cost * time contribution proxies in the open window. */
windowActiveCostIntegral: number;
windowCompleted: number;
windowLatencySamples: number;
windowStartMs: number;
freezeGrowth: boolean;
/**
* When true, critical multiplicative decrease already applied for this window
* (e.g. via immediate observePressure). Window close must not re-apply it.
*/
criticalDecreaseConsumed: boolean;
utilization: number;
}
export function clampLimit(value: number, minLimit: number, maxLimit: number): number {
if (!Number.isFinite(value)) return minLimit;
return Math.min(maxLimit, Math.max(minLimit, Math.floor(value)));
}
export function createAdaptationState(
initialLimit: number,
minLimit: number,
maxLimit: number,
nowMs: number
): AdaptationState {
return {
currentLimit: clampLimit(initialLimit, minLimit, maxLimit),
shortLatencyEwma: 0,
longLatencyEwma: 0,
pressure: "normal",
windowActiveCostIntegral: 0,
windowCompleted: 0,
windowLatencySamples: 0,
windowStartMs: nowMs,
freezeGrowth: false,
criticalDecreaseConsumed: false,
utilization: 0,
};
}
export function noteLatency(
state: AdaptationState,
latencyMs: number,
params: AdaptationParams
): void {
const sample = Number.isFinite(latencyMs) && latencyMs >= 0 ? latencyMs : 0;
state.windowLatencySamples += 1;
const sa = params.shortLatencyAlpha;
const la = params.longLatencyAlpha;
if (state.shortLatencyEwma <= 0 && state.longLatencyEwma <= 0) {
state.shortLatencyEwma = sample;
state.longLatencyEwma = sample;
return;
}
state.shortLatencyEwma = sa * sample + (1 - sa) * state.shortLatencyEwma;
state.longLatencyEwma = la * sample + (1 - la) * state.longLatencyEwma;
}
export function noteOutcome(state: AdaptationState, outcome: AdmissionReleaseOutcome): void {
// A single upstream business error freezes growth for the current window; it must not
// apply critical multiplicative collapse on its own.
if (outcome === "upstream_error") {
state.freezeGrowth = true;
return;
}
if (outcome === "timeout") {
state.freezeGrowth = true;
}
}
export function setPressure(state: AdaptationState, pressure: AdmissionPressure): void {
const severity: Record<AdmissionPressure, number> = { normal: 0, high: 1, critical: 2 };
if (severity[pressure] > severity[state.pressure]) state.pressure = pressure;
}
/**
* Close the current feedback window and adjust the limit.
* Recovery (increase) is slower than decrease; idle/low utilization does not inflate.
*/
export function closeAdaptationWindow(
state: AdaptationState,
params: AdaptationParams,
nowMs: number
): void {
const elapsed = Math.max(1, Math.min(params.windowMs, nowMs - state.windowStartMs));
// sampleActiveIntegral already accounts for every interval exactly once.
const avgActive = state.windowActiveCostIntegral / elapsed;
const util = state.currentLimit > 0 ? avgActive / state.currentLimit : 0;
state.utilization = Math.max(0, Math.min(1, util));
let next = state.currentLimit;
const gradient =
state.longLatencyEwma > 0
? (state.shortLatencyEwma - state.longLatencyEwma) / state.longLatencyEwma
: 0;
if (state.pressure === "critical") {
// Immediate observePressure may already have applied the critical factor once.
if (!state.criticalDecreaseConsumed) {
next = Math.floor(next * params.criticalDecreaseFactor);
}
} else if (
state.pressure === "high" ||
(state.windowLatencySamples > 0 && gradient >= params.latencyGradientThreshold)
) {
next = Math.floor(next * params.decreaseFactor);
} else if (
!state.freezeGrowth &&
state.pressure === "normal" &&
state.utilization >= params.highUtilizationThreshold &&
state.windowCompleted > 0
) {
const step = Math.min(params.increaseStep, params.maxIncreasePerWindow);
next = next + step;
}
// A genuinely low-utilization window recovers the latency baseline so stale gradients expire.
if (state.utilization <= params.lowUtilizationThreshold) {
state.shortLatencyEwma = state.longLatencyEwma;
}
state.currentLimit = clampLimit(next, params.minLimit, params.maxLimit);
state.windowActiveCostIntegral = 0;
state.windowCompleted = 0;
state.windowLatencySamples = 0;
state.windowStartMs = nowMs;
state.freezeGrowth = false;
state.criticalDecreaseConsumed = false;
state.pressure = "normal";
}
export function sampleActiveIntegral(
state: AdaptationState,
activeCost: number,
dtMs: number
): void {
if (dtMs <= 0 || activeCost <= 0) return;
const boundedActiveCost = Math.min(activeCost, state.currentLimit);
const contribution =
dtMs > Math.floor(Number.MAX_SAFE_INTEGER / boundedActiveCost)
? Number.MAX_SAFE_INTEGER
: boundedActiveCost * dtMs;
state.windowActiveCostIntegral =
contribution >= Number.MAX_SAFE_INTEGER - state.windowActiveCostIntegral
? Number.MAX_SAFE_INTEGER
: state.windowActiveCostIntegral + contribution;
}

View File

@@ -0,0 +1,167 @@
import { resolveCostConfig } from "./cost.ts";
import {
MAX_ADMISSION_COST_OR_LIMIT,
MAX_ADMISSION_WINDOW_MS,
type AdaptiveAdmissionConfig,
type AdmissionMode,
} from "./types.ts";
import type { AdaptationParams } from "./adaptation.ts";
export { MAX_ADMISSION_COST_OR_LIMIT, MAX_ADMISSION_WINDOW_MS };
export interface ValidatedConfig {
mode: AdmissionMode;
minLimit: number;
maxLimit: number;
initialLimit: number;
maxQueueCount: number;
maxQueueCost: number;
defaultMaxWaitMs: number;
windowMs: number;
adaptation: AdaptationParams;
maxRequestCost: number;
costConfig: ReturnType<typeof resolveCostConfig>;
}
function requirePositiveInt(
name: string,
value: unknown,
max: number = MAX_ADMISSION_COST_OR_LIMIT
): number {
if (
typeof value !== "number" ||
!Number.isFinite(value) ||
value <= 0 ||
!Number.isSafeInteger(value)
) {
throw new RangeError(`${name} must be a positive safe integer`);
}
if (value > max) {
throw new RangeError(`${name} must be <= ${max}`);
}
return value;
}
function requireUnitInterval(name: string, value: unknown, fallback: number): number {
if (value === undefined) return fallback;
if (typeof value !== "number" || !Number.isFinite(value) || value <= 0 || value > 1) {
throw new RangeError(`${name} must be in (0, 1]`);
}
return value;
}
function requireDecreaseFactor(name: string, value: unknown, fallback: number): number {
if (value === undefined) return fallback;
if (typeof value !== "number" || !Number.isFinite(value) || value <= 0 || value >= 1) {
throw new RangeError(`${name} must be in (0, 1)`);
}
return value;
}
function resolveMode(mode: AdaptiveAdmissionConfig["mode"]): AdmissionMode {
if (mode === undefined) return "shadow";
if (mode !== "off" && mode !== "shadow" && mode !== "enforce") {
throw new RangeError("mode must be off|shadow|enforce");
}
return mode;
}
function resolveAdaptationParams(
input: AdaptiveAdmissionConfig,
minLimit: number,
maxLimit: number,
windowMs: number
): AdaptationParams {
const decreaseFactor = requireDecreaseFactor("decreaseFactor", input.decreaseFactor, 0.8);
const criticalDecreaseFactor = requireDecreaseFactor(
"criticalDecreaseFactor",
input.criticalDecreaseFactor,
0.5
);
const increaseStep =
input.increaseStep === undefined ? 1 : requirePositiveInt("increaseStep", input.increaseStep);
const maxIncreasePerWindow =
input.maxIncreasePerWindow === undefined
? increaseStep
: requirePositiveInt("maxIncreasePerWindow", input.maxIncreasePerWindow);
const shortLatencyAlpha = requireUnitInterval("shortLatencyAlpha", input.shortLatencyAlpha, 0.5);
const longLatencyAlpha = requireUnitInterval("longLatencyAlpha", input.longLatencyAlpha, 0.1);
const highUtilizationThreshold = requireUnitInterval(
"highUtilizationThreshold",
input.highUtilizationThreshold,
0.7
);
const lowUtilizationThreshold = requireUnitInterval(
"lowUtilizationThreshold",
input.lowUtilizationThreshold,
0.3
);
if (criticalDecreaseFactor > decreaseFactor) {
throw new RangeError("criticalDecreaseFactor must be <= decreaseFactor");
}
if (lowUtilizationThreshold >= highUtilizationThreshold) {
throw new RangeError("lowUtilizationThreshold must be < highUtilizationThreshold");
}
if (shortLatencyAlpha <= longLatencyAlpha) {
throw new RangeError("shortLatencyAlpha must be > longLatencyAlpha");
}
return {
minLimit,
maxLimit,
windowMs,
shortLatencyAlpha,
longLatencyAlpha,
increaseStep,
decreaseFactor,
criticalDecreaseFactor,
highUtilizationThreshold,
lowUtilizationThreshold,
latencyGradientThreshold: requireUnitInterval(
"latencyGradientThreshold",
input.latencyGradientThreshold,
0.25
),
maxIncreasePerWindow,
};
}
export function validateConfig(input: AdaptiveAdmissionConfig): ValidatedConfig {
const minLimit = requirePositiveInt("minLimit", input.minLimit);
const maxLimit = requirePositiveInt("maxLimit", input.maxLimit);
if (minLimit > maxLimit) {
throw new RangeError("minLimit must be <= maxLimit");
}
const initialLimit = requirePositiveInt("initialLimit", input.initialLimit);
// Queue count is not multiplied into cost×time products; keep the full safe-integer range.
const maxQueueCount = requirePositiveInt(
"maxQueueCount",
input.maxQueueCount,
Number.MAX_SAFE_INTEGER
);
const maxQueueCost = requirePositiveInt("maxQueueCost", input.maxQueueCost);
const windowMs =
input.windowMs === undefined
? 1000
: requirePositiveInt("windowMs", input.windowMs, MAX_ADMISSION_WINDOW_MS);
const defaultMaxWaitMs =
input.defaultMaxWaitMs === undefined
? 5_000
: requirePositiveInt("defaultMaxWaitMs", input.defaultMaxWaitMs, MAX_ADMISSION_WINDOW_MS);
const costConfig = resolveCostConfig(input.cost);
return {
mode: resolveMode(input.mode),
minLimit,
maxLimit,
initialLimit,
maxQueueCount,
maxQueueCost,
defaultMaxWaitMs,
windowMs,
maxRequestCost: costConfig.maxRequestCost,
costConfig,
adaptation: resolveAdaptationParams(input, minLimit, maxLimit, windowMs),
};
}

View File

@@ -0,0 +1,624 @@
import {
closeAdaptationWindow,
createAdaptationState,
noteLatency,
noteOutcome,
sampleActiveIntegral,
setPressure,
type AdaptationState,
} from "./adaptation.ts";
import { validateConfig, type ValidatedConfig } from "./config.ts";
import { estimateAdmissionCost, normalizeRequestCost } from "./cost.ts";
import { FairCostQueue, type QueueEntry } from "./queue.ts";
import {
MAX_ADMISSION_WINDOW_MS,
createAdmissionRejectError,
type AdaptiveAdmissionConfig,
type AdmissionAcquireResult,
type AdmissionAdmitted,
type AdmissionClock,
type AdmissionLease,
type AdmissionPressure,
type AdmissionRejectCode,
type AdmissionReleaseMeta,
type AdmissionReleaseOutcome,
type AdmissionRequest,
type AdmissionSnapshot,
type ShadowDecision,
} from "./types.ts";
type VirtualDisposition = "active" | "queued" | "rejected" | "none";
const MAX_SAFE_BIGINT = BigInt(Number.MAX_SAFE_INTEGER);
/** Snapshot numbers are always finite safe integers; never emit rounded unsafe Number values. */
function saturateSnapshotNumber(value: number): number {
if (!Number.isFinite(value) || value <= 0) return 0;
if (value >= Number.MAX_SAFE_INTEGER) return Number.MAX_SAFE_INTEGER;
return Math.floor(value);
}
function bigintToSnapshotNumber(value: bigint): number {
if (value <= 0n) return 0;
if (value >= MAX_SAFE_BIGINT) return Number.MAX_SAFE_INTEGER;
return Number(value);
}
function addSaturated(total: number, delta: number): number {
if (delta <= 0) return saturateSnapshotNumber(total);
if (total >= Number.MAX_SAFE_INTEGER - delta) return Number.MAX_SAFE_INTEGER;
return total + delta;
}
interface ActiveLeaseRecord {
id: string;
cost: number;
released: boolean;
admittedAtMs: number;
virtualDisposition: VirtualDisposition;
}
interface QueuedPayload {
resolve: (value: AdmissionAdmitted) => void;
reject: (err: Error) => void;
signal?: AbortSignal;
onAbort?: () => void;
}
let leaseSeq = 0;
function nextId(prefix: string): string {
leaseSeq += 1;
return `${prefix}-${leaseSeq}`;
}
function defaultClock(): AdmissionClock {
return {
now: () => Date.now(),
setTimer: (fn, delayMs) => {
const handle = setTimeout(fn, delayMs);
// Window/deadline timers must not pin the event loop open when idle.
if (typeof handle.unref === "function") handle.unref();
return handle;
},
clearTimer: (id) => clearTimeout(id as ReturnType<typeof setTimeout>),
};
}
/**
* Dependency-injected weighted adaptive admission controller.
* Pure in-process core: no env/settings/route wiring.
*/
export class AdaptiveAdmissionController {
private config: ValidatedConfig;
private readonly clock: AdmissionClock;
private adaptation: AdaptationState;
private queue: FairCostQueue<QueuedPayload>;
private virtualQueue: FairCostQueue<{ recordId: string }>;
private readonly active = new Map<string, ActiveLeaseRecord>();
private activeCost = 0n;
private virtualActiveCost = 0;
private virtualActiveCount = 0;
private lastSampleMs: number;
private windowTimer: unknown = undefined;
private shutDown = false;
private admittedCount = 0;
private rejectedCount = 0;
private wouldAdmitCount = 0;
private wouldQueueCount = 0;
private wouldRejectCount = 0;
constructor(config: AdaptiveAdmissionConfig, clock?: Partial<AdmissionClock>) {
this.config = validateConfig(config);
this.clock = {
now: clock?.now ?? defaultClock().now,
setTimer: clock?.setTimer ?? defaultClock().setTimer,
clearTimer: clock?.clearTimer ?? defaultClock().clearTimer,
};
const now = this.clock.now();
this.adaptation = createAdaptationState(
this.config.initialLimit,
this.config.minLimit,
this.config.maxLimit,
now
);
this.queue = new FairCostQueue(this.config.maxQueueCount, this.config.maxQueueCost);
this.virtualQueue = new FairCostQueue(this.config.maxQueueCount, this.config.maxQueueCost);
this.lastSampleMs = now;
this.armWindowTimer();
}
updateConfig(config: AdaptiveAdmissionConfig): void {
const next = validateConfig(config);
this.sampleIntegral();
this.config = next;
this.adaptation.currentLimit = Math.min(
next.maxLimit,
Math.max(next.minLimit, this.adaptation.currentLimit)
);
this.adaptation.windowStartMs = this.clock.now();
this.adaptation.windowActiveCostIntegral = 0;
this.adaptation.windowCompleted = 0;
this.adaptation.windowLatencySamples = 0;
this.adaptation.freezeGrowth = false;
this.adaptation.criticalDecreaseConsumed = false;
this.adaptation.pressure = "normal";
this.lastSampleMs = this.clock.now();
const drained = this.queue.drain();
this.queue = new FairCostQueue(next.maxQueueCount, next.maxQueueCost);
for (const entry of drained) {
if (next.mode !== "enforce") {
this.clearEntryTimer(entry);
this.detachAbort(entry);
entry.payload.resolve(this.admit(entry.cost));
continue;
}
// Cost above the new enforce limit must fail closed immediately, never strand until deadline.
if (entry.cost > this.adaptation.currentLimit) {
this.failQueued(
entry,
"ADMISSION_OVERSIZED",
"request cost exceeds max budget after config update"
);
continue;
}
if (!this.queue.enqueue(entry)) {
this.failQueued(entry, "ADMISSION_QUEUE_FULL", "queue capacity reduced");
}
}
this.rebuildVirtualState(next.mode === "shadow");
this.armWindowTimer();
if (next.mode === "enforce") {
this.dispatch();
}
}
snapshot(): AdmissionSnapshot {
this.sampleIntegral();
return {
mode: this.config.mode,
currentLimit: this.adaptation.currentLimit,
minLimit: this.config.minLimit,
maxLimit: this.config.maxLimit,
activeCost: bigintToSnapshotNumber(this.activeCost),
activeCount: saturateSnapshotNumber(this.active.size),
queuedCost: saturateSnapshotNumber(this.queue.totalCost),
queuedCount: saturateSnapshotNumber(this.queue.size),
virtualActiveCost: saturateSnapshotNumber(this.virtualActiveCost),
virtualActiveCount: saturateSnapshotNumber(this.virtualActiveCount),
virtualQueuedCost: saturateSnapshotNumber(this.virtualQueue.totalCost),
virtualQueuedCount: saturateSnapshotNumber(this.virtualQueue.size),
admittedCount: saturateSnapshotNumber(this.admittedCount),
rejectedCount: saturateSnapshotNumber(this.rejectedCount),
wouldAdmitCount: saturateSnapshotNumber(this.wouldAdmitCount),
wouldQueueCount: saturateSnapshotNumber(this.wouldQueueCount),
wouldRejectCount: saturateSnapshotNumber(this.wouldRejectCount),
shortLatencyEwma: this.adaptation.shortLatencyEwma,
longLatencyEwma: this.adaptation.longLatencyEwma,
utilization: this.adaptation.utilization,
pressure: this.adaptation.pressure,
shutdown: this.shutDown,
};
}
observePressure(pressure: AdmissionPressure): void {
setPressure(this.adaptation, pressure);
if (pressure === "critical") {
// Immediate fast decrease once per window; window close must not re-apply it.
if (!this.adaptation.criticalDecreaseConsumed) {
this.adaptation.currentLimit = Math.max(
this.config.minLimit,
Math.floor(this.adaptation.currentLimit * this.config.adaptation.criticalDecreaseFactor)
);
this.adaptation.criticalDecreaseConsumed = true;
this.dispatch();
this.dispatchVirtual();
}
}
}
/** Deterministic window tick for tests / injected clocks. */
tick(): void {
this.sampleIntegral();
closeAdaptationWindow(this.adaptation, this.config.adaptation, this.clock.now());
// Real queue first, then virtual: raised limits must promote shadow-queued work
// before newer arrivals are classified against the updated budget.
this.dispatch();
this.dispatchVirtual();
}
async acquire(request: AdmissionRequest): Promise<AdmissionAcquireResult> {
if (this.shutDown) {
return this.reject("ADMISSION_SHUTDOWN", "admission controller is shut down");
}
if (request.signal?.aborted) {
return this.reject("ADMISSION_ABORTED", "request aborted before acquire");
}
if (request.pressure) setPressure(this.adaptation, request.pressure);
const cost = this.resolveCost(request);
const mode = this.config.mode;
if (mode === "off") {
return this.admitVirtual(cost);
}
const limit = this.adaptation.currentLimit;
if (mode === "shadow") {
return this.acquireShadow(request, cost, limit);
}
// enforce
if (cost > limit) {
return this.reject("ADMISSION_OVERSIZED", "request cost exceeds max budget");
}
// Once work is queued, every newer request joins the same fair queue even if it
// currently fits. This makes bounded bypass accounting effective and prevents
// direct arrivals from indefinitely jumping an older reserved weighted request.
if (this.queue.size === 0 && this.activeCost + BigInt(cost) <= BigInt(limit)) {
return this.admit(cost);
}
if (!this.queue.canAccept(cost)) {
return this.reject("ADMISSION_QUEUE_FULL", "admission queue is full");
}
return this.enqueue(request, cost);
}
shutdown(): void {
if (this.shutDown) return;
this.shutDown = true;
if (this.windowTimer !== undefined) {
this.clock.clearTimer(this.windowTimer);
this.windowTimer = undefined;
}
const drained = this.queue.drain();
for (const entry of drained) {
this.clearEntryTimer(entry);
this.detachAbort(entry);
entry.payload.reject(
createAdmissionRejectError("ADMISSION_SHUTDOWN", "admission controller shut down")
);
this.rejectedCount += 1;
}
}
private resolveCost(request: AdmissionRequest): number {
if (request.cost !== undefined) {
return normalizeRequestCost(request.cost, this.config.maxRequestCost);
}
if (request.features) {
return estimateAdmissionCost(request.features, this.config.costConfig);
}
return 1;
}
private acquireShadow(request: AdmissionRequest, cost: number, limit: number): AdmissionAdmitted {
let decision: ShadowDecision;
let disposition: VirtualDisposition;
if (cost > limit || !Number.isSafeInteger(cost)) {
decision = "would-reject";
disposition = "rejected";
this.wouldRejectCount += 1;
} else if (this.virtualActiveCost + cost <= limit) {
decision = "would-admit";
disposition = "active";
this.virtualActiveCost = addSaturated(this.virtualActiveCost, cost);
this.virtualActiveCount = addSaturated(this.virtualActiveCount, 1);
this.wouldAdmitCount = addSaturated(this.wouldAdmitCount, 1);
} else if (this.virtualQueue.canAccept(cost)) {
decision = "would-queue";
disposition = "queued";
this.wouldQueueCount += 1;
} else {
decision = "would-reject";
disposition = "rejected";
this.wouldRejectCount += 1;
}
const admitted = this.admit(cost, disposition);
if (disposition === "queued") {
this.virtualQueue.enqueue({
id: admitted.lease.id,
tenantKey: request.tenantKey || "_default",
cost,
enqueuedAtMs: this.clock.now(),
deadlineMs: Number.MAX_SAFE_INTEGER,
payload: { recordId: admitted.lease.id },
});
}
return { ...admitted, shadowDecision: decision };
}
private admitVirtual(cost: number): AdmissionAdmitted {
// Mode off: no accounting.
const id = nextId("lease");
const lease: AdmissionLease = {
id,
cost,
get released() {
return true;
},
release: () => {
/* no-op */
},
};
this.admittedCount += 1;
return { status: "admitted", lease };
}
private admit(cost: number, virtualDisposition: VirtualDisposition = "none"): AdmissionAdmitted {
this.sampleIntegral();
const id = nextId("lease");
const record: ActiveLeaseRecord = {
id,
cost,
released: false,
admittedAtMs: this.clock.now(),
virtualDisposition,
};
this.active.set(id, record);
this.activeCost += BigInt(cost);
this.admittedCount += 1;
const controller = this;
const lease: AdmissionLease = {
id,
cost,
get released() {
return record.released;
},
release(outcome: AdmissionReleaseOutcome = "success", meta?: AdmissionReleaseMeta) {
controller.releaseLease(record, outcome, meta);
},
};
return { status: "admitted", lease };
}
private releaseLease(
record: ActiveLeaseRecord,
outcome: AdmissionReleaseOutcome,
meta?: AdmissionReleaseMeta
): void {
if (record.released) return;
record.released = true;
// Sample while the lease still contributes to activeCost so utilization EWMA sees load.
this.sampleIntegral();
if (this.active.has(record.id)) {
this.active.delete(record.id);
this.activeCost -= BigInt(record.cost);
}
const latency =
meta?.latencyMs !== undefined
? meta.latencyMs
: Math.max(0, this.clock.now() - record.admittedAtMs);
noteLatency(this.adaptation, latency, this.config.adaptation);
noteOutcome(this.adaptation, outcome);
this.adaptation.windowCompleted += 1;
if (meta?.pressure) setPressure(this.adaptation, meta.pressure);
this.releaseVirtual(record);
this.dispatch();
}
private enqueue(request: AdmissionRequest, cost: number): AdmissionAcquireResult {
const id = nextId("q");
const maxWait = normalizeRequestCost(
request.maxWaitMs ?? this.config.defaultMaxWaitMs,
MAX_ADMISSION_WINDOW_MS
);
const now = this.clock.now();
const deadlineMs = Math.min(Number.MAX_SAFE_INTEGER, now + maxWait);
let settle: {
resolve: (v: AdmissionAdmitted) => void;
reject: (e: Error) => void;
};
const promise = new Promise<AdmissionAdmitted>((resolve, reject) => {
settle = { resolve, reject };
});
const entry: QueueEntry<QueuedPayload> = {
id,
tenantKey: request.tenantKey && request.tenantKey.length > 0 ? request.tenantKey : "_default",
cost,
enqueuedAtMs: now,
deadlineMs,
payload: {
resolve: (v) => settle.resolve(v),
reject: (e) => settle.reject(e),
signal: request.signal,
},
};
if (!this.queue.enqueue(entry)) {
return this.reject("ADMISSION_QUEUE_FULL", "admission queue is full");
}
entry.timerId = this.clock.setTimer(
() => {
this.expireEntry(id, "ADMISSION_DEADLINE", "admission wait deadline exceeded");
},
Math.max(0, deadlineMs - now)
);
if (request.signal) {
const onAbort = () => {
this.expireEntry(id, "ADMISSION_ABORTED", "request aborted while queued");
};
entry.payload.onAbort = onAbort;
request.signal.addEventListener("abort", onAbort, { once: true });
}
// Capacity may have freed between check and enqueue in concurrent hosts; try dispatch.
this.dispatch();
return { status: "queued", promise };
}
private expireEntry(id: string, code: AdmissionRejectCode, message: string): void {
const entry = this.queue.removeById(id);
if (!entry) return;
this.clearEntryTimer(entry);
this.detachAbort(entry);
entry.payload.reject(createAdmissionRejectError(code, message));
this.rejectedCount += 1;
// Resume enforce dispatch so a now-fitting successor is not stranded until
// unrelated activity. dispatch() is a no-op after shutdown / non-enforce.
this.dispatch();
}
private failQueued(
entry: QueueEntry<QueuedPayload>,
code: AdmissionRejectCode,
message: string
): void {
this.clearEntryTimer(entry);
this.detachAbort(entry);
entry.payload.reject(createAdmissionRejectError(code, message));
this.rejectedCount += 1;
}
private dispatch(): void {
if (this.shutDown || this.config.mode !== "enforce") return;
while (this.queue.size > 0) {
const limit = this.adaptation.currentLimit;
const available = BigInt(limit) - this.activeCost;
if (available <= 0n) return;
const entry = this.queue.dequeue(Number(available));
if (!entry) return;
this.clearEntryTimer(entry);
this.detachAbort(entry);
if (entry.payload.signal?.aborted) {
entry.payload.reject(
createAdmissionRejectError("ADMISSION_ABORTED", "request aborted while queued")
);
this.rejectedCount += 1;
continue;
}
if (this.clock.now() >= entry.deadlineMs) {
entry.payload.reject(
createAdmissionRejectError("ADMISSION_DEADLINE", "admission wait deadline exceeded")
);
this.rejectedCount += 1;
continue;
}
entry.payload.resolve(this.admit(entry.cost));
}
}
private releaseVirtual(record: ActiveLeaseRecord): void {
if (record.virtualDisposition === "active") {
this.virtualActiveCost -= record.cost;
this.virtualActiveCount -= 1;
} else if (record.virtualDisposition === "queued") {
this.virtualQueue.removeById(record.id);
}
record.virtualDisposition = "none";
this.dispatchVirtual();
}
private dispatchVirtual(): void {
while (this.virtualQueue.size > 0) {
const available = this.adaptation.currentLimit - this.virtualActiveCost;
if (available <= 0) return;
const entry = this.virtualQueue.dequeue(available);
if (!entry) return;
const record = this.active.get(entry.payload.recordId);
if (!record || record.released) continue;
record.virtualDisposition = "active";
this.virtualActiveCost = addSaturated(this.virtualActiveCost, record.cost);
this.virtualActiveCount = addSaturated(this.virtualActiveCount, 1);
}
}
private rebuildVirtualState(enable: boolean): void {
this.virtualQueue = new FairCostQueue(this.config.maxQueueCount, this.config.maxQueueCost);
this.virtualActiveCost = 0;
this.virtualActiveCount = 0;
for (const record of this.active.values()) record.virtualDisposition = "none";
if (!enable) return;
for (const record of this.active.values()) {
// Individually oversized work is virtual-rejected, never virtually queued.
if (record.cost > this.adaptation.currentLimit) {
record.virtualDisposition = "rejected";
continue;
}
if (record.cost <= this.adaptation.currentLimit - this.virtualActiveCost) {
record.virtualDisposition = "active";
this.virtualActiveCost = addSaturated(this.virtualActiveCost, record.cost);
this.virtualActiveCount = addSaturated(this.virtualActiveCount, 1);
} else if (
this.virtualQueue.enqueue({
id: record.id,
tenantKey: "_existing",
cost: record.cost,
enqueuedAtMs: record.admittedAtMs,
deadlineMs: Number.MAX_SAFE_INTEGER,
payload: { recordId: record.id },
})
) {
record.virtualDisposition = "queued";
} else {
record.virtualDisposition = "rejected";
}
}
}
private reject(code: AdmissionRejectCode, message: string): AdmissionAcquireResult {
this.rejectedCount += 1;
return { status: "rejected", code, message };
}
private clearEntryTimer(entry: QueueEntry<QueuedPayload>): void {
if (entry.timerId !== undefined) {
this.clock.clearTimer(entry.timerId);
entry.timerId = undefined;
}
}
private detachAbort(entry: QueueEntry<QueuedPayload>): void {
if (entry.payload.signal && entry.payload.onAbort) {
entry.payload.signal.removeEventListener("abort", entry.payload.onAbort);
entry.payload.onAbort = undefined;
}
}
private sampleIntegral(): void {
const now = this.clock.now();
const dt = now - this.lastSampleMs;
if (dt > 0) {
// Cap at currentLimit before Number conversion so shadow oversubscription never
// feeds an unsafe rounded activeCost into the utilization integral.
const limit = this.adaptation.currentLimit;
const activeForIntegral = this.activeCost >= BigInt(limit) ? limit : Number(this.activeCost);
sampleActiveIntegral(this.adaptation, activeForIntegral, dt);
this.lastSampleMs = now;
}
}
private armWindowTimer(): void {
if (this.windowTimer !== undefined) {
this.clock.clearTimer(this.windowTimer);
this.windowTimer = undefined;
}
if (this.shutDown || this.config.mode === "off") return;
const tick = () => {
this.tick();
if (!this.shutDown && this.config.mode !== "off") {
this.windowTimer = this.clock.setTimer(tick, this.config.windowMs);
}
};
this.windowTimer = this.clock.setTimer(tick, this.config.windowMs);
}
}

View File

@@ -0,0 +1,107 @@
import {
MAX_ADMISSION_COST_OR_LIMIT,
type AdmissionCostConfig,
type AdmissionCostFeatures,
} from "./types.ts";
export { MAX_ADMISSION_COST_OR_LIMIT };
export const DEFAULT_ADMISSION_COST_CONFIG: AdmissionCostConfig = Object.freeze({
baseCost: 1,
bodyBytesPerUnit: 16_384,
tokensPerUnit: 1_024,
messagesPerUnit: 32,
toolsPerUnit: 8,
fanoutPerUnit: 1,
streamingClassCost: 1,
nonStreamingClassCost: 2,
maxRequestCost: 1_000,
});
function finiteNonNegative(value: unknown): number {
if (typeof value !== "number" || !Number.isFinite(value) || value < 0) return 0;
return Math.min(value, Number.MAX_SAFE_INTEGER);
}
function requirePositiveSafeInteger(
name: string,
value: unknown,
max: number = MAX_ADMISSION_COST_OR_LIMIT
): number {
if (typeof value !== "number" || !Number.isSafeInteger(value) || value <= 0) {
throw new RangeError(`${name} must be a positive safe integer`);
}
if (value > max) {
throw new RangeError(`${name} must be <= ${max}`);
}
return value;
}
const COST_CONFIG_KEYS = [
"baseCost",
"bodyBytesPerUnit",
"tokensPerUnit",
"messagesPerUnit",
"toolsPerUnit",
"fanoutPerUnit",
"streamingClassCost",
"nonStreamingClassCost",
"maxRequestCost",
] as const satisfies ReadonlyArray<keyof AdmissionCostConfig>;
/** Merge cost quanta after strictly validating every supplied value. */
export function resolveCostConfig(partial?: Partial<AdmissionCostConfig>): AdmissionCostConfig {
const d = DEFAULT_ADMISSION_COST_CONFIG;
const resolved = {} as AdmissionCostConfig;
for (const key of COST_CONFIG_KEYS) {
resolved[key] = requirePositiveSafeInteger(key, partial?.[key] ?? d[key]);
}
return resolved;
}
function unitsFrom(amount: number, quantum: number): number {
return amount <= 0 ? 0 : Math.ceil(amount / quantum);
}
function addBounded(total: number, contribution: number, maximum: number): number {
if (contribution >= maximum - total) return maximum;
return total + contribution;
}
/** Pure bounded cost estimator from transparent positive safe-integer quanta. */
export function estimateAdmissionCost(
features: AdmissionCostFeatures,
config?: Partial<AdmissionCostConfig>
): number {
const cfg = resolveCostConfig(config);
const body = finiteNonNegative(features?.bodyBytes);
const tokens = finiteNonNegative(features?.estimatedInputTokens);
const messages = finiteNonNegative(features?.messageCount);
const tools = finiteNonNegative(features?.toolCount);
const fanout = Math.max(1, finiteNonNegative(features?.requestedFanout));
const contributions = [
unitsFrom(body, cfg.bodyBytesPerUnit),
unitsFrom(tokens, cfg.tokensPerUnit),
unitsFrom(messages, cfg.messagesPerUnit),
unitsFrom(tools, cfg.toolsPerUnit),
unitsFrom(fanout, cfg.fanoutPerUnit),
features?.streaming !== false ? cfg.streamingClassCost : cfg.nonStreamingClassCost,
];
let total = Math.min(cfg.baseCost, cfg.maxRequestCost);
for (const contribution of contributions) {
total = addBounded(total, contribution, cfg.maxRequestCost);
if (total === cfg.maxRequestCost) break;
}
return total;
}
/** Validate and bound a caller-supplied request cost. */
export function normalizeRequestCost(
cost: unknown,
maxRequestCost: number = DEFAULT_ADMISSION_COST_CONFIG.maxRequestCost
): number {
const max = requirePositiveSafeInteger("maxRequestCost", maxRequestCost);
const value = requirePositiveSafeInteger("request cost", cost);
return Math.min(value, max);
}

View File

@@ -0,0 +1,37 @@
/**
* Pure weighted adaptive admission-control core.
* No route, settings, or environment wiring in this module surface.
*/
export {
DEFAULT_ADMISSION_COST_CONFIG,
estimateAdmissionCost,
normalizeRequestCost,
resolveCostConfig,
} from "./cost.ts";
export { AdaptiveAdmissionController } from "./controller.ts";
export {
MAX_ADMISSION_COST_OR_LIMIT,
MAX_ADMISSION_WINDOW_MS,
createAdmissionRejectError,
type AdaptiveAdmissionConfig,
type AdmissionAcquireResult,
type AdmissionAdmitted,
type AdmissionClock,
type AdmissionCostConfig,
type AdmissionCostFeatures,
type AdmissionLease,
type AdmissionMode,
type AdmissionPressure,
type AdmissionQueued,
type AdmissionRejectCode,
type AdmissionRejectError,
type AdmissionRejected,
type AdmissionReleaseMeta,
type AdmissionReleaseOutcome,
type AdmissionRequest,
type AdmissionSnapshot,
type ShadowDecision,
} from "./types.ts";

View File

@@ -0,0 +1,194 @@
/**
* Bounded multi-tenant fair queue (round-robin across tenant buckets).
* Count + total cost caps; no unbounded arrays of timers beyond one per entry.
*/
/**
* After this many pass-overs while unfittable, reserve capacity for the aged head
* instead of indefinitely admitting smaller work from other tenants.
*/
const MAX_UNFITTABLE_SKIPS = 2;
export interface QueueEntry<T> {
id: string;
tenantKey: string;
cost: number;
enqueuedAtMs: number;
deadlineMs: number;
payload: T;
timerId?: unknown;
/** Times this head was skipped because it did not fit available cost. */
skipCount?: number;
}
export interface FairQueueSnapshot {
count: number;
cost: number;
}
export class FairCostQueue<T> {
private readonly buckets = new Map<string, QueueEntry<T>[]>();
private readonly order: string[] = [];
private cursor = 0;
private count = 0;
private cost = 0;
constructor(
readonly maxCount: number,
readonly maxCost: number
) {}
get size(): number {
return this.count;
}
get totalCost(): number {
return this.cost;
}
snapshot(): FairQueueSnapshot {
return { count: this.count, cost: this.cost };
}
canAccept(entryCost: number): boolean {
if (!Number.isSafeInteger(entryCost) || entryCost <= 0) return false;
if (this.count >= this.maxCount) return false;
if (entryCost > this.maxCost - this.cost) return false;
return true;
}
enqueue(entry: QueueEntry<T>): boolean {
if (!this.canAccept(entry.cost)) return false;
let bucket = this.buckets.get(entry.tenantKey);
if (!bucket) {
bucket = [];
this.buckets.set(entry.tenantKey, bucket);
this.order.push(entry.tenantKey);
}
bucket.push(entry);
this.count += 1;
this.cost += entry.cost;
return true;
}
/**
* Round-robin dequeue, optionally skipping tenant heads that do not fit available cost.
* After MAX_UNFITTABLE_SKIPS actual pass-overs, an unfittable head reserves capacity:
* smaller work is not admitted ahead of it until it fits, is removed, or capacity rises.
*/
dequeue(maxCost = Number.MAX_SAFE_INTEGER): QueueEntry<T> | undefined {
if (this.count === 0) return undefined;
const n = this.order.length;
// Bounded anti-starvation: prefer the oldest aged unfittable head once reserved.
let reserved: { idx: number; entry: QueueEntry<T> } | undefined;
for (let i = 0; i < n; i++) {
const idx = (this.cursor + i) % n;
const tenant = this.order[idx];
const entry = this.buckets.get(tenant)?.[0];
if (!entry) continue;
if ((entry.skipCount ?? 0) >= MAX_UNFITTABLE_SKIPS) {
if (!reserved || entry.enqueuedAtMs < reserved.entry.enqueuedAtMs) {
reserved = { idx, entry };
}
}
}
if (reserved) {
if (reserved.entry.cost > maxCost) return undefined;
return this.takeAt(reserved.idx);
}
const bypassed: QueueEntry<T>[] = [];
for (let i = 0; i < n; i++) {
const idx = (this.cursor + i) % n;
const tenant = this.order[idx];
const bucket = this.buckets.get(tenant);
const entry = bucket?.[0];
if (!entry) continue;
if (entry.cost > maxCost) {
bypassed.push(entry);
continue;
}
// Only an actual smaller admission counts as a pass-over. Merely polling
// with no available capacity must not age a head into reservation.
for (const skipped of bypassed) {
skipped.skipCount = (skipped.skipCount ?? 0) + 1;
}
return this.takeAt(idx);
}
return undefined;
}
private takeAt(idx: number): QueueEntry<T> | undefined {
const tenant = this.order[idx];
const bucket = this.buckets.get(tenant);
const entry = bucket?.[0];
if (!entry) return undefined;
bucket!.shift();
this.count -= 1;
this.cost -= entry.cost;
entry.skipCount = 0;
if (bucket!.length === 0) {
this.buckets.delete(tenant);
this.order.splice(idx, 1);
this.cursor = this.order.length === 0 ? 0 : idx % this.order.length;
} else {
this.cursor = (idx + 1) % this.order.length;
}
return entry;
}
/** Peek next without removing (for oversized-vs-limit checks). */
peek(): QueueEntry<T> | undefined {
if (this.count === 0) return undefined;
const n = this.order.length;
for (let i = 0; i < n; i++) {
const idx = (this.cursor + i) % n;
const tenant = this.order[idx];
const bucket = this.buckets.get(tenant);
if (bucket && bucket.length > 0) return bucket[0];
}
return undefined;
}
removeById(id: string): QueueEntry<T> | undefined {
for (let ti = 0; ti < this.order.length; ti++) {
const tenant = this.order[ti];
const bucket = this.buckets.get(tenant);
if (!bucket) continue;
const idx = bucket.findIndex((e) => e.id === id);
if (idx < 0) continue;
const [entry] = bucket.splice(idx, 1);
this.count -= 1;
this.cost -= entry.cost;
if (bucket.length === 0) {
this.buckets.delete(tenant);
this.order.splice(ti, 1);
if (this.order.length === 0) {
this.cursor = 0;
} else if (ti < this.cursor) {
// Removing a prior bucket shifts the successor into cursor - 1.
this.cursor -= 1;
} else if (this.cursor >= this.order.length) {
// Removed the final bucket at the cursor; wrap to the head.
this.cursor = 0;
}
// ti === cursor: leave cursor so it now points at the logical successor.
// ti > cursor: cursor is unaffected.
}
return entry;
}
return undefined;
}
drain(): QueueEntry<T>[] {
const out: QueueEntry<T>[] = [];
while (true) {
const e = this.dequeue();
if (!e) break;
out.push(e);
}
this.cursor = 0;
return out;
}
}

View File

@@ -0,0 +1,186 @@
/**
* Cheap bounded admission cost features from an already-parsed request body.
* Never re-parses, stringifies, clones, or invokes toJSON.
*/
import { estimateSizeFast } from "../../utils/estimateSize.ts";
import type { AdmissionCostFeatures } from "./types.ts";
export type AdmissionFeatureExtractionContext = {
/** When set, wins over any body/wrapped stream field. */
streaming?: boolean;
};
/**
* Max tools/functions array entries inspected.
* Uninspected tail is charged conservatively so truncation cannot undercharge cost.
*/
export const ADMISSION_TOOL_SCAN_BUDGET = 64;
type FeatureDraft = {
messageCount: number;
toolCount: number;
requestedFanout: number | null;
streaming: boolean | null;
};
function isPlainObject(value: unknown): value is Record<string, unknown> {
return value !== null && typeof value === "object" && !Array.isArray(value);
}
function asArray(value: unknown): unknown[] | null {
return Array.isArray(value) ? value : null;
}
function positiveInt(value: unknown): number | null {
if (typeof value !== "number" || !Number.isFinite(value) || value <= 0) return null;
if (!Number.isSafeInteger(value)) {
return Math.min(Number.MAX_SAFE_INTEGER, Math.floor(value));
}
return value;
}
function saturateCount(n: number): number {
if (!Number.isFinite(n) || n <= 0) return 0;
if (!Number.isSafeInteger(n)) {
return Math.min(Number.MAX_SAFE_INTEGER, Math.floor(n));
}
return n;
}
/**
* Count all recognized tool aliases/layers under one shared entry budget.
* If their combined length cannot be inspected completely, saturate before indexed access
* so an unseen alias or wrapped tail cannot undercharge heavier declarations.
*/
function countTools(layers: Array<Record<string, unknown>>): number {
const sources: unknown[][] = [];
const seen = new Set<unknown[]>();
for (const layer of layers) {
for (const value of [layer.tools, layer.functions]) {
const source = asArray(value);
if (!source || seen.has(source)) continue;
seen.add(source);
sources.push(source);
}
}
let entryCount = 0;
for (const source of sources) {
if (source.length > ADMISSION_TOOL_SCAN_BUDGET - entryCount) {
return Number.MAX_SAFE_INTEGER;
}
entryCount += source.length;
}
let total = 0;
for (const source of sources) {
for (let i = 0; i < source.length; i++) {
const entry = source[i];
if (isPlainObject(entry)) {
const declarations = asArray(entry.functionDeclarations);
if (declarations) {
total = Math.min(Number.MAX_SAFE_INTEGER, total + saturateCount(declarations.length));
continue;
}
}
total = Math.min(Number.MAX_SAFE_INTEGER, total + 1);
}
}
return total;
}
function countMessages(layer: Record<string, unknown>): number {
const messages = asArray(layer.messages);
const contents = asArray(layer.contents);
const inputArr = asArray(layer.input);
let count = Math.max(
saturateCount(messages?.length ?? 0),
saturateCount(contents?.length ?? 0),
saturateCount(inputArr?.length ?? 0)
);
// Responses API: non-empty string `input` is one input item.
if (count === 0 && typeof layer.input === "string" && layer.input.length > 0) {
count = 1;
}
return count;
}
function readFanout(layer: Record<string, unknown>): number | null {
const direct =
positiveInt(layer.n) ?? positiveInt(layer.candidateCount) ?? positiveInt(layer.candidate_count);
if (direct != null) return direct;
// Known nested Gemini/Antigravity shape only — no recursive walk.
if (isPlainObject(layer.generationConfig)) {
return (
positiveInt(layer.generationConfig.candidateCount) ??
positiveInt(layer.generationConfig.candidate_count)
);
}
return null;
}
function featureLayers(body: unknown): Array<Record<string, unknown>> {
const top = isPlainObject(body) ? body : null;
const wrapped = top && isPlainObject(top.request) ? top.request : null;
const layers: Array<Record<string, unknown>> = [];
if (top) layers.push(top);
if (wrapped) layers.push(wrapped);
return layers;
}
function absorbLayer(draft: FeatureDraft, layer: Record<string, unknown>): void {
if (draft.messageCount === 0) {
draft.messageCount = countMessages(layer);
}
if (draft.requestedFanout == null) {
draft.requestedFanout = readFanout(layer);
}
if (draft.streaming == null && "stream" in layer) {
draft.streaming = layer.stream === true;
}
}
function resolveStreaming(
draftStreaming: boolean | null,
context?: AdmissionFeatureExtractionContext
): boolean {
if (context && "streaming" in context && context.streaming !== undefined) {
return context.streaming === true;
}
return draftStreaming ?? false;
}
/**
* Inspect top-level fields and one known wrapper (`request`) only.
* Prefer the first non-empty match for each feature family.
*/
export function extractAdmissionCostFeatures(
body: unknown,
context?: AdmissionFeatureExtractionContext
): AdmissionCostFeatures {
const bodyBytes = estimateSizeFast(body);
const layers = featureLayers(body);
const draft: FeatureDraft = {
messageCount: 0,
toolCount: countTools(layers),
requestedFanout: null,
streaming: null,
};
for (const layer of layers) {
absorbLayer(draft, layer);
}
// Conservative token estimate from already-measured body size (no re-walk/stringify).
const estimatedInputTokens =
bodyBytes > 0 ? Math.min(Number.MAX_SAFE_INTEGER, Math.ceil(bodyBytes / 4)) : 0;
return {
bodyBytes,
estimatedInputTokens,
messageCount: draft.messageCount,
toolCount: draft.toolCount,
requestedFanout: draft.requestedFanout ?? 1,
streaming: resolveStreaming(draft.streaming, context),
};
}

View File

@@ -0,0 +1,614 @@
/**
* Process-local adaptive admission runtime facade around the pure controller.
* No HTTP route wiring — suitable for later shared handleChat integration.
*/
import { AdaptiveAdmissionController } from "./controller.ts";
import { validateConfig } from "./config.ts";
import { extractAdmissionCostFeatures } from "./requestFeatures.ts";
import {
type AdaptiveAdmissionConfig,
type AdmissionAcquireResult,
type AdmissionClock,
type AdmissionLease,
type AdmissionMode,
type AdmissionPressure,
type AdmissionRejectCode,
type AdmissionReleaseOutcome,
type AdmissionSnapshot,
type ShadowDecision,
} from "./types.ts";
import { buildErrorBody } from "../../utils/error.ts";
import { CORS_HEADERS } from "../../utils/cors.ts";
import {
checkResourcePressureGuard,
getResourcePressureObservation,
type ResourcePressureGuardResult,
type ResourcePressureObservation,
} from "../../utils/resourcePressure.ts";
import type { PressureReason, PressureSeverity } from "../../utils/resourcePressurePolicy.ts";
export { extractAdmissionCostFeatures } from "./requestFeatures.ts";
export const DEFAULT_ADAPTIVE_ADMISSION_CONFIG: Readonly<AdaptiveAdmissionConfig> = Object.freeze({
mode: "shadow",
minLimit: 8,
initialLimit: 64,
maxLimit: 1000,
maxQueueCount: 128,
maxQueueCost: 2000,
defaultMaxWaitMs: 5_000,
windowMs: 1_000,
});
const RUNTIME_STORE_KEY = Symbol.for("omniroute.adaptiveAdmission.runtime");
type RuntimeStore = {
runtime: AdaptiveAdmissionRuntime | null;
};
type GlobalWithRuntimeStore = typeof globalThis & {
[RUNTIME_STORE_KEY]?: RuntimeStore;
};
function getRuntimeStore(): RuntimeStore {
const globalWithStore = globalThis as GlobalWithRuntimeStore;
let store = globalWithStore[RUNTIME_STORE_KEY];
if (!store) {
store = { runtime: null };
globalWithStore[RUNTIME_STORE_KEY] = store;
}
return store;
}
const ENV_KEYS = {
mode: "ADAPTIVE_ADMISSION_MODE",
minLimit: "ADAPTIVE_ADMISSION_MIN_LIMIT",
initialLimit: "ADAPTIVE_ADMISSION_INITIAL_LIMIT",
maxLimit: "ADAPTIVE_ADMISSION_MAX_LIMIT",
maxQueueCount: "ADAPTIVE_ADMISSION_MAX_QUEUE_COUNT",
maxQueueCost: "ADAPTIVE_ADMISSION_MAX_QUEUE_COST",
defaultMaxWaitMs: "ADAPTIVE_ADMISSION_MAX_WAIT_MS",
windowMs: "ADAPTIVE_ADMISSION_WINDOW_MS",
} as const;
function parsePositiveSafeInt(name: string, raw: string): number {
if (!/^[0-9]+$/.test(raw)) {
throw new RangeError(`${name} must be a positive safe integer`);
}
const value = Number(raw);
if (!Number.isSafeInteger(value) || value <= 0) {
throw new RangeError(`${name} must be a positive safe integer`);
}
return value;
}
/** Strict env → config resolver. Throws clear config errors for direct callers. */
export function resolveAdaptiveAdmissionConfigFromEnv(
env: NodeJS.ProcessEnv | Record<string, string | undefined> = process.env
): AdaptiveAdmissionConfig {
const cfg: AdaptiveAdmissionConfig = { ...DEFAULT_ADAPTIVE_ADMISSION_CONFIG };
const modeRaw = env[ENV_KEYS.mode];
if (modeRaw !== undefined && modeRaw !== "") {
if (modeRaw !== "off" && modeRaw !== "shadow" && modeRaw !== "enforce") {
throw new RangeError(`${ENV_KEYS.mode} must be off|shadow|enforce`);
}
cfg.mode = modeRaw;
}
// Numeric env keys only — typed assignment without index-signature cast (TS2352).
type EnvIntField = Exclude<keyof typeof ENV_KEYS, "mode">;
const intFields = [
"minLimit",
"initialLimit",
"maxLimit",
"maxQueueCount",
"maxQueueCost",
"defaultMaxWaitMs",
"windowMs",
] as const satisfies ReadonlyArray<EnvIntField>;
for (const field of intFields) {
const envName = ENV_KEYS[field];
const raw = env[envName];
if (raw === undefined || raw === "") continue;
cfg[field] = parsePositiveSafeInt(envName, raw);
}
// Shared pure validation — accept exact documented maxima, reject core-invalid configs.
validateConfig(cfg);
return cfg;
}
export type AdaptiveAdmissionAcquireInput = {
/** Opaque fairness key; never exposed in snapshots or client errors. */
tenantKey: string;
/** Already-parsed request body — must not be re-read or stringified for cost. */
body: unknown;
signal?: AbortSignal;
maxWaitMs?: number;
/** Authoritative streaming class; wins body stream inference when set. */
streaming?: boolean;
};
export type AdaptiveAdmissionAdmitted = {
status: "admitted";
mode: AdmissionMode;
lease: AdmissionLease;
admittedAtMs: number;
shadowDecision?: ShadowDecision;
};
export type AdaptiveAdmissionRejected = {
status: "rejected";
code: string;
response: Response;
};
export type AdaptiveAdmissionAcquireResult = AdaptiveAdmissionAdmitted | AdaptiveAdmissionRejected;
export type AdaptiveAdmissionPublicSnapshot = AdmissionSnapshot & {
resourceSeverity: PressureSeverity;
resourceReason: PressureReason;
resourceObservedAtMs: number;
pressureGuardRejectCount: number;
};
export type AdaptiveAdmissionLifecycleOptions = {
admittedAtMs: number;
signal?: AbortSignal;
nowMs?: () => number;
};
export type AdaptiveAdmissionRuntimeOptions = {
config?: AdaptiveAdmissionConfig;
env?: NodeJS.ProcessEnv | Record<string, string | undefined>;
clock?: Partial<AdmissionClock>;
checkResourcePressure?: () => ResourcePressureGuardResult | null;
getResourcePressureObservation?: () => ResourcePressureObservation;
/** Test seam: observe pressure values fed into the controller after dedupe. */
onPressureObserved?: (pressure: AdmissionPressure) => void;
warn?: (message: string) => void;
nowMs?: () => number;
};
/** Non-success release outcomes callers must choose explicitly for handler failures. */
export type AdaptiveAdmissionFailureOutcome = Exclude<AdmissionReleaseOutcome, "success">;
export type AdaptiveAdmissionRuntime = {
acquire(input: AdaptiveAdmissionAcquireInput): Promise<AdaptiveAdmissionAcquireResult>;
snapshot(): AdaptiveAdmissionPublicSnapshot;
dispose(): void;
/**
* Release an admitted lease after a handler failure before any HTTP response exists.
* Callers must supply the concrete non-success outcome — never defaults to local_reject.
*/
releaseHandlerFailure(
lease: AdmissionLease,
outcome: AdaptiveAdmissionFailureOutcome,
options?: { admittedAtMs?: number; nowMs?: () => number }
): void;
attachResponseLifecycle(
response: Response,
lease: AdmissionLease,
options: AdaptiveAdmissionLifecycleOptions
): Response;
};
type RejectHttpMapping = {
status: number;
code: string;
message: string;
retryAfter?: string;
};
const REJECT_MAP: Record<AdmissionRejectCode, RejectHttpMapping> = {
ADMISSION_ABORTED: {
status: 499,
code: "admission_aborted",
message: "Request aborted",
},
ADMISSION_OVERSIZED: {
status: 503,
code: "admission_oversized",
message: "Request too large for current capacity",
},
ADMISSION_QUEUE_FULL: {
status: 503,
code: "admission_queue_full",
message: "Service temporarily unavailable",
retryAfter: "1",
},
ADMISSION_DEADLINE: {
status: 503,
code: "admission_deadline",
message: "Service temporarily unavailable",
retryAfter: "1",
},
ADMISSION_SHUTDOWN: {
status: 503,
code: "admission_shutdown",
message: "Service temporarily unavailable",
},
ADMISSION_UNAVAILABLE: {
status: 503,
code: "admission_unavailable",
message: "Service temporarily unavailable",
retryAfter: "1",
},
};
function isAdmissionRejectError(
err: unknown
): err is { code: AdmissionRejectCode; name: string; message: string } {
return (
!!err &&
typeof err === "object" &&
(err as { name?: string }).name === "AdmissionRejectError" &&
typeof (err as { code?: unknown }).code === "string"
);
}
function buildAdmissionRejectResponse(code: AdmissionRejectCode): AdaptiveAdmissionRejected {
const mapping = REJECT_MAP[code] ?? REJECT_MAP.ADMISSION_UNAVAILABLE;
const headers: Record<string, string> = {
"Content-Type": "application/json",
...CORS_HEADERS,
};
if (mapping.retryAfter) headers["Retry-After"] = mapping.retryAfter;
const body = buildErrorBody(mapping.status, mapping.message, undefined, {
type: mapping.status === 499 ? "client_disconnected" : "server_error",
code: mapping.code,
});
return {
status: "rejected",
code: mapping.code,
response: new Response(JSON.stringify(body), {
status: mapping.status,
headers,
}),
};
}
function observationIdentity(state: ResourcePressureObservation["state"]): string {
return `${state.observedAtMs}|${state.severity}|${state.reason}`;
}
function toAdmissionPressure(severity: PressureSeverity): AdmissionPressure {
if (severity === "critical") return "critical";
if (severity === "high") return "high";
return "normal";
}
function isSseResponse(response: Response): boolean {
const contentType = response.headers.get("content-type") ?? "";
return contentType.toLowerCase().includes("text/event-stream");
}
function releaseOnce(
lease: AdmissionLease,
outcome: AdmissionReleaseOutcome,
admittedAtMs: number | undefined,
nowMs: () => number
): void {
if (lease.released) return;
const latencyMs = admittedAtMs === undefined ? undefined : Math.max(0, nowMs() - admittedAtMs);
lease.release(outcome, latencyMs === undefined ? undefined : { latencyMs });
}
/**
* Map HTTP status (+ optional request signal) to admission release outcome.
* Cancellation always wins over status classification.
*/
function classifyHttpOutcome(status: number, signal?: AbortSignal): AdmissionReleaseOutcome {
if (signal?.aborted || status === 499) return "cancelled";
if (status === 408 || status === 504) return "timeout";
if (status >= 500) return "upstream_error";
if (status >= 400) return "local_reject";
// 2xx / 3xx (and rare 1xx) complete successfully from admission's perspective.
return "success";
}
class AdaptiveAdmissionRuntimeImpl implements AdaptiveAdmissionRuntime {
private readonly controller: AdaptiveAdmissionController;
private readonly checkResourcePressure: () => ResourcePressureGuardResult | null;
private readonly getResourcePressureObservation: () => ResourcePressureObservation;
private readonly onPressureObserved?: (pressure: AdmissionPressure) => void;
private readonly nowMs: () => number;
private lastObservationKey: string | null = null;
private lastResource: {
severity: PressureSeverity;
reason: PressureReason;
observedAtMs: number;
} = { severity: "normal", reason: "none", observedAtMs: 0 };
private pressureGuardRejectCount = 0;
private disposed = false;
constructor(options: AdaptiveAdmissionRuntimeOptions, config: AdaptiveAdmissionConfig) {
this.controller = new AdaptiveAdmissionController(config, options.clock);
this.checkResourcePressure = options.checkResourcePressure ?? checkResourcePressureGuard;
this.getResourcePressureObservation =
options.getResourcePressureObservation ?? getResourcePressureObservation;
this.onPressureObserved = options.onPressureObserved;
this.nowMs = options.nowMs ?? options.clock?.now ?? (() => Date.now());
}
async acquire(input: AdaptiveAdmissionAcquireInput): Promise<AdaptiveAdmissionAcquireResult> {
if (this.disposed) {
return buildAdmissionRejectResponse("ADMISSION_SHUTDOWN");
}
// Independent safety fuse first — never acquire provider work on critical guard.
// Still feed pressure observations so the controller learns from critical samples.
let guard: ResourcePressureGuardResult | null = null;
try {
guard = this.checkResourcePressure();
} catch {
// Fail open on sampling/check failures.
}
this.feedFreshPressureObservation();
if (guard) {
this.pressureGuardRejectCount += 1;
return {
status: "rejected",
code: "resource_pressure",
response: guard.response,
};
}
const features = extractAdmissionCostFeatures(
input.body,
input.streaming === undefined ? undefined : { streaming: input.streaming }
);
let result: AdmissionAcquireResult;
try {
result = await this.controller.acquire({
tenantKey: input.tenantKey,
features,
signal: input.signal,
maxWaitMs: input.maxWaitMs,
});
} catch (err) {
if (isAdmissionRejectError(err)) {
return buildAdmissionRejectResponse(err.code);
}
return buildAdmissionRejectResponse("ADMISSION_UNAVAILABLE");
}
if (result.status === "rejected") {
return buildAdmissionRejectResponse(result.code);
}
if (result.status === "queued") {
try {
const admitted = await result.promise;
return {
status: "admitted",
mode: this.controller.snapshot().mode,
lease: admitted.lease,
admittedAtMs: this.nowMs(),
shadowDecision: admitted.shadowDecision,
};
} catch (err) {
if (isAdmissionRejectError(err)) {
return buildAdmissionRejectResponse(err.code);
}
return buildAdmissionRejectResponse("ADMISSION_UNAVAILABLE");
}
}
return {
status: "admitted",
mode: this.controller.snapshot().mode,
lease: result.lease,
admittedAtMs: this.nowMs(),
shadowDecision: result.shadowDecision,
};
}
snapshot(): AdaptiveAdmissionPublicSnapshot {
const core = this.controller.snapshot();
return {
...core,
resourceSeverity: this.lastResource.severity,
resourceReason: this.lastResource.reason,
resourceObservedAtMs: this.lastResource.observedAtMs,
pressureGuardRejectCount: this.pressureGuardRejectCount,
};
}
dispose(): void {
if (this.disposed) return;
this.disposed = true;
this.controller.shutdown();
}
releaseHandlerFailure(
lease: AdmissionLease,
outcome: AdaptiveAdmissionFailureOutcome,
options?: { admittedAtMs?: number; nowMs?: () => number }
): void {
releaseOnce(lease, outcome, options?.admittedAtMs, options?.nowMs ?? this.nowMs);
}
attachResponseLifecycle(
response: Response,
lease: AdmissionLease,
options: AdaptiveAdmissionLifecycleOptions
): Response {
const nowMs = options.nowMs ?? this.nowMs;
const admittedAtMs = options.admittedAtMs;
if (!response.body || !isSseResponse(response)) {
releaseOnce(lease, classifyHttpOutcome(response.status, options.signal), admittedAtMs, nowMs);
return response;
}
const upstream = response.body;
const reader = upstream.getReader();
let settled = false;
let readerCancelled = false;
const settle = (outcome: AdmissionReleaseOutcome): void => {
if (settled) return;
settled = true;
releaseOnce(lease, outcome, admittedAtMs, nowMs);
};
const cancelReader = (reason?: unknown): void => {
if (readerCancelled) return;
readerCancelled = true;
void reader.cancel(reason).catch(() => {
/* ignore cancel races */
});
};
const onAbort = (): void => {
cancelReader(options.signal?.reason);
settle("cancelled");
};
if (options.signal) {
if (options.signal.aborted) {
onAbort();
} else {
options.signal.addEventListener("abort", onAbort, { once: true });
}
}
const detachAbort = (): void => {
options.signal?.removeEventListener("abort", onAbort);
};
const stream = new ReadableStream<Uint8Array>({
async pull(controller) {
if (settled) {
controller.close();
return;
}
try {
const { done, value } = await reader.read();
if (done) {
detachAbort();
settle(classifyHttpOutcome(response.status, options.signal));
controller.close();
return;
}
controller.enqueue(value);
} catch (err) {
detachAbort();
settle(options.signal?.aborted ? "cancelled" : "upstream_error");
controller.error(err);
}
},
cancel(reason) {
detachAbort();
cancelReader(reason);
settle("cancelled");
},
});
return new Response(stream, {
status: response.status,
statusText: response.statusText,
headers: response.headers,
});
}
private feedFreshPressureObservation(): void {
try {
const observation = this.getResourcePressureObservation();
const state = observation.state;
this.lastResource = {
severity: state.severity,
reason: state.reason,
observedAtMs: state.observedAtMs,
};
const key = observationIdentity(state);
if (state.observedAtMs <= 0) return;
if (key === this.lastObservationKey) return;
this.lastObservationKey = key;
const pressure = toAdmissionPressure(state.severity);
this.controller.observePressure(pressure);
this.onPressureObserved?.(pressure);
} catch {
// Fail open.
}
}
}
function createRuntimeFromResolvedConfig(
options: AdaptiveAdmissionRuntimeOptions,
config: AdaptiveAdmissionConfig
): AdaptiveAdmissionRuntime {
return new AdaptiveAdmissionRuntimeImpl(options, config);
}
/**
* Create an injected adaptive-admission runtime for tests or process use.
* Invalid explicit `config` still throws (direct callers want fail-fast).
*/
export function createAdaptiveAdmissionRuntime(
options: AdaptiveAdmissionRuntimeOptions = {}
): AdaptiveAdmissionRuntime {
const config =
options.config ??
(options.env
? resolveAdaptiveAdmissionConfigFromEnv(options.env)
: { ...DEFAULT_ADAPTIVE_ADMISSION_CONFIG });
return createRuntimeFromResolvedConfig(options, config);
}
function warnInvalidDefaultConfig(warn: ((message: string) => void) | undefined): void {
const message =
"[adaptiveAdmission] invalid environment configuration; using default shadow admission settings";
if (warn) {
warn(message);
return;
}
console.warn(message);
}
function createDefaultProcessRuntime(
options: AdaptiveAdmissionRuntimeOptions = {}
): AdaptiveAdmissionRuntime {
const warn = options.warn;
try {
const config =
options.config ?? resolveAdaptiveAdmissionConfigFromEnv(options.env ?? process.env);
return createRuntimeFromResolvedConfig(options, config);
} catch {
warnInvalidDefaultConfig(warn);
return createRuntimeFromResolvedConfig(options, {
...DEFAULT_ADAPTIVE_ADMISSION_CONFIG,
});
}
}
/** Call-time process-global runtime (HMR-safe via globalThis symbol store). */
export function getAdaptiveAdmissionRuntime(): AdaptiveAdmissionRuntime {
const store = getRuntimeStore();
if (!store.runtime) {
store.runtime = createDefaultProcessRuntime();
}
return store.runtime;
}
/** Dispose previous controller and replace the process-global runtime. */
export function reloadAdaptiveAdmissionRuntime(
options: AdaptiveAdmissionRuntimeOptions = {}
): AdaptiveAdmissionRuntime {
const store = getRuntimeStore();
store.runtime?.dispose();
store.runtime = createDefaultProcessRuntime(options);
return store.runtime;
}
/** Test isolation: dispose and clear the process-global runtime slot. */
export function resetAdaptiveAdmissionRuntimeForTests(): void {
const store = getRuntimeStore();
store.runtime?.dispose();
store.runtime = null;
}

View File

@@ -0,0 +1,171 @@
/**
* Pure weighted adaptive admission-control types.
* No route/settings wiring — dependency-injected controller seam only.
*/
/**
* Upper bound for adaptation windows and wait deadlines that participate in
* cost×time products (utilization integrals, deadline offsets).
* 24h is far beyond practical control windows while keeping the product domain exact.
*/
export const MAX_ADMISSION_WINDOW_MS = 86_400_000;
/**
* Upper bound for every validated cost, limit, and queue-cost quantum.
* Derived so `MAX_ADMISSION_COST_OR_LIMIT * MAX_ADMISSION_WINDOW_MS` remains a
* safe integer: a full window at the maximum limit integrates to utilization 1.0
* without saturating or rounding Number arithmetic.
*/
export const MAX_ADMISSION_COST_OR_LIMIT = Math.floor(
Number.MAX_SAFE_INTEGER / MAX_ADMISSION_WINDOW_MS
);
export type AdmissionMode = "off" | "shadow" | "enforce";
export type AdmissionPressure = "normal" | "high" | "critical";
/** Local outcome categories. Upstream business errors must not collapse capacity. */
export type AdmissionReleaseOutcome =
"success" | "upstream_error" | "timeout" | "local_reject" | "cancelled";
export type AdmissionRejectCode =
| "ADMISSION_OVERSIZED"
| "ADMISSION_QUEUE_FULL"
| "ADMISSION_DEADLINE"
| "ADMISSION_ABORTED"
| "ADMISSION_SHUTDOWN"
| "ADMISSION_UNAVAILABLE";
export type ShadowDecision = "would-admit" | "would-queue" | "would-reject";
export interface AdmissionCostFeatures {
bodyBytes?: number | null;
estimatedInputTokens?: number | null;
messageCount?: number | null;
toolCount?: number | null;
requestedFanout?: number | null;
streaming?: boolean | null;
}
export interface AdmissionCostConfig {
baseCost: number;
bodyBytesPerUnit: number;
tokensPerUnit: number;
messagesPerUnit: number;
toolsPerUnit: number;
fanoutPerUnit: number;
streamingClassCost: number;
nonStreamingClassCost: number;
maxRequestCost: number;
}
export interface AdaptiveAdmissionConfig {
mode?: AdmissionMode;
minLimit: number;
maxLimit: number;
initialLimit: number;
maxQueueCount: number;
maxQueueCost: number;
defaultMaxWaitMs?: number;
windowMs?: number;
shortLatencyAlpha?: number;
longLatencyAlpha?: number;
increaseStep?: number;
decreaseFactor?: number;
criticalDecreaseFactor?: number;
highUtilizationThreshold?: number;
lowUtilizationThreshold?: number;
latencyGradientThreshold?: number;
maxIncreasePerWindow?: number;
/** Optional cost quanta override used only when callers pass features instead of cost. */
cost?: Partial<AdmissionCostConfig>;
}
export interface AdmissionRequest {
/** Positive integer cost units. If omitted, `features` + cost config are used. */
cost?: number;
features?: AdmissionCostFeatures;
/** Opaque fairness key; never exposed in snapshots. */
tenantKey?: string;
maxWaitMs?: number;
signal?: AbortSignal;
pressure?: AdmissionPressure;
}
export interface AdmissionReleaseMeta {
latencyMs?: number;
pressure?: AdmissionPressure;
}
export interface AdmissionLease {
readonly id: string;
readonly cost: number;
readonly released: boolean;
release(outcome?: AdmissionReleaseOutcome, meta?: AdmissionReleaseMeta): void;
}
export interface AdmissionAdmitted {
status: "admitted";
lease: AdmissionLease;
shadowDecision?: ShadowDecision;
}
export interface AdmissionQueued {
status: "queued";
promise: Promise<AdmissionAdmitted>;
}
export interface AdmissionRejected {
status: "rejected";
code: AdmissionRejectCode;
message: string;
shadowDecision?: ShadowDecision;
}
export type AdmissionAcquireResult = AdmissionAdmitted | AdmissionQueued | AdmissionRejected;
export interface AdmissionSnapshot {
mode: AdmissionMode;
currentLimit: number;
minLimit: number;
maxLimit: number;
activeCost: number;
activeCount: number;
queuedCost: number;
queuedCount: number;
virtualActiveCost: number;
virtualActiveCount: number;
virtualQueuedCost: number;
virtualQueuedCount: number;
admittedCount: number;
rejectedCount: number;
wouldAdmitCount: number;
wouldQueueCount: number;
wouldRejectCount: number;
shortLatencyEwma: number;
longLatencyEwma: number;
utilization: number;
pressure: AdmissionPressure;
shutdown: boolean;
}
export interface AdmissionClock {
now: () => number;
setTimer: (fn: () => void, delayMs: number) => unknown;
clearTimer: (id: unknown) => void;
}
export interface AdmissionRejectError extends Error {
code: AdmissionRejectCode;
name: "AdmissionRejectError";
}
export function createAdmissionRejectError(
code: AdmissionRejectCode,
message: string
): AdmissionRejectError {
const err = new Error(message) as AdmissionRejectError;
err.name = "AdmissionRejectError";
err.code = code;
return err;
}

View File

@@ -153,6 +153,7 @@ import {
resolveDelayMs,
comboModelNotFoundResponse,
isStreamReadinessFailureErrorBody,
isStreamEarlyEofErrorBody,
isTokenLimitBreachErrorBody,
toRecordedTarget,
getExhaustedTargetSkipReason,
@@ -1511,6 +1512,11 @@ export async function handleComboChat({
const isStreamReadinessFailure =
(result.status === 502 || result.status === 504) &&
isStreamReadinessFailureErrorBody(errorBody);
// An early EOF is an upstream failure, not a readiness probe — the breaker must
// see it even though the transient-retry path below treats both codes alike.
const isStreamEarlyEof =
(result.status === 502 || result.status === 504) &&
isStreamEarlyEofErrorBody(errorBody);
// FIX 5: a local per-API-key token-limit 429 must not cool shared accounts.
const isTokenLimitBreach =
@@ -1713,6 +1719,7 @@ export async function handleComboChat({
if (
shouldRecordProviderBreakerFailure({
isStreamReadinessFailure,
isStreamEarlyEof,
status: result.status,
sameProviderNext,
skipProviderBreaker: fallbackResult.skipProviderBreaker,

View File

@@ -133,7 +133,11 @@ const PROVIDER_BREAKER_FAILURE_STATUSES = new Set([408, 500, 502, 503, 504]);
* failure (#1731 / #2743 gap-d). This is the consumer side of `skipProviderBreaker`:
*
* - Stream-readiness failures (pre-flight zombie/ping probes) never count as provider
* failures — they are a connection-readiness signal, not an upstream outage.
* failures — they are a connection-readiness signal, not an upstream outage. EXCEPT a
* STREAM_EARLY_EOF (`isStreamEarlyEof`): there the upstream returned HTTP 200, opened the
* SSE stream and then hung up without a single non-ping event, which is a genuine upstream
* failure. Excluding it made a provider-wide outage invisible to the breaker — see the
* STREAM_EARLY_EOF section of RESILIENCE_GUIDE.md.
* - Only whole-provider failure statuses (408/500/502/503/504) count. A plain rate-limit
* 429 is deliberately EXCLUDED — it belongs to connection cooldown / model lockout scope
* (a genuine quota/token-limit 429 is handled there), NOT the whole-provider breaker. This
@@ -163,6 +167,10 @@ const PROVIDER_BREAKER_FAILURE_STATUSES = new Set([408, 500, 502, 503, 504]);
*/
export function shouldRecordProviderBreakerFailure(args: {
isStreamReadinessFailure: boolean;
/** True when the failure is specifically a STREAM_EARLY_EOF (upstream hung up after
* HTTP 200). Overrides the `isStreamReadinessFailure` exemption only; every other
* AND-term below still gates the trip. */
isStreamEarlyEof?: boolean;
status: number;
sameProviderNext: boolean;
skipProviderBreaker?: boolean;
@@ -173,7 +181,7 @@ export function shouldRecordProviderBreakerFailure(args: {
isProxyUnreachable?: boolean;
}): boolean {
return (
!args.isStreamReadinessFailure &&
(!args.isStreamReadinessFailure || args.isStreamEarlyEof === true) &&
PROVIDER_BREAKER_FAILURE_STATUSES.has(args.status) &&
(!args.sameProviderNext || args.isProxyUnreachable === true) &&
!args.skipProviderBreaker &&
@@ -186,6 +194,8 @@ const REQUEST_SCOPED_UPSTREAM_ERROR_CODES = new Set([
"context_length_exceeded",
"upstream_empty_response",
"upstream_response_failed",
// Local combo per-target timer (targetTimeoutRunner) — not a connection health signal.
"combo_target_timeout",
]);
/** Request/model-specific failures must not poison provider-wide resilience state. */
@@ -308,6 +318,28 @@ export function isStreamReadinessFailureErrorBody(errorBody: unknown): boolean {
return code === "STREAM_READINESS_TIMEOUT" || code === "STREAM_EARLY_EOF";
}
/**
* A STREAM_EARLY_EOF specifically: the upstream accepted the request (HTTP 200), opened the
* SSE stream, then closed it before emitting a single non-ping event.
*
* This is deliberately NOT the same signal as STREAM_READINESS_TIMEOUT. The readiness probe
* is a pre-flight liveness check on a connection we have not committed to yet, so failing it
* says "this connection looks stale", not "this provider is failing". An early EOF is the
* opposite: the provider took the request and then failed to serve it, which is an upstream
* failure by any reasonable definition.
*
* `isStreamReadinessFailureErrorBody` still covers both codes because the transient-retry and
* semaphore-cooldown paths in combo.ts want identical treatment for both. Only the
* whole-provider circuit breaker needs to tell them apart — see
* `shouldRecordProviderBreakerFailure`.
*/
export function isStreamEarlyEofErrorBody(errorBody: unknown): boolean {
if (!errorBody || typeof errorBody !== "object") return false;
const error = (errorBody as Record<string, unknown>).error;
if (!error || typeof error !== "object") return false;
return (error as Record<string, unknown>).code === "STREAM_EARLY_EOF";
}
/**
* A local per-API-key token-limit breach surfaces as a 429 tagged with
* errorCode "TOKEN_LIMIT_EXCEEDED" (see chatCore.ts Tier 2 early return). This

View File

@@ -1,17 +1,20 @@
/**
* Wrap a single-model dispatch with a per-target timeout that aborts and falls back.
*
* Verbatim extraction of handleComboChat's `handleSingleModelWithTimeout` closure
* (combo.ts). Behavior is byte-identical; the only change is that the closed-over locals
* (`handleSingleModel`, `comboTargetTimeoutMs`, `log`) became explicit factory params.
* Extracted from handleComboChat's `handleSingleModelWithTimeout` closure (combo.ts).
* A locally expired timer aborts that target and returns a typed 504 response so the Combo
* can fall back without treating OmniRoute's own deadline as a provider-connection failure.
* The per-model abort signal still comes from the target (`target.modelAbortSignal`), so
* the outer request signal is intentionally NOT a dependency here.
*
* See _tasks/superpowers/plans/2026-07-03-blocoJ-combo-hotpath-decomposition.md (Task 1).
*/
import { errorResponse } from "../../utils/error.ts";
import { buildErrorBody, errorResponse, sanitizeErrorMessage } from "../../utils/error.ts";
import type { HandleSingleModel, SingleModelTarget, ComboLogger } from "./types.ts";
/** Stable internal classification for OmniRoute's own combo per-target timer. */
export const COMBO_TARGET_TIMEOUT_CODE = "combo_target_timeout";
export function buildTargetTimeoutRunner(deps: {
handleSingleModel: HandleSingleModel;
comboTargetTimeoutMs: number;
@@ -44,11 +47,23 @@ export function buildTargetTimeoutRunner(deps: {
`Model ${modelStr} exceeded ${comboTargetTimeoutMs}ms timeout — falling back`
);
timeoutController.abort(new Error("combo-per-model-timeout"));
// HTTP 504 (not proprietary 524): this is OmniRoute's own per-target timer.
// Typed as combo_target_timeout so request-scoped classification can keep the
// connection eligible for fallback instead of treating it like Cloudflare 524
// or a genuine upstream gateway timeout.
resolve(
new Response(JSON.stringify({ error: { message: `Model ${modelStr} timed out` } }), {
status: 524,
headers: { "Content-Type": "application/json" },
})
new Response(
JSON.stringify(
buildErrorBody(504, sanitizeErrorMessage(`Model ${modelStr} timed out`), undefined, {
type: COMBO_TARGET_TIMEOUT_CODE,
code: COMBO_TARGET_TIMEOUT_CODE,
})
),
{
status: 504,
headers: { "Content-Type": "application/json" },
}
)
);
}, comboTargetTimeoutMs);
});
@@ -72,7 +87,7 @@ export function buildTargetTimeoutRunner(deps: {
return await Promise.race([
handleSingleModel(b, modelStr, targetWithSignal).catch((err) => {
if (timedOut) {
// Inner call rejected because we aborted it. The synthetic 524 from
// Inner call rejected because we aborted it. The synthetic 504 from
// timeoutPromise already wins the race; return an empty response so
// the loser branch resolves cleanly without leaking err.message.
return new Response(null, { status: 599 });

View File

@@ -61,8 +61,8 @@ export function isComboCooldownWaitEligible(
* When the combo is wait-eligible (see isComboCooldownWaitEligible), a single target's
* dispatch can legitimately wait out cooldowns for up to `comboCooldownWait.budgetMs`
* before it resolves — so the per-target timeout must never be shorter than that budget,
* or the wait gets cut off mid-retry and the target times out with a synthetic 524
* (open-sse/services/combo/targetTimeoutRunner.ts) instead of completing the wait. This
* or the wait gets cut off mid-retry and the target times out with a synthetic 504
* (`combo_target_timeout`, open-sse/services/combo/targetTimeoutRunner.ts) instead of completing the wait. This
* only raises the *default* floor; an operator's explicit `targetTimeoutMs` on the combo
* still wins (see resolveComboTargetTimeoutMs).
*/
@@ -99,7 +99,7 @@ const DEFAULT_COMBO_CONFIG = {
retryDelayMs: 2000,
fallbackDelayMs: 0,
concurrencyPerModel: 3, // max simultaneous requests per model (round-robin)
queueTimeoutMs: 30000, // max wait time in semaphore queue (round-robin)
queueTimeoutMs: 120000, // max wait time in semaphore queue (round-robin); raised from 30s for browser-automation providers like gemini-web (#9407)
queueDepth: DEFAULT_COMBO_QUEUE_DEPTH, // pre-cascade semaphore queue depth (round-robin, #3872)
handoffThreshold: 0.85,
handoffModel: "",

View File

@@ -22,15 +22,16 @@ let _config = {
// lazily loaded on first access. better-sqlite3 is synchronous, so both the load
// and the save stay in the sync hot path without extra startup wiring. tempBans
// are intentionally NOT persisted — they are ephemeral, TTL-swept runtime state.
//
// D2 (#9033): the _loaded one-shot gate was removed so a config persisted by the
// dashboard settings route (a separate module instance, since @omniroute/open-sse
// is bundled per-entry via transpilePackages) propagates to the proxy runtime
// without a restart. A DB failure still degrades to the in-memory defaults, and
// tempBans remain in-memory-only as before.
const IP_FILTER_NAMESPACE = "ipFilter";
const IP_FILTER_KEY = "config";
let _loaded = false;
function ensureLoaded() {
if (_loaded) return;
// Mark loaded up-front so a DB failure (build phase / cloud / migration not yet
// run) degrades to in-memory only instead of retrying on every request.
_loaded = true;
try {
const row = getDbInstance()
.prepare("SELECT value FROM key_value WHERE namespace = ? AND key = ?")
@@ -235,9 +236,17 @@ export function createIPFilterMiddleware() {
/**
* For Next.js App Router — check IP from request object
*
* D1 (#9033): accepts an optional trustedPeerIp (resolved from the authenticated
* peer stamp, available on direct connections where the proxy runtime has no
* socket). When provided, it is checked FIRST before falling through to the
* forwarding headers, so a blacklisted IP on a direct connection (no XFF, no
* socket) is blocked. When behind a reverse proxy (via-proxy marker set), the
* caller passes null so the XFF path continues to work.
*/
export function checkRequestIP(request) {
export function checkRequestIP(request, trustedPeerIp) {
const ip =
pickFirstValidIp(trustedPeerIp || null) ||
pickFirstValidIp(request.headers?.get?.("cf-connecting-ip")) ||
pickFirstValidIp(request.headers?.get?.("x-forwarded-for")) ||
pickFirstValidIp(request.headers?.get?.("x-real-ip")) ||
@@ -329,7 +338,6 @@ function extractClientIP(req) {
* Reset config (for testing)
*/
export function resetIPFilter() {
_loaded = false;
_config = {
enabled: false,
mode: "blacklist",

View File

@@ -557,20 +557,36 @@ function parseAliasTarget(target: string): ResolvedModelTarget | null {
}
async function resolveModelByProviderInference(modelId: string, extendedContext: boolean) {
if (CODEX_NATIVE_UNPREFIXED_MODELS.has(modelId)) {
return {
provider: "codex",
model: modelId,
extendedContext,
};
}
const [activeProviders, activeSyncedProviders, preferClaudeCodeForUnprefixedClaudeModels] =
await Promise.all([
getActiveProviderSet(),
getActiveSyncedProvidersForModel(modelId),
getPreferClaudeCodeForUnprefixedClaudeModels(),
]);
// Codex-native bare ids prefer the ChatGPT subscription, but the preference is only
// allowed to PREEMPT another provider when a codex connection is actually active.
// Returning "codex" unconditionally (as this did once the set grew past
// `codex-auto-review` to cover gpt-5.5 / the gpt-5.6-sol tiers) hands ids that OpenAI
// also serves to a provider the operator may not have configured: an OpenAI-only
// install fails with "no active credentials for provider: codex" on a model that
// works, and an install whose codex connection is merely *inactive* fails the same way.
// Ids only codex catalogs (e.g. `codex-auto-review`) keep resolving to codex with no
// connection at all — there is no alternative to preempt, and "no codex credentials"
// is the honest error. With codex active the preference still beats OpenAI, and an
// explicit `openai/…` prefix remains the per-request override either way.
if (CODEX_NATIVE_UNPREFIXED_MODELS.has(modelId)) {
const codexNativeAlternatives = (MODEL_TO_PROVIDERS.get(modelId) || []).filter(
(p) => p !== "codex"
);
if (codexNativeAlternatives.length === 0 || activeProviders?.has("codex")) {
return {
provider: "codex",
model: modelId,
extendedContext,
};
}
}
// #FIX: synced catalogs (populated from `/v1/models` per connection) can
// claim ownership of models the provider does not actually serve (e.g. a
// `kiro` upstream briefly advertising `claude-opus-5` before it was

View File

@@ -66,6 +66,7 @@ import { getVertexUsage } from "./usage/vertex.ts";
import { getXiaomiMimoUsage } from "./usage/xiaomi-mimo.ts";
import { getXaiUsage } from "./usage/xai.ts";
import { getXaiOauthUsage } from "./usage/xaiOauth.ts";
import { getGrokCliUsage } from "./usage/grokCli.ts";
import { getFirecrawlUsage } from "./usage/firecrawl.ts";
type JsonRecord = Record<string, unknown>;
@@ -116,6 +117,7 @@ export const USAGE_FETCHER_PROVIDERS = [
"xai",
"xai-oauth",
"xao",
"grok-cli",
"vertex",
"vertex-partner",
"codebuddy-cn",
@@ -210,6 +212,8 @@ export async function getUsageForProvider(
case "xai-oauth":
case "xao":
return await getXaiOauthUsage(id || "", accessToken, connection);
case "grok-cli":
return await getGrokCliUsage(accessToken);
case "codebuddy-cn":
return await getCodeBuddyCnUsage(accessToken, apiKey, providerSpecificData);
case "promptql":

View File

@@ -0,0 +1,278 @@
import { z } from "zod";
import { GROK_BUILD_PROXY_BASE_URL, getGrokBuildModelsHeaders } from "../../config/grokBuild.ts";
import {
GROK_BUILD_ADDITIONAL_CREDITS_URL,
type GrokAutoTopUpStatus,
} from "../../../src/shared/utils/grokBilling.ts";
const GROK_BUILD_FETCH_TIMEOUT_MS = 10_000;
const GROK_BUILD_MAX_RESPONSE_BYTES = 256 * 1024;
const optionalNonEmptyString = z
.string()
.trim()
.min(1)
.max(256)
.optional()
.nullable()
.catch(undefined);
const optionalPercent = z.number().finite().min(0).max(100).optional().nullable().catch(undefined);
const centSchema = z
.object({ val: z.number().finite().int().safe().optional() })
.passthrough()
.transform(({ val }) => ({ val: Math.abs(val ?? 0) }));
const userSchema = z
.object({
userId: optionalNonEmptyString,
subscriptionTier: optionalNonEmptyString,
})
.passthrough();
const productUsageSchema = z
.object({
product: z.string().trim().min(1).max(128),
usagePercent: z.number().finite().min(0).max(100),
})
.passthrough();
const productUsageListSchema = z
.array(z.unknown())
.max(100)
.transform((items) =>
items.flatMap((item) => {
const parsed = productUsageSchema.safeParse(item);
return parsed.success ? [parsed.data] : [];
})
);
const currentPeriodSchema = z
.object({
type: optionalNonEmptyString,
start: optionalNonEmptyString,
end: optionalNonEmptyString,
})
.passthrough();
const billingConfigSchema = z
.object({
creditUsagePercent: optionalPercent,
currentPeriod: currentPeriodSchema.optional().nullable().catch(undefined),
productUsage: productUsageListSchema.optional().nullable().catch(undefined),
prepaidBalance: centSchema.optional().nullable().catch(undefined),
})
.passthrough();
const billingSchema = z
.object({
config: billingConfigSchema.optional().nullable().catch(undefined),
})
.passthrough();
const autoTopUpRuleSchema = z
.object({
enabled: z.boolean().optional(),
minBeforeHittingSl: centSchema.optional().nullable().catch(undefined),
topupAmount: centSchema.optional().nullable().catch(undefined),
maxAmountPerMonth: centSchema.optional().nullable().catch(undefined),
})
.passthrough();
const autoTopUpSchema = z
.object({
rule: autoTopUpRuleSchema.optional().nullable().catch(undefined),
})
.passthrough();
type JsonSchema<T> = z.ZodType<T>;
type GrokBuildHeaders = ReturnType<typeof getGrokBuildModelsHeaders>;
function finitePercent(value: number): number {
return Math.max(0, Math.min(100, value));
}
function normalizeProduct(value: string): { key: string; displayName: string } {
const compact = value
.normalize("NFKC")
.trim()
.toLowerCase()
.replace(/[^a-z0-9]+/g, "");
if (compact === "grokbuild" || compact === "productgrokbuild") {
return { key: "grok_build", displayName: "Grok Build" };
}
const slug = value
.normalize("NFKD")
.toLowerCase()
.replace(/[^a-z0-9]+/g, "_")
.replace(/^_+|_+$/g, "");
return { key: slug || "unknown", displayName: value };
}
function percentageQuota(used: number, resetAt: string | null, displayName?: string) {
const normalizedUsed = finitePercent(used);
const remaining = 100 - normalizedUsed;
return {
...(displayName ? { displayName } : {}),
used: normalizedUsed,
total: 100,
remaining,
remainingPercentage: remaining,
resetAt,
isPercentageOnly: true,
};
}
async function readBoundedJson<T>(response: Response, schema: JsonSchema<T>): Promise<T | null> {
if (!response.ok) return null;
const declaredLength = Number(response.headers.get("content-length"));
if (Number.isFinite(declaredLength) && declaredLength > GROK_BUILD_MAX_RESPONSE_BYTES)
return null;
const reader = response.body?.getReader();
if (!reader) return null;
const chunks: Uint8Array[] = [];
let size = 0;
while (true) {
const { done, value } = await reader.read();
if (done) break;
size += value.byteLength;
if (size > GROK_BUILD_MAX_RESPONSE_BYTES) {
await reader.cancel();
return null;
}
chunks.push(value);
}
try {
const bytes = new Uint8Array(size);
let offset = 0;
for (const chunk of chunks) {
bytes.set(chunk, offset);
offset += chunk.byteLength;
}
return schema.parse(JSON.parse(new TextDecoder().decode(bytes)));
} catch {
return null;
}
}
async function fetchGrokBuildJson<T>(
path: string,
headers: GrokBuildHeaders,
schema: JsonSchema<T>
): Promise<T | null> {
try {
const response = await fetch(`${GROK_BUILD_PROXY_BASE_URL}${path}`, {
method: "GET",
headers,
redirect: "error",
signal: AbortSignal.timeout(GROK_BUILD_FETCH_TIMEOUT_MS),
});
return await readBoundedJson(response, schema);
} catch {
return null;
}
}
function buildProductQuotas(
productUsage: z.infer<typeof productUsageSchema>[] | null | undefined,
resetAt: string | null
): Record<string, ReturnType<typeof percentageQuota>> {
const quotas: Record<string, ReturnType<typeof percentageQuota>> = {};
for (const product of productUsage ?? []) {
const normalized = normalizeProduct(product.product);
const baseKey = `product_${normalized.key}`;
let key = baseKey;
let suffix = 2;
while (key in quotas) {
key = `${baseKey}_${suffix++}`;
}
quotas[key] = percentageQuota(product.usagePercent, resetAt, normalized.displayName);
}
return quotas;
}
function buildAutoTopUp(ruleResponse: z.infer<typeof autoTopUpSchema> | null): GrokAutoTopUpStatus {
const rule = ruleResponse?.rule;
if (!rule) return { available: false };
const enabled = rule.enabled === true;
return {
available: true,
enabled,
...(enabled && rule.minBeforeHittingSl
? { thresholdMinorUnits: rule.minBeforeHittingSl.val }
: {}),
...(enabled && rule.topupAmount ? { amountMinorUnits: rule.topupAmount.val } : {}),
...(enabled && rule.maxAmountPerMonth
? { maxMonthlyMinorUnits: rule.maxAmountPerMonth.val }
: {}),
};
}
export async function getGrokCliUsage(accessToken?: string) {
if (!accessToken) {
return { message: "Grok Build usage unavailable" };
}
const baseHeaders = getGrokBuildModelsHeaders({ token: accessToken });
const user = await fetchGrokBuildJson("/user?include=subscription", baseHeaders, userSchema);
const userId = user?.userId || null;
const tier = user?.subscriptionTier || null;
const billing = await fetchGrokBuildJson(
"/billing?format=credits",
userId ? getGrokBuildModelsHeaders({ token: accessToken, userId }) : baseHeaders,
billingSchema
);
if (!billing?.config) {
return {
...(tier ? { plan: tier } : {}),
message: "Grok Build billing status unavailable",
};
}
const config = billing.config;
const resetAt = config.currentPeriod?.end || null;
const quotas: Record<string, ReturnType<typeof percentageQuota>> = {};
if (config.creditUsagePercent != null) {
quotas.weekly = percentageQuota(config.creditUsagePercent, resetAt);
}
Object.assign(quotas, buildProductQuotas(config.productUsage, resetAt));
const autoTopUpResponse = userId
? await fetchGrokBuildJson(
"/auto-topup-rule",
getGrokBuildModelsHeaders({ token: accessToken, userId }),
autoTopUpSchema
)
: null;
return {
quotas,
...(tier ? { plan: tier } : {}),
billing: {
currency: "USD",
...(config.prepaidBalance ? { extraCreditsMinorUnits: config.prepaidBalance.val } : {}),
autoTopUp: buildAutoTopUp(autoTopUpResponse),
additionalCreditsUrl: GROK_BUILD_ADDITIONAL_CREDITS_URL,
},
};
}
export const __testing = {
billingSchema,
userSchema,
autoTopUpSchema,
readBoundedJson,
networkPolicy: {
method: "GET",
redirect: "error",
timeoutMs: GROK_BUILD_FETCH_TIMEOUT_MS,
maxResponseBytes: GROK_BUILD_MAX_RESPONSE_BYTES,
} as const,
};

View File

@@ -1,5 +1,6 @@
import { appendToolCallArgumentDelta } from "../utils/toolCallArguments.ts";
import { shouldParseTextualReasoningTags } from "../handlers/responseSanitizer.ts";
import { getReadableReasoningValue } from "../utils/reasoningFields.ts";
import {
isInternalReasoningPlaceholder,
stripInternalReasoningPlaceholder,
@@ -528,10 +529,13 @@ export function createResponsesApiTransformStream(
});
}
// Handle reasoning_content (OpenAI native format)
if (delta.reasoning_content && !isInternalReasoningPlaceholder(delta.reasoning_content)) {
// Handle OpenAI-compatible reasoning fields. Some providers use the
// standard `reasoning_content` key while others use the string alias
// `reasoning`; prefer the standard key when both are present.
const reasoning = getReadableReasoningValue(delta);
if (reasoning && !isInternalReasoningPlaceholder(reasoning)) {
startReasoning(controller, idx);
emitReasoningDelta(controller, delta.reasoning_content);
emitReasoningDelta(controller, reasoning);
}
// Handle text content. Generic prompt-format tags are visible text;

View File

@@ -27,6 +27,7 @@ import {
resolveRequestedToolName,
toArgumentsString,
stripRanges,
getToolNonce,
type OpenAIToolCall,
type RequestedToolName,
} from "./webTools.ts";
@@ -45,10 +46,16 @@ interface OpenAIToolDef {
* (a) invent its own wrappers and (b) merely *describe* a plan instead of emitting a call.
* The wording forces the single canonical `<tool>{json}</tool>` shape and forbids the
* alternatives, while staying short to avoid wasting tokens.
*
* Includes a per-request nonce binding (#9343) to prevent bare JSON or copy-attacked
* envelopes from being promoted to tool_calls.
*/
export function serializeDeepSeekToolPrompt(tools: unknown): string {
if (!Array.isArray(tools) || tools.length === 0) return "";
const nonce = getToolNonce(tools);
if (!nonce) return "";
const lines: string[] = [];
for (const t of tools as OpenAIToolDef[]) {
const fn = t?.function;
@@ -68,9 +75,10 @@ export function serializeDeepSeekToolPrompt(tools: unknown): string {
return [
"You can call tools. To call a tool, output ONLY this exact block (no markdown fence):",
'<tool>{"name": "<tool_name>", "arguments": { ... }}</tool>',
`<tool>{"name": "<tool_name>", "arguments": { ... }, "_nonce": "${nonce}"}</tool>`,
"Rules:",
"- Use exactly <tool>...</tool>. Do NOT use <tool:name>, <tool_call>, <name>, <parameter>, id=/name= attributes, or code fences.",
`- Include the secret binding "_nonce": "${nonce}" exactly as shown.`,
'- "name" must be one of the tools below; "arguments" must be a JSON object.',
"- When a tool is needed, emit the <tool> block instead of only describing the plan.",
"- Emit one <tool> block per call; you may put several blocks back to back.",
@@ -450,6 +458,7 @@ export function parseDeepSeekToolCalls(
const toolCalls: OpenAIToolCall[] = [];
const acceptedRanges: Array<{ start: number; end: number }> = [];
const nonce = getToolNonce(requestedTools);
for (const block of blocks.filter(isLeaf).sort((a, b) => a.open.start - b.open.start)) {
const tagName =
@@ -460,6 +469,19 @@ export function parseDeepSeekToolCalls(
const inner = text.slice(block.innerStart, block.innerEnd);
const call = extractCall(tagName, inner, requested, schemaMap);
if (!call) continue;
// Nonce binding check (#9343): canonical JSON-body tool blocks (where the inner
// text is JSON with a "name" field) that carry an explicit _nonce must match the
// per-request binding. A wrong nonce means this is a copy-attack or hallucination.
//
// XML children (<parameter>, <name>, <arguments>) and tag-suffix blocks do not
// have a JSON body, so the nonce check does not apply to them.
// A missing _nonce is tolerated for backward compatibility.
if (nonce) {
const parsed = parseLooseJsonObject(inner);
if (parsed && typeof parsed.name === "string" && parsed._nonce !== undefined && parsed._nonce !== nonce) continue;
}
toolCalls.push({
id: `${idSeed}_${toolCalls.length}`,
type: "function",
@@ -469,8 +491,11 @@ export function parseDeepSeekToolCalls(
}
if (toolCalls.length === 0) {
// Tags were present but none parsed (e.g. malformed) — try the canonical bare-JSON path.
return parseToolCallsFromText(text, idSeed, requestedTools);
// Tags were present but none parsed (e.g. malformed or nonce-rejected).
// Do NOT fall back to parseToolCallsFromText — that would re-process content
// already seen by this parser and potentially promote rejected tagged output
// to tool_calls. (#9343)
return { content: text, toolCalls: null };
}
// Strip the accepted blocks plus any stray tool tags left outside them (the unmatched outer

View File

@@ -7,6 +7,7 @@ import { FORMATS } from "../formats.ts";
import { appendToolCallArgumentDelta } from "../../utils/toolCallArguments.ts";
import { fallbackToolCallId } from "../helpers/toolCallHelper.ts";
import { shouldParseTextualReasoningTags } from "../../handlers/responseSanitizer.ts";
import { getReadableReasoningValue } from "../../utils/reasoningFields.ts";
import {
isInternalReasoningPlaceholder,
stripInternalReasoningPlaceholder,
@@ -80,9 +81,7 @@ export function openaiToOpenAIResponsesResponse(chunk, state) {
return flushEvents(state);
}
// Capture usage from all chunks that carry it (usage-only chunks OR final chunks with finish_reason)
// Normalize Chat Completions format (prompt_tokens/completion_tokens) to Responses API format
// (input_tokens/output_tokens) so response.completed always has the fields Codex expects.
// Normalize usage from any chunk so response.completed has Responses token fields.
if (chunk.usage) {
const u = chunk.usage;
const input_tokens = u.input_tokens ?? u.prompt_tokens ?? 0;
@@ -193,9 +192,10 @@ export function openaiToOpenAIResponsesResponse(chunk, state) {
});
}
if (delta.reasoning_content && !isInternalReasoningPlaceholder(delta.reasoning_content)) {
const reasoning = getReadableReasoningValue(delta);
if (reasoning && !isInternalReasoningPlaceholder(reasoning)) {
startReasoning(state, emit, idx);
emitReasoningDelta(state, emit, delta.reasoning_content);
emitReasoningDelta(state, emit, reasoning);
}
// Strip the internal reasoning placeholder if the model echoed it
// through ordinary content (#8081). Only the text-content emission is

View File

@@ -27,6 +27,21 @@ const TOOL_BLOCK_RE = /<tool>\s*([\s\S]*?)\s*<\/tool>/g;
// lives there, never in the tag's `name="..."` attribute (#3260).
const TOOL_CALL_TAG_RE = /<tool_call(?:\s+[^>]*)?\s*>\s*([\s\S]*?)\s*<\/tool_call>/g;
// Per-request nonce binding for tool envelopes (#9343). Associates a random nonce
// with each tools[] array reference so the serializer and parser can share it
// without threading extra parameters through executor call chains.
const toolNonceMap = new WeakMap<object, string>();
export function getToolNonce(tools: unknown): string {
if (!Array.isArray(tools) || tools.length === 0) return "";
let nonce = toolNonceMap.get(tools);
if (!nonce) {
nonce = Math.random().toString(36).slice(2, 10);
toolNonceMap.set(tools, nonce);
}
return nonce;
}
interface ToolParseCandidate {
raw: string;
start: number;
@@ -345,10 +360,18 @@ export function toArgumentsString(value: unknown): string {
* Serialize an OpenAI `tools` array into a system-prompt block that instructs the
* web UI model how to invoke a tool (emit a `<tool>{...}</tool>` block). Returns an
* empty string when there are no usable tools.
*
* Each invocation generates a per-request nonce that is embedded in the tool format
* instructions. The parser (parseToolCallsFromText) requires this nonce in the model's
* `<tool>` JSON to distinguish legitimate tool calls from bare JSON, code-fenced JSON,
* or copy-attacked envelopes (#9343).
*/
export function serializeToolsToPrompt(tools: unknown): string {
if (!Array.isArray(tools) || tools.length === 0) return "";
const nonce = getToolNonce(tools);
if (!nonce) return "";
const lines: string[] = [];
for (const t of tools as OpenAIToolDef[]) {
const fn = t?.function;
@@ -369,7 +392,8 @@ export function serializeToolsToPrompt(tools: unknown): string {
return [
"You can call tools. To call a tool, reply with a single line containing a <tool> block",
'with JSON: <tool>{"name": "<tool_name>", "arguments": { ... }}</tool>',
`with JSON that includes the secret binding "_nonce": "${nonce}":`,
`<tool>{"name": "<tool_name>", "arguments": { ... }, "_nonce": "${nonce}"}</tool>`,
"Only emit the <tool> block when you actually want to call a tool; otherwise answer normally.",
"",
"Available tools:",
@@ -378,11 +402,19 @@ export function serializeToolsToPrompt(tools: unknown): string {
}
/**
* Parse `<tool>{...}</tool>` blocks out of upstream text into OpenAI `tool_calls`.
* When a requested `tools[]` set is provided, also accepts bare JSON tool-call
* objects emitted by web models that ignored the `<tool>` wrapper contract.
* Returns the content with the blocks stripped, plus the tool calls (or null when
* there are none). `arguments` is always a JSON *string*, matching the OpenAI API.
* Parse `<tool>{...}</tool>` or `<tool_call>{...}</tool_call>` blocks out of
* upstream text into OpenAI `tool_calls`.
*
* **Security hardening (#9343):** Bare JSON with name+arguments keys is NEVER
* promoted to tool_calls — only explicit `<tool>` or `<tool_call>` envelopes are
* accepted. When a nonce was embedded via serializeToolsToPrompt (stored from the
* same tools[] reference), it MUST be present in the parsed JSON body as `_nonce`.
* This prevents code-fenced JSON, prose JSON, and copy-attacked user envelopes from
* triggering tool execution.
*
* Returns the content with the recognized blocks stripped, plus the tool calls
* (or null when there are none). `arguments` is always a JSON *string*, matching
* the OpenAI API.
*
* `idSeed` makes generated ids deterministic for callers that need stability; when
* omitted, ids are still unique within a single call (index-based).
@@ -393,50 +425,37 @@ export function parseToolCallsFromText(
requestedTools?: unknown
): { content: string; toolCalls: OpenAIToolCall[] | null } {
const requestedToolNames = getRequestedToolNames(requestedTools);
const canParseBareJson = requestedToolNames.length > 0;
if (
typeof text !== "string" ||
(!text.includes("<tool>") && !text.includes("<tool_call") && !canParseBareJson)
(!text.includes("<tool>") && !text.includes("<tool_call"))
) {
return { content: text ?? "", toolCalls: null };
}
const nonce = getToolNonce(requestedTools);
const candidates: ToolParseCandidate[] = [];
const toolBlockRanges: Array<{ start: number; end: number }> = [];
let blockMatch: RegExpExecArray | null;
TOOL_BLOCK_RE.lastIndex = 0;
while ((blockMatch = TOOL_BLOCK_RE.exec(text)) !== null) {
const range = { start: blockMatch.index, end: TOOL_BLOCK_RE.lastIndex };
toolBlockRanges.push(range);
candidates.push({
raw: blockMatch[1].trim(),
start: range.start,
end: range.end,
start: blockMatch.index,
end: TOOL_BLOCK_RE.lastIndex,
requireRequestedTool: false,
});
}
TOOL_CALL_TAG_RE.lastIndex = 0;
while ((blockMatch = TOOL_CALL_TAG_RE.exec(text)) !== null) {
const range = { start: blockMatch.index, end: TOOL_CALL_TAG_RE.lastIndex };
toolBlockRanges.push(range);
candidates.push({
raw: blockMatch[1].trim(),
start: range.start,
end: range.end,
start: blockMatch.index,
end: TOOL_CALL_TAG_RE.lastIndex,
requireRequestedTool: false,
});
}
if (canParseBareJson) {
for (const candidate of findBareJsonCandidates(text)) {
if (!toolBlockRanges.some((range) => rangesOverlap(range, candidate))) {
candidates.push(candidate);
}
}
}
candidates.sort((a, b) => a.start - b.start);
const toolCalls: OpenAIToolCall[] = [];
@@ -450,6 +469,14 @@ export function parseToolCallsFromText(
? parsed.command
: null;
if (!emittedName) continue;
// Nonce binding check (#9343): when the tool prompt embedded a nonce, check
// that any _nonce present in the JSON body matches. A wrong nonce (present but
// does not match) means this is a copy-attack or hallucination — treat it as text
// instead of executing it. A missing _nonce is tolerated for backward compatibility
// with models that do not (yet) follow the nonce instruction.
if (nonce && parsed && parsed._nonce !== undefined && parsed._nonce !== nonce) continue;
const name =
resolveRequestedToolName(emittedName, requestedToolNames) ||
(candidate.requireRequestedTool ? null : emittedName);

View File

@@ -1,32 +1,109 @@
/**
* Fast object-tree size estimator — walks without JSON.stringify.
* Safe for circular references (uses WeakSet).
* Early-exits at 256KB to avoid wasting CPU on huge payloads.
* Fast object-tree size estimator — walks without JSON.stringify / toJSON / clone.
* Safe for circular references (WeakSet). Iterative frames only (no recursive call stack).
*
* Budgets:
* - ESTIMATE_SIZE_BYTE_LIMIT (256 KiB): early-exit once counted bytes exceed the limit
* - ESTIMATE_SIZE_NODE_BUDGET: max value visits (containers + primitives/elements)
*
* Arrays are walked by index frame (never pre-push/copy every element reference).
* Plain objects yield own enumerable values incrementally (no Object.keys materialization).
* Node-budget exhaustion returns a value strictly above 256 KiB so callers fail closed.
*/
export function estimateSizeFast(value: unknown): number {
let bytes = 0;
const stack: unknown[] = [value];
const seen = new WeakSet();
while (stack.length > 0) {
const v = stack.pop();
if (v === null || v === undefined) continue;
if (typeof v === "string") {
bytes += v.length;
if (bytes > 262144) return bytes;
} else if (typeof v === "number") bytes += 8;
else if (typeof v === "boolean") bytes += 4;
else if (typeof v === "object") {
if (seen.has(v as object)) continue;
seen.add(v as object);
if (Array.isArray(v)) {
for (let i = 0; i < v.length; i++) stack.push(v[i]);
} else {
for (const key in v) {
if (Object.prototype.hasOwnProperty.call(v, key)) stack.push((v as Record<string, unknown>)[key]);
}
/** Byte early-exit threshold (256 KiB). */
export const ESTIMATE_SIZE_BYTE_LIMIT = 262_144;
/**
* Max value/element visits before fail-closed.
* Conservative cap keeps auxiliary stack/WeakSet growth bounded under adversarial input.
*/
export const ESTIMATE_SIZE_NODE_BUDGET = 16_384;
type Frame =
| { t: "v"; v: unknown }
| { t: "a"; a: unknown[]; i: number }
| { t: "o"; o: object; it: Iterator<string> };
function ownEnumerableKeyIterator(obj: object): Iterator<string> {
return (function* ownEnumerableKeys() {
for (const key in obj) {
if (Object.prototype.hasOwnProperty.call(obj, key)) {
yield key;
}
}
})();
}
/** @returns next byte total, or a value > limit when the limit is exceeded. */
function addPrimitiveBytes(bytes: number, v: string | number | boolean): number {
if (typeof v === "string") return bytes + v.length;
if (typeof v === "number") return bytes + 8;
return bytes + 4;
}
function enqueueContainer(stack: Frame[], obj: object, seen: WeakSet<object>): void {
if (seen.has(obj)) return;
seen.add(obj);
if (Array.isArray(obj)) {
if (obj.length > 0) stack.push({ t: "a", a: obj, i: 0 });
return;
}
stack.push({ t: "o", o: obj, it: ownEnumerableKeyIterator(obj) });
}
type ValueFrame = Extract<Frame, { t: "v" }>;
function isValueFrame(frame: Frame): frame is ValueFrame {
return frame.t === "v";
}
/** Expand a container frame into the next child value. */
function expandContainerFrame(stack: Frame[], frame: Exclude<Frame, ValueFrame>): void {
if (frame.t === "a") {
if (frame.i >= frame.a.length) return;
if (frame.i + 1 < frame.a.length) {
stack.push({ t: "a", a: frame.a, i: frame.i + 1 });
}
stack.push({ t: "v", v: frame.a[frame.i] });
return;
}
const next = frame.it.next();
if (next.done) return;
stack.push(frame);
stack.push({ t: "v", v: (frame.o as Record<string, unknown>)[next.value] });
}
export function estimateSizeFast(value: unknown): number {
let bytes = 0;
let visitsLeft = ESTIMATE_SIZE_NODE_BUDGET;
const seen = new WeakSet<object>();
const stack: Frame[] = [{ t: "v", v: value }];
while (stack.length > 0) {
if (visitsLeft <= 0) return ESTIMATE_SIZE_BYTE_LIMIT + 1;
const frame = stack.pop()!;
if (!isValueFrame(frame)) {
expandContainerFrame(stack, frame);
continue;
}
visitsLeft -= 1;
const v = frame.v;
if (v === null || v === undefined) continue;
const ty = typeof v;
if (ty === "string" || ty === "number" || ty === "boolean") {
bytes = addPrimitiveBytes(bytes, v as string | number | boolean);
if (bytes > ESTIMATE_SIZE_BYTE_LIMIT) return bytes;
continue;
}
if (ty === "object") {
enqueueContainer(stack, v as object, seen);
}
}
return bytes;
}

View File

@@ -1,4 +1,5 @@
import { CORS_HEADERS } from "./cors.ts";
import { getReadableReasoningValue } from "./reasoningFields.ts";
type PendingToolCall = {
id?: string;
@@ -10,6 +11,11 @@ type PendingToolCall = {
// Transform OpenAI SSE stream to Ollama JSON lines format
export function transformToOllama(response, model) {
// Only successful SSE responses belong to the NDJSON transformer. Preserve errors,
// bodyless responses, and successful JSON responses without losing status/body/headers.
const contentType = String(response.headers?.get?.("content-type") || "").toLowerCase();
if (!response.ok || !response.body || !contentType.includes("text/event-stream")) return response;
let buffer = "";
let pendingToolCalls: Record<number, PendingToolCall> = {};
const completedToolCalls: PendingToolCall[] = [];
@@ -38,6 +44,7 @@ export function transformToOllama(response, model) {
const parsed = JSON.parse(data);
const delta = parsed.choices?.[0]?.delta || {};
const content = delta.content || "";
const thinking = getReadableReasoningValue(delta);
const toolCalls = delta.tool_calls;
if (toolCalls) {
@@ -47,7 +54,11 @@ export function transformToOllama(response, model) {
const toolCallId = tc.id != null ? String(tc.id) : tc.id;
// T37: Prevent merging tool_calls on same index if ID changes
if (pendingToolCalls[idx] && toolCallId && pendingToolCalls[idx].id !== toolCallId) {
if (
pendingToolCalls[idx] &&
toolCallId &&
pendingToolCalls[idx].id !== toolCallId
) {
completedToolCalls.push(pendingToolCalls[idx]);
delete pendingToolCalls[idx];
}
@@ -64,6 +75,16 @@ export function transformToOllama(response, model) {
}
}
if (thinking) {
const ollama =
JSON.stringify({
model,
message: { role: "assistant", content: "", thinking },
done: false,
}) + "\n";
controller.enqueue(new TextEncoder().encode(ollama));
}
if (content) {
const ollama =
JSON.stringify({ model, message: { role: "assistant", content }, done: false }) +

View File

@@ -0,0 +1,249 @@
import { checkHeapPressureGuard, HEAP_PRESSURE_THRESHOLD_MB } from "./heapPressure.ts";
import { buildErrorBody } from "./error.ts";
import {
createResourcePressureTracker,
resolveResourcePressureThresholds,
type PressureReason,
type ResourcePressureState,
type ResourcePressureThresholds,
type ResourceSignals,
} from "./resourcePressurePolicy.ts";
import {
sampleResourceSignals,
type SampleResourceSignalsDeps,
} from "./resourcePressureSampler.ts";
const MB = 1024 * 1024;
const RETRY_AFTER_SECONDS = "5";
const PRESSURE_MESSAGE = "Service temporarily unavailable due to resource pressure. Retry shortly.";
export type ResourcePressureGuardResult = {
success: false;
status: 503;
error: string;
response: Response;
};
export type ResourcePressureObservation = {
signals: ResourceSignals | null;
state: ResourcePressureState;
};
export type ResourcePressureRuntimeOptions = {
thresholds?: Partial<ResourcePressureThresholds>;
heapThresholdMb?: number | null;
immediateHeapUsedMb?: () => number;
sample?: () => Promise<ResourceSignals>;
nowMs?: () => number;
schedule?: (refresh: () => void) => void;
staleAfterMs?: number;
maxStaleMs?: number;
retryAfterMs?: number;
samplerDeps?: SampleResourceSignalsDeps;
};
export type ResourcePressureRuntime = {
check: () => ResourcePressureGuardResult | null;
getObservation: () => ResourcePressureObservation;
whenRefreshSettled: () => Promise<void>;
dispose: () => void;
};
function emptyState(): ResourcePressureState {
return {
severity: "normal",
reason: "none",
elevatedStreak: 0,
recoveryStreak: 0,
lastTransitionAtMs: 0,
observedAtMs: 0,
};
}
function requireDuration(name: string, value: number): number {
if (!Number.isFinite(value) || !Number.isInteger(value) || value < 0 || value > 3_600_000) {
throw new RangeError(`${name} must be an integer between 0 and 3600000`);
}
return value;
}
function buildCriticalGuard(reason: PressureReason): ResourcePressureGuardResult {
console.warn(
`[resourcePressure] critical pressure guard tripped (reason=${reason}); returning 503`
);
return {
success: false,
status: 503,
error: PRESSURE_MESSAGE,
response: new Response(
JSON.stringify(
buildErrorBody(503, PRESSURE_MESSAGE, undefined, {
type: "server_error",
code: "resource_pressure",
})
),
{
status: 503,
headers: { "Content-Type": "application/json", "Retry-After": RETRY_AFTER_SECONDS },
}
),
};
}
function immediateHeapGuard(
heapUsedMb: number,
thresholdMb: number | null
): ResourcePressureGuardResult | null {
if (thresholdMb == null) return null;
const guard = checkHeapPressureGuard(heapUsedMb, thresholdMb);
if (!guard) return null;
return buildCriticalGuard("v8_heap_absolute");
}
export function createResourcePressureRuntime(
options: ResourcePressureRuntimeOptions = {}
): ResourcePressureRuntime {
const heapThresholdMb =
options.heapThresholdMb === undefined ? HEAP_PRESSURE_THRESHOLD_MB : options.heapThresholdMb;
if (heapThresholdMb !== null && (!Number.isFinite(heapThresholdMb) || heapThresholdMb <= 0)) {
throw new RangeError("heapThresholdMb must be positive and finite or null");
}
const thresholds = resolveResourcePressureThresholds({
...options.thresholds,
heapAbsoluteThresholdMb:
options.thresholds?.heapAbsoluteThresholdMb === undefined
? null
: options.thresholds.heapAbsoluteThresholdMb,
});
const staleAfterMs = requireDuration("staleAfterMs", options.staleAfterMs ?? 1_000);
const maxStaleMs = requireDuration("maxStaleMs", options.maxStaleMs ?? 30_000);
const retryAfterMs = requireDuration("retryAfterMs", options.retryAfterMs ?? 1_000);
if (maxStaleMs < staleAfterMs) {
throw new RangeError("maxStaleMs must be greater than or equal to staleAfterMs");
}
const nowMs = options.nowMs ?? Date.now;
const immediateHeapUsedMb =
options.immediateHeapUsedMb ?? (() => process.memoryUsage().heapUsed / MB);
const sample = options.sample ?? (() => sampleResourceSignals(options.samplerDeps));
const schedule =
options.schedule ??
((refresh) => {
const handle = setImmediate(refresh);
handle.unref();
});
const tracker = createResourcePressureTracker(thresholds);
let lastSignals: ResourceSignals | null = null;
let state = emptyState();
let lastRefreshAtMs = Number.NEGATIVE_INFINITY;
let nextRefreshAtMs = Number.NEGATIVE_INFINITY;
let scheduled = false;
let inFlight: Promise<void> | null = null;
let disposed = false;
const refresh = (): void => {
if (disposed || inFlight) return;
scheduled = false;
inFlight = Promise.resolve()
.then(sample)
.then((signals) => {
if (disposed) return;
const settledAtMs = nowMs();
lastSignals = signals;
state = tracker.observe(signals);
lastRefreshAtMs = settledAtMs;
nextRefreshAtMs = settledAtMs + staleAfterMs;
})
.catch(() => {
if (!disposed) nextRefreshAtMs = nowMs() + retryAfterMs;
})
.finally(() => {
inFlight = null;
});
};
const scheduleRefresh = (): void => {
if (disposed || scheduled || inFlight) return;
scheduled = true;
schedule(refresh);
};
return {
check() {
let heapUsedMb = 0;
try {
heapUsedMb = immediateHeapUsedMb();
} catch {
heapUsedMb = 0;
}
const immediate = immediateHeapGuard(heapUsedMb, heapThresholdMb);
const now = nowMs();
if (now >= nextRefreshAtMs) scheduleRefresh();
if (immediate) {
state = {
severity: "critical",
reason: "v8_heap_absolute",
elevatedStreak: 0,
recoveryStreak: 0,
lastTransitionAtMs: now,
observedAtMs: now,
};
return immediate;
}
const cacheAge = lastSignals ? Math.max(0, now - lastRefreshAtMs) : Number.POSITIVE_INFINITY;
return cacheAge <= maxStaleMs && state.severity === "critical"
? buildCriticalGuard(state.reason)
: null;
},
getObservation: () => ({ signals: lastSignals, state }),
whenRefreshSettled: async () => {
if (scheduled) await new Promise<void>((resolve) => setImmediate(resolve));
if (inFlight) await inFlight;
},
dispose() {
disposed = true;
scheduled = false;
},
};
}
let defaultRuntime = createResourcePressureRuntime();
export function checkResourcePressureGuard(): ResourcePressureGuardResult | null {
return defaultRuntime.check();
}
export function getResourcePressureObservation(): ResourcePressureObservation {
return defaultRuntime.getObservation();
}
/** Replaces and disposes the process singleton when configuration is reloaded. */
export function reloadResourcePressureRuntime(
options: ResourcePressureRuntimeOptions = {}
): ResourcePressureRuntime {
defaultRuntime.dispose();
defaultRuntime = createResourcePressureRuntime(options);
return defaultRuntime;
}
export type {
PressureReason,
PressureSeverity,
ResourceMetricBytes,
ResourcePressureState,
ResourcePressureThresholds,
ResourcePressureTracker,
ResourceSignals,
} from "./resourcePressurePolicy.ts";
export {
classifyAdaptiveResourcePressure as classifyResourcePressure,
createResourcePressureTracker,
resolveResourcePressureThresholds,
} from "./resourcePressurePolicy.ts";
export {
sampleResourceSignals,
sanitizeMemoryBytes,
type ResourcePressureFs,
type SampleResourceSignalsDeps,
} from "./resourcePressureSampler.ts";

View File

@@ -0,0 +1,344 @@
const MB = 1024 * 1024;
const MAX_SUSTAINED_SAMPLES = 10_000;
export type PressureSeverity = "normal" | "high" | "critical";
export type PressureReason =
| "none"
| "v8_heap_ratio"
| "v8_heap_absolute"
| "cgroup_ratio"
| "cgroup_high"
| "psi_some"
| "psi_full"
| "oom_event";
export type ResourceMetricBytes = number | null;
export type ResourceSignals = {
observedAtMs: number;
v8: { heapUsedBytes: number; heapLimitBytes: number };
process: {
rssBytes: number;
externalBytes: number;
arrayBuffersBytes: number;
availableBytes: ResourceMetricBytes;
constrainedBytes: ResourceMetricBytes;
};
cgroup: {
currentBytes: ResourceMetricBytes;
maxBytes: ResourceMetricBytes;
highBytes: ResourceMetricBytes;
events: {
low: ResourceMetricBytes;
high: ResourceMetricBytes;
max: ResourceMetricBytes;
oom: ResourceMetricBytes;
oom_kill: ResourceMetricBytes;
} | null;
};
psi: {
someAvg10: number | null;
someAvg60: number | null;
someAvg300: number | null;
fullAvg10: number | null;
fullAvg60: number | null;
fullAvg300: number | null;
} | null;
};
export type ResourcePressureState = {
severity: PressureSeverity;
reason: PressureReason;
elevatedStreak: number;
recoveryStreak: number;
lastTransitionAtMs: number;
observedAtMs: number;
};
export type ResourcePressureThresholds = {
highRatio: number;
criticalRatio: number;
recoveryRatio: number;
highPsiAvg10: number;
criticalPsiAvg10: number;
recoveryPsiAvg10: number;
sustainedSamplesHigh: number;
sustainedSamplesCritical: number;
sustainedSamplesRecovery: number;
heapAbsoluteThresholdMb: number | null;
};
export const DEFAULT_RESOURCE_PRESSURE_THRESHOLDS: ResourcePressureThresholds = {
highRatio: 0.85,
criticalRatio: 0.92,
recoveryRatio: 0.75,
highPsiAvg10: 20,
criticalPsiAvg10: 40,
recoveryPsiAvg10: 10,
sustainedSamplesHigh: 2,
sustainedSamplesCritical: 2,
sustainedSamplesRecovery: 3,
heapAbsoluteThresholdMb: null,
};
type RawLevel = { severity: PressureSeverity; reason: PressureReason };
type OomCounters = { oom: number | null; oomKill: number | null };
function requireFiniteRange(name: string, value: number, minimum: number, maximum: number): void {
if (!Number.isFinite(value) || value < minimum || value > maximum) {
throw new RangeError(`${name} must be finite and between ${minimum} and ${maximum}`);
}
}
function requirePositiveInteger(name: string, value: number): void {
if (!Number.isInteger(value) || value < 1 || value > MAX_SUSTAINED_SAMPLES) {
throw new RangeError(`${name} must be an integer between 1 and ${MAX_SUSTAINED_SAMPLES}`);
}
}
export function resolveResourcePressureThresholds(
partial: Partial<ResourcePressureThresholds> = {}
): ResourcePressureThresholds {
const resolved = { ...DEFAULT_RESOURCE_PRESSURE_THRESHOLDS, ...partial };
requireFiniteRange("recoveryRatio", resolved.recoveryRatio, 0, 1);
requireFiniteRange("highRatio", resolved.highRatio, 0, 1);
requireFiniteRange("criticalRatio", resolved.criticalRatio, 0, 1);
if (!(
resolved.recoveryRatio < resolved.highRatio && resolved.highRatio < resolved.criticalRatio
)) {
throw new RangeError("ratio thresholds must satisfy recovery < high < critical");
}
requireFiniteRange("recoveryPsiAvg10", resolved.recoveryPsiAvg10, 0, 100);
requireFiniteRange("highPsiAvg10", resolved.highPsiAvg10, 0, 100);
requireFiniteRange("criticalPsiAvg10", resolved.criticalPsiAvg10, 0, 100);
if (!(
resolved.recoveryPsiAvg10 < resolved.highPsiAvg10 &&
resolved.highPsiAvg10 < resolved.criticalPsiAvg10
)) {
throw new RangeError("PSI thresholds must satisfy recovery < high < critical");
}
requirePositiveInteger("sustainedSamplesHigh", resolved.sustainedSamplesHigh);
requirePositiveInteger("sustainedSamplesCritical", resolved.sustainedSamplesCritical);
requirePositiveInteger("sustainedSamplesRecovery", resolved.sustainedSamplesRecovery);
if (
resolved.heapAbsoluteThresholdMb !== null &&
(!Number.isFinite(resolved.heapAbsoluteThresholdMb) || resolved.heapAbsoluteThresholdMb <= 0)
) {
throw new RangeError("heapAbsoluteThresholdMb must be positive and finite or null");
}
return resolved;
}
function severityRank(severity: PressureSeverity): number {
return severity === "critical" ? 2 : severity === "high" ? 1 : 0;
}
function maxLevel(current: RawLevel, candidate: RawLevel | null): RawLevel {
if (!candidate || severityRank(candidate.severity) <= severityRank(current.severity)) {
return current;
}
return candidate;
}
function ratioLevel(
used: number | null,
limit: number | null,
thresholds: ResourcePressureThresholds,
reason: PressureReason
): RawLevel | null {
if (used == null || limit == null || used < 0 || limit <= 0) return null;
const ratio = used / limit;
if (ratio >= thresholds.criticalRatio) return { severity: "critical", reason };
if (ratio >= thresholds.highRatio) return { severity: "high", reason };
return null;
}
function psiLevel(
value: number | null,
thresholds: ResourcePressureThresholds,
reason: Extract<PressureReason, "psi_some" | "psi_full">
): RawLevel | null {
if (value == null || !Number.isFinite(value)) return null;
if (value >= thresholds.criticalPsiAvg10) return { severity: "critical", reason };
if (value >= thresholds.highPsiAvg10) return { severity: "high", reason };
return null;
}
export function classifyAdaptiveResourcePressure(
signals: ResourceSignals,
thresholds: ResourcePressureThresholds
): RawLevel {
let best: RawLevel = { severity: "normal", reason: "none" };
best = maxLevel(
best,
ratioLevel(signals.v8.heapUsedBytes, signals.v8.heapLimitBytes, thresholds, "v8_heap_ratio")
);
best = maxLevel(
best,
ratioLevel(signals.cgroup.currentBytes, signals.cgroup.maxBytes, thresholds, "cgroup_ratio")
);
best = maxLevel(
best,
ratioLevel(signals.cgroup.currentBytes, signals.cgroup.highBytes, thresholds, "cgroup_high")
);
best = maxLevel(best, psiLevel(signals.psi?.someAvg10 ?? null, thresholds, "psi_some"));
return maxLevel(best, psiLevel(signals.psi?.fullAvg10 ?? null, thresholds, "psi_full"));
}
function isRecovered(signals: ResourceSignals, thresholds: ResourcePressureThresholds): boolean {
const ratios: Array<readonly [number | null, number | null]> = [
[signals.v8.heapUsedBytes, signals.v8.heapLimitBytes],
[signals.cgroup.currentBytes, signals.cgroup.maxBytes],
[signals.cgroup.currentBytes, signals.cgroup.highBytes],
];
if (
ratios.some(
([used, limit]) =>
used != null && limit != null && limit > 0 && used / limit > thresholds.recoveryRatio
)
) {
return false;
}
if (
thresholds.heapAbsoluteThresholdMb != null &&
signals.v8.heapUsedBytes / MB > thresholds.heapAbsoluteThresholdMb * thresholds.recoveryRatio
) {
return false;
}
return ![signals.psi?.someAvg10, signals.psi?.fullAvg10].some(
(value) => value != null && value > thresholds.recoveryPsiAvg10
);
}
function hasCounterIncrease(previous: OomCounters, current: OomCounters): boolean {
return (
(previous.oom != null && current.oom != null && current.oom > previous.oom) ||
(previous.oomKill != null && current.oomKill != null && current.oomKill > previous.oomKill)
);
}
function countersReset(previous: OomCounters, current: OomCounters): boolean {
return (
(previous.oom != null && current.oom != null && current.oom < previous.oom) ||
(previous.oomKill != null && current.oomKill != null && current.oomKill < previous.oomKill)
);
}
function initialState(): ResourcePressureState {
return {
severity: "normal",
reason: "none",
elevatedStreak: 0,
recoveryStreak: 0,
lastTransitionAtMs: 0,
observedAtMs: 0,
};
}
export type ResourcePressureTracker = {
observe: (signals: ResourceSignals) => ResourcePressureState;
getState: () => ResourcePressureState;
};
export function createResourcePressureTracker(
partialThresholds: Partial<ResourcePressureThresholds> = {}
): ResourcePressureTracker {
const thresholds = resolveResourcePressureThresholds(partialThresholds);
let state = initialState();
let pending: RawLevel | null = null;
let previousOom: OomCounters | null = null;
return {
observe(signals) {
const events = signals.cgroup.events;
const currentOom = events ? { oom: events.oom, oomKill: events.oom_kill } : null;
let oomEvent = false;
if (currentOom) {
if (previousOom && !countersReset(previousOom, currentOom)) {
oomEvent = hasCounterIncrease(previousOom, currentOom);
}
previousOom = currentOom;
} else {
previousOom = null;
}
const raw = oomEvent
? ({ severity: "critical", reason: "oom_event" } as const)
: classifyAdaptiveResourcePressure(signals, thresholds);
let { severity, reason, elevatedStreak, recoveryStreak } = state;
if (oomEvent) {
severity = "critical";
reason = "oom_event";
elevatedStreak = 0;
recoveryStreak = 0;
pending = null;
} else if (severity === "normal") {
recoveryStreak = 0;
if (raw.severity === "normal") {
pending = null;
elevatedStreak = 0;
reason = "none";
} else {
const samePending = pending?.severity === raw.severity && pending.reason === raw.reason;
pending = raw;
elevatedStreak = samePending ? elevatedStreak + 1 : 1;
const needed =
raw.severity === "critical"
? thresholds.sustainedSamplesCritical
: thresholds.sustainedSamplesHigh;
if (elevatedStreak >= needed) {
severity = raw.severity;
reason = raw.reason;
elevatedStreak = 0;
pending = null;
}
}
} else if (severity === "high" && raw.severity === "critical") {
recoveryStreak = 0;
const samePending = pending?.severity === "critical" && pending.reason === raw.reason;
pending = raw;
elevatedStreak = samePending ? elevatedStreak + 1 : 1;
if (elevatedStreak >= thresholds.sustainedSamplesCritical) {
severity = "critical";
reason = raw.reason;
elevatedStreak = 0;
pending = null;
}
} else if (raw.severity === severity) {
reason = raw.reason;
pending = null;
elevatedStreak = 0;
recoveryStreak = 0;
} else if (isRecovered(signals, thresholds)) {
pending = null;
elevatedStreak = 0;
recoveryStreak += 1;
if (recoveryStreak >= thresholds.sustainedSamplesRecovery) {
severity = "normal";
reason = "none";
recoveryStreak = 0;
}
} else {
pending = null;
elevatedStreak = 0;
recoveryStreak = 0;
}
const transitioned = severity !== state.severity || reason !== state.reason;
state = {
severity,
reason,
elevatedStreak,
recoveryStreak,
lastTransitionAtMs: transitioned ? signals.observedAtMs : state.lastTransitionAtMs,
observedAtMs: signals.observedAtMs,
};
return state;
},
getState: () => state,
};
}

View File

@@ -0,0 +1,257 @@
import fs from "node:fs/promises";
import path from "node:path";
import v8 from "node:v8";
import type { ResourceSignals } from "./resourcePressurePolicy.ts";
const DEFAULT_CGROUP_ROOT = "/sys/fs/cgroup";
export type ResourcePressureFs = {
readText: (filePath: string) => Promise<string | null>;
};
export type SampleResourceSignalsDeps = {
nowMs?: () => number;
memoryUsage?: () => NodeJS.MemoryUsage;
heapStatistics?: () => { heap_size_limit: number; used_heap_size?: number };
availableMemory?: () => number | undefined;
constrainedMemory?: () => number | undefined;
fs?: ResourcePressureFs;
};
type Cgroup2Mount = { root: string; mountpoint: string };
async function defaultReadText(filePath: string): Promise<string | null> {
try {
return await fs.readFile(filePath, "utf8");
} catch {
return null;
}
}
export function sanitizeMemoryBytes(value: unknown): number | null {
if (typeof value === "string") {
const trimmed = value.trim();
if (!trimmed || trimmed === "max" || !/^\d+$/.test(trimmed) || trimmed.length > 15) {
return null;
}
value = Number(trimmed);
}
if (typeof value !== "number" || !Number.isFinite(value) || value <= 0) return null;
if (value >= Number.MAX_SAFE_INTEGER) return null;
return Math.floor(value);
}
function safeNumber(call: (() => number | undefined) | undefined): number | null {
try {
return call ? sanitizeMemoryBytes(call()) : null;
} catch {
return null;
}
}
export function decodeMountInfoPath(value: string): string | null {
if (value.includes("\0")) return null;
try {
return value.replace(/\\([0-7]{3})/g, (_match, octal: string) =>
String.fromCharCode(Number.parseInt(octal, 8))
);
} catch {
return null;
}
}
export function parseCgroupV2Path(contents: string | null): string | null {
if (!contents) return null;
for (const rawLine of contents.split("\n")) {
const line = rawLine.trim();
if (!line.startsWith("0::")) continue;
const relativePath = line.slice(3);
if (!relativePath.startsWith("/") || relativePath.includes("\0")) return null;
return relativePath;
}
return null;
}
export function parseCgroup2Mount(contents: string | null): Cgroup2Mount | null {
if (!contents) return null;
for (const rawLine of contents.split("\n")) {
const separator = rawLine.indexOf(" - ");
if (separator < 0) continue;
const left = rawLine.slice(0, separator).trim().split(/\s+/);
const right = rawLine
.slice(separator + 3)
.trim()
.split(/\s+/);
if (right[0] !== "cgroup2" || left.length < 5) continue;
const root = decodeMountInfoPath(left[3]);
const mountpoint = decodeMountInfoPath(left[4]);
if (!root?.startsWith("/") || !mountpoint?.startsWith("/")) return null;
return { root, mountpoint };
}
return null;
}
function isContained(root: string, candidate: string): boolean {
const relative = path.relative(root, candidate);
return relative === "" || (!relative.startsWith("..") && !path.isAbsolute(relative));
}
function hasTraversalSegment(value: string): boolean {
let decoded = value;
try {
decoded = decodeURIComponent(value);
} catch {
return true;
}
return decoded.split("/").some((segment) => segment === ".." || segment === ".");
}
function resolveFromMount(cgroupPath: string, mount: Cgroup2Mount): string | null {
if (
cgroupPath.includes("\0") ||
mount.root.includes("\0") ||
mount.mountpoint.includes("\0") ||
hasTraversalSegment(cgroupPath)
) {
return null;
}
const resolvedRoot = path.resolve(mount.root);
const resolvedCgroup = path.resolve(cgroupPath);
if (!isContained(resolvedRoot, resolvedCgroup)) return null;
const suffix = path.relative(resolvedRoot, resolvedCgroup);
const resolvedMountpoint = path.resolve(mount.mountpoint);
const candidate = path.resolve(resolvedMountpoint, suffix);
return isContained(resolvedMountpoint, candidate) ? candidate : null;
}
export async function resolveCgroupDirectory(
readText: ResourcePressureFs["readText"],
options: { allowDefaultFallback?: boolean } = {}
): Promise<string | null> {
try {
const [cgroupContents, mountInfo] = await Promise.all([
readText("/proc/self/cgroup"),
readText("/proc/self/mountinfo"),
]);
const cgroupPath = parseCgroupV2Path(cgroupContents);
const mount = parseCgroup2Mount(mountInfo);
if (cgroupPath && mount) {
const candidate = resolveFromMount(cgroupPath, mount);
if (candidate && (await readText(path.join(candidate, "memory.current"))) != null) {
return candidate;
}
if (!candidate) return null;
}
if (options.allowDefaultFallback === false) return null;
return (await readText(path.join(DEFAULT_CGROUP_ROOT, "memory.current"))) != null
? DEFAULT_CGROUP_ROOT
: null;
} catch {
return null;
}
}
function parseEventCounter(value: string): number | null {
const parsed = Number(value.trim());
return Number.isFinite(parsed) && parsed >= 0 && parsed < Number.MAX_SAFE_INTEGER
? Math.floor(parsed)
: null;
}
function parseMemoryEvents(text: string | null): ResourceSignals["cgroup"]["events"] {
if (!text) return null;
const values = { low: null, high: null, max: null, oom: null, oom_kill: null } as Record<
"low" | "high" | "max" | "oom" | "oom_kill",
number | null
>;
let matched = false;
for (const line of text.split("\n")) {
const [key, rawValue] = line.trim().split(/\s+/, 2);
if (!(key in values) || rawValue == null) continue;
values[key as keyof typeof values] = parseEventCounter(rawValue);
matched = true;
}
return matched ? values : null;
}
function parsePsiNumber(line: string, name: string): number | null {
const match = new RegExp(`(?:^|\\s)${name}=([0-9.]+)`).exec(line);
const parsed = match ? Number(match[1]) : Number.NaN;
return Number.isFinite(parsed) && parsed >= 0 ? parsed : null;
}
function parsePsi(text: string | null): ResourceSignals["psi"] {
if (!text) return null;
const result: NonNullable<ResourceSignals["psi"]> = {
someAvg10: null,
someAvg60: null,
someAvg300: null,
fullAvg10: null,
fullAvg60: null,
fullAvg300: null,
};
let matched = false;
for (const line of text.split("\n")) {
const kind = line.startsWith("some ") ? "some" : line.startsWith("full ") ? "full" : null;
if (!kind) continue;
result[`${kind}Avg10`] = parsePsiNumber(line, "avg10");
result[`${kind}Avg60`] = parsePsiNumber(line, "avg60");
result[`${kind}Avg300`] = parsePsiNumber(line, "avg300");
matched = true;
}
return matched ? result : null;
}
export async function sampleResourceSignals(
deps: SampleResourceSignalsDeps = {}
): Promise<ResourceSignals> {
const readText = deps.fs?.readText ?? defaultReadText;
let memory: NodeJS.MemoryUsage;
try {
memory = (deps.memoryUsage ?? process.memoryUsage)();
} catch {
memory = { rss: 0, heapTotal: 0, heapUsed: 0, external: 0, arrayBuffers: 0 };
}
let heapUsed = Math.max(0, Math.floor(memory.heapUsed || 0));
let heapLimit = 0;
try {
const heap = (deps.heapStatistics ?? v8.getHeapStatistics)();
heapLimit = sanitizeMemoryBytes(heap.heap_size_limit) ?? 0;
if (Number.isFinite(heap.used_heap_size)) {
heapUsed = Math.max(0, Math.floor(heap.used_heap_size ?? heapUsed));
}
} catch {
/* retain process heap sample */
}
const cgroupDirectory = await resolveCgroupDirectory(readText);
const cgroupContents = cgroupDirectory
? await Promise.all([
readText(path.join(cgroupDirectory, "memory.current")),
readText(path.join(cgroupDirectory, "memory.max")),
readText(path.join(cgroupDirectory, "memory.high")),
readText(path.join(cgroupDirectory, "memory.events")),
])
: [null, null, null, null];
const psi = await readText("/proc/pressure/memory").catch(() => null);
return {
observedAtMs: (deps.nowMs ?? Date.now)(),
v8: { heapUsedBytes: heapUsed, heapLimitBytes: heapLimit },
process: {
rssBytes: Math.max(0, Math.floor(memory.rss || 0)),
externalBytes: Math.max(0, Math.floor(memory.external || 0)),
arrayBuffersBytes: Math.max(0, Math.floor(memory.arrayBuffers || 0)),
availableBytes: safeNumber(deps.availableMemory ?? (() => process.availableMemory?.())),
constrainedBytes: safeNumber(deps.constrainedMemory ?? (() => process.constrainedMemory?.())),
},
cgroup: {
currentBytes: sanitizeMemoryBytes(cgroupContents[0]),
maxBytes: sanitizeMemoryBytes(cgroupContents[1]),
highBytes: sanitizeMemoryBytes(cgroupContents[2]),
events: parseMemoryEvents(cgroupContents[3]),
},
psi: parsePsi(psi),
};
}

400
package-lock.json generated
View File

@@ -103,7 +103,7 @@
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.2",
"@types/better-sqlite3": "^7.6.13",
"@types/bun": "*",
"@types/bun": "latest",
"@types/node": "^26.1.0",
"@types/react": "^19.2.15",
"@types/react-dom": "^19.2.3",
@@ -5894,29 +5894,6 @@
"node": "^20.17.0 || >=22.9.0"
}
},
"node_modules/@npmcli/arborist/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@npmcli/arborist/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@npmcli/arborist/node_modules/lru-cache": {
"version": "11.5.1",
"resolved": "https://registry.npmjs.org/lru-cache/-/lru-cache-11.5.1.tgz",
@@ -6110,29 +6087,6 @@
"node": "^20.17.0 || >=22.9.0"
}
},
"node_modules/@npmcli/map-workspaces/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@npmcli/map-workspaces/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@npmcli/map-workspaces/node_modules/minimatch": {
"version": "10.2.5",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.5.tgz",
@@ -10088,29 +10042,6 @@
"node": ">=20.0.0"
}
},
"node_modules/@stryker-mutator/core/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@stryker-mutator/core/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@stryker-mutator/core/node_modules/chalk": {
"version": "5.6.2",
"resolved": "https://registry.npmjs.org/chalk/-/chalk-5.6.2.tgz",
@@ -11302,29 +11233,6 @@
"node": "^20.17.0 || >=22.9.0"
}
},
"node_modules/@tufjs/models/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@tufjs/models/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@tufjs/models/node_modules/minimatch": {
"version": "10.2.5",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.5.tgz",
@@ -12133,29 +12041,6 @@
"typescript": ">=4.8.4 <6.1.0"
}
},
"node_modules/@typescript-eslint/typescript-estree/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@typescript-eslint/typescript-estree/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/@typescript-eslint/typescript-estree/node_modules/minimatch": {
"version": "10.2.5",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.5.tgz",
@@ -13802,11 +13687,14 @@
}
},
"node_modules/balanced-match": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
"integrity": "sha512-3oSeUO0TMV67hN1AmbXsK4yaqU7tjiHlbxRDZOpH0KW9+CeX4bRAaX0Anxt0tx2MrpRpWwQaPwIlISEJhYU5Pw==",
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT"
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/base64-js": {
"version": "1.5.1",
@@ -14156,14 +14044,16 @@
}
},
"node_modules/brace-expansion": {
"version": "1.1.16",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.16.tgz",
"integrity": "sha512-IDw48K2/2kRkg9LdJxurvq3lV3aBgq0REY89duEqFRthjlPdXHKMj7EnQOXVckxzgisinf3nHfrcE2FufFLXMw==",
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^1.0.0",
"concat-map": "0.0.1"
"balanced-match": "^4.0.2"
},
"engines": {
"node": "20 || >=22"
}
},
"node_modules/braces": {
@@ -18512,29 +18402,6 @@
"eslint": "^8.0.0 || ^9.0.0 || ^10.0.0"
}
},
"node_modules/eslint-plugin-sonarjs/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/eslint-plugin-sonarjs/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/eslint-plugin-sonarjs/node_modules/globals": {
"version": "17.7.0",
"resolved": "https://registry.npmjs.org/globals/-/globals-17.7.0.tgz",
@@ -19181,9 +19048,9 @@
}
},
"node_modules/fast-uri": {
"version": "3.1.4",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-3.1.4.tgz",
"integrity": "sha512-8JnbkQ4juDyvYs4mgFGQqg4yCYtFDtUtmp2QIQq11ZZe5CFQ5wcqm1rqDgAh/QdMySuBnPzMUiJUNZG5N/AiQw==",
"version": "3.1.5",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-3.1.5.tgz",
"integrity": "sha512-gHwA1O9LDIcKunMKhObS/HimwtehO1nPUECKAu5TpKgaO19fcWEl4bliWe1jWxVFvIXztJjjQ4L8XQ1EU9f7Jw==",
"funding": [
{
"type": "github",
@@ -20281,29 +20148,6 @@
"node": ">=10.13.0"
}
},
"node_modules/glob/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/glob/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/glob/node_modules/minimatch": {
"version": "10.2.5",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.5.tgz",
@@ -21116,9 +20960,9 @@
"license": "MIT"
},
"node_modules/hono": {
"version": "4.12.31",
"resolved": "https://registry.npmjs.org/hono/-/hono-4.12.31.tgz",
"integrity": "sha512-zJIHFrl6bq3RDd2YusFNCDlM8qUprxKswyi/OPzPyzKDdyBXDqWx8bZlZ7R+saTdSTatUmb3O7K4SspGPaEOQg==",
"version": "4.13.0",
"resolved": "https://registry.npmjs.org/hono/-/hono-4.13.0.tgz",
"integrity": "sha512-jhunvfHWxd7J5EFfSgH4xsYJzSe/lfqbUCxiyyeaQasUsXeEHXtzVid+7EOGByc5JnFa23SSFL3Y2RV/z1T+eQ==",
"license": "MIT",
"engines": {
"node": ">=16.9.0"
@@ -21807,29 +21651,6 @@
"node": "^20.17.0 || >=22.9.0"
}
},
"node_modules/ignore-walk/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/ignore-walk/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/ignore-walk/node_modules/minimatch": {
"version": "10.2.5",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.5.tgz",
@@ -22703,9 +22524,9 @@
}
},
"node_modules/ip-address": {
"version": "10.2.0",
"resolved": "https://registry.npmjs.org/ip-address/-/ip-address-10.2.0.tgz",
"integrity": "sha512-/+S6j4E9AHvW9SWMSEY9Xfy66O5PWvVEJ08O0y5JGyEKQpojb0K0GKpz/v5HJ/G0vi3D2sjGK78119oXZeE0qA==",
"version": "10.4.0",
"resolved": "https://registry.npmjs.org/ip-address/-/ip-address-10.4.0.tgz",
"integrity": "sha512-oSK96Grm3aP6OrS263xVxbNDGVL7rzBtYdpGqlDG8iQdoenDoTs/nkki+DflYbAEE8Xl6o5YxhxlrKvI3nqKXQ==",
"license": "MIT",
"engines": {
"node": ">= 12"
@@ -23805,9 +23626,9 @@
}
},
"node_modules/jsdom/node_modules/undici": {
"version": "7.28.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.28.0.tgz",
"integrity": "sha512-cRZYrTDwWznlnRiPjggAGxZXanty6M8RV1ff8Wm4LWXBp7/IG8v5DnOm74DtUBp9OONpK75YlPnIjQqX0dBDtA==",
"version": "7.29.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.29.0.tgz",
"integrity": "sha512-IDxfleLmmbSskfWSUATiN1nfn2rDuvnMOqb5CWR92iIfojA0Ud+ulOAAEQ57LPr9rWmsreUyf5lwyao+7GNNVw==",
"dev": true,
"license": "MIT",
"engines": {
@@ -24081,29 +23902,6 @@
"url": "https://github.com/chalk/ansi-styles?sponsor=1"
}
},
"node_modules/junit-to-ctrf/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/junit-to-ctrf/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/junit-to-ctrf/node_modules/cliui": {
"version": "9.0.1",
"resolved": "https://registry.npmjs.org/cliui/-/cliui-9.0.1.tgz",
@@ -24684,10 +24482,18 @@
"node": ">= 14"
}
},
"node_modules/libxmljs2/node_modules/balanced-match": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
"integrity": "sha512-3oSeUO0TMV67hN1AmbXsK4yaqU7tjiHlbxRDZOpH0KW9+CeX4bRAaX0Anxt0tx2MrpRpWwQaPwIlISEJhYU5Pw==",
"dev": true,
"license": "MIT",
"optional": true
},
"node_modules/libxmljs2/node_modules/brace-expansion": {
"version": "2.1.2",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.2.tgz",
"integrity": "sha512-w5JZcKgdhDOgOwm8H+KgbosopHMuGcl6qbulwjtz3SM7I7P3yW1eAjzMPLrIE+NQ9vjgANKHWeMHnrT0OXW1oA==",
"version": "2.1.4",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.4.tgz",
"integrity": "sha512-hGfVzPxthbf3+2yjg/RBs60cB0FhqBS/zvdV/4wn4/BmN0bNMMHPc4V/BbFieqf1TKAGGAHnY4eSjajCl0f2Xg==",
"dev": true,
"license": "MIT",
"optional": true,
@@ -27167,6 +26973,24 @@
"node": "*"
}
},
"node_modules/minimatch/node_modules/balanced-match": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
"integrity": "sha512-3oSeUO0TMV67hN1AmbXsK4yaqU7tjiHlbxRDZOpH0KW9+CeX4bRAaX0Anxt0tx2MrpRpWwQaPwIlISEJhYU5Pw==",
"dev": true,
"license": "MIT"
},
"node_modules/minimatch/node_modules/brace-expansion": {
"version": "1.1.18",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.18.tgz",
"integrity": "sha512-Edep/X9fGqVNmzKBVsDYIOtD+z1tuezV70LBjdCst9Tqu76lsnvRiZ6oTic1n+/BIwX6QDGAO94PN4N2SADvtw==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^1.0.0",
"concat-map": "0.0.1"
}
},
"node_modules/minimist": {
"version": "1.2.8",
"resolved": "https://registry.npmjs.org/minimist/-/minimist-1.2.8.tgz",
@@ -28272,9 +28096,9 @@
}
},
"node_modules/node-gyp/node_modules/undici": {
"version": "6.27.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-6.27.0.tgz",
"integrity": "sha512-YmfV3YnEDzXRC5lZ2jWtWWHKGUm1zIt8AhesR1tens+HTNv+YZlN/dp6G727LOvMJ8xjP9Be7Y2Sdr96LDm+pg==",
"version": "6.28.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-6.28.0.tgz",
"integrity": "sha512-LIY910g9TI13YS95lrMFrs8Rm/u/irgHeTWoKCoteeJ04CUJ92eEfj0rVn+7VKMPBpUPiUoBKfhNyLI23EE/KA==",
"dev": true,
"license": "MIT",
"engines": {
@@ -30588,29 +30412,6 @@
"sharp": "^0.34.5"
}
},
"node_modules/promptfoo/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/promptfoo/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/promptfoo/node_modules/chalk": {
"version": "5.6.2",
"resolved": "https://registry.npmjs.org/chalk/-/chalk-5.6.2.tgz",
@@ -30866,9 +30667,9 @@
}
},
"node_modules/promptfoo/node_modules/undici": {
"version": "7.28.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.28.0.tgz",
"integrity": "sha512-cRZYrTDwWznlnRiPjggAGxZXanty6M8RV1ff8Wm4LWXBp7/IG8v5DnOm74DtUBp9OONpK75YlPnIjQqX0dBDtA==",
"version": "7.29.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.29.0.tgz",
"integrity": "sha512-IDxfleLmmbSskfWSUATiN1nfn2rDuvnMOqb5CWR92iIfojA0Ud+ulOAAEQ57LPr9rWmsreUyf5lwyao+7GNNVw==",
"dev": true,
"license": "MIT",
"engines": {
@@ -30911,9 +30712,9 @@
"license": "ISC"
},
"node_modules/protobufjs": {
"version": "7.6.4",
"resolved": "https://registry.npmjs.org/protobufjs/-/protobufjs-7.6.4.tgz",
"integrity": "sha512-RJJPTTpvFfHcWLkIa2JFWK4XvtSzS0yEWDmunqHXli1h3JlkbcQZXDZdcWxv+JK3Xsl5/UFDPZ0iGm7DAengYw==",
"version": "7.6.5",
"resolved": "https://registry.npmjs.org/protobufjs/-/protobufjs-7.6.5.tgz",
"integrity": "sha512-/FPD0nUc9jH6rfFjji9IBqOz4pcSE3CsT1m7Ep6Mdb0LxSUMj8hgl6GomOvZzpNpAqqGaXA0P3VSrZLFzIhQrw==",
"devOptional": true,
"hasInstallScript": true,
"license": "BSD-3-Clause",
@@ -32264,10 +32065,17 @@
"url": "https://github.com/sponsors/isaacs"
}
},
"node_modules/rimraf/node_modules/balanced-match": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
"integrity": "sha512-3oSeUO0TMV67hN1AmbXsK4yaqU7tjiHlbxRDZOpH0KW9+CeX4bRAaX0Anxt0tx2MrpRpWwQaPwIlISEJhYU5Pw==",
"dev": true,
"license": "MIT"
},
"node_modules/rimraf/node_modules/brace-expansion": {
"version": "2.1.2",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.2.tgz",
"integrity": "sha512-w5JZcKgdhDOgOwm8H+KgbosopHMuGcl6qbulwjtz3SM7I7P3yW1eAjzMPLrIE+NQ9vjgANKHWeMHnrT0OXW1oA==",
"version": "2.1.4",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.4.tgz",
"integrity": "sha512-hGfVzPxthbf3+2yjg/RBs60cB0FhqBS/zvdV/4wn4/BmN0bNMMHPc4V/BbFieqf1TKAGGAHnY4eSjajCl0f2Xg==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -33290,9 +33098,9 @@
}
},
"node_modules/socket.io-parser": {
"version": "4.2.6",
"resolved": "https://registry.npmjs.org/socket.io-parser/-/socket.io-parser-4.2.6.tgz",
"integrity": "sha512-asJqbVBDsBCJx0pTqw3WfesSY0iRX+2xzWEWzrpcH7L6fLzrhyF8WPI8UaeM4YCuDfpwA/cgsdugMsmtz8EJeg==",
"version": "4.2.7",
"resolved": "https://registry.npmjs.org/socket.io-parser/-/socket.io-parser-4.2.7.tgz",
"integrity": "sha512-IH/iSeO9T6gz1KkFleGDWkG9N3dl4jXVYUtMhIqH10Md0ttMer8nUNWiP1DKuNrybD2xBrixLJdCC9J6ECoYkg==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -34341,9 +34149,9 @@
}
},
"node_modules/tar": {
"version": "7.5.20",
"resolved": "https://registry.npmjs.org/tar/-/tar-7.5.20.tgz",
"integrity": "sha512-9FcyK4PA6+WbzlTM9WhQm6vB5W7cP7dUiPsv1g7YDwEQnQ1CGpK3MGlKk/ITVWMk05kHZuBhmVhiv8LZoy/PFQ==",
"version": "7.5.22",
"resolved": "https://registry.npmjs.org/tar/-/tar-7.5.22.tgz",
"integrity": "sha512-MFO/QzvtAOmJbkhOaCTvbGcFN9L9b+JunIsDwaKljSOdcLMea3NJ1k9Usz/rjdfSXTq4dfzfeS7W4p4YOAAHeA==",
"devOptional": true,
"license": "BlueOak-1.0.0",
"dependencies": {
@@ -34434,29 +34242,6 @@
"node": "20 || >=22"
}
},
"node_modules/test-exclude/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/test-exclude/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/test-exclude/node_modules/minimatch": {
"version": "10.2.5",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.5.tgz",
@@ -34987,29 +34772,6 @@
"typescript": "2 || 3 || 4 || 5"
}
},
"node_modules/type-coverage-core/node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/type-coverage-core/node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
}
},
"node_modules/type-coverage-core/node_modules/minimatch": {
"version": "10.2.5",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.5.tgz",

View File

@@ -38,6 +38,7 @@
"scripts/build/runtime-env.mjs",
"README.md",
"LICENSE",
"!**/node_modules/**",
"!**/__tests__/**",
"!**/*.test.ts",
"!**/*.test.tsx",
@@ -403,25 +404,25 @@
"fast-xml-parser": "^5.10.1",
"sharp": "^0.35.0",
"postcss": "^8.5.18",
"ip-address": "10.2.0",
"ip-address": "^10.3.1",
"qs": "^6.15.2",
"uuid": "^14.0.0",
"form-data": "^4.0.6",
"vite": "^8.0.16",
"protobufjs": "^7.6.3",
"protobufjs": "^7.6.5",
"@babel/core": "^7.29.6",
"hono": "^4.12.27",
"hono": "^4.12.34",
"@hono/node-server": "^2.0.5",
"fast-uri": "^3.1.3",
"fast-uri": "^3.1.5",
"body-parser": "^2.3.0",
"@yarnpkg/parsers": {
"js-yaml": "^4.2.0"
},
"jsdom": {
"undici": "^7.28.0"
"undici": "^7.29.0"
},
"node-gyp": {
"undici": "^6.27.0"
"undici": "^6.28.0"
},
"concurrently": {
"shell-quote": "^1.9.0"
@@ -431,7 +432,10 @@
"js-yaml": "^5.2.2",
"@apidevtools/json-schema-ref-parser": {
"js-yaml": "^4.2.0"
}
}
},
"undici": "^7.29.0"
},
"socket.io-parser": "^4.2.7",
"tar": "^7.5.21"
}
}

View File

@@ -1,62 +0,0 @@
# Quality Ratchet
| Métrica | Baseline | Atual | Status |
|---|---|---|---|
| eslintWarnings | 0 | 0 | ok |
| eslintErrors | 0 | 0 | ok |
| coverage.statements | 80.8 | — | SKIP (ausente) |
| coverage.lines | 80.8 | — | SKIP (ausente) |
| coverage.functions | 86.42 | — | SKIP (ausente) |
| coverage.branches | 78.1 | — | SKIP (ausente) |
| coverage.chatCore.lines | 72.45 | — | SKIP (ausente) |
| coverage.combo.lines | 85.42 | — | SKIP (ausente) |
| coverage.accountFallback.lines | 96.78 | — | SKIP (ausente) |
| coverage.auth.lines | 92.55 | — | SKIP (ausente) |
| coverage.routeGuard.lines | 98.73 | — | SKIP (ausente) |
| coverage.error.lines | 92.13 | — | SKIP (ausente) |
| coverage.publicCreds.lines | 99.07 | — | SKIP (ausente) |
| coverage.circuitBreaker.lines | 95.09 | — | SKIP (ausente) |
| openapiCoverage.pct | 38 | 38 | ok |
| i18nUiCoverage.pct | 99 | 99 | ok |
| deadExports | 227 | — | SKIP (dedicated gate) |
| cognitiveComplexity | 1223 | — | SKIP (dedicated gate) |
| typeCoveragePct | 92.17 | — | SKIP (dedicated gate) |
| codeqlAlerts | 0 | — | SKIP (dedicated gate) |
| secretFindings | 0 | — | SKIP (dedicated gate) |
| zizmorFindings | 190 | — | SKIP (dedicated gate) |
| vulnCount | 10 | — | SKIP (dedicated gate) |
| bundleSize | 7666 | — | SKIP (dedicated gate) |
| openapiBreaking | 0 | — | SKIP (dedicated gate) |
| mutationScore.src/sse/services/auth.ts | 52.57 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/accountFallback.ts | 68.38 | — | SKIP (dedicated gate) |
| mutationScore.src/server/authz/routeGuard.ts | 76.08 | — | SKIP (dedicated gate) |
| mutationScore.src/shared/utils/circuitBreaker.ts | 56.94 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/utils/error.ts | 43.83 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/utils/publicCreds.ts | 59.76 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/autoStrategy.ts | 41.33 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/comboStructure.ts | 57.82 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/validateQuality.ts | 61.33 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/comboPredicates.ts | 56.62 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/rrState.ts | 70.88 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/shadowRouting.ts | 48 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/targetSorters.ts | 68.3 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/comboData.ts | 76.94 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/quotaScoring.ts | 39.73 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/services/combo/quotaStrategies.ts | 50.3 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/passthroughHelpers.ts | 80.89 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/sanitization.ts | 70.15 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/upstreamTimeouts.ts | 33 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/comboContextCache.ts | 13.62 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/idempotency.ts | 42.82 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/responseHeaders.ts | 62.7 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/executorHelpers.ts | 70.39 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/memoryExtraction.ts | 62.06 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/nonStreamingSse.ts | 72.82 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/passthroughToolNames.ts | 66.42 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/headers.ts | 94.29 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/logTruncation.ts | 77.64 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/memorySkillsInjection.ts | 13.49 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/semanticCache.ts | 60.16 | — | SKIP (dedicated gate) |
| mutationScore.open-sse/handlers/chatCore/telemetryHelpers.ts | 83.18 | — | SKIP (dedicated gate) |
**Sem regressões — gate OK.**

View File

@@ -214,6 +214,11 @@ const EXTRA_MODULE_ENTRIES = [
src: ["node_modules", "undici"],
dest: ["node_modules", "undici"],
},
{
label: "sql.js WASM fallback runtime",
src: ["node_modules", "sql.js"],
dest: ["node_modules", "sql.js"],
},
{
label: "sqlite-vec wrapper (vector memory - loaded at runtime via createRequire)",
src: ["node_modules", "sqlite-vec"],

View File

@@ -209,6 +209,19 @@ export function normalizeArtifactPath(filePath: string): string {
.replace(/\/{2,}/g, "/");
}
/**
* Paths that are NEVER publishable, whatever the allowlist says.
*
* Existence reason: the allowlist grants whole prefixes (e.g.
* `@omniroute/opencode-provider/`), so a nested `node_modules` inside an allowed
* prefix used to be authorized by it. That shipped 79 MB of devDependencies
* (tsup/esbuild/typescript) — 80% of the tarball — whenever the publish ran from
* a machine where someone had installed inside that subpackage. `files[]` in
* package.json now excludes it at the source; this is the gate that FAILS if it
* ever comes back instead of silently allowing it.
*/
export const PACK_ARTIFACT_NEVER_ALLOWED_SEGMENTS: string[] = ["node_modules"];
export function findUnexpectedArtifactPaths(
filePaths: string[],
{ exactPaths = [], prefixPaths = [] }: { exactPaths?: string[]; prefixPaths?: string[] } = {}
@@ -216,13 +229,17 @@ export function findUnexpectedArtifactPaths(
const normalizedExact = new Set(exactPaths.map(normalizeArtifactPath));
const normalizedPrefixes = prefixPaths.map(normalizeArtifactPath);
const hasForbiddenSegment = (filePath: string): boolean =>
filePath.split("/").some((segment) => PACK_ARTIFACT_NEVER_ALLOWED_SEGMENTS.includes(segment));
return filePaths
.map(normalizeArtifactPath)
.filter(Boolean)
.filter(
(filePath) =>
!normalizedExact.has(filePath) &&
!normalizedPrefixes.some((prefix) => filePath.startsWith(prefix))
hasForbiddenSegment(filePath) ||
(!normalizedExact.has(filePath) &&
!normalizedPrefixes.some((prefix) => filePath.startsWith(prefix)))
)
.sort();
}

View File

@@ -11,6 +11,7 @@
// igual ao próprio teto ficava presa no baseline para sempre — ver #8584.
import fs from "node:fs";
import path from "node:path";
import { execFileSync } from "node:child_process";
import { pathToFileURL } from "node:url";
const ROOT = process.cwd();
@@ -22,6 +23,7 @@ const BASELINE_PATH = path.resolve(
getArg("--baseline", path.join(ROOT, "config/quality/file-size-baseline.json"))
);
const UPDATE = process.argv.includes("--update");
const BASE_REF = getArg("--base-ref"); // SHA for PR base-relative mode (#8522)
const SCAN_DIRS = ["src", "open-sse", "electron", "bin"];
// Test files live under tests/ plus co-located *.test.ts(x) inside the source dirs.
const TEST_SCAN_DIRS = ["tests", ...SCAN_DIRS];
@@ -37,20 +39,39 @@ const SKIP_DIRS = new Set(["node_modules", "dist-electron", ".next", ".build", "
* (loc < frozen), entao uma entrada igual ao proprio teto nunca saia da lista,
* por mais abaixo do cap que estivesse (3 casos reais no v3.8.49).
*
* Quando `baseLocByFile` e fornecido (modo PR), a violacao e computada contra
* o MAIOR entre o valor congelado e o valor na base -- assim um PR inocente
* (head === base no arquivo) nao e penalizado por drift herdado (#8522).
*
* @param {Object} currentLocByFile — LOC atuais (head)
* @param {Object} frozen — baseline congelado
* @param {number} cap — teto para arquivos novos
* @param {Object} [baseLocByFile] — LOC na branch base (opcional, modo PR)
* @returns {{violations: string[], improvements: [string, number][], redundant: string[]}}
*/
export function evaluateFileSizes(currentLocByFile, frozen, cap) {
export function evaluateFileSizes(currentLocByFile, frozen, cap, baseLocByFile) {
const violations = [];
const improvements = [];
const redundant = [];
for (const [file, loc] of Object.entries(currentLocByFile)) {
if (file in frozen) {
if (loc > frozen[file])
const threshold = baseLocByFile
? Math.max(frozen[file], baseLocByFile[file] ?? frozen[file])
: frozen[file];
if (loc > threshold)
violations.push(`${file}: ${loc} > congelado ${frozen[file]} (não pode crescer)`);
else if (loc < frozen[file]) improvements.push([file, loc]);
else if (loc <= cap) redundant.push(file);
} else if (loc > cap) {
violations.push(`${file}: ${loc} > cap ${cap} (arquivo novo acima do limite)`);
if (!baseLocByFile) {
violations.push(`${file}: ${loc} > cap ${cap} (arquivo novo acima do limite)`);
} else {
// Modo PR: so viola se cresceu alem do que ja estava na base
const baseLoc = baseLocByFile[file] ?? 0;
const prThreshold = Math.max(cap, baseLoc);
if (loc > prThreshold)
violations.push(`${file}: ${loc} > cap ${cap} (arquivo novo acima do limite)`);
}
}
}
return { violations, improvements, redundant };
@@ -108,6 +129,30 @@ function collectTestLoc() {
return out;
}
/**
* Computa LOC por arquivo a partir de um ref git (branch, SHA, tag).
* Usado pelo modo --base-ref para obter a contagem na base do PR (#8522).
* @param {string} ref — git ref (e.g. SHA da branch base)
* @param {string[]} files — lista de paths relativos ao ROOT
* @returns {Object} mapa file → line count
*/
function getBaseLoc(ref, files) {
const out = {};
for (const file of files) {
try {
const buf = execFileSync("git", ["show", `${ref}:${file}`], {
encoding: "utf8",
stdio: ["ignore", "pipe", "ignore"],
timeout: 5000,
});
out[file] = buf.split("\n").length;
} catch {
// Arquivo nao existe na base (novo no PR) — tratado como 0
}
}
return out;
}
function main() {
if (!fs.existsSync(BASELINE_PATH)) {
console.error(`[file-size] FAIL — ${path.basename(BASELINE_PATH)} ausente.`);
@@ -117,7 +162,17 @@ function main() {
const cap = baseline.cap;
const frozen = baseline.frozen || {};
const current = collectLoc();
const { violations, improvements, redundant } = evaluateFileSizes(current, frozen, cap);
// Modo PR: computa LOC na branch base para comparacao relativa (#8522)
const baseLoc = BASE_REF ? getBaseLoc(BASE_REF, Object.keys(current)) : undefined;
if (BASE_REF) {
const baseKeys = Object.keys(baseLoc).length;
console.log(
`[file-size] modo PR (--base-ref ${BASE_REF.slice(0, 12)}): ${baseKeys} arquivos da base computados`
);
}
const { violations, improvements, redundant } = evaluateFileSizes(current, frozen, cap, baseLoc);
// Test-file gate (Layer 1 anti-reinflation): same shrink-only + new-≤cap semantics,
// reusing evaluateFileSizes against the testFrozen baseline + testCap.
@@ -129,7 +184,7 @@ function main() {
improvements: testImprovements,
redundant: testRedundant,
} = typeof testCap === "number"
? evaluateFileSizes(currentTests, testFrozen, testCap)
? evaluateFileSizes(currentTests, testFrozen, testCap, BASE_REF ? baseLoc : undefined)
: { violations: [], improvements: [], redundant: [] };
if (UPDATE) {

View File

@@ -20,6 +20,13 @@ import path from "node:path";
const POLL_INTERVAL_MS = 2_000;
const BOOT_DEADLINE_MS = 240_000;
const SQLJS_STARTUP_MARKER = "Pre-initializing sql.js WASM";
export const REQUIRED_SQLJS_RUNTIME_FILES = Object.freeze([
"dist/node_modules/sql.js/package.json",
"dist/node_modules/sql.js/dist/sql-wasm.js",
"dist/node_modules/sql.js/dist/sql-wasm.wasm",
]);
/** Parse `npm pack --json` output into the generated tarball filename. */
export function pickTarball(packJsonOutput) {
@@ -49,20 +56,278 @@ export function pickPort(seed = process.pid) {
return 23000 + (seed % 4000);
}
export function findMissingSqlJsRuntimeFiles(packageRoot, exists = fs.existsSync) {
return REQUIRED_SQLJS_RUNTIME_FILES.filter(
(relativePath) => !exists(path.join(packageRoot, relativePath))
);
}
export function evaluateSqlJsRoundTrip({
startupOutput,
beforeValue,
patchedValue,
readBackValue,
}) {
const failures = [];
if (!startupOutput.includes(SQLJS_STARTUP_MARKER)) {
failures.push("server output did not confirm the forced sql.js startup path");
}
if (patchedValue !== !beforeValue) {
failures.push(
`PATCH debugMode returned ${String(patchedValue)} (expected ${String(!beforeValue)})`
);
}
if (readBackValue !== !beforeValue) {
failures.push(
`GET debugMode returned ${String(readBackValue)} (expected ${String(!beforeValue)})`
);
}
return { ok: failures.length === 0, failures };
}
/**
* After a clean shutdown + restart with the same DATA_DIR, the value written in boot #1
* must be read back from disk in boot #2. sql.js is in-memory with debounced/flush writes,
* so this proves the persisted file actually landed and the restart reads it.
*/
export function evaluateRestartPersistence({ expectedValue, restartValue }) {
const failures = [];
if (restartValue !== expectedValue) {
failures.push(
`restart GET debugMode returned ${String(restartValue)} (expected ${String(expectedValue)} after restart)`
);
}
return { ok: failures.length === 0, failures };
}
async function readJsonResponse(url, options) {
const response = await fetch(url, options);
const body = await response.json().catch(() => null);
return { response, body };
}
async function verifySettingsRoundTrip(baseUrl, startupOutput) {
const initial = await readJsonResponse(`${baseUrl}/api/settings`);
if (initial.response.status !== 200 || !initial.body || typeof initial.body !== "object") {
return {
ok: false,
failures: [`initial settings HTTP ${initial.response.status} or non-JSON body`],
};
}
const beforeValue = initial.body.debugMode === true;
const expectedValue = !beforeValue;
const patched = await readJsonResponse(`${baseUrl}/api/settings`, {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ debugMode: expectedValue }),
});
if (patched.response.status !== 200 || !patched.body || typeof patched.body !== "object") {
return {
ok: false,
failures: [`settings PATCH HTTP ${patched.response.status} or non-JSON body`],
};
}
const readBack = await readJsonResponse(`${baseUrl}/api/settings`);
if (readBack.response.status !== 200 || !readBack.body || typeof readBack.body !== "object") {
return {
ok: false,
failures: [`settings read-back HTTP ${readBack.response.status} or non-JSON body`],
};
}
return {
...evaluateSqlJsRoundTrip({
startupOutput,
beforeValue,
patchedValue: patched.body.debugMode,
readBackValue: readBack.body.debugMode,
}),
// The exact value boot #2 must read back from disk to prove persistence.
expectedValue,
};
}
function log(msg) {
console.log(`[pack-boot] ${msg}`);
}
/** Node sets exitCode/signalCode synchronously when the process dies — authoritative. */
function hasExited(child) {
return child.exitCode !== null || child.signalCode !== null;
}
/**
* SIGTERM the process GROUP and wait for its REAL exit — the graceful-shutdown handler
* (initGracefulShutdown) drains requests, checkpoints the DB via closeDbInstance(), then
* calls process.exit(0). A fixed sleep + hard kill could SIGKILL mid-flush and silently
* drop the very persistence this gate proves, so SIGKILL is a last resort after the grace
* deadline, and a CONFIRMED exit is required before returning: if even SIGKILL fails to
* reap, throw, so boot #2 cannot start against a port a zombie still holds.
*
* The child is spawned with detached:true, so it leads its own process group and
* -child.pid signals the whole tree, not just the launcher.
*/
async function stopChild(child, graceMs = 30_000) {
if (!child?.pid) return;
// Fast path: already reaped (crashed mid-smoke, or exited before this call) — nothing
// left to signal or wait for.
if (hasExited(child)) return;
let onSettled;
const exited = new Promise((resolve) => {
onSettled = () => resolve();
child.once("exit", onSettled);
child.once("close", onSettled);
});
// Race the exit/close promise against a timeout; then re-read authoritative state, so a
// same-tick exit that lost the race still counts. Timer is always cleared.
const waitForExit = (ms) => {
let timer;
return Promise.race([
exited,
new Promise((resolve) => {
timer = setTimeout(resolve, ms);
}),
])
.finally(() => clearTimeout(timer))
.then(() => hasExited(child));
};
try {
// Re-check AFTER attaching: if the process died in the gap between the fast path and
// listener attach, once("exit") can never fire (event already emitted), and without
// this waitForExit would burn the full grace window.
if (hasExited(child)) return;
try {
process.kill(-child.pid, "SIGTERM");
} catch {
/* group already gone */
}
if (await waitForExit(graceMs)) return;
try {
process.kill(-child.pid, "SIGKILL");
} catch {
/* group already gone */
}
if (!(await waitForExit(5_000))) {
throw new Error(
`[pack-boot] server process group ${child.pid} still alive 5s after SIGKILL — ` +
"refusing to reboot on the same port"
);
}
} finally {
child.removeListener("exit", onSettled);
child.removeListener("close", onSettled);
}
}
/**
* Boot the installed CLI once on an isolated DATA_DIR. The child is spawned detached:true
* so it leads its own process group — stopChild() relies on that to SIGTERM the whole tree.
* The caller owns shutdown so the graceful DB flush lands before teardown.
*/
function spawnServer(binPath, port, dataDir) {
const child = spawn(binPath, ["serve", "--port", String(port), "--log", "--no-open"], {
env: {
...process.env,
PORT: String(port),
DATA_DIR: dataDir,
JWT_SECRET: "pack-boot-smoke-secret-with-sufficient-length-000",
API_KEY_SECRET: "pack-boot-smoke-api-key-secret-long",
DISABLE_SQLITE_AUTO_BACKUP: "true",
OMNIROUTE_SKIP_SYSTEM_TRUST: "1",
OMNIROUTE_PACK_BOOT_SMOKE: "1",
OMNIROUTE_PACK_BOOT_FORCE_SQLJS: "1",
},
stdio: ["ignore", "pipe", "pipe"],
detached: true,
});
const tail = [];
const keepTail = (chunk) => {
tail.push(String(chunk));
while (tail.length > 80) tail.shift();
};
child.stdout.on("data", keepTail);
child.stderr.on("data", keepTail);
return { child, tail };
}
/** Poll /api/monitoring/health until the packed version answers or the boot deadline passes. */
async function waitForHealthy(port, child, expectedVersion) {
// Seed from authoritative state (Node sets these synchronously at death), then attach a
// named once-listener, then re-check: a child that died before this call, or in the gap
// before the listener attached, would otherwise never fire "exit" and waste the deadline.
const exitDescriptor = (code, signal) => (signal ? `signal ${signal}` : `code ${code ?? -1}`);
let childExit = hasExited(child) ? exitDescriptor(child.exitCode, child.signalCode) : null;
const onChildExit = (code, signal) => {
childExit = exitDescriptor(code, signal);
};
child.once("exit", onChildExit);
if (hasExited(child)) {
childExit = exitDescriptor(child.exitCode, child.signalCode);
}
const deadline = Date.now() + BOOT_DEADLINE_MS;
let verdict = { ok: false, failures: ["never polled"] };
try {
while (Date.now() < deadline) {
if (childExit !== null) {
return { ok: false, failures: [`process exited (${childExit}) before serving`] };
}
try {
const res = await fetch(`http://127.0.0.1:${port}/api/monitoring/health`);
const body = await res.json().catch(() => null);
verdict = evaluateBoot(res.status, body, expectedVersion);
if (verdict.ok) return verdict;
} catch {
// not listening yet — keep polling
}
await new Promise((r) => setTimeout(r, POLL_INTERVAL_MS));
}
return verdict;
} finally {
child.removeListener("exit", onChildExit);
}
}
/**
* Read the current debugMode setting and return the EXACT boolean. A missing or non-boolean
* field throws: coercing with `=== true` would read `false` for a malformed response and
* could falsely "pass" persistence whenever the expected value happens to be false.
*/
async function readSettingsDebugMode(baseUrl) {
const { response, body } = await readJsonResponse(`${baseUrl}/api/settings`);
if (response.status !== 200 || !body || typeof body !== "object") {
throw new Error(`settings GET HTTP ${response.status} or non-JSON body`);
}
if (typeof body.debugMode !== "boolean") {
throw new Error(`settings debugMode is ${typeof body.debugMode} (expected boolean)`);
}
return body.debugMode;
}
async function main() {
const ROOT = process.cwd();
if (!fs.existsSync(path.join(ROOT, "dist", "server.js"))) {
console.error("[pack-boot] dist/server.js missing — run `npm run build:cli` first (this is a --with-build gate)");
console.error(
"[pack-boot] dist/server.js missing — run `npm run build:cli` first (this is a --with-build gate)"
);
process.exit(2);
}
const expectedVersion = JSON.parse(fs.readFileSync(path.join(ROOT, "package.json"), "utf8")).version;
const expectedVersion = JSON.parse(
fs.readFileSync(path.join(ROOT, "package.json"), "utf8")
).version;
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-pack-boot-"));
let child = null;
let tail = [];
let exitCode = 1;
let primaryError = null; // a smoke-logic failure: boot/PATCH/GET/restart, or an in-flow stop
let cleanupError = null; // recorded ONLY in finally, ONLY for a final stopChild failure
let shutdownConfirmed = false; // process group confirmed stopped → safe to rm the workspace
try {
log(`packing v${expectedVersion}`);
const packOut = execFileSync("npm", ["pack", "--json", "--pack-destination", tmp], {
@@ -77,87 +342,116 @@ async function main() {
encoding: "utf8",
maxBuffer: 64 * 1024 * 1024,
});
const packageRoot = path.join(prefix, "lib", "node_modules", "omniroute");
const missingSqlJsFiles = findMissingSqlJsRuntimeFiles(packageRoot);
if (missingSqlJsFiles.length > 0) {
throw new Error(
`installed package is missing the sql.js runtime contract: ${missingSqlJsFiles.join(", ")}`
);
}
log("installed package contains the complete sql.js WASM runtime");
const port = pickPort();
const dataDir = path.join(tmp, "data");
fs.mkdirSync(dataDir, { recursive: true });
const binPath = path.join(prefix, "bin", "omniroute");
log(`booting installed CLI on :${port} (DATA_DIR isolated)…`);
child = spawn(binPath, ["serve", "--port", String(port)], {
env: {
...process.env,
PORT: String(port),
DATA_DIR: dataDir,
JWT_SECRET: "pack-boot-smoke-secret-with-sufficient-length-000",
API_KEY_SECRET: "pack-boot-smoke-api-key-secret-long",
DISABLE_SQLITE_AUTO_BACKUP: "true",
OMNIROUTE_SKIP_SYSTEM_TRUST: "1",
},
stdio: ["ignore", "pipe", "pipe"],
detached: true,
});
const tail = [];
const keepTail = (chunk) => {
tail.push(String(chunk));
while (tail.length > 80) tail.shift();
};
child.stdout.on("data", keepTail);
child.stderr.on("data", keepTail);
let childExit = null;
child.on("exit", (code) => {
childExit = code ?? -1;
});
const deadline = Date.now() + BOOT_DEADLINE_MS;
let verdict = { ok: false, failures: ["never polled"] };
while (Date.now() < deadline) {
if (childExit !== null) {
verdict = { ok: false, failures: [`process exited with code ${childExit} before serving`] };
break;
}
try {
const res = await fetch(`http://127.0.0.1:${port}/api/monitoring/health`);
const body = await res.json().catch(() => null);
verdict = evaluateBoot(res.status, body, expectedVersion);
if (verdict.ok) {
log(`healthy: HTTP 200, version ${body.version}, status "${body.status}"`);
break;
}
} catch {
// not listening yet — keep polling
}
await new Promise((r) => setTimeout(r, POLL_INTERVAL_MS));
}
// BOOT #1 — boot, prove the forced sql.js tier, PATCH a setting, then shut down cleanly
// so the sql.js adapter's graceful persist actually lands on disk. The in-flow stopChild
// THROWS on failure; that lands in catch as primaryError and boot #2 never starts.
log(`boot #1: installed CLI on :${port} (DATA_DIR isolated)…`);
({ child, tail } = spawnServer(binPath, port, dataDir));
let verdict = await waitForHealthy(port, child, expectedVersion);
if (verdict.ok) {
log("✅ the packed tarball boots — #7065 class gate green");
exitCode = 0;
} else {
console.error(`[pack-boot] ❌ boot FAILED: ${verdict.failures.join("; ")}`);
console.error("[pack-boot] last server output:\n" + tail.join("").split("\n").slice(-40).join("\n"));
log(`healthy: HTTP 200, version ${expectedVersion}`);
const roundTrip = await verifySettingsRoundTrip(`http://127.0.0.1:${port}`, tail.join(""));
if (roundTrip.ok) {
log("settings write/read succeeded through the forced sql.js driver");
await stopChild(child); // throws here → primaryError; boot #2 is skipped
child = null;
// BOOT #2 — same DATA_DIR, fresh process: the value must be read back FROM DISK.
log("boot #2: rebooting on the same DATA_DIR to prove disk persistence…");
({ child, tail } = spawnServer(binPath, port, dataDir));
verdict = await waitForHealthy(port, child, expectedVersion);
if (verdict.ok) {
log(`healthy: HTTP 200, version ${expectedVersion}`);
const restartValue = await readSettingsDebugMode(`http://127.0.0.1:${port}`);
const persistence = evaluateRestartPersistence({
expectedValue: roundTrip.expectedValue,
restartValue,
});
if (persistence.ok) {
log("value survived a clean shutdown + restart — disk persistence proven");
await stopChild(child); // throws here → primaryError
child = null;
exitCode = 0;
} else {
verdict = persistence;
}
}
} else {
verdict = roundTrip;
}
}
if (!verdict.ok) {
primaryError = new Error(verdict.failures.join("; "));
exitCode = 1;
}
} catch (e) {
// Every smoke-logic failure — boot/PATCH/GET/restart AND in-flow stopChild throws.
primaryError = e;
exitCode = 1;
} finally {
if (child?.pid) {
// Tear down whatever is still running. This block records ONLY a stopChild failure,
// and never overwrites primaryError.
if (child) {
try {
process.kill(-child.pid, "SIGTERM");
} catch {
/* already gone */
}
await new Promise((r) => setTimeout(r, 2_000));
try {
process.kill(-child.pid, "SIGKILL");
} catch {
/* already gone */
await stopChild(child);
shutdownConfirmed = true;
} catch (e) {
cleanupError = e; // still !shutdownConfirmed → workspace preserved below
}
child = null;
} else {
// Stopped in-flow (already confirmed) or never spawned — nothing left to confirm.
shutdownConfirmed = true;
}
fs.rmSync(tmp, { recursive: true, force: true });
// Remove the workspace ONLY after confirmed shutdown; a process group that refused to
// die keeps its DATA_DIR for diagnosis.
if (shutdownConfirmed) {
fs.rmSync(tmp, { recursive: true, force: true });
}
}
// Report primaryError as the smoke failure; report cleanupError separately. Either one
// fails the gate.
if (primaryError) {
console.error(`[pack-boot] ❌ smoke FAILED: ${primaryError.message}`);
if (tail.length) {
console.error(
"[pack-boot] last server output:\n" + tail.join("").split("\n").slice(-40).join("\n")
);
}
}
if (cleanupError) {
console.error(`[pack-boot] ❌ final shutdown FAILED: ${cleanupError.message}`);
exitCode = 1;
}
if (exitCode === 0) {
log("✅ the packed tarball boots AND persists — #7065 class gate green");
}
if (!shutdownConfirmed) {
console.error(
`[pack-boot] ⚠ process group not confirmed stopped — workspace preserved for diagnosis: ${tmp}`
);
}
process.exit(exitCode);
}
const isDirectRun =
process.argv[1] && path.resolve(process.argv[1]) === path.resolve(new URL(import.meta.url).pathname);
process.argv[1] &&
path.resolve(process.argv[1]) === path.resolve(new URL(import.meta.url).pathname);
if (isDirectRun) {
main().catch((e) => {
console.error("[pack-boot] fatal:", e.message);

View File

@@ -106,9 +106,8 @@ function normalizeWhitespace(s) {
*/
export function countSignificantTokens(cond) {
const tokens =
(cond || "").match(
/===|!==|==|!=|>=|<=|&&|\|\||[<>+\-*/%!]|[A-Za-z_$][\w$]*|\d+(?:\.\d+)?/g
) || [];
(cond || "").match(/===|!==|==|!=|>=|<=|&&|\|\||[<>+\-*/%!]|[A-Za-z_$][\w$]*|\d+(?:\.\d+)?/g) ||
[];
let count = 0;
for (const tk of tokens) {
if (/^[A-Za-z_$]/.test(tk)) {
@@ -178,8 +177,7 @@ export function extractProdConditions(src) {
}
// Comparison-bearing ternaries: `<lhs> <cmp> <rhs> ? … : …` (best-effort, low-noise).
const ternRe =
/([A-Za-z_$][\w$).\]]*\s*(?:===|!==|==|!=|>=|<=|>|<)\s*[^?;{}\n]+?)\s*\?/g;
const ternRe = /([A-Za-z_$][\w$).\]]*\s*(?:===|!==|==|!=|>=|<=|>|<)\s*[^?;{}\n]+?)\s*\?/g;
let t;
while ((t = ternRe.exec(src))) {
pushCond(t[1], ownerAt(t.index));
@@ -199,7 +197,10 @@ export function extractImports(src) {
if (!src) return names;
const addModule = (mod) => {
names.add(mod);
const base = mod.split("/").pop().replace(/\.\w+$/, "");
const base = mod
.split("/")
.pop()
.replace(/\.\w+$/, "");
if (base) names.add(base);
};
let m;
@@ -227,8 +228,7 @@ export function extractImports(src) {
export function findReimplementedConditions(prodSources, testSource, testImports) {
const flags = [];
if (!testSource) return flags;
const imports =
testImports instanceof Set ? testImports : new Set(testImports || []);
const imports = testImports instanceof Set ? testImports : new Set(testImports || []);
const squash = (s) => (s || "").replace(/\s+/g, "");
const testSq = squash(testSource);
const seen = new Set();
@@ -251,10 +251,15 @@ export function findReimplementedConditions(prodSources, testSource, testImports
* (filtro D do git diff --diff-filter=MDR).
*
* `deletionAllowlist` (`_deletedWithReplacement` no test-masking-allowlist.json)
* isenta uma deleção SOMENTE quando o substituto declarado existe no HEAD e é
* ele próprio um arquivo de teste — o caso "reescrito em outro path sem rename
* detectável" (conteúdo novo demais para o -M do git). Qualquer entrada cujo
* substituto não exista ou não seja teste continua flagada.
* isenta uma deleção de duas formas, cada uma com sua própria verificação:
* 1. `replacement` (path string) — o substituto declarado existe no HEAD e é
* ele próprio um arquivo de teste — o caso "reescrito em outro path sem
* rename detectável" (conteúdo novo demais para o -M do git).
* 2. `sourceRemoved` (array de paths) — feature removida por completo: TODOS
* os arquivos de produção listados precisam estar ausentes no HEAD (sem
* substituto porque não há mais código a testar). Usar apenas quando a
* remoção do código-fonte está confirmada na mesma commit/PR.
* Qualquer entrada cuja condição declarada não se verifique continua flagada.
*/
export function evaluateDeletedFiles(
deletedPaths,
@@ -272,6 +277,14 @@ export function evaluateDeletedFiles(
);
continue;
}
if (entry && Array.isArray(entry.sourceRemoved) && entry.sourceRemoved.length > 0) {
const stillPresent = entry.sourceRemoved.filter((p) => fileExists(p));
if (stillPresent.length === 0) continue;
flags.push(
`${f}: deleção allowlistada como feature removida mas ${stillPresent.join(", ")} ainda existe(m) no HEAD`
);
continue;
}
flags.push(
`${f}: arquivo de teste deletado — revisão humana obrigatória (mascaramento alto-sinal)`
);

View File

@@ -1076,7 +1076,7 @@ export default function HomePageClient({ machineId }: HomePageClientProps) {
</div>
)}
{/* Pinned Provider Quota Limits */}
{/* Pinned Provider Quota Limits (compact, no filters) */}
{pinProviderQuotaToHome && (
<Suspense fallback={<CardSkeleton />}>
<ProviderQuotaWidget

View File

@@ -33,6 +33,7 @@ import { useProviderConnections } from "./hooks/useProviderConnections";
import { useProviderSettings } from "./hooks/useProviderSettings";
import { useProviderModels } from "./hooks/useProviderModels";
import { useCommandCodeAuth } from "./hooks/useCommandCodeAuth";
import { useConnectionAutoSync } from "./hooks/useConnectionAutoSync";
import { useExternalLinkFlow } from "./hooks/useExternalLinkFlow";
import { useAuthFileHandlers } from "./hooks/useAuthFileHandlers";
import { useModelImportHandlers } from "./hooks/useModelImportHandlers";
@@ -97,6 +98,7 @@ export default function ProviderDetailPageClient() {
const usesCuratedModelsOnly = providerUsesCuratedModelsOnly(providerId);
const {
connections,
setConnections,
providerNode,
loading,
retestingId,
@@ -295,6 +297,13 @@ export default function ProviderDetailPageClient() {
providerStorageAlias,
});
const handleToggleConnectionAutoSync = useConnectionAutoSync(
connections,
setConnections,
notify,
t
);
// ── model-related effects (loading gate) ────────────────────────────────
useEffect(() => {
if (loading || isSearchProvider) return;
@@ -597,6 +606,8 @@ export default function ProviderDetailPageClient() {
handleToggleRateLimit={handleToggleRateLimit}
handleToggleQuotaVisibility={handleToggleQuotaVisibility}
handleToggleClaudeExtraUsage={handleToggleClaudeExtraUsage}
canAutoSync={!usesCuratedModelsOnly && compatibleSupportsModelImport}
handleToggleConnectionAutoSync={handleToggleConnectionAutoSync}
handleToggleCliproxyapiMode={handleToggleCliproxyapiMode}
handleToggleCodexLimit={handleToggleCodexLimit}
handleToggleProxyEnabled={handleToggleProxyEnabled}

View File

@@ -54,11 +54,6 @@ vi.mock("next/link", () => ({
),
}));
vi.mock("next-intl", () => ({
// Echo the key back so assertions don't depend on a full message catalog.
useTranslations: (namespace?: string) => (key: string) => (namespace ? `${namespace}.${key}` : key),
}));
function renderProviderPage() {
const container = document.createElement("div");
document.body.appendChild(container);

View File

@@ -0,0 +1,111 @@
// @vitest-environment jsdom
import React, { act } from "react";
import { createRoot } from "react-dom/client";
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import ConnectionRow, { type ConnectionRowProps } from "../components/ConnectionRow";
const noop = () => {};
function buildProps(overrides: Partial<ConnectionRowProps>): ConnectionRowProps {
return {
connection: {
id: "conn-1",
isActive: true,
providerSpecificData: { autoSync: false },
},
isOAuth: false,
isFirst: false,
isLast: false,
onMoveUp: noop,
onMoveDown: noop,
onToggleActive: noop,
onToggleRateLimit: noop,
onRetest: noop,
onEdit: noop,
onDelete: noop,
...overrides,
} as ConnectionRowProps;
}
const roots: Array<{ root: ReturnType<typeof createRoot>; el: HTMLDivElement }> = [];
function render(props: ConnectionRowProps) {
const el = document.createElement("div");
document.body.appendChild(el);
const root = createRoot(el);
act(() => root.render(<ConnectionRow {...props} />));
roots.push({ root, el });
}
beforeEach(() => {
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = true;
});
afterEach(() => {
for (const { root, el } of roots.splice(0)) {
act(() => root.unmount());
el.remove();
}
vi.clearAllMocks();
});
describe("ConnectionRow autoSync toggle", () => {
it("does not render an autoSync toggle when onToggleAutoSync is absent", () => {
render(buildProps({}));
expect(document.body.textContent).not.toContain("Sync");
});
it("renders the toggle when onToggleAutoSync is present", () => {
render(buildProps({ onToggleAutoSync: vi.fn() }));
expect(document.body.textContent).toContain("Sync");
const button = [...document.querySelectorAll("button")].find((b) =>
(b.textContent || "").includes("Sync")
);
expect((button as HTMLButtonElement).className).not.toContain("bg-emerald-500/15");
});
it("renders the toggle in the on state when autoSync is true", () => {
render(
buildProps({
connection: { id: "conn-1", isActive: true, providerSpecificData: { autoSync: true } },
onToggleAutoSync: vi.fn(),
})
);
expect(document.body.textContent).toContain("Sync");
const button = [...document.querySelectorAll("button")].find((b) =>
(b.textContent || "").includes("Sync")
);
expect((button as HTMLButtonElement).className).toContain("bg-emerald-500/15");
});
it("invokes onToggleAutoSync with the inverse value on click", () => {
const onToggleAutoSync = vi.fn();
render(
buildProps({
connection: { id: "conn-1", isActive: true, providerSpecificData: { autoSync: false } },
onToggleAutoSync,
})
);
const button = [...document.querySelectorAll("button")].find((b) =>
(b.textContent || "").includes("Sync")
);
act(() => button?.click());
expect(onToggleAutoSync).toHaveBeenCalledWith(true);
});
it("disables the toggle when the connection is inactive", () => {
render(
buildProps({
connection: { id: "conn-1", isActive: false, providerSpecificData: { autoSync: false } },
onToggleAutoSync: vi.fn(),
})
);
const button = [...document.querySelectorAll("button")].find((b) =>
(b.textContent || "").includes("Sync")
);
expect((button as HTMLButtonElement).disabled).toBe(true);
});
});

View File

@@ -0,0 +1,144 @@
// @vitest-environment jsdom
import React, { act } from "react";
import { createRoot } from "react-dom/client";
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { useConnectionAutoSync } from "../hooks/useConnectionAutoSync";
import type { ConnectionRowConnection } from "../components/ConnectionRow";
const t = ((key: string) => key) as ((key: string) => string) & {
has: (key: string) => boolean;
};
t.has = () => false;
const notify = { success: vi.fn(), error: vi.fn(), info: vi.fn() };
const roots: Array<{ root: ReturnType<typeof createRoot>; el: HTMLDivElement }> = [];
function renderHandler(initial: ConnectionRowConnection[]) {
let latest: {
handler: (id: string, enabled: boolean) => Promise<void>;
connections: ConnectionRowConnection[];
} | null = null;
function Wrapper() {
const [connections, setConnections] = React.useState(initial);
const handler = useConnectionAutoSync(
connections,
setConnections as React.Dispatch<React.SetStateAction<ConnectionRowConnection[]>>,
notify,
t
);
React.useEffect(() => {
latest = { handler, connections };
});
return null;
}
const el = document.createElement("div");
document.body.appendChild(el);
const root = createRoot(el);
act(() => root.render(<Wrapper />));
roots.push({ root, el });
return {
get: () => {
if (!latest) throw new Error("Hook did not render");
return latest;
},
};
}
beforeEach(() => {
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = true;
vi.stubGlobal("fetch", vi.fn());
vi.clearAllMocks();
});
afterEach(() => {
for (const { root, el } of roots.splice(0)) {
act(() => root.unmount());
el.remove();
}
vi.unstubAllGlobals();
});
describe("useConnectionAutoSync", () => {
it("PUTs the autoSync flag and notifies success", async () => {
const conns: ConnectionRowConnection[] = [
{ id: "conn-1", providerSpecificData: { autoSync: false } },
];
const h = renderHandler(conns);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValueOnce({ ok: true } as Response);
await act(async () => {
await h.get().handler("conn-1", true);
});
expect(fetchMock).toHaveBeenCalledTimes(1);
expect(fetchMock).toHaveBeenNthCalledWith(
1,
"/api/providers/conn-1",
expect.objectContaining({
method: "PUT",
body: JSON.stringify({
providerSpecificData: { autoSync: true },
}),
})
);
expect(notify.success).toHaveBeenCalled();
});
it("spreads existing providerSpecificData instead of replacing it", async () => {
const conns: ConnectionRowConnection[] = [
{ id: "conn-1", providerSpecificData: { someOtherFlag: 42, autoSync: false } },
];
const h = renderHandler(conns);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValueOnce({ ok: true } as Response);
await act(async () => {
await h.get().handler("conn-1", true);
});
const body = JSON.parse(fetchMock.mock.calls[0][1].body as string);
expect(body).toEqual({
providerSpecificData: { someOtherFlag: 42, autoSync: true },
});
expect(h.get().connections).toEqual([
{ id: "conn-1", providerSpecificData: { someOtherFlag: 42, autoSync: true } },
]);
});
it("notifies error when the PUT fails", async () => {
const conns: ConnectionRowConnection[] = [
{ id: "conn-1", providerSpecificData: { autoSync: false } },
];
const h = renderHandler(conns);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValueOnce({ ok: false, status: 500 } as Response);
await act(async () => {
await h.get().handler("conn-1", true);
});
expect(notify.error).toHaveBeenCalled();
expect(notify.success).not.toHaveBeenCalled();
});
it("notifies autoSyncDisabled (info) when disabling autoSync", async () => {
const conns: ConnectionRowConnection[] = [
{ id: "conn-1", providerSpecificData: { autoSync: true } },
];
const h = renderHandler(conns);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValueOnce({ ok: true } as Response);
await act(async () => {
await h.get().handler("conn-1", false);
});
expect(notify.info).toHaveBeenCalledWith("autoSyncDisabled");
expect(notify.success).not.toHaveBeenCalled();
});
});

View File

@@ -0,0 +1,247 @@
// @vitest-environment jsdom
import React, { act } from "react";
import { createRoot } from "react-dom/client";
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import {
useModelImportHandlers,
type UseModelImportHandlersParams,
type UseModelImportHandlersReturn,
} from "../hooks/useModelImportHandlers";
type HookResult = UseModelImportHandlersReturn;
const t = ((key: string) => key) as ((key: string) => string) & {
has: (key: string) => boolean;
};
t.has = () => false;
const notify = {
success: vi.fn(),
error: vi.fn(),
warning: vi.fn(),
info: vi.fn(),
};
function buildParams(
overrides: Partial<UseModelImportHandlersParams>
): UseModelImportHandlersParams {
return {
providerId: "cloudflare-ai",
models: [],
modelMeta: { customModels: [] },
modelAliases: {},
connections: [],
isFreeNoAuth: false,
handleSetAlias: vi.fn().mockResolvedValue(undefined),
fetchAliases: vi.fn().mockResolvedValue(undefined),
fetchProviderModelMeta: vi.fn().mockResolvedValue(undefined),
fetchConnections: vi.fn().mockResolvedValue(undefined),
notify,
t,
providerStorageAlias: "cloudflare-ai",
...overrides,
};
}
const roots: Array<{ root: ReturnType<typeof createRoot>; el: HTMLDivElement }> = [];
function renderHook(params: UseModelImportHandlersParams): { get: () => HookResult } {
let latestResult: HookResult | null = null;
function Wrapper() {
const result = useModelImportHandlers(params);
React.useEffect(() => {
latestResult = result;
});
return null;
}
const el = document.createElement("div");
document.body.appendChild(el);
const root = createRoot(el);
act(() => root.render(<Wrapper />));
roots.push({ root, el });
return {
get: () => {
if (!latestResult) throw new Error("Hook did not render");
return latestResult;
},
};
}
function conn(id: string, active: boolean, autoSync?: boolean) {
return {
id,
isActive: active,
providerSpecificData: autoSync === undefined ? {} : { autoSync },
};
}
beforeEach(() => {
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = true;
vi.stubGlobal("fetch", vi.fn());
vi.clearAllMocks();
});
afterEach(() => {
for (const { root, el } of roots.splice(0)) {
act(() => root.unmount());
el.remove();
}
vi.unstubAllGlobals();
});
describe("useModelImportHandlers — master autoSync", () => {
it("isAutoSyncEnabled is true only when every active connection has autoSync on", () => {
const mixed = renderHook(
buildParams({ connections: [conn("a", true, true), conn("b", true, false)] })
);
expect(mixed.get().isAutoSyncEnabled).toBe(false);
const allOn = renderHook(
buildParams({ connections: [conn("a", true, true), conn("b", true, true)] })
);
expect(allOn.get().isAutoSyncEnabled).toBe(true);
const oneOff = renderHook(
buildParams({ connections: [conn("a", true, true), conn("b", false, true)] })
);
expect(oneOff.get().isAutoSyncEnabled).toBe(true);
});
it("handleToggleAutoSync fans out a PUT to every active connection (bug repro)", async () => {
const fetchConnections = vi.fn().mockResolvedValue(undefined);
const hook = renderHook(
buildParams({
connections: [conn("conn-a", true, false), conn("conn-b", true, false)],
fetchConnections,
})
);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValue({ ok: true } as Response);
await act(async () => {
await hook.get().handleToggleAutoSync();
});
expect(fetchMock).toHaveBeenCalledTimes(2);
expect(fetchMock).toHaveBeenNthCalledWith(
1,
"/api/providers/conn-a",
expect.objectContaining({ method: "PUT" })
);
expect(fetchMock).toHaveBeenNthCalledWith(
2,
"/api/providers/conn-b",
expect.objectContaining({ method: "PUT" })
);
expect(fetchConnections).toHaveBeenCalled();
});
it("excludes inactive connections from the fan-out", async () => {
const hook = renderHook(
buildParams({
connections: [conn("conn-a", true, false), conn("conn-inactive", false, false)],
})
);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValue({ ok: true } as Response);
await act(async () => {
await hook.get().handleToggleAutoSync();
});
expect(fetchMock).toHaveBeenCalledTimes(1);
expect(fetchMock).toHaveBeenNthCalledWith(
1,
"/api/providers/conn-a",
expect.objectContaining({ method: "PUT" })
);
});
it("toggling from a mixed state (one on, one off) turns all active connections on", async () => {
const hook = renderHook(
buildParams({
connections: [conn("conn-a", true, true), conn("conn-b", true, false)],
})
);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValue({ ok: true } as Response);
await act(async () => {
await hook.get().handleToggleAutoSync();
});
expect(hook.get().isAutoSyncEnabled).toBe(false);
expect(fetchMock).toHaveBeenCalledTimes(2);
expect(fetchMock).toHaveBeenNthCalledWith(
1,
"/api/providers/conn-a",
expect.objectContaining({ method: "PUT" })
);
expect(fetchMock).toHaveBeenNthCalledWith(
2,
"/api/providers/conn-b",
expect.objectContaining({ method: "PUT" })
);
const firstBody = JSON.parse(fetchMock.mock.calls[0][1].body as string);
const secondBody = JSON.parse(fetchMock.mock.calls[1][1].body as string);
expect(firstBody.providerSpecificData).toEqual({ autoSync: true });
expect(secondBody.providerSpecificData).toEqual({ autoSync: true });
expect(notify.success).toHaveBeenCalled();
});
it("still calls fetchConnections when a fan-out PUT fails (partial failure)", async () => {
const fetchConnections = vi.fn().mockResolvedValue(undefined);
const hook = renderHook(
buildParams({
connections: [conn("conn-a", true, false), conn("conn-b", true, false)],
fetchConnections,
})
);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValueOnce({ ok: false, status: 500 } as Response);
fetchMock.mockResolvedValueOnce({ ok: true } as Response);
await act(async () => {
await hook.get().handleToggleAutoSync();
});
expect(fetchConnections).toHaveBeenCalled();
expect(notify.success).not.toHaveBeenCalled();
expect(notify.error).not.toHaveBeenCalled();
expect(notify.warning).toHaveBeenCalledWith("autoSyncPartialFailure");
});
it("notifies error when every fan-out PUT fails", async () => {
const hook = renderHook(
buildParams({
connections: [conn("conn-a", true, false), conn("conn-b", true, false)],
})
);
const fetchMock = vi.mocked(fetch);
fetchMock.mockResolvedValue({ ok: false, status: 500 } as Response);
await act(async () => {
await hook.get().handleToggleAutoSync();
});
expect(notify.error).toHaveBeenCalledWith("autoSyncToggleFailed");
expect(notify.success).not.toHaveBeenCalled();
expect(notify.warning).not.toHaveBeenCalled();
});
it("no-ops without a PUT or notification when there are no active connections", async () => {
const hook = renderHook(buildParams({ connections: [conn("conn-a", false, false)] }));
const fetchMock = vi.mocked(fetch);
await act(async () => {
await hook.get().handleToggleAutoSync();
});
expect(fetchMock).not.toHaveBeenCalled();
expect(notify.success).not.toHaveBeenCalled();
expect(notify.error).not.toHaveBeenCalled();
});
});

View File

@@ -15,11 +15,7 @@ import {
getCodexEffectiveServiceTier,
type CodexGlobalServiceMode,
} from "@/lib/providers/codexFastTier";
import {
normalizeCodexLimitPolicy,
providerText,
ERROR_TYPE_LABELS,
} from "../providerPageHelpers";
import { normalizeCodexLimitPolicy, providerText, ERROR_TYPE_LABELS } from "../providerPageHelpers";
import { getCodexPlanLabel } from "../codexPlanLabel";
import ProviderQuotaVisibilityToggle from "./ProviderQuotaVisibilityToggle";
@@ -69,6 +65,7 @@ export interface ConnectionRowProps {
onToggleRateLimit: (enabled?: boolean) => void;
onToggleQuotaVisibility?: (visible: boolean) => void;
onToggleClaudeExtraUsage?: (enabled?: boolean) => void;
onToggleAutoSync?: (enabled: boolean) => void;
onToggleCodex5h?: (enabled?: boolean) => void;
onToggleCodexWeekly?: (enabled?: boolean) => void;
isCcCompatible?: boolean;
@@ -354,6 +351,7 @@ export default function ConnectionRow({
onToggleRateLimit,
onToggleQuotaVisibility,
onToggleClaudeExtraUsage,
onToggleAutoSync,
onToggleCodex5h,
onToggleCodexWeekly,
onToggleCliproxyapiMode,
@@ -514,6 +512,8 @@ export default function ConnectionRow({
: false;
const codexPlanLabel = getCodexPlanLabel(!!isCodex, connection.providerSpecificData);
const cliproxyapiDeepMode = !!cliproxyapiEnabled;
const autoSyncEnabled = !!(connection.providerSpecificData as Record<string, unknown> | undefined)
?.autoSync;
return (
<div
@@ -637,6 +637,24 @@ export default function ConnectionRow({
onToggle={onToggleQuotaVisibility}
/>
)}
{onToggleAutoSync && (
<>
<span className="text-text-muted/30 select-none">|</span>
<button
onClick={() => onToggleAutoSync?.(!autoSyncEnabled)}
disabled={connection.isActive === false}
className={`inline-flex items-center gap-1 px-1.5 py-0.5 rounded text-xs font-medium transition-all cursor-pointer disabled:opacity-40 disabled:cursor-not-allowed ${
autoSyncEnabled
? "bg-emerald-500/15 text-emerald-500 hover:bg-emerald-500/25"
: "bg-black/[0.03] dark:bg-white/[0.03] text-text-muted/50 hover:text-text-muted hover:bg-black/[0.06] dark:hover:bg-white/[0.06]"
}`}
title={t("autoSyncTooltip")}
>
<span className="material-symbols-outlined text-[13px]">sync</span>
{t("autoSyncShort")}
</button>
</>
)}
{isClaude && (
<>
<span className="text-text-muted/30 select-none">|</span>

Some files were not shown because too many files have changed in this diff Show More