Compare commits

...

66 Commits

Author SHA1 Message Date
Diego Rodrigues de Sa e Souza
382449d593 maint: follow-up cherry-pick fix-in-place #9711 (conflict-resolved fallback) (#9891)
* fix(sse): grace period before finalizing a client disconnect as 499 (#9653)

A client that closes its connection right after reading a fully-completed
SSE stream can race OmniRoute's own completion bookkeeping: the bytes
already reached the client, but the transform stream's own completion
callback (onStreamComplete, which flips streamCompletionRecorded) hasn't
finished bubbling up when the disconnect handler fires, so the request gets
persisted as a false 499 with zero token usage even though it delivered its
full response.

Confirmed live on real traffic before this fix: a request whose server log
showed "disconnect: request_signal_aborted" at 18236ms was persisted with
status 200 and full token usage (82814/1292) once the grace period let the
real completion win the race, matching what the client actually received.

createClientDisconnectGraceHandler (new leaf in
streamFailureFinalization.ts) polls isStreamCompletionRecorded() for up to
STREAM_DISCONNECT_GRACE_PERIOD_MS (default 10s, env-configurable, 0
disables) before finalizing as a failure. If a real completion lands within
the window, handleStreamFailure's own guard is a no-op and the genuine 200
stands.

Covered by tests/unit/stream-disconnect-grace-period-9653.test.ts (fake-timer
driven: already-recorded completion short-circuits, disabled-grace-period
finalizes immediately, a completion landing mid-window skips finalize
entirely, and no completion ever landing finalizes once the deadline
passes).

(cherry picked from commit 5d0fe28c42)

* chore(quality): rebaseline chatCore.ts for the disconnect grace-period fix

Own growth from the disconnect grace-period fix: 5030->5039 (+9, the
createClientDisconnectGraceHandler wiring at the existing
onClientDisconnectFinalize call site).

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 10:08:25 -03:00
Diego Rodrigues de Sa e Souza
807a0d2022 maint: follow-up cherry-pick fix-in-place #9704 (conflict-resolved fallback) (#9889)
* fix(sse): persist per-tool-call JSON escape state across SSE delta chunks

escapeJsonStringValues() reset its inString/pendingEscape state on every
call instead of carrying it forward per tool-call index, so a raw newline
byte (or an already-escaped \n) split across two delta chunks got corrupted
in transit — the model's own output was correctly escaped, OmniRoute broke
it. Root-caused via a dispatched investigation into real OpenClaw traffic
that looked like model-generation quality but wasn't.

Fix: escapeJsonStringValues now takes and mutates a persistent per-call
state object (JsonStringEscapeState), keyed per tool-call index in the
translator's init state and cleared when a tool call is superseded.

* chore(quality): rebaseline openai-responses.ts for the escape-state fix

Own growth from the extracted per-tool-call JSON escape-state fix
(previous commit): open-sse/translator/response/openai-responses.ts
1204->1249 (+45).

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 10:07:36 -03:00
Diego Rodrigues de Sa e Souza
54bba33e2f maint: follow-up cherry-pick fix-in-place #9629 (conflict-resolved fallback) (#9885)
* fix(compression): add Lite tool truncation toggle

* fix(antigravity): add missing antigravityProjectPersistence.ts module

The quota-strategy engine (quotaStrategies.ts) imports from
antigravityProjectPersistence.ts, but only antigravityProjectPersist.ts
existed in the tree.  Add the missing module with the expected
preferAntigravityConnectionsWithStoredProject() helper and re-export
the existing persistDiscoveredAntigravityProjectId().

Co-authored-by: diegosouzapw <diegosouza.pw@outlook.com>

* fix(file-size): rebaseline strategySelector.ts for Lite truncation toggle

The PR adds one line to threading options?.config?.lite into
applyLiteCompression. Update the frozen size from 1060 to 1061.

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>

Refs #9629

---------

Co-authored-by: Xiangzhe <xiangzhedev@gmail.com>
Co-authored-by: xz-dev <xz-dev@users.noreply.github.com>
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-09 10:07:30 -03:00
Diego Rodrigues de Sa e Souza
a54c1f73af fix(db): resolve ccr migration version collision (#9884)
Renumber the CCR block-store migration from 134 to 139, reconcile databases that already applied the legacy slot, and add regression coverage for both upgrade paths.

Co-authored-by: fenix007 <fenix007@users.noreply.github.com>
2026-08-09 10:07:23 -03:00
Diego Rodrigues de Sa e Souza
57fb90d734 maint: follow-up cherry-pick fix-in-place #9549 (conflict-resolved fallback) (#9881)
* fix(adobe-firefly): open browser sign-in and resolve provider slug in /login

POST /api/providers/[id]/login passed the connection DB id to
inAppLoginService.startLogin, but that service looks up the provider by
slug in TOKEN_EXTRACTION_CONFIGS. The lookup always missed and returned
"No extraction config" without launching a browser — so the VibeProxy
"Sign in" button for Adobe Firefly (and every other web-cookie provider)
never opened a browser.

Adobe Firefly additionally had no extraction config because its IMS JWT
is never in cookies/localStorage — it only rides on the Authorization:
Bearer header of firefly-3p.ff.adobe.io XHRs.

- Resolve the provider slug from the connection row and pass the slug
  (not the DB id) to inAppLoginService.startLogin.
- Add open-sse/services/adobeFireflyBrowserLogin.ts: a Playwright
  service that launches a visible browser at firefly.adobe.com and
  intercepts firefly-3p requests to capture the IMS JWT + sherlockToken
  cookie. Wire it into the /login route for the adobe-firefly slug.
- Fix latent bug: updateProviderConnection reads camelCase keys
  (apiKey, providerSpecificData), so the previous snake_case call never
  persisted extracted credentials.

* fix(adobe-firefly): open browser sign-in and resolve provider slug in /login

POST /api/providers/[id]/login passed the connection DB id to
inAppLoginService.startLogin, but TOKEN_EXTRACTION_CONFIGS is keyed by
provider slug — so browser login never launched for web-cookie providers.

Adobe Firefly also cannot use cookie extraction: the IMS JWT only appears
on Authorization headers to firefly-3p.ff.adobe.io. Add a dedicated
Playwright interceptor and persist credentials with camelCase keys that
updateProviderConnection actually reads.

* fix(adobe-firefly): use system Chrome/Edge CDP for browser sign-in

Playwright is not available inside the pkg-packaged VibeProxyServices.exe,
so import('playwright') always failed with 'Playwright not installed' and
never opened a window. Launch Chrome/Edge with --remote-debugging-port and
capture the firefly-3p Authorization Bearer via pure CDP WebSocket instead.

* fix(adobe-firefly): live x-arp-session-id / Arkose wire (stop 408 under load)

Browser generate-async requires x-arp-session-id as base64({sid,ark,ftr}) with a
real Arkose blob (sherlockToken). JWT alone frequently returns colligo HTTP 408
system under load while credits still work.

- Match live ftr magic __UDF43-m4_31ck + Arkose pk in synthetic ARP fallback
- Ranked extract of sherlockToken / x-arp from Cookie, HAR, fetch() paste, and
  space-joined JWT+ARP (PasswordBox newline collapse)
- Reuse one ARP for storage upload + generate-async
- Clearer 408 errors when browser ARP is missing vs stale
- Unit suite 42/42

* fix(adobe-firefly): durable session ARP rebuild and aux_sid false-positive

Rebuild x-arp-session-id from forterToken/arkose/ff_session_guid instead of
ranking long Cookie pairs (e.g. aux_sid=…) as opaque ARP, which caused colligo
HTTP 408. Cache IMS JWT + cookie sessions, rotate ARP on 408 retries, and keep
Playwright warm-up opt-in only (headless Forter is rejected).

Also expand synthetic ARP shape with bfp/fpjs to match live successful captures.

* fix(adobe-firefly): durable session, off-screen Chrome recovery, browser sign-in

Rebuild x-arp-session-id from Cookie pieces (sid/ark/forter) so aux_sid is never
sent as ARP. Sticky ARP + submit spacing reduce mid-batch colligo 408 thrash.

Add optional managed Chrome warm (off-screen headed by default; Forter rejects
headless) and POST /api/providers/{id}/login browser sign-in that returns JWT+Cookie
after a fresh SSO. Visible sign-in resets off-screen window placement and clears
prior Adobe session when adding another account.

* fix(adobe-firefly): renew sessions through durable CDP

* fix(adobe-firefly): isolate browser sessions per account

* fix(adobe-firefly): make account login fresh and deterministic

* chore(adobe-firefly): remove obsolete browser fallback

* docs(adobe-firefly): document renewal controls

* fix(adobe-firefly): harden CDP warm, risk session, and browser sign-in

Stop colligo 408 thrash from stale Forter and frozen Google login during
Sign in with browser:

- CDP warm: clear Firefly origin storage + risk cookies (keep SSO); require
  forter age under 10 minutes on loop and timeout paths; dual CDP queues;
  await Runtime.runIfWaitingForDebugger; profile-lock launch retries
- Session: connectionId fingerprint; write-back JWT+Cookie; warm-fail
  cooldown; fail closed risk_session_stale when forter is known-stale
- Client: submit gate around generate-async; max 2 attempts when forter
  known-stale; poll 401 one refresh; pass sessionBrowserKey through handlers
- Login route: pure system Chrome/Edge CDP only; camelCase credential persist
- Unit: browser-login + firefly suites green (60)

---------

Co-authored-by: artickc <artur1992123@mail.ru>
2026-08-09 10:07:17 -03:00
Diego Rodrigues de Sa e Souza
5eba045175 maint: follow-up cherry-pick fix-in-place #9510 (fallback resolution) (#9880)
* feat(api): add GET /api/resilience/connections for per-account state

The three temporary-failure mechanisms each have their own scope -- the
provider circuit breaker covers a whole provider, connection cooldown covers
one account, model lockout covers a provider/connection/model triple -- and
until now nothing showed them side by side. Diagnosing "why is this key being
skipped" meant reading three separate surfaces and correlating by hand, which
is exactly what the docs' own debugging guidance asks an operator to do.

The route returns all three keyed by connection, plus the breaker's transition
history so a flapping provider is visible as a sequence rather than a single
current state. getStatus() already assembled everything except that history;
it now returns a copy of it and carries an explicit CircuitBreakerStatus type
instead of an inferred one.

Reading raw connection rows for this meant widening getRawProviderConnections'
column projection, so the existing allowlist is exported and the route selects
through it. A test asserts every column the route names is in that allowlist,
which turns a future typo into a failure here rather than a silent empty field.

Each of the three data sources is wrapped independently: one of them throwing
degrades that section and sets meta.degraded rather than failing the whole
response, since a partial view still answers most of the questions the page
exists for.

Loopback-gated. It spawns nothing, unlike every other entry on that list, but
it exposes per-account operational state and the comment says so to keep it
from being read as precedent for gating read-only routes generally.

Tests are real isolated-DB integration tests rather than mocks -- ESM mocking
is unavailable here (no mock.module, non-configurable exports) and the
codebase already has the isolated-DB pattern, which exercises more than a mock
would anyway.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* feat(dashboard): add the per-account resilience connections page

Renders what the API added: every connection with its cooldown, its provider
breaker, and its model lockouts in one table, with a detail view per connection
and the breaker's transitions drawn as a timeline. The timeline is the part that
is hard to get from the existing surfaces -- a breaker sitting at CLOSED right
now looks healthy, and only the sequence shows it has opened four times in the
last hour.

Polls rather than streams. The state it displays changes on the order of
seconds to minutes and the page is loopback-gated, so an SSE channel would buy
nothing over an interval.

ModelCooldownsCard had its own formatRemaining. The new table needs the same
countdown format and two copies would drift, so it moves to
shared/utils/formatRemaining.ts and both import it -- behaviour unchanged, the
extracted version differs from the deleted one only in local variable names.
DataTable's column and row interfaces are exported for the same reason: the new
table types against them rather than restating their shape.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* fix(i18n): translate new resilience-connections screen strings

PR #9510 added the "Connection Resilience" dashboard screen but the
sync-added i18n keys (sidebar.resilienceConnections/Subtitle and the
full resilienceConnections namespace) were left as __MISSING__: in
every non-English locale, dropping i18nUiCoverage.pct below the 99
ratchet baseline.

Translate all ~78 new leaf strings into all 41 non-English locales.
Pre-existing unrelated __MISSING__ debt (hermesRole*, apiProtocol*,
grokAutoTopUp*, featureFlagExposeFunctionalGatewayMirrorsDescription)
is left untouched — out of scope for this fix.

Co-authored-by: HouMinXi <HouMinXi@users.noreply.github.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: HouMinXi <HouMinXi@users.noreply.github.com>
2026-08-09 10:07:10 -03:00
Diego Rodrigues de Sa e Souza
06727f0e74 cherry-pick(pr-9556): fix(translator): preserve Kimi K3 Responses reasoning (#9879)
* fix(translator): preserve Kimi K3 Responses reasoning

* fix(translator): make K3 reasoning preservation model-driven

* fix(translator): replay cached Kimi reasoning before fallback

* fix(translator): keep authentic K3 reasoning through cleanup

* refactor(reasoning): use replay policy for K3

---------

Co-authored-by: jackjinke <jack.kejin@gmail.com>
2026-08-09 10:07:03 -03:00
Diego Rodrigues de Sa e Souza
5f75abe4a2 cherry-pick(pr-9634): fix(test): reconcile base-drifted test expectations on release/v3.8.50 (#9874)
* fix(combo): restore routing module load

* fix(db): resolve ccr migration version collision

Renumber the CCR block-store migration from 134 to 139, reconcile databases that already applied the legacy slot, and add regression coverage for both upgrade paths.

Co-Authored-By: GPT-5 <noreply@openai.com>

* fix(changelog): format the aggregator balance fragment as a bullet

The fragment landed with YAML frontmatter rather than the bullet the
aggregator reads, so check:changelog-integrity exits 1 on every branch and
takes the merge-integrity job down with it regardless of what the branch
changed.

Only the format changes. The entry text is the author's, unedited, and now
carries the link to the pull request that shipped it.

* fix(test): update expected auth/vision/provider schema for base-drifted expectations

* fix(test): narrow this branch to the drifted test expectations

Three other PRs already cover what this one was carrying. #9618 renumbers the
colliding ccr_blocks migration, #9632 repairs the malformed aggregator changelog
fragment, and #9676 restores the combo module load by implementing the selection
helper the import was reaching for, rather than deleting the caller the way this
branch did. Keeping any of it here would put two files back on the same migration
slot and overwrite a better fix with a worse one.

What survives is the part none of them touch. Once the combo barrel loads again,
three assertions in the context-window filter suite start failing: they demand
that catalog-too-small targets be dropped, while the file's own header and its
four neighbouring tests say those targets stay available as runtime fallback.
The unresolved import was masking them. A new case pins the output-token limit
as a genuine hard requirement so the relaxation cannot drift further.

The provider count assertion kept one literal at the old value after the rest of
the file moved to 198, so the partition check failed on a sum that was correct.

* chore(quality): re-time migrationRunner for the 139 guard on the new tip

---------

Co-authored-by: alexey.nazarov@softmg.ru <alexey.nazarov@softmg.ru>
Co-authored-by: GPT-5 <noreply@openai.com>
Co-authored-by: Minxi Hou <houminxi@gmail.com>
2026-08-09 10:06:57 -03:00
Diego Rodrigues de Sa e Souza
58f0ff1b41 cherry-pick(pr-9675): fix(providers): per-provider opt-out for anonymous no-auth fallback (opencode-go/zen 401s) (#9873)
* fix(providers): add per-provider opt-out for anonymous no-auth fallback

API-key providers with anonymousFallback: true (opencode-go, opencode-zen,
pollinations, kilocode) receive a synthetic "noauth" connection whenever all
real connections are terminal (credits_exhausted/banned/expired) or
unavailable. The opencode upstream now rejects anonymous requests with
401 Missing API key, so the fallback adds a guaranteed-failing round trip
and health/reconnect noise before the combo moves on.

Add a noAuthFallbackDisabledProviders settings array (zod-validated,
persisted via /api/settings, following the blockedProviders pattern).
When a provider is listed, maybeSyntheticNoAuthFallback returns null for
anonymousFallback-only providers, so exhausted providers are skipped
immediately as allExpired/allRateLimited while real keyed connections keep
working and recover automatically once quota state clears. True no-auth
providers are unaffected; blockedProviders remains their disable mechanism.
Default (absent/empty list) preserves current behavior.

Provider detail pages for anonymousFallback providers gain an
"Anonymous fallback" toggle (default ON) backed by the new setting.

Refs #9674

* fix(auth): reduce file size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Hermes Agent <hermes@hermes-chloe.hyades.io>
2026-08-09 10:06:50 -03:00
Diego Rodrigues de Sa e Souza
05ab06f3a1 fix: address self-review findings (#9900)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
2026-08-09 09:55:22 -03:00
Diego Rodrigues de Sa e Souza
431fc02e75 cherry-pick(pr-9569): fix(settings): use provider prefixes in model overrides (#9878)
* fix(settings): use provider prefixes in model overrides

* refactor(settings): extract pricing tab helpers

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Xiangzhe <xiangzhedev@gmail.com>
2026-08-09 09:55:16 -03:00
Diego Rodrigues de Sa e Souza
c1d951e5c5 cherry-pick(pr-9572): fix(providers): reject the dashboard password as a connection API key (#9877)
* fix(providers): refuse to store the dashboard password as a connection API key

A browser autofilled the management password into a connection's API-key field.
The resulting credential authenticates against nothing, so every request routed
through that connection came back 401, and because the field looks like any
other password input the same autofill fired again while the connection was
being repaired by hand.

The refusal belongs on the write path rather than in the form. Twenty routes
create or update connections and all of them funnel through
createProviderConnection and updateProviderConnection, so one check there covers
every entry point including a future one. The two other places that write
api_key are left alone on purpose: one re-encrypts rows that already exist and
the other is the one-time db.json import, and neither takes a value an operator
just typed.

Update checks the incoming value, never the merged one. A connection that
already holds the password has to stay editable or the operator cannot repair
the exact state this prevents, and re-checking the merged value would spend a
bcrypt round on every unrelated field edit.

Only a real match blocks the write. An unreadable settings row or a throwing
bcrypt call logs and allows, because a guard against one specific mistake must
not turn into a way to lock out every connection write.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* fix(providers): compare the untrimmed credential, and cover the guard's branches

The guard trimmed the incoming value before comparing it, which catches a paste
carrying whitespace the password does not have. It missed the mirror case:
neither the login route nor the set-password route trims, so a dashboard
password may itself begin or end with a space, and an autofill reproducing it
exactly was trimmed into a value that no longer matched the stored hash. The
write then went through, which is the state this guard exists to prevent. Both
forms are compared now, the second only when the first fails on a string that
differs, so an ordinary key still costs a single bcrypt round.

Two branches carried no coverage and both are load-bearing. The catch that logs
and allows is the only path that lets a write through; a stored hash bcrypt
cannot parse reaches it without needing a mock, since the shape check accepts an
impossible cost factor that the comparison then rejects. The early return is
what keeps a token renewal -- a write carrying tokens but no apiKey -- from
paying for a settings read and a bcrypt round every time it fires, and the same
unparseable hash makes that path observable, so an absent warning is proof the
return happened.

The narrower scope is deliberate and now says so in the code: the OAuth tokens
arrive from a provider's token endpoint rather than from a form, so extending
the comparison to them would charge every renewal for a field no autofill can
reach.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: Minxi Hou <houminxi@gmail.com>
2026-08-09 09:55:09 -03:00
Diego Rodrigues de Sa e Souza
740e16c3e2 cherry-pick(pr-9601): feat(responses): add encrypted reasoning replay opt-in (#9876)
* feat(codex): add encrypted reasoning replay opt-in

* feat(responses): generalize encrypted reasoning replay

* docs: clarify encrypted reasoning provider scope

* fix(ui): group reasoning replay with connection controls

* fix(logs): omit encrypted reasoning payloads

* fix(chat): reduce file size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(chat): reduce combined file size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: jackjinke <jack.kejin@gmail.com>
2026-08-09 09:55:02 -03:00
Diego Rodrigues de Sa e Souza
d8967efc6c cherry-pick(pr-9605): ci(test): route orphaned Vitest tests through blocking CI (#9875)
* ci(test): route orphaned Vitest tests through blocking CI

* docs: fix advisory status in AGENTS.md and refresh baseline note

* fix(changelog): fix fragment format for #9415

* fix(changelog): preserve upstream fragment format

---------

Co-authored-by: MohitRawat017 <rawatmohit17906@gmail.com>
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-09 09:54:56 -03:00
Diego Rodrigues de Sa e Souza
3a66761cb7 fix: pass max reasoning effort through by default, add global model registry fallback (#8057) (#9883)
Co-authored-by: Mo'men Qatr <momen.qatr04@eng-st.cu.edu.eg>
2026-08-09 09:54:52 -03:00
Diego Rodrigues de Sa e Souza
5e5919dcc0 maint: follow-up cherry-pick fix-in-place #9631 (conflict-resolved fallback) (#9886)
* feat(db): add a job registry for scheduled background work

Background jobs each ship their own timer today, so there is no list of what
is scheduled, no history of what ran, and no way to pause one without an
environment variable and a restart. The registry gives them one home: a jobs
table holding the schedule, a job_runs table holding the outcomes, and a
loopback-only API to inspect and control both.

Cron jobs read their expression through an optional cronGetter rather than the
stored column, so an operator changing OMNIROUTE_WARMUP_CRON does not need the
row rewritten. register() is an idempotent upsert that refreshes the schedule
but never overwrites `enabled` or `created_at`, which is what lets a job be
re-registered on every boot without discarding the operator's toggle.

Run history is pruned per job rather than globally, and safeRun records a
failure for a handler that throws as well as one that returns success:false,
so a crashing job leaves a trail instead of a gap.

The API is under /api/jobs and gated to loopback in the route guard. It can
trigger a run and flip a job off, which is runtime administration and does not
belong on a remotely reachable surface.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* feat(jobs): move the budget reset and token health check onto the registry

Both jobs owned their own timer and started themselves as an import side effect,
so nothing could report whether they were running, when they last ran, or why a
run failed. They now register with the job registry and are started from it, which
also means their schedule and run history are visible through /api/jobs.

startAll() runs each interval job's first tick synchronously, so both entry points
start the registry only after initializeCloudSync() has been awaited. The old
wiring reached that ordering two different ways: the budget reset was started
after the init call, and the health check's first sweep sat behind a 10s timer.
Replacing both with one startAll() would otherwise have moved the two handlers
in front of the initialisation they run against.

Both entry points also register the same pair of jobs. Registering one and not
the other is how a background job goes missing without anything failing.

sweep() now returns how many connections it swept, so the health check can record
a real records_affected the way the budget reset does. The migration documents
that column as a per-job count, and hardcoding zero would have left one of the two
jobs reporting a number the schema promises but the code never produces. A skipped
or empty sweep reports zero. Every existing caller ignores the return value.

The token health check keeps its own disable semantics: the handler still calls
isHealthCheckDisabled() before sweeping, so OMNIROUTE_DISABLE_TOKEN_HEALTHCHECK,
the production-build phase and the automated-test guard behave as before. Its
registry adapter lives in src/lib/jobs/ next to the budget reset rather than in
tokenHealthCheck.ts, which is already above its frozen size ceiling on the base
branch and should not grow further. The adapter lets a failing sweep throw rather
than reporting it itself, matching the budget reset: safeRun records a thrown
error as a failure run with its message.

The warmup job is seeded disabled. Its handler arrives with the warmup scheduler,
and startAll() filters on enabled before it looks for a handler, so seeding it
enabled here would warn about the missing handler on every boot.

* fix: allowlist cron-parser dep and document OMNIROUTE_RUNNOW_TIMEOUT_MS env var

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-09 09:54:47 -03:00
Diego Rodrigues de Sa e Souza
f68856695d maint: follow-up cherry-pick fix-in-place #9693 (conflict-resolved fallback) (#9887)
* fix(web-tools): anchor tool contract at prompt tail + user-turn reminder

The <tool> contract from prepareToolMessages was prepended as the first
system message. Web executors fold all system messages into one block, so
with agentic clients whose system prompts exceed ~28K chars the contract
sat at the head of a huge block and web models ignored it, refusing tool
calls with "tool X is not in my tool set" (chatgpt-web, 0/3 at 30K chars).

Two changes, both required in testing:

- Dual placement: the full contract now rides as a trailing system
  message (folds to the tail of the system block) and a one-line
  reminder naming the tools is appended to the latest user message.
- Rewording: the contract now frames injected tools as client tools
  invoked via a plain-text protocol, distinct from the model's native
  tool registry (web.run, python.exec, ...), and instructs the model to
  never claim they are unavailable. Without this the model resolved
  tool names against its native registry and refused even when it had
  seen the contract.

Measured on cgpt-web gpt-5.5-thinking/gpt-5.6-thinking/o3: prepend 0/3
tool calls at 30K chars; dual placement 16/17 across 30K-250K system
prompts, 30-tool sets, multi-turn tool history, streaming, and 3-way
concurrency, with no spurious calls on no-tool prompts. Known limit:
~40K-char single user messages still flake (2/3) due to the upstream
model's own injection heuristics.

All prepareToolMessages consumers parse system messages
position-independently and select the current user turn by role scan,
so the trailing system message is shape-safe for every web executor.

* test(web-tools): cover contract placement edge cases

---------

Co-authored-by: Ryan Brosas <ryanjoserbrosas@gmail.com>
2026-08-09 09:54:40 -03:00
Diego Rodrigues de Sa e Souza
be9f43e1ee cherry-pick(pr-9695): fix(docker): make the webpack build-arg escape hatch actually work (#9872)
* build(docker): make the bundler build-arg actually take effect

A bare ENV shadows a same-named ARG for the rest of the stage, so
--build-arg OMNIROUTE_USE_TURBOPACK=0 was silently ignored and the
webpack escape hatch the surrounding comment advertises only ever
worked through -e at runtime, never at build time.

That mattered because Turbopack compiles in native Rust memory living
outside the V8 heap, so OMNIROUTE_BUILD_MEMORY_MB cannot bound it. A
build host with a memory ceiling gets SIGKILLed by the cgroup OOM
killer with no error text at all, which reads like a hung build rather
than an out-of-memory one.

* docs(docker): correct the builder stage facts and document its cost

The stage table described a builder that no longer exists: it named
node:24.15.0-trixie-slim where every stage now derives from
node:26-trixie-slim, and said the stage runs `npm run build -- --webpack`
where it runs plain `npm run build`, which is Turbopack by default.

That second one is worse than stale. A reader who needs the webpack
fallback would conclude the Docker build already uses it and never look
for the switch.

Adds a Build-time resources section covering the two build args, why the
V8 heap arg cannot bound Turbopack, and measured ceilings for both
bundlers. The runtime paragraphs that followed get their own heading so
they no longer read as part of the build-time story.

* docs(docker): correct the runtime heap defaults

Same drift as the builder stage, in the paragraphs just below it. The
image exports OMNIROUTE_MEMORY_MB=1024 and derives NODE_OPTIONS from it,
but the guide reported 512 in three places, including the environment
variable table.

The "if unset, the launcher uses 512" line was misleading in both
readings: the image always sets the variable so that branch cannot fire
under Docker, and outside Docker the launcher calibrates from host RAM
rather than using a flat 512.

* docs(changelog): add fragment for #9695

---------

Co-authored-by: Minxi Hou <houminxi@gmail.com>
2026-08-09 09:54:34 -03:00
Diego Rodrigues de Sa e Souza
a39f78c6f8 maint: follow-up cherry-pick fix-in-place #9707 (conflict-resolved fallback) (#9890)
* fix(db): renumber ccr_blocks migration 134 -> 139

134 was taken by 134_proxy_logs_egress_ip, so two migrations shared the
same numeric prefix and check-migration-numbering failed. Move ccr_blocks
to the next free slot and add the retroactive isSchemaAlreadyApplied guard
so a DB that already applied it under 134 skips the re-run.

* fix(combo): restore missing preferAntigravityConnectionsWithStoredProject

quotaStrategies imported the reset-aware pool filter from
../antigravityProjectPersistence.ts, a module that does not exist — the
helper belongs in antigravityProjectPersist.ts and was never added there,
breaking typecheck. Add the helper alongside the persist path, point the
import at the real module, and cover the filter with unit tests.

* chore: add Makefile wrapping the canonical npm scripts

* fix(compression): remove duplicate Antigravity project helper

The release branch already includes the generic project-aware connection
selection helper. Keep that implementation and remove the duplicate introduced
while cherry-picking #9707.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Matias Baglieri <168452313+matiasbaglieri@users.noreply.github.com>
2026-08-09 09:54:30 -03:00
Diego Rodrigues de Sa e Souza
bc38965088 maint: follow-up cherry-pick fix-in-place #9712 (conflict-resolved fallback) (#9892)
* fix(build): colocateLlmlinguaOptionals skip-check treated a Next-traced stub as fully copied

Debugging the omniroute-beta Docker rebuild: `npm run build` (and the
Dockerfile's own post-build verification) failed with
`Cannot find module '.../node_modules/@atjsh/llmlingua-2/dist/index.js'`.

Root cause, reproduced directly (both against a live Docker builder image
and in a unit test): Next.js's own standalone trace creates a stub
directory for `@atjsh/llmlingua-2` containing only `package.json` — it
references the package (a dynamically-imported optional dependency) but
can't fully bundle it. colocateLlmlinguaOptionals's skip checks (both the
closure-level early return and the per-package loop) only tested
`existsSync(dest)`, so that stub was indistinguishable from "already fully
co-located" — the function skipped copying the real `dist/` output
entirely, silently shipping a package with a manifest but no code.

Fix: check for the package's declared `main` entry file when it has one
(the real-world case for every actual SLM optional). Packages with no
`main` field fall back to comparing the destination's top-level entries
against the source's — correct both for genuinely multi-file packages and
for a metadata-only source (package.json is then its complete, faithfully-
copied contents), which the existing idempotency test exercises.

Covered by tests/unit/colocate-optionals.test.ts's new stub-reproduction
case (fails against the pre-fix code, passes after — confirmed directly)
plus the 6 pre-existing cases, all still green.

(cherry picked from commit 359aba59c7)

* fix(build): register onnxruntime-node's native bin/ as a standalone asset (#9687)

Docker/standalone builds of the LLMLingua SLM compression tier failed at
runtime with "Error: libonnxruntime.so.1: cannot open shared object file:
No such file or directory" (open-sse/services/compression/engines/llmlingua's
worker, via @huggingface/transformers -> onnxruntime-node).

onnxruntime-node's dist/binding.js is a normal JS file Next.js's standalone
trace bundles correctly, but binding.js dlopen()s a platform-specific native
library shipped under bin/napi-v3/<platform>/<arch>/libonnxruntime.so.1 — a
dynamic native load static file tracing can't see (same blind-spot class as
the separate colocateLlmlinguaOptionals stub bug, just for a .so instead of
a JS import, via NATIVE_ASSET_ENTRIES instead). That directory was simply
never registered, unlike better-sqlite3's native binary, which already goes
through the exact same mechanism correctly.

Fix: add an entry for onnxruntime-node/bin, mirroring the existing
better-sqlite3 entry. Confirmed against a real Docker build of the
Dockerfile's own post-build verification step: this was the very next
failure once the separate llmlingua-2 stub bug was fixed and the build
progressed far enough to reach it.

Covered by tests/unit/assemble-standalone-onnxruntime-native-asset.test.ts
(fails against the pre-fix code on both assertions, passes after).

(cherry picked from commit 8c98a59f26)

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 09:54:24 -03:00
Diego Rodrigues de Sa e Souza
bd33b4589a feat(resilience): expose providerQuotaOverrides via /api/resilience (#9871)
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
2026-08-09 09:54:17 -03:00
Diego Rodrigues de Sa e Souza
c8e6b07df5 cherry-pick(pr-9718): feat(src): proxy-pool-toolbar-minor-improvements (#9870)
* feat(proxy-pool): streamline pool actions

* test(proxy-pool): cover toolbar layout

* refactor(settings): extract proxy registry helpers

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* refactor(settings): reduce proxy registry component size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Agnes <linkscrazy2@gmail.com>
2026-08-09 09:54:12 -03:00
Diego Rodrigues de Sa e Souza
ed7a68e1a9 maint: follow-up cherry-pick fix-in-place #9719 (conflict-resolved fallback) (#9893)
* fix(db): clear combo pins when connections are deleted

* docs: add changelog entry for #9719

---------

Co-authored-by: Zartharas <1402357+Zartharas@users.noreply.github.com>
2026-08-09 09:54:07 -03:00
Diego Rodrigues de Sa e Souza
6f3738b009 feat(oauth): add Openference OAuth and API key provider integration (#9869)
Wire Openference as a first-party OAuth gateway (PKCE, rotating refresh)
and an API-key catalog entry on api.openference.com, with live model
discovery, connection testing, free-tier badges, and regression tests.

Co-authored-by: Anh Tran <anhlead@outlook.com>
2026-08-09 09:54:01 -03:00
Diego Rodrigues de Sa e Souza
b254890c07 fix(combo): remove stray brace from #9630 error handling (#9894)
Co-authored-by: Zartharas <1402357+Zartharas@users.noreply.github.com>
2026-08-09 09:53:56 -03:00
Diego Rodrigues de Sa e Souza
09520785f8 fix(dashboard): unregister leftover service workers in dev mode (#9868)
A phone that previously loaded a production build on this origin (or
an old dev build from before the registration was gated) kept an
active service worker across dev restarts. It intercepted every
navigation/asset fetch, occasionally serving a JS chunk that didn't
match the running dev server, which tripped Next's dev-client
chunk-mismatch auto-reload — visible as an unexplained, unstoppable
refresh loop on that device only (confirmed via a clean private tab
on the same phone/URL not looping).

PwaRegister now actively unregisters any existing service worker
registrations and clears their caches outside production, instead of
just skipping a new registration.

(cherry picked from commit 66a2515cbc)

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 09:53:50 -03:00
Diego Rodrigues de Sa e Souza
04b4690f84 cherry-pick(pr-9730): fix(compression): persist RTK renderer configuration (#9867)
* fix(compression): persist RTK renderer configuration

* docs(changelog): add fragment for #9730

Adds the changelog.d/fixes/9730-persist-rtk-renderers.md fragment
required by check:changelog-integrity for the RTK enableRenderers
persistence fix in PR #9730.

---------

Co-authored-by: Isaac <isaaclyons98@gmail.com>
2026-08-09 09:53:45 -03:00
Diego Rodrigues de Sa e Souza
356fd5d606 fix(executors): repair DuckDuckGo AI Chat challenge solver (418 ERR_CHALLENGE) (#9866)
Every duckduckgo-web chat request failed with HTTP 418 ERR_CHALLENGE while
duck.ai worked normally in a browser from the same IP. Ground truth was
established by driving a real headful Chromium at duck.ai from that IP (it
returned 200), so the environment was never the problem — the anti-abuse
challenge solver was. Six independent defects were found; the first alone
disabled the solver completely.

1. Module syntax inside the vm sandbox source.
   CHALLENGE_STUBS is executed with vm.runInContext, which compiles in SCRIPT
   mode. A refactor mass-added `export` to the five `function` declarations
   inside that template literal (they read as ordinary top-level TS functions),
   so every solve threw SyntaxError. The executor swallows solve failures and
   posts the raw unsolved challenge, which upstream answers with 418.

2. Double-escaped regex in a String.raw template.
   `\\s` in __parseCssDisplay reached the sandbox as a literal backslash, so the
   display regex never matched and a getComputedStyle probe silently read empty.

3. buildHtmlLookup undercounted descendants by one.
   `count` backs el.querySelectorAll('*').length; that returns DESCENDANTS and
   countHtmlElements already skips the #document-fragment root, so the `- 1` was
   wrong. Chromium reports 3 for '<li><div></li><li></div'; we reported 2, and a
   variant multiplies innerHTML.length by that count.

4. Browser-fidelity probes.
   Newer challenge variants assert JS/DOM invariants a flat stub cannot satisfy:
   real prototype chains (HTMLDivElement -> HTMLElement -> Element), NodeList
   identity, a live body.children HTMLCollection, native-code toString, and
   sloppy-mode `this === window`. Nine of thirteen failed. Notably Math must NOT
   be sealed — Chromium reports Object.isSealed(Math) === false, and sealing it
   made our vector differ by one.

5. The solved payload dropped meta.origin / meta.stack / meta.duration.
   The duck.ai bundle always sends all three; captured browser requests confirm
   it. Without them upstream returns 418 even when every client_hash is correct.

6. reasoningEffort is now mandatory on duckchat/v1/chat.
   An otherwise byte-identical payload returns 200 with the field and 400
   ERR_BAD_REQUEST without it (A/B verified live, repeated).

Also removes the throwaway "seed" chat POST that ran before every real request.
It existed to coax a usable challenge out of the upstream while the solver was
broken; it only doubled chat calls against an IP-rate-limited endpoint, showing
up as spurious 429 ERR_RATE_LIMIT.

Verification: the solver now reproduces real Chromium's probe vectors exactly
for all 8 captured challenge variants, and the executor returns 200 end-to-end
live (non-streaming, streaming, claude-haiku-4-5, and a math prompt returning
"42").

Tests: tests/unit/duckduckgo-challenge-solver-regression.test.ts (32 tests) and
tests/unit/duckduckgo-reasoning-effort-required.test.ts (5 tests), backed by
tests/fixtures/duckduckgo/challenge-variants.json — real captured challenge
programs plus the probe vectors a real browser produced for them, so the suite
asserts against recorded browser behaviour rather than our own output. Each fix
was confirmed to fail its test when individually reverted.

Co-authored-by: Mynacol <git@mynacol.xyz>
2026-08-09 09:53:39 -03:00
Diego Rodrigues de Sa e Souza
61cb52399e fix(logging): use configurable max-depth when bounding logged tool_calls (#9865)
requestLogger.ts's cloneBoundedForLog had its own hardcoded depth cap of 6,
independent of the existing configurable getChatLogMaxDepth(). A typical
Chat Completions response body's responseBody.choices[0].message.tool_calls[0].function
sits at exactly depth 6, so every logged tool call's function field
(name+arguments) was silently replaced with the literal string "[MaxDepth]"
before ever being stored — corrupting the data, not just how it renders.
Bumped the shared default 6->20 and switched requestLogger.ts to read it
instead of using its own literal.

(cherry picked from commit a2df6cf289)

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 09:53:32 -03:00
Diego Rodrigues de Sa e Souza
9fb7d6a493 cherry-pick(pr-9735): feat(logging): bump CHAT_LOG_ARRAY_TAIL_ITEMS default 24 -> 128 (#9864)
* feat(logging): bump CHAT_LOG_ARRAY_TAIL_ITEMS default 24 -> 128

Real agentic CLIs with many MCP servers routinely declare 40-50+ tools in
a single request — a live OpenClaw session logged 47. The tail-24 default
silently dropped the array's earlier entries behind an
_omniroute_truncated_array marker, so investigating why a specific tool
call (apply_patch) behaved oddly turned up nothing: its declared shape
(function vs custom type) was unrecoverable from the call log across 40
recent requests, even though the calls themselves succeeded.

Bumped the configurable default to comfortably cover real large tool
lists with headroom. Updated .env.example and docs/reference/
ENVIRONMENT.md to match (env-doc-sync check passes).

* test(logging): pin CHAT_LOG_ARRAY_TAIL_ITEMS default at 128

The bump commit had no dedicated test asserting the literal default
value; the existing chatcore-log-truncation.test.ts derives its
expectations from getChatLogArrayTailItems() itself, so it can't
discriminate a regression back toward the old, too-small 24 default.

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 09:53:26 -03:00
Diego Rodrigues de Sa e Souza
e117249baa cherry-pick(pr-9738): feat(logging): make the chat-log truncation limit configurable, bumped default 128x (#9863)
* feat(logging): make the chat-log truncation limit configurable, bumped default 128x

The 8KB cap on logged request/response bodies
(open-sse/handlers/chatCore/logTruncation.ts::truncateForLog()) was
hardcoded — trivially exceeded by any real multi-turn agentic
conversation, meaning the dashboard's "Full Conversation" panel could
only ever show a placeholder instead of the actual messages for nearly
every logged row of any conversation with real substance.

- Added CHAT_LOG_MAX_BODY_KB env var (src/lib/logEnv.ts::
  getChatLogMaxBodyBytes()), default 1024 KB (1MB) — a 128x bump from
  the old hardcoded 8KB — following the same configurable-limit pattern
  as the sibling CHAT_LOG_TEXT_LIMIT/CHAT_LOG_ARRAY_TAIL_ITEMS/etc. vars.
- Documented in .env.example and docs/reference/ENVIRONMENT.md.

estimateSizeFast() (open-sse/utils/estimateSize.ts) has been
substantially rewritten upstream since this bug was first found (now an
iterative Frame-based walker with a separate node-visit budget, not the
simple stack loop originally patched) — re-implemented the fix against
the current algorithm rather than porting the old diff: the byte
early-exit was unconditionally the module-level ESTIMATE_SIZE_BYTE_LIMIT
(256 KiB) with no way for a caller to raise it, so any caller comparing
against a bigger configured threshold could never see a size above
~256 KiB — every payload between 256 KiB and the caller's real limit
looked "under threshold" and truncation never fired, the opposite of
intended. Added an optional byteLimit parameter (default unchanged at
ESTIMATE_SIZE_BYTE_LIMIT, so isSmallEnoughForSemanticCache's existing
behavior is untouched) threaded through both the byte-check early-exit
and the node-budget-exhaustion fail-closed fallback, with
truncateForLog() now passing its own configured getChatLogMaxBodyBytes()
value through.

* feat(dashboard): show conversation session tag in request detail metadata

Adds a "Conversation" field to the request detail panel's metadata
grid (after "Combo"), showing the request's conversation id
(sessionTag) for quick reference/copy.

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 09:53:21 -03:00
Diego Rodrigues de Sa e Souza
a524fdeaf0 maint: follow-up cherry-pick fix-in-place #9741 (conflict-resolved fallback) (#9895)
* fix(responses-api): sync reasoning-cache write index with the fixed read side

The turn-index-hardcoding fix updated the reasoning-cache read side
(translator/index.ts's main replay loop) to key lookups by the assistant
message's real position in the messages array, but two other spots still
used the old hardcoded convention:

- chatCore.ts's write side (both the streaming and non-streaming
  completion paths) still cached every response under a hardcoded
  messageIndex: 0.
- translator/index.ts's own plain-turn (non-tool-call) cache-key lookup
  ALSO still hardcoded messageIndex 0 at its call site — a second,
  previously undiscovered instance of the same class of bug, found while
  re-verifying this fix against the current upstream tip (the original
  fix only addressed the write side).

Past the first assistant turn these conventions no longer matched, so
DeepSeek/Xiaomi-mimo plain-turn reasoning replay silently missed the
cache and fell back to the placeholder (or, once #9573 removed the
placeholder fallback, to an absent field) in ordinary multi-turn
conversations.

Compute the write-side index from the incoming request's message count
instead, and use the real loop-provided messageIndex on the read-side
lookup, both matching the position the response occupies once the
client appends it to history for the next turn.

Note: this was originally part of a larger squashed fix (output_index
collision prevention across reasoning/message/tool_call items,
reasoning-content-alias generalization) that has since been superseded
by upstream's own independent fix — translator/response/openai-responses.ts
now has its own dense-output-index-sort + getReadableReasoningValue
implementation (own comment: "mirrors upstream PR #721"). Only this
narrower, still-genuinely-broken write/read index sync survives as a
distinct bug.

Test plan:
- TDD: tests/unit/reasoning-cache.test.ts's new end-to-end
  "write side (chatCore's messageIndex) and read side (translateRequest)
  agree on the same key end-to-end" test, plus the pre-existing
  "should inject placeholder for a plain (non-tool-call) DeepSeek turn"
  and "should replay cached reasoning for a plain (non-tool-call)
  DeepSeek turn when available" tests — confirmed failing against the
  pre-fix code on a clean release/v3.8.50 checkout (both the
  hardcoded-0 write side AND the hardcoded-0 read-side lookup
  independently reproduce the mismatch), passing after both fixes
- npm run typecheck:core — clean
- npm run lint — clean
- npm run check:file-size — clean (chatCore.ts rebaselined 5034->5042
  for the messageIndex computation at both call sites;
  reasoning-cache.test.ts frozen at 1035, matching the original fix's
  own rebaseline)
- 2 pre-existing, unrelated test failures in the same file
  ("should replace empty-string reasoning_content with
  NON_ANTHROPIC_THINKING_PLACEHOLDER on cache miss",
  "should inject placeholder for a plain (non-tool-call) DeepSeek turn
  missing reasoning_content") confirmed present on a completely clean,
  untouched release/v3.8.50 checkout — these test obsolete
  placeholder-injection behavior the code deliberately removed per
  #9573 (see the code's own comment); not touched by this PR

* fix(chat): reduce file size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(chat): reconcile file-size baseline

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 09:53:15 -03:00
Diego Rodrigues de Sa e Souza
a448b146bf cherry-pick(pr-9744): test(integration): add general live-test tool for the real "default" combo + rootless wire capture (#9862)
* test(integration): add general live-test tool for the real "default" combo

Temporary WIP commit on this deferred branch — lands in its own separate
PR once the bug-fix extraction batch is done (never bundled into a
bug-fix PR). Unlike liveGeminiShared.ts (provisions its own narrow
2-model Gemini-only combo), this reads the REAL "default" combo
currently configured on the target instance directly from the DB and
exercises every provider/model step in it directly, bypassing combo
routing, so live-test coverage always matches whatever is actually
configured instead of a hardcoded snapshot.

Live-verified against omniroute-beta (seeded with the real 18-model,
5-provider default combo): 14/18 models pass consistently across
non-streaming + streaming Chat Completions and streaming Responses API.
The 4 consistent failures are real external state (cerebras
credits_exhausted, one deprecated openrouter free-tier model), not code
regressions.

(cherry picked from commit c40b13a48fd897259c56f5122e9e57a3dc7654ba)

* test(integration): add rootless wire-capture correlation to the live-test tool

Temporary WIP commit on this deferred branch — lands in the same final
live-test-tool PR as the general default-combo suite, never bundled into
a bug-fix PR.

liveContainerHarness.ts spins up a dedicated, throwaway podman container
(same runner-base image target as the operator's local dev/beta
containers) so wire-capture tests are fully self-contained: builds the
image if missing, starts the container with a persistent data dir, waits
for health, seeds the real "default" combo + provider connections from
the operator's local omniroute-dev instance (idempotent — only runs once
per data dir), and provisions API keys via the running instance's own
auth flow.

wireCapture.ts captures the container's actual network traffic via
`podman unshare nsenter --net=<container netns> -- tcpdump` — no root
needed, verified working live (this generalizes the root-requiring
`sudo nsenter -t $PID` command scripts/sre/tcp-close-analyzer.py already
documented for the same rootless-Podman netns problem; that script's
docstring now documents both). Capture and analysis needed two real fixes
found only by running the pipeline live: `-U` (unbuffered tcpdump writes)
plus a `pkill -f <pcap path>` fallback, since `podman unshare -> nsenter
-> tcpdump` is a 3-level subprocess chain and SIGTERM to the top-level
process doesn't reach the tcpdump grandchild, leaving an orphaned process
and a truncated/unreadable pcap; and filtering on the container's
internal listening port (20128) rather than the dynamically-assigned host
port, since capture happens inside the container's own network namespace
where only the internal port is meaningful.

live-default-combo-wire-capture.test.ts (gated on RUN_LIVE_WIRE_CAPTURE=1)
ties it together: sends a small representative sample of requests through
the real default combo, then cross-checks each one's app-level JSON
status against the actual HTTP status line observed on the wire via
scripts/sre/tcp-close-analyzer.py's stream reassembly — catching bugs
where the app layer claims success but the wire shows a
truncated/reset stream, not just what liveDefaultComboShared.ts's
existing breadth suite already covers.

Live-verified end-to-end: 4/4 sampled requests correlated correctly
across 8 captured TCP streams, container + capture process fully torn
down afterward (verified no orphaned podman container or tcpdump
process left running).

sendModelRequest/filterActiveModelTargets (liveDefaultComboShared.ts) gain
optional baseUrl/apiKey overrides, defaulting to the existing module-level
omniroute-beta target, so the wire-capture suite can point the same
request-sending logic at its own dedicated container instead.

(cherry picked from commit 914a7e42cbe914f257db9f72eedc902ee1532083)

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 09:53:07 -03:00
Diego Rodrigues de Sa e Souza
efbc7a7ba2 fix(perf): memoize synced pricing reads (#9861)
Co-authored-by: chloeassistant <279834366+chloeassistant@users.noreply.github.com>
2026-08-09 09:52:59 -03:00
Diego Rodrigues de Sa e Souza
a7d2dba1eb fix(types): narrow DeepSeek tool calls (#9860)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:52:54 -03:00
Diego Rodrigues de Sa e Souza
a1833b1159 fix(skills): normalize web fetch credentials (#9859)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:52:48 -03:00
Diego Rodrigues de Sa e Souza
c4c39b1a4a cherry-pick(pr-9770): chore(repo): ignore Electron build output unpacked into repo root (#9858)
* chore(repo): ignore Electron build output unpacked into repo root

electron-builder (squirrel-windows target) unpacks the packaged app -- the
entire Chromium runtime, ~24k files -- directly into the repository root:
OmniRoute.exe, chrome_*.pak, *.dll, locales/, resources/, icudtl.dat,
snapshot blobs and the Chromium license files.

None of it was covered by .gitignore, so `git add -A` would commit the whole
runtime. Every rule is root-anchored (leading `/`) because a bare `locales/`
or `resources/` would also swallow tracked sources -- notably the CLI
translations in bin/cli/locales/*.json.

Verified with `git check-ignore`: all artifact paths ignored, and
bin/cli/locales/{en,de}.json remain tracked.

* chore(electron): sync package-lock for windows installer deps

Adds the lockfile entries for the Windows installer/signing toolchain that
the electron build now pulls in: electron-builder-squirrel-windows,
electron-winstaller and @electron/windows-sign (plus their transitive
fs-extra/jsonfile/universalify/mkdirp pins), and bumps app-builder-lib and
builder-util-runtime.

Lockfile-only change; no source or runtime behaviour is affected.

---------

Co-authored-by: Mihaly Bodo <michael@proton-quantum.com>
2026-08-09 09:52:42 -03:00
Diego Rodrigues de Sa e Souza
8a17f43849 fix(i18n): translate validation model keys in 34 locales (#9857)
The provider-connection dialog (AddApiKeyModal / EditConnectionModal)
rendered humanized key names instead of real copy for
providers.validationModelId{Label,Placeholder,Hint} in 34 of 43 locales —
the values read "Validation Model Id Label", "Validation Model Id
Placeholder" and "Validation Model Id Hint" verbatim.

Each translation follows the terminology and register already used by the
neighbouring provider keys in its own file — e.g. de Anbieter/API-Schlüssel
with formal Sie, fr fournisseur/clé API, ru провайдер/ключ API — and each
locale's own "e.g." convention (z. B., 例:, напр., ör., cth., hal.).

Source of truth is en.json, which labels the field "Validation Model"
(no "ID"); a few older locales say "validation model ID" and were left
untouched rather than propagating that divergence.

Co-authored-by: Mihaly Bodo <michael@proton-quantum.com>
2026-08-09 09:52:34 -03:00
Diego Rodrigues de Sa e Souza
332c738844 fix(sse): route claude/<provider>/<model> aliases for catalog-only providers (#9856)
The /v1/models catalog mirrors `claude/<provider>/<model>` ids purely from the
alias gate -- ccAliasPredicate.ts consults no provider registry. The request
path additionally required the prefix to be an open-sse REGISTRY entry or an
operator-defined custom node.

Enterprise-cloud providers such as azure-ai / azure-openai live only in the
provider catalog (src/shared/constants/providers/apikey/enterprise-cloud.ts).
They route fine directly -- `azure-ai/Phi-4` returns 200 -- but have no
open-sse registry entry, so the two sides disagreed: the catalog advertised
`claude/azure-ai/<model>` while stripCcDiscoveryAlias refused to strip it.

The unstripped id then fell through to normal resolution, which splits on the
first / and parsed `claude` as the provider. Every Claude Code request for an
Azure model was routed to the Claude provider instead:

  ROUTING: Provider: claude, Model: azure-ai/DeepSeek-V4-Flash

Extract the predicate as `isRoutableProviderPrefix()` and widen it to the
provider catalog (id + alias) alongside the open-sse registry, so the request
path recognises exactly what the catalog can advertise.

Regression guard: tests/unit/cc-discovery-alias-routable-prefix.test.ts pins
azure-ai/azure-openai/azure as routable, keeps openai/anthropic routable, and
keeps an unknown prefix non-routable. Verified failing before the widening.

Co-authored-by: Mihaly Bodo <michael@proton-quantum.com>
2026-08-09 09:52:28 -03:00
Diego Rodrigues de Sa e Souza
a102a2d773 maint: final follow-up cherry-pick #9783 (#9904)
* fix(deps): bump transitive deps for 6 Dependabot + remaining audit vulns on main

Same overrides as #9464 (ip-address, hono, fast-uri, socket.io-parser, undici)
applied directly to main. Also covers brace-expansion (scoped), js-yaml v4 copies,
and mermaid.

npm audit: 6→0 vulnerabilities.
Closes Dependabot #161-#166.

* fix(deps): bump nanoid, dompurify for 2 new Dependabot alerts (#189, #190)

Bumps: nanoid ^3.3.17 (was transitive, now overridden), dompurify ^3.4.13
(with monaco-editor scoped override). Closes Dependabot #189, #190.

Remaining #182-#188 (js-yaml + mermaid) already closed by #9651 merge —
awaiting Dependabot re-scan.

npm audit → 0 vulnerabilities.

* fix(repo): harden .gitignore to also ignore a _tasks symlink (/_tasks)

_tasks is a SEPARATE nested git repo (gitignored). The pattern _tasks/ (trailing
slash) ignores only a directory, not a SYMLINK named _tasks. A self-referential
_tasks symlink can slip in via git add -A and, once pulled, checkout materializes
it over the real _tasks repo (destroying plans/specs/hands-off). Anchored /_tasks
ignores the symlink too, preventing re-capture.

* fix(translator): keep Responses namespace identity across the hub-and-spoke pivot

Step 1 of the pivot (openai-responses -> openai) flattens namespace sub-tools
to a qualified wire name (#8295) and records the `{namespace, name}` pair on a
non-enumerable `_toolNameMap`. Step 2 (openai -> target) returns a brand-new
object, so the property was dropped for every non-OpenAI target. chatCore then
handed `null` to the #7936 response seam and namespace sub-tool calls reached
the client under their flattened name, which Codex rejects with
`unsupported call: <name>` — the symptom #7936 was opened to fix.

Copying `_toolNameMap` through is not viable: openai-to-claude and
openai-to-gemini publish their own `Map<string, string>` alias map on that same
property during step 2, so it carries two incompatible types. This adds a
dedicated `_namespaceToolIdentityMap`, propagated by translateRequest across
the pivot; chatCore prefers it and falls back to `_toolNameMap` for the
non-pivot producers. Both keys are stripped from the cliproxyapi wire body.

Fixes #9780

* fix(chat): reduce file size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(chat): reduce combined file size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(chat): reduce combined file size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: VXNCXNX <vincent@preuve.ai>
2026-08-09 09:52:23 -03:00
Diego Rodrigues de Sa e Souza
5926d35758 cherry-pick(pr-9787): fix(sse): apply Azure param rules on azure-ai and clamp gpt-4o-mini output tokens (#9855)
* fix(sse): apply Azure request-param rules on the azure-ai wire path

Azure rejects several stock Chat Completions params on its newer deployments
and returns HTTP 400 rather than ignoring them:

  max_tokens       -> 'max_tokens' is not supported with this model.
                      Use 'max_completion_tokens' instead.
  reasoning_effort -> Function tools with reasoning_effort are not supported.

Those rules lived inline in AzureOpenAIExecutor, so they only covered the
azure-openai provider. azure-ai (Azure AI Foundry) had no executor entry and
fell through to the bare DefaultExecutor, so the SAME Azure deployment
succeeded on one connection and 400'd on the other. Every agentic client sends
tools on every turn, so azure-ai failed on the first request.

Extract the rules to open-sse/executors/azureParamRules.ts, add an
AzureAiExecutor that inherits DefaultExecutor's azure-ai URL/header/apiType
handling unchanged and applies the shared rules, and register it for azure-ai.

Also widen the deployment pattern to cover gpt-chat-latest: it is a moving
alias that resolves to a GPT-5-era model and rejects max_tokens, but carries no
version number for the token-boundary pattern to key on. Verified against the
base regex - gpt-chat-latest did not match, which is exactly the observed 400.

Regression guard: tests/unit/azure-param-rules.test.ts, including an assertion
that getExecutor("azure-ai") no longer resolves to a bare DefaultExecutor.

* fix(sse): clamp Azure gpt-4o-mini completion tokens to its 16384 ceiling

Azure gpt-4o-mini deployments accept at most 16384 completion tokens and 400 on
anything larger:

  max_tokens is too large: 32000. This model supports at most 16384 completion
  tokens, whereas you provided 32000.

The 32000 is OmniRoute's own doing: adjustMaxTokens raises any smaller
max_tokens to DEFAULT_MIN_TOKENS (32000) whenever tools are present, to avoid
truncated tool arguments. That floor has no upper bound, so an agentic client
asking for far less still trips the model ceiling on its first turn.

Add scoped maxOutputCap rules in paramSupport.ts for both Azure wire paths.
PROVIDER_MAX_TOKENS is the wrong lever here - it is provider-wide, and the same
Azure resource also serves GPT-5 deployments with a much higher ceiling.

Regression guard: tests/unit/azure-max-output-clamp.test.ts, which also pins
that the clamp does not leak to gpt-5.1 or to gpt-4o-mini on other providers.

---------

Co-authored-by: Mihaly Bodo <michael@proton-quantum.com>
2026-08-09 09:52:16 -03:00
Diego Rodrigues de Sa e Souza
48d43240f4 fix(api): enforce model permissions on gateway mirrors (#9854)
Co-authored-by: Xiangzhe <xiangzhedev@gmail.com>
2026-08-09 09:52:09 -03:00
Diego Rodrigues de Sa e Souza
87c145a2be fix(response): strip internal reasoning placeholder from all reasoning fields (#9853)
copyOpenAICompatibleReasoningFields only stripped the sentinel
(NON_ANTHROPIC_THINKING_PLACEHOLDER = "(prior reasoning summary
unavailable)") from reasoning_content and reasoning. Non-standard
reasoning fields (reasoning_text, thinking, thought) and
reasoning_details items passed through raw, leaking the internal
replay sentinel to clients on providers that use those fields
(e.g. Venice), where the model echo surfaces as a bogus thought block
and can degrade into empty turns.

Strip the sentinel from every forwarded reasoning field, including
per-item text/content inside reasoning_details; drop items/fields that
strip to nothing while preserving non-text details such as
reasoning.encrypted.

Fixes #9765
Refs #8081, #9606

Co-authored-by: safeer <asafeer1994@gmail.com>
2026-08-09 09:52:02 -03:00
Diego Rodrigues de Sa e Souza
4fe0fffb31 fix(types): preserve Claude thinking body contracts (#9852)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:51:56 -03:00
Diego Rodrigues de Sa e Souza
65dae70403 fix(types): normalize Gemini Business credentials (#9851)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:51:49 -03:00
Diego Rodrigues de Sa e Souza
0b5ab6570d fix(types): preserve The Old LLM proxy contracts (#9850)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:51:44 -03:00
Diego Rodrigues de Sa e Souza
97a1355037 fix(types): validate default executor pool config (#9849)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:51:37 -03:00
Diego Rodrigues de Sa e Souza
a2eab58dde fix(types): expose SQLite transaction state (#9848)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:51:30 -03:00
Diego Rodrigues de Sa e Souza
fed05a3207 fix(types): normalize DuckDuckGo request messages (#9847)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:51:22 -03:00
Diego Rodrigues de Sa e Souza
2c21f292cd fix(types): accept synced catalog model rows (#9846)
Co-authored-by: backryun <bakryun0718@proton.me>
2026-08-09 09:51:15 -03:00
Diego Rodrigues de Sa e Souza
d9df8bb512 maint: final follow-up cherry-pick #9810 (#9906)
* fix(deps): bump transitive deps for 6 Dependabot + remaining audit vulns on main

Same overrides as #9464 (ip-address, hono, fast-uri, socket.io-parser, undici)
applied directly to main. Also covers brace-expansion (scoped), js-yaml v4 copies,
and mermaid.

npm audit: 6→0 vulnerabilities.
Closes Dependabot #161-#166.

* fix(deps): bump nanoid, dompurify for 2 new Dependabot alerts (#189, #190)

Bumps: nanoid ^3.3.17 (was transitive, now overridden), dompurify ^3.4.13
(with monaco-editor scoped override). Closes Dependabot #189, #190.

Remaining #182-#188 (js-yaml + mermaid) already closed by #9651 merge —
awaiting Dependabot re-scan.

npm audit → 0 vulnerabilities.

* fix(repo): harden .gitignore to also ignore a _tasks symlink (/_tasks)

_tasks is a SEPARATE nested git repo (gitignored). The pattern _tasks/ (trailing
slash) ignores only a directory, not a SYMLINK named _tasks. A self-referential
_tasks symlink can slip in via git add -A and, once pulled, checkout materializes
it over the real _tasks repo (destroying plans/specs/hands-off). Anchored /_tasks
ignores the symlink too, preventing re-capture.

* docs(proposals): Telegram Mini App integration feasibility analysis

Assess adding a Telegram Mini App chat surface to OmniRoute. Verifies
against current main (918fba5e3) what exists (outbound telegram webhook
integration, bot-token validation + encryption gate) and what is missing
(inbound Bot API listener, WebApp initData HMAC verification, mini app
hosting, per-user API key mapping).

Concludes: feasible with moderate effort (2-4 dev-days for a working
slice). Identifies constraints (public HTTPS webhook, no native
streaming to Telegram, server-side initData trust, encryption gate) and
a phased next-steps plan (spike, minimal chat slice, hardening).

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: benzntech <bensonkbmca@gmail.com>
2026-08-09 09:51:06 -03:00
Diego Rodrigues de Sa e Souza
0bb17b91c6 maint: final follow-up cherry-pick #9812 (#9907)
* fix(deps): bump transitive deps for 6 Dependabot + remaining audit vulns on main

Same overrides as #9464 (ip-address, hono, fast-uri, socket.io-parser, undici)
applied directly to main. Also covers brace-expansion (scoped), js-yaml v4 copies,
and mermaid.

npm audit: 6→0 vulnerabilities.
Closes Dependabot #161-#166.

* fix(deps): bump nanoid, dompurify for 2 new Dependabot alerts (#189, #190)

Bumps: nanoid ^3.3.17 (was transitive, now overridden), dompurify ^3.4.13
(with monaco-editor scoped override). Closes Dependabot #189, #190.

Remaining #182-#188 (js-yaml + mermaid) already closed by #9651 merge —
awaiting Dependabot re-scan.

npm audit → 0 vulnerabilities.

* fix(repo): harden .gitignore to also ignore a _tasks symlink (/_tasks)

_tasks is a SEPARATE nested git repo (gitignored). The pattern _tasks/ (trailing
slash) ignores only a directory, not a SYMLINK named _tasks. A self-referential
_tasks symlink can slip in via git add -A and, once pulled, checkout materializes
it over the real _tasks repo (destroying plans/specs/hands-off). Anchored /_tasks
ignores the symlink too, preventing re-capture.

* feat(telegram): Mini App chat bridge — initData auth, update webhook, chat proxy

Implements the Phase-1 slice of the Telegram Mini App integration
(docs/proposals/TELEGRAM-MINIAPP.md):

- src/lib/telegram/initData.ts — dependency-free WebApp initData HMAC-SHA256
  verification (Telegram Bot API spec), with auth_date freshness check.
- src/lib/telegram/config.ts — TELEGRAM_BOT_TOKEN / model / API base / timeout
  env config; token format validation; enabled gate.
- src/lib/telegram/botApi.ts — minimal fetch-based Bot API client
  (sendMessage, editMessageText, setWebhook) + update shape helpers.
- src/lib/telegram/chatProxy.ts — maps a Telegram user to a per-user
  OmniRoute API key (createApiKey, name telegram:<userId>) and proxies
  prompts through the existing handleChat pipeline.
- src/app/api/telegram/update/route.ts — inbound endpoint serving both the
  Bot API update webhook (/start + chat replies) and the Mini App direct
  path (initData HMAC verified → 401 on mismatch). Public route prefix;
  own auth only.
- src/app/miniapp/page.tsx — Telegram WebApp SDK chat UI.
- Tests: telegram-init-data (7), telegram-botapi (5) — 12/12 pass.
- Env docs: TELEGRAM_* vars in .env.example + ENVIRONMENT.md (sync ✓).
- Route-validation check: PASS (body validated via Zod).

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: benzntech <bensonkbmca@gmail.com>
2026-08-09 09:50:58 -03:00
Diego Rodrigues de Sa e Souza
3d590c310b fix(ci): repair release lint test regressions (#9896)
Co-authored-by: Alex Jordan <60003097+alex-jordan547@users.noreply.github.com>
2026-08-09 09:50:49 -03:00
Diego Rodrigues de Sa e Souza
247f2606cd fix(admission): queue heavyweight chat requests before 503 busy (#9845)
Agent clients (OpenCode, Claude Code, Cursor) fan out heavy sub-requests
that land on the admission gate together. With the single heavyweight
slot, concurrent heavy requests were rejected immediately with a
retryable 503; clients burn their retry budget in seconds and the agent
dies mid-task.

Heavy requests now wait up to OMNIROUTE_CHAT_ADMISSION_QUEUE_MS (default
5000ms) for a slot before the 503, served FIFO; 0 restores the legacy
immediate-reject behaviour. Applied to both the byte-based path
(admitChatRequest) and the structure-based path (admitChatStructure, now
async).

Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
2026-08-09 09:50:37 -03:00
Diego Rodrigues de Sa e Souza
580162548a cherry-pick(pr-9818): feat: generic OpenAI-compatible video custom provider (#9844)
* feat: generic OpenAI-compatible video custom provider

Adds a generic OpenAI-compatible video generation path so users can add
custom video providers (base URL + API key) without per-provider code.

Changes:
- open-sse/handlers/videoGeneration/openai.ts (new): generic handler
  with resolveVideoEndpoint, fetchVideoEndpoint, handleOpenAIVideoGeneration
- open-sse/handlers/videoGeneration.ts: added resolveVideoBaseUrl(),
  dispatch for 'openai-video' format before 'vertex-veo', synthetic config
  for custom providers, fallback for resolvedProvider
- src/app/api/v1/videos/generations/route.ts: scans custom models for
  supportedEndpoints.includes('videos'), resolves credentials via
  getProviderCredentialsWithQuotaPreflight, passes resolvedProvider
- src/shared/validation/schemas/provider.ts: added 'videos' to
  supportedEndpoints enum
- tests/unit/video-generation-handler.test.ts: handler-level test for
  custom provider
- tests/unit/video-custom-provider-route.test.ts (new): route-level tests
  covering custom provider with/without videos endpoint, unknown provider

All verification:
- typecheck:core passes
- 17 video tests pass (3 new route tests + 1 new handler test)
- no regressions in image generation tests

* test(video): drop duplicated test.after cleanup in custom-provider route test

* feat(video): declarative job presets + dispatcher, route, and handler test coverage (#9818)

- job.ts: presets (agnes-video-job, muapi-video-job) with submit→poll→done executor
- videoGeneration.ts: generationConfig.preset dispatch branch + mediaGenerationRoute pass-through
- provider-models/route.ts + models.ts: generationConfig persisted on addCustomModel
- provider schema: generationConfig optional field
- tests: resolvedProvider bare-model + job preset happy/failed/unknown paths
- docs/video-preset-generation.md

* fix(video): restore dashscope + novita handler imports dropped in refactor

* refactor(video): extract Runway helpers

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: oyi77 <oyi77@users.noreply.github.com>
2026-08-09 09:50:27 -03:00
Diego Rodrigues de Sa e Souza
79b8c8351c fix(command-code): include tool call arguments (#9897)
Co-authored-by: Choti Wongbussakorn <126886556+Chewji9875@users.noreply.github.com>
2026-08-09 09:50:17 -03:00
Diego Rodrigues de Sa e Souza
05940f4c7f fix(responses-api): tool call after a text message collided on the same output_index (#9843)
Live incident (2026-08-08): an OpenClaw agent sent a short preamble line
("Kör nu, på riktigt — apply_patch på vibe-scriptet:") followed by an
apply_patch tool call in the same turn. The client only spoke the preamble
and never executed the patch, even though OmniRoute's own recorded
responseBody had a complete, valid tool_calls entry.

Root cause: emitToolCall/closeToolCall computed a tool call's output_index
as `reasoningIndex + 1 + tcIdx`, assuming reasoningIndex + 1 was free for
the first tool call (tcIdx=0). But a text message emitted in the same turn
ALSO claims reasoningIndex + 1 (or index 0 with no reasoning) — so a
turn with reasoning + text content + a tool call collided the tool call's
added/delta/done events onto the same output_index as the just-closed
message. A client that tracks response items by output_index (as expected
for the Responses API) sees the tool call events land on an index it
already marked complete and can silently drop them.

Fix: track whether a message item was actually emitted at that index
(state.msgItemAdded) and, if so, tool calls start one slot after it.
Extracted a shared toolCallOutputIndexBase() helper so emitToolCall and
closeToolCall can no longer compute this independently and drift apart.

Confirmed via the live call log artifact (id 1786223153235-770a1c):
response.output_item.done for the text message and response.output_item.added
for the tool call both carried output_index=1 in the raw SSE stream, 1.84s
apart, exactly matching the reported symptom.

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-09 09:50:09 -03:00
Diego Rodrigues de Sa e Souza
0f5699165b fix(providers): remove retired NVIDIA NIM catalog entries (#9898)
Co-authored-by: Zartharas <1402357+Zartharas@users.noreply.github.com>
2026-08-09 09:47:45 -03:00
Diego Rodrigues de Sa e Souza
714a36cf99 cherry-pick(pr-9826): fix(executors): preserve Command Code usage in Responses streams (#9842)
* fix(executors): preserve Command Code usage in Responses streams

* fix(executors): add Command Code usage changelog fragment

---------

Co-authored-by: MrShitFox <qwert2006gleb@gmail.com>
2026-08-09 09:47:37 -03:00
Diego Rodrigues de Sa e Souza
bbe5c78f4d cherry-pick(pr-9828): fix(executors): strip redundant oneOf matching sibling enum (#9841)
* fix(executors): strip redundant oneOf matching sibling enum

The Codex private Responses endpoint intermittently returns a 502 upstream_empty_response for tool parameters that combine oneOf:[{const,...}] with a sibling enum containing the same value set.

When the const and enum sets match exactly, oneOf adds no constraint beyond enum. Add stripRedundantOneOfConstEnum to normalizeCodexTools to remove only this semantically redundant form.

The schema-aware recursive walker requires non-empty, unique string const branches containing annotations only, string enum values, and an exact set match. It preserves bare oneOf[const], narrowing or non-matching sets, type-discriminated oneOf, empty oneOf, non-string values, and anyOf/allOf.

Run the normalization after stripUnsupportedRegexPatterns and before assigning tool.parameters. Add focused regression coverage for matching, non-matching, nested, immutable, and Chat-to-Responses cases.

* docs(changelog): update PR number in changelog fragment

---------

Co-authored-by: Vasily Larin <larin.vas@outlook.com>
2026-08-09 09:47:28 -03:00
Diego Rodrigues de Sa e Souza
2d49f1c743 maint: follow-up cherry-pick fix-in-place #9833 (conflict-resolved fallback) (#9899)
* fix(nvidia): keep 410 failures model-scoped

* test: register NVIDIA 410 regression for mutation coverage

* chore: preserve Stryker config formatting

* fix(auth): reduce file size

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Zartharas <1402357+Zartharas@users.noreply.github.com>
2026-08-09 09:47:17 -03:00
Diego Rodrigues de Sa e Souza
d11b99f6cc cherry-pick(pr-9834): fix(cursor): SelectedImage blobIdWithData + JPEG soft-cap prep (#9840)
* fix(cursor): hydrate SelectedImage via blobIdWithData + JPEG soft-cap

Cursor vision expects SelectedImage.blob_id_with_data (field 9) backed by
the session blobStore, and large clipboard PNGs need JPEG soft-cap prep
rather than a hard 1 MiB reject before encode.

* docs(changelog): add fragment for Cursor SelectedImage blobIdWithData fix

* refactor(cursor): split image protobuf encoding

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: SB Yoon <44089734+yansigit@users.noreply.github.com>
2026-08-09 09:47:07 -03:00
Diego Rodrigues de Sa e Souza
065fa67f63 chore(quality): reconcile final v3.8.50 ratchets (#9839)
* chore(quality): reconcile final v3.8.50 ratchets

* chore(changelog): record v3.8.50 ratchet reconciliation

* chore(ci): retrigger base-reds reconciliation checks for #9839

* chore(ci): retrigger base-red sweep run for #9839

* chore(ci): retrigger base-red checks after queued-cancel

* chore(ci): retrigger base-reds checks #9839 (queue clear)

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-09 09:28:12 -03:00
Diego Rodrigues de Sa e Souza
4d8506c2c5 fix(i18n): restore Vietnamese locale parity (#9925)
* fix(i18n): restore Vietnamese locale parity

* fix(changelog): follow fragment convention

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-09 09:06:05 -03:00
dionjoshualobo
918647af69 fix(i18n): restore escaped entities in gatesDescription 2026-08-09 02:18:40 -03:00
dionjoshualobo
f39d74b83a fix(i18n): unescape HTML entities in UI strings 2026-08-09 02:18:33 -03:00
350 changed files with 49885 additions and 22771 deletions

View File

@@ -67,14 +67,6 @@ DISABLE_SQLITE_AUTO_BACKUP=false
# Used by: src/shared/utils/rateLimiter.ts
# Example: redis://localhost:6379 (or redis://redis:6379 in Docker)
# REDIS_URL=redis://localhost:6379
# Host interface docker-compose publishes the Redis sidecar on.
# Default: 127.0.0.1 (loopback only). The compose Redis runs WITHOUT
# `requirepass`, and app containers reach it over the compose network
# (redis:6379) — the published port is only for host-side tooling. Setting this
# to 0.0.0.0 exposes an unauthenticated Redis to your whole LAN.
# REDIS_BIND_HOST=127.0.0.1
# Host port for the compose Redis sidecar. Default: 6379.
# REDIS_PORT=6379
# ═══════════════════════════════════════════════════════════════════════════════
# 3. NETWORK & PORTS
@@ -345,18 +337,14 @@ ALLOW_API_KEY_REVEAL=false
# OMNIROUTE_CHAT_HEAVY_TOOL_COUNT=64
# Conservative string-size token estimate that classifies a request as heavyweight. Default 32000.
# OMNIROUTE_CHAT_HEAVY_ESTIMATED_TOKENS=32000
# Optional opt-in hard message-count cap; excess receives compact-required 413 before
# compression can run. Unset/0 (the default) means no history cap: heap growth is bounded
# by OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT and the heap-pressure shed instead. Set a positive
# value only on memory-constrained deployments that need a hard ceiling.
# OMNIROUTE_CHAT_HARD_MAX_MESSAGES=0
# Hard message-count cap; excess receives compact-required 413. Default 800.
# OMNIROUTE_CHAT_HARD_MAX_MESSAGES=800
# Hard cap (bytes) for a non-streaming upstream response buffered fully into memory
# (#5152). Past this the upstream reader is cancelled and the request fails fast
# instead of growing an unbounded string until the V8 heap is exhausted.
# Used by: open-sse/handlers/chatCore/nonStreamingResponseBody.ts
# Default: 67108864 (64 MB)
# OMNIROUTE_FORWARDING_HEADER_BUDGET_BYTES=768
# OMNIROUTE_MAX_NONSTREAMING_RESPONSE_BYTES=67108864
# CORS configuration — controls which cross-origin browser clients can call the API.
@@ -457,13 +445,6 @@ ALLOW_API_KEY_REVEAL=false
# Default: false
# OMNIROUTE_PREFER_CLAUDE_CODE_FOR_UNPREFIXED_CLAUDE_MODELS=false
# Per-model concurrency cap for round-robin combos (#9100).
# Used by: open-sse/services/comboConfig.ts — the round-robin combo semaphore
# was hard-capped at 3 concurrent requests per model with no override, which
# serialized higher-concurrency traffic behind that cap.
# Validated to >= 1, clamped to <= 32. | Default: 3
# COMBO_CONCURRENCY_PER_MODEL=3
# ═══════════════════════════════════════════════════════════════════════════════
# 7. URLS & CLOUD SYNC
# ═══════════════════════════════════════════════════════════════════════════════
@@ -539,6 +520,16 @@ NEXT_PUBLIC_BASE_URL=http://localhost:20128
# cost of more upstream polling; raise to reduce request volume.
# OMNIROUTE_CGPT_WEB_PRO_POLL_INTERVAL_MS=4000
# Timeout for the /api/jobs/:id/run-now endpoint, in milliseconds.
# This bounds the CALL, not the job. runNow() dispatches the handler with
# `void` and returns as soon as it has decided to start, so on the normal
# path it resolves in milliseconds. It only matters when the job is already
# running: runNow() then waits for the in-flight run before starting the
# queued one, and this timeout prevents that wait from hanging forever.
# Used by: src/app/api/jobs/[id]/run-now/route.ts
# Default: 30000 (30 seconds)
# OMNIROUTE_RUNNOW_TIMEOUT_MS=30000
# Public cloud URL — client-side mirror of CLOUD_URL.
NEXT_PUBLIC_CLOUD_URL=
@@ -797,16 +788,6 @@ PROVIDER_LIMITS_SYNC_SPACING_MS=1500
# Disable the proactive recovery scheduler entirely (default: false).
# OMNIROUTE_DISABLE_CONNECTION_RECOVERY=false
# Proactive Claude warmup scheduler (#8848): fires a trivial request to opted-in
# OAuth connections on a cron schedule (America/Los_Angeles) so accounts do not
# hit the 5-hour sliding window cold. Off by default — set ENABLED=1 and flip
# per-connection flags in settings.claudeWarmup.connections to activate.
# Used by: src/lib/warmupScheduler.ts.
# OMNIROUTE_WARMUP_ENABLED=false
# OMNIROUTE_WARMUP_CRON="0 7 * * *"
# OMNIROUTE_WARMUP_CONCURRENCY=3
# OMNIROUTE_WARMUP_MODEL=
# Background job interval for budget reset checks (ms). Default: 600000 (10m).
# Used by: src/lib/jobs/budgetResetJob.ts. Floor: 10000.
#OMNIROUTE_BUDGET_RESET_JOB_INTERVAL_MS=600000
@@ -871,12 +852,6 @@ PROVIDER_LIMITS_SYNC_SPACING_MS=1500
# (>= 3 retrievals = never compressed). 1 disables the ramp (binary skip at the threshold only).
# Used by: open-sse/services/compression/engines/ccr/index.ts. Default: 2.
#COMPRESSION_CCR_RETRIEVAL_RAMP_FACTOR=2
# CCR durable block store (#9061). The in-memory store loses blocks to LRU eviction, the TTL, a
# restart, or a retrieve landing on another instance, while the model is told it can retrieve them
# verbatim. Set to false to keep blocks in memory only, at the cost of that promise. Blocks over
# 512KB and cloud runtimes are memory-only regardless.
# Used by: open-sse/services/compression/engines/ccr/index.ts. Default: true.
#COMPRESSION_CCR_DURABLE_STORE=true
# T08/H5 — usage-observed prefix freeze (OPT-IN, default off). When enabled, a system prompt seen
# >= THRESHOLD times is treated as a stable cacheable prefix and preserved from compression even
# for providers the static cache-aware heuristic does not recognize (freeze = preserve, never
@@ -1061,17 +1036,6 @@ GITHUB_OAUTH_CLIENT_ID=Iv1.b507a08c87ecfe98
# VISION_BRIDGE_BASE_URL=
# VISION_BRIDGE_API_KEY=
# ── Raycast Pro (local auto-import) ──
# Raycast Pro AI is a reverse-engineered, unofficial API — local/personal use
# only (no OAuth client_id/secret; token is captured via macOS Auto-Import
# from the Keychain + local Raycast SQLite DB, or pasted manually). These
# vars are optional manual overrides used by open-sse/services/raycast.ts
# and the direct-probe benchmark script scripts/raycast/usage-benchmark.mjs.
# RAYCAST_BEARER_TOKEN=
# RAYCAST_DEVICE_ID=
# RAYCAST_AID=
# RAYCAST_SIG_SECRET=
# ─────────────────────────────────────────────────────────────────────────────
# ⚠️ GOOGLE OAUTH (Antigravity) & OTHER PROVIDERS — REMOTE SERVERS
# ─────────────────────────────────────────────────────────────────────────────
@@ -1212,17 +1176,6 @@ CURSOR_USER_AGENT="Cursor/3.4"
# fallback when FETCH_TIMEOUT_MS is unset. Default: 120000 (2 min).
# OMNIROUTE_DEFAULT_FETCH_TIMEOUT_MS=120000
# ── Proxy/relay fetch (connection pooling, #9158) ──
# Used by: open-sse/utils/proxyFetch.ts.
# A hung relay must fail BEFORE the client/agent timeout (typically 30s) so the
# caller sees a relay-specific failure instead of a generic upstream timeout.
# Capped at 29000ms so this timeout always fires first. Default: 25000 (25s).
# OMNIROUTE_RELAY_FETCH_TIMEOUT_MS=25000
# Shared retry backoff (ms) for the direct/relay/proxy retry-once paths.
# 0 = retry immediately. Default: 10.
# OMNIROUTE_RETRY_BACKOFF_MS=10
# ── Firecrawl web-fetch executor ──
# Point at a self-hosted Firecrawl instance (defaults to the public cloud API).
# When set to a non-cloud base URL, the API key becomes optional.
@@ -1284,14 +1237,6 @@ CURSOR_USER_AGENT="Cursor/3.4"
# OMNIROUTE_BROWSER_POOL=on
# WEB_COOKIE_USE_BROWSER=0
# ── Adobe Firefly browser sign-in (system Chrome/Edge CDP) ──
# Used by: open-sse/services/adobeFireflyBrowserLogin.ts. The Firefly login
# flow drives a real, system-installed Chrome or Microsoft Edge via CDP so the
# user can sign in interactively; the executable is auto-detected from common
# install paths per OS. Set this to override that detection (e.g. a portable
# install or a non-standard path) when auto-detection fails.
# OMNIROUTE_LOGIN_BROWSER_PATH=
# ── Circuit breaker thresholds and reset windows ──
# Used by: open-sse/config/constants.ts → src/lib/resilience/settings.ts.
# Defaults match historical PROVIDER_PROFILES values (post-scaling for
@@ -1403,10 +1348,6 @@ APP_LOG_TO_FILE=true
# Default: 100000
# CALL_LOGS_TABLE_MAX_ROWS=100000
# Force detailed request logging on or off, overriding the dashboard setting.
# Values: true | false | Default: unset (follow dashboard setting)
# ENABLE_REQUEST_LOGS=false
# Maximum age for orphaned active request log entries before the in-memory
# pending-request reaper removes them. Accepts milliseconds.
# Default: 3600000 (1 hour)
@@ -1426,9 +1367,10 @@ APP_LOG_TO_FILE=true
# bodies is retained in the database.
# Used by: open-sse/handlers/chatCore.ts — cloneBoundedChatLogPayload()
# CHAT_LOG_TEXT_LIMIT=65536 # Max string length before truncation (default: 64 KB)
# CHAT_LOG_ARRAY_TAIL_ITEMS=24 # Number of array items retained from tail (default: 24)
# CHAT_LOG_ARRAY_TAIL_ITEMS=128 # Number of array items retained from tail (default: 128)
# CHAT_LOG_MAX_DEPTH=6 # Max nesting depth before truncation (default: 6)
# CHAT_LOG_MAX_OBJECT_KEYS=80 # Max object keys retained (default: 80, 0 = no limit)
# CHAT_LOG_MAX_BODY_KB=1024 # Max request/response body size before summarizing, in KB (default: 1024)
# Maximum rows in the proxy_logs SQLite table.
# Default: 100000
@@ -1483,6 +1425,10 @@ APP_LOG_TO_FILE=true
# Default: ~/.omniroute/plugins/ Override in dev/CI to point at a local plugin tree.
# OMNIROUTE_PLUGIN_PATH=
# Allow plugins to request the 'exec' permission (spawn child processes from the
# plugin worker sandbox). Disabled by default; set to 1 to enable (local operator only).
# OMNIROUTE_PLUGINS_ALLOW_EXEC=0
# ── Prompt cache (system prompt deduplication) ──
# Used by: open-sse/services — caches identical system prompts across requests.
# PROMPT_CACHE_MAX_SIZE=50 # Max cached entries (default: 50)
@@ -1540,15 +1486,6 @@ APP_LOG_TO_FILE=true
# ═══════════════════════════════════════════════════════════════════════════════
# 19. MODEL SYNC (Dev)
# ═══════════════════════════════════════════════════════════════════════════════
# Enable the models.dev capability sync. Default: false (opt-in only).
# Also settable from Dashboard > Settings > AI. This variable wins over that
# setting whenever it is set to anything non-empty, in either direction, so a
# deployment can pin the sync on or off without depending on database state
# surviving a rebuild. Leave it unset to let the dashboard toggle decide.
# On: 1, true, yes or on (any casing). Any other value is off.
# Used by: src/lib/modelsDevSync.ts
# MODELS_DEV_SYNC_ENABLED=false
# Development-time model catalog sync interval in seconds.
# Used by: src/lib/modelsDevSync.ts
# Default: 86400 (24 hours)
@@ -1571,14 +1508,6 @@ APP_LOG_TO_FILE=true
# Default: 86400000 (24 hours)
# OPENROUTER_CATALOG_TTL_MS=86400000
# Enrich the dashboard providers list with OpenRouter weekly ranking stats.
# ON by default; set false to skip the background fetch entirely (#9324).
# Used by: src/lib/catalog/openrouterProviderStats.ts
# OPENROUTER_PROVIDER_STATS_ENABLED=true
# Cache TTL for the OpenRouter provider stats snapshot, in ms.
# Default: 86400000 (24 hours)
# OPENROUTER_PROVIDER_STATS_TTL_MS=86400000
# ── Model catalog response shape ──
# Include display-friendly name fields in /v1/models responses.
# Disable for clients that expect model IDs only.
@@ -1593,19 +1522,25 @@ APP_LOG_TO_FILE=true
# NANOBANANA_POLL_TIMEOUT_MS=120000 # Max wait for job completion (default: 120s)
# NANOBANANA_POLL_INTERVAL_MS=2500 # Poll frequency (default: 2.5s)
# ── Adobe Firefly (Image / Video Generation) ──
# Optional absolute path to a system Chrome or Edge executable used for interactive sign-in
# and off-screen risk-session renewal. Auto-detected when unset.
# OMNIROUTE_LOGIN_BROWSER_PATH=
# Browser renewal and durable session cache are enabled by default; set either to 0 to opt out.
# ADOBE_FIREFLY_BROWSER_REFRESH=1
# ADOBE_FIREFLY_SESSION_DISK=1
# Minimum gap between generate submissions and extra gap after every third success (ms).
# ADOBE_FIREFLY_MIN_SUBMIT_GAP_MS=12000
# ADOBE_FIREFLY_BATCH_EXTRA_GAP_MS=15000
# Base backoff after a transient 408 response (ms); five attempts maximum.
# ADOBE_FIREFLY_SUBMIT_BASE_DELAY_MS=8000
# ── Microsoft Designer Web (Image Generation) ──
# Polling config for the microsoft-designer-web submit-then-poll image job.
# Used by: open-sse/handlers/imageGeneration/providers/designerWeb.ts
# DESIGNER_WEB_POLL_TIMEOUT_MS=60000 # Max wait for job completion (default: 60s)
# DESIGNER_WEB_POLL_INTERVAL_MS=2000 # Poll frequency (default: 2s)
# ── Adobe Firefly (Image Upscale) ──
# Base delay (ms) for the submit-retry exponential backoff when Adobe Firefly's
# upscale job submission is rate-limited. Used by:
# open-sse/services/adobeFireflyUpscale.ts::submitRetryDelayMs.
# Default: 8000 (20 under NODE_ENV=test/VITEST/NODE_TEST_CONTEXT).
# ADOBE_FIREFLY_SUBMIT_BASE_DELAY_MS=8000
# ── AWS Bedrock (Kiro / Audio) ──
# Region used to construct AWS Bedrock endpoints. Used by:
# src/lib/providers/validation.ts and open-sse/handlers/audioSpeech.ts.
@@ -1700,26 +1635,6 @@ APP_LOG_TO_FILE=true
# Used by: src/lib/services/bootstrap.ts, src/app/api/services/mux/_lib.ts
# MUX_SERVICE_PORT=8322
# ── Dario embedded service ──
# Override the host/port the embedded Dario (Claude Code subscription proxy)
# daemon binds to and is reached at. Always bound to 127.0.0.1 — never
# configurable to 0.0.0.0. Rarely needed — defaults to 127.0.0.1:3456.
# Used by: src/lib/services/installers/dario.ts, src/lib/services/bootstrap.ts,
# src/app/api/services/dario/_lib.ts, src/app/api/services/dario/admin/_lib.ts,
# open-sse/executors/dario.ts
# DARIO_HOST=127.0.0.1
# DARIO_PORT=3456
# ── Dario embedded service ──
# Override the host/port the embedded Dario (Claude Code subscription proxy)
# daemon binds to and is reached at. Always bound to 127.0.0.1 — never
# configurable to 0.0.0.0. Rarely needed — defaults to 127.0.0.1:3456.
# Used by: src/lib/services/installers/dario.ts, src/lib/services/bootstrap.ts,
# src/app/api/services/dario/_lib.ts, src/app/api/services/dario/admin/_lib.ts,
# open-sse/executors/dario.ts
# DARIO_HOST=127.0.0.1
# DARIO_PORT=3456
# ── Local hostnames (Docker networking) ──
# Comma-separated additional hostnames treated as "local" for provider routing.
# Used by: open-sse/config/providerRegistry.ts — allows Docker service names.
@@ -1952,18 +1867,6 @@ APP_LOG_TO_FILE=true
# ── Devin CLI binary path ──
# Used by: open-sse/executors/devin-cli.ts. Default: looked up via PATH.
# CLI_DEVIN_BIN=devin
# Agentic bridge-only binary override. The bridge still executes ACP stdio only.
# CLI_DEVIN_AGENTIC_BIN=devin
# Required isolated HOME for the agentic Devin child process.
# DEVIN_AGENTIC_HOME=/home/bridge
# Bounded ACP turn timeout in milliseconds. Default: 120000.
# DEVIN_AGENTIC_ACP_TIMEOUT_MS=120000
# Agentic bridge model aliases. Values must keep the devin-cli-agentic/ prefix.
# DEVIN_BRIDGE_MODEL=devin-cli-agentic/swe-1-7
# DEVIN_BRIDGE_SONNET_MODEL=devin-cli-agentic/swe-1-7
# DEVIN_BRIDGE_OPUS_MODEL=devin-cli-agentic/swe-1-7
# DEVIN_BRIDGE_HAIKU_MODEL=devin-cli-agentic/swe-1-7
# DEVIN_BRIDGE_SUBAGENT_MODEL=devin-cli-agentic/swe-1-7
# ── Command Code (custom CLI) callback ──
# Local port used for OAuth-style callbacks from the Command Code CLI helper.
@@ -2018,15 +1921,6 @@ APP_LOG_TO_FILE=true
# CHANGELOG_BASE_REF=origin/release/v0.0.0
# ALLOW_CHANGELOG_REMOVALS=1
# ── Remote audio provider nodes ──
# Used by: src/app/api/v1/_shared/audioProviderNodes.ts — lets the /v1/audio/*
# routes use an OpenAI-compatible provider node hosted outside localhost.
# OFF by default: routing audio to a remote host changes egress identity, so it
# must be an explicit operator decision. Loopback/private nodes (localhost,
# 127.0.0.1, 172.16-31.x) are always allowed and unaffected by this flag.
# When enabled, the node authenticates with the API key stored on its connection.
# AUDIO_REMOTE_PROVIDER_NODES=false
# ── 1Proxy egress pool ──
# Used by: src/lib/oneproxySync.ts — fetches proxy nodes from the OmniRoute
# CrofAI 1Proxy service. Disable, override URL, or tune the import quality.
@@ -2234,11 +2128,6 @@ PLAYGROUND_COMPARE_MAX_COLUMNS=4
# MEMORY_TYPED_DECAY_EPISODIC_DAYS=30 # episodic TTL in days; 0 = episodic immune too
# MEMORY_TYPED_DECAY_ACCESS_IMMUNITY=3 # access_count >= N → immune; 0 disables access immunity
# MEMORY_TYPED_DECAY_SWEEP_INTERVAL=0 # periodic sweep interval (seconds); 0 = no periodic sweep
# ─── Memory Backend Connectors (Generic HTTP) ──────────────────────────────
# NOTION_API_KEY=
# NOTION_API_URL=
# OBSIDIAN_API_KEY=
# OBSIDIAN_API_URL=
# AgentBridge + Traffic Inspector (Group A)
# AgentBridge
@@ -2254,15 +2143,6 @@ INSPECTOR_MAX_BODY_KB=1024
INSPECTOR_MASK_SECRETS=true
INSPECTOR_LLM_HOSTS_EXTRA=
INSPECTOR_INTERNAL_INGEST_TOKEN=
# Shared secret for identity-preserving internal REST hops (#9260): when an
# OmniRoute component calls another local OmniRoute route, this token (sent as
# x-omniroute-internal-service-token) marks the request as internal so the
# original caller identity is preserved. OPT-IN: unset disables the mechanism.
# Used by: src/lib/api/internalServiceAuth.ts
# OMNIROUTE_INTERNAL_SERVICE_TOKEN=
# File-based variant (secret-file pattern; wins only when the inline var is
# unset): path to a file whose trimmed content is the token.
# OMNIROUTE_INTERNAL_SERVICE_TOKEN_FILE=
# Quota Sharing (Group B — planos 16+22)
QUOTA_STORE_DRIVER=sqlite # sqlite | redis
# QUOTA_STORE_REDIS_URL= # ex.: redis://localhost:6379 (apenas quando driver=redis)
@@ -2373,11 +2253,6 @@ QUOTA_STORE_DRIVER=sqlite # sqlite | redis
# Host port for the 1-click Redis launcher. Default: 6379. Bump if the host
# already binds 6379. The container's internal port stays 6379.
# OMNIROUTE_REDIS_HOST_PORT=
# Host interface the 1-click Redis launcher publishes on. Default: 127.0.0.1
# (loopback only). The launcher starts Redis WITHOUT a password, so binding
# 0.0.0.0 hands every host on your LAN an unauthenticated Redis — only widen
# this if you also set a password on the instance yourself.
# OMNIROUTE_REDIS_BIND_HOST=
# Redis image used by the 1-click Redis launcher. Default: redis:7-alpine.
# Override to redis:8-alpine or a private registry mirror as needed.
# OMNIROUTE_REDIS_IMAGE=
@@ -2480,37 +2355,20 @@ QUOTA_STORE_DRIVER=sqlite # sqlite | redis
# ─────────────────────────────────────────────────────────────────────────────
# VIBEPROXY_DATA_DIR=
# ── Internal service auth (management-plane service-to-service calls) ─────────
# Inline token for internal service authentication; prefer the _FILE variant in
# containerized deployments so the secret never lands in the environment table.
# OMNIROUTE_INTERNAL_SERVICE_TOKEN=
# Path to a file containing the internal service token (overrides the inline var).
# OMNIROUTE_INTERNAL_SERVICE_TOKEN_FILE=
# ─────────────────────────────────────────────────────────────────────────────
# Telegram Mini App (inbound bot webhook + Mini App chat)
# Used by: src/lib/telegram/*, src/app/api/telegram/update/route.ts
# ─────────────────────────────────────────────────────────────────────────────
# Bot token from @BotFather (<numeric_id>:<secret>). Enables the inbound
# update webhook and doubles as the HMAC secret for Mini App initData
# verification. When unset, /api/telegram/update returns 503.
# TELEGRAM_BOT_TOKEN=
# ═══════════════════════════════════════════════════════════════════════════════
# 26. RADAR FEED (SELF-HOSTING)
# ═══════════════════════════════════════════════════════════════════════════════
# Optional add-on (feature flag RADAR_ENABLED, default off — see feature flag
# settings, not an env var) that overlays a signed, freshly-curated free-model
# catalog on top of the release baseline. All four variables below are optional
# and only needed to point the client at a self-hosted/forked feed or
# supporter-key flow instead of the default OmniRoute Radar service. Used by:
# src/lib/radar/sync.ts, src/lib/radar/pinnedKeys.ts, src/lib/radar/links.ts.
# Model used for Telegram chat replies (default: auto/chat).
# TELEGRAM_DEFAULT_MODEL=auto/chat
# Base URL of the Radar feed service. Overrides the built-in default so forks
# and self-hosters can point at their own signed feed.
# RADAR_FEED_URL=https://radar.omniroute.online
# Bot API base URL override (for proxies/self-hosted Bot API servers).
# TELEGRAM_BOT_API_BASE=https://api.telegram.org
# Ed25519 public key (base64-DER SPKI or PEM) used to verify the feed
# signature, replacing the pinned default key. Required when self-hosting a
# feed signed with a different key pair.
# RADAR_FEED_PUBKEY=
# URL the dashboard's "I'm a contributor" button opens (GitHub OAuth
# supporter-key claim flow). No pricing/value lives in this repo — only the
# link.
# RADAR_CONTRIBUTOR_CLAIM_URL=https://radar.omniroute.online/auth/github
# URL the dashboard's "Support the project" button opens (payment/plans
# page). No pricing/value lives in this repo — only the link.
# RADAR_SUPPORTER_PLANS_URL=https://radar.omniroute.online/planos
# Timeout (ms) for outbound Bot API calls (sendMessage/setWebhook).
# TELEGRAM_WEBHOOK_TIMEOUT_MS=60000

46
.gitignore vendored
View File

@@ -72,7 +72,6 @@ yarn-error.log*
# env files (can opt-in for committing if needed)
.env*
!.env.example
!.env.devin-bridge.example
!.env.homolog.example
# Provider API keys (never commit)
*.api-key
@@ -172,6 +171,7 @@ config/quality/test-impact-map.json
# GitNexus local index
.gitnexus
.worktrees
bin/omniroute.mjs
# Consistent with .dockerignore / .npmignore
.omc/
@@ -201,17 +201,12 @@ scripts/i18n/_pending-keys.json
.codegraph/
# Fumadocs generated source
/.source/
# Temporary local worktrees used to build unpublished npm tarballs
/.deploy-build-*/
.source/
# AI agent local settings and configs
.agents/
.antigravitycli/
.claude/
!tests/fixtures/devin-bridge/e2e-workspace/.claude/
!tests/fixtures/devin-bridge/e2e-workspace/.claude/**
# PR Reviews and local feedback files
pr_reviews*.json
@@ -226,6 +221,26 @@ CODEX-SETUP-PROMPT.md
# Quality ratchet — métricas efêmeras (baseline commitado em config/quality/; métricas não)
config/quality/quality-metrics.json
# Electron desktop build output unpacked into the repo root.
# `electron-builder` (squirrel-windows target) unpacks the packaged app — the
# entire Chromium runtime, ~24k files — directly into the repository root.
# Every rule below is ROOT-ANCHORED (leading `/`) on purpose: a bare `locales/`
# or `resources/` would also swallow tracked sources such as the CLI
# translations in `bin/cli/locales/*.json`.
/OmniRoute.exe
/Uninstall OmniRoute.exe
/uninstallerIcon.ico
/locales/
/resources/
/*.pak
/*.dll
/icudtl.dat
/snapshot_blob.bin
/v8_context_snapshot.bin
/vk_swiftshader_icd.json
/LICENSE.electron.txt
/LICENSES.chromium.html
# Runtime logs (diretório local, nunca versionado)
/logs/
-home-diegosouzapw-dev-automações-bots-yt-downloader-20260504 .txt
@@ -238,10 +253,7 @@ omniroute.md
# mise configuration
mise.toml
# release-green artifacts (.gitignore has no inline comments — a trailing
# `# ...` becomes part of the pattern, so it must sit on its own line).
# Already covered by /_*/ above; kept explicit for discoverability.
_artifacts/
_artifacts/ # release-green artifacts
.claude-flow/
# ESLint file cache (npm run lint --cache / complexity ratchets)
@@ -251,8 +263,6 @@ _artifacts/
# CI/local quality artifacts (eslint-results.json, quality-ratchet.md, etc.)
.artifacts/
# Isolated Devin bridge workspaces, evidence, and test databases
.sandbox/
# Homologation E2E suite (npm run homolog) — real-environment credentials + report output
.env.homolog
@@ -260,12 +270,8 @@ tests/homolog/.auth/
tests/homolog/ui/.auth/
homolog-report/
docker-compose.yml.bak
.playwright-cli/
# Playwright screenshot/log output. Today every artifact happens to land inside
# output/**/.playwright-cli/ (covered above), but anything written directly to
# output/ would otherwise show up as untracked.
/output/
# _tasks e um repo git SEPARADO (ver AGENTS.md). _tasks/ (com barra) NAO ignora um
# SYMLINK _tasks; /_tasks (ancorado) cobre symlink/dir na raiz (incidente 2026-08-08).
# _tasks e um repo git SEPARADO (ver AGENTS.md). A linha _tasks/ (com barra) NAO
# ignora um SYMLINK chamado _tasks; /_tasks (ancorado) cobre arquivo/symlink/dir na raiz
# e impede que um git add -A recapture o symlink (incidente 2026-08-08).
/_tasks

View File

@@ -627,7 +627,7 @@ procedures are in [`docs/architecture/QUALITY_GATES.md`](docs/architecture/QUALI
complexity) must not regress vs `quality-baseline.json`. Update via
`npm run quality:ratchet -- --update` when a metric genuinely improves.
- Job `test-vitest` runs `npm run test:vitest` (MCP tools, autoCombo, cache) — blocking.
`test:vitest:ui` is advisory until UI component tests are triaged.
`test:vitest:ui` has been blocking since PR #7127.
**Allowlist policy (short form):** Fix the cause; use the allowlist only for pre-existing
violations you cannot fix in the same PR. Add a comment with justification + issue number.

View File

@@ -93,7 +93,15 @@ RUN --mount=type=cache,id=npm-cache,target=/root/.npm \
# build from 17min to 9min on the same 32-core box. Webpack stays available as the
# escape hatch: `--build-arg`/-e OMNIROUTE_USE_TURBOPACK=0.
# See docs/ops/QUALITY_GATE_PLAYBOOK.md Parte 6.
ENV OMNIROUTE_USE_TURBOPACK=1
#
# Declared as ARG+ENV, not a bare ENV: a bare ENV shadows any same-named ARG for
# the rest of the stage, so `--build-arg OMNIROUTE_USE_TURBOPACK=0` was silently
# ignored and the escape hatch above only ever worked via `-e` at runtime, never
# at build time. Turbopack compiles in native Rust memory that lives outside the
# V8 heap, so OMNIROUTE_BUILD_MEMORY_MB cannot bound it and a memory-constrained
# build host gets SIGKILLed by the cgroup OOM killer with no error message.
ARG OMNIROUTE_USE_TURBOPACK=1
ENV OMNIROUTE_USE_TURBOPACK="${OMNIROUTE_USE_TURBOPACK}"
# Next.js basePath is fixed at build time; pass OMNIROUTE_BASE_PATH here when the
# image should serve under a reverse-proxy subpath without a runtime patch.

69
Makefile Normal file
View File

@@ -0,0 +1,69 @@
.PHONY: help install dev start build build-release lint typecheck typecheck-strict \
test test-unit test-vitest test-coverage test-all test-integration test-e2e \
check check-cycles check-docs env-sync clean
# OmniRoute — convenience wrapper around the npm scripts.
# All targets delegate to the canonical package.json scripts (single source of truth).
help: ## Show this help
@grep -E '^[a-zA-Z_-]+:.*?## .*$$' $(MAKEFILE_LIST) | awk 'BEGIN {FS = ":.*?## "}; {printf " \033[36m%-18s\033[0m %s\n", $$1, $$2}'
install: ## Install dependencies (auto-generates .env from .env.example)
npm install
dev: ## Dev server at http://localhost:20128
npm run dev
start: ## Production server (requires a prior build)
npm run start
build: ## Production build (Next.js 16 standalone)
npm run build
build-release: ## Release build
npm run build:release
lint: ## ESLint (0 errors expected)
npm run lint
typecheck: ## TypeScript check (core)
npm run typecheck:core
typecheck-strict: ## Strict check (no implicit any)
npm run typecheck:noimplicit:core
test: ## Unit tests (Node native runner)
npm run test:unit
test-unit: ## Alias for `test`
npm run test:unit
test-vitest: ## Vitest (MCP server, autoCombo, cache)
npm run test:vitest
test-coverage: ## Unit tests + coverage gate (60/60/60/60)
npm run test:coverage
test-all: ## All suites (unit + vitest + ecosystem + e2e)
npm run test:all
test-integration: ## Integration tests
npm run test:integration
test-e2e: ## E2E (Playwright)
npm run test:e2e
check: ## lint + test combined
npm run check
check-cycles: ## Detect circular dependencies
npm run check:cycles
check-docs: ## Validate documentation (incl. fabricated-docs)
npm run check:docs-all
env-sync: ## Sync .env from .env.example
npm run env:sync
clean: ## Remove build artifacts
rm -rf .build dist coverage .eslintcache

View File

@@ -0,0 +1,2 @@
- Add a default-off connection setting for Codex, OpenAI, and OpenAI-compatible Responses API providers that preserves client-supplied `reasoning.encrypted_content` items for replay, including per-target combo routing.
- Omit opaque encrypted reasoning values from persisted call logs while retaining compact diagnostic markers.

View File

@@ -0,0 +1 @@
- fix(translator): preserve authentic K3 Responses reasoning by model across providers, keep it on the matching assistant turn, and make Kimi Coding prefer client reasoning then cached replay before its empty-marker fallback (#9496)

View File

@@ -0,0 +1 @@
- **fix(providers):** new per-provider `noAuthFallbackDisabledProviders` setting lets operators disable the synthetic anonymous (no-auth) credential fallback for API-key providers whose static definition declares `anonymousFallback: true` (e.g. `opencode-go`, `opencode-zen`) — upstream endpoints now reject anonymous requests with `401 Missing API key`, so the fallback added latency and caused UI health/reconnect churn. Real keyed connections keep working and recover automatically once quota state clears; true no-auth providers (`opencode`, `mimocode`, …) are unaffected, with `blockedProviders` remaining their disable mechanism. Default behavior is unchanged ([#9675](https://github.com/diegosouzapw/OmniRoute/pull/9675))

View File

@@ -0,0 +1 @@
- **fix(docker):** `--build-arg OMNIROUTE_USE_TURBOPACK=0` now reaches the builder stage — a bare `ENV` was shadowing the `ARG`, so the documented webpack escape hatch was silently ignored and memory-constrained hosts were OOM-killed with no error output ([#9695](https://github.com/diegosouzapw/OmniRoute/pull/9695))

View File

@@ -0,0 +1 @@
- fix(db): clear stale combo connection pins when provider connections are deleted (#9719)

View File

@@ -0,0 +1 @@
- fix(compression): persist `enableRenderers` through `normalizeRtkConfig` so RTK renderer settings survive a DB round-trip ([#9730](https://github.com/diegosouzapw/OmniRoute/pull/9730))

View File

@@ -0,0 +1 @@
- Restore Vietnamese locale parity after the entity-normalization sync dropped Radar, provider, and mini-playground messages.

View File

@@ -0,0 +1 @@
- **fix(api):** Model catalogs no longer expose functional gateway mirrors unless the API key permits the mirror's final public model ID ([#9788](https://github.com/diegosouzapw/OmniRoute/pull/9788)) — thanks @xz-dev

View File

@@ -0,0 +1 @@
- **fix(executors):** preserve Command Code usage in `/v1/responses` streams so Codex clients receive real input, output, cache, and reasoning token counts ([#9826](https://github.com/diegosouzapw/OmniRoute/pull/9826)) — thanks @MrShitFox

View File

@@ -0,0 +1 @@
- **fix(executors):** prevent intermittent Codex `upstream_empty_response` errors for tool schemas that combine `oneOf` const branches with a matching sibling `enum` by removing only the semantically redundant `oneOf`; bare, narrowing, non-matching, and type-discriminated `oneOf` schemas remain unchanged. ([#9828](https://github.com/diegosouzapw/OmniRoute/pull/9828))

View File

@@ -0,0 +1 @@
- **fix(cursor):** SelectedImage uses `blobIdWithData` + session blobStore, with JPEG soft-cap prep via sharp ([#9834](https://github.com/diegosouzapw/OmniRoute/pull/9834)) — thanks @yansigit

View File

@@ -0,0 +1 @@
- **fix(test):** reconcile test expectations that drifted from the code they guard on `release/v3.8.50` — auth/vision/provider schema snapshots, and three context-aware combo compatibility assertions that contradicted the same file's own stated contract (catalog-too-small targets stay available as runtime fallback rather than being dropped). The combo assertions were masked by an unresolved import that stopped `combo.ts` from loading at all, so they only become reachable once that import is repaired.

View File

@@ -0,0 +1 @@
- Reconcile the final v3.8.50 bundle-size and file-size ratchets against the measured release tip, preserving exact direction-down ceilings and their source attribution.

View File

@@ -44,6 +44,7 @@
"commander",
"concurrently",
"cross-env",
"cron-parser",
"csv-stringify",
"ctrf",
"dompurify",

View File

@@ -1673,7 +1673,7 @@
},
"tests/unit/base-executor-sanitize-effort.test.ts": {
"@typescript-eslint/no-explicit-any": {
"count": 48
"count": 6
}
},
"tests/unit/batch-deletion.test.ts": {
@@ -3339,4 +3339,4 @@
"count": 5
}
}
}
}

View File

@@ -1,5 +1,9 @@
{
"_rebaseline_2026_08_09_9296_adobe_media_capabilities": "PR #9296 (artickc, fix/adobe-firefly-model-capabilities) own growth: src/app/api/v1/models/catalog.ts 1590->1597 (+7). The image and video catalog serializers now expose the already-normalized Adobe Firefly discovery capability data (media_capabilities, plus the existing video modality/size fields) at their only response-emission chokepoints. The discovery parser and capability normalization remain in open-sse/services/adobeFireflyModels.ts; extracting these seven serialization fields would obscure the catalog contract. Covered by tests/unit/adobe-firefly.test.ts and tests/unit/image-upscale.test.ts.",
"_rebaseline_2026_08_08_v3850_base_drift_batch_9757": "Base drift on release/v3.8.50, not own growth: the 08-06..08-08 merge batches grew 12 already-frozen (or newly-landed) files without carrying their rebaselines — the dedicated rebaseline PR #9616 was closed as 'superseded' but its file-size entries never actually reached the base, and later merges (#8894 combos page, #9539 EditConnectionModal, #8895 models route, #9294/#9293 catalog, #9541 db/core, #8970 tokenHealthCheck, #8925 mcp schemas+server, #8890 accountFallback, #9467 chat.ts, #8931 openai-to-kiro, ProxyRegistryManager) kept growing them. All 12 values re-measured on THIS branch's tree (= pure tip + this PR's 1-line chat.ts fix, which adds zero lines). This PR's own source changes (chat.ts identifier restore, stream.ts format carve-out) do not grow any frozen file past these values.",
"_rebaseline_2026_08_08_migration_135_collision": "fix(db): resolve migration version 135 numbering collision — #9449's 135_connection_runtime_state.sql and #8908's 135_migrate_model_capability_max_token.sql both claimed version 135 (#9449 branched before #8908 merged and never got renumbered before landing on release/v3.8.50), which threw 'Migration version collision detected' the moment ANY code touched the database — a fresh install/deploy from this tip cannot even boot. Renumbered the later-landing file to 140 (next free slot) and added the matching isSchemaAlreadyApplied('140') retroactive guard, matching the established pattern already used for the prior 135/136 -> 137/138 renumber in the same file. Own growth: src/lib/db/migrationRunner.ts 1084->1094 (+10, the new case block) — irreducible, matches the existing per-case guard pattern exactly. Covered by tests/unit/migration-135-numbering-collision.test.ts (2/2), confirmed failing (reproducing the exact live crash) against the pre-fix colliding filenames, passing after.",
"_rebaseline_2026_08_08_9183_reasoning_cache_index_sync": "Extracted fix(responses-api): sync reasoning-cache write index with the fixed read side (from the originally-authored #9183) — chatCore.ts's write side cached every response under a hardcoded messageIndex:0, and translator/index.ts's plain-turn (non-tool-call) cache-key lookup ALSO still hardcoded messageIndex 0 at its call site (a second, previously-undiscovered instance of the same hardcoding bug, found while re-verifying this fix against the current upstream tip — the two never agreed once a conversation went past its first assistant turn, so DeepSeek/Xiaomi-mimo plain-turn reasoning replay silently missed the cache). Own growth: open-sse/handlers/chatCore.ts 5034->5042 (+8, computing messageIndex from the incoming request's message count at both the streaming and non-streaming cache-write call sites) — irreducible call-site wiring. Covered by tests/unit/reasoning-cache.test.ts (new end-to-end write/read regression test, rebaselined below) and tests/unit/translator-helper-branches.test.ts fixture updates. Other #9183 sub-fixes (output_index collision prevention, reasoning-content-alias generalization) were originally assumed already superseded by upstream's own independent fix — a live incident 2026-08-08 disproved that for the message-vs-tool-call collision case specifically (fixed separately in #9822); not re-extracted here since this PR's own scope is the narrower messageIndex sync only.",
"_rebaseline_2026_08_02_9259_rolling_rpm": "PR #9259 (issue #8733) own growth: open-sse/services/rateLimitManager.ts baseline 1060->1167 (+107; final source 1153). The existing withRateLimit chokepoint now composes process-local rolling RPM leases with Bottleneck admission, releases pre-dispatch leases on queue timeout/abort/connection disable, preserves caller abort reasons, and wires 429/header state into the extracted rollingRpmGate.ts. The remaining growth is irreducible lifecycle wiring at the dispatch boundary plus the real watchdog test hooks needed to verify queued-wedge recovery; moving it further would obscure lease ownership and Bottleneck cleanup. Covered by the focused rate-limit manager/sliding-window suite (33/33); distributed multi-instance coordination remains explicitly out of scope.",
"_rebaseline_2026_07_24_8470_hyperagent_sticky_thread": "PR #8470 (artickc, fix/hyperagent-tool-loop-thread-sticky) own growth: open-sse/executors/hyperagent.ts 936->1025 (wc -l; check-file-size.mjs counts via split(\"\\n\").length so the gate sees 937->1026, +89, crosses the 1000 cap). Fixes a real bug where a reverse-conversion proxy (text-Intent/JSON to Claude Code native tool_calls) rewrites assistant messages between agentic tool-loop turns, breaking HyperAgents conversation-prefix fingerprint and cold-starting the thread mid tool-loop. Adds Anthropic tool_use/tool_result flattening to extractMessageText() plus a new rootUserFingerprint()/root-key lookup tier in resolveHyperAgentThreadBinding()/storeHyperAgentThreadAfterTurn() so the thread stays sticky across the tool loop. Cohesive additions inside the existing single-file executor; not extractable without splitting the executor mid-request-flow. Covered by tests/unit/executor-hyperagent.test.ts (19/19, +5 new cases for tool_result/tool_use flattening + root-key stickiness). Pre-merge review flagged a cross-conversation root-key collision risk (tracked in the PRs own mandatory pre-merge checklist, not yet addressed) — unrelated to this file-size ratchet, tracked separately by /fix-prs.",
"_rebaseline_2026_07_25_8494_capability_filter_fail_closed": "PR #8494 (fix/capability-filters-fail-closed, #8488) own growth: open-sse/services/combo.ts 3640->3693 (+53) adds a fail-closed guard after filterTargetsByRequestCompatibility() — when every eligible target is excluded by request-capability filtering (vision/tools/etc) instead of quota/health, the combo now returns an explicit `capability_mismatch` 400 (describeCapabilityFilterExhaustion, imported from combo/comboStructure.ts) rather than silently falling through to a generic no-targets error, plus a `compatFilterFailOpen` escape hatch (combo config OR settings) mirrored at both the main/auto and round-robin call sites for symmetry. combo/comboStructure.ts (previously under cap, un-frozen) grows 794->918 (+124) — new home for describeCapabilityFilterExhaustion + providerSupportsEmulatedToolCalling (#5240 emulated tool-calling exemption so fail-closed does not regress prompt-emulation-only combos like all-chatgpt-web). Irreducible orchestration wiring at the existing filter chokepoint (same precedent as #7301's universal-cooldown-retry generalization). Companion test tests/unit/combo-routing-engine.test.ts 3409->3449 (+40, fail-closed/fail-open coverage across both call sites) also rebaselined. Covered by tests/unit/8488-capability-filter-fail-closed.test.ts (new) + 95/95 passing across both files. Structural shrink of combo.ts tracked in #3501.",
"_rebaseline_2026_07_25_8499_ts7_result_union_predicates": "PR #8499 (backryun, chore/ts7-types-executor-scattered) own growth: muse-spark-web.ts 1396->1405 (+9, irreducible). Under this workspace's `strictNullChecks: false`, the boolean-literal discriminant on `GraphqlResult` (`{ ok: true } | { ok: false; error: string }`) narrows the positive `.ok===true` branch but leaves `!result.ok` at the full union under TS7, making `.error` unreachable to the checker at the two call sites (warmup, mode-switch). Fixed by adding a single `isGraphqlFailure()` type-predicate helper (doc comment + 3-line body) reused at both call sites instead of duplicating the predicate inline — not extractable to a shared module without splitting a single-file executor's local narrowing helper out of its own file. Covered by the existing muse-spark-web executor test suite (no behavior change, pure narrowing fix).",
@@ -162,6 +166,7 @@
"cap": 1000,
"testCap": 1000,
"testFrozen": {
"tests/unit/reasoning-cache.test.ts": 1035,
"_rebaseline_2026_06_27_5193_antigravity_test": "#5193 own test growth: oauth-providers-config.test.ts 870->873 (+3: antigravity projectId assertion + 50ms tick for the now fire-and-forget onboarding, matching the no-PKCE/no-openid flow).",
"_rebaseline_2026_07_02_5928_base_red": "web-cookie-providers-new.test.ts 845->850: #5928 (test(security) Kimi Web URL host parse, CodeQL #689) grew the file +5 lines and merged into release/v3.8.44 WITHOUT rebaselining, leaving a fast-gates base-red that blocked every subsequent PR->release. Test growth is legitimate (a security regression test); maintainer absorbs the drift here. Frozen at 850.",
"_rebaseline_2026_07_09_6126_clinepass_dualauth": "#6126 (ClinePass dual-auth) own test growth: oauth-providers-config.test.ts 842->845 (+3: clinepass key/config/required-fields entries reusing the Cline WorkOS flow config, needed after registering clinepass in the oauth.ts PROVIDERS enum).",
@@ -350,7 +355,7 @@
"open-sse/executors/deepseek-web.ts": 1148,
"open-sse/executors/grok-web.ts": 1044,
"open-sse/executors/muse-spark-web.ts": 1405,
"open-sse/handlers/chatCore.ts": 5034,
"open-sse/handlers/chatCore.ts": 5042,
"open-sse/handlers/imageGeneration.ts": 3101,
"open-sse/handlers/responseSanitizer.ts": 1128,
"open-sse/handlers/search.ts": 1536,
@@ -364,7 +369,7 @@
"open-sse/services/combo.ts": 3648,
"open-sse/services/compression/strategySelector.ts": 1060,
"open-sse/services/rateLimitManager.ts": 1167,
"open-sse/translator/response/openai-responses.ts": 1204,
"open-sse/translator/response/openai-responses.ts": 1224,
"open-sse/utils/cursorAgentProtobuf.ts": 1505,
"open-sse/utils/stream.ts": 2889,
"src/app/(dashboard)/dashboard/HomePageClient.tsx": 1388,
@@ -388,7 +393,7 @@
"src/app/(dashboard)/dashboard/usage/components/EvalsTab.tsx": 2148,
"src/app/(dashboard)/dashboard/usage/components/ProviderLimits/index.tsx": 1119,
"src/app/api/providers/[id]/models/route.ts": 2361,
"src/app/api/v1/models/catalog.ts": 1597,
"src/app/api/v1/models/catalog.ts": 1590,
"src/lib/db/apiKeys.ts": 1529,
"src/lib/db/core.ts": 1639,
"src/lib/db/migrationRunner.ts": 1094,
@@ -402,7 +407,7 @@
"src/shared/components/analytics/charts.tsx": 1035,
"src/shared/services/cliRuntime.ts": 1122,
"src/sse/handlers/chat.ts": 1904,
"src/sse/services/auth.ts": 2520,
"src/sse/services/auth.ts": 2508,
"tests/unit/account-fallback-service.test.ts": 1572,
"tests/unit/provider-validation-specialty.test.ts": 2985,
"open-sse/executors/hyperagent.ts": 1026,
@@ -421,9 +426,6 @@
"_rebaseline_2026_07_28_8863_firefly_detail_level": "PR #8863 (fix/adobe-firefly-gpt-detail-level-max) own growth: adobeFireflyClient.ts 2317->2322 (+5 = gpt-image detailLevel defaulting to maximal at the existing payload-build site). Covered by tests/unit/adobe-firefly.test.ts.",
"_rebaseline_2026_07_29_8281_home_quickstart_prefetch": "Release v3.8.49 base-red fix (no PR — captain sweep): src/app/(dashboard)/dashboard/HomePageClient.tsx 1377->1381 (+4). #8292 added prefetch={false} to the sidebar but left /home's five quick-start Links prefetching, so first paint still fired 12 speculative RSC requests — caught by navigation.spec.ts only after the e2e helper bug (APP_ROUTE_PATTERN missing /home) was repaired in the same cycle. Growth is the five prefetch attributes; it was offset first by extracting the repeated className literals (INLINE_LINK x4, DOCS_LINK x1), which collapsed five wrapped <Link> blocks back to one line each — a naive fix measured 1391. Guard: tests/unit/sidebar-prefetch-policy-8281.test.ts.",
"_rebaseline_2026_08_02_v3850_agentrouter_responses": "Release v3.8.50 AgentRouter/Codex compatibility reconciliation. open-sse/executors/base.ts 1562->1578: #9190 wires AgentRouter's selected Claude/OpenAI/Responses protocol through the existing executor URL, auth, identity-header and fingerprint chokepoints; the reusable alternate resolver remains outside base.ts. open-sse/utils/stream.ts 2887->2889: #9213 evaluates Responses ID and usage normalization independently so response.completed always receives finite usage.total_tokens instead of short-circuiting after an ID rewrite. tests/unit/chatcore-translation-paths.test.ts 2769->2776: #9191 updates the existing Claude-Code bridge assertions for the dynamic AgentRouter wire image. PR #9224 offsets its own chatCore growth by extracting the AgentRouter protocol decisions into chatCore/agentRouterProtocol.ts, leaving chatCore below its frozen ceiling. Covered by agentrouter executor/chatCore protocol tests, chatcore translation-path tests, and responses-commentary-passthrough tests.",
"_rebaseline_2026_08_08_v3850_base_drift_batch_9757": "Base drift on release/v3.8.50, not own growth: the 08-06..08-08 merge batches grew 12 already-frozen (or newly-landed) files without carrying their rebaselines — the dedicated rebaseline PR #9616 was closed as 'superseded' but its file-size entries never actually reached the base, and later merges (#8894 combos page, #9539 EditConnectionModal, #8895 models route, #9294/#9293 catalog, #9541 db/core, #8970 tokenHealthCheck, #8925 mcp schemas+server, #8890 accountFallback, #9467 chat.ts, #8931 openai-to-kiro, ProxyRegistryManager) kept growing them. All 12 values re-measured on THIS branch's tree (= pure tip + this PR's 1-line chat.ts fix, which adds zero lines). This PR's own source changes (chat.ts identifier restore, stream.ts format carve-out) do not grow any frozen file past these values.",
"_rebaseline_2026_08_08_migration_135_collision": "fix(db): resolve migration version 135 numbering collision — #9449's 135_connection_runtime_state.sql and #8908's 135_migrate_model_capability_max_token.sql both claimed version 135 (#9449 branched before #8908 merged and never got renumbered before landing on release/v3.8.50), which threw 'Migration version collision detected' the moment ANY code touched the database — a fresh install/deploy from this tip cannot even boot. Renumbered the later-landing file to 140 (next free slot) and added the matching isSchemaAlreadyApplied('140') retroactive guard, matching the established pattern already used for the prior 135/136 -> 137/138 renumber in the same file. Own growth: src/lib/db/migrationRunner.ts 1084->1094 (+10, the new case block) — irreducible, matches the existing per-case guard pattern exactly. Covered by tests/unit/migration-135-numbering-collision.test.ts (2/2), confirmed failing (reproducing the exact live crash) against the pre-fix colliding filenames, passing after.",
"_rebaseline_2026_08_02_9259_rolling_rpm": "PR #9259 (issue #8733) own growth: open-sse/services/rateLimitManager.ts baseline 1060->1167 (+107; final source 1153). The existing withRateLimit chokepoint now composes process-local rolling RPM leases with Bottleneck admission, releases pre-dispatch leases on queue timeout/abort/connection disable, preserves caller abort reasons, and wires 429/header state into the extracted rollingRpmGate.ts. The remaining growth is irreducible lifecycle wiring at the dispatch boundary plus the real watchdog test hooks needed to verify queued-wedge recovery; moving it further would obscure lease ownership and Bottleneck cleanup. Covered by the focused rate-limit manager/sliding-window suite (33/33); distributed multi-instance coordination remains explicitly out of scope.",
"_rebaseline_2026_07_25_dario_upstream_proxy_selector": "PR #8523 (Dario embedded service): upstream-proxy mode selector replaces the binary CLIProxyAPI toggle with Native/CLIProxyAPI/Dario/Fallback + a fallback-backend picker. ProviderDetailPageClient.tsx 798->804 (+6, new hook fields threaded through to ConnectionsListPanel), ConnectionRow.tsx 942->958 (+16, the mode <select> + conditional fallback-backend <select> replacing a single pill button), useProviderConnections.ts 954->986 (+32, upstreamProxyMode/upstreamProxyFallbackBackend state + handleSetUpstreamProxyMode, handleToggleCliproxyapiMode kept as a thin backward-compat wrapper for the existing hook-shape test). All additive UI/state for the new modes — no unrelated refactor.",
"_rebaseline_2026_08_02_9242_token_health_transient": "PR #9242 (fix/refresh-circuit-transient): src/lib/tokenHealthCheck.ts 1021 (new file, above cap 1000). The file consolidates token-refresh health checking logic that was previously scattered across auth.ts and tokenRefresh.ts. Cohesive single-responsibility module for refresh circuit state management; not extractable without splitting the refresh state machine. Covered by tests/unit/tokenHealthCheck-transient.test.ts.",
"_rebaseline_2026_07_28_8870_firefly_ref_cap_timeout": "PR #8870 (fix/adobe-firefly-gpt-ref-cap-timeout) own growth: adobeFireflyClient.ts 2322->2385 (+63 = gpt-image subject-ref hard cap at 2 + adaptive poll timeout budget (base 300s + 60s/ref, max 600s) + defensive .slice on referenceBlobs for gpt/nano/generic families). Fixes live 504s on multi-screenshot listing jobs (Featured Promo / Box Art) where 34+ subject refs stall colligo until the old 180s poll budget expires. Helpers adobeFireflyMaxImageRefs/adobeFireflyImageTimeoutMs live next to the existing payload/poll chokepoint (not extractable without splitting the wire recipe mid-PR). Covered by tests/unit/adobe-firefly.test.ts (ref-cap + timeout cases). Structural shrink tracked in #3501.",
@@ -435,6 +437,7 @@
"_rebaseline_2026_08_06b_v3850_sweepreds_drift": "Segunda reconciliacao de 2026-08-06 (/sweep-reds sobre o tip puro 2ddbbc61a6): 3 arquivos voltaram a passar do frozen apos os merges do mesmo dia, com atribuicao 1:1 por commit. (1) src/app/(dashboard)/dashboard/providers/page.tsx 1928->1944 e (2) open-sse/executors/base.ts 1635->1640, ambos do #9515 (feat(radar): flag-gated signed free-model catalog overlay, commit e7f6b1d130) — o overlay do Radar entra por wiring nos chokepoints ja existentes (a resolucao/verificacao do catalogo assinado mora fora destes dois arquivos); +16 e +5 linhas liquidas nao sao extraiveis sem inventar um leaf por callsite. (3) open-sse/services/accountFallback.ts 1966->1972 do #8704 (commit c4527f97bd), +6 linhas de dados em CREDITS_EXHAUSTED_SIGNALS ('has been exhausted', fixes #8631). src/sse/handlers/chat.ts 1880>1877 tambem estava violando e NAO entra aqui de proposito: e drenado por encolhimento na PR #9598, sem rebaseline. Crescimento proprio DESTA PR: src/lib/db/migrationRunner.ts 1077->1084 (+7) — o guard retroativo em isSchemaAlreadyApplied para os arquivos renumerados 137/138, exigido pela propria mensagem de erro de colisao do runner (ambas as migracoes sao ALTER TABLE ADD COLUMN puro, nao idempotente). Dois `case` + dois `return hasColumn(...)` + 3 linhas de comentario dentro do switch existente; nao extraivel.",
"_rebaseline_2026_08_06c_v3850_sweepreds_pr2": "Segunda PR do /sweep-reds (fix/release-v3.8.50-basereds-0806b): tests/unit/provider-models-route.test.ts 1784->1787 (medido pelo gate, que conta split(\"\\n\").length) (+2 apos compressao de comentarios) — alinhamento de contrato forcado por dois merges do dia: #9106 tornou gemini-3.1-pro-high user-callable (a entry do alias entra na lista esperada do teste de discovery-retry, +1 linha de dado + 1 de comentario) e ff012ff420 adicionou onboardUser como bootstrap hop (exclusao no mock, ja comprimida a 1 linha). Nao ha o que encolher sem apagar o comentario que explica o porque.",
"_rebaseline_2026_08_07_v3850_sweepreds_pr2_toolnamemap": "tests/unit/translator-openai-to-gemini.test.ts 1616->1619 (+3). O frozen estava EXATAMENTE no tamanho da base, entao qualquer linha nova viola. #9568 (c9a3361e5a) fez buildChangedToolNameMap emitir entradas IDENTIDADE (o Gemini minusculiza nomes de tool nas respostas, entao o tradutor de resposta precisa da chave para mapear de volta), o que passou a incluir `_toolNameMap` no envelope Antigravity de qualquer request com tools. As 3 linhas sao: a chave nova na lista esperada de Object.keys, 1 comentario explicando POR QUE ela aparece (sem ele o proximo leitor tenta remove-la de novo) e 1 assert do CONTEUDO do map — presenca de chave sozinha nao provaria a entrada identidade, que e justamente o comportamento novo. Nao ha o que extrair: e alinhamento de contrato dentro de um teste existente.",
"_rebaseline_2026_08_08_9634_migration_139_guard": "PR #9634 (fix/release-v3850-basereds) own growth, re-measured on e0ce95c59 after rebase: src/lib/db/migrationRunner.ts 1094->1096 (+2, the isSchemaAlreadyApplied case-139 retroactive guard for the renumbered ccr migration). Irreducible, matches the per-case guard pattern exactly. Covered by tests/unit/migration-135-numbering-collision.test.ts.",
"_rebaseline_2026_06_22_4644_deepseek_web_tools": "PR #4644 (BugsBag/robust deepseek-web tool-call parsing): open-sse/executors/deepseek-web.ts 1117->1125 (+8). The new agentic tool-call path emits surrounding text + reasoning before tool_calls and swaps to the dedicated deepseekWebTools.ts parser; the +8 lines are cohesive wiring at the existing transformSSE chokepoint (the parser itself lives in the new deepseekWebTools.ts file, already under cap). The PR's own fast-gate (PR->release) does not run check:file-size, so this surfaced only at release reconcile. Covered by tests/unit/deepseek-web-tools-variants.test.ts + deepseek-web-tools-execute.test.ts.",
"_rebaseline_2026_06_23_4712_deepseek_web_tool_results": "PR for #4712 (deepseek-web drops role:tool): open-sse/executors/deepseek-web.ts 1125->1148 (+23). messagesToPrompt() now folds role:\\\"tool\\\" results into the single-prompt transcript (recovering the tool name from the preceding assistant tool_calls by tool_call_id) instead of silently dropping them; the lines are cohesive wiring inside the existing function. Covered by tests/unit/deepseek-web-tool-result-prompt-4712.test.ts.",
"_rebaseline_2026_06_24_headroom_strategy": "Headroom-aware connection selection (dario technique): combo.ts 3168->3180 (+12 = a new `else if (strategy === \\\"headroom\\\")` dispatch branch in handleComboChat that delegates to orderTargetsByHeadroom + its log line, plus the import). The actual logic lives OUT of the god-file: the pure ranker rankByHeadroom/computeHeadroom is the new leaf open-sse/services/combo/headroomRanking.ts (91 LOC, <cap) and the async orderer orderTargetsByHeadroom is appended to the existing open-sse/services/combo/quotaStrategies.ts (<cap) next to its sibling reset-aware/reset-window orderers (reuses their connection-expansion machinery). headroom = 1 - max(util_5h, util_7d) from getSaturation (src/lib/quota/saturationSignals.ts), prefers the connection with the most free capacity. Only the dispatch wiring is irreducible at the existing combo strategy chokepoint (mirrors the reset-aware/reset-window/context-optimized branches); not extractable without hiding the call site. fill-first stays default; all existing strategies untouched. Covered by tests/unit/combo-headroom-ranking.test.ts (pure helper) + tests/unit/combo-headroom-strategy.test.ts (orderer, saturation injected). Structural shrink of combo.ts tracked in #3501.",
@@ -541,7 +544,7 @@
"src/lib/tokenHealthCheck.ts": "1053",
"src/lib/db/apiKeys.ts": "1529",
"src/lib/db/core.ts": "1639",
"src/lib/db/migrationRunner.ts": "1094",
"src/lib/db/migrationRunner.ts": "1096",
"src/lib/db/models.ts": "1097",
"src/lib/db/providers.ts": "1034",
"src/lib/memory/retrieval.ts": "1073",
@@ -560,5 +563,9 @@
"open-sse/executors/kiro.ts": "1069",
"open-sse/translator/request/openai-to-kiro.ts": "1057",
"open-sse/utils/sseHeartbeat.ts": "142",
"_rebaseline_2026_08_04_9305_sse_comments": "#9305 fix: broadened sseCommentsEnabled()"
"_rebaseline_2026_08_04_9305_sse_comments": "#9305 fix: broadened sseCommentsEnabled()",
"_rebaseline_2026_08_09_v3850_release_close": "Release v3.8.50 close reconciliation on e0ce95c592: src/sse/handlers/chat.ts 1904->1918 is the irreducible request-pipeline wiring from #9759 that invokes the Modality Bridge guardrail without moving its implementation into the handler; covered by the 17 Vision Bridge canaries plus the PR-1 focused suite. open-sse/translator/response/openai-responses.ts 1204->1215 is #9168's Responses tool-call argument delta buffering/normalization at the existing translator state-machine chokepoint; covered by its dedicated translator regression tests. Both values are measured by check:file-size (split-newline semantics), and the gate remains frozen at the new exact sizes.",
"_rebaseline_2026_08_08_toolcall_message_index_collision": "fix(responses-api): tool call after a text message collided on the same output_index. own growth: open-sse/translator/response/openai-responses.ts 1204->1224 (+20, extracted toolCallOutputIndexBase() shared helper so emitToolCall/closeToolCall can no longer compute a tool call's output_index independently and collide with a text message emitted in the same turn). Live incident (2026-08-08, OpenClaw agent): a client that tracks response items by output_index saw the tool call's added/delta/done events land on an index it had already marked complete (the just-closed text message), and silently dropped them — the agent spoke its preamble and never executed the tool call, even though OmniRoute's own recorded responseBody had a complete, valid tool_calls entry. Covered by the new regression test in tests/unit/translator-resp-openai-responses.test.ts reproducing the exact live scenario.",
"_rebaseline_2026_08_03_9255_adobe_firefly_durable_sessions": "PR #9255 own cohesive growth: open-sse/services/adobeFireflyClient.ts 2322->2894 adds authenticated-vs-guest IMS classification, browser-risk ARP validation/rebuild, bounded 408 retry/recovery, sticky accepted-session handling, and matching image/video submit recovery at the existing Adobe upstream client chokepoints. This client was already explicitly frozen as a single self-contained upstream integration by #8006/#8510; splitting only the retry/auth helpers now would scatter one request state machine while structural shrink remains tracked in #3501. tests/unit/adobe-firefly.test.ts 871->1136 adds direct regression coverage for guest-token rejection, cookie/ARP rebuilding, 408 retries, sticky accepted ARP reuse, forced auth recovery, and cookie-to-IMS exchange. The obsolete 1179-line managed-Chrome fallback module was deleted rather than rebaselined after the packaged-safe pure-CDP path became authoritative. Focused Adobe suite: 61/61.",
"_rebaseline_2026_08_07_9653_disconnect_grace_period": "Extracted fix(sse): grace period before finalizing a client disconnect as 499 (#9653) — a client that closes its connection right after reading a fully-completed SSE stream can race OmniRoute's own completion bookkeeping, getting persisted as a false 499/0-tokens even though it delivered the full response (live-confirmed: a real disconnect at 18236ms was corrected to 200/82814+1292 tokens). Own growth: open-sse/handlers/chatCore.ts 5030->5039 (+9, wiring createClientDisconnectGraceHandler at the existing onClientDisconnectFinalize call site) — irreducible call-site wiring, the actual grace-period logic lives in the new leaf createClientDisconnectGraceHandler (open-sse/utils/streamFailureFinalization.ts, not frozen). Re-measured to 5042 after rebasing onto a newer release/v3.8.50 tip: the file carries an unrelated +3 base drift from already-merged upstream commits between this PR's original branch point and the rebase target, not covered by this entry. Covered by tests/unit/stream-disconnect-grace-period-9653.test.ts (4/4, fake-timer driven). Other file-size gate violations present on this base tip are pre-existing/unrelated to this change (base-red #9679, re-verify current issue number at merge time)."
}

View File

@@ -177,12 +177,13 @@
"dedicatedGate": true
},
"bundleSize": {
"value": 7666,
"value": 8045,
"direction": "down",
"dedicatedGate": true,
"_rebaseline_2026_07_07_v3846_release_close": "5601->6534 (+933). v3.8.46 release close: gzip of the 4 bin/*.mjs entrypoints (size-limit + @size-limit/file) grew from this cycle's feature/fix merges pulled transitively into the CLI entrypoints (new providers, combo pipeline strategy #6396, effort/thinking standardization #6241, catalog cache-invalidation #6408). Measured 6534 locally via `check:bundle-size --ratchet` (deterministic gzip, matches CI). Legitimate cycle growth; shrink is separate debt.",
"_rebaseline_2026_07_19_7808_codeql_alias_resolver_hook": "6534->6762 (+228). PR #7808 (CodeQL js/incomplete-url-substring-sanitization fix): the ESM loader hook source moved out of the inline `HOOK_SOURCE` template literal in bin/aliasResolver.mjs into a real file bin/aliasResolverHook.mjs, loaded via pathToFileURL() instead of a dynamically-built `data:text/javascript,...` URL. The new file is now counted by size-limit as a 5th bin/*.mjs entrypoint. Net +228 = the hook's gzip size (previously hidden inside aliasResolver.mjs because the template literal was compressed away). Security-driven; no shrink opportunity.",
"_rebaseline_2026_07_28_v3849_release_preflight": "6762 -> 7666 (+904). Fechamento do ciclo v3.8.49: gzip dos entrypoints bin/*.mjs (size-limit + @size-limit/file) cresceu com o que os merges do ciclo puxam transitivamente para o CLI (novos provedores — 271->290, seletor de protocolo por conexão #8861, catálogos de busca #8814, resiliência). Crescimento legítimo de ciclo, medido localmente com `npm run check:bundle-size` = 7666 (gzip determinístico, bate com o CI). Encolher é dívida separada."
"_rebaseline_2026_07_28_v3849_release_preflight": "6762 -> 7666 (+904). Fechamento do ciclo v3.8.49: gzip dos entrypoints bin/*.mjs (size-limit + @size-limit/file) cresceu com o que os merges do ciclo puxam transitivamente para o CLI (novos provedores — 271->290, seletor de protocolo por conexão #8861, catálogos de busca #8814, resiliência). Crescimento legítimo de ciclo, medido localmente com `npm run check:bundle-size` = 7666 (gzip determinístico, bate com o CI). Encolher é dívida separada.",
"_rebaseline_2026_08_09_v3850_release_close": "7666 -> 8045 (+379 gzip bytes, +4.9%). Release v3.8.50 close reconciliation measured twice with the real size-limit + @size-limit/file path on tip e0ce95c592. Per-entry measurements remain below their absolute budgets: omniroute.mjs 4380/15000, mcp-server.mjs 1195/5000, nodeRuntimeSupport.mjs 887/8000, reset-password.mjs 1583/6000. The growth accumulated through legitimate CLI/runtime work in this cycle, including global-install ESM alias resolution, Termux cache preparation, and MCP stdio startup hardening; no entrypoint is near its absolute ceiling. The direction:down ratchet stays blocking from this exact measured tip."
},
"openapiBreaking": {
"value": 0,

View File

@@ -1,24 +1,10 @@
{
"_comment": "Catraca de test-discovery (check-test-discovery.mjs). Cada entrada e um arquivo de teste que NENHUM runner coleta (ele nunca roda) — divida congelada na auditoria 6A.1 (2026-06-09; 195 originais, 135 religados no node runner em 6A.1c). So pode DIMINUIR: religue o teste (ajustando o glob do runner ou movendo o arquivo) e remova a entrada via --update. NAO adicione novos orfaos — corrija o runner.",
"_remaining_60": "Categorias: 33 .test.tsx de tests/unit (religaveis via vitest.config root, MAS o experimento 2026-06-09 mostrou 24 arquivos vermelhos — triagem de drift de UI na janela 2026-06-16, junto com os 14 fails do proprio test:vitest:ui atual); 9 open-sse __tests__ + 8 src __tests__ (includes de vitest.config que NENHUM script executa sem filtro); 4 golden-set + 1 benchmarks + 1 live + 1 stress (deliberadamente manuais — decidir runner/gating); 3 integration/services (gated RUN_SERVICES_INT=1, sem runner CI).",
"_remaining_13": "13 orfaos restantes: 2 testes de API em settings + 1 snapshot de quota do DB; 4 golden-set + 1 benchmark + 1 teste live + 1 stress (deliberadamente manuais — decidir runner/gating); 3 integration/services (gated RUN_SERVICES_INT=1, sem runner CI).",
"orphans": [
"open-sse/services/__tests__/chatgptTlsClient.test.ts",
"open-sse/services/__tests__/claudeTlsClient.test.ts",
"open-sse/services/__tests__/grokTlsClient.test.ts",
"open-sse/services/__tests__/manifestAdapter.test.ts",
"open-sse/services/__tests__/specificityDetector.test.ts",
"open-sse/services/__tests__/tierResolver.test.ts",
"open-sse/services/__tests__/volumeDetector.test.ts",
"open-sse/translator/helpers/__tests__/maxTokensHelper.test.ts",
"open-sse/translator/helpers/__tests__/schemaCoercion.test.ts",
"src/app/api/settings/__tests__/memory.test.ts",
"src/app/api/settings/__tests__/settings.test.ts",
"src/lib/db/__tests__/quotaSnapshots.test.ts",
"src/lib/memory/__tests__/injection.test.ts",
"src/lib/memory/__tests__/qdrant-wiring.test.ts",
"src/lib/memory/__tests__/retrieval.test.ts",
"src/lib/memory/__tests__/schemas.test.ts",
"src/lib/skills/__tests__/integration.test.ts",
"tests/benchmarks/pipeline-accuracy.test.ts",
"tests/golden-set/compression-caveman-v2.test.ts",
"tests/golden-set/compression-quality.test.ts",
@@ -28,36 +14,6 @@
"tests/integration/services/full-lifecycle.int.test.ts",
"tests/integration/services/route-guard-services.int.test.ts",
"tests/live/deepseek-web-live.test.ts",
"tests/theoldllm-stress.test.ts",
"tests/unit/AutoComboCatalog.test.tsx",
"tests/unit/SkillsConceptCard.test.tsx",
"tests/unit/agent-skills-page.test.tsx",
"tests/unit/dashboard/batch/components/BatchDetailModal.test.tsx",
"tests/unit/dashboard/batch/components/ExpirationBadge.test.tsx",
"tests/unit/dashboard/batch/components/NewBatchWizard.test.tsx",
"tests/unit/dashboard/batch/components/ProgressBarBicolor.test.tsx",
"tests/unit/dashboard/batch/components/UploadFileModal.test.tsx",
"tests/unit/dashboard/batch/components/useBatchActions.test.tsx",
"tests/unit/dashboard/batch/concept-cards.test.tsx",
"tests/unit/dashboard/batch/list-regression.test.tsx",
"tests/unit/dashboard/batch/sanitization.test.tsx",
"tests/unit/omni-skills-page.test.tsx",
"tests/unit/shared-clipboard.test.tsx",
"tests/unit/shared/components/AutoRoutingBanner.test.tsx",
"tests/unit/shared/components/KiroAuthModal.test.tsx",
"tests/unit/shared/components/ProxyConfigModal.test.tsx",
"tests/unit/translator-friendly-advanced-section.test.tsx",
"tests/unit/translator-friendly-compression.test.tsx",
"tests/unit/translator-friendly-concept-card.test.tsx",
"tests/unit/translator-friendly-integration.test.tsx",
"tests/unit/translator-friendly-monitor-tab.test.tsx",
"tests/unit/translator-friendly-page-client.test.tsx",
"tests/unit/translator-friendly-pipeline-view.test.tsx",
"tests/unit/translator-friendly-raw-json-panel.test.tsx",
"tests/unit/translator-friendly-result-narrated.test.tsx",
"tests/unit/translator-friendly-simple-controls.test.tsx",
"tests/unit/translator-friendly-stream-transformer.test.tsx",
"tests/unit/translator-friendly-test-bench.test.tsx",
"tests/unit/translator-friendly-translate-tab.test.tsx"
"tests/theoldllm-stress.test.ts"
]
}

View File

@@ -186,10 +186,10 @@ Runs on pull requests only.
Runs after `build`. Blocks merge on failure.
| Suite | Validates | Blocking |
| ---------------- | ------------------------------------------------------- | -------------------------------------------------------------------------- |
| `test:vitest` | MCP server (94 tools), autoCombo, cache — vitest runner | Yes |
| `test:vitest:ui` | UI component tests — vitest runner | **Advisory** (`continue-on-error: true`) — failing until Fase 6A UI triage |
| Suite | Validates | Blocking |
| ---------------- | ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| `test:vitest` | MCP server (94 tools), autoCombo, cache — vitest runner | Yes |
| `test:vitest:ui` | UI component tests — vitest runner | **Blocking** — pre-existing failures are explicitly excluded in `vitest.config.ts`; new failures fail the job |
### Nightly workflows (scheduled, advisory)
@@ -401,7 +401,7 @@ several "obvious" merges turned out to hide debt and are **not** clean drop-ins.
- `check:openapi-security-tiers` (advisory) — ❌ **NOT cleanly flippable.** It exits 0 but warns that several `traffic-inspector` routes under `LOCAL_ONLY_API_PREFIXES` lack the `x-loopback-only: true` annotation. Enforcing it requires adding those annotations to `openapi.yaml` first.
- `typecheck:noimplicit:core` (advisory) — largely subsumed by the blocking `check:type-coverage` ratchet. Flip to a ratchet or drop the redundant second `tsc` pass.
- `test:vitest:ui` (advisory, 14 parked fails) — fix-and-block or delete; don't leave rotting.
- `test:vitest:ui` (now **blocking**) — pre-existing failures are explicitly excluded in `vitest.config.ts` with `// #8618` tracking comments; new failures fail the job.
- `check:secrets` (gitleaks, blocking ratchet frozen at 3 documented false-positives) — allowlist the 3 to reach 0, or demote to advisory. Overlaps GitHub native secret-scanning + `check:public-creds`.
- `check:pr-evidence` (blocking, greps PR-body prose) — high false-positive risk; weakens Hard Rule #18 enforcement if dropped, so this is a genuine policy call.
- `semgrep` (advisory standalone) — overlaps CodeQL for the OWASP families; wire its baseline to a ratchet or drop.

View File

@@ -147,11 +147,11 @@ The prod stack runs in parallel with the dev compose (different container names,
The repository ships a multi-stage Dockerfile (`Dockerfile`). Three stages are exposed; pick the right `target` for your use case.
| Stage | Base image | Purpose |
| ------------- | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `builder` | `node:24.15.0-trixie-slim` | Installs deps (`npm ci --legacy-peer-deps`) and runs `npm run build -- --webpack` |
| `runner-base` | `node:24.15.0-trixie-slim` | Production runtime with the Next.js standalone output. **No provider CLIs bundled.** |
| `runner-cli` | `runner-base` | Adds `git`, `docker.io`, `docker-compose` and global CLIs: `@openai/codex`, `@anthropic-ai/claude-code`, `droid`, `openclaw`. **Pick this for agentic workflows.** |
| Stage | Base image | Purpose |
| ------------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `builder` | `node:26-trixie-slim` | Installs deps (`npm ci --legacy-peer-deps`) and runs `npm run build` (Turbopack by default — see Build-time resources below) |
| `runner-base` | `node:26-trixie-slim` | Production runtime with the Next.js standalone output. **No provider CLIs bundled.** |
| `runner-cli` | `runner-base` | Adds `git`, `docker.io`, `docker-compose` and global CLIs: `@openai/codex`, `@anthropic-ai/claude-code`, `droid`, `openclaw`. **Pick this for agentic workflows.** |
Build a specific target manually:
@@ -160,14 +160,50 @@ docker build --target runner-base -t omniroute:base .
docker build --target runner-cli -t omniroute:cli .
```
Defaults exported by `runner-base`: `PORT=20128`, `HOSTNAME=0.0.0.0`, `NODE_OPTIONS=--max-old-space-size=512`, `DATA_DIR=/app/data`, `OMNIROUTE_MIGRATIONS_DIR=/app/migrations`.
### Build-time resources
Two build args control what the `builder` stage costs. They are build-time only —
`OMNIROUTE_MEMORY_MB` (below) is a separate, runtime knob.
| Build arg | Default | Effect |
| --------------------------- | ------- | ---------------------------------------------------------------------- |
| `OMNIROUTE_USE_TURBOPACK` | `1` | `0` builds with webpack instead. Lower peak memory, slower. |
| `OMNIROUTE_BUILD_MEMORY_MB` | `4096` | V8 heap ceiling (`--max-old-space-size`) for the spawned `next build`. |
Turbopack compiles in native Rust memory that lives **outside** the V8 heap, so
`OMNIROUTE_BUILD_MEMORY_MB` does not bound it. On a host with a memory ceiling the
build is then SIGKILLed by the OOM killer with no error text at all — it simply
stops mid-`Creating an optimized production build`, which reads like a hang rather
than an out-of-memory. If the build host is constrained, switch bundlers:
```bash
docker build --target runner-base \
--build-arg OMNIROUTE_USE_TURBOPACK=0 \
-t omniroute:base .
```
`webpackBuildWorker` is enabled, so `next build` runs a parent **and** a worker
process and each honours `OMNIROUTE_BUILD_MEMORY_MB` separately. Size the container
ceiling above roughly twice that value, not once.
Measured on this tree (`--target runner-base`, `OMNIROUTE_BUILD_MEMORY_MB=6144`):
| Bundler | Container ceiling | Result |
| --------- | ----------------- | ----------------------------- |
| Turbopack | 8 GiB / 16 GiB | OOM-killed at both, silently |
| webpack | 8 GiB | build worker SIGKILLed |
| webpack | 12 GiB | succeeded, peaked at 11.1 GiB |
### Runtime defaults
Defaults exported by `runner-base`: `PORT=20128`, `HOSTNAME=0.0.0.0`, `OMNIROUTE_MEMORY_MB=1024`, `NODE_OPTIONS=--max-old-space-size=1024`, `DATA_DIR=/app/data`, `OMNIROUTE_MIGRATIONS_DIR=/app/migrations`.
Memory behavior in Docker:
- `NODE_OPTIONS=--max-old-space-size=512` is baked into the image as a fallback.
- The image sets `OMNIROUTE_MEMORY_MB=1024` and derives `NODE_OPTIONS=--max-old-space-size=1024` from it.
- The actual server process is started by the standalone launcher, which reads `OMNIROUTE_MEMORY_MB` and appends `--max-old-space-size=<OMNIROUTE_MEMORY_MB>`.
- Node uses the last repeated `--max-old-space-size` value, so setting `OMNIROUTE_MEMORY_MB` controls the effective Docker heap limit.
- If `OMNIROUTE_MEMORY_MB` is unset, the launcher uses `512`.
- Because the image always sets it, the launcher's own RAM-calibrated fallback never applies under Docker. Raise it explicitly (`-e OMNIROUTE_MEMORY_MB=2048`) on a host with headroom.
## Critical Environment Variables
@@ -180,7 +216,7 @@ Beyond the defaults documented in [ENVIRONMENT.md](../reference/ENVIRONMENT.md),
| `REDIS_PORT` | Host-side port for the bundled Redis container | `6379` |
| `REDIS_BIND_HOST` | Host interface the bundled Redis port is published on (loopback unless you add AUTH) | `127.0.0.1` |
| `AUTO_UPDATE_HOST_REPO_DIR` | Host path mounted into `cli` profile at `/workspace/omniroute` for self-update workflows | `.` (current directory) |
| `OMNIROUTE_MEMORY_MB` | Runtime Node heap ceiling for the Docker standalone server; overrides the image fallback above | `512` |
| `OMNIROUTE_MEMORY_MB` | Runtime Node heap ceiling for the Docker standalone server; overrides the image default above | `1024` |
| `DASHBOARD_PORT` / `API_PORT` | Override exposed ports for dashboard (20128) and API (20129) | `20128` / `20129` |
| `OMNIROUTE_BASE_PATH` | URL subpath when the app is published behind a reverse proxy (e.g. `/omniroute`) | _(empty = root)_ |
| `NEXT_PUBLIC_BASE_URL` | Public browser origin including the subpath (e.g. `https://host/omniroute`) | unset |

View File

@@ -523,6 +523,12 @@ exhausts its bounds of `10,000` visited nodes or depth `12`.
Each process uses a process-local guard to reserve limited heavyweight capacity before retaining
and parsing a large request body. A heavyweight lease remains held for the lifetime of an SSE
response.
When capacity is busy, a heavyweight request first waits up to
`OMNIROUTE_CHAT_ADMISSION_QUEUE_MS` (default `5000`, `0` disables the wait) for a slot to free up
before answering the retryable `503`. The bounded wait exists so agent-style clients
(OpenCode, Claude Code, Cursor) that fan out heavy sub-requests concurrently serialize the burst
instead of burning their whole retry budget on immediate rejections and dying mid-task.
Current heavyweight lease occupancy is not surfaced in the dashboard.
Settings → Resilience → Request Queue → Concurrent Requests does not control this; that setting
governs a separate provider request-queue mechanism.
@@ -530,13 +536,18 @@ governs a separate provider request-queue mechanism.
**Fix:**
1. Retry first. Clients should honor `Retry-After` and use backoff rather than immediately
repeating the request.
repeating the request. Note that with the default `OMNIROUTE_CHAT_ADMISSION_QUEUE_MS=5000`
a heavy request already waited up to 5 seconds before the `503`, so a client retry loop should
back off beyond that instead of hammering.
2. If normal deployment traffic repeatedly exhausts the guard, you can cautiously raise
`OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT` from its default of `1`. Increase it one step at a time,
restart OmniRoute after each change, and observe memory headroom under representative load.
Every additional heavyweight request can increase concurrent V8 heap use and container or
host OOM risk. No value is safe for every deployment; validate the setting against your own
traffic and memory limits rather than assuming that `2` is universally safe.
3. Prefer widening the wait (`OMNIROUTE_CHAT_ADMISSION_QUEUE_MS`) over raising the in-flight
limit when bursts are short: waiting costs latency, while an extra concurrent heavyweight
request costs heap residency for the whole request lifetime.
See the [environment-variable reference](../reference/ENVIRONMENT.md#4-security--authentication)
for the authoritative admission settings. Loosening the heavyweight classification thresholds

View File

@@ -0,0 +1,143 @@
---
title: "Feasibility — Telegram Mini App Integration"
version: 3.8.49
lastUpdated: 2026-08-08
---
# Telegram Mini App Integration — Feasibility Analysis
**Status: FEASIBLE with moderate effort (estimated 24 dev-days for a working slice)**
## 1. What "Telegram Mini App" means here
A Telegram Mini App is an iframe-hosted web app opened inside Telegram (via
inline buttons / bot menu buttons) that talks to a bot backend through the
[Telegram WebApp SDK](https://core.telegram.org/bots/webapps). For OmniRoute
the natural shape is:
- **Bot backend** (new): receives Telegram updates (webhook), validates the
Mini App's `initData` signature, and proxies chat requests to OmniRoute's
existing OpenAI-compatible `/v1/chat/completions` surface.
- **Mini App frontend** (new): a small chat UI served by OmniRoute (Next.js
route or `public/` static bundle), using the Telegram WebApp JS SDK.
## 2. Current state of the codebase (verified against `main` @ 918fba5e3)
### Already present — outbound notifications only
| Piece | Location | What it does |
| ---------------------------- | ------------------------------------------- | ----------------------------------------------------------------------------------------------- |
| Telegram webhook integration | `src/lib/webhooks/integrations/telegram.ts` | Builds `sendMessage` payloads for **outbound** gateway events (model, provider, latency, error) |
| Webhook dispatcher | `src/lib/webhookDispatcher.ts` | Routes by kind; decrypts `botToken` from DB metadata for telegram |
| Webhook kinds | `src/lib/db/webhooks.ts` | `slack \| telegram \| discord \| custom` |
| Webhook CRUD + test | `src/app/api/webhooks/*` | Create/update/test; telegram kind skips `url` (uses bot token + chat_id) |
| Bot token validation | `telegram.ts:18` | `BOT_TOKEN_RE = /^\d+:[A-Za-z0-9_-]{35,}$/` |
| Encryption requirement | `webhooks/route.ts:77` | Telegram webhooks require DB encryption enabled (bot tokens stored at rest) |
### Missing — what a Mini App needs that does not exist yet
| Gap | Detail |
| -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Inbound Bot API listener** | No `setWebhook` registration, no `/bot<token>/getUpdates` polling, no update handling anywhere. Only the `sendMessage` direction exists. |
| **WebApp `initData` validation** | No HMAC-SHA256 check of `initData` against the bot token (`WebAppData` hash validation from the Bot API docs). |
| **Telegram bot library** | `package.json` has no `telegraf`/`grammy`/`telegram-bot-api` dependency. Would need to add one or hand-roll the (small) HMAC + fetch logic. |
| **Mini App hosting surface** | `public/` exists (static assets) and Next.js routes exist; no `/miniapp` route or static bundle yet. |
| **Session → API key mapping** | Mini App users need to authenticate to `/v1/chat/completions`. Two options: per-user generated OmniRoute API keys (via `src/lib/db/apiKeys`) or a bot-side proxy that injects a shared key. |
## 3. Constraints
### 3.1 Architectural
- **No existing inbound-bot layer.** The webhook system is strictly
event→outbound. A Mini App needs a _new_ Bot API webhook endpoint
(`POST /api/telegram/webhook/<botToken-prefix>` or a dedicated route) plus
update dispatch. This is additive — no conflicts with the existing
`webhooks/` subsystem, but the two must not share the `botToken` storage
semantics blindly (webhooks store bot tokens for _outbound_; the Mini App
needs the same token for _inbound_ signature checks — same token, new use).
- **Public HTTPS required.** Telegram only delivers updates to an HTTPS
endpoint with a valid cert. Self-hosted OmniRoute behind Tailscale/ngrok
needs a public tunnel or Cloudflare Tunnel for the webhook path
(`TELEGRAM_WEBHOOK_URL`-style env). The dashboard can render the current
public origin (`OMNIROUTE_PUBLIC_BASE_URL`) but no webhook registration
helper exists.
- **Encryption gate.** `webhooks/route.ts:77` already refuses telegram
kinds without DB encryption. The Mini App bot token has the same
sensitivity (it _is_ the HMAC secret for initData validation) — same gate
applies, which is a _good_ constraint (no plaintext tokens).
### 3.2 Telegram platform
- **initData is the only trust anchor.** Mini App auth = verify
`hash` field of `initData` using HMAC-SHA256(key = SHA256(bot_token),
data = sorted `key=value` pairs minus `hash`). Must be implemented
server-side; never trust the client.
- **No inbound push to arbitrary users.** Telegram bots cannot initiate
conversations. The Mini App works for users who _already_ have the bot —
or you add a `/start` command handler + deep-link (`t.me/bot?startapp=`).
- **Rate limits.** Bot API ~30 msg/s per bot, 20 msg/min per chat group.
Chat responses via `sendMessage`/`answerWebAppQuery` are fine at gateway
scale, but streaming must be emulated (send progressive edits or chunked
messages) — no native SSE into Telegram.
- **WebApp SDK quirks.** `Telegram.WebApp.ready()` must be called; theme
params come from the SDK; the mini app is sandboxed iframe (no
`window.open` to external, clipboard limited). For a chat UI this is fine.
### 3.3 Security / policy
- **Per-user key issuance is the clean model.** Rather than exposing the
admin's own API keys, mint a scoped OmniRoute API key per Telegram user
(`apiKeys` table + `isModelAllowedForKey` policy), or proxy with a single
gateway key and map `user_id` → account. Recommendation: per-user keys so
existing rate-limit / model-allowlist / policy code applies unchanged.
- **initData expiry.** `auth_date` in initData must be checked (Telegram
recommends < 24h; short TTLs for chat flows).
- **Secret handling.** Bot token must stay in the encrypted DB / env —
mirror the existing `isEncryptionEnabled()` gate.
## 4. Required next steps (implementation plan)
### Phase 0 — Spike (½1 dev-day)
1. Add `grammy` or `telegraf` (or ~60 lines of hand-rolled HMAC + fetch).
2. Implement `src/lib/telegram/initData.ts``verifyInitData(initData, botToken)`.
3. Stand up a throwaway `POST /api/telegram/miniapp/webhook` route behind
`TELEGRAM_WEBHOOK_SECRET`; register via `setWebhook` once, locally.
### Phase 1 — Minimal chat slice (12 dev-days)
1. **Webhook endpoint** `POST /api/telegram/bot/update` (or
`/api/telegram/miniapp/update`): parse Update, verify initData, dispatch.
2. **Command handler**: `/start` → reply with deep link
`https://t.me/<bot>?startapp=<userKey>`; `startapp` param carries a
one-time token that maps to a generated OmniRoute API key.
3. **Chat proxy**: map `initData.user.id` → API key → call
`handleChat` (same path as `/v1/chat/completions`) → reply via
`sendMessage` (non-stream) or chunked edits (fake streaming).
4. **Mini App page**: `src/app/(dashboard)/miniapp/page.tsx` (or static
bundle in `public/miniapp/`) — Telegram WebApp SDK init + minimal chat
UI posting to the bot webhook.
5. **Config**: `TELEGRAM_BOT_TOKEN` env (or reuse webhook metadata),
`OMNIROUTE_PUBLIC_BASE_URL` for webhook URL display; doc in
`.env.example` + `ENVIRONMENT.md` (env-doc-sync check).
### Phase 2 — Production hardening (1 dev-day)
- Streaming emulation (message edits), error/backpressure mapping to Bot API
limits, per-user key revocation (`/logout` command → revoke API key),
usage/rate-limit surfacing (reuse `enforceApiKeyPolicy`), webhook
registration helper in dashboard settings, i18n for the mini app UI.
## 5. Verdict
**Feasible.** The gateway already exposes the exact API a Mini App chat
needs (`/v1/chat/completions` with per-key policy), and the outbound
Telegram webhook shows the team already handles bot tokens safely
(encryption gate + token format validation). The genuinely new surface is
small: an inbound update webhook + initData HMAC verification + a thin
chat proxy + a static Mini App page. No changes to the core SSE/relay
pipeline are required.
**Primary risks:** (1) public HTTPS requirement for the webhook (tunnel
needed on self-hosted installs), (2) no native streaming to Telegram
(UX tradeoff), (3) initData trust must be strictly server-side.

View File

@@ -43,7 +43,6 @@ lastUpdated: 2026-06-28
- [22. Debugging](#22-debugging)
- [23. GitHub Integration](#23-github-integration)
- [24. Skills Sandbox (v3.8.0+)](#24-skills-sandbox-v380)
- [27. Radar Feed (Self-Hosting)](#27-radar-feed-self-hosting)
- [Deployment Scenarios](#deployment-scenarios)
- [Audit: Removed / Dead Variables](#audit-removed--dead-variables)
@@ -190,20 +189,19 @@ OmniRoute uses **SQLite** (via `better-sqlite3`) for all persistence. These vari
| `MAX_BODY_SIZE_BYTES` | `10485760` (10 MB) | `src/shared/middleware/bodySizeGuard.ts` | Maximum allowed request body size. Rejects payloads exceeding this limit. |
| `OMNIROUTE_CHAT_LARGE_BODY_BYTES` | `262144` (256 KB) | `src/shared/middleware/chatBodyAdmission.ts` | Actual request bodies at or above this threshold require an atomic process-local heavyweight admission lease before JSON parsing. |
| `OMNIROUTE_CHAT_HARD_MAX_BODY_BYTES` | `52428800` (50 MB) | `src/shared/middleware/chatBodyAdmission.ts` | Chat-route hard cap enforced against bytes read during bounded ingestion, including requests with missing, invalid, or dishonest `Content-Length`; excess receives `413`. |
| `OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT` | `1` | `src/shared/middleware/chatBodyAdmission.ts` | Maximum heavyweight chat requests admitted concurrently in one process. When capacity is unavailable, OmniRoute returns retryable `503` with `Retry-After`. |
| `OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT` | `1` | `src/shared/middleware/chatBodyAdmission.ts` | Maximum heavyweight chat requests admitted concurrently in one process. When capacity is unavailable, OmniRoute waits up to `OMNIROUTE_CHAT_ADMISSION_QUEUE_MS` for a slot, then returns retryable `503` with `Retry-After`. |
| `OMNIROUTE_CHAT_ADMISSION_QUEUE_MS` | `5000` | `src/shared/middleware/chatBodyAdmission.ts` | How long a heavyweight chat request waits for an admission slot before the retryable `503`. A bounded wait serializes agent bursts (OpenCode, Claude Code, Cursor sub-requests) that would otherwise burn their client retry budget on immediate rejections; `0` restores the legacy immediate-reject behaviour. |
| `OMNIROUTE_CHAT_HEAVY_MESSAGE_COUNT` | `200` | `src/shared/middleware/chatBodyAdmission.ts` | Message count that classifies a chat request as heavyweight even when its body is below the byte threshold. |
| `OMNIROUTE_CHAT_HEAVY_TOOL_COUNT` | `64` | `src/shared/middleware/chatBodyAdmission.ts` | Tool count that classifies a chat request as heavyweight even when its body is below the byte threshold. |
| `OMNIROUTE_CHAT_HEAVY_ESTIMATED_TOKENS` | `32000` | `src/shared/middleware/chatBodyAdmission.ts` | Conservative string-size token estimate that classifies a request as heavyweight; this is an admission-cost proxy, not provider billing tokenization. |
| `OMNIROUTE_CHAT_HARD_MAX_MESSAGES` | `0` (disabled) | `src/shared/middleware/chatBodyAdmission.ts` | Optional opt-in chat history cap. Disabled by default: a message count is deployment policy, not a universal property of a request, and capping here rejects conversations with a terminal `413` before the compression pipeline can make them servable. Heap growth is bounded by `OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT` and the heap-pressure shed. Set a positive value on memory-constrained deployments that need a hard ceiling; excess then receives structured compact-required `413`. |
| `OMNIROUTE_CHAT_HARD_MAX_MESSAGES` | `800` | `src/shared/middleware/chatBodyAdmission.ts` | Hard chat history cap. Requests above it receive structured compact-required `413` before compression, translation, or provider dispatch. |
| `OMNIROUTE_MAX_NONSTREAMING_RESPONSE_BYTES` | `67108864` (64 MB) | `open-sse/handlers/chatCore/nonStreamingResponseBody.ts` | Hard cap for a non-streaming upstream response buffered fully into memory. Past this the upstream reader is cancelled and the request fails fast instead of growing an unbounded string until the heap is exhausted. |
| `OMNIROUTE_FORWARDING_HEADER_BUDGET_BYTES` | `768` | `open-sse/handlers/chatCore/responseHeaders.ts` | Max wire bytes forwarded from upstream response headers. When the budget is exceeded, lower-priority headers (e.g., custom `x-codex-*`, `x-oai-request-id`) are dropped to stay within common reverse-proxy header limits. Set higher to forward more upstream metadata at the cost of larger response header size. |
| `CORS_ORIGIN` | _(unset)_ | `src/server/cors/origins.ts` | Legacy single-origin CORS allowlist. Prefer `CORS_ALLOWED_ORIGINS` for new deployments. CORS is only for cross-origin browser API clients; authenticated dashboard writes use same-origin requests plus session-bound CSRF protection instead. |
| `CORS_ALLOWED_ORIGINS` | _(unset)_ | `src/server/cors/origins.ts` | Comma-separated CORS allowlist. No wildcard is sent unless `CORS_ALLOW_ALL=true` is explicitly configured. |
| `CORS_ALLOW_ALL` | `false` | `src/server/cors/origins.ts` | Development-only escape hatch to echo any browser `Origin`. Do not enable on shared or production deployments. |
| `OUTBOUND_SSRF_GUARD_ENABLED` | `true` | `src/shared/network/outboundUrlGuard.ts` | Block provider calls targeting private/loopback/link-local IP ranges. Disable only in isolated test envs. |
| `OMNIROUTE_ALLOW_PRIVATE_PROVIDER_URLS` | `false` | `src/shared/network/outboundUrlGuard.ts` | Allow provider URLs pointing to private/local networks (localhost, 192.168.x.x, 10.x.x.x, etc.). **REQUIRED for self-hosted providers** (LM Studio, Ollama, vLLM, Llamafile, Triton, SearXNG). When `false`, the dashboard rejects validation of local URLs. |
| `OMNIROUTE_ALLOW_LOCAL_PROVIDER_URLS` | `true` | `src/shared/network/outboundUrlGuard.ts` | Allow adding/validating providers on local/private addresses (127.0.0.1, localhost, LAN, private ranges) — scoped to the provider validation path. **Default `true`** (local-first); set `false` to enforce strict public-only blocking. Cloud-metadata endpoints (169.254.169.254, metadata.google.internal) stay blocked regardless. (#5066) |
| `AUDIO_REMOTE_PROVIDER_NODES` | `false` | `src/app/api/v1/_shared/audioProviderNodes.ts` | Let the `/v1/audio/*` routes (transcriptions, speech, translations) use an OpenAI-compatible provider node hosted outside localhost. Off by default — routing audio to a remote host changes egress identity and must be an explicit operator decision. Loopback/private nodes (localhost, 127.0.0.1, 172.16-31.x) are always allowed and unaffected. (#3963) |
### Hardening Checklist
@@ -267,7 +265,6 @@ OmniRoute provides a two-layer defense: request-side injection scanning and resp
| `OMNIROUTE_PAYLOAD_RULES_PATH` | `./config/payloadRules.json` | `open-sse/services/payloadRules.ts` | Path to payload manipulation rules JSON file (per-model/protocol upstream tweaks). |
| `OMNIROUTE_PAYLOAD_RULES_RELOAD_MS` | `5000` | `open-sse/services/payloadRules.ts` | Reload interval (ms) for hot-reloading the payload rules file. Minimum `1000`. |
| `OMNIROUTE_PREFER_CLAUDE_CODE_FOR_UNPREFIXED_CLAUDE_MODELS` | `false` | `open-sse/services/model.ts` | Opt-in: route bare `claude-*` model IDs from Claude Code clients through the Claude Code OAuth account instead of requiring a provider prefix. Explicit provider prefixes still win. Also configurable via a dashboard toggle on the Claude provider page. |
| `COMBO_CONCURRENCY_PER_MODEL` | `3` | `open-sse/services/comboConfig.ts` | Per-model concurrency cap for round-robin combos (#9100). The round-robin combo semaphore was hard-capped at 3 concurrent requests per model with no override, serializing higher-concurrency traffic behind that cap. Validated to `>= 1`, clamped to `<= 32`. |
---
@@ -381,14 +378,6 @@ Controls how OmniRoute discovers and launches CLI sidecars (Claude Code, Codex,
| `CLI_QODER_BIN` | `qoder` | `src/shared/services/cliRuntime.ts` | Custom path to Qoder CLI binary. |
| `CLI_QWEN_BIN` | `qwen` | `src/shared/services/cliRuntime.ts` | Custom path to the Qwen Code CLI binary. |
| `CLI_DEVIN_BIN` | `devin` | `open-sse/executors/devin-cli.ts` | Custom path to the Devin CLI binary (v3.8.0). Used by the Windsurf/Devin executor. |
| `CLI_DEVIN_AGENTIC_BIN` | `devin` | `open-sse/executors/devin-cli-agentic.ts` | Agentic bridge-only Devin CLI override. The executor accepts only the local ACP stdio upstream. |
| `DEVIN_AGENTIC_HOME` | _(required)_ | `open-sse/executors/devin-cli-agentic.ts` | Absolute isolated home for the agentic Devin subprocess; accepted bridge paths are `/home/bridge` and task-local `.sandbox` paths. |
| `DEVIN_AGENTIC_ACP_TIMEOUT_MS` | `120000` | `open-sse/executors/devin-cli-agentic.ts` | Maximum duration of one Devin ACP turn before the bridge terminates the child and returns an explicit timeout. |
| `DEVIN_BRIDGE_MODEL` | `devin-cli-agentic/swe-1-7` | `docker/devin-bridge/compose.yml` | Main Claude Code model alias for the isolated bridge. The live harness replaces the example with a model returned by the current Devin account. |
| `DEVIN_BRIDGE_SONNET_MODEL` | `DEVIN_BRIDGE_MODEL` | `docker/devin-bridge/compose.yml` | Isolated bridge alias used when Claude Code requests its Sonnet default. |
| `DEVIN_BRIDGE_OPUS_MODEL` | `DEVIN_BRIDGE_MODEL` | `docker/devin-bridge/compose.yml` | Isolated bridge alias used when Claude Code requests its Opus default. |
| `DEVIN_BRIDGE_HAIKU_MODEL` | `DEVIN_BRIDGE_MODEL` | `docker/devin-bridge/compose.yml` | Isolated bridge alias used when Claude Code requests its Haiku default. |
| `DEVIN_BRIDGE_SUBAGENT_MODEL` | `DEVIN_BRIDGE_MODEL` | `docker/devin-bridge/compose.yml` | Isolated bridge alias used for Claude Code subagents. |
| `AUGGIE_BIN` | `auggie` | `open-sse/executors/auggie.ts` | Absolute-path override for the Augment (Auggie) CLI binary used by the local `auggie` provider. Falls back to `CLI_AUGGIE_BIN`, then a PATH lookup. |
| `CLI_AUGGIE_BIN` | `auggie` | `open-sse/executors/auggie.ts` | Alias override for the Augment (Auggie) CLI binary path (checked after `AUGGIE_BIN`). |
| `HERMES_HOME` | `~/.hermes` | `src/lib/cli-helper/config-generator/hermesHome.ts` | Hermes Agent home directory where OmniRoute reads/writes the Hermes CLI config. Matches the env var the Hermes PowerShell installer sets on Windows (`%LOCALAPPDATA%\hermes`). |
@@ -426,6 +415,7 @@ detection above).
| `OMNIROUTE_HTTP_TIMEOUT_MS` | `30000` | `bin/cli/api.mjs` | Per-attempt HTTP timeout (ms) for CLI → server requests. |
| `OMNIROUTE_VERBOSE` | `0` | `bin/cli/api.mjs` | Set to `1` to print retry/backoff diagnostics to stderr during CLI commands. |
| `OMNIROUTE_PLUGIN_PATH` | _(unset)_ | `bin/cli/plugins.mjs` | Custom directory for CLI plugin discovery (`omniroute-cmd-*` packages). Defaults to `~/.omniroute/plugins/` when unset. |
| `OMNIROUTE_PLUGINS_ALLOW_EXEC` | `0` | `src/lib/plugins/pluginWorker.ts` | Set to `1` to allow plugins to request the `exec` permission (spawn child processes from the worker sandbox). Local operator only. |
---
@@ -449,6 +439,12 @@ detection above).
| `PROVIDER_LIMITS_SYNC_SPACING_MS` | `1500` | `src/lib/usage/providerLimits.ts` | Gap (ms) between consecutive OAuth quota fetches in a bulk sync; OAuth connections are fetched one at a time to avoid bursting an upstream. `0` opts out (concurrent). |
| `OMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS` | `250` | `open-sse/services/quotaFetchThrottle.ts` | Min interval (ms) between consecutive upstream quota fetches on the per-request preflight/monitor path; spaces concurrent network calls so many accounts on one IP don't burst the upstream. Wired into the Codex (`/wham/usage`), DeepSeek, Bailian (both fetch sites), OpenCode, and Crof quota fetchers (#6009, #6911). The generic `usage.ts::getUsageForProvider` dispatch path (github/glm/minimax/nanogpt/xai/etc.) is not yet covered — tracked separately. Cache hits unaffected. `0` disables; clamped `0..5000`. |
| `PROVIDER_LIMITS_POST_USAGE_REFRESH_DELAY_MS` | `5000` | `src/lib/usage/providerLimits.ts` | Delay (ms) before refreshing provider limits after a real usage event, giving the upstream quota API time to register consumption. |
| `OMNIROUTE_LOGIN_BROWSER_PATH` | auto-detect | `open-sse/services/adobeFireflyBrowserLogin.ts` | Absolute path to a system Chrome or Edge executable used for interactive Adobe Firefly sign-in and off-screen renewal. |
| `ADOBE_FIREFLY_BROWSER_REFRESH` | enabled | `open-sse/services/adobeFireflySession.ts` | Keeps IMS and browser-risk state fresh with account-scoped Chrome CDP sessions. Set to `0` to disable browser renewal. |
| `ADOBE_FIREFLY_SESSION_DISK` | enabled | `open-sse/services/adobeFireflySession.ts` | Persists repaired Adobe sessions under `DATA_DIR` across process restarts. Set to `0` to keep sessions memory-only. |
| `ADOBE_FIREFLY_MIN_SUBMIT_GAP_MS` | `12000` | `open-sse/services/adobeFireflySession.ts` | Minimum spacing in milliseconds between Adobe Firefly generate submissions; `0` disables spacing. |
| `ADOBE_FIREFLY_BATCH_EXTRA_GAP_MS` | `15000` | `open-sse/services/adobeFireflySession.ts` | Extra quiet period in milliseconds after every third successful Adobe submission. |
| `ADOBE_FIREFLY_SUBMIT_BASE_DELAY_MS` | `8000` | `open-sse/services/adobeFireflyClient.ts` | Base backoff in milliseconds after transient Adobe 408 responses; combined with submit spacing across at most five attempts. |
| `OMNIROUTE_DISABLE_BACKGROUND_SERVICES` | `false` | `src/instrumentation-node.ts` | Disable all background services (sync, pricing, model refresh). Useful for CI/test. |
| `OMNIROUTE_ENABLE_RUNTIME_BACKGROUND_TASKS` | _(unset)_ | `src/lib/config/runtimeSettings.ts` | Force background tasks on under automated test detection. Set `1` to override the test heuristic. |
| `OMNIROUTE_BUDGET_RESET_JOB_INTERVAL_MS` | `600000` | `src/lib/jobs/budgetResetJob.ts` | Budget reset check cadence (ms). Floor `10000`. |
@@ -462,7 +458,6 @@ detection above).
| `COMPRESSION_PIPELINE_BREAKER_THRESHOLD` | `3` | `open-sse/services/compression/pipelineEngineBreaker.ts` | Consecutive cross-request failures before an engine's breaker opens. |
| `COMPRESSION_PIPELINE_BREAKER_COOLDOWN_MS` | `30000` | `open-sse/services/compression/pipelineEngineBreaker.ts` | Milliseconds an opened engine stays skipped before a half-open probe. |
| `COMPRESSION_CCR_RETRIEVAL_RAMP_FACTOR` | `2` | `open-sse/services/compression/engines/ccr/index.ts` | T08/H8 CCR retrieval-feedback ramp: each prior retrieval of a stored block raises its effective `minChars` linearly (frequently-retrieved content compresses less; `>=3` retrievals = never compressed). `1` disables the ramp (binary skip at the threshold only). |
| `COMPRESSION_CCR_DURABLE_STORE` | `true` | `open-sse/services/compression/engines/ccr/index.ts` | CCR durable block store (#9061). Backs the in-memory store with SQLite so a block survives LRU eviction, the TTL, a restart, or a retrieve landing on another instance. Set `false` to keep blocks in memory only. Blocks over 512KB and cloud runtimes stay memory-only regardless. |
| `COMPRESSION_PREFIX_FREEZE_ENABLED` | `false` | `open-sse/services/compression/prefixFreeze.ts` | T08/H5 usage-observed prefix freeze master switch. **Opt-in (default off)** — when on, a system prompt observed `>=` the threshold is treated as a stable cacheable prefix and preserved from compression even for providers the static cache heuristic misses (freeze only *preserves*, never mutates). |
| `COMPRESSION_PREFIX_FREEZE_THRESHOLD` | `3` | `open-sse/services/compression/prefixFreeze.ts` | Observations of a system prompt before it is treated as a frozen stable prefix. |
| `OMNIROUTE_BOOTSTRAPPED` | `false` | `src/app/(dashboard)/dashboard/page.tsx` | Set `true` by bootstrap script after initial setup. Controls setup wizard visibility. |
@@ -518,12 +513,8 @@ Built-in credentials for **localhost development**. For remote deployments, regi
| `OMNIROUTE_QODER_WORKSPACE` | Qoder | Alias for `QODER_CLI_WORKSPACE`. |
| `QODER_CLI_CONFIG_DIR` | Qoder | Override the Qoder CLI config dir (isolated PAT session, avoids clobbering a browser login). |
| `BLACKBOX_WEB_VALIDATED_TOKEN` | Blackbox Web | Frontend `tk` token to send as `validated` on `/api/chat`. Required when Blackbox enforces token matching; otherwise OmniRoute falls back to a random UUID. See issue #2252. |
| `VISION_BRIDGE_BASE_URL` | Vision Bridge guardrail | OpenAI-compatible base URL for non-Anthropic vision-bridge calls. Defaults to the legacy OpenAI URL env or api.openai.com. Point at OmniRoute's `/v1` self-loop or any OpenAI-compat endpoint (Gemini OpenAI-compat, OpenRouter). Issue #2232. When the URL is OmniRoute's own `/v1`, the describe sub-request sends `x-omniroute-admission-bypass: internal` and authenticates with the resolved self-loop credential (`sk_omniroute` sentinel in local mode, or `OMNIROUTE_API_KEY` / `ROUTER_API_KEY`#1350) so `REQUIRE_API_KEY=true` deployments work. |
| `VISION_BRIDGE_BASE_URL` | Vision Bridge guardrail | OpenAI-compatible base URL for non-Anthropic vision-bridge calls. Defaults to the legacy OpenAI URL env or api.openai.com. Point at OmniRoute's `/v1` self-loop or any OpenAI-compat endpoint (Gemini OpenAI-compat, OpenRouter). Issue #2232. |
| `VISION_BRIDGE_API_KEY` | Vision Bridge guardrail | API key for the URL above. Overrides per-provider OpenAI / Google env vars for non-Anthropic vision-bridge calls. Anthropic models keep their dedicated Anthropic key path. Issue #2232. |
| `RAYCAST_BEARER_TOKEN` | Raycast Pro | Optional manual override for the Raycast access token (normally captured via macOS Auto-Import). No OAuth client_id/secret — reverse-engineered, local/personal use only. |
| `RAYCAST_DEVICE_ID` | Raycast Pro | Optional manual override for the Raycast device ID used to sign requests. |
| `RAYCAST_AID` | Raycast Pro | Optional manual override for the Raycast account/app ID; falls back to the device ID when unset. |
| `RAYCAST_SIG_SECRET` | Raycast Pro | Optional override for the request-signing HMAC secret. Defaults to a community-extracted value in `open-sse/services/raycast.ts`. |
> [!WARNING]
>
@@ -671,13 +662,12 @@ REQUEST_TIMEOUT_MS (global override)
| `API_BRIDGE_SERVER_SOCKET_TIMEOUT_MS` | `0` | Raw socket timeout (0 = disabled). |
| `SHUTDOWN_TIMEOUT_MS` | `30000` | Grace period on SIGTERM/SIGINT before force-exit. |
| `OMNIROUTE_DEFAULT_FETCH_TIMEOUT_MS` | `120000` | Fallback used by `src/shared/utils/fetchTimeout.ts` when `FETCH_TIMEOUT_MS` is unset. |
| `OMNIROUTE_RELAY_FETCH_TIMEOUT_MS` | `25000` | Relay-specific fetch timeout in `open-sse/utils/proxyFetch.ts` (#9158). A hung relay must fail before the client/agent timeout (~30s) so callers see a relay-specific failure instead of a generic upstream timeout. Capped at `29000` so it always fires first. |
| `OMNIROUTE_RETRY_BACKOFF_MS` | `10` | Shared retry backoff for the direct/relay/proxy retry-once paths in `open-sse/utils/proxyFetch.ts` (#9158). `0` = retry immediately. |
| `OMNIROUTE_CHATGPT_TLS_TIMEOUT_MS` | `60000` | Wire-level timeout for the bogdanfinn/tls-client koffi binding (`chatgptTlsClient.ts`). |
| `OMNIROUTE_CHATGPT_TLS_GRACE_MS` | `10000` | JS-side grace added on top of the wire timeout when the native binding is wedged. |
| `OMNIROUTE_CHATGPT_STREAM_FIRST_BYTE_TIMEOUT_MS` | `30000` (30s) | Max wait for the first streamed byte from the ChatGPT TLS sidecar (`chatgptTlsClient.ts`) before aborting a dead stream. Raise if upstream cold-starts exceed the window. |
| `OMNIROUTE_CLAUDE_TLS_TIMEOUT_MS` | `60000` | Wire-level timeout for the bogdanfinn/tls-client koffi binding (`claudeTlsClient.ts`). |
| `OMNIROUTE_CLAUDE_TLS_GRACE_MS` | `10000` | JS-side grace added on top of the wire timeout when the native binding is wedged. |
| `OMNIROUTE_CHATGPT_STREAM_FIRST_BYTE_TIMEOUT_MS` | `30000` | Max wait for the first streamed byte from the ChatGPT TLS sidecar. |
| `OMNIROUTE_CLAUDE_TLS_TIMEOUT_MS` | `60000` | Wire-level timeout for the bogdanfinn/tls-client koffi binding. |
| `OMNIROUTE_CLAUDE_TLS_GRACE_MS` | `10000` | JS-side grace added on top of the wire timeout when the native binding is wedged. |
| `OMNIROUTE_RUNNOW_TIMEOUT_MS` | `30000` | Timeout for the `/api/jobs/:id/run-now` endpoint. Bounds how long a run-now call waits for an in-flight job to finish before starting the queued run. See `src/app/api/jobs/[id]/run-now/route.ts`. |
| `OMNIROUTE_PPLX_TLS_TIMEOUT_MS` | `30000` | Wire-level timeout for the bogdanfinn/tls-client koffi binding (`perplexityTlsClient.ts`). |
| `OMNIROUTE_PPLX_TLS_GRACE_MS` | `10000` | JS-side grace added on top of the wire timeout when the native binding is wedged. |
| `OMNIROUTE_GROK_TLS_TIMEOUT_MS` | `60000` | Wire-level timeout for the bogdanfinn/tls-client koffi binding (`grokTlsClient.ts`). |
@@ -686,7 +676,6 @@ REQUEST_TIMEOUT_MS (global override)
| `OMNIROUTE_NOTION_TLS_GRACE_MS` | `10000` | JS-side grace added on top of the wire timeout when the native binding is wedged. |
| `OMNIROUTE_BROWSER_POOL` | `on` | Shared Playwright browser pool for browser-backed web-cookie chat (`browserPool.ts`); set `off` to disable. |
| `WEB_COOKIE_USE_BROWSER` | `0` | Opt a web-cookie chat request into the browser-backed path (`browserBackedChat.ts`); `1` to enable. |
| `OMNIROUTE_LOGIN_BROWSER_PATH` | _(auto-detected)_ | Path to a system Chrome/Edge executable for the Adobe Firefly interactive browser sign-in (`adobeFireflyBrowserLogin.ts`); overrides per-OS auto-detection. |
Combo target attempts inherit the resolved upstream request timeout (`FETCH_TIMEOUT_MS`, or
`REQUEST_TIMEOUT_MS` when it supplies the fetch default). Set `targetTimeoutMs` in a combo,
@@ -734,16 +723,16 @@ The logging system writes to both stdout and rotated log files. All configuratio
| `CALL_LOG_RETENTION_DAYS` | `7` | Days to keep request/call log entries in the database. |
| `CALL_LOG_MAX_ENTRIES` | `10000` | Max call log entries in the in-memory buffer. |
| `CALL_LOGS_TABLE_MAX_ROWS` | `100000` | Max rows in the `call_logs` SQLite table before pruning. |
| `ENABLE_REQUEST_LOGS` | _(unset)_ | Force detailed request logging on or off, overriding the dashboard setting. |
| `MAX_PENDING_REQUEST_AGE_MS` | `3600000` (1 hour) | Max age for orphaned active request log entries before in-memory cleanup. |
| `CALL_LOG_PIPELINE_CAPTURE_STREAM_CHUNKS` | `true` | Store stream chunks in pipeline artifacts when `call_log_pipeline_enabled=true`. |
| `CALL_LOG_PIPELINE_MAX_SIZE_KB` | `512` | Max pipeline call log artifact size in KB when `call_log_pipeline_enabled=true`. |
| `PROXY_LOGS_TABLE_MAX_ROWS` | `100000` | Max rows in the `proxy_logs` SQLite table before pruning. |
| `APP_LOG_ROTATION_CHECK_INTERVAL_MS` | `60000` (1 min) | How often `src/lib/logRotation.ts` re-checks the active log file size. |
| `CHAT_LOG_TEXT_LIMIT` | `65536` | Max string length retained in chat log artifacts (default 64 KB). |
| `CHAT_LOG_ARRAY_TAIL_ITEMS` | `24` | Number of array items retained from the tail when truncating chat log payloads. |
| `CHAT_LOG_ARRAY_TAIL_ITEMS` | `128` | Number of array items retained from the tail when truncating chat log payloads. |
| `CHAT_LOG_MAX_DEPTH` | `6` | Max nesting depth before chat log payloads are truncated. |
| `CHAT_LOG_MAX_OBJECT_KEYS` | `80` | Max object keys retained in chat log payloads (0 = unlimited). |
| `CHAT_LOG_MAX_BODY_KB` | `1024` | Max request/response body size before `truncateForLog()` summarizes it, in KB. |
| `CHAT_DEBUG_FILE` | `false` | When true, `serializeArtifactForStorage` skips size-based truncation. Debug only. |
---
@@ -780,13 +769,9 @@ Embedding layer, vector store and reranking knobs for the persistent memory subs
| `MEMORY_TRANSFORMERS_MODEL` | `Xenova/all-MiniLM-L6-v2` | HF repo id for the opt-in `@huggingface/transformers` local MiniLM pipeline (~23 MB int8, ~400 MB RAM). |
| `MEMORY_STATIC_MODEL` | `minishlab/potion-base-8M` | HF repo id for the static potion/Model2Vec lookup-table embedder. Downloaded lazily into the cache dir. |
| `MEMORY_STATIC_CACHE_DIR` | `<DATA_DIR>/embeddings` | Directory used to cache the static potion model files. Defaults under `DATA_DIR` when unset. |
| `HF_HUB_ENDPOINT` | `https://huggingface.co` | Override Hugging Face Hub base URL used by `staticPotion.ts` (e.g. mirror endpoint for air-gapped setups). |
| `MEMORY_VEC_TOP_K` | `20` | Default top-K used by the `sqlite-vec` brute-force vector search inside `src/lib/memory/vectorStore.ts`. |
| `MEMORY_RRF_K` | `60` | Reciprocal Rank Fusion constant `k` for hybrid FTS5 + vector retrieval (sqlite-vec recipe). |
| `NOTION_API_KEY` | _(unset)_ | API key for Notion backend (used by `genericBackend.ts` known backend preset). |
| `NOTION_API_URL` | `https://api.notion.com/v1`| Base URL for Notion API (can override for self-hosted Notion alternatives). |
| `OBSIDIAN_API_KEY` | _(unset)_ | API key for Obsidian Vault backend (used by `genericBackend.ts` known backend preset). |
| `OBSIDIAN_API_URL` | `http://localhost:27123` | Base URL for Obsidian Vault API (can override for remote vault). |
| `HF_HUB_ENDPOINT` | `https://huggingface.co` | Override Hugging Face Hub base URL used by `staticPotion.ts` (e.g. mirror endpoint for air-gapped setups). |
| `MEMORY_TYPED_DECAY_ENABLED` | `false` | TV6 typed memory decay master switch. **Opt-in (default off)** — the sweep **deletes** decayed memories. With it off, `access_count`/`last_accessed_at` are pure telemetry and nothing is ever deleted. |
| `MEMORY_TYPED_DECAY_EPISODIC_DAYS` | `30` | TTL (days) after which an unused `episodic` memory decays. `0` makes episodic immune too. Durable types (`factual`/`procedural`/`semantic`) are always immune. The decay clock re-bases on `last_accessed_at`. |
| `MEMORY_TYPED_DECAY_ACCESS_IMMUNITY` | `3` | A memory injected `>=` this many times becomes immune to decay regardless of type. `0` disables access immunity. |
@@ -853,7 +838,6 @@ Reverse-engineered session bridge for hyperagent.com (`src/shared/constants/prov
| Variable | Default | Source File | Description |
| ----------------------------------- | ------------- | ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `MODELS_DEV_SYNC_ENABLED` | `false` | `src/lib/modelsDevSync.ts` | Opt-in switch for the models.dev capability sync. Set to anything non-empty it wins over the `modelsDevSyncEnabled` setting (Dashboard > Settings > AI) in either direction, so a deployment can pin the sync on or off without depending on database state surviving a rebuild; unset, it defers to that setting. On for `1`, `true`, `yes` or `on` in any casing; any other value is off. |
| `MODELS_DEV_SYNC_INTERVAL` | `86400` (24h) | `src/lib/modelsDevSync.ts` | Development-time model catalog sync interval in seconds. |
| `CONTEXT_WINDOW_RECONCILE_INTERVAL` | `86400` (24h) | `src/lib/contextWindowResolver.ts` | Interval (seconds) for the self-correcting context-window reconciler (5004): pins provider-declared windows from `/models` discovery as `auto:discovery` overrides when they diverge from the catalog. Set to `0` to disable. Reuses already-synced data (no new fetch); never overwrites `manual` overrides. |
@@ -869,7 +853,6 @@ Reverse-engineered session bridge for hyperagent.com (`src/shared/constants/prov
| `NANOBANANA_POLL_INTERVAL_MS` | `2500` | `open-sse/handlers/imageGeneration.ts` | NanoBanana job polling frequency. |
| `DESIGNER_WEB_POLL_TIMEOUT_MS` | `60000` | `open-sse/handlers/imageGeneration/providers/designerWeb.ts` | Max wait for microsoft-designer-web image generation jobs. |
| `DESIGNER_WEB_POLL_INTERVAL_MS` | `2000` | `open-sse/handlers/imageGeneration/providers/designerWeb.ts` | microsoft-designer-web job polling frequency. |
| `ADOBE_FIREFLY_SUBMIT_BASE_DELAY_MS` | `8000` | `open-sse/services/adobeFireflyUpscale.ts` | Base delay for the Adobe Firefly upscale submit-retry exponential backoff. |
| `AWS_REGION` | _(unset)_ | `src/lib/providers/validation.ts`, `open-sse/handlers/audioSpeech.ts` | Region used to construct AWS Bedrock endpoints (Kiro, audio). |
| `AWS_DEFAULT_REGION` | _(unset)_ | `src/lib/providers/validation.ts`, `open-sse/handlers/audioSpeech.ts` | Fallback when `AWS_REGION` is not set. |
| `CLOUDFLARE_ACCOUNT_ID` | _(unset)_ | `open-sse/executors/cloudflare-ai.ts` | Account ID for Cloudflare Workers AI. |
@@ -891,10 +874,6 @@ Reverse-engineered session bridge for hyperagent.com (`src/shared/constants/prov
| `CLIPROXYAPI_PORT` | `5544` | `open-sse/executors/cliproxyapi.ts` | CLIProxyAPI bridge port. |
| `CLIPROXYAPI_CONFIG_DIR` | `~/.cli-proxy-api` | `src/lib/versionManager/processManager.ts` | CLIProxyAPI config directory. |
| `MUX_SERVICE_PORT` | `8322` | `src/lib/services/bootstrap.ts` | Override the port where the embedded Mux (coder/mux) agent-orchestration daemon listens (always 127.0.0.1). |
| `DARIO_HOST` | `127.0.0.1` | `open-sse/executors/dario.ts` | Dario embedded-service bind/connect host (loopback only by default). |
| `DARIO_PORT` | `3456` | `open-sse/executors/dario.ts` | Dario embedded-service port. |
| `DARIO_HOST` | `127.0.0.1` | `open-sse/executors/dario.ts` | Dario embedded-service bind/connect host (loopback only by default). |
| `DARIO_PORT` | `3456` | `open-sse/executors/dario.ts` | Dario embedded-service port. |
| `LOCAL_HOSTNAMES` | _(empty)_ | `open-sse/config/providerRegistry.ts` | Comma-separated additional hostnames treated as "local" (Docker service names, etc.). |
`ENABLE_CC_COMPATIBLE_PROVIDER` is only for third-party relays that accept Claude Code clients
@@ -1175,13 +1154,6 @@ Provider quota endpoints, network tunnels (Tailscale, Ngrok, MITM debug proxy),
| `OMNIROUTE_LOCAL_ENDPOINTS_TOKEN` | _(unset)_ | `src/lib/security/localEndpoints.ts` | Bearer token for `/api/local/*` callers that aren't on loopback (e.g. the desktop app). When set, requests from non-loopback IPs must carry `Authorization: Bearer <token>`. Required when `OMNIROUTE_LOCAL_ENDPOINTS_ENABLED=1` in non-loopback deployments. |
| `OMNIROUTE_REDIS_CONTAINER_NAME` | `omniroute-redis` | `bin/cli/commands/redis.mjs` | Container name for the 1-click Redis launcher (`omniroute redis up`). Used by both the CLI and the `RedisLauncherPanel` GUI. |
| `OMNIROUTE_REDIS_HOST_PORT` | `6379` | `bin/cli/commands/redis.mjs` | Host port for the 1-click Redis launcher. Bump if the host already binds 6379. The container's internal port stays 6379. |
| `OMNIROUTE_REDIS_BIND_HOST` | `127.0.0.1` | `bin/cli/commands/redis.mjs` | Host interface the 1-click Redis launcher publishes on. The launcher starts Redis WITHOUT a password, so binding `0.0.0.0` hands every host on your LAN an unauthenticated Redis — only widen this if you also set a password on the instance yourself. |
| `REDIS_BIND_HOST` | `127.0.0.1` | `docker-compose.yml` | Host interface docker-compose publishes the Redis sidecar on (#9286). The compose Redis runs without `requirepass`; app containers reach it over the compose network (`redis:6379`) — the published port exists only for host-side tooling. `0.0.0.0` exposes an unauthenticated Redis to the whole LAN. |
| `REDIS_PORT` | `6379` | `docker-compose.yml` | Host port for the compose Redis sidecar. |
| `OMNIROUTE_INTERNAL_SERVICE_TOKEN` | _(unset — mechanism disabled)_ | `src/lib/api/internalServiceAuth.ts` | Shared secret for identity-preserving internal REST hops (#9260): OmniRoute components calling other local OmniRoute routes send it as `x-omniroute-internal-service-token` so the original caller identity is preserved. Compared with `timingSafeEqual`. |
| `OMNIROUTE_INTERNAL_SERVICE_TOKEN_FILE` | _(unset)_ | `src/lib/api/internalServiceAuth.ts` | Secret-file variant of the internal service token: path to a file whose trimmed content is the token. Only consulted when the inline var is unset. |
| `OPENROUTER_PROVIDER_STATS_ENABLED` | `true` | `src/lib/catalog/openrouterProviderStats.ts` | Enrich the dashboard providers list with OpenRouter weekly ranking stats (#9324). On by default; set `false` to skip the background fetch entirely (non-blocking, never fatal). |
| `OPENROUTER_PROVIDER_STATS_TTL_MS` | `86400000` (24h) | `src/lib/catalog/openrouterProviderStats.ts` | Cache TTL for the OpenRouter provider-stats snapshot, in milliseconds. |
| `OMNIROUTE_REDIS_IMAGE` | `redis:7-alpine` | `bin/cli/commands/redis.mjs` | Redis image used by the 1-click Redis launcher. Override to `redis:8-alpine` or a private registry mirror as needed. |
| `QDRANT_HOST` | `qdrant` | _(opt-in cluster profile)_ | Hostname of the Qdrant sidecar when `--profile memory` is active. Default points to the in-network qdrant service name; override for an external deployment. Only consumed when `qdrantEnabled` is `true` in code (`src/lib/memory/vectorStore.ts:108`). |
| `QDRANT_PORT` | `6333` | _(opt-in cluster profile)_ | REST port of the Qdrant sidecar. |
@@ -1207,17 +1179,6 @@ Provider quota endpoints, network tunnels (Tailscale, Ngrok, MITM debug proxy),
| `OMNIROUTE_ROTATE_400_THRESHOLD` | `1` | `open-sse/services/rotationConfig.ts` | Number of `400` errors within `OMNIROUTE_ROTATE_400_WINDOW_SECONDS` required before the account is rotated (only consulted when `OMNIROUTE_ROTATE_ON_400=true`). |
| `OMNIROUTE_ROTATE_400_WINDOW_SECONDS` | `120` | `open-sse/services/rotationConfig.ts` | Sliding window (seconds) over which `400` errors are counted toward `OMNIROUTE_ROTATE_400_THRESHOLD`. |
### Claude Warmup Scheduler
Cron-driven warmup for opted-in Anthropic OAuth connections, so the 5-hour rate-limit window is opened by a trivial scheduled request instead of by the first real one (#8848). The scheduler is off unless `OMNIROUTE_WARMUP_ENABLED` is truthy **and** the connection is flagged in `settings.claudeWarmup.connections`; an empty connection list means nothing is warmed even with the env var on.
| Variable | Default | Source File | Description |
| ----------------------------- | -------------------------------------- | ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `OMNIROUTE_WARMUP_ENABLED` | _(unset → off)_ | `src/lib/warmupScheduler.ts` | Master switch for the warmup scheduler. Accepts `1`/`true`/`yes`/`on` (case-insensitive, trimmed). Any other value, or unset, leaves the scheduler off. |
| `OMNIROUTE_WARMUP_CRON` | `0 7 * * *` | `src/lib/warmupScheduler.ts` | Five-field cron expression for the warmup tick, evaluated in `America/Los_Angeles` (Anthropic's reset timezone) regardless of the host clock. |
| `OMNIROUTE_WARMUP_CONCURRENCY` | `3` | `src/lib/warmupScheduler.ts` | How many connections are warmed in parallel per tick. Clamped to `1`-`10`; a non-numeric value falls back to `3`. |
| `OMNIROUTE_WARMUP_MODEL` | `claude-3-5-haiku-20241022` | `src/lib/warmupScheduler.ts` | Model used for the warmup request. Override only if the default is unavailable on your plan; pick the cheapest model that still opens the window. |
### Browser-Login VNC Sessions & Data-Dir Alias
Containerized Chromium+VNC used for interactive browser-login credential capture (`/api/vnc-session`), plus a legacy `DATA_DIR` alias. All optional — the VNC defaults target the bundled `omniroute-vnc-chromium:local` image and are only overridden for a custom container image, ports, or lifecycle tuning.
@@ -1282,25 +1243,6 @@ that should be able to run the docs translator.
---
## 27. Radar Feed (Self-Hosting)
Optional add-on gated by the RADAR_ENABLED feature flag (default off — a feature
flag toggled via Settings/DB, not an env var; see
[docs/frameworks/RADAR.md](../frameworks/RADAR.md#flag-radar_enabled-default-off)).
The four variables below are optional overrides used only to point the client at a
self-hosted or forked feed / supporter-key flow instead of the default OmniRoute
Radar service. See [docs/frameworks/RADAR.md](../frameworks/RADAR.md) for the full
module doc.
| Variable | Default | Source File | Description |
| -------------------------------- | --------------------------------------------------- | ------------------------------ | ------------------------------------------------------------------------------------------------ |
| `RADAR_FEED_URL` | `https://radar.omniroute.online` | `src/lib/radar/sync.ts` | Base URL of the Radar feed service. Override to point at a self-hosted or forked feed. |
| `RADAR_FEED_PUBKEY` | _(pinned default key)_ | `src/lib/radar/pinnedKeys.ts` | Ed25519 public key (base64-DER SPKI or PEM) used to verify feed signatures from a custom feed. |
| `RADAR_CONTRIBUTOR_CLAIM_URL` | `https://radar.omniroute.online/auth/github` | `src/lib/radar/links.ts` | URL the "I'm a contributor" dashboard button opens (GitHub OAuth supporter-key claim flow). |
| `RADAR_SUPPORTER_PLANS_URL` | `https://radar.omniroute.online/planos` | `src/lib/radar/links.ts` | URL the "Support the project" dashboard button opens (payment/plans page). |
---
## Audit: Removed / Dead Variables
The following variables appeared in previous versions of `.env.example` but have **no runtime references** in the current codebase. They have been removed:
@@ -1368,24 +1310,13 @@ Used by `src/lib/vncSession/manifest.ts` to configure Docker-based headless Chro
| `OMNIROUTE_VNC_HARVEST_MS` | `20000` | `src/lib/vncSession/manifest.ts` | Harvest/cleanup timeout (ms). |
| `VIBEPROXY_DATA_DIR` | _(unset)_ | `open-sse/services/notionThreadSessions.ts` | Directory for Notion thread session persistence. |
### Internal service auth
### Telegram Mini App
| Variable | Default | Description |
| --- | --- | --- |
| `OMNIROUTE_INTERNAL_SERVICE_TOKEN` | | Inline token for management-plane service-to-service authentication. |
| `OMNIROUTE_INTERNAL_SERVICE_TOKEN_FILE` | | Path to a file containing the internal service token (preferred in containers; overrides the inline variable). |
Used by `src/lib/telegram/*` and `src/app/api/telegram/update/route.ts` for the inbound bot webhook and Mini App chat proxy. All optional — the endpoint returns 503 when `TELEGRAM_BOT_TOKEN` is unset.
### OpenRouter provider stats
| Variable | Default | Description |
| --- | --- | --- |
| `OPENROUTER_PROVIDER_STATS_ENABLED` | `true` | Set to `false` to skip fetching OpenRouter per-provider stats for catalog enrichment. |
| `OPENROUTER_PROVIDER_STATS_TTL_MS` | `3600000` | Cache TTL (ms) for the fetched OpenRouter provider stats. |
### Embedded Redis binding
| Variable | Default | Description |
| --- | --- | --- |
| `REDIS_BIND_HOST` | `127.0.0.1` | Bind address for the embedded Redis service. |
| `REDIS_PORT` | `6379` | Port for the embedded Redis service. |
| `OMNIROUTE_REDIS_BIND_HOST` | | OmniRoute-scoped override for the embedded Redis bind address. |
| Variable | Default | Source File | Description |
| ------------------------------ | -------------------------- | ---------------------------------------- | --------------------------------------------------------------------------------------- |
| `TELEGRAM_BOT_TOKEN` | _(unset)_ | `src/lib/telegram/config.ts` | Bot token from @BotFather (`<numeric_id>:<secret>`). Enables the inbound webhook; doubles as the HMAC secret for Mini App `initData` verification. |
| `TELEGRAM_DEFAULT_MODEL` | `auto/chat` | `src/lib/telegram/chatProxy.ts` | Model used for Telegram chat replies. |
| `TELEGRAM_BOT_API_BASE` | `https://api.telegram.org` | `src/lib/telegram/config.ts` | Bot API base URL override (proxies / self-hosted Bot API servers). |
| `TELEGRAM_WEBHOOK_TIMEOUT_MS` | `60000` | `src/lib/telegram/config.ts` | Timeout (ms) for outbound Bot API calls (`sendMessage`/`setWebhook`). |

View File

@@ -0,0 +1,113 @@
# Video Generation Through Preset Jobs
Custom provider nodes whose `/videos` surface is an **async submit → poll → fetch-result API** (instead of a synchronous generation endpoint) can be wired into the `/api/v1/videos/generations` route without any new provider code. The model row carries a `generationConfig.preset`, and the dispatcher routes the request through a single job executor that is configured entirely by declarative preset data.
## How dispatch works
1. The route parses `model` as `provider/model` and resolves the provider node's credentials (`POST /api/v1/videos/generations`).
2. `handleVideoGeneration` (in `open-sse/handlers/videoGeneration.ts`) checks whether the provider is a **custom provider node** (no entry in the static video registry).
3. For custom nodes it reads the custom model row via `getCustomModelVideoPreset(provider, model)`:
- The model row has `generationConfig.preset` set (e.g. `"agnes-video-job"`) → dispatch through the **job executor** (`open-sse/handlers/videoGeneration/job.ts`).
- The preset name does not match any known preset → **502** `Unknown video job preset: <preset>` (server-side misconfiguration).
- No preset configured → fall back to the generic OpenAI-compatible sync handler, mirroring the images route.
4. The job executor runs the preset pipeline: **submit** the job, **poll** for terminal status, **read** the finished video URL, and return the standard OpenAI-compatible response shape.
The executor is one handler family; every provider-specific detail (paths, auth, body shape, status/result fields, poll cadence) is data in the preset definition.
## Response contract
Both the sync and job paths return the same shape:
```json
{
"created": 1234567890,
"data": [{ "url": "https://…", "format": "mp4" }]
}
```
This is the shape the media-generation consumer reads (`data.data[0].url`), so preset-job providers are drop-in replacements for sync providers.
## Presets
Presets live in `open-sse/handlers/videoGeneration/job.ts` (`VIDEO_JOB_PRESETS`). Each preset declares:
| Field | Meaning |
| -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `authHeaderName` / `authScheme` | `x-api-key` with `raw` value (Agnes, muapi) or `Authorization` with `Bearer` prefix (Sora). Missing credentials → request goes out without an auth header. |
| `baseUrlFallback` | Default base URL. Overridden by the provider connection's `providerSpecificData.baseUrl` (or top-level `baseUrl`), which wins when set. |
| `submit.path` / `submit.buildBody` | Where and how the job is submitted. `{model}` in the path is substituted with the encoded model id; the body is built from `model`/`prompt`/`duration` plus pass-through of every other request field. |
| `taskIdPath` | Dot path into the submit response identifying the job (e.g. `task_id`, `request_id`, `id`). Missing job id → **502**. |
| `poll.pathTemplate` | Poll URL template; `{taskId}` is substituted. |
| `statusPath` / `statusDone` / `statusFailed` | Where the job status lives and which values are terminal. |
| `resultPath` | Dot path into the poll response holding the finished video URL: a string, a string array, or an array of `{ url }` objects are all accepted. Completed job with no readable URL → **502**. |
| `maxPolls` / `pollIntervalMs` | Poll budget (default 60 polls × 2000 ms). Exhausted → **504** `Video job timed out`. |
### `agnes-video-job` — Agnes Video V2.0
- Auth: `x-api-key: <key>` (raw).
- Base URL fallback: `https://apihub.agnes-ai.com`.
- Submit: `POST /v1/videos` with `{ model, prompt, ...extras }` — image, mode, `num_frames`, `frame_rate` and other provider knobs pass through untouched.
- Job id: `task_id` from the submit response.
- Poll: `GET /v1/videos/{taskId}`; status at `status` (`completed` / `failed`).
- Result: `metadata.url` — the completed video URL is returned as JSON metadata, not a binary body.
### `muapi-video-job` — muapi.ai
- Auth: `x-api-key: <key>` (raw).
- Base URL fallback: `https://api.muapi.ai`.
- Submit: `POST /api/v1/{model}` with `{ prompt, duration?, ...extras }`.
- Job id: `request_id` from the submit response.
- Poll: `GET /api/v1/predictions/{taskId}/result`; status at `status` (`completed` / `failed`).
- Result: `outputs` — an array of video URLs.
### `sora-job` — OpenAI Sora
- Auth: `Authorization: Bearer <key>`.
- Base URL fallback: `https://api.openai.com`.
- Submit: `POST /v1/videos` with `{ model, prompt, seconds?, ...extras }`. `seconds` is a **string** enum (`"4" | "8" | "12"`) in the Sora API, so a numeric `duration` is stringified; size mapping is intentionally not forced.
- Job id: `id` from the submit response.
- Poll: `GET /v1/videos/{taskId}`; status at `status` (`completed` / `failed`).
- Result: `data` — an array whose entries are either a URL string or `{ url: "…" }`.
## Setup
1. **Register the provider node** as an OpenAI-compatible custom provider (`providerSpecificData.baseUrl` optional — the preset's `baseUrlFallback` is used when absent).
2. **Register a custom model** tagged with the `videos` endpoint and a `generationConfig`:
```json
{
"id": "super-video-v1",
"name": "Super Video v1",
"source": "manual",
"apiFormat": "chat-completions",
"supportedEndpoints": ["videos"],
"generationConfig": { "preset": "agnes-video-job" }
}
```
`addCustomModel` (in `src/lib/db/models.ts`) accepts `generationConfig?: { preset: string }` as its final parameter and persists it on the model row; `updateCustomModel` forwards it the same way. The provider-models API accepts `generationConfig` on create and update.
3. **Call the route** as usual:
```bash
curl -X POST http://localhost:8787/api/v1/videos/generations \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $API_KEY" \
-d '{
"model": "my-custom-provider/super-video-v1",
"prompt": "a cat playing piano",
"duration": 5
}'
```
## Troubleshooting
| Symptom | Cause |
| ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------ |
| `400 Unknown video provider: …` | Non-custom provider not in the static registry; preset jobs only apply to custom provider nodes. |
| `502 Unknown video job preset: …` | `generationConfig.preset` does not match any preset in `VIDEO_JOB_PRESETS`. Fix the model row. |
| `502 Video provider did not return a job id (…)` | Submit succeeded but the response had no readable value at `taskIdPath`. |
| `502 Video job failed (…)` / `Video job completed but no result URL found (…)` | Poll reached a terminal `statusFailed` state, or `resultPath` held no readable URL. |
| `504 Video job timed out after 60 polls (…)` | Job never reached a terminal status within the poll budget. |
| Upstream 4xx/5xx passthrough | `fetchJson` returns the upstream status when the submit/poll request itself is not OK. |
| Requests go out without auth | No `apiKey`/`accessToken` on the provider connection; the executor sends `Content-Type` only. |

View File

@@ -55,9 +55,9 @@
"license": "MIT"
},
"node_modules/@electron/asar/node_modules/brace-expansion": {
"version": "1.1.16",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.16.tgz",
"integrity": "sha512-IDw48K2/2kRkg9LdJxurvq3lV3aBgq0REY89duEqFRthjlPdXHKMj7EnQOXVckxzgisinf3nHfrcE2FufFLXMw==",
"version": "1.1.18",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.18.tgz",
"integrity": "sha512-Edep/X9fGqVNmzKBVsDYIOtD+z1tuezV70LBjdCst9Tqu76lsnvRiZ6oTic1n+/BIwX6QDGAO94PN4N2SADvtw==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -257,9 +257,9 @@
"license": "MIT"
},
"node_modules/@electron/universal/node_modules/brace-expansion": {
"version": "2.1.2",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.2.tgz",
"integrity": "sha512-w5JZcKgdhDOgOwm8H+KgbosopHMuGcl6qbulwjtz3SM7I7P3yW1eAjzMPLrIE+NQ9vjgANKHWeMHnrT0OXW1oA==",
"version": "2.1.4",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.4.tgz",
"integrity": "sha512-hGfVzPxthbf3+2yjg/RBs60cB0FhqBS/zvdV/4wn4/BmN0bNMMHPc4V/BbFieqf1TKAGGAHnY4eSjajCl0f2Xg==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -297,6 +297,45 @@
"url": "https://github.com/sponsors/isaacs"
}
},
"node_modules/@electron/windows-sign": {
"version": "1.2.2",
"resolved": "https://registry.npmjs.org/@electron/windows-sign/-/windows-sign-1.2.2.tgz",
"integrity": "sha512-dfZeox66AvdPtb2lD8OsIIQh12Tp0GNCRUDfBHIKGpbmopZto2/A8nSpYYLoedPIHpqkeblZ/k8OV0Gy7PYuyQ==",
"dev": true,
"license": "BSD-2-Clause",
"optional": true,
"peer": true,
"dependencies": {
"cross-dirname": "^0.1.0",
"debug": "^4.3.4",
"fs-extra": "^11.1.1",
"minimist": "^1.2.8",
"postject": "^1.0.0-alpha.6"
},
"bin": {
"electron-windows-sign": "bin/electron-windows-sign.js"
},
"engines": {
"node": ">=14.14"
}
},
"node_modules/@electron/windows-sign/node_modules/fs-extra": {
"version": "11.4.0",
"resolved": "https://registry.npmjs.org/fs-extra/-/fs-extra-11.4.0.tgz",
"integrity": "sha512-EQsFzMUJkCKGr1ePqlYADkIUmHW1s3ZXr5Yqy6wbGrfUCphpl2maM/kyOIRA2HpP3AaFQTZXD4ldjek+nccddA==",
"dev": true,
"license": "MIT",
"optional": true,
"peer": true,
"dependencies": {
"graceful-fs": "^4.2.0",
"jsonfile": "^6.0.1",
"universalify": "^2.0.0"
},
"engines": {
"node": ">=14.14"
}
},
"node_modules/@isaacs/fs-minipass": {
"version": "4.0.1",
"resolved": "https://registry.npmjs.org/@isaacs/fs-minipass/-/fs-minipass-4.0.1.tgz",
@@ -835,16 +874,16 @@
"optional": true
},
"node_modules/brace-expansion": {
"version": "5.0.7",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.7.tgz",
"integrity": "sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==",
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
"node": "20 || >=22"
}
},
"node_modules/buffer-from": {
@@ -1091,6 +1130,15 @@
"dev": true,
"license": "MIT"
},
"node_modules/cross-dirname": {
"version": "0.1.0",
"resolved": "https://registry.npmjs.org/cross-dirname/-/cross-dirname-0.1.0.tgz",
"integrity": "sha512-+R08/oI0nl3vfPcqftZRpytksBXDzOUveBq/NBVx0sUp1axwzPQrKinNx5yd5sxPu8j1wIy8AfnVQ+5eFdha6Q==",
"dev": true,
"license": "MIT",
"optional": true,
"peer": true
},
"node_modules/cross-spawn": {
"version": "7.0.6",
"resolved": "https://registry.npmjs.org/cross-spawn/-/cross-spawn-7.0.6.tgz",
@@ -1260,9 +1308,9 @@
"license": "MIT"
},
"node_modules/dir-compare/node_modules/brace-expansion": {
"version": "1.1.16",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.16.tgz",
"integrity": "sha512-IDw48K2/2kRkg9LdJxurvq3lV3aBgq0REY89duEqFRthjlPdXHKMj7EnQOXVckxzgisinf3nHfrcE2FufFLXMw==",
"version": "1.1.18",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.18.tgz",
"integrity": "sha512-Edep/X9fGqVNmzKBVsDYIOtD+z1tuezV70LBjdCst9Tqu76lsnvRiZ6oTic1n+/BIwX6QDGAO94PN4N2SADvtw==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -1411,6 +1459,19 @@
"node": ">=14.0.0"
}
},
"node_modules/electron-builder-squirrel-windows": {
"version": "26.15.3",
"resolved": "https://registry.npmjs.org/electron-builder-squirrel-windows/-/electron-builder-squirrel-windows-26.15.3.tgz",
"integrity": "sha512-Jc19XPV9y9+2bAdZPkXuVNGNIEFBq9poHC61l8Kv6FdK7DRG3+Ic0rerC0DXOaeHNz8yW0fg/JnF8GQROOF5MA==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"app-builder-lib": "26.15.3",
"builder-util": "26.15.3",
"electron-winstaller": "5.4.0"
}
},
"node_modules/electron-publish": {
"version": "26.15.3",
"resolved": "https://registry.npmjs.org/electron-publish/-/electron-publish-26.15.3.tgz",
@@ -1445,6 +1506,66 @@
"tiny-typed-emitter": "^2.1.0"
}
},
"node_modules/electron-winstaller": {
"version": "5.4.0",
"resolved": "https://registry.npmjs.org/electron-winstaller/-/electron-winstaller-5.4.0.tgz",
"integrity": "sha512-bO3y10YikuUwUuDUQRM4KfwNkKhnpVO7IPdbsrejwN9/AABJzzTQ4GeHwyzNSrVO+tEH3/Np255a3sVZpZDjvg==",
"dev": true,
"hasInstallScript": true,
"license": "MIT",
"peer": true,
"dependencies": {
"@electron/asar": "^3.2.1",
"debug": "^4.1.1",
"fs-extra": "^7.0.1",
"lodash": "^4.17.21",
"temp": "^0.9.0"
},
"engines": {
"node": ">=8.0.0"
},
"optionalDependencies": {
"@electron/windows-sign": "^1.1.2"
}
},
"node_modules/electron-winstaller/node_modules/fs-extra": {
"version": "7.0.1",
"resolved": "https://registry.npmjs.org/fs-extra/-/fs-extra-7.0.1.tgz",
"integrity": "sha512-YJDaCJZEnBmcbw13fvdAM9AwNOJwOzrE4pqMqBq5nFiEqXUqHwlK4B+3pUw6JNvfSPtX05xFHtYy/1ni01eGCw==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"graceful-fs": "^4.1.2",
"jsonfile": "^4.0.0",
"universalify": "^0.1.0"
},
"engines": {
"node": ">=6 <7 || >=8"
}
},
"node_modules/electron-winstaller/node_modules/jsonfile": {
"version": "4.0.0",
"resolved": "https://registry.npmjs.org/jsonfile/-/jsonfile-4.0.0.tgz",
"integrity": "sha512-m6F1R3z8jjlf2imQHS2Qez5sjKWQzbuuhuJ/FKYFRZvPE3PuHcSMVZzfsLhGVOkfd20obL5SWEBew5ShlquNxg==",
"dev": true,
"license": "MIT",
"peer": true,
"optionalDependencies": {
"graceful-fs": "^4.1.6"
}
},
"node_modules/electron-winstaller/node_modules/universalify": {
"version": "0.1.2",
"resolved": "https://registry.npmjs.org/universalify/-/universalify-0.1.2.tgz",
"integrity": "sha512-rBJeI5CXAlmy1pV+617WB9J63U6XcazHHF2f2dbJix4XzpUF0RS3Zbj0FGIOCAva5P/d/GBOYaACQ1w+0azUkg==",
"dev": true,
"license": "MIT",
"peer": true,
"engines": {
"node": ">= 4.0.0"
}
},
"node_modules/emoji-regex": {
"version": "8.0.0",
"resolved": "https://registry.npmjs.org/emoji-regex/-/emoji-regex-8.0.0.tgz",
@@ -1627,9 +1748,9 @@
"license": "MIT"
},
"node_modules/filelist/node_modules/brace-expansion": {
"version": "2.1.2",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.2.tgz",
"integrity": "sha512-w5JZcKgdhDOgOwm8H+KgbosopHMuGcl6qbulwjtz3SM7I7P3yW1eAjzMPLrIE+NQ9vjgANKHWeMHnrT0OXW1oA==",
"version": "2.1.4",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.4.tgz",
"integrity": "sha512-hGfVzPxthbf3+2yjg/RBs60cB0FhqBS/zvdV/4wn4/BmN0bNMMHPc4V/BbFieqf1TKAGGAHnY4eSjajCl0f2Xg==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -1792,9 +1913,9 @@
"license": "MIT"
},
"node_modules/glob/node_modules/brace-expansion": {
"version": "1.1.16",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.16.tgz",
"integrity": "sha512-IDw48K2/2kRkg9LdJxurvq3lV3aBgq0REY89duEqFRthjlPdXHKMj7EnQOXVckxzgisinf3nHfrcE2FufFLXMw==",
"version": "1.1.18",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.18.tgz",
"integrity": "sha512-Edep/X9fGqVNmzKBVsDYIOtD+z1tuezV70LBjdCst9Tqu76lsnvRiZ6oTic1n+/BIwX6QDGAO94PN4N2SADvtw==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -2113,9 +2234,9 @@
}
},
"node_modules/js-yaml": {
"version": "4.3.0",
"resolved": "https://registry.npmjs.org/js-yaml/-/js-yaml-4.3.0.tgz",
"integrity": "sha512-1td788aAnnZ5qs7V2QIRl1owjtYpbKt749Y3xauqQgwIIGF/xXWz1wMTEBx5O3LK3lXLVuqXPdPxj2BoFHaW9Q==",
"version": "4.3.1",
"resolved": "https://registry.npmjs.org/js-yaml/-/js-yaml-4.3.1.tgz",
"integrity": "sha512-CY6crGq313MX8GkwvB7tzgp99vjQxY1++5y10/BKN/GUfHqWaOGQMNZkBvqSzsZKWk/ijwHlWzzkLulsGHhjWQ==",
"funding": [
{
"type": "github",
@@ -2359,6 +2480,20 @@
"node": ">= 18"
}
},
"node_modules/mkdirp": {
"version": "0.5.6",
"resolved": "https://registry.npmjs.org/mkdirp/-/mkdirp-0.5.6.tgz",
"integrity": "sha512-FP+p8RB8OWpF3YZBCrP5gtADmtXApB5AMLn+vdyA+PyxCjrCs00mjyUozssO33cwDeT3wNGdLxJ5M//YqtHAJw==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"minimist": "^1.2.6"
},
"bin": {
"mkdirp": "bin/cmd.js"
}
},
"node_modules/ms": {
"version": "2.1.3",
"resolved": "https://registry.npmjs.org/ms/-/ms-2.1.3.tgz",
@@ -2622,6 +2757,36 @@
"node": ">=18"
}
},
"node_modules/postject": {
"version": "1.0.0-alpha.6",
"resolved": "https://registry.npmjs.org/postject/-/postject-1.0.0-alpha.6.tgz",
"integrity": "sha512-b9Eb8h2eVqNE8edvKdwqkrY6O7kAwmI8kcnBv1NScolYJbo59XUF0noFq+lxbC1yN20bmC0WBEbDC5H/7ASb0A==",
"dev": true,
"license": "MIT",
"optional": true,
"peer": true,
"dependencies": {
"commander": "^9.4.0"
},
"bin": {
"postject": "dist/cli.js"
},
"engines": {
"node": ">=14.0.0"
}
},
"node_modules/postject/node_modules/commander": {
"version": "9.5.0",
"resolved": "https://registry.npmjs.org/commander/-/commander-9.5.0.tgz",
"integrity": "sha512-KRs7WVDKg86PWiuAqhDrAQnTXZKraVcCc6vFdL14qrZ/DcWwuRo7VoiYXalXO7S5GKpqYiVEwCbgFDfxNHKJBQ==",
"dev": true,
"license": "MIT",
"optional": true,
"peer": true,
"engines": {
"node": "^12.20.0 || >=14"
}
},
"node_modules/proc-log": {
"version": "6.1.0",
"resolved": "https://registry.npmjs.org/proc-log/-/proc-log-6.1.0.tgz",
@@ -2816,6 +2981,21 @@
"node": ">= 4"
}
},
"node_modules/rimraf": {
"version": "2.6.3",
"resolved": "https://registry.npmjs.org/rimraf/-/rimraf-2.6.3.tgz",
"integrity": "sha512-mwqeW5XsA2qAejG46gYdENaxXjx9onRNCfn7L0duuP4hCuTIi/QO7PDK07KJfp1d+izWPrzEJDcSqBa0OZQriA==",
"deprecated": "Rimraf versions prior to v4 are no longer supported",
"dev": true,
"license": "ISC",
"peer": true,
"dependencies": {
"glob": "^7.1.3"
},
"bin": {
"rimraf": "bin.js"
}
},
"node_modules/roarr": {
"version": "2.15.4",
"resolved": "https://registry.npmjs.org/roarr/-/roarr-2.15.4.tgz",
@@ -3045,9 +3225,9 @@
}
},
"node_modules/tar": {
"version": "7.5.20",
"resolved": "https://registry.npmjs.org/tar/-/tar-7.5.20.tgz",
"integrity": "sha512-9FcyK4PA6+WbzlTM9WhQm6vB5W7cP7dUiPsv1g7YDwEQnQ1CGpK3MGlKk/ITVWMk05kHZuBhmVhiv8LZoy/PFQ==",
"version": "7.5.22",
"resolved": "https://registry.npmjs.org/tar/-/tar-7.5.22.tgz",
"integrity": "sha512-MFO/QzvtAOmJbkhOaCTvbGcFN9L9b+JunIsDwaKljSOdcLMea3NJ1k9Usz/rjdfSXTq4dfzfeS7W4p4YOAAHeA==",
"dev": true,
"license": "BlueOak-1.0.0",
"dependencies": {
@@ -3071,6 +3251,21 @@
"node": ">=18"
}
},
"node_modules/temp": {
"version": "0.9.4",
"resolved": "https://registry.npmjs.org/temp/-/temp-0.9.4.tgz",
"integrity": "sha512-yYrrsWnrXMcdsnu/7YMYAofM1ktpL5By7vZhf15CrXijWWrEYZks5AXBudalfSWJLlnen/QUJUB5aoB0kqZUGA==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"mkdirp": "^0.5.1",
"rimraf": "~2.6.2"
},
"engines": {
"node": ">=6.0.0"
}
},
"node_modules/temp-file": {
"version": "3.4.0",
"resolved": "https://registry.npmjs.org/temp-file/-/temp-file-3.4.0.tgz",

View File

@@ -18,6 +18,15 @@ export const FETCH_TIMEOUT_MS = upstreamTimeouts.fetchTimeoutMs;
// idle for this duration. Override with STREAM_IDLE_TIMEOUT_MS env var.
export const STREAM_IDLE_TIMEOUT_MS = upstreamTimeouts.streamIdleTimeoutMs;
// Grace period (ms) a client-disconnect finalization waits for the stream's own
// completion bookkeeping to land before persisting a 499. See #9653 — a client
// that closes right after reading a fully-completed SSE stream can otherwise
// race OmniRoute's own completion callback, resulting in a false 499 with zero
// token usage for a request that actually delivered its full response. Set
// STREAM_DISCONNECT_GRACE_PERIOD_MS=0 to disable and restore the old
// immediate-fail behavior.
export const STREAM_DISCONNECT_GRACE_PERIOD_MS = upstreamTimeouts.streamDisconnectGracePeriodMs;
// Timeout for the first non-ping SSE event. Inherits REQUEST_TIMEOUT_MS when
// set, unless STREAM_READINESS_TIMEOUT_MS is specified directly. This must stay
// conservative for large prompts and slow first-byte reasoning providers.
@@ -65,27 +74,27 @@ export const PROVIDERS: Record<string, LegacyProvider> = new Proxy(
{} as Record<string, LegacyProvider>,
{
get(_, prop) {
if (typeof prop === 'symbol') return undefined;
if (typeof prop === "symbol") return undefined;
return Reflect.get(initProviders(), prop, _providers);
},
has(_, prop) {
if (typeof prop === 'symbol') return false;
if (typeof prop === "symbol") return false;
return Reflect.has(initProviders(), prop);
},
ownKeys() {
return Reflect.ownKeys(initProviders());
},
getOwnPropertyDescriptor(_, prop) {
if (typeof prop === 'symbol') return undefined;
if (typeof prop === "symbol") return undefined;
return Object.getOwnPropertyDescriptor(initProviders(), prop);
},
set(_, prop, value) {
if (typeof prop === 'symbol') return false;
if (typeof prop === "symbol") return false;
(initProviders() as Record<string, LegacyProvider>)[prop] = value;
return true;
},
deleteProperty(_, prop) {
if (typeof prop === 'symbol') return false;
if (typeof prop === "symbol") return false;
return Reflect.deleteProperty(initProviders(), prop);
},
}
@@ -124,6 +133,11 @@ export const OAUTH_ENDPOINTS = {
auth: "https://github.com/login/oauth/authorize",
deviceCode: "https://github.com/login/device/code",
},
openference: {
token: "https://openference.com/oauth/token",
auth: "https://openference.com/app/oauth/authorize",
clientId: "omniroute",
},
};
// Cache TTLs (seconds)

View File

@@ -311,7 +311,6 @@ export const FREE_MODEL_BUDGETS: FreeModelBudget[] = [
{ provider: "nscale", modelId: "openai/gpt-oss-20b", displayName: "openai/gpt-oss-20b", monthlyTokens: 0, creditTokens: 5000000, freeType: "one-time-initial", poolKey: "nscale", tos: "caution" },
{ provider: "nscale", modelId: "meta-llama/Llama-4-Scout-17B-16E-Instruct", displayName: "meta-llama/Llama-4-Scout-17B-16E-Instruct", monthlyTokens: 0, creditTokens: 5000000, freeType: "one-time-initial", poolKey: "nscale", tos: "caution" },
{ provider: "nscale", modelId: "meta-llama/Llama-3.3-70B-Instruct", displayName: "meta-llama/Llama-3.3-70B-Instruct", monthlyTokens: 0, creditTokens: 5000000, freeType: "one-time-initial", poolKey: "nscale", tos: "caution" },
{ provider: "nvidia", modelId: "z-ai/glm-5.1", displayName: "GLM 5.1", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "z-ai/glm-5.2", displayName: "GLM 5.2", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "minimaxai/minimax-m2.7", displayName: "MiniMax M2.7", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "google/gemma-4-31b-it", displayName: "Gemma 4 31B", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
@@ -321,7 +320,6 @@ export const FREE_MODEL_BUDGETS: FreeModelBudget[] = [
{ provider: "nvidia", modelId: "qwen/qwen3.5-397b-a17b", displayName: "Qwen3.5-397B-A17B", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "qwen/qwen3.5-122b-a10b", displayName: "Qwen3.5-122B-A10B", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "stepfun-ai/step-3.5-flash", displayName: "Step 3.5 Flash", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "deepseek-ai/deepseek-v4-pro", displayName: "DeepSeek V4 Pro", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "openai/gpt-oss-120b", displayName: "GPT OSS 120B", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "openai/gpt-oss-20b", displayName: "GPT OSS 20B", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },
{ provider: "nvidia", modelId: "nvidia/nemotron-3-super-120b-a12b", displayName: "Nemotron 3 Super 120B A12B", monthlyTokens: 0, creditTokens: 0, freeType: "one-time-initial", poolKey: "nvidia", tos: "caution" },

View File

@@ -1,5 +1,4 @@
[
"deepseek-ai/deepseek-v4-pro",
"google/gemma-4-31b-it",
"minimaxai/minimax-m2.7",
"mistralai/devstral-2-123b-instruct-2512",
@@ -13,6 +12,5 @@
"qwen/qwen3.5-397b-a17b",
"stepfun-ai/step-3.5-flash",
"thinkingmachines/inkling",
"z-ai/glm-5.1",
"z-ai/glm-5.2"
]

View File

@@ -90,10 +90,65 @@ export function getDefaultModel(aliasOrId: string): string | null {
return models?.[0]?.id || null;
}
/** Score a registry entry by how many capability flags it defines. */
function modelRichness(m: RegistryModel): number {
let score = 0;
if (m.supportsXHighEffort !== undefined) score += 10; // critical for effort routing
if (m.supportsReasoning !== undefined) score += 5;
if (m.contextLength !== undefined) score += 3;
if (m.maxOutputTokens !== undefined) score += 2;
if (m.supportsVision !== undefined) score += 2;
if (m.toolCalling !== undefined) score += 2;
if (m.interleavedField !== undefined) score += 1;
if (m.unsupportedParams !== undefined) score += 1;
return score;
}
function getGlobalModel(modelId: string): RegistryModel | undefined {
// 1. Exact match — collect all, pick the richest
let candidates: RegistryModel[] = [];
for (const models of Object.values(PROVIDER_MODELS)) {
const found = models.find((m) => m.id === modelId);
if (found) candidates.push(found);
}
if (candidates.length > 0) {
return candidates.sort((a, b) => modelRichness(b) - modelRichness(a))[0];
}
// 2. Strip provider prefix (e.g. moonshotai/kimi-k3-free -> kimi-k3-free)
const basename = modelId.split("/").pop() || modelId;
candidates = [];
for (const models of Object.values(PROVIDER_MODELS)) {
const found = models.find((m) => m.id === basename);
if (found) candidates.push(found);
}
if (candidates.length > 0) {
return candidates.sort((a, b) => modelRichness(b) - modelRichness(a))[0];
}
// 3. Substring match for base model name (e.g. kimi-k3-free -> kimi-k3)
// Finds the longest matching base model ID; on ties, prefers the richer entry.
let bestMatch: RegistryModel | undefined;
for (const models of Object.values(PROVIDER_MODELS)) {
for (const m of models) {
if (basename.startsWith(m.id)) {
if (
!bestMatch ||
m.id.length > bestMatch.id.length ||
(m.id.length === bestMatch.id.length && modelRichness(m) > modelRichness(bestMatch))
) {
bestMatch = m;
}
}
}
}
return bestMatch;
}
export function getProviderModel(aliasOrId: string, modelId: string): RegistryModel | undefined {
const models = PROVIDER_MODELS[aliasOrId];
if (!models) return undefined;
return models.find((model) => model.id === modelId);
if (!models) return getGlobalModel(modelId);
return models.find((model) => model.id === modelId) || getGlobalModel(modelId);
}
export function isValidModel(
@@ -103,26 +158,20 @@ export function isValidModel(
): boolean {
if (passthroughProviders.has(aliasOrId)) return true;
const models = PROVIDER_MODELS[aliasOrId];
if (!models) return false;
return models.some((m) => m.id === modelId);
if (!models) return !!getGlobalModel(modelId);
return models.some((m) => m.id === modelId) || !!getGlobalModel(modelId);
}
export function findModelName(aliasOrId: string, modelId: string): string {
const models = PROVIDER_MODELS[aliasOrId];
if (!models) return modelId;
const found = models.find((m) => m.id === modelId);
if (!models) return getGlobalModel(modelId)?.name || modelId;
const found = models.find((m) => m.id === modelId) || getGlobalModel(modelId);
return found?.name || modelId;
}
export function getModelTargetFormat(aliasOrId: string, modelId: string): string | null {
const models = PROVIDER_MODELS[aliasOrId];
// Strip provider prefix if present: "openai/gpt-5.6-luna" → "gpt-5.6-luna"
const prefix = aliasOrId + "/";
const bareModelId =
typeof modelId === "string" && modelId.startsWith(prefix)
? modelId.slice(prefix.length)
: modelId;
const found = models?.find((m) => m.id === bareModelId);
const found = models?.find((m) => m.id === modelId) || getGlobalModel(modelId);
if (found?.targetFormat) return found.targetFormat;
// #5842: OpenAI "*-pro" reasoning models (o1-pro, gpt-5.x-pro) are only served by
// the native /v1/responses endpoint — /v1/chat/completions 404s ("only supported
@@ -130,14 +179,17 @@ export function getModelTargetFormat(aliasOrId: string, modelId: string): string
// covers dynamically-synced ids that post-date the catalog (same spirit as the gh
// executor's /codex/i routing, 9router#102). Scoped to the openai alias so other
// providers shipping *-pro ids keep their own endpoint semantics.
if (aliasOrId === "openai" && /-pro$/i.test(bareModelId)) return "openai-responses";
if (aliasOrId === "openai" && /-pro$/i.test(modelId)) return "openai-responses";
return null;
}
export function getModelStripTypes(aliasOrId: string, modelId: string): string[] {
const models = PROVIDER_MODELS[aliasOrId];
if (!models) return [];
const found = models.find((m) => m.id === modelId);
if (!models)
return Array.isArray(getGlobalModel(modelId)?.strip)
? [...getGlobalModel(modelId)!.strip!]
: [];
const found = models.find((m) => m.id === modelId) || getGlobalModel(modelId);
return Array.isArray(found?.strip) ? [...found.strip] : [];
}
@@ -262,7 +314,7 @@ function resolveProviderModelList(aliasOrId: string): {
export function supportsXHighEffort(aliasOrId: string, modelId: string): boolean {
const { models: providerModels } = resolveProviderModelList(aliasOrId);
const model = providerModels?.find((entry) => entry.id === modelId);
const model = providerModels?.find((entry) => entry.id === modelId) || getGlobalModel(modelId);
if (model?.supportsXHighEffort !== undefined) {
return model.supportsXHighEffort !== false;
}

View File

@@ -121,6 +121,8 @@ import { chatgpt_webProvider } from "./registry/chatgpt-web/index.ts";
import { openrouterProvider } from "./registry/openrouter/index.ts";
import { cheaperinferenceProvider } from "./registry/cheaperinference/index.ts";
import { openvectaProvider } from "./registry/openvecta/index.ts";
import { openferenceProvider } from "./registry/openference/index.ts";
import { openference_apiProvider } from "./registry/openference-api/index.ts";
import { orcarouterProvider } from "./registry/orcarouter/index.ts";
import { copilot_webProvider } from "./registry/copilot-web/index.ts";
import { copilot_m365_webProvider } from "./registry/copilot-m365-web/index.ts";
@@ -345,6 +347,8 @@ export const REGISTRY: Record<string, RegistryEntry> = {
openrouter: openrouterProvider,
cheaperinference: cheaperinferenceProvider,
openvecta: openvectaProvider,
openference: openferenceProvider,
"openference-api": openference_apiProvider,
orcarouter: orcarouterProvider,
"copilot-web": copilot_webProvider,
"copilot-m365-web": copilot_m365_webProvider,

View File

@@ -32,8 +32,6 @@ export const nvidiaProvider: RegistryEntry = {
{ id: "qwen/qwen3.5-122b-a10b", name: "Qwen3.5-122B-A10B" },
{ id: "stepfun-ai/step-3.5-flash", name: "Step 3.5 Flash" },
{ id: "stepfun-ai/step-3.7-flash", name: "Step 3.7 Flash" },
{ id: "deepseek-ai/deepseek-v4-pro", name: "DeepSeek V4 Pro", supportsReasoning: true },
{ id: "deepseek-ai/deepseek-v4-flash", name: "DeepSeek V4 Flash", supportsReasoning: true },
// Sweep 2026-06-19: verified present in the live NVIDIA NIM /v1/models catalog.
{ id: "moonshotai/kimi-k2.6", name: "Kimi K2.6" },
{ id: "openai/gpt-oss-120b", name: "GPT OSS 120B", toolCalling: false },

View File

@@ -0,0 +1,18 @@
import type { RegistryEntry } from "../../shared.ts";
import { buildOpenAiCompatibleRegistryEntry } from "../../shared.ts";
/**
* Openference API key — OpenAI-compatible gateway (https://openference.com/).
*
* Bearer API keys (`sk-…`) hit the same api.openference.com/v1/* surface as OAuth
* JWTs. Live model discovery uses NAMED_OPENAI_STYLE_PROVIDERS; the seed below is
* the offline fallback when the live fetch fails.
*/
export const openference_apiProvider: RegistryEntry = buildOpenAiCompatibleRegistryEntry({
id: "openference-api",
alias: "ofa",
baseUrl: "https://api.openference.com/v1/chat/completions",
responsesBaseUrl: "https://api.openference.com/v1/responses",
passthroughModels: true,
models: [{ id: "GLM-5.2", name: "GLM 5.2", contextLength: 850000 }],
});

View File

@@ -0,0 +1,25 @@
import type { RegistryEntry } from "../../shared.ts";
/**
* Openference — OpenAI-compatible AI inference gateway (https://openference.com/).
*
* OAuth access tokens are ES256 JWTs accepted as Bearer credentials on
* api.openference.com/v1/*. Live model discovery uses NAMED_OPENAI_STYLE_PROVIDERS;
* seed models below are the offline fallback when the live fetch fails.
*/
export const openferenceProvider: RegistryEntry = {
id: "openference",
alias: "of",
format: "openai",
executor: "default",
baseUrl: "https://api.openference.com/v1/chat/completions",
responsesBaseUrl: "https://api.openference.com/v1/responses",
authType: "oauth",
authHeader: "bearer",
passthroughModels: true,
oauth: {
clientIdDefault: "omniroute",
tokenUrl: "https://openference.com/oauth/token",
},
models: [{ id: "GLM-5.2", name: "GLM 5.2", contextLength: 850000 }],
};

View File

@@ -0,0 +1,35 @@
import { DefaultExecutor } from "./default.ts";
import type { ProviderCredentials } from "./base.ts";
import { applyAzureParamRules } from "./azureParamRules.ts";
/**
* Azure AI Foundry (`azure-ai`).
*
* URL building, auth headers and the `responses` vs `chat` apiType switch all
* live in `DefaultExecutor`, keyed on the `azure-ai` provider id — this subclass
* inherits them unchanged and adds only the Azure request-param rules.
*
* Before this existed, `azure-ai` fell through to the bare `DefaultExecutor`
* while `azure-openai` had the rules inline, so the same Azure deployment
* behaved differently depending on which connection served it: `azure-openai`
* succeeded and `azure-ai` returned HTTP 400 for `max_tokens` /
* `reasoning_effort`.
*/
export class AzureAiExecutor extends DefaultExecutor {
constructor() {
super("azure-ai");
}
override transformRequest(
model: string,
body: unknown,
stream: boolean,
credentials: ProviderCredentials
): unknown {
return applyAzureParamRules(
model,
body,
super.transformRequest(model, body, stream, credentials)
);
}
}

View File

@@ -1,9 +1,9 @@
import { DefaultExecutor } from "./default.ts";
import type { ProviderCredentials } from "./base.ts";
import { stripTrailingSlashes } from "../utils/urlSanitize.ts";
import { applyAzureParamRules } from "./azureParamRules.ts";
const DEFAULT_API_VERSION = "2024-12-01-preview";
const GPT5_OR_REASONING_DEPLOYMENT = /(?:^|[/_-])(?:gpt-5|o(?:1|3|4))(?:[._-]|$)/i;
function normalizeAzureBaseUrl(rawBaseUrl?: string | null): string {
const normalized = stripTrailingSlashes((rawBaseUrl || "").trim());
@@ -57,37 +57,10 @@ export class AzureOpenAIExecutor extends DefaultExecutor {
stream: boolean,
credentials: ProviderCredentials
): unknown {
const transformed = super.transformRequest(model, body, stream, credentials);
if (!GPT5_OR_REASONING_DEPLOYMENT.test(model)) return transformed;
if (!transformed || typeof transformed !== "object" || Array.isArray(transformed)) {
return transformed;
}
const original =
body && typeof body === "object" && !Array.isArray(body)
? (body as Record<string, unknown>)
: null;
const normalized = { ...(transformed as Record<string, unknown>) };
if (original?.max_completion_tokens !== undefined) {
normalized.max_completion_tokens = original.max_completion_tokens;
} else if (
normalized.max_completion_tokens === undefined &&
original?.max_tokens !== undefined
) {
normalized.max_completion_tokens = original.max_tokens;
}
delete normalized.max_tokens;
if (normalized.temperature !== undefined && normalized.temperature !== 1) {
delete normalized.temperature;
}
const hasTools = Array.isArray(normalized.tools) && normalized.tools.length > 0;
if (hasTools || normalized.reasoning_effort === "none") {
delete normalized.reasoning_effort;
}
return normalized;
return applyAzureParamRules(
model,
body,
super.transformRequest(model, body, stream, credentials)
);
}
}

View File

@@ -0,0 +1,76 @@
/**
* Azure Chat Completions param rules, shared by every Azure wire path.
*
* Azure's newer deployments reject a handful of stock OpenAI Chat Completions
* params and return HTTP 400 rather than ignoring them:
*
* - `max_tokens` -> "Unsupported parameter: 'max_tokens' is not supported
* with this model. Use 'max_completion_tokens' instead."
* - `temperature` -> only the default (1) is accepted.
* - `reasoning_effort` -> "Function tools with reasoning_effort are not
* supported ... Please use /v1/responses instead."
*
* This logic previously lived inline in `AzureOpenAIExecutor`, so it only
* covered the `azure-openai` provider. `azure-ai` (Azure AI Foundry) routes
* through `DefaultExecutor` and inherited none of it, which meant an identical
* deployment 400'd on one connection and succeeded on the other. Extracted here
* so both executors apply exactly the same rules.
*/
/**
* Deployments that require `max_completion_tokens` instead of `max_tokens`.
*
* Matches the GPT-5 family and the o1/o3/o4 reasoning series at a token
* boundary, so a deployment named `my-gpt-5-prod` matches while an unrelated
* `piston-o4-legacy`-style name does not match by accident. `gpt-chat-latest`
* is listed explicitly: it is a moving alias that currently resolves to a
* GPT-5-era model and rejects `max_tokens`, but carries no version number for
* the boundary pattern to key on.
*/
export const AZURE_COMPLETION_TOKEN_DEPLOYMENT =
/(?:^|[/_-])(?:gpt-5|o(?:1|3|4))(?:[._-]|$)|^gpt-chat-latest$/i;
/**
* Apply the Azure param rules to an already-translated Chat Completions body.
*
* `originalBody` is the pre-translation request, consulted only to recover a
* caller-supplied token budget that translation may have moved or dropped.
* Returns `transformed` untouched when the deployment is unaffected or the body
* is not a plain object, and never mutates either input.
*/
export function applyAzureParamRules(
model: string,
originalBody: unknown,
transformed: unknown
): unknown {
if (!AZURE_COMPLETION_TOKEN_DEPLOYMENT.test(model)) return transformed;
if (!transformed || typeof transformed !== "object" || Array.isArray(transformed)) {
return transformed;
}
const original =
originalBody && typeof originalBody === "object" && !Array.isArray(originalBody)
? (originalBody as Record<string, unknown>)
: null;
const normalized = { ...(transformed as Record<string, unknown>) };
if (original?.max_completion_tokens !== undefined) {
normalized.max_completion_tokens = original.max_completion_tokens;
} else if (normalized.max_completion_tokens === undefined && original?.max_tokens !== undefined) {
normalized.max_completion_tokens = original.max_tokens;
}
delete normalized.max_tokens;
if (normalized.temperature !== undefined && normalized.temperature !== 1) {
delete normalized.temperature;
}
// Azure 400s on reasoning_effort as soon as tools are present, which is every
// agentic client (Claude Code, Cursor agent) on every turn.
const hasTools = Array.isArray(normalized.tools) && normalized.tools.length > 0;
if (hasTools || normalized.reasoning_effort === "none") {
delete normalized.reasoning_effort;
}
return normalized;
}

View File

@@ -2,16 +2,19 @@
// Extracted verbatim from base.ts. Deps are config/services only (no host import → no cycle).
import { PROVIDER_CLAUDE } from "../../services/systemTransforms.ts";
import { isClaudeCodeCompatible } from "../../services/provider.ts";
import { supportsClaudeMaxEffort, supportsXHighEffort } from "../../config/providerModels.ts";
import {
supportsClaudeMaxEffort,
supportsXHighEffort,
getProviderModel,
} from "../../config/providerModels.ts";
/**
* Sanitize reasoning_effort for providers that don't accept all values.
*
* The claude→openai translator passes output_config.effort through verbatim
* (including max) and only performs form conversion; provider-aware effort
* policy is owned here. Combined with runtime alias remapping (e.g.
* claude-opus-4-6 → mimo/mimo-v2.5-pro), this routes a client's effort value
* to OpenAI-shape providers that don't accept it:
* The claude→openai translator may emit reasoning_effort=max/xhigh when the
* client sends output_config.effort=max on a Claude-shape request. Combined with
* runtime alias remapping (e.g. claude-opus-4-6 → mimo/mimo-v2.5-pro), this
* routes xhigh to OpenAI-shape providers that don't accept the value:
*
* xiaomi-mimo : low|medium|high only — 400 literal_error on xhigh
* mistral : devstral models reject reasoning_effort entirely
@@ -140,9 +143,11 @@ export function mapNvidiaGlm52ReasoningParams(
}
export function supportsMaxEffortForProvider(provider: string, model: string): boolean {
const resolvedModelId = getProviderModel(provider, model)?.id || model;
const isClaude =
(provider === PROVIDER_CLAUDE || isClaudeCodeCompatible(provider)) &&
supportsClaudeMaxEffort(model);
supportsClaudeMaxEffort(resolvedModelId);
// opencode-go proxies DeepSeek with the native DeepSeek API contract, which
// accepts {high, max} literally. Without this opt-in, max would be
// normalized to xhigh (the OmniRoute-internal top tier) and rejected by the
@@ -151,11 +156,12 @@ export function supportsMaxEffortForProvider(provider: string, model: string): b
// Ollama Cloud also accepts literal max (for example GLM 5.2 supports
// low|medium|high|max|none) and rejects xhigh.
const isOpencodeGoDeepSeek =
(provider === "opencode-go" || provider === "opencode-zen") &&
model.toLowerCase().includes("deepseek");
provider === "opencode-go" && resolvedModelId.toLowerCase().includes("deepseek");
const isOllamaCloud = provider === "ollama-cloud";
const isMoonshotK3 =
(provider === "moonshot" || provider === "kimi") && /^kimi-k3(?:$|-)/i.test(model);
// Kimi K3 only accepts literal max and rejects xhigh natively. Apply this mapping
// regardless of provider so that OpenAI-compatible proxies (e.g. TokenRouter)
// correctly pass max instead of the internal xhigh top tier.
const isMoonshotK3 = /^kimi-k3(?:$|-)/i.test(resolvedModelId);
return isClaude || isOpencodeGoDeepSeek || isOllamaCloud || isMoonshotK3;
}
@@ -253,16 +259,6 @@ export function sanitizeReasoningEffortForProvider(
const effortStr = typeof c.effort === "string" ? c.effort.toLowerCase() : "";
const modelStr = model || "";
// Oh My Pi exposes `minimal`, while Codex's Responses API starts at `low`.
// Normalize every carrier before the Codex executor sends the upstream request.
if (provider === "codex" && effortStr === "minimal") {
log?.info?.(
"REASONING_SANITIZE",
`${provider}/${modelStr}: normalized reasoning_effort minimal → low`
);
return writeEffortValue(b, "low", c);
}
const githubOptIn =
provider === "github" && GITHUB_REASONING_EFFORT_OPT_IN_PATTERN.test(modelStr);
const rejecting =
@@ -298,27 +294,48 @@ export function sanitizeReasoningEffortForProvider(
}
const supportsXHigh = supportsXHighEffort(provider, modelStr);
const shouldDowngradeXHigh = effortStr === "xhigh" && !supportsXHigh;
const supportsXHighForMax = supportsXHigh;
const supportsMax = supportsMaxEffortForProvider(provider, modelStr);
const shouldNormalizeMaxToXHigh = effortStr === "max" && !supportsMax && supportsXHighForMax;
const shouldDowngradeMax = effortStr === "max" && !supportsMax && !supportsXHighForMax;
if (shouldNormalizeMaxToXHigh) {
// ── xhigh handling ──────────────────────────────────────────────────────
// xhigh is OmniRoute-internal. Map it to the best effort the model accepts.
if (effortStr === "xhigh") {
if (supportsXHigh) return body; // model accepts xhigh natively
if (supportsMax) {
log?.info?.(
"REASONING_SANITIZE",
`${provider}/${modelStr}: mapped reasoning_effort xhigh → max`
);
return writeEffortValue(b, "max", c);
}
// Model explicitly rejects xhigh — gracefully degrade to high (its highest standard tier)
log?.info?.(
"REASONING_SANITIZE",
`${provider}/${modelStr}: normalized reasoning_effort maxxhigh`
);
return writeEffortValue(b, "xhigh", c);
}
if (shouldDowngradeXHigh || shouldDowngradeMax) {
log?.info?.(
"REASONING_SANITIZE",
`${provider}/${modelStr}: downgraded reasoning_effort ${effortStr} → high`
`${provider}/${modelStr}: downgraded reasoning_effort xhigh → high`
);
return writeEffortValue(b, "high", c);
}
// ── max handling ────────────────────────────────────────────────────────
// NEW DEFAULT: pass max through unchanged. Most reasoning-capable APIs
// accept max natively. Only degrade when we KNOW the model rejects it
// (registry has supportsXHighEffort explicitly set to false AND it's not
// in the supportsMax whitelist). Unknown models pass through — trust the
// upstream, and if it 400s the user gets a clear signal. This prevents
// new models from being unusable for weeks until they're whitelisted (#8057).
if (effortStr === "max") {
if (supportsMax) return body; // explicitly known to accept max
if (!supportsXHigh) {
// Model is explicitly flagged as rejecting xhigh (and not in supportsMax) —
// it likely only accepts standard tiers. Degrade to its highest: high.
log?.info?.(
"REASONING_SANITIZE",
`${provider}/${modelStr}: downgraded reasoning_effort max → high (model rejects max/xhigh)`
);
return writeEffortValue(b, "high", c);
}
// Default: pass max through unchanged — trust the upstream
return body;
}
return body;
}

View File

@@ -408,12 +408,13 @@ export class CliproxyapiExecutor extends BaseExecutor {
input.log?.info?.("CPA", `CLIProxyAPI → ${url} (model: ${input.model}, shape: ${shape})`);
// _toolNameMap is an in-memory channel to chatCore for response-side
// tool name restoration; never send it over the wire.
// _toolNameMap and _namespaceToolIdentityMap are in-memory channels to
// chatCore for response-side tool name restoration; never send them over
// the wire.
const wireBody =
transformedBody && typeof transformedBody === "object"
? JSON.stringify(transformedBody, (key, value) =>
key === "_toolNameMap" ? undefined : value
key === "_toolNameMap" || key === "_namespaceToolIdentityMap" ? undefined : value
)
: JSON.stringify(transformedBody);

View File

@@ -1,5 +1,82 @@
import { DefaultExecutor } from "./default.ts";
import type { ProviderCredentials } from "./base.ts";
import type { ExecuteInput, ExecutorExecuteResult, ProviderCredentials } from "./base.ts";
const SENSITIVE_CONTENT_REJECTION =
"抱歉,系统检测到您当前输入的信息存在敏感内容,我无法响应您的请求,请检查后重新输入";
const LARGE_TOOL_METADATA_BYTES = 64 * 1024;
function responseFromResult(result: ExecutorExecuteResult): Response {
return result instanceof Response ? result : result.response;
}
function credentialsFromResult(
result: ExecutorExecuteResult,
fallback: ProviderCredentials
): ProviderCredentials {
if (result instanceof Response || !result.headers) return fallback;
const authorization = Object.entries(result.headers).find(
([name]) => name.toLowerCase() === "authorization"
)?.[1];
if (!authorization?.startsWith("Bearer ")) return fallback;
return {
...fallback,
accessToken: authorization.slice("Bearer ".length),
expiresAt: undefined,
};
}
function compactToolDescriptions(body: unknown): unknown | null {
if (!body || typeof body !== "object" || Array.isArray(body)) return null;
const request = body as Record<string, unknown>;
if (!Array.isArray(request.tools) || request.tools.length === 0) return null;
const originalTools = request.tools;
try {
const serializedTools = JSON.stringify(originalTools);
if (new TextEncoder().encode(serializedTools).byteLength < LARGE_TOOL_METADATA_BYTES) {
return null;
}
} catch {
return null;
}
let tools: unknown[] | null = null;
originalTools.forEach((tool, index) => {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return;
const declaration = tool as Record<string, unknown>;
if (
declaration.type !== "function" ||
!declaration.function ||
typeof declaration.function !== "object" ||
Array.isArray(declaration.function)
) {
return;
}
const toolFunction = declaration.function as Record<string, unknown>;
if (!Object.prototype.hasOwnProperty.call(toolFunction, "description")) return;
const compactFunction = { ...toolFunction };
delete compactFunction.description;
tools ??= originalTools.slice();
tools[index] = { ...declaration, function: compactFunction };
});
return tools ? { ...request, tools } : null;
}
async function isSensitiveContentRejection(response: Response): Promise<boolean> {
if (response.status !== 400) return false;
const responseText = await response
.clone()
.text()
.catch(() => "");
return responseText.includes(SENSITIVE_CONTENT_REJECTION);
}
/**
* CodeBuddyCnExecutor — talks to https://copilot.tencent.com/v2/chat/completions
@@ -21,6 +98,26 @@ export class CodeBuddyCnExecutor extends DefaultExecutor {
super("codebuddy-cn");
}
async execute(input: ExecuteInput): Promise<ExecutorExecuteResult> {
const result = await super.execute(input);
if (!(await isSensitiveContentRejection(responseFromResult(result)))) {
return result;
}
const compactBody = compactToolDescriptions(input.body);
if (!compactBody) return result;
input.log?.debug?.(
"CODEBUDDY_CN",
"Upstream rejected an oversized tool request as sensitive content; retrying with compact tool descriptions"
);
return super.execute({
...input,
body: compactBody,
credentials: credentialsFromResult(result, input.credentials),
});
}
transformRequest(
model: string,
body: unknown,

View File

@@ -32,6 +32,7 @@ import {
} from "../config/codexIdentity.ts";
import { getAccessToken } from "../services/tokenRefresh.ts";
import { sanitizeResponsesInputItems } from "../services/responsesInputSanitizer.ts";
import { applyResponsesInputPolicy } from "../services/responsesInputPolicy.ts";
import { normalizeCodexVerbosity } from "../services/codexVerbosity.ts";
import { getThinkingBudgetConfig, ThinkingMode } from "../services/thinkingBudget.ts";
import { CORS_HEADERS } from "../utils/cors.ts";
@@ -222,90 +223,6 @@ function convertSystemToDeveloperRole(body: Record<string, unknown>): void {
}
}
/**
* Strip server-generated item IDs from the input array.
*
* The Codex /codex/responses endpoint does not persist response items even when
* store=true is sent. When proxy clients (e.g. OpenClaw) include response items
* from previous turns in the input array, those items carry server-assigned IDs
* (prefixed with "rs_", "fc_", "resp_", "msg_"). The Codex backend tries to
* validate these IDs against its persistence store and returns 404 when the items
* are not found (because store was effectively false).
*
* This function:
* 1. Removes bare string references ("rs_abc123") from the input array
* 2. Removes object items with type "item_reference" (explicit stored-item refs)
* 3. Strips the "id" field from any object in input whose id matches a
* server-generated prefix (rs_, fc_, resp_, msg_) — so the content is
* preserved but the backend won't try to look it up
*/
export function stripStoredItemReferences(body: Record<string, unknown>): void {
if (Array.isArray(body.input) && body.input.length === 0) {
body.input = [
{
type: "message",
role: "user",
content: [{ type: "input_text", text: "continue" }],
},
];
}
if (!Array.isArray(body.input)) return;
const SERVER_ID_PATTERN = /^(rs|fc|resp|msg)_/;
let strippedCount = 0;
body.input = body.input.filter((item) => {
// Bare string references: "rs_abc123", "resp_abc123"
if (typeof item === "string" && SERVER_ID_PATTERN.test(item)) {
strippedCount++;
return false;
}
// Object references: { type: "item_reference", id: "rs_..." }
if (
item &&
typeof item === "object" &&
!Array.isArray(item) &&
(item as Record<string, unknown>).type === "item_reference"
) {
strippedCount++;
return false;
}
// Reasoning blobs (encrypted_content) are unusable with store=false since
// previous_response_id is deleted — strip them to avoid wasting context
// tokens (O(n^2) growth across agentic turns).
if (
item &&
typeof item === "object" &&
!Array.isArray(item) &&
(item as Record<string, unknown>).type === "reasoning"
) {
strippedCount++;
return false;
}
// Object items with server-generated IDs: strip the id field but keep the item.
// e.g. { id: "rs_...", type: "reasoning", summary: [...] } → keep content, remove id
// e.g. { id: "fc_...", type: "function_call", ... } → keep content, remove id
if (item && typeof item === "object" && !Array.isArray(item)) {
const record = item as Record<string, unknown>;
if (typeof record.id === "string" && SERVER_ID_PATTERN.test(record.id)) {
delete record.id;
strippedCount++;
}
}
return true;
});
if (strippedCount > 0) {
console.debug(
`[Codex] stripStoredItemReferences: sanitized ${strippedCount} server-generated ID(s) from input`
);
}
}
function stripOrphanedCodexFunctionCallOutputs(body: Record<string, unknown>): void {
if (!Array.isArray(body.input)) return;
@@ -1296,7 +1213,7 @@ export class CodexExecutor extends BaseExecutor {
}
// Issue #1832 & #1853: Map messages to input for clients like Cursor 5.5 that use responses/compact but send messages instead of input.
// This MUST run before convertSystemToDeveloperRole and stripStoredItemReferences.
// This MUST run before convertSystemToDeveloperRole.
if (!body.input && Array.isArray(body.messages)) {
body.input = body.messages.map((msg: ResponsesMessageInput) => ({
type: "message",
@@ -1419,11 +1336,6 @@ export class CodexExecutor extends BaseExecutor {
preserveCustomTools: nativeCodexPassthrough,
});
// Strip stored response item references (rs_, resp_, msg_ IDs) from input.
// The /codex/responses endpoint does not persist responses even with store=true,
// so any references to previous response items would cause 404 errors.
stripStoredItemReferences(body);
// Issue #806: Even for native passthrough, some clients (purist completions) might indiscriminately inject
// a `messages` or `prompt` array which the strict Codex Responses schema rejects.
delete body.messages;
@@ -1515,6 +1427,11 @@ export class CodexExecutor extends BaseExecutor {
delete body.session_id;
delete body.conversation_id;
applyResponsesInputPolicy(
body,
credentials?.providerSpecificData?.preserveEncryptedReasoning === true
);
if (nativeCodexPassthrough) {
return body;
}

View File

@@ -30,6 +30,114 @@ export function isCodexFreePlan(providerSpecificData: unknown): boolean {
return typeof plan === "string" && plan.trim().toLowerCase() === "free";
}
type JsonRecord = Record<string, unknown>;
const REDUNDANT_ONEOF_OBJECT_MAP_FIELDS = [
"properties",
"patternProperties",
"$defs",
"definitions",
] as const;
const REDUNDANT_ONEOF_ARRAY_SCHEMA_FIELDS = ["prefixItems", "oneOf", "anyOf", "allOf"] as const;
const REDUNDANT_ONEOF_SINGLE_SCHEMA_FIELDS = [
"items",
"additionalProperties",
"not",
"if",
"then",
"else",
] as const;
const REDUNDANT_ONEOF_ANNOTATION_KEYS = new Set(["const", "description", "title", "$comment"]);
/**
* Remove a redundant `oneOf` when it is fully covered by a sibling `enum`.
*
* The Codex private Responses endpoint (`chatgpt.com/backend-api/codex/responses`)
* intermittently returns a 502 `upstream_empty_response` when a tool parameter
* carries the JSON-Schema pattern `oneOf: [{const, ...annotations}]` together
* with a sibling `enum` whose value set exactly matches the `const` set. In that
* case `oneOf` adds no constraint beyond `enum`, so dropping it is semantically
* safe and eliminates the trigger.
*
* Only the exact-match redundant case is stripped. Bare `oneOf[const]` without
* a sibling `enum`, narrowing const sets, non-matching enums, type-discriminated
* `oneOf`, and `anyOf`/`allOf` are all preserved.
*/
export function stripRedundantOneOfConstEnum(schema: unknown): unknown {
if (Array.isArray(schema)) {
return schema.map((entry) => stripRedundantOneOfConstEnum(entry));
}
if (!isPlainObject(schema)) return schema;
const result: JsonRecord = { ...schema };
maybeStripRedundantOneOf(result);
for (const field of REDUNDANT_ONEOF_OBJECT_MAP_FIELDS) {
const map = result[field];
if (isPlainObject(map)) {
result[field] = Object.fromEntries(
Object.entries(map).map(([key, value]) => [key, stripRedundantOneOfConstEnum(value)])
);
}
}
for (const field of REDUNDANT_ONEOF_ARRAY_SCHEMA_FIELDS) {
if (Array.isArray(result[field])) {
result[field] = (result[field] as unknown[]).map((entry) =>
stripRedundantOneOfConstEnum(entry)
);
}
}
for (const field of REDUNDANT_ONEOF_SINGLE_SCHEMA_FIELDS) {
if (result[field] !== undefined) {
result[field] = stripRedundantOneOfConstEnum(result[field]);
}
}
return result;
}
function maybeStripRedundantOneOf(node: JsonRecord): void {
const branches = node.oneOf;
if (!Array.isArray(branches) || branches.length === 0) return;
const enumValues = Array.isArray(node.enum) ? node.enum : null;
if (!enumValues || enumValues.length === 0) return;
// Every branch must be {const, ...annotations only}.
const constValues: unknown[] = [];
for (const branch of branches) {
if (!isPlainObject(branch)) return;
const branchKeys = Object.keys(branch);
if (!branchKeys.includes("const")) return;
if (!branchKeys.every((key) => REDUNDANT_ONEOF_ANNOTATION_KEYS.has(key))) return;
constValues.push((branch as JsonRecord).const);
}
// Restrict to string consts and string enums (confirmed production shape).
if (!constValues.every((value) => typeof value === "string")) return;
if (!enumValues.every((value) => typeof value === "string")) return;
// All const values must be unique.
if (new Set(constValues).size !== constValues.length) return;
// The const set must exactly match the enum set.
const enumSet = new Set(enumValues);
if (enumSet.size !== constValues.length) return;
if (!constValues.every((value) => enumSet.has(value))) return;
delete node.oneOf;
}
function isPlainObject(value: unknown): value is JsonRecord {
return typeof value === "object" && value !== null && !Array.isArray(value);
}
export function normalizeCodexTools(
body: Record<string, unknown>,
options?: { dropImageGeneration?: boolean; preserveCustomTools?: boolean }
@@ -138,7 +246,9 @@ export function normalizeCodexTools(
// Codex/OpenAI Responses API rejects `pattern` fields using regex lookaround
// (e.g. `^(?=.*@).+$`) with a 400 "regex lookaround is not supported" error.
// Strip those before the schema reaches upstream (9router#1556).
const sanitizedParameters = stripUnsupportedRegexPatterns(parameters);
const sanitizedParameters = stripRedundantOneOfConstEnum(
stripUnsupportedRegexPatterns(parameters)
);
// Rewrite in-place to Responses format
for (const key of Object.keys(tool)) {

View File

@@ -48,6 +48,21 @@ function recordOrEmpty(value: unknown): JsonRecord {
return {};
}
/**
* Build the `arguments` field for an assistant tool-call part that Command
* Code's /alpha/generate schema REQUIRES (rejects a missing field with
* `missing required field 'arguments'`). Valid source values round-trip:
* - object arguments -> JSON string of the object
* - string arguments -> the string as-is (already valid JSON)
* - missing / empty / invalid JSON -> "{}" (a valid empty-object string)
*/
function toolCallArgumentsString(value: unknown): string {
const parsed = recordOrEmpty(value);
if (isRecord(value)) return JSON.stringify(parsed);
if (typeof value === "string" && value.trim()) return value;
return JSON.stringify(parsed);
}
function normalizeContentText(content: unknown): string {
if (typeof content === "string") return content;
return asRecordArray(content)
@@ -244,11 +259,15 @@ function convertMessages(
const id = stringValue(call.id) || "";
if (!id || !pairedToolCallIds.has(id)) continue;
const fn = isRecord(call.function) ? call.function : {};
const parsedInput = recordOrEmpty(fn.arguments);
parts.push({
type: "tool-call",
toolCallId: id,
toolName: stringValue(fn.name) || "",
input: recordOrEmpty(fn.arguments),
input: parsedInput,
// /alpha/generate requires this field on assistant tool-call parts;
// a missing one is rejected with `missing required field 'arguments'`.
arguments: toolCallArgumentsString(fn.arguments),
});
}
@@ -420,7 +439,61 @@ type AggregateState = {
usage: JsonRecord | null;
};
function firstRecord(record: JsonRecord, keys: readonly string[]): JsonRecord {
for (const key of keys) {
const value = record[key];
if (isRecord(value)) return value;
}
return {};
}
function firstNumber(record: JsonRecord, keys: readonly string[]): number | undefined {
for (const key of keys) {
const value = numberValue(record[key]);
if (value !== undefined) return value;
}
return undefined;
}
/** Keep earlier finish-step usage when the terminal finish event omits it. */
function mergeCommandCodeUsage(previous: JsonRecord | null, next: unknown): JsonRecord | null {
if (!isRecord(next)) return previous;
const merged: JsonRecord = { ...(previous || {}), ...next };
for (const key of [
"inputTokenDetails",
"input_token_details",
"input_tokens_details",
"prompt_tokens_details",
"outputTokenDetails",
"output_token_details",
"output_tokens_details",
"completion_tokens_details",
"reasoningTokenDetails",
"reasoning_token_details",
]) {
const before = isRecord(previous?.[key]) ? previous[key] : {};
const after = isRecord(next[key]) ? next[key] : {};
if (Object.keys(before).length > 0 || Object.keys(after).length > 0) {
merged[key] = { ...before, ...after };
}
}
return merged;
}
function rememberCommandCodeUsage(state: AggregateState, event: JsonRecord): void {
const usage =
event.type === "finish-step"
? (event.usage ?? event.totalUsage)
: (event.totalUsage ?? event.usage);
state.usage = mergeCommandCodeUsage(state.usage, usage);
}
function applyEventToAggregate(event: JsonRecord, state: AggregateState): void {
// Some Command Code protocol revisions attach usage to the terminal payload
// without preserving the event type. Capture it before event-specific handling.
rememberCommandCodeUsage(state, event);
switch (event.type) {
case "text-delta":
state.content += stringValue(event.text) || "";
@@ -440,9 +513,10 @@ function applyEventToAggregate(event: JsonRecord, state: AggregateState): void {
});
break;
}
case "finish-step":
break;
case "finish":
state.finishReason = mapFinishReason(event.finishReason);
state.usage = isRecord(event.totalUsage) ? event.totalUsage : null;
break;
}
}
@@ -460,30 +534,72 @@ function applyEventToAggregateOrThrow(event: JsonRecord, state: AggregateState):
function usageFromCommandCode(usage: JsonRecord | null) {
if (!usage) return undefined;
const details = isRecord(usage.inputTokenDetails) ? usage.inputTokenDetails : {};
const cacheRead = numberValue(details.cacheReadTokens) || 0;
const noCache = numberValue(details.noCacheTokens) || 0;
const inputDetails = firstRecord(usage, [
"inputTokenDetails",
"input_token_details",
"input_tokens_details",
"prompt_tokens_details",
]);
const outputDetails = firstRecord(usage, [
"outputTokenDetails",
"output_token_details",
"output_tokens_details",
"completion_tokens_details",
]);
const reasoningDetails = firstRecord(usage, [
"reasoningTokenDetails",
"reasoning_token_details",
"reasoning_tokens_details",
]);
const cacheRead =
firstNumber(usage, [
"cachedInputTokens",
"cached_input_tokens",
"cacheReadInputTokens",
"cache_read_input_tokens",
"cacheReadTokens",
"cache_read_tokens",
"cached_tokens",
]) ??
firstNumber(inputDetails, [
"cachedTokens",
"cached_tokens",
"cacheReadTokens",
"cache_read_tokens",
]);
const noCache = firstNumber(inputDetails, ["noCacheTokens", "no_cache_tokens"]);
// Command Code's totalUsage.inputTokens is the FULL prompt total and already
// includes the cached portion (noCacheTokens + cacheReadTokens = inputTokens),
// so we must NOT add cacheRead back — that would double-count. There is no
// cache-write field in the upstream payload, so cache creation stays unset.
const inputTokens = numberValue(usage.inputTokens) || 0;
const prompt = inputTokens;
const completion = numberValue(usage.outputTokens) || 0;
const prompt =
firstNumber(usage, ["inputTokens", "input_tokens", "promptTokens", "prompt_tokens"]) ??
(noCache ?? 0) + (cacheRead ?? 0);
const reasoning =
firstNumber(usage, ["reasoningTokens", "reasoning_tokens"]) ??
firstNumber(outputDetails, ["reasoningTokens", "reasoning_tokens"]) ??
firstNumber(reasoningDetails, ["reasoningTokens", "reasoning_tokens"]);
const textOutput = firstNumber(outputDetails, ["textTokens", "text_tokens"]);
const completion =
firstNumber(usage, [
"outputTokens",
"output_tokens",
"completionTokens",
"completion_tokens",
]) ?? (textOutput ?? 0) + (reasoning ?? 0);
const total = firstNumber(usage, ["totalTokens", "total_tokens"]) ?? prompt + completion;
const result: JsonRecord = {
prompt_tokens: prompt,
prompt_tokens_details: { cached_tokens: cacheRead ?? 0 },
completion_tokens: completion,
total_tokens: prompt + completion,
completion_tokens_details: { reasoning_tokens: reasoning ?? 0 },
total_tokens: total,
};
// Surface the cache breakdown as informational fields so logUsage prints
// `| cache_read=X | no_cache=Y` and appendRequestLog persists them. These are
// NOT added to prompt_tokens (already included) — metering stays accurate.
if (cacheRead > 0) result.cache_read_input_tokens = cacheRead;
if (noCache > 0) result.no_cache_tokens = noCache;
// Keep reasoning_token_details (reasoningTokens) when present so stream.ts's
// extractUsage can surface it as reasoning_tokens.
const reasoningDetails = isRecord(usage.reasoningTokenDetails) ? usage.reasoningTokenDetails : {};
const reasoning = numberValue(reasoningDetails.reasoningTokens);
if (cacheRead !== undefined && cacheRead > 0) result.cache_read_input_tokens = cacheRead;
if (noCache !== undefined && noCache > 0) result.no_cache_tokens = noCache;
if (reasoning !== undefined && reasoning > 0) result.reasoning_tokens = reasoning;
return result;
}
@@ -523,6 +639,7 @@ function createStreamResponse(
const emitEvent = (event: unknown) => {
if (!isRecord(event) || closed) return;
rememberCommandCodeUsage(state, event);
if (!sentRole) {
sentRole = true;
controller.enqueue(sse(chatCompletionChunk(id, model, { role: "assistant" })));
@@ -562,9 +679,10 @@ function createStreamResponse(
}
case "reasoning-end":
break;
case "finish-step":
break;
case "finish": {
state.finishReason = mapFinishReason(event.finishReason);
state.usage = isRecord(event.totalUsage) ? event.totalUsage : null;
controller.enqueue(sse(chatCompletionChunk(id, model, {}, state.finishReason)));
// Emit a standards-compliant usage-only chunk (choices: []) before
// [DONE] when upstream reported usage. stream.ts's extractUsage

View File

@@ -515,7 +515,6 @@ export function messagesToPrompt(
historyWindow = 0
): string {
if (messages.length === 0) return "";
const systemParts: string[] = [];
const conversation: Array<{ role: string; text: string }> = [];
const callNameById = new Map<string, string>();
@@ -527,8 +526,9 @@ export function messagesToPrompt(
} else if (m.role === "user" || m.role === "assistant") {
if (text) conversation.push({ role: m.role, text });
if (m.role === "user") lastUserContent = text;
const calls = Array.isArray((m as { tool_calls?: unknown }).tool_calls)
? (m as { tool_calls: Array<{ id?: string; function?: { name?: string } }> }).tool_calls
const toolCalls = (m as { tool_calls?: unknown }).tool_calls;
const calls = Array.isArray(toolCalls)
? (toolCalls as Array<{ id?: string; function?: { name?: string } }>)
: [];
for (const c of calls) {
if (c?.id && typeof c.function?.name === "string") callNameById.set(c.id, c.function.name);

View File

@@ -61,12 +61,11 @@ import {
} from "@/lib/providers/validation/urlHelpers";
import { forwardOpencodeClientHeaders } from "../utils/opencodeHeaders.ts";
import { resolveZaiUrl } from "./default/zaiFormatOverride.ts";
import { normalizePoolConfig } from "./default/poolConfig.ts";
import { acquireNvidiaConcurrencySlot } from "./default/nvidiaConcurrencyGate.ts";
import { resolveAlibabaProviderBaseUrl } from "@/shared/constants/alibabaProviderRegions";
import { usesCcWireImage } from "../services/ccWireImageBuiltins.ts";
import type { PoolConfig } from "../services/sessionPool/types.ts";
const NVIDIA_TOOL_CALL_ID_PATTERN = /^[A-Za-z0-9]{9}$/;
function normalizeNvidiaToolCallId(id: unknown): unknown {
@@ -146,7 +145,7 @@ export class DefaultExecutor extends BaseExecutor {
super(provider, PROVIDERS[provider] || PROVIDERS.openai);
const registryEntry = getRegistryEntry(provider);
if (registryEntry?.poolConfig) {
this.poolConfig = registryEntry.poolConfig as PoolConfig;
this.poolConfig = normalizePoolConfig(registryEntry.poolConfig) ?? undefined;
}
}

View File

@@ -0,0 +1,33 @@
import type { PoolConfig } from "../../services/sessionPool/types.ts";
export function normalizePoolConfig(value: Record<string, unknown>): PoolConfig | null {
const {
minSessions,
maxSessions,
cooldownBase,
cooldownMax,
cooldownJitter,
requestTimeout,
requestJitter,
} = value;
if (
typeof minSessions !== "number" ||
typeof maxSessions !== "number" ||
typeof cooldownBase !== "number" ||
typeof cooldownMax !== "number" ||
typeof cooldownJitter !== "number" ||
typeof requestTimeout !== "number" ||
typeof requestJitter !== "number"
) {
return null;
}
return {
minSessions,
maxSessions,
cooldownBase,
cooldownMax,
cooldownJitter,
requestTimeout,
requestJitter,
};
}

View File

@@ -137,8 +137,23 @@ interface DuckDuckGoModelCapabilities {
reasoningEffort: string | null;
}
type DuckDuckGoRequestMessage = Record<string, unknown> & {
role: string;
content: unknown;
};
let durablePublicKey: JsonWebKey | null = null;
export function normalizeDuckDuckGoMessages(value: unknown): DuckDuckGoRequestMessage[] {
if (!Array.isArray(value)) return [];
return value.flatMap((message) => {
if (!message || typeof message !== "object" || Array.isArray(message)) return [];
const record = message as Record<string, unknown>;
if (typeof record.role !== "string") return [];
return [{ ...record, role: record.role, content: record.content }];
});
}
function extractDuckDuckGoContent(data: unknown): string {
if (!data || typeof data !== "object") return "";
const record = data as Record<string, unknown>;
@@ -251,11 +266,14 @@ export function normalizeDuckDuckGoModel(model: string | undefined): string {
}
function getDuckDuckGoModelCapabilities(model: string): DuckDuckGoModelCapabilities {
// Per duckchat/v1/models (2026-07-22): claude-haiku-4-5 and gpt-oss-120b take a "low"
// reasoningEffort on the free tier; the others omit it (duck.ai applies its own default).
// `reasoningEffort` is REQUIRED on every duckchat/v1/chat request. Omitting it
// returns 400 ERR_BAD_REQUEST — A/B verified live against duck.ai with an
// otherwise byte-identical payload (200 with the field, 400 without, repeated).
// The live duck.ai bundle always sends one, so there is no "let the server
// pick a default" path any more.
if (model === "claude-haiku-4-5") return { reasoningEffort: "low" };
if (model === "tinfoil/gpt-oss-120b") return { reasoningEffort: "low" };
return { reasoningEffort: null };
return { reasoningEffort: "none" };
}
function extractDuckDuckGoFeVersion(html: string): string | null {
@@ -353,7 +371,6 @@ export class DuckDuckGoWebExecutor extends BaseExecutor {
}
private warmed = false;
private seeded = false;
private feVersion = DEFAULT_FE_VERSION;
private pendingVqdHash1: string | null = null;
private readonly cookieJar = new Map<string, string>();
@@ -440,14 +457,12 @@ export class DuckDuckGoWebExecutor extends BaseExecutor {
const { model, body, stream, signal, upstreamExtraHeaders } = input;
const upstreamModel = normalizeDuckDuckGoModel(model);
const bodyObj = (body || {}) as Record<string, unknown>;
const rawMessages = Array.isArray((body as { messages?: unknown[] } | null)?.messages)
? ((body as { messages: unknown[] }).messages as Array<Record<string, unknown>>)
: [];
const rawMessages = normalizeDuckDuckGoMessages(bodyObj.messages);
const { hasTools, requestedTools, effectiveMessages } = prepareToolMessages(
bodyObj,
rawMessages
);
const messages = effectiveMessages as Array<Record<string, unknown>>;
const messages = effectiveMessages;
const isStreaming = stream !== false;
const upstreamHeaders = upstreamExtraHeaders || {};
@@ -561,7 +576,12 @@ export class DuckDuckGoWebExecutor extends BaseExecutor {
}
await this.warmSession(mergedSignal);
await this.seedChallengeChain(upstreamModel, mergedSignal);
// NOTE: the throwaway "seed" chat POST that used to run here has been removed.
// It existed to coax a usable challenge out of the upstream while the solver
// was broken; now that the solver reproduces a real browser's probe vectors
// exactly, the first real request succeeds on its own. Keeping it only doubled
// the chat calls per user request against an IP-rate-limited endpoint, which
// showed up as spurious 429 ERR_RATE_LIMIT.
const vqdHeaders = await this.acquireAuthHeaders(mergedSignal);
if (!vqdHeaders.vqd4 && !vqdHeaders.vqdHash1) {
clearTimeout(timeout);
@@ -770,41 +790,6 @@ export class DuckDuckGoWebExecutor extends BaseExecutor {
);
}
private async seedChallengeChain(model: string, signal: AbortSignal): Promise<void> {
if (this.seeded || signal.aborted) return;
this.seeded = true;
const seedMessages = [{ role: "user", content: "hi" }];
const previousPending = this.pendingVqdHash1;
try {
const vqdHeaders = await this.acquireAuthHeaders(signal);
if (!vqdHeaders.vqd4 && !vqdHeaders.vqdHash1) {
this.pendingVqdHash1 = previousPending;
return;
}
const response = await fetch(CHAT_URL, {
method: "POST",
headers: mergeHeadersCaseInsensitive(this.buildRequestHeaders(), {
Accept: "text/event-stream",
"Content-Type": "application/json",
"x-ddg-journey-id": randomUUID().replaceAll("-", ""),
"x-fe-signals": makeDuckDuckGoFeSignals(),
"x-fe-version": this.feVersion,
...(vqdHeaders.vqd4 ? { "x-vqd-4": vqdHeaders.vqd4 } : {}),
...(vqdHeaders.vqdHash1 ? { "x-vqd-hash-1": vqdHeaders.vqdHash1 } : {}),
}),
body: JSON.stringify(buildDuckDuckGoPayload(model, seedMessages, false)),
signal,
});
this.rememberResponseCookies(response);
if (response.ok) this.rememberChallengeHeader(response);
else this.pendingVqdHash1 = previousPending;
await response.body?.cancel().catch(() => {});
} catch (error) {
void error;
this.pendingVqdHash1 = previousPending;
}
}
private async processResponse(
response: Response,
streaming: boolean,

View File

@@ -5,12 +5,38 @@ import { createHash } from "node:crypto";
import vm from "node:vm";
import { parseFragment, serialize } from "parse5";
// WARNING: the contents of this template literal are NOT TypeScript — they are plain
// script-mode JavaScript executed via `vm.runInContext`. `vm.runInContext` compiles in
// script (non-module) mode, so an `export` keyword anywhere in here is a hard
// SyntaxError that kills the whole solver. A refactor that mass-added `export` to the
// five `function` declarations below silently broke every DuckDuckGo chat request
// (solve threw -> unsolved challenge sent -> HTTP 418 ERR_CHALLENGE). Do not add
// `export`/`import` to this string; `duckduckgo-challenge-split.test.ts` guards this.
export const CHALLENGE_STUBS = String.raw`
var __ua = __DDG_REAL_UA__;
var __HTML_LOOKUP = __DDG_HTML_LOOKUP__;
export function __makeHtmlElement(tag) {
// Browser-fidelity shims for the DDG "am I a real browser" probes.
// In a browser every built-in stringifies as native code; under a plain vm
// context the user-land re-declarations below would otherwise leak their source.
function __nativeFn(fn, name){
Object.defineProperty(fn, 'name', { value: name, configurable: true });
fn.toString = function(){ return 'function ' + name + '() { [native code] }'; };
return fn;
}
__nativeFn(parseInt, 'parseInt');
__nativeFn(parseFloat, 'parseFloat');
__nativeFn(isNaN, 'isNaN');
__nativeFn(encodeURIComponent, 'encodeURIComponent');
__nativeFn(decodeURIComponent, 'decodeURIComponent');
// NOTE: do NOT seal Math. Real Chromium reports Object.isSealed(Math) === false,
// and at least one challenge variant probes exactly that; sealing it here made
// the vector differ from the browser by one and failed the challenge.
function __makeHtmlElement(tag) {
var state = { _innerHTML: '', _qsaCount: 0, _cssText: '' };
var el = {
// Instantiate against the real per-tag constructor so
// document.createElement('div') instanceof HTMLDivElement holds.
var el = Object.create(__ctorForTag(tag).prototype);
Object.assign(el, {
tagName: String(tag).toUpperCase(), nodeName: String(tag).toUpperCase(), nodeType: 1,
children: [], childNodes: [], classList: [], dataset: {},
offsetWidth: 1, offsetHeight: 1, clientWidth: 1, clientHeight: 1, scrollHeight: 1, scrollWidth: 1,
@@ -19,9 +45,9 @@ export function __makeHtmlElement(tag) {
getAttribute: function(a){ if(a==='srcdoc') return state._srcdoc||''; return null; },
hasAttribute: function(){ return false; }, appendChild: function(c){ return c; }, removeChild: function(c){ return c; },
addEventListener: function(){}, removeEventListener: function(){}, querySelector: function(){ return null; },
querySelectorAll: function(s){ if (s === '*') { var arr = []; arr.length = state._qsaCount; return arr; } return []; },
querySelectorAll: function(s){ if (s === '*') { return __makeNodeList(state._qsaCount); } return __makeNodeList(0); },
cloneNode: function(){ return __makeHtmlElement(tag); }
};
});
Object.defineProperty(el, 'style', { value: new Proxy({}, { set: function(t, k, v){ t[k] = v; if (k === 'cssText') state._cssText = String(v); return true; }, get: function(t, k){ if (k === 'cssText') return state._cssText; return t[k] || ''; } }), enumerable: true, configurable: true });
Object.defineProperty(el, 'innerHTML', { get: function(){ return state._innerHTML; }, set: function(v){ var key = String(v); var entry = __HTML_LOOKUP && __HTML_LOOKUP[key]; if (entry) { state._innerHTML = String(entry.html); state._qsaCount = entry.count|0; } else { state._innerHTML = key; state._qsaCount = 0; } }, enumerable: true, configurable: true });
Object.defineProperty(el, 'outerHTML', { get: function(){ return '<' + tag + '>' + state._innerHTML + '</' + tag + '>'; }, enumerable: true });
@@ -30,7 +56,7 @@ export function __makeHtmlElement(tag) {
Object.defineProperty(el, 'contentDocument', { get: function(){ return __ifDoc; }, enumerable: true });
return el;
}
export function __mkObj(name, base) {
function __mkObj(name, base) {
base = base || {};
return new Proxy(base, {
get: function(t, k) {
@@ -54,18 +80,105 @@ export function __mkObj(name, base) {
has: function(t, k){ return k in t; }, set: function(t, k, v){ t[k] = v; return true; }
});
}
export function __parseCssDisplay(cssText){ if(!cssText) return ''; var m = String(cssText).match(/(?:^|;)\\s*display\\s*:\\s*([^;]+)/i); return m ? String(m[1]).trim() : ''; }
export function __getComputedStyle(el){ var cssText = el && el.style && el.style.cssText || ''; var display = __parseCssDisplay(cssText); return { getPropertyValue: function(name){ if(String(name).toLowerCase()==='display') return display; return ''; }, cssText: cssText, display: display }; }
function __parseCssDisplay(cssText){ if(!cssText) return ''; var m = String(cssText).match(/(?:^|;)\s*display\s*:\s*([^;]+)/i); return m ? String(m[1]).trim() : ''; }
function __getComputedStyle(el){ var cssText = el && el.style && el.style.cssText || ''; var display = __parseCssDisplay(cssText); return { getPropertyValue: function(name){ if(String(name).toLowerCase()==='display') return display; return ''; }, cssText: cssText, display: display }; }
var __ifMeta = __mkObj('meta', { getAttribute: function(a){ return a==='content' ? "default-src 'none'; script-src 'unsafe-inline';" : null; }, hasAttribute: function(a){ return a==='content'; }, tagName: 'META', nodeName: 'META' });
var __ifDoc = __mkObj('iframeDoc', { querySelector: function(s){ if (s && s.indexOf('Content-Security-Policy') !== -1) return __ifMeta; if (s === 'meta') return __ifMeta; return null; }, querySelectorAll: function(s){ if (s && s.indexOf('Content-Security-Policy') !== -1) return [__ifMeta]; if (s === 'meta') return [__ifMeta]; return []; }, getElementsByTagName: function(t){ return t && t.toLowerCase()==='meta' ? [__ifMeta] : []; }, body: __mkObj('iframeBody'), head: __mkObj('iframeHead'), documentElement: __mkObj('iframeRoot'), createElement: function(){ return __mkObj('elem', {setAttribute:function(){}, appendChild:function(){}, removeChild:function(){}, getAttribute:function(){return null;}, hasAttribute:function(){return false;}}); }, cookie: '', readyState: 'complete' });
var __iframeEl = __mkObj('iframe', { contentDocument: __ifDoc, contentWindow: __mkObj('iframeWin', { document: __ifDoc, top: undefined, parent: undefined }), document: __ifDoc, getAttribute: function(a){ if (a==='sandbox') return 'allow-scripts allow-same-origin'; if (a==='srcdoc') return ''; if (a==='id') return 'jsa'; return null; }, hasAttribute: function(a){ return a==='sandbox'||a==='id'; }, tagName: 'IFRAME', nodeName: 'IFRAME', id: 'jsa' });
var document = __mkObj('document', { querySelector: function(s){ if (s === '#jsa') return __iframeEl; if (s && s.indexOf('Content-Security-Policy') !== -1) return __ifMeta; return null; }, querySelectorAll: function(s){ if (s === '#jsa') return [__iframeEl]; if (s && s.indexOf('Content-Security-Policy') !== -1) return [__ifMeta]; return []; }, getElementById: function(id){ return id==='jsa' ? __iframeEl : null; }, getElementsByTagName: function(t){ if(t&&t.toLowerCase()==='iframe') return [__iframeEl]; return []; }, getElementsByClassName: function(){ return []; }, body: __mkObj('body', {appendChild:function(){}, removeChild:function(){}, querySelector:function(s){return s==='#jsa'?__iframeEl:null;}, querySelectorAll:function(s){return s==='#jsa'?[__iframeEl]:[];}}), head: __mkObj('head'), documentElement: __mkObj('root'), createElement: function(tag){ return __makeHtmlElement(tag||'div'); }, createTextNode: function(t){ return {nodeType:3, nodeValue:String(t||''), textContent:String(t||'')}; }, cookie: '', readyState: 'complete', title: '', addEventListener: function(){}, removeEventListener: function(){} });
// document.body keeps a LIVE children collection: challenges append a node and
// assert body.children.length grew by exactly 1, then remove it again.
var __bodyKids = [];
Object.defineProperty(__bodyKids, 'constructor', { value: HTMLCollection, enumerable: false, configurable: true });
var __body = __mkObj('body', {
appendChild: function(c){ __bodyKids.push(c); return c; },
removeChild: function(c){ var i = __bodyKids.indexOf(c); if (i !== -1) __bodyKids.splice(i, 1); return c; },
contains: function(c){ return __bodyKids.indexOf(c) !== -1; },
querySelector: function(s){ return s === '#jsa' ? __iframeEl : null; },
querySelectorAll: function(s){ return s === '#jsa' ? [__iframeEl] : __makeNodeList(0); },
children: __bodyKids, childNodes: __bodyKids,
tagName: 'BODY', nodeName: 'BODY', nodeType: 1
});
var document = __mkObj('document', { querySelector: function(s){ if (s === '#jsa') return __iframeEl; if (s && s.indexOf('Content-Security-Policy') !== -1) return __ifMeta; return null; }, querySelectorAll: function(s){ if (s === '#jsa') return [__iframeEl]; if (s && s.indexOf('Content-Security-Policy') !== -1) return [__ifMeta]; return __makeNodeList(__bodyKids.length + 3); }, getElementById: function(id){ return id==='jsa' ? __iframeEl : null; }, getElementsByTagName: function(t){ if(t&&t.toLowerCase()==='iframe') return [__iframeEl]; return []; }, getElementsByClassName: function(){ return []; }, body: __body, head: __mkObj('head'), documentElement: __mkObj('root'), createElement: function(tag){ return __makeHtmlElement(tag||'div'); }, createTextNode: function(t){ return {nodeType:3, nodeValue:String(t||''), textContent:String(t||'')}; }, cookie: '', readyState: 'complete', title: '', addEventListener: function(){}, removeEventListener: function(){} });
var window = __mkObj('window', { document: document, __DDG_BE_VERSION__: 1, __DDG_FE_CHAT_HASH__: 1, navigator: __mkObj('navigator', { userAgent: __ua, webdriver: false, language: 'en-US', languages: ['en-US','en'], platform: 'Linux x86_64', vendor: 'Google Inc.', appVersion: '5.0 (X11)', cookieEnabled: true, onLine: true, hardwareConcurrency: 8, deviceMemory: 8 }), innerWidth: 1280, innerHeight: 800, outerWidth: 1280, outerHeight: 800, devicePixelRatio: 1, screen: __mkObj('screen', { width:1920, height:1080, availWidth:1920, availHeight:1080, colorDepth:24, pixelDepth:24 }), location: __mkObj('location', { href:'https://duck.ai/', origin:'https://duck.ai', host:'duck.ai', hostname:'duck.ai', protocol:'https:', pathname:'/' }), performance: __mkObj('perf', { now: function(){ return 0; }, timeOrigin: 0 }), history: __mkObj('history', { length: 1, state: null }), addEventListener: function(){}, removeEventListener: function(){}, dispatchEvent: function(){return true;}, setTimeout: function(fn){ try{fn();}catch(e){} return 0; }, clearTimeout: function(){}, hasOwnProperty: function(k){ if (k==='__DDG_BE_VERSION__'||k==='__DDG_FE_CHAT_HASH__') return true; return Object.prototype.hasOwnProperty.call(this,k); } });
window.top = window; window.self = window; window.window = window; window.parent = window; window.globalThis = window;
// Object.prototype.toString.call(window) must be "[object Window]".
try { window[Symbol.toStringTag] = 'Window'; } catch (e) {}
// In a browser a sloppy-mode function called with no receiver gets the global
// object, and challenges assert (function(){return this;})() === window.
// In a vm context that is the context's own global, so alias it to window.
try {
var __g = (function(){ return this; })();
if (__g && __g !== window) {
Object.defineProperty(__g, Symbol.toStringTag, { value: 'Window', configurable: true });
// Copy by VALUE, not via accessors. Two reasons:
// 1) the var top/self/navigator/... declarations further down are hoisted,
// so those names already exist on the vm global and an "in" guard would
// skip them, leaving window.navigator undefined;
// 2) accessors closing over the window binding would recurse once it is
// rebound to __g below.
// The stub window is static, so a value copy is equivalent.
var __winStub = window;
for (var __k in __winStub) {
try { __g[__k] = __winStub[__k]; } catch (e) {}
}
// hasOwnProperty is probed for the __DDG_* markers; keep the stub's version.
try { __g.hasOwnProperty = function(k){ return __winStub.hasOwnProperty(k); }; } catch (e) {}
window = __g;
window.top = window; window.self = window; window.window = window; window.parent = window; window.globalThis = window;
}
} catch (e) {}
var top = window, self = window, parent = window, navigator = window.navigator, location = window.location, screen = window.screen, performance = window.performance, history = window.history;
var __R = null, __E = null;
export function __HTMLClass(name){ var c = function(){}; c.prototype = __mkObj(name+'.proto'); return c; }
var HTMLElement = __HTMLClass('HTMLElement'), HTMLDivElement = __HTMLClass('HTMLDivElement'), HTMLIFrameElement = __HTMLClass('HTMLIFrameElement'), HTMLDocument = __HTMLClass('HTMLDocument'), Document = __HTMLClass('Document'), Element = __HTMLClass('Element'), Node = __HTMLClass('Node'), Window = __HTMLClass('Window'), Event = __HTMLClass('Event'), MouseEvent = __HTMLClass('MouseEvent'), KeyboardEvent = __HTMLClass('KeyboardEvent'), TouchEvent = __HTMLClass('TouchEvent'), XMLHttpRequest = __HTMLClass('XMLHttpRequest'), WebSocket = __HTMLClass('WebSocket'), Image = __HTMLClass('Image'), FormData = __HTMLClass('FormData'), Blob = __HTMLClass('Blob'), File = __HTMLClass('File'), FileReader = __HTMLClass('FileReader'), URL = __HTMLClass('URL'), URLSearchParams = __HTMLClass('URLSearchParams'), Headers = __HTMLClass('Headers'), Request = __HTMLClass('Request'), Response = __HTMLClass('Response');
// Real DOM constructor chain. Some DDG challenge variants assert
// HTMLDivElement.prototype instanceof HTMLElement and
// HTMLElement.prototype instanceof Element, so these cannot be flat
// unrelated stubs — the prototype links have to be real.
function __DomClass(name, parent){
var c = function(){};
if (parent) c.prototype = Object.create(parent.prototype);
c.prototype.constructor = c;
Object.defineProperty(c, 'name', { value: name, configurable: true });
c.toString = function(){ return 'function ' + name + '() { [native code] }'; };
return c;
}
var EventTarget = __DomClass('EventTarget', null);
var Node = __DomClass('Node', EventTarget);
var Element = __DomClass('Element', Node);
var HTMLElement = __DomClass('HTMLElement', Element);
var HTMLDivElement = __DomClass('HTMLDivElement', HTMLElement);
var HTMLIFrameElement = __DomClass('HTMLIFrameElement', HTMLElement);
var HTMLLIElement = __DomClass('HTMLLIElement', HTMLElement);
var HTMLUnknownElement = __DomClass('HTMLUnknownElement', HTMLElement);
var Document = __DomClass('Document', Node);
var HTMLDocument = __DomClass('HTMLDocument', Document);
var NodeList = __DomClass('NodeList', null);
var HTMLCollection = __DomClass('HTMLCollection', null);
// Map a tag name to the constructor a browser would use, so
// document.createElement('div') instanceof HTMLDivElement holds.
function __ctorForTag(tag){
var t = String(tag||'div').toLowerCase();
if (t === 'div') return HTMLDivElement;
if (t === 'iframe') return HTMLIFrameElement;
if (t === 'li') return HTMLLIElement;
return HTMLElement;
}
// A NodeList-like: array-shaped but NOT a real Array, with .constructor.name
// === 'NodeList' — challenges check both !Array.isArray(x) and the ctor name.
function __makeNodeList(length){
var nl = Object.create(NodeList.prototype);
var n = length|0;
for (var i = 0; i < n; i++) nl[i] = __makeHtmlElement('div');
Object.defineProperty(nl, 'length', { value: n, enumerable: false, configurable: true });
nl.item = function(i){ return this[i] || null; };
nl.forEach = function(fn, thisArg){ for (var i = 0; i < n; i++) fn.call(thisArg, this[i], i, this); };
nl[Symbol.iterator] = function(){ var i = 0, self = this; return { next: function(){ return i < n ? { value: self[i++], done: false } : { value: undefined, done: true }; } }; };
return nl;
}
function __HTMLClass(name){ var c = function(){}; c.prototype = __mkObj(name+'.proto'); return c; }
// NOTE: HTMLElement / HTMLDivElement / HTMLIFrameElement / Element / Node /
// Document / HTMLDocument / NodeList are defined above via __DomClass with a
// REAL prototype chain — do not redeclare them here or the instanceof probes break.
var Window = __HTMLClass('Window'), Event = __HTMLClass('Event'), MouseEvent = __HTMLClass('MouseEvent'), KeyboardEvent = __HTMLClass('KeyboardEvent'), TouchEvent = __HTMLClass('TouchEvent'), XMLHttpRequest = __HTMLClass('XMLHttpRequest'), WebSocket = __HTMLClass('WebSocket'), Image = __HTMLClass('Image'), FormData = __HTMLClass('FormData'), Blob = __HTMLClass('Blob'), File = __HTMLClass('File'), FileReader = __HTMLClass('FileReader'), URL = __HTMLClass('URL'), URLSearchParams = __HTMLClass('URLSearchParams'), Headers = __HTMLClass('Headers'), Request = __HTMLClass('Request'), Response = __HTMLClass('Response');
var fetch = function(){ return Promise.resolve(__mkObj('resp', {ok:true, status:200, json:function(){return Promise.resolve({});}, text:function(){return Promise.resolve('');}})); };
var getComputedStyle = __getComputedStyle;
`;
@@ -90,9 +203,16 @@ export function buildHtmlLookup(js: string): Record<string, { html: string; coun
if (seen.has(html)) continue;
seen.add(html);
const fragment = parseFragment(html);
// `count` backs `element.querySelectorAll('*').length` for an element whose
// innerHTML is `html`. `querySelectorAll('*')` on a container returns its
// DESCENDANTS, and `countHtmlElements` already excludes the `#document-fragment`
// root, so the fragment's element count IS the descendant count. The former
// `- 1` undercounted by one (verified against a real browser: for
// `<li><div></li><li></div` Chromium reports 3, this returned 2), which
// corrupted every probe that multiplies by that length.
lookup[html] = {
html: serialize(fragment),
count: Math.max(0, countHtmlElements(fragment) - 1),
count: countHtmlElements(fragment),
};
}
return lookup;
@@ -102,14 +222,33 @@ export function sha256Base64(value: string): string {
return createHash("sha256").update(value, "utf8").digest("base64");
}
// Shape of the object a DDG challenge program resolves to.
type DuckDuckGoChallengeResult = {
client_hashes?: unknown;
meta?: unknown;
[key: string]: unknown;
};
/**
* Origin the solved challenge claims to come from. The duck.ai frontend stamps
* `meta.origin` with its own origin and the upstream cross-checks it.
*/
export const DUCKDUCKGO_CHALLENGE_ORIGIN = "https://duck.ai";
/**
* `meta.stack` mimics the frontend's captured Error stack. The upstream only
* requires a plausible stack that points at the duck.ai bundle — verified by
* ablation: a generic bundle path is accepted, omitting the field is not.
*/
function buildChallengeStack(origin: string, bundlePath: string): string {
const url = `${origin}${bundlePath}`;
return `Error\nat l (${url}:2:1695625)\nat async ${url}:2:1519117`;
}
export async function solveDuckDuckGoChallenge(
challenge: string,
userAgent: string
userAgent: string,
options: { origin?: string; bundlePath?: string } = {}
): Promise<string> {
// SECURITY NOTE: This function executes base64-decoded JavaScript from duck.ai via vm.runInContext.
// The challenge code is upstream-supplied (supply-chain surface). It is sandboxed with a 5s timeout
@@ -121,14 +260,31 @@ export async function solveDuckDuckGoChallenge(
);
const context = vm.createContext({});
vm.runInContext(stubs, context, { timeout: 5000 });
const startedAt = Date.now();
const result = (await vm.runInContext(js, context, {
timeout: 5000,
})) as DuckDuckGoChallengeResult;
const elapsedMs = Date.now() - startedAt;
const clientHashes = Array.isArray(result.client_hashes) ? result.client_hashes : [];
if (clientHashes.length === 0)
throw new Error("DuckDuckGo challenge returned empty client_hashes");
clientHashes[0] = userAgent;
result.client_hashes = clientHashes.map((hash) => sha256Base64(String(hash)));
// The real frontend augments the challenge's own `meta` with origin / stack /
// duration before sending it back. Omitting them yields 418 ERR_CHALLENGE even
// when every client_hash is correct (confirmed by capturing a real browser's
// x-vqd-hash-1 header, which always carries all three).
const origin = options.origin ?? DUCKDUCKGO_CHALLENGE_ORIGIN;
const bundlePath = options.bundlePath ?? "/dist/duckai-dist/entry.duckai.js";
const meta = (result.meta ?? {}) as Record<string, unknown>;
result.meta = {
...meta,
origin,
stack: buildChallengeStack(origin, bundlePath),
duration: String(elapsedMs),
};
return Buffer.from(JSON.stringify(result), "utf8").toString("base64");
}

View File

@@ -80,16 +80,7 @@ export class GeminiBusinessExecutor extends BaseExecutor {
// Extract cookies from credentials — check apiKey/cookie first, then
// try each __Secure-1PSID* key in providerSpecificData individually.
// A user with only __Secure-1PSID (no PSIDTS) is still valid.
const directCookie =
readCredentialString(credentials?.apiKey) || readCredentialString(credentials?.cookie);
const psid = readProviderSpecificString(credentials?.providerSpecificData, [
"__Secure-1PSID",
"cookie",
]);
const psidts = readProviderSpecificString(credentials?.providerSpecificData, [
"__Secure-1PSIDTS",
]);
const cookie = directCookie || [psid, psidts].filter(Boolean).join("; ");
const cookie = resolveGeminiBusinessCookie(credentials);
if (!cookie) {
return makeErrorResult(
@@ -380,6 +371,15 @@ function readProviderSpecificString(providerSpecificData: unknown, keys: string[
return "";
}
export function resolveGeminiBusinessCookie(credentials: unknown): string {
if (!credentials || typeof credentials !== "object") return "";
const data = credentials as Record<string, unknown>;
const directCookie = readCredentialString(data.apiKey) || readCredentialString(data.cookie);
const psid = readProviderSpecificString(data.providerSpecificData, ["__Secure-1PSID", "cookie"]);
const psidts = readProviderSpecificString(data.providerSpecificData, ["__Secure-1PSIDTS"]);
return directCookie || [psid, psidts].filter(Boolean).join("; ");
}
function extractTextContent(content: unknown): string {
if (typeof content === "string") return content.trim();
if (Array.isArray(content)) {

View File

@@ -25,6 +25,7 @@ import { ChatGptWebExecutor } from "./chatgpt-web.ts";
import { BlackboxWebExecutor } from "./blackbox-web.ts";
import { MuseSparkWebExecutor } from "./muse-spark-web.ts";
import { AzureOpenAIExecutor } from "./azure-openai.ts";
import { AzureAiExecutor } from "./azure-ai.ts";
import { CommandCodeExecutor } from "./commandCode.ts";
import { GitlabExecutor } from "./gitlab.ts";
import { NlpCloudExecutor } from "./nlpcloud.ts";
@@ -89,6 +90,7 @@ const executors = {
glmt: new GlmExecutor("glmt"),
cu: new CursorExecutor(), // Alias for cursor
"azure-openai": new AzureOpenAIExecutor(),
"azure-ai": new AzureAiExecutor(),
"command-code": new CommandCodeExecutor(),
cmd: new CommandCodeExecutor(), // Alias
gitlab: new GitlabExecutor(),
@@ -263,6 +265,7 @@ export { ChatGptWebExecutor } from "./chatgpt-web.ts";
export { BlackboxWebExecutor } from "./blackbox-web.ts";
export { MuseSparkWebExecutor } from "./muse-spark-web.ts";
export { AzureOpenAIExecutor } from "./azure-openai.ts";
export { AzureAiExecutor } from "./azure-ai.ts";
export { CommandCodeExecutor } from "./commandCode.ts";
export { GitlabExecutor } from "./gitlab.ts";
export { NlpCloudExecutor } from "./nlpcloud.ts";

View File

@@ -108,13 +108,9 @@ export function mapModel(model: string): string {
const TOKEN_SEED = "oldllm-client-2026";
const UA_PREFIX = CHROME_UA.slice(0, 20); // "Mozilla/5.0 (Windows"
type TheOldLlmProxy = {
type?: string;
host: string;
port: number;
username?: string | null;
password?: string | null;
} | null;
type TheOldLlmProxy = Awaited<
ReturnType<typeof import("../../src/lib/db/proxies").resolveProxyForProvider>
>;
interface TheOldLlmFetchDependencies {
resolveProxy: () => Promise<TheOldLlmProxy>;

View File

@@ -21,6 +21,7 @@ import { assembleStreamingResponseHeaders } from "./chatCore/streamingResponseHe
import { storeStreamingSemanticCacheResponse } from "./chatCore/streamingSemanticCacheStore.ts";
import { assembleStreamingPipeline } from "./chatCore/streamingPipeline.ts";
import { sanitizeChatRequestBody } from "./chatCore/sanitization.ts";
import { applyResponsesInputPolicy } from "../services/responsesInputPolicy.ts";
import {
getHeaderValueCaseInsensitive,
isNoMemoryRequested,
@@ -170,6 +171,7 @@ import {
ANTIGRAVITY_PRE_RESPONSE_TIMEOUT_CODE,
STREAM_RECOVERY,
DEFAULT_MAX_TOKENS,
STREAM_DISCONNECT_GRACE_PERIOD_MS,
} from "../config/constants.ts";
import { createRecoverableStream, makeContinuationBody } from "../services/streamRecovery.ts";
import {
@@ -207,7 +209,6 @@ import { stageTrace } from "./chatCore/stageTrace.ts";
import { attachCompressionUsageReceiptAfterAnalytics as attachCompressionUsageReceiptAfterAnalyticsFor } from "./chatCore/compressionUsageReceipt.ts";
import { prepareUpstreamBody } from "./chatCore/upstreamBody.ts";
import { getQuotaScopeLabelForProvider } from "../services/antigravityQuotaFamily.ts";
import {
getCallLogPipelineCaptureStreamChunks,
getCallLogPipelineMaxSizeBytes,
@@ -367,9 +368,7 @@ import {
isTpmExhausted,
isRpmExhausted,
} from "../services/geminiRateLimitTracker.ts";
import { isSmallEnoughForSemanticCache } from "../utils/estimateSize.ts";
/**
* Core chat handler - shared between SSE and Worker
* Returns { success, response, status, error } for caller to handle fallback
@@ -389,10 +388,8 @@ import { isSmallEnoughForSemanticCache } from "../utils/estimateSize.ts";
* @param {boolean} options.isCombo - Whether this request is from a combo
* @param {string} options.connectionId - Connection ID for settings lookup
*/
// extractSystemRoleMessages extracted to chatCore/claudeSystemRole.ts (#3501); re-exported above so
// existing importers (e.g. tests/unit/system-role-extraction.test.ts) keep resolving it from here.
export async function handleChatCore({
body,
modelInfo,
@@ -428,7 +425,6 @@ export async function handleChatCore({
/* fail open */
}
}
// Per-request model-routing metadata (first extracted slice of the request-setup phase).
const { apiFormat, customModelTargetFormat, requestedModel } = resolveChatCoreRequestSetup(
modelInfo,
@@ -442,7 +438,6 @@ export async function handleChatCore({
// (not Math.random) purely to satisfy CodeQL js/insecure-randomness — this id
// is a log-correlation token, not a security secret.
const traceId = globalThis.crypto.randomUUID().slice(0, 6);
// Emit request.started event for real-time dashboard
setImmediate(() => {
emit("request.started", {
@@ -526,7 +521,6 @@ export async function handleChatCore({
`long-running goal mode enabled: readinessMax=${agentGoalPolicy.readinessMaxTimeoutMs}ms streamRecovery=${agentGoalPolicy.streamRecoveryEnabled}`
);
}
let effectiveServiceTier: EffectiveServiceTier = "standard";
// Codex service-tier resolvers extracted to chatCore/serviceTier.ts (#3501); bind the per-request
// provider/credentials once and delegate so the existing call sites stay byte-identical.
@@ -555,7 +549,6 @@ export async function handleChatCore({
})
).catch(() => {});
};
// Key-health updater extracted to chatCore/keyHealth.ts (#3501); bind the per-request log once
// and delegate so the existing call sites stay byte-identical.
const recordKeyHealthStatus = (
@@ -563,11 +556,9 @@ export async function handleChatCore({
creds: Record<string, unknown> | null | undefined,
transport?: string
): void => recordKeyHealthStatusFor(status, creds, log, transport);
const persistCodexQuotaState = async (headers: Record<string, string> | null, status = 0) => {
const currentConnectionId = getCurrentConnectionId();
if (provider !== "codex" || !currentConnectionId || !headers) return;
try {
const existingProviderData =
credentials?.providerSpecificData && typeof credentials.providerSpecificData === "object"
@@ -582,28 +573,23 @@ export async function handleChatCore({
status,
});
if (!built) return;
if (built.exhaustionLog) {
log?.debug?.("CODEX", built.exhaustionLog);
}
// Invalidate the preflight cache for this connection so the next
// isModelAvailable check fetches fresh quota data.
if (status === 429) {
invalidateCodexQuotaCache(currentConnectionId);
}
await updateProviderConnection(currentConnectionId, {
providerSpecificData: built.nextProviderData,
});
credentials.providerSpecificData = built.nextProviderData;
} catch (err) {
const errMessage = err instanceof Error ? err.message : String(err);
log?.debug?.("CODEX", `Failed to persist codex quota state: ${errMessage}`);
}
};
// ── Phase 9.2: Idempotency check ──
// Resolve the idempotency key once here and reuse it at the Phase 9.2 save site below,
// rather than re-deriving it. (#3821-review LEDGER-6)
@@ -622,13 +608,11 @@ export async function handleChatCore({
if (idempotencyHit) {
return idempotencyHit;
}
// T07: Inject connectionId into credentials so executors can rotate API keys
// using providerSpecificData.extraApiKeys (API Key Round-Robin feature)
if (connectionId && credentials && !credentials.connectionId) {
credentials.connectionId = connectionId;
}
// Endpoint/format resolution extracted to chatCore/requestFormat.ts (#3501); pure derivation
// from the inbound request, destructured so every downstream use stays byte-identical.
const {
@@ -1071,6 +1055,13 @@ export async function handleChatCore({
return cacheHit;
}
if (targetFormat === FORMATS.OPENAI_RESPONSES && body && typeof body === "object") {
applyResponsesInputPolicy(
body as Record<string, unknown>,
credentials?.providerSpecificData?.preserveEncryptedReasoning === true
);
}
body = sanitizeChatRequestBody(body, sourceFormat, targetFormat);
// Per-request opt-out: clients that manage their own context send
// `x-omniroute-no-memory: true` to skip memory+skills injection (a null owner
@@ -2264,8 +2255,19 @@ export async function handleChatCore({
// the latter is a Kiro/Claude passthrough alias channel with string values,
// while namespace identities carry `{namespace, name}` for the #7936 response
// seam. Extract first because Kiro merge may reuse `_toolNameMap` below.
//
// #9780 — prefer the dedicated channel: on a pivot the openai->claude/gemini
// step publishes its own alias map on `_toolNameMap`, so that property alone
// yields aliases here. The `_toolNameMap` read stays as the fallback for the
// non-pivot producers (executors/base.ts, cliproxyapi.ts, antigravity).
const namespaceIdentityMap = translatedBody._namespaceToolIdentityMap;
const requestToolIdentityMap =
translatedBody._toolNameMap instanceof Map ? translatedBody._toolNameMap : null;
namespaceIdentityMap instanceof Map
? namespaceIdentityMap
: translatedBody._toolNameMap instanceof Map
? translatedBody._toolNameMap
: null;
delete translatedBody._namespaceToolIdentityMap;
delete translatedBody._toolNameMap;
// Kiro: sanitize tool schemas before dispatch. Kiro returns 400 "Improperly
@@ -4333,9 +4335,14 @@ export async function handleChatCore({
try {
const firstChoice = translatedResponse?.choices?.[0];
const msg = firstChoice?.message;
// The response being cached now will be replayed as history on the *next*
// turn, where the read side (translator/index.ts) keys the lookup by the
// message's real position in that future `messages` array — i.e. right
// after everything the client sent this turn.
const bodyMessages = (body as { messages?: unknown[] } | null | undefined)?.messages;
cacheReasoningFromAssistantMessage(msg, provider, model, {
requestId: skillRequestId,
messageIndex: 0,
messageIndex: Array.isArray(bodyMessages) ? bodyMessages.length : 0,
});
} catch {
// Cache capture is non-critical — never block the response
@@ -4760,12 +4767,15 @@ export async function handleChatCore({
// with tool_calls so it can be replayed on subsequent turns (DeepSeek V4, Kimi K2, etc.)
if (normalizedStreamStatus === 200 && streamResponseBody) {
try {
const body = streamResponseBody as Record<string, unknown>;
const choices = body.choices as { message?: Record<string, unknown> }[] | undefined;
const streamBody = streamResponseBody as Record<string, unknown>;
const choices = streamBody.choices as { message?: Record<string, unknown> }[] | undefined;
const msg = choices?.[0]?.message;
// See the non-streaming capture above: messageIndex must match the
// position this message will occupy in the *next* turn's history.
const bodyMessages = (body as { messages?: unknown[] } | null | undefined)?.messages;
cacheReasoningFromAssistantMessage(msg, provider, model, {
requestId: skillRequestId,
messageIndex: 0,
messageIndex: Array.isArray(bodyMessages) ? bodyMessages.length : 0,
});
} catch {
// Cache capture is non-critical — never block the stream
@@ -4908,13 +4918,20 @@ export async function handleChatCore({
});
const handleStreamFailure = streamFailureFinalizers.handleStreamFailure;
onPipelineStreamError = streamFailureFinalizers.onPipelineStreamError;
onClientDisconnectFinalize = (event) =>
handleStreamFailure({
status: 499,
message: `Client disconnected: ${event.reason}`,
code: "client_disconnected",
type: "client_disconnected",
});
// #9653: gives a genuine, race-delayed completion a chance to land (see
// createClientDisconnectGraceHandler's doc comment) before persisting a false
// 499/0-tokens for a request that actually delivered its full response.
onClientDisconnectFinalize = streamFailure.createClientDisconnectGraceHandler({
isStreamCompletionRecorded: () => streamCompletionRecorded,
gracePeriodMs: STREAM_DISCONNECT_GRACE_PERIOD_MS,
finalize: (event) =>
handleStreamFailure({
status: 499,
message: `Client disconnected: ${event.reason}`,
code: "client_disconnected",
type: "client_disconnected",
}),
});
// For providers using Responses API format, translate stream back to openai (Chat Completions) format
// UNLESS client is Droid CLI which expects openai-responses format back
@@ -5025,7 +5042,6 @@ export async function handleChatCore({
}),
};
}
export function isTokenExpiringSoon(expiresAt, bufferMs = 5 * 60 * 1000) {
if (!expiresAt) return false;
const expiresAtMs = new Date(expiresAt).getTime();

View File

@@ -3,11 +3,11 @@ import {
getChatLogMaxDepth,
getChatLogArrayTailItems,
getChatLogMaxObjectKeys,
getChatLogMaxBodyBytes,
} from "@/lib/logEnv";
import { estimateSizeFast } from "../../utils/estimateSize.ts";
export const MEMORY_EXTRACTION_TEXT_LIMIT = 64 * 1024;
const MAX_LOG_BODY_CHARS = 8 * 1024; // 8KB cap for logged request/response bodies
export function capMemoryExtractionText(value: string): string {
if (value.length <= MEMORY_EXTRACTION_TEXT_LIMIT) return value;
@@ -60,9 +60,10 @@ export function cloneBoundedChatLogPayload(value: unknown, depth = 0): unknown {
/**
* Truncate a large object for logging. If its JSON representation exceeds
* MAX_LOG_BODY_CHARS, return a lightweight summary instead of the full clone.
* This prevents persistAttemptLogs from holding multi-MB references to
* translatedBody across 17 call sites per request.
* the configured max body size (getChatLogMaxBodyBytes()), return a
* lightweight summary instead of the full clone. This prevents
* persistAttemptLogs from holding multi-MB references to translatedBody
* across 17 call sites per request.
*
* When the summarized object carries a `tools` definition, re-attach it
* (bounded via `cloneBoundedChatLogPayload`) so the request-details view can
@@ -75,8 +76,9 @@ export function cloneBoundedChatLogPayload(value: unknown, depth = 0): unknown {
export function truncateForLog(value: unknown): Record<string, unknown> | null | undefined {
if (value === null || value === undefined) return value as null | undefined;
if (typeof value !== "object") return value as unknown as Record<string, unknown>;
const estimatedSize = estimateSizeFast(value);
if (estimatedSize <= MAX_LOG_BODY_CHARS) return value as Record<string, unknown>;
const maxBodyBytes = getChatLogMaxBodyBytes();
const estimatedSize = estimateSizeFast(value, maxBodyBytes);
if (estimatedSize <= maxBodyBytes) return value as Record<string, unknown>;
// Object is too large — return a summary instead of a deep clone
const obj = value as Record<string, unknown>;
const summary: Record<string, unknown> = {

View File

@@ -15,14 +15,15 @@ import { saveImageErrorResult, saveImageSuccessResult } from "../../imageGenerat
import {
AdobeFireflyError,
adobeFireflyGenerateImage,
adobeFireflyImageTimeoutMs,
resolveAdobeAccessToken,
resolveAdobeSourceImageReferences,
resolveAdobeSourceImageIds,
resolveAdobeImageModel,
} from "../../../services/adobeFireflyClient.ts";
import { getAdobeReferenceUploadLimit } from "../../../services/adobeFireflyModels.ts";
import { isAdobeFireflyUpscaleModel } from "../../../services/adobeFireflyUpscale.ts";
import { handleAdobeFireflyImageUpscale } from "../../imageUpscale/adobeFirefly.ts";
import { ensureAdobeFireflySession } from "../../../services/adobeFireflySession.ts";
function normalizePositiveNumber(value: unknown, fallback: number): number {
const n = Number(value);
return Number.isFinite(n) && n > 0 ? n : fallback;
}
export async function handleAdobeFireflyImageGeneration({
model,
@@ -50,25 +51,22 @@ export async function handleAdobeFireflyImageGeneration({
images?: unknown;
[key: string]: unknown;
};
credentials: { apiKey?: string; accessToken?: string };
credentials: {
apiKey?: string;
accessToken?: string;
connectionId?: string;
providerSpecificData?: {
cookie?: unknown;
access_token?: unknown;
accessToken?: unknown;
browserSessionKey?: unknown;
} | null;
};
log?: { info?: (...args: unknown[]) => void; error?: (...args: unknown[]) => void };
fetchImpl?: typeof fetch;
}) {
const startTime = Date.now();
const prompt = typeof body.prompt === "string" ? body.prompt.trim() : "";
// Topaz upscalers share adobe-firefly but use /v2/3p-images/upsample (no prompt).
if (isAdobeFireflyUpscaleModel(model)) {
return handleAdobeFireflyImageUpscale({
model,
provider,
body: body as Record<string, unknown>,
credentials,
log,
fetchImpl,
});
}
if (!prompt) {
return saveImageErrorResult({
provider,
@@ -80,7 +78,17 @@ export async function handleAdobeFireflyImageGeneration({
}
try {
const accessToken = await resolveAdobeAccessToken(credentials, fetchImpl);
// Durable session: JWT + Cookie once → auto-rebuild ARP from forter/arkose,
// cache, optional Playwright warm-up. Submit path rotates ARP on 408.
const session = await ensureAdobeFireflySession({
credentials,
fetchImpl,
log,
});
const accessToken = session.accessToken;
const sessionCookie = session.cookie || undefined;
const arpSessionId = session.arpSessionId;
const timeoutMs = normalizePositiveNumber(body.timeout_ms, 180_000);
const seed =
typeof body.seed === "number"
? body.seed
@@ -88,44 +96,26 @@ export async function handleAdobeFireflyImageGeneration({
? Number(body.seed)
: undefined;
// Keep the raw credential blob for Cookie + sherlockToken (x-arp-session-id).
// JWT may be embedded in the same paste as cookies (HAR / multi-line).
const psd = (credentials as { providerSpecificData?: { cookie?: string } })
?.providerSpecificData;
const sessionCookie =
(typeof psd?.cookie === "string" && psd.cookie.trim()) ||
(typeof credentials?.apiKey === "string" && credentials.apiKey.trim()) ||
(typeof credentials?.accessToken === "string" && credentials.accessToken.includes(";")
? credentials.accessToken
: undefined);
// Cap uploads by model family (matches MediaViewModel GetSourceImageLimit).
const { id: resolvedId } = resolveAdobeImageModel(model);
const maxRefs = resolvedId.includes("nano-banana") || resolvedId.includes("gpt-image") ? 4 : 2;
const { spec } = resolveAdobeImageModel(model);
const references = await resolveAdobeSourceImageReferences({
const sourceImageIds = await resolveAdobeSourceImageIds({
accessToken,
body,
max: getAdobeReferenceUploadLimit(spec, "image"),
max: maxRefs,
sessionCookie,
arpSessionId,
prompt,
fetchImpl,
log,
});
const explicitTimeout =
typeof body.timeout_ms === "number"
? body.timeout_ms
: typeof body.timeout_ms === "string" && body.timeout_ms.trim()
? Number(body.timeout_ms)
: undefined;
const timeoutMs = adobeFireflyImageTimeoutMs({
timeoutMs: explicitTimeout,
refCount: references.length,
});
log?.info?.(
"IMAGE",
`${provider}/${model} (adobe-firefly) | prompt: "${prompt.slice(0, 60)}${prompt.length > 60 ? "..." : ""}"` +
(references.length ? ` | refs: ${references.length}` : "") +
` | pollTimeoutMs=${timeoutMs}`
(sourceImageIds.length ? ` | refs: ${sourceImageIds.length}` : "") +
` | session=${session.source}`
);
const result = await adobeFireflyGenerateImage({
@@ -137,8 +127,11 @@ export async function handleAdobeFireflyImageGeneration({
quality: body.quality,
seed: Number.isFinite(seed as number) ? (seed as number) : undefined,
negativePrompt: typeof body.negative_prompt === "string" ? body.negative_prompt : undefined,
references: references.length ? references : undefined,
sourceImageIds: sourceImageIds.length ? sourceImageIds : undefined,
sessionCookie,
arpSessionId,
sessionFingerprint: session.fingerprint,
sessionBrowserKey: session.browserSessionKey,
timeoutMs,
fetchImpl,
log,

View File

@@ -40,7 +40,12 @@ export async function handleResponsesCore({
const customToolNames = collectResponsesCustomToolNames(body?.tools, inputItems);
// Convert Responses API format to Chat Completions format
const convertedBody = convertResponsesApiFormat(body, credentials, modelInfo?.provider);
const convertedBody = convertResponsesApiFormat(
body,
credentials,
modelInfo?.provider,
modelInfo?.model
);
// Ensure stream is enabled
convertedBody.stream = true;

View File

@@ -4,7 +4,7 @@
* Handles POST /v1/videos/generations requests. Proxies to upstream video
* generation providers (ComfyUI AnimateDiff/SVD, SD WebUI AnimateDiff, and
* more — see the per-format handlers below). Response format (OpenAI-like):
* { "created": 1234567890, "data": [{ "b64_json": "...", "format": "mp4" }] }
* { "created": 1234567890, "data": [{ "url": "https://…", "format": "mp4" }] }
*/
import { getVideoProvider, parseVideoModel } from "../config/videoRegistry.ts";
@@ -18,6 +18,16 @@ import { handleNovitaVideoGeneration } from "./videoGeneration/novitaHandler.ts"
import { handleXaiVideoGeneration } from "./videoGeneration/xaiGrokImagineHandler.ts";
import { handleSegmindVideoGeneration } from "./videoGeneration/providers/segmind.ts";
import { handleAdobeFireflyVideoGeneration } from "./videoGeneration/adobeFireflyHandler.ts";
import { handleOpenAIVideoGeneration } from "./videoGeneration/openai.ts";
import { getVideoJobPreset, handleVideoJobGeneration } from "./videoGeneration/job.ts";
import {
extractRunwayFailureMessage,
normalizeRunwayVideoResult,
resolvePositiveInteger,
resolveRunwayDuration,
resolveRunwayPromptImage,
resolveRunwayRatio,
} from "./videoGeneration/runwayHelpers.ts";
import { getExecutor } from "../executors/index.ts";
import { getKieTaskId, isJsonObject, parseKieResultJson } from "../utils/kieTask.ts";
import {
@@ -33,13 +43,94 @@ import {
resolveComfyUiBaseUrl,
} from "../utils/comfyuiClient.ts";
import { saveCallLog } from "@/lib/usageDb";
import { getAllCustomModels } from "@/lib/db/models";
import { sanitizeErrorMessage } from "../utils/error.ts";
import {
FetchTimeoutError,
fetchWithTimeout,
getConfiguredTimeout,
} from "@/shared/utils/fetchTimeout";
/**
* Resolve the base URL for OpenAI-compatible video generation endpoints.
* Prefers providerSpecificData.baseUrl (from custom node config), falls back to
* top-level credentials.baseUrl, then to the provided fallback.
*/
export function resolveVideoBaseUrl(
credentials:
{ baseUrl?: unknown; providerSpecificData?: { baseUrl?: unknown } | null } | null | undefined,
fallback: string
): string {
const psd = credentials?.providerSpecificData;
const psdBaseUrl =
psd && typeof psd === "object" && typeof psd.baseUrl === "string" && psd.baseUrl.trim()
? psd.baseUrl.trim()
: null;
const topLevelBaseUrl =
typeof credentials?.baseUrl === "string" && credentials.baseUrl.trim()
? credentials.baseUrl.trim()
: null;
const nodeBaseUrl = psdBaseUrl || topLevelBaseUrl;
if (!nodeBaseUrl) return fallback;
// Trim trailing slashes
let normalized = nodeBaseUrl;
while (normalized.endsWith("/")) normalized = normalized.slice(0, -1);
if (normalized.endsWith("/videos/generations")) return normalized;
const stripped = normalized.replace(/\/videos\/generations$/, "");
return `${stripped}/videos/generations`;
}
/**
* Read generationConfig.preset from the custom model row for the given
* provider/model id. Returns null when the model has no preset configured (or
* the registry is unreadable), so callers can fall back to the sync path.
*/
async function getCustomModelVideoPreset(
providerId: string,
modelId: string
): Promise<string | null> {
try {
const customModelsMap = (await getAllCustomModels()) as Record<
string,
Array<Record<string, unknown>>
>;
const models = customModelsMap[providerId];
if (!Array.isArray(models)) return null;
for (const model of models) {
if (!model || typeof model !== "object" || model.id !== modelId) continue;
const generationConfig = model.generationConfig;
if (
generationConfig &&
typeof generationConfig === "object" &&
typeof (generationConfig as Record<string, unknown>).preset === "string"
) {
return (generationConfig as Record<string, unknown>).preset as string;
}
return null;
}
return null;
} catch {
return null;
}
}
/**
* Handle video generation request
*/
export async function handleVideoGeneration({ body, credentials, log }) {
const { provider, model } = parseVideoModel(body.model);
/**
* Handle video generation request
*/
export async function handleVideoGeneration({ body, credentials, log, resolvedProvider = null }) {
let { provider, model } = parseVideoModel(body.model);
if (resolvedProvider) {
provider = resolvedProvider;
model = body.model.startsWith(provider + "/")
? body.model.slice(provider.length + 1)
: body.model;
}
if (!provider) {
return {
@@ -51,11 +142,59 @@ export async function handleVideoGeneration({ body, credentials, log }) {
const providerConfig = getVideoProvider(provider);
if (!providerConfig) {
return {
success: false,
status: 400,
error: `Unknown video provider: ${provider}`,
if (!resolvedProvider) {
return {
success: false,
status: 400,
error: `Unknown video provider: ${provider}`,
};
}
// Custom provider node. When the custom model row carries a
// generationConfig.preset (e.g. "agnes-video-job"), dispatch through the
// submit → poll job pipeline; otherwise mirror the images route and use the
// generic OpenAI-compatible handler with a synthetic config.
const presetName = await getCustomModelVideoPreset(provider, model);
if (presetName !== null) {
if (!getVideoJobPreset(presetName)) {
return {
success: false,
status: 502,
error: `Unknown video job preset: ${presetName}`,
};
}
if (log)
log.info("VIDEO", `Custom model ${provider}/${model} — using job preset ${presetName}`);
return handleVideoJobGeneration({
model,
presetName,
body,
credentials,
log,
});
}
if (log)
log.info("VIDEO", `Custom model ${provider}/${model} — using OpenAI-compatible handler`);
const syntheticConfig = {
id: provider,
baseUrl: resolveVideoBaseUrl(
credentials,
"http://generative.language.googleapis.com/v1beta/openai/videos/generations"
),
authType: "apikey",
authHeader: "bearer",
format: "openai-video",
};
return handleOpenAIVideoGeneration({
model,
body,
credentials,
provider,
providerConfig: syntheticConfig,
log,
});
}
if (providerConfig.format === "openai-video") {
return handleOpenAIVideoGeneration({ model, provider, providerConfig, body, credentials, log });
}
if (providerConfig.format === "vertex-veo") {
@@ -158,7 +297,10 @@ export async function handleVideoGeneration({ body, credentials, log }) {
log,
});
}
if (resolvedProvider) {
// Custom provider with no matching built-in format — use OpenAI-compatible fallback
return handleOpenAIVideoGeneration({ model, provider, providerConfig, body, credentials, log });
}
return {
success: false,
status: 400,
@@ -832,148 +974,6 @@ const RUNWAY_TERMINAL_FAILURE_STATUSES = new Set([
"DELETED",
]);
function resolveRunwayPromptImage(body) {
const directCandidates = [
body.promptImage,
body.prompt_image,
body.image,
body.image_url,
body.imageUrl,
body.provider_options?.promptImage,
body.provider_options?.prompt_image,
];
for (const candidate of directCandidates) {
if (typeof candidate === "string" && candidate.trim()) return candidate.trim();
if (candidate && typeof candidate === "object") return candidate;
if (Array.isArray(candidate) && candidate.length > 0) return candidate;
}
const arrayCandidates = [
body.imageUrls,
body.image_urls,
body.provider_options?.imageUrls,
body.provider_options?.image_urls,
];
for (const candidate of arrayCandidates) {
if (Array.isArray(candidate) && candidate.length > 0) return candidate;
}
return null;
}
function resolveRunwayRatio(body) {
const aspectRatio = typeof body.aspect_ratio === "string" ? body.aspect_ratio : body.aspectRatio;
if (aspectRatio === "1280:720" || aspectRatio === "720:1280") return aspectRatio;
if (aspectRatio === "16:9") return "1280:720";
if (aspectRatio === "9:16") return "720:1280";
const size = typeof body.size === "string" ? body.size : "";
const [widthRaw, heightRaw] = size.split("x");
const width = Number(widthRaw);
const height = Number(heightRaw);
if (Number.isFinite(width) && Number.isFinite(height) && width > 0 && height > 0) {
return width >= height ? "1280:720" : "720:1280";
}
return "1280:720";
}
function resolveRunwayDuration(body) {
if (Number.isFinite(body.duration)) {
return clampRunwayDuration(body.duration);
}
if (Number.isFinite(body.frames) && Number.isFinite(body.fps) && Number(body.fps) > 0) {
return clampRunwayDuration(Number(body.frames) / Number(body.fps));
}
return 5;
}
function clampRunwayDuration(value) {
const duration = Math.round(Number(value));
if (!Number.isFinite(duration)) return 5;
return Math.min(10, Math.max(2, duration));
}
function resolvePositiveInteger(value, fallback) {
const numeric = Number(value);
if (!Number.isFinite(numeric) || numeric <= 0) return fallback;
return Math.floor(numeric);
}
function extractRunwayOutputUrls(task) {
const rawOutput = Array.isArray(task?.output)
? task.output
: Array.isArray(task?.result)
? task.result
: [];
return rawOutput
.map((entry) => {
if (typeof entry === "string") return entry;
if (!entry || typeof entry !== "object") return null;
return entry.url || entry.uri || entry.videoUrl || entry.video_url || null;
})
.filter((value) => typeof value === "string" && value.length > 0);
}
function extractRunwayFailureMessage(task) {
const directCandidates = [
task?.failure,
task?.failureReason,
task?.error,
task?.errorMessage,
task?.message,
];
for (const candidate of directCandidates) {
if (typeof candidate === "string" && candidate.trim()) return candidate.trim();
}
if (task?.failure && typeof task.failure === "object") {
const nestedCandidates = [
task.failure.message,
task.failure.reason,
task.failure.error,
task.failure.code,
];
for (const candidate of nestedCandidates) {
if (typeof candidate === "string" && candidate.trim()) return candidate.trim();
}
}
return null;
}
async function normalizeRunwayVideoResult(task, body) {
const urls = extractRunwayOutputUrls(task);
if (urls.length === 0) {
throw new Error(
`Runway task completed without output URLs: ${JSON.stringify(task).slice(0, 400)}`
);
}
if (body.response_format === "url") {
return urls.map((url) => ({ url, format: "mp4" }));
}
const videos = [];
for (const url of urls) {
const response = await fetch(url);
if (!response.ok) {
throw new Error(`Runway output fetch failed (${response.status})`);
}
const arrayBuffer = await response.arrayBuffer();
videos.push({
b64_json: Buffer.from(arrayBuffer).toString("base64"),
format: "mp4",
});
}
return videos;
}
async function handleHaiperVideoGeneration({
model,
provider,

View File

@@ -9,11 +9,10 @@ import { sanitizeErrorMessage } from "../../utils/error.ts";
import {
AdobeFireflyError,
adobeFireflyGenerateVideo,
resolveAdobeAccessToken,
resolveAdobeSourceImageReferences,
resolveAdobeSourceImageIds,
resolveAdobeVideoModel,
} from "../../services/adobeFireflyClient.ts";
import { getAdobeReferenceUploadLimit } from "../../services/adobeFireflyModels.ts";
import { ensureAdobeFireflySession } from "../../services/adobeFireflySession.ts";
function normalizePositiveNumber(value: unknown, fallback: number): number {
const n = Number(value);
@@ -32,7 +31,17 @@ export async function handleAdobeFireflyVideoGeneration({
provider: string;
providerConfig?: { baseUrl?: string };
body: Record<string, unknown>;
credentials?: { apiKey?: string; accessToken?: string } | null;
credentials?: {
apiKey?: string;
accessToken?: string;
connectionId?: string;
providerSpecificData?: {
cookie?: unknown;
access_token?: unknown;
accessToken?: unknown;
browserSessionKey?: unknown;
} | null;
} | null;
log?: { info?: (...args: unknown[]) => void; error?: (...args: unknown[]) => void };
fetchImpl?: typeof fetch;
}) {
@@ -47,7 +56,14 @@ export async function handleAdobeFireflyVideoGeneration({
}
try {
const accessToken = await resolveAdobeAccessToken(credentials, fetchImpl);
const session = await ensureAdobeFireflySession({
credentials,
fetchImpl,
log,
});
const accessToken = session.accessToken;
const sessionCookie = session.cookie || undefined;
const arpSessionId = session.arpSessionId;
const timeoutMs = normalizePositiveNumber(body.timeout_ms, 300_000);
const seed =
typeof body.seed === "number"
@@ -55,22 +71,16 @@ export async function handleAdobeFireflyVideoGeneration({
: typeof body.seed === "string" && String(body.seed).trim()
? Number(body.seed)
: undefined;
// Keep raw paste for Cookie + sherlockToken (x-arp-session-id).
const psd = (credentials as { providerSpecificData?: { cookie?: string } })
?.providerSpecificData;
const sessionCookie =
(typeof psd?.cookie === "string" && psd.cookie.trim()) ||
(typeof credentials?.apiKey === "string" && credentials.apiKey.trim()) ||
(typeof credentials?.accessToken === "string" && credentials.accessToken.includes(";")
? credentials.accessToken
: undefined);
const { spec } = resolveAdobeVideoModel(String(model));
const references = await resolveAdobeSourceImageReferences({
// Kling i2v / Veo ref / Sora frame: upload reference images first.
const { id: videoModelId } = resolveAdobeVideoModel(String(model));
const maxFrames = videoModelId.includes("kling") || videoModelId.includes("sora") ? 2 : 3;
const sourceImageIds = await resolveAdobeSourceImageIds({
accessToken,
body,
max: getAdobeReferenceUploadLimit(spec, "image"),
max: maxFrames,
sessionCookie,
arpSessionId,
prompt,
fetchImpl,
log,
@@ -79,7 +89,8 @@ export async function handleAdobeFireflyVideoGeneration({
log?.info?.(
"VIDEO",
`${provider}/${model} (adobe-firefly) | prompt: "${prompt.slice(0, 60)}${prompt.length > 60 ? "..." : ""}"` +
(references.length ? ` | refs: ${references.length}` : "")
(sourceImageIds.length ? ` | frames: ${sourceImageIds.length}` : "") +
` | session=${session.source}`
);
const result = await adobeFireflyGenerateVideo({
@@ -99,8 +110,11 @@ export async function handleAdobeFireflyVideoGeneration({
? body.negativePrompt
: undefined,
generateAudio: body.generate_audio !== false && body.generateAudio !== false,
references: references.length ? references : undefined,
sourceImageIds: sourceImageIds.length ? sourceImageIds : undefined,
sessionCookie,
arpSessionId,
sessionFingerprint: session.fingerprint,
sessionBrowserKey: session.browserSessionKey,
timeoutMs,
fetchImpl,
log,

View File

@@ -0,0 +1,418 @@
/**
* Async job/poll video generation for custom OpenAI-compatible provider nodes
* whose /videos surface is a submit → poll → fetch-result API (e.g. Agnes
* Video V2.0, muapi.ai, OpenAI Sora). Presets are declarative data — the
* handler here is one family; everything else is per-preset config.
*
* Response shape stays OpenAI-like: { created, data: [{ url, format: "mp4" }] } so the
* /v1/videos/generations route returns the same contract as the synchronous
* path.
*/
import {
fetchWithTimeout,
FetchTimeoutError,
getConfiguredTimeout,
} from "@/shared/utils/fetchTimeout";
import { sanitizeErrorMessage } from "../../utils/error.ts";
import { sleep } from "../../utils/sleep.ts";
interface LogLike {
info?: (tag: string, msg: string, meta?: unknown) => void;
warn?: (tag: string, msg: string, meta?: unknown) => void;
error?: (tag: string, msg: string, meta?: unknown) => void;
}
interface CredentialsLike {
providerSpecificData?: { baseUrl?: unknown } | null;
baseUrl?: unknown;
apiKey?: unknown;
accessToken?: unknown;
}
/** Dot-path reader restricted to plain objects/arrays (no prototypes). */
function readPath(value: unknown, path: string): unknown {
if (!path) return value;
let current: unknown = value;
for (const segment of path.split(".")) {
if (current === null || current === undefined) return undefined;
if (typeof current !== "object") return undefined;
if (Array.isArray(current)) {
const index = Number(segment);
if (!Number.isInteger(index) || index < 0 || index >= current.length) return undefined;
current = current[index];
continue;
}
if (!Object.prototype.hasOwnProperty.call(current, segment)) return undefined;
current = (current as Record<string, unknown>)[segment];
}
return current;
}
/** Non-empty string from a dot path, or null. */
function readStringPath(value: unknown, path: string): string | null {
const found = readPath(value, path);
return typeof found === "string" && found.trim() ? found : null;
}
function isDoneStatus(
status: unknown,
done: string[],
failed: string[]
): "done" | "failed" | "pending" {
if (typeof status !== "string") return "pending";
if (failed.includes(status)) return "failed";
if (done.includes(status)) return "done";
return "pending";
}
export type VideoJobPreset = {
id: string;
displayName: string;
/** auth header name plus value scheme */
authHeaderName: "x-api-key" | "Authorization";
authScheme: "bearer" | "raw";
baseUrlFallback: string;
submit: {
method: "POST";
/** may contain {model} — substituted before POST */
path: string;
buildBody: (params: {
model?: string;
prompt?: string;
duration?: number;
extras: Record<string, unknown>;
}) => Record<string, unknown>;
};
/** dot path into the submit response identifying the job */
taskIdPath: string;
poll: {
/** contains {taskId} */
pathTemplate: string;
};
statusPath: string;
statusDone: string[];
statusFailed: string[];
/** dot path into the poll response holding the finished video URL/array */
resultPath: string;
maxPolls: number;
pollIntervalMs: number;
};
// #9820: declarative presets for the shipping async job/poll video providers.
const VIDEO_JOB_PRESETS: Record<string, VideoJobPreset> = {
"agnes-video-job": {
id: "agnes-video-job",
displayName: "Agnes Video V2.0",
authHeaderName: "x-api-key",
authScheme: "raw",
// Real default, matching the Agnes Video V2.0 reference: POST /v1/videos with
// x-api-key auth; GET /v1/videos/{task_id} returns status/progress/metadata.
baseUrlFallback: "https://apihub.agnes-ai.com",
submit: {
method: "POST",
path: "/v1/videos",
buildBody: ({ model, prompt, extras }) => ({
model,
prompt,
// passthrough of image/mode/num_frames/frame_rate/… — the generic
// route body uses .catchall, so provider-specific knobs survive.
...extras,
}),
},
taskIdPath: "task_id",
poll: { pathTemplate: "/v1/videos/{taskId}" },
statusPath: "status",
statusDone: ["completed"],
statusFailed: ["failed"],
resultPath: "metadata.url",
maxPolls: 60,
pollIntervalMs: 2000,
},
"muapi-video-job": {
id: "muapi-video-job",
displayName: "muapi.ai",
authHeaderName: "x-api-key",
authScheme: "raw",
// muapi.ai video/audio surface is Replicate-style: POST /api/v1/{model}
// returns { request_id }; poll GET /api/v1/predictions/{id}/result.
baseUrlFallback: "https://api.muapi.ai",
submit: {
method: "POST",
path: "/api/v1/{model}",
buildBody: (params) => {
const { prompt, duration, extras } = params;
return {
prompt,
...(typeof duration === "number" ? { duration } : {}),
...extras,
};
},
},
taskIdPath: "request_id",
poll: { pathTemplate: "/api/v1/predictions/{taskId}/result" },
statusPath: "status",
statusDone: ["completed"],
statusFailed: ["failed"],
resultPath: "outputs",
maxPolls: 60,
pollIntervalMs: 2000,
},
"sora-job": {
id: "sora-job",
displayName: "OpenAI Sora",
authHeaderName: "Authorization",
authScheme: "bearer",
baseUrlFallback: "https://api.openai.com",
submit: {
method: "POST",
path: "/v1/videos",
buildBody: (params) => {
const { model, prompt, duration, extras } = params;
// seconds is a STRING enum ("4"|"8"|"12") in the Sora API; absolute
// size mapping is intentionally not forced here.
return {
model,
prompt,
...(typeof duration === "number" ? { seconds: String(duration) } : {}),
...extras,
};
},
},
taskIdPath: "id",
poll: { pathTemplate: "/v1/videos/{taskId}" },
statusPath: "status",
statusDone: ["completed"],
statusFailed: ["failed"],
resultPath: "data",
maxPolls: 60,
pollIntervalMs: 2000,
},
};
/** Resolve a configured job preset; null when the preset is unknown/none. */
export function getVideoJobPreset(presetName: unknown): VideoJobPreset | null {
if (typeof presetName !== "string") return null;
const preset = VIDEO_JOB_PRESETS[presetName];
return preset ?? null;
}
/**
* Handle a video-generation job via the submit→poll preset pipeline.
* Returns the same shape as the sync handlers: { success, data?: …, status?, error? }.
*/
export async function handleVideoJobGeneration({
model,
presetName,
body,
credentials,
log,
maxPolls: maxPollsOverride,
pollIntervalMs: pollIntervalOverride,
}: {
model: string;
presetName: string;
body: Record<string, unknown>;
credentials?: unknown;
log?: {
info?: (tag: string, msg: string, meta?: unknown) => void;
error?: (tag: string, msg: string) => void;
};
maxPolls?: number;
pollIntervalMs?: number;
}) {
const preset = getVideoJobPreset(presetName);
if (!preset) {
return {
success: false,
status: 400,
error: `Unknown video job preset: ${presetName}`,
};
}
const baseUrl = resolveJobBaseUrl(credentials, preset.baseUrlFallback);
log?.info?.("VIDEO", `Job preset ${presetName} submitting ${model}`);
log?.info?.("VIDEO", JSON.stringify({ baseUrl }));
const bodyForPreset = preset.submit.buildBody({
model: model,
prompt: typeof body.prompt === "string" ? body.prompt : undefined,
duration: typeof body.duration === "number" ? body.duration : undefined,
// passthrough of the remainder — the API keeps catchall extras
extras: Object.fromEntries(
Object.entries(body ?? {}).filter(
([key]) => key !== "model" && key !== "prompt" && key !== "duration"
)
),
});
const submitPath = preset.submit.path.replace("{model}", encodeURIComponent(model));
const submitUrl = `${baseUrl}${submitPath}`; // baseUrl never ends with "/"
const submitResult = await fetchJson(submitUrl, {
method: preset.submit.method,
headers: buildJobHeaders(preset, credentials),
body: JSON.stringify(bodyForPreset),
log,
});
if (!submitResult.ok) {
return { success: false, status: submitResult.status, error: submitResult.error };
}
const taskId = readStringPath(submitResult.data, preset.taskIdPath);
if (!taskId) {
return {
success: false,
status: 502,
error: `Video provider did not return a job id (${presetName})`,
};
}
// Poll loop.
const maxPolls = maxPollsOverride ?? preset.maxPolls;
const pollInterval = pollIntervalOverride ?? preset.pollIntervalMs;
for (let attempt = 1; attempt <= maxPolls; attempt += 1) {
await sleep(pollInterval);
const pollUrl = `${baseUrl}${preset.poll.pathTemplate.replace("{taskId}", encodeURIComponent(taskId))}`;
const pollResult = await fetchJson(pollUrl, {
method: "GET",
headers: buildJobHeaders(preset, credentials),
log,
});
if (!pollResult.ok) {
return { success: false, status: pollResult.status, error: pollResult.error };
}
const status = readPath(pollResult.data, preset.statusPath);
const jobState = isDoneStatus(status, preset.statusDone, preset.statusFailed);
if (jobState === "done") {
const url = readResultUrl(pollResult.data, preset.resultPath);
if (!url) {
return {
success: false,
status: 502,
error: `Video job completed but no result URL found (${presetName})`,
};
}
log?.info?.("VIDEO", `Job completed after ${attempt} poll(s)`);
return {
success: true,
data: {
created: Math.floor(Date.now() / 1000),
data: [{ url, format: "mp4" }],
},
};
}
if (jobState === "failed") {
return {
success: false,
status: 502,
error: `Video job failed (${presetName})`,
};
}
}
return {
success: false,
status: 504,
error: `Video job timed out after ${maxPolls} polls (${presetName})`,
};
}
function buildJobHeaders(preset: VideoJobPreset, credentials?: unknown): Record<string, string> {
const creds = credentials as CredentialsLike | null | undefined;
const apiKey =
typeof creds?.apiKey === "string" && creds.apiKey
? creds.apiKey
: typeof creds?.accessToken === "string" && creds.accessToken
? creds.accessToken
: "";
const headers: Record<string, string> = { "Content-Type": "application/json" };
if (!apiKey) return headers;
if (preset.authScheme === "raw") {
headers[preset.authHeaderName] = apiKey;
} else {
headers[preset.authHeaderName] = `Bearer ${apiKey}`;
}
return headers;
}
function resolveJobBaseUrl(credentials: unknown, fallback: string): string {
const creds = credentials as CredentialsLike | null | undefined;
const psdBaseUrl =
creds?.providerSpecificData?.baseUrl != null &&
typeof creds.providerSpecificData.baseUrl === "string" &&
creds.providerSpecificData.baseUrl.trim()
? (creds.providerSpecificData.baseUrl as string).trim()
: null;
const topLevelBaseUrl =
creds?.baseUrl != null && typeof creds.baseUrl === "string" && creds.baseUrl.trim()
? (creds.baseUrl as string).trim()
: null;
const nodeBaseUrl = psdBaseUrl || topLevelBaseUrl;
if (!nodeBaseUrl) return fallback.replace(/\/+$/, "");
let normalized = nodeBaseUrl;
while (normalized.endsWith("/")) normalized = normalized.slice(0, -1);
return normalized;
}
async function fetchJson(
url: string,
{
method,
headers,
body,
log,
}: {
method: string;
headers: Record<string, string>;
body?: string;
log?: LogLike;
}
): Promise<{ ok: true; data: unknown } | { ok: false; status: number; error: string }> {
try {
const response = await fetchWithTimeout(url, {
method,
headers,
...(body !== undefined ? { body } : {}),
timeoutMs: getConfiguredTimeout(),
});
if (!response.ok) {
const errorText = await response.text();
log?.error?.("VIDEO", `Upstream ${response.status} for ${url}: ${errorText.slice(0, 200)}`);
return { ok: false, status: response.status, error: errorText };
}
const data = await response.json();
return { ok: true, data };
} catch (err: unknown) {
const message = err instanceof Error ? err.message : String(err);
const isTimeout =
err instanceof FetchTimeoutError || (err instanceof Error && err.name === "AbortError");
log?.error?.(
"VIDEO",
`${isTimeout ? "Timeout" : "Request error"} for ${url}: ${sanitizeErrorMessage(message)}`
);
return {
ok: false,
status: isTimeout ? 504 : 502,
error: `Video provider error: ${sanitizeErrorMessage(message)}`,
};
}
}
function readResultUrl(data: unknown, resultPath: string): string | null {
const found = readPath(data, resultPath);
if (typeof found === "string" && found.trim()) return found.trim();
if (Array.isArray(found)) {
const first = found[0];
// muapi-style: resultPath "outputs" resolves to ["https://…"].
if (typeof first === "string" && first.trim()) return first.trim();
// sora-style: resultPath "data" resolves to [{ url: "https://…" }].
if (first && typeof first === "object" && !Array.isArray(first)) {
const urlEntry = (first as Record<string, unknown>).url;
if (typeof urlEntry === "string" && urlEntry.trim()) return urlEntry.trim();
}
return null;
}
return null;
}

View File

@@ -0,0 +1,156 @@
import {
fetchWithTimeout,
FetchTimeoutError,
getConfiguredTimeout,
} from "@/shared/utils/fetchTimeout";
import { saveCallLog } from "@/lib/usageDb";
import { sanitizeErrorMessage } from "../../utils/error.ts";
interface LogLike {
info?: (tag: string, msg: string, meta?: unknown) => void;
error?: (tag: string, msg: string) => void;
}
interface CredentialsLike {
providerSpecificData?: { baseUrl?: unknown } | null;
baseUrl?: unknown;
apiKey?: unknown;
accessToken?: unknown;
}
/**
* Resolve the video generation endpoint URL from credentials and fallback.
* Handles baseUrl from providerSpecificData or top-level credentials.
*/
function resolveVideoEndpoint(credentials: unknown, fallback: string): string {
const creds = credentials as CredentialsLike | null | undefined;
const psdBaseUrl =
creds?.providerSpecificData?.baseUrl != null &&
typeof creds.providerSpecificData.baseUrl === "string" &&
creds.providerSpecificData.baseUrl.trim()
? creds.providerSpecificData.baseUrl.trim()
: null;
const topLevelBaseUrl =
creds?.baseUrl != null && typeof creds.baseUrl === "string" && creds.baseUrl.trim()
? creds.baseUrl.trim()
: null;
const nodeBaseUrl = psdBaseUrl || topLevelBaseUrl;
let n = nodeBaseUrl;
while (n.endsWith("/")) n = n.slice(0, -1);
if (n.endsWith("/videos/generations")) return n;
return `${n}/videos/generations`;
}
/**
* Fetch the video generation endpoint with timeout and error handling.
*/
async function fetchVideoEndpoint(
url: string,
{ headers, body, log }: { headers: Record<string, string>; body: string; log?: LogLike }
) {
try {
const response = await fetchWithTimeout(url, {
method: "POST",
headers,
body,
timeoutMs: getConfiguredTimeout(),
});
if (!response.ok) {
const errorText = await response.text();
log?.error?.("VIDEO", `Upstream ${response.status} for ${url}: ${errorText}`);
return { success: false, status: response.status, error: errorText };
}
const data = await response.json();
return {
success: true,
data: { created: data.created || Math.floor(Date.now() / 1000), data: data.data || [] },
};
} catch (err) {
const message = err?.message;
const isTimeout = err instanceof FetchTimeoutError || err?.name === "AbortError";
log?.error?.(
"VIDEO",
`${isTimeout ? "Timeout" : "Request error"} for ${url}: ${sanitizeErrorMessage(message || err)}`
);
return {
success: false,
status: isTimeout ? 504 : 502,
error: `Video provider error: ${sanitizeErrorMessage(message || err)}`,
};
}
}
/**
* Handle OpenAI-compatible video generation.
* This handler is dispatched for custom providers with format "openai-video".
*/
export async function handleOpenAIVideoGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}: {
model: string;
provider: string;
providerConfig: { baseUrl: string; authHeader: string };
body: unknown;
credentials: unknown;
log?: LogLike;
}) {
const startTime = Date.now();
const creds = credentials as CredentialsLike | null | undefined;
const apiToken = creds?.apiKey || creds?.accessToken;
const endpoint = resolveVideoEndpoint(credentials, providerConfig.baseUrl);
const headers = {
"Content-Type": "application/json",
...(providerConfig.authHeader === "x-api-key"
? { "x-api-key": String(apiToken) }
: { Authorization: `Bearer ${apiToken}` }),
};
const bodyObj = body as Record<string, unknown>;
const upstreamBody = {
model,
prompt: (bodyObj.prompt ?? "") as string,
...(typeof bodyObj.duration === "number" && { duration: bodyObj.duration }),
};
const logRequestBody = {
model: bodyObj.model,
prompt:
typeof bodyObj.prompt === "string"
? bodyObj.prompt.slice(0, 200)
: String(bodyObj.prompt ?? ""),
duration: bodyObj.duration,
};
log?.info?.("VIDEO", `OpenAI-compatible video generation: ${provider}/${model} -> ${endpoint}`, {
body: logRequestBody,
});
const fetchResult = await fetchVideoEndpoint(endpoint, {
headers,
body: JSON.stringify(upstreamBody),
log,
});
if (!fetchResult.success) {
return { success: false, status: fetchResult.status, error: fetchResult.error };
}
// Save call log for billing/tracking
await saveCallLog({
provider,
model: String(bodyObj.model),
endpoint: "video",
status: fetchResult.status,
durationMs: Date.now() - startTime,
tokensIn: 0,
tokensOut: 0,
requestId: null,
});
return {
success: true,
data: fetchResult.data,
};
}

View File

@@ -0,0 +1,125 @@
export function resolveRunwayPromptImage(body) {
const directCandidates = [
body.promptImage,
body.prompt_image,
body.image,
body.image_url,
body.imageUrl,
body.provider_options?.promptImage,
body.provider_options?.prompt_image,
];
for (const candidate of directCandidates) {
if (typeof candidate === "string" && candidate.trim()) return candidate.trim();
if (candidate && typeof candidate === "object") return candidate;
if (Array.isArray(candidate) && candidate.length > 0) return candidate;
}
const arrayCandidates = [
body.imageUrls,
body.image_urls,
body.provider_options?.imageUrls,
body.provider_options?.image_urls,
];
for (const candidate of arrayCandidates) {
if (Array.isArray(candidate) && candidate.length > 0) return candidate;
}
return null;
}
export function resolveRunwayRatio(body) {
const aspectRatio = typeof body.aspect_ratio === "string" ? body.aspect_ratio : body.aspectRatio;
if (aspectRatio === "1280:720" || aspectRatio === "720:1280") return aspectRatio;
if (aspectRatio === "16:9") return "1280:720";
if (aspectRatio === "9:16") return "720:1280";
const size = typeof body.size === "string" ? body.size : "";
const [widthRaw, heightRaw] = size.split("x");
const width = Number(widthRaw);
const height = Number(heightRaw);
if (Number.isFinite(width) && Number.isFinite(height) && width > 0 && height > 0) {
return width >= height ? "1280:720" : "720:1280";
}
return "1280:720";
}
export function resolveRunwayDuration(body) {
if (Number.isFinite(body.duration)) return clampRunwayDuration(body.duration);
if (Number.isFinite(body.frames) && Number.isFinite(body.fps) && Number(body.fps) > 0) {
return clampRunwayDuration(Number(body.frames) / Number(body.fps));
}
return 5;
}
function clampRunwayDuration(value) {
const duration = Math.round(Number(value));
if (!Number.isFinite(duration)) return 5;
return Math.min(10, Math.max(2, duration));
}
export function resolvePositiveInteger(value, fallback) {
const numeric = Number(value);
if (!Number.isFinite(numeric) || numeric <= 0) return fallback;
return Math.floor(numeric);
}
function extractRunwayOutputUrls(task) {
const rawOutput = Array.isArray(task?.output)
? task.output
: Array.isArray(task?.result)
? task.result
: [];
return rawOutput
.map((entry) => {
if (typeof entry === "string") return entry;
if (!entry || typeof entry !== "object") return null;
return entry.url || entry.uri || entry.videoUrl || entry.video_url || null;
})
.filter((value) => typeof value === "string" && value.length > 0);
}
export function extractRunwayFailureMessage(task) {
const directCandidates = [
task?.failure,
task?.failureReason,
task?.error,
task?.errorMessage,
task?.message,
];
for (const candidate of directCandidates) {
if (typeof candidate === "string" && candidate.trim()) return candidate.trim();
}
if (task?.failure && typeof task.failure === "object") {
const nestedCandidates = [
task.failure.message,
task.failure.reason,
task.failure.error,
task.failure.code,
];
for (const candidate of nestedCandidates) {
if (typeof candidate === "string" && candidate.trim()) return candidate.trim();
}
}
return null;
}
export async function normalizeRunwayVideoResult(task, body) {
const urls = extractRunwayOutputUrls(task);
if (urls.length === 0) {
throw new Error(
`Runway task completed without output URLs: ${JSON.stringify(task).slice(0, 400)}`
);
}
if (body.response_format === "url") return urls.map((url) => ({ url, format: "mp4" }));
const videos = [];
for (const url of urls) {
const response = await fetch(url);
if (!response.ok) throw new Error(`Runway output fetch failed (${response.status})`);
const arrayBuffer = await response.arrayBuffer();
videos.push({ b64_json: Buffer.from(arrayBuffer).toString("base64"), format: "mp4" });
}
return videos;
}

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

View File

@@ -1,9 +1,24 @@
/**
* Re-export from `antigravityProjectPersist.ts` plus a connection-preference helper.
*
* The persistence layer for a runtime-discovered Antigravity projectId lives in
* the sibling file `antigravityProjectPersist.ts` (named by its core function).
* This module adds `preferAntigravityConnectionsWithStoredProject()`, used by the
* quota-strategy engine to give priority to connections whose projectId has
* already been discovered and persisted.
*/
import { persistDiscoveredAntigravityProjectId } from "./antigravityProjectPersist.ts";
export { persistDiscoveredAntigravityProjectId };
/**
* Return only the Antigravity connections that already have a stored projectId.
*
* A connection whose projectId has been discovered and persisted can be used
* immediately; a connection without one would need to go through the Code Assist
* bootstrap first, which is handled by the calling strategy's fallback path.
*/
export function preferAntigravityConnectionsWithStoredProject(
connections: Array<Record<string, unknown>>
): Array<Record<string, unknown>> {

View File

@@ -54,7 +54,7 @@ export function normalizeClaudeAdaptiveThinking<T extends Record<string, unknown
delete nextThinking.budget_tokens;
delete nextThinking.max_tokens;
return { ...record, thinking: nextThinking } as T;
return { ...body, thinking: nextThinking };
}
/**
@@ -84,7 +84,7 @@ export function normalizeClaudeDisabledThinkingEffort<T extends Record<string, u
}
return {
...record,
...body,
output_config: { ...outputConfig, effort: disabledEffortCap },
} as T;
};
}

View File

@@ -3020,7 +3020,8 @@ async function handleRoundRobinCombo({
return new Response(
JSON.stringify({
error: {
message: "Service temporarily unavailable: all targets were skipped by pre-dispatch filters",
message:
"Service temporarily unavailable: all targets were skipped by pre-dispatch filters",
type: "service_unavailable",
code: "ALL_TARGETS_SKIPPED",
},

View File

@@ -615,6 +615,12 @@ export type CompatFilterOptions = {
failOpen?: boolean;
};
function hasHardCapabilityFailure(reasons: string[]): boolean {
return reasons.some((reason) => HARD_COMPAT_REASONS.has(reason));
}
/**
* Summarize a capability-filter exhaustion for a 400-class combo error (#8488).
* Returns null when the empty pool is not attributable to hard requirements.

View File

@@ -221,6 +221,14 @@ const LITE_SCHEMA: EngineConfigField[] = [
label: "Preserve system prompt",
defaultValue: true,
},
{
key: "compressToolResults",
type: "boolean",
label: "Proactively truncate long tool results",
description:
"Truncates tool results over 2,000 characters during Lite compression. Emergency overflow protection may still trim content when the context exceeds the model budget.",
defaultValue: true,
},
];
function validateLiteConfig(config: Record<string, unknown>): EngineValidationResult {
@@ -231,6 +239,7 @@ function validateLiteConfig(config: Record<string, unknown>): EngineValidationRe
) {
errors.push("preserveSystemPrompt must be a boolean");
}
validateBoolean(config, "compressToolResults", errors);
return { valid: errors.length === 0, errors };
}
@@ -256,6 +265,13 @@ export const liteEngine: CompressionEngine = {
const result = applyLiteCompression(adapter.body, {
...options,
preserveSystemPrompt: options?.config?.preserveSystemPrompt !== false,
// buildStepOptions() already merges global config.lite with explicit step.config
// (step wins) into stepConfig, so consume that single effective value instead of
// AND-ing root and step values — an explicit step `true` must override a global `false`.
compressToolResults:
options?.stepConfig?.compressToolResults ??
options?.config?.lite?.compressToolResults ??
true,
});
return adapter.adapted ? { ...result, body: adapter.restore(result.body) } : result;
},

View File

@@ -17,6 +17,7 @@ interface LiteCompressionOptions {
model?: string;
supportsVision?: boolean | null;
preserveSystemPrompt?: boolean;
compressToolResults?: boolean;
}
function trimTrailingHorizontalWhitespace(line: string): string {
@@ -253,9 +254,11 @@ export function applyLiteCompression(
current = r2.body;
if (r2.applied) techniquesApplied.push("system-dedup");
const r3 = compressToolResults(current);
current = r3.body;
if (r3.applied) techniquesApplied.push("tool-compress");
if (options?.compressToolResults !== false) {
const r3 = compressToolResults(current);
current = r3.body;
if (r3.applied) techniquesApplied.push("tool-compress");
}
const r4 = removeRedundantContent(current, options);
current = r4.body;

View File

@@ -14,6 +14,8 @@ export function resolveStepDetailConfig(
config: CompressionConfig | undefined
) {
switch (engine) {
case "lite":
return config?.lite ?? {};
case "headroom":
return config?.headroom ?? {};
case "session-dedup":

View File

@@ -349,6 +349,7 @@ function runCompression(
const result = applyLiteCompression(compressionBody, {
...options,
preserveSystemPrompt: options?.config?.preserveSystemPrompt !== false,
...options?.config?.lite,
});
return adapter.adapted ? { ...result, body: adapter.restore(result.body) } : result;
}

View File

@@ -157,6 +157,12 @@ export interface LiveZoneConfig {
enabled: boolean;
}
/** Lite detail settings for proactive request-time transformations. */
export interface LiteConfig {
/** Truncate tool-result strings over 2,000 characters before provider dispatch. */
compressToolResults: boolean;
}
export interface CompressionPipelineStep {
engine: CompressionEngineId;
intensity?: CavemanIntensity | RtkIntensity;
@@ -218,6 +224,8 @@ export interface CompressionConfig {
languageConfig?: CompressionLanguageConfig;
aggressive?: AggressiveConfig;
ultra?: UltraConfig;
/** Lite proactive transformation detail settings. */
lite?: LiteConfig;
/** Headroom SmartCrusher detail settings (minRows gate). */
headroom?: HeadroomConfig;
/** Session Dedup detail settings (minBlockChars / fuzzy, #8388). */
@@ -395,6 +403,7 @@ export const DEFAULT_COMPRESSION_CONFIG: CompressionConfig = {
ultraEngine: "heuristic",
ultraSlmPrewarm: false,
liveZone: { enabled: false },
lite: { compressToolResults: true },
codexResponsesConfig: { ...DEFAULT_CODEX_RESPONSES_CONFIG },
};

View File

@@ -64,6 +64,8 @@ const REASONING_REPLAY_MODEL_PATTERNS = [
];
const DEEPSEEK_V4_MODEL_PATTERN = /deepseek[-/]v4[-.](flash|pro)/i;
const K3_REASONING_REPLAY_MODEL_PATTERN = /(?:^|\/)(?:kimi-)?k3(?:$|-)/i;
const NATIVE_K27_REASONING_REPLAY_MODEL_PATTERN = /(?:^|\/)kimi-k2\.7-code(?:$|-)/i;
export function isDeepSeekReasoningModel(params: {
provider: string;
@@ -94,6 +96,14 @@ export function requiresReasoningReplay(params: {
if (normalizedInterleavedField === "reasoning_content") return true;
if (normalizedInterleavedField === "reasoning_details") return false;
if (K3_REASONING_REPLAY_MODEL_PATTERN.test(normalizedModel)) return true;
if (
(normalizedProvider === "moonshot" || normalizedProvider === "kimi") &&
NATIVE_K27_REASONING_REPLAY_MODEL_PATTERN.test(normalizedModel)
) {
return true;
}
// DeepSeek legacy reasoner family has an inverse contract: do not replay.
if (/deepseek-reasoner/i.test(normalizedModel) || /deepseek-r1/i.test(normalizedModel)) {
return false;

View File

@@ -0,0 +1,55 @@
type JsonRecord = Record<string, unknown>;
const SERVER_ITEM_ID_PATTERN = /^(rs|fc|resp|msg)_/;
/**
* Applies the persistence-independent policy for replayed Responses input items.
* Stored references can only be resolved by the upstream that created them, so
* they are always removed. Self-contained encrypted reasoning is retained only
* when the selected connection explicitly opts in.
*/
export function applyResponsesInputPolicy(
body: Record<string, unknown>,
preserveEncryptedReasoning = false
): void {
if (Array.isArray(body.input) && body.input.length === 0) {
body.input = [
{
type: "message",
role: "user",
content: [{ type: "input_text", text: "continue" }],
},
];
}
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (typeof item === "string" && SERVER_ITEM_ID_PATTERN.test(item)) {
return false;
}
const record =
item && typeof item === "object" && !Array.isArray(item) ? (item as JsonRecord) : null;
if (!record) return true;
if (record.type === "item_reference") {
return false;
}
if (
record.type === "reasoning" &&
(!preserveEncryptedReasoning ||
typeof record.encrypted_content !== "string" ||
record.encrypted_content.trim().length === 0)
) {
return false;
}
if (typeof record.id === "string" && SERVER_ITEM_ID_PATTERN.test(record.id)) {
delete record.id;
}
return true;
});
}

View File

@@ -48,6 +48,7 @@ import { refreshGoogleToken } from "./tokenRefresh/providers/google.ts";
import { ensureAntigravityProjectAssigned } from "./antigravityProjectBootstrap.ts";
import { persistDiscoveredAntigravityProjectId } from "./antigravityProjectPersist.ts";
import { refreshCodexToken } from "./tokenRefresh/providers/codex.ts";
import { refreshOpenferenceToken } from "./tokenRefresh/providers/openference.ts";
import { refreshKiroToken } from "./tokenRefresh/providers/kiro.ts";
import { refreshQoderToken } from "./tokenRefresh/providers/qoder.ts";
import { refreshGitHubToken } from "./tokenRefresh/providers/github.ts";
@@ -62,6 +63,7 @@ export {
refreshClaudeOAuthToken,
refreshGoogleToken,
refreshCodexToken,
refreshOpenferenceToken,
refreshKiroToken,
refreshQoderToken,
refreshGitHubToken,
@@ -339,10 +341,7 @@ async function _getAccessTokenInternal(provider, credentials, log, proxyConfig:
!(credentials.projectId || credentials.providerSpecificData?.projectId)
) {
try {
const discovered = await ensureAntigravityProjectAssigned(
result.accessToken,
fetch
);
const discovered = await ensureAntigravityProjectAssigned(result.accessToken, fetch);
if (discovered) {
result.projectId = discovered;
result.providerSpecificData = {
@@ -362,7 +361,8 @@ async function _getAccessTokenInternal(provider, credentials, log, proxyConfig:
});
}
} catch (discoveryError) {
const msg = discoveryError instanceof Error ? discoveryError.message : String(discoveryError);
const msg =
discoveryError instanceof Error ? discoveryError.message : String(discoveryError);
log?.warn?.("TOKEN", `Antigravity projectId discovery failed: ${msg}`);
}
}
@@ -376,6 +376,9 @@ async function _getAccessTokenInternal(provider, credentials, log, proxyConfig:
case "codex":
return await refreshCodexToken(credentials.refreshToken, log, proxyConfig);
case "openference":
return await refreshOpenferenceToken(credentials.refreshToken, log, proxyConfig);
case "qoder":
return await refreshQoderToken(credentials.refreshToken, log, proxyConfig);
@@ -439,6 +442,7 @@ export function supportsTokenRefresh(provider) {
"agy",
"claude",
"codex",
"openference",
"qoder",
"github",
"kiro",

View File

@@ -0,0 +1,92 @@
// @ts-nocheck
import { OAUTH_ENDPOINTS } from "../../../config/constants.ts";
import { runWithProxyContext } from "../../../utils/proxyFetch.ts";
import { buildFormParams } from "../shared.ts";
/**
* Specialized refresh for Openference OAuth tokens.
* Openference uses rotating (one-time-use) oar_* refresh tokens.
*/
export async function refreshOpenferenceToken(refreshToken, log, proxyConfig: unknown = null) {
try {
const response = await runWithProxyContext(proxyConfig, () =>
fetch(OAUTH_ENDPOINTS.openference.token, {
method: "POST",
headers: {
"Content-Type": "application/x-www-form-urlencoded",
Accept: "application/json",
},
body: buildFormParams({
grant_type: "refresh_token",
refresh_token: refreshToken,
client_id: OAUTH_ENDPOINTS.openference.clientId,
}),
})
);
if (!response.ok) {
const errorText = await response.text();
let errorCode = null;
try {
const parsed = JSON.parse(errorText);
errorCode =
parsed?.error?.code || (typeof parsed?.error === "string" ? parsed.error : null);
} catch {
// not JSON, ignore
}
if (
errorCode === "invalid_grant" ||
errorCode === "token_expired" ||
errorCode === "invalid_token"
) {
log?.error?.(
"TOKEN_REFRESH",
"Openference refresh token already used or invalid. Re-authentication required.",
{
status: response.status,
errorCode,
}
);
return { error: "unrecoverable_refresh_error", code: errorCode };
}
if (response.status === 401) {
const code = errorCode || "unauthorized";
log?.error?.(
"TOKEN_REFRESH",
"Openference OAuth token endpoint returned 401. Re-authentication required.",
{
status: response.status,
errorCode: code,
}
);
return { error: "unrecoverable_refresh_error", code };
}
log?.error?.("TOKEN_REFRESH", "Failed to refresh Openference token", {
status: response.status,
error: errorText,
});
return null;
}
const tokens = await response.json();
log?.info?.("TOKEN_REFRESH", "Successfully refreshed Openference token", {
hasNewAccessToken: !!tokens.access_token,
hasNewRefreshToken: !!tokens.refresh_token,
expiresIn: tokens.expires_in,
});
return {
accessToken: tokens.access_token,
refreshToken: tokens.refresh_token || refreshToken,
expiresIn: tokens.expires_in,
};
} catch (error) {
log?.error?.("TOKEN_REFRESH", `Network error refreshing Openference token: ${error.message}`);
return null;
}
}

View File

@@ -37,6 +37,102 @@ async function getPath() {
return _path || null;
}
type UsageRecord = Record<string, unknown>;
function usageRecord(value: unknown): UsageRecord {
return typeof value === "object" && value !== null && !Array.isArray(value)
? (value as UsageRecord)
: {};
}
function usageNumber(value: unknown): number | undefined {
return typeof value === "number" && Number.isFinite(value) ? value : undefined;
}
function usageDetails(record: UsageRecord, ...keys: string[]): UsageRecord {
for (const key of keys) {
const value = usageRecord(record[key]);
if (Object.keys(value).length > 0) return value;
}
return {};
}
/** Normalize Chat Completions and Responses usage into the Responses API shape. */
function normalizeResponsesUsage(previous: unknown, raw: unknown): UsageRecord | null {
const source = usageRecord(raw);
if (Object.keys(source).length === 0) return usageRecord(previous);
const before = usageRecord(previous);
const beforeInputDetails = usageDetails(before, "input_tokens_details", "prompt_tokens_details");
const beforeOutputDetails = usageDetails(
before,
"output_tokens_details",
"completion_tokens_details"
);
const inputDetails = usageDetails(
source,
"input_tokens_details",
"prompt_tokens_details",
"inputTokenDetails",
"input_token_details"
);
const outputDetails = usageDetails(
source,
"output_tokens_details",
"completion_tokens_details",
"outputTokenDetails",
"output_token_details",
"reasoningTokenDetails",
"reasoning_token_details"
);
const inputTokens =
usageNumber(source.input_tokens) ??
usageNumber(source.prompt_tokens) ??
usageNumber(source.inputTokens) ??
usageNumber(source.promptTokens) ??
usageNumber(before.input_tokens) ??
usageNumber(before.prompt_tokens) ??
0;
const cachedTokens =
usageNumber(source.cache_read_input_tokens) ??
usageNumber(source.cached_input_tokens) ??
usageNumber(source.cachedInputTokens) ??
usageNumber(source.cached_tokens) ??
usageNumber(inputDetails.cached_tokens) ??
usageNumber(inputDetails.cachedTokens) ??
usageNumber(inputDetails.cacheReadTokens) ??
usageNumber(beforeInputDetails.cached_tokens) ??
0;
const outputTokens =
usageNumber(source.output_tokens) ??
usageNumber(source.completion_tokens) ??
usageNumber(source.outputTokens) ??
usageNumber(source.completionTokens) ??
usageNumber(before.output_tokens) ??
usageNumber(before.completion_tokens) ??
0;
const reasoningTokens =
usageNumber(source.reasoning_tokens) ??
usageNumber(source.reasoningTokens) ??
usageNumber(outputDetails.reasoning_tokens) ??
usageNumber(outputDetails.reasoningTokens) ??
usageNumber(beforeOutputDetails.reasoning_tokens) ??
0;
const totalTokens =
usageNumber(source.total_tokens) ??
usageNumber(source.totalTokens) ??
inputTokens + outputTokens;
return {
input_tokens: inputTokens,
input_tokens_details: { cached_tokens: cachedTokens },
output_tokens: outputTokens,
output_tokens_details: { reasoning_tokens: reasoningTokens },
total_tokens: totalTokens,
};
}
// Create log directory for responses (Node.js only)
export function createResponsesLogger(model, logsDir = null) {
// Skip logging in worker environment (no fs)
@@ -477,10 +573,11 @@ export function createResponsesApiTransformStream(
continue;
}
if (parsed.usage) {
state.usage = normalizeResponsesUsage(state.usage, parsed.usage);
}
if (!parsed.choices?.length) {
if (parsed.usage) {
state.usage = parsed.usage;
}
// #6906: trailing usage-only chunk after finish_reason already deferred
// completion — send it now with the usage just captured above.
if (state.awaitingTrailingUsage && !state.completedSent) {

View File

@@ -84,6 +84,10 @@ export function hasValidContent(msg: ClaudeMessage): boolean {
return msg.content.some(
(block) =>
(block.type === "text" && block.text?.trim()) ||
(block.type === "thinking" && block.thinking?.trim()) ||
(block.type === "redacted_thinking" &&
typeof block.data === "string" &&
block.data.trim()) ||
block.type === "tool_use" ||
block.type === "tool_result" ||
// #7777: media-only user turns are real content — dropping them

View File

@@ -2,10 +2,16 @@
* Convert OpenAI Responses API format to standard chat completions format.
* Delegates to the canonical translator to avoid logic duplication.
*/
import { requiresReasoningReplay } from "../../services/reasoningCache.ts";
import { openaiResponsesToOpenAIRequest } from "../request/openai-responses.ts";
import { toRecord } from "../request/openai-responses/helpers.ts";
export function convertResponsesApiFormat(body, credentials = null, provider = null) {
export function convertResponsesApiFormat(
body: Record<string, unknown>,
credentials: unknown = null,
provider: unknown = null,
model: unknown = null
): Record<string, unknown> {
const bodyModel = toRecord(body).model;
const requestedModel =
typeof bodyModel === "string" && bodyModel.trim().length > 0
@@ -13,5 +19,25 @@ export function convertResponsesApiFormat(body, credentials = null, provider = n
? bodyModel
: `${provider}/${bodyModel}`
: provider;
return openaiResponsesToOpenAIRequest(requestedModel, body, null, credentials);
const credentialRecord =
credentials && typeof credentials === "object" && !Array.isArray(credentials)
? (credentials as Record<string, unknown>)
: {};
const translationCredentials = requiresReasoningReplay({
provider: String(provider ?? ""),
model: String(model ?? ""),
allowLegacyFallback: false,
})
? { ...credentialRecord, _preserveReasoningContent: true }
: credentials;
const converted = openaiResponsesToOpenAIRequest(
requestedModel,
body,
null,
translationCredentials
);
if (!converted || typeof converted !== "object" || Array.isArray(converted)) {
throw new TypeError("Responses request conversion must produce an object");
}
return converted as Record<string, unknown>;
}

View File

@@ -13,7 +13,6 @@ import {
providerHonorsOpenAIFormatCacheControl,
resolveConnectionCacheOverride,
} from "../utils/cacheControlPolicy.ts";
import { requiresAuthenticReasoningContent } from "../utils/reasoningContentInjector.ts";
import { isInternalReasoningPlaceholder } from "../utils/reasoningPlaceholder.ts";
import {
coerceToolSchemas,
@@ -162,7 +161,11 @@ function isReasoningOnlyReplayTarget(provider: unknown, model: unknown): boolean
/(^|\/)deepseek/i.test(normalizedModel) ||
normalizedProvider === "xiaomi-mimo" ||
/(^|\/)mimo/i.test(normalizedModel) ||
requiresAuthenticReasoningContent(normalizedProvider, normalizedModel)
requiresReasoningReplay({
provider: normalizedProvider,
model: normalizedModel,
allowLegacyFallback: false,
})
);
}
@@ -233,6 +236,17 @@ export function translateRequest(
const connectionCacheOverride = resolveConnectionCacheOverride(
(credentials as { providerSpecificData?: unknown } | null)?.providerSpecificData
);
const normalizedProvider = String(provider ?? "");
const normalizedModel = String(model ?? "");
const isKimiCoding =
normalizedProvider === "kimi-coding" || normalizedProvider === "kimi-coding-apikey";
const requiresExplicitReasoningReplay = requiresReasoningReplay({
provider: normalizedProvider,
model: normalizedModel,
allowLegacyFallback: false,
});
const preserveResponsesReasoning =
sourceFormat === FORMATS.OPENAI_RESPONSES && requiresExplicitReasoningReplay;
// Phase 2: Apply thinking budget control before normalization
result = applyThinkingBudget(result);
@@ -318,12 +332,16 @@ export function translateRequest(
options?.preserveCacheControl === true &&
providerHonorsOpenAIFormatCacheControl(provider, connectionCacheOverride);
const step1Credentials =
options?.copilotClient || hasTargetHint || preserveCacheControl
options?.copilotClient ||
hasTargetHint ||
preserveCacheControl ||
preserveResponsesReasoning
? {
...(credentials && typeof credentials === "object" ? credentials : {}),
...(options?.copilotClient ? { _copilotClient: true } : {}),
...(hasTargetHint ? { _targetFormat: targetFormat } : {}),
...(preserveCacheControl ? { _preserveCacheControl: true } : {}),
...(preserveResponsesReasoning ? { _preserveReasoningContent: true } : {}),
}
: credentials;
result = toOpenAI(model, result, stream, step1Credentials);
@@ -352,7 +370,27 @@ export function translateRequest(
...(hasProvider ? { _provider: provider } : {}),
}
: credentials;
result = fromOpenAI(model, result, stream, translationCredentials);
// #9780 — carry the Responses namespace identity map across the pivot.
// Target translators return a brand-new object (buildKiroPayload et
// al.), dropping the non-enumerable property step 1 attached; the
// #7936 seam then gets null and namespace sub-tool calls come back
// flattened, which Codex rejects with `unsupported call: <name>`.
const identityMap = (result as Record<string, unknown>)._namespaceToolIdentityMap;
const translated = fromOpenAI(model, result, stream, translationCredentials);
if (
identityMap instanceof Map &&
translated &&
typeof translated === "object" &&
!((translated as Record<string, unknown>)._namespaceToolIdentityMap instanceof Map)
) {
Object.defineProperty(translated, "_namespaceToolIdentityMap", {
value: identityMap,
enumerable: false,
configurable: true,
writable: true,
});
}
result = translated;
}
}
}
@@ -361,14 +399,6 @@ export function translateRequest(
// Resolve reasoning-replay status up-front: it gates both the reasoning_content
// strip in filterToOpenAIFormat below (#4849 must NOT strip client reasoning for
// replay providers) and the cache re-injection further down.
const normalizedProvider = String(provider ?? "");
const normalizedModel = String(model ?? "");
const isKimiCoding =
normalizedProvider === "kimi-coding" || normalizedProvider === "kimi-coding-apikey";
const requiresAuthenticReasoning = requiresAuthenticReasoningContent(
normalizedProvider,
normalizedModel
);
const resolvedCapabilities = getResolvedModelCapabilities({
provider: normalizedProvider,
model: normalizedModel,
@@ -377,7 +407,10 @@ export function translateRequest(
provider: normalizedProvider,
model: normalizedModel,
thinkingEnabled: hasThinkingConfig(result),
supportsReasoning: supportsReasoning({ provider: normalizedProvider, model: normalizedModel }),
supportsReasoning: supportsReasoning({
provider: normalizedProvider,
model: normalizedModel,
}),
interleavedField: resolvedCapabilities?.interleavedField ?? null,
});
@@ -450,7 +483,7 @@ export function translateRequest(
if (
targetFormat === FORMATS.OPENAI &&
!requiresAuthenticReasoning &&
!requiresExplicitReasoningReplay &&
result.messages &&
Array.isArray(result.messages)
) {
@@ -475,7 +508,7 @@ export function translateRequest(
// isReasoner / normalizedProvider / normalizedModel / resolvedCapabilities were
// resolved up-front (before the OpenAI-format filter) so the #4849 reasoning strip
// could honor reasoning-replay providers.
if (isReasoner && !isKimiCoding && result.messages && Array.isArray(result.messages)) {
if (isReasoner && result.messages && Array.isArray(result.messages)) {
const canReplayReasoningOnly = isReasoningOnlyReplayTarget(normalizedProvider, normalizedModel);
for (const [messageIndex, msg] of result.messages.entries()) {
@@ -524,29 +557,51 @@ export function translateRequest(
// Has tool_use blocks but no thinking block yet.
// Reasoning models (Kimi K2, etc.) require a thinking block before tool_use
// on multi-turn or they regenerate the same tool call infinitely.
const hasThinkingBlock = msg.content.some(
const thinkingBlock = msg.content.find(
(b) => b?.type === "thinking" || b?.type === "redacted_thinking"
);
if (hasThinkingBlock) continue;
const hasNonEmptyClientThinking =
thinkingBlock?.type === "thinking" &&
typeof thinkingBlock.thinking === "string" &&
thinkingBlock.thinking.trim().length > 0;
if (thinkingBlock && (!isKimiCoding || hasNonEmptyClientThinking)) continue;
const toolUseBlocks = msg.content.filter((b) => b?.type === "tool_use");
const firstToolUseId = toolUseBlocks[0]?.id;
const firstToolUseIdx = msg.content.findIndex((b) => b?.type === "tool_use");
// Try reasoning cache first
// Client reasoning wins above. Otherwise try authentic replay before
// retaining Kimi Code's empty protocol marker as the final fallback.
if (firstToolUseId) {
const cached = lookupReasoning(firstToolUseId);
if (cached) {
msg.content.splice(firstToolUseIdx, 0, {
type: "thinking",
thinking: cached,
});
if (thinkingBlock) {
thinkingBlock.type = "thinking";
thinkingBlock.thinking = cached;
delete thinkingBlock.data;
delete thinkingBlock.signature;
} else {
msg.content.splice(firstToolUseIdx, 0, {
type: "thinking",
thinking: cached,
});
}
recordReplay();
continue;
}
}
if (requiresAuthenticReasoning) continue;
// Fallback: inject placeholder (must be non-empty for kimi-coding)
if (isKimiCoding) {
if (thinkingBlock) {
thinkingBlock.type = "thinking";
thinkingBlock.thinking = "";
delete thinkingBlock.data;
delete thinkingBlock.signature;
} else {
msg.content.splice(firstToolUseIdx, 0, { type: "thinking", thinking: "" });
}
continue;
}
if (requiresExplicitReasoningReplay) continue;
msg.content.splice(firstToolUseIdx, 0, {
type: "thinking",
thinking: NON_ANTHROPIC_THINKING_PLACEHOLDER,
@@ -570,7 +625,7 @@ export function translateRequest(
const cacheKey = hasToolCalls
? msg.tool_calls[0]?.id
: getAssistantMessageCacheKey(result, 0);
: getAssistantMessageCacheKey(result, messageIndex);
if (cacheKey) {
const cached = lookupReasoning(cacheKey);
if (cached) {
@@ -583,7 +638,7 @@ export function translateRequest(
// Native Moonshot K3/K2.7 accepts only the real prior reasoning. If it
// was not supplied and the cache missed, leave it absent so upstream can
// enforce its contract instead of corrupting history with a placeholder.
if (requiresAuthenticReasoning) {
if (requiresExplicitReasoningReplay) {
if (msg.reasoning_content === "") delete msg.reasoning_content;
continue;
}
@@ -596,7 +651,7 @@ export function translateRequest(
// deepseek-v4-flash accepts an ABSENT reasoning_content field (the 400 is
// specific to empty-string, and even that is endpoint-dependent). Omit
// the field instead; providers that genuinely enforce the contract
// (kimi-coding, moonshot authentic-reasoning) have their own paths above.
// (kimi-coding, moonshot reasoning replay) have their own paths above.
if ((hasToolCalls || shouldReplayReasoningOnly) && !msg.reasoning_content) {
if (requiresReasoningContentPresence(normalizedProvider, normalizedModel)) {
msg.reasoning_content = NON_ANTHROPIC_THINKING_PLACEHOLDER;
@@ -735,6 +790,7 @@ export function initState(sourceFormat) {
inThinking: false,
parseTextualReasoningTags: false,
funcArgsBuf: {},
funcArgsEscapeState: {},
funcNames: {},
funcCallIds: {},
funcArgsDone: {},

View File

@@ -63,7 +63,12 @@ const STRIP_RULES: StripRule[] = [
// MoonshotAI/kimi-cli#1124), and by upstream decolua/9router#2460. Scoped to
// OmniRoute's actual volcengine Kimi id (not a broad /kimi/i regex) so it
// never clamps an unrelated future Kimi listing whose Ark cap may differ.
{ provider: "volcengine", match: /^kimi-k2-5-260127$/, maxOutputCap: 32768, clampToModelMaxOutput: true },
{
provider: "volcengine",
match: /^kimi-k2-5-260127$/,
maxOutputCap: 32768,
clampToModelMaxOutput: true,
},
// #7364: Z.AI's glm-4.6v vision endpoint enforces a 32768 max_tokens ceiling
// server-side and 400s when a client sends a larger explicit max_tokens (e.g. a
// client defaulting to 65536). Scoped to both wire paths that can reach this
@@ -75,6 +80,19 @@ const STRIP_RULES: StripRule[] = [
// glmProvider.ts, maxOutputTokens: 32768, so clampToModelMaxOutput suffices).
{ provider: "zai", match: /^glm-4\.6v$/i, maxOutputCap: 32768 },
{ provider: "glm", match: /^glm-4\.6v$/i, clampToModelMaxOutput: true },
// Azure gpt-4o-mini deployments cap completion tokens at 16384 and 400 on
// anything larger: "max_tokens is too large: 32000. This model supports at
// most 16384 completion tokens". OmniRoute's own tool-calling floor
// (DEFAULT_MIN_TOKENS = 32000, applied by adjustMaxTokens) raises even a tiny
// explicit max_tokens to 32000 whenever tools are present, so every agentic
// client trips this on its first turn. PROVIDER_MAX_TOKENS is not the right
// lever here: it is provider-wide, and the same Azure resource also serves
// GPT-5 deployments whose ceiling is far higher. Azure deployment names are
// operator-chosen, hence a prefix match rather than an exact id, and the
// models are passthrough (no catalog maxOutputTokens for clampToModelMaxOutput
// to read), hence the fixed cap.
{ provider: "azure-openai", match: /^gpt-4o-mini/i, maxOutputCap: 16384 },
{ provider: "azure-ai", match: /^gpt-4o-mini/i, maxOutputCap: 16384 },
];
function matches(rule: StripRule, model: string): boolean {

View File

@@ -73,6 +73,19 @@ function toolOutputContentToString(output: unknown): string {
return parts.join("\n");
}
function getReasoningSummaryText(item: JsonRecord): string {
if (!Array.isArray(item.summary)) return "";
return item.summary
.map((part) => toString(toRecord(part).text))
.filter((text) => text.length > 0)
.join("\n\n");
}
function appendReasoningContent(current: unknown, next: string): string {
const existing = typeof current === "string" ? current : "";
return existing ? `${existing}\n\n${next}` : next;
}
/**
* Convert OpenAI Responses API request to OpenAI Chat Completions format
*/
@@ -83,13 +96,13 @@ export function openaiResponsesToOpenAIRequest(
credentials: unknown
): unknown {
void stream;
void credentials;
const collapseToPlainString = requiresPlainStringContent(extractProviderHint(model));
const root = toRecord(body);
if (root.input === undefined) return body;
const credentialRecord = toRecord(credentials);
const storeEnabled = isOpenAIResponsesStoreEnabled(credentialRecord.providerSpecificData);
const preserveReasoningContent = credentialRecord._preserveReasoningContent === true;
const rawInputItems = normalizeResponsesInputForChat(root.input);
// Tools may be declared at the Responses top level or in one or more
@@ -204,6 +217,7 @@ export function openaiResponsesToOpenAIRequest(
// Group items by conversation turn
let currentAssistantMsg: JsonRecord | null = null;
let pendingToolResults: JsonRecord[] = [];
let pendingReasoningContent = "";
// Upstream providers reject messages:[] with "400: at least one message is required".
// When the client sends input:[] (empty), inject a placeholder user message — mirrors
@@ -220,11 +234,20 @@ export function openaiResponsesToOpenAIRequest(
const itemType = toString(item.type) || (item.role ? "message" : "");
if (itemType === "message") {
const role = toString(item.role);
// Flush pending assistant message with tool calls
if (currentAssistantMsg) {
messages.push(currentAssistantMsg);
currentAssistantMsg = null;
}
if (role !== "assistant" && pendingReasoningContent) {
messages.push({
role: "assistant",
content: null,
reasoning_content: pendingReasoningContent,
});
pendingReasoningContent = "";
}
// Flush pending tool results
if (pendingToolResults.length > 0) {
@@ -269,7 +292,12 @@ export function openaiResponsesToOpenAIRequest(
})
: item.content;
messages.push({ role: toString(item.role), content });
const message: JsonRecord = { role, content };
if (role === "assistant" && pendingReasoningContent) {
message.reasoning_content = pendingReasoningContent;
pendingReasoningContent = "";
}
messages.push(message);
continue;
}
@@ -294,6 +322,10 @@ export function openaiResponsesToOpenAIRequest(
content: null,
tool_calls: [],
};
if (pendingReasoningContent) {
currentAssistantMsg.reasoning_content = pendingReasoningContent;
pendingReasoningContent = "";
}
}
const toolCalls = Array.isArray(currentAssistantMsg.tool_calls)
@@ -353,6 +385,10 @@ export function openaiResponsesToOpenAIRequest(
content: null,
tool_calls: [],
};
if (pendingReasoningContent) {
currentAssistantMsg.reasoning_content = pendingReasoningContent;
pendingReasoningContent = "";
}
}
const toolCalls = Array.isArray(currentAssistantMsg.tool_calls)
? currentAssistantMsg.tool_calls
@@ -401,7 +437,21 @@ export function openaiResponsesToOpenAIRequest(
}
if (itemType === "reasoning") {
// Skip reasoning items - they are display-only metadata
// Responses reasoning summaries are normally display metadata. Preserve them only
// when the routed upstream explicitly requires prior reasoning to continue a turn.
if (preserveReasoningContent) {
const reasoning = getReasoningSummaryText(item);
if (reasoning) {
if (currentAssistantMsg) {
currentAssistantMsg.reasoning_content = appendReasoningContent(
currentAssistantMsg.reasoning_content,
reasoning
);
} else {
pendingReasoningContent = appendReasoningContent(pendingReasoningContent, reasoning);
}
}
}
continue;
}
@@ -430,6 +480,13 @@ export function openaiResponsesToOpenAIRequest(
if (currentAssistantMsg) {
messages.push(currentAssistantMsg);
}
if (pendingReasoningContent) {
messages.push({
role: "assistant",
content: null,
reasoning_content: pendingReasoningContent,
});
}
if (pendingToolResults.length > 0) {
for (const toolResult of pendingToolResults) {
messages.push(toolResult);
@@ -752,8 +809,19 @@ export function openaiResponsesToOpenAIRequest(
delete result.prompt_cache_retention;
if (namespaceToolIdentityMap.size > 0) {
// chatCore extracts and deletes this transient side channel before dispatch.
// chatCore extracts and deletes these transient side channels before dispatch.
// Non-enumerability keeps internal request metadata off the upstream wire.
//
// Two properties on purpose (#9780): `_toolNameMap` is also the alias
// channel for openai-to-claude/gemini, which overwrite it on a pivot, so
// the identity map needs a name of its own. `_toolNameMap` stays populated
// for the existing consumers (executors/base.ts, cliproxyapi, antigravity).
Object.defineProperty(result, "_namespaceToolIdentityMap", {
value: namespaceToolIdentityMap,
enumerable: false,
configurable: true,
writable: true,
});
Object.defineProperty(result, "_toolNameMap", {
value: namespaceToolIdentityMap,
enumerable: false,

View File

@@ -31,6 +31,19 @@ import {
// normalizeUpstreamFailure is re-exported for external importers (tests).
export { normalizeUpstreamFailure } from "./openai-responses/pureHelpers.ts";
/** Carries escapeJsonStringValues's scan state (whether we're inside a JSON
* string, and whether the fragment ended mid-escape-sequence) across calls
* for the SAME tool call — see escapeJsonStringValues's own doc comment for
* why this must persist across chunks rather than reset per call. */
interface JsonStringEscapeState {
inString: boolean;
pendingEscape: boolean;
}
function createJsonStringEscapeState(): JsonStringEscapeState {
return { inString: false, pendingEscape: false };
}
/**
* Escape control characters (newlines, tabs, carriage returns) that appear
* inside JSON string values, ensuring the resulting string is valid JSON.
@@ -38,18 +51,42 @@ export { normalizeUpstreamFailure } from "./openai-responses/pureHelpers.ts";
* newlines (0x0A) instead of \n escapes inside tool call argument JSON.
* Only escapes characters inside string contexts to avoid double-escaping
* already-proper JSON or corrupting structural newlines.
*
* `arguments` deltas arrive as arbitrary fragments of one continuous JSON
* string (OpenAI's Chat Completions streaming contract only guarantees each
* `tool_calls[].function.arguments` delta is the next slice, not that it
* starts/ends on a quote or escape boundary) — a large multi-line argument
* value routinely gets split mid-string. `escapeState` must therefore be the
* SAME object passed in on every call for a given tool call index, not a
* fresh `{inString: false}` each time: resetting per call made the
* in-string/out-of-string decision (and therefore whether a raw newline
* gets escaped) depend on where a chunk boundary happened to fall, which
* produced a real, reported bug — a single reassembled arguments string
* with a mix of real newlines and literal two-character `\n` sequences,
* breaking generated code (e.g. Python) that embeds multi-line content.
*/
function escapeJsonStringValues(json: string): string {
function escapeJsonStringValues(json: string, escapeState: JsonStringEscapeState): string {
let result = "";
let inString = false;
let { inString, pendingEscape } = escapeState;
for (let i = 0; i < json.length; i++) {
const ch = json[i];
// Inside a string, skip over escape sequences
// This char is the one immediately following a backslash from a
// previous iteration (possibly in a prior fragment) — it's already
// "consumed" by that escape sequence, pass it through untouched.
if (pendingEscape) {
result += ch;
pendingEscape = false;
continue;
}
// Inside a string, an unescaped backslash starts an escape sequence —
// the char AFTER it (next iteration, possibly in the next fragment)
// must not be reinterpreted as a quote/control-char in its own right.
if (inString && ch === "\\") {
result += ch + (json[i + 1] ?? "");
i++;
result += ch;
pendingEscape = true;
continue;
}
@@ -69,6 +106,8 @@ function escapeJsonStringValues(json: string): string {
result += ch;
}
escapeState.inString = inString;
escapeState.pendingEscape = pendingEscape;
return result;
}
@@ -451,11 +490,22 @@ function closeMessage(state, emit, idx) {
}
}
// Tool calls sit after reasoning (if any) AND after a text message (if one was
// actually emitted this turn) — a model commonly emits a short preamble before
// calling a tool (e.g. "Kör nu, på riktigt — apply_patch..."), and that message
// claims the same reasoningIndex+1 slot the old per-call math (`reasoningIndex
// + 1 + tcIdx`) assumed was free for tcIdx=0. Not accounting for the message
// item collided the tool call's added/delta/done events onto the same
// output_index as the just-closed message, which a client keying per-item
// state by output_index can silently drop (live incident 2026-08-08).
function toolCallOutputIndexBase(state) {
const msgIdx = state.reasoningId ? normalizeOutputIndex(state.reasoningIndex) + 1 : 0;
return state.msgItemAdded[msgIdx] ? msgIdx + 1 : msgIdx;
}
function emitToolCall(state, emit, tc) {
const tcIdx = tc.index ?? 0;
const outputIndex = state.reasoningId
? normalizeOutputIndex(state.reasoningIndex) + 1 + normalizeOutputIndex(tcIdx)
: normalizeOutputIndex(tcIdx);
const outputIndex = toolCallOutputIndexBase(state) + normalizeOutputIndex(tcIdx);
const newCallId = tc.id;
const funcName = tc.function?.name;
@@ -471,6 +521,7 @@ function emitToolCall(state, emit, tc) {
delete state.funcArgsDone[tcIdx];
delete state.funcItemAdded[tcIdx];
delete state.funcItemDone[tcIdx];
delete state.funcArgsEscapeState?.[tcIdx];
}
if (funcName) state.funcNames[tcIdx] = funcName;
@@ -517,7 +568,14 @@ function emitToolCall(state, emit, tc) {
if (tc.function?.arguments) {
const refCallId = state.funcCallIds[tcIdx] || newCallId;
const existingArgs = state.funcArgsBuf[tcIdx] || "";
const sanitized = escapeJsonStringValues(tc.function.arguments);
if (!state.funcArgsEscapeState) state.funcArgsEscapeState = {};
if (!state.funcArgsEscapeState[tcIdx]) {
state.funcArgsEscapeState[tcIdx] = createJsonStringEscapeState();
}
const sanitized = escapeJsonStringValues(
tc.function.arguments,
state.funcArgsEscapeState[tcIdx]
);
const nextArgs = appendToolCallArgumentDelta(existingArgs, sanitized);
const emittedDelta = nextArgs.slice(existingArgs.length);
state.funcArgsBuf[tcIdx] = nextArgs;
@@ -536,9 +594,7 @@ function emitToolCall(state, emit, tc) {
function closeToolCall(state, emit, idx, recordAsCompleted = true) {
const callId = state.funcCallIds[idx];
if (callId && !state.funcItemDone[idx]) {
const normalizedIndex = state.reasoningId
? normalizeOutputIndex(state.reasoningIndex) + 1 + normalizeOutputIndex(idx)
: normalizeOutputIndex(idx);
const normalizedIndex = toolCallOutputIndexBase(state) + normalizeOutputIndex(idx);
const args = state.funcArgsBuf[idx] || "{}";
const toolName = state.funcNames[idx] || "";
const isCustomTool =

View File

@@ -356,18 +356,24 @@ export function toArgumentsString(value: unknown): string {
}
}
export interface SerializeToolOptions {
/** Hardened mode for thinking/reasoning models: repeat the instruction
* both before AND after the tool list, use a more distinctive tag format,
* and explicitly tell the model not to claim tools are unavailable. */
hardened?: boolean;
}
/**
* Serialize an OpenAI `tools` array into a system-prompt block that instructs the
* web UI model how to invoke a tool (emit a `<tool>{...}</tool>` block). Returns an
* empty string when there are no usable tools.
*
* Each invocation generates a per-request nonce that is embedded in the tool format
* instructions. The parser (parseToolCallsFromText) requires this nonce in the model's
* `<tool>` JSON to distinguish legitimate tool calls from bare JSON, code-fenced JSON,
* or copy-attacked envelopes (#9343).
*/
export function serializeToolsToPrompt(tools: unknown): string {
if (!Array.isArray(tools) || tools.length === 0) return "";
// ── Tool list rendering (shared between standard and hardened) ─────────────────
const nonce = getToolNonce(tools);
if (!nonce) return "";
function renderToolList(tools: OpenAIToolDef[]): string[] {
const lines: string[] = [];
for (const t of tools) {
for (const t of tools as OpenAIToolDef[]) {
const fn = t?.function;
if (!fn?.name) continue;
const desc = typeof fn.description === "string" && fn.description ? fn.description : "";
@@ -381,52 +387,19 @@ function renderToolList(tools: OpenAIToolDef[]): string[] {
`- ${fn.name}${desc ? `: ${desc}` : ""}${params ? `\n parameters: ${params}` : ""}`
);
}
return lines;
}
/**
* Serialize an OpenAI `tools` array into a system-prompt block that instructs the
* web UI model how to invoke a tool (emit a `<tool>{...}</tool>` block). Returns an
* empty string when there are no usable tools.
*
* When `options.hardened` is set (intended for thinking/reasoning models), the
* contract is more emphatic: the `<tool>` format example is shown before the tool
* list, an explicit "IMPORTANT" directive is appended after the list, and the
* model is told not to claim tools are unavailable.
*/
export function serializeToolsToPrompt(tools: unknown, options?: SerializeToolOptions): string {
if (!Array.isArray(tools) || tools.length === 0) return "";
// #9343: the per-request nonce is mandatory in BOTH modes — the parser rejects
// any <tool> JSON without the matching `_nonce` binding.
const nonce = getToolNonce(tools);
if (!nonce) return "";
const defs = tools as OpenAIToolDef[];
const lines = renderToolList(defs);
if (lines.length === 0) return "";
if (options?.hardened) {
return [
"You have access to the following tools and you MUST use them when appropriate.",
"",
`<tool>{"name": "<tool_name>", "arguments": { ... }, "_nonce": "${nonce}"}</tool>`,
`Every tool call MUST include the secret binding "_nonce": "${nonce}" exactly as shown.`,
"",
"Available tools:",
...lines,
"",
"IMPORTANT: You CAN and MUST use these tools. Do NOT say you cannot use tools or that",
"tools are unavailable — you have them and they are ready. If a task requires a tool,",
"call it using the TOOL block format described above.",
].join("\n");
}
return [
"You can call tools. To call a tool, reply with a single line containing a <tool> block",
"The client application provides tools beyond your built-in ones. They are NOT in your " +
"native tool registry; they are invoked via a plain-text protocol: the client parses " +
"your reply and executes the tool on the user machine. Treat these client tools as " +
"fully available to you; never claim they are unavailable. To invoke one, reply with " +
"a single line containing a <tool> block",
`with JSON that includes the secret binding "_nonce": "${nonce}":`,
`<tool>{"name": "<tool_name>", "arguments": { ... }, "_nonce": "${nonce}"}</tool>`,
"Only emit the <tool> block when you actually want to call a tool; otherwise answer normally.",
"These client tools ARE available to you in this conversation. Only emit the <tool> " +
"block when you actually want to call a tool; otherwise answer normally.",
"",
"Available tools:",
...lines,
@@ -457,10 +430,7 @@ export function parseToolCallsFromText(
requestedTools?: unknown
): { content: string; toolCalls: OpenAIToolCall[] | null } {
const requestedToolNames = getRequestedToolNames(requestedTools);
if (
typeof text !== "string" ||
(!text.includes("<tool>") && !text.includes("<tool_call"))
) {
if (typeof text !== "string" || (!text.includes("<tool>") && !text.includes("<tool_call"))) {
return { content: text ?? "", toolCalls: null };
}
@@ -538,27 +508,66 @@ interface ToolPrepResult {
effectiveMessages: Array<{ role: string; content: unknown }>;
}
/** One-line nudge appended to the latest user message. Web-UI models weigh the
* current user turn far more heavily than a large system block, and ChatGPT's
* injection heuristics distrust long instructions embedded in user content —
* so the full contract stays in the system block (trailing, see below) and the
* user turn only carries a short pointer back to it, naming the tools. */
function buildToolReminder(toolPrompt: string): string {
const names = (toolPrompt.match(/^- [^:\n]+/gm) || []).map((s) => s.slice(2).trim()).join(", ");
return (
"\n\n[Client protocol reminder: the client-tool contract in the system instructions " +
"is active in this conversation. These client tools ARE available via the <tool> " +
"block protocol" +
(names ? ": " + names : "") +
".]"
);
}
/**
* Extract tools from an OpenAI request body and prepend a tool-system-prompt
* to the messages array when tools are present. Every web-cookie executor
* that wants tool-call support calls this once before building its upstream
* request body.
* Extract tools from an OpenAI request body and inject the tool contract when
* tools are present. Every web-cookie executor that wants tool-call support
* calls this once before building its upstream request body.
*
* Placement matters: the contract used to be PREPENDED as the first system
* message. Executors fold all system messages into one block, so with agentic
* clients whose system prompts exceed ~28K chars the contract sat at the head
* of a huge block and web models (chatgpt-web observed) ignored it, answering
* "tool X is not in my tool set" instead of emitting <tool> blocks. Dual
* placement fixes it: the full contract goes AFTER the client messages (folds
* to the tail of the system block) and a one-line reminder rides at the end of
* the latest user message. Measured on cgpt-web/gpt-5.5-thinking with a
* 30K-char system prompt: prepend 0/3 tool calls, dual placement 16/17 across
* 30K-250K prompts, 30-tool sets, multi-turn tool history, and streaming.
*/
export function prepareToolMessages(
bodyObj: Record<string, unknown>,
messages: Array<{ role: string; content: unknown }>,
options?: SerializeToolOptions
messages: Array<{ role: string; content: unknown }>
): ToolPrepResult {
const requestedTools = bodyObj.tools;
const hasTools = Array.isArray(requestedTools) && requestedTools.length > 0;
if (!hasTools) return { hasTools: false, requestedTools, effectiveMessages: messages };
const toolPrompt = serializeToolsToPrompt(requestedTools, options);
return {
hasTools: true,
requestedTools,
effectiveMessages: [{ role: "system", content: toolPrompt }, ...messages],
};
const toolPrompt = serializeToolsToPrompt(requestedTools);
if (!toolPrompt) return { hasTools: true, requestedTools, effectiveMessages: messages };
const effectiveMessages = [...messages];
const reminder = buildToolReminder(toolPrompt);
for (let i = effectiveMessages.length - 1; i >= 0; i--) {
const msg = effectiveMessages[i];
if (msg?.role !== "user") continue;
if (typeof msg.content === "string") {
effectiveMessages[i] = { ...msg, content: msg.content + reminder };
} else if (Array.isArray(msg.content)) {
effectiveMessages[i] = {
...msg,
content: [...msg.content, { type: "text", text: reminder }],
};
}
break;
}
effectiveMessages.push({ role: "system", content: toolPrompt });
return { hasTools: true, requestedTools, effectiveMessages };
}
interface ToolCompletionResult {

View File

@@ -19,6 +19,11 @@
import zlib from "node:zlib";
import crypto from "node:crypto";
import { decodeNativeTodoWriteCompletion } from "./cursorAgentProtobuf/nativeTodoWrite.ts";
import {
cursorImageAttachmentPath,
encodeSelectedImageBody,
type EncodedImage,
} from "./cursorAgentProtobuf/imageEncoding.ts";
import {
WT_VARINT,
WT_LEN,
@@ -63,25 +68,8 @@ const UM_MESSAGE_ID = 2; // UserMessage.message_id
const UM_SELECTED_CONTEXT = 3; // UserMessage.selected_context (empty placeholder required)
const UM_MODE = 4; // UserMessage.mode (cursor-agent sends 1)
// ─── Vision input (image) field numbers ────────────────────────────────────
// Pinned from cursor-agent's agent.v1 protobuf descriptor (bundle version
// 2026.06.02-8c11d9f, cross-checked against composer-api's older-endpoint
// encoder for shape). Images attach to the current UserMessage through its
// selected_context (field 3): UserMessage.selected_context is a SelectedContext
// whose `selected_images` (field 1) is a repeated SelectedImage. Each
// SelectedImage carries the raw bytes inline in its `data_or_blob_id` oneof
// (the `data` case, field 8) — cursor-agent's CLI instead sends a local file
// `path`, which a proxy cannot use, so we inline the bytes like composer-api.
const SC_SELECTED_IMAGES = 1; // SelectedContext.selected_images [repeated SelectedImage]
const SI_UUID = 2; // SelectedImage.uuid
const SI_DIMENSION = 4; // SelectedImage.dimension (SelectedImage.Dimension)
const SI_MIME_TYPE = 7; // SelectedImage.mime_type
const SI_DATA = 8; // SelectedImage.data (oneof data_or_blob_id) — inline image bytes
const DIM_WIDTH = 1; // SelectedImage.Dimension.width (int32)
const DIM_HEIGHT = 2; // SelectedImage.Dimension.height (int32)
const RM_MODEL_ID = 1; // RequestedModel.model_id
const RM_PARAMETERS = 3; // RequestedModel.parameters [repeated]
@@ -413,58 +401,15 @@ export type AgentRunInput = {
// which the executor's processFrame replies to with the stored bytes.
systemPrompt?: string;
blobStore?: Map<string, Buffer>;
// Vision input: images attached to the current user turn. Encoded inline as
// SelectedContext.selected_images[] (see encodeSelectedImageBody). Empty /
// undefined keeps the request byte-identical to the text-only path.
// Vision input: images attached to the current user turn. Encoded as
// SelectedContext.selected_images[] via blobIdWithData (see
// encodeSelectedImageBody). Empty / undefined keeps the request
// byte-identical to the text-only path.
images?: EncodedImage[];
};
/**
* A resolved image ready to embed in a cursor request. `data` is the raw
* decoded image bytes (already SSRF-checked / size-capped by the executor's
* resolveCursorImages helper). `mimeType` (e.g. "image/png") helps cursor
* decode the inline bytes; `width`/`height` populate the optional Dimension
* sub-message when cheaply known; `uuid` is a stable per-image id.
*/
export type EncodedImage = {
data: Buffer;
mimeType?: string;
width?: number;
height?: number;
uuid: string;
};
/**
* Encode the body of a SelectedImage message (no outer field tag — the caller
* wraps it via encodeMessage(SC_SELECTED_IMAGES, [body])). Sets the inline
* `data` oneof case plus uuid, optional dimension, and mime_type. Fields are
* written in ascending field-number order (canonical protobuf layout).
*/
export function encodeSelectedImageBody(img: EncodedImage): Buffer {
const parts: Buffer[] = [encodeString(SI_UUID, img.uuid)];
if (
typeof img.width === "number" &&
typeof img.height === "number" &&
Number.isFinite(img.width) &&
Number.isFinite(img.height) &&
img.width > 0 &&
img.height > 0
) {
parts.push(
encodeMessage(SI_DIMENSION, [
encodeUInt32Field(DIM_WIDTH, Math.floor(img.width)),
encodeUInt32Field(DIM_HEIGHT, Math.floor(img.height)),
])
);
}
if (img.mimeType) {
parts.push(encodeString(SI_MIME_TYPE, img.mimeType));
}
// data_or_blob_id oneof = data (inline bytes) — field 8, written last to
// keep ascending field order.
parts.push(encodeBytes(SI_DATA, img.data));
return Buffer.concat(parts);
}
export { cursorImageAttachmentPath, encodeSelectedImageBody };
export type { EncodedImage };
/**
* Convert OpenAI tool definitions to cursor McpToolDefinition bodies. Used
@@ -493,12 +438,15 @@ export function encodeAgentRunRequest(input: AgentRunInput): Buffer {
// UserMessage { text, message_id, selected_context, mode=1 }.
// selected_context is normally an empty placeholder (required by the server
// even when empty — see below), but when the turn carries vision input we
// populate its selected_images[] with the inline-encoded images. The
// empty-images path produces byte-identical output to the text-only request.
// populate its selected_images[] with blobIdWithData-encoded images (and
// store the bytes in blobStore for getBlob). The empty-images path produces
// byte-identical output to the text-only request.
const selectedContextParts: Buffer[] = [];
if (input.images && input.images.length > 0) {
for (const img of input.images) {
selectedContextParts.push(encodeMessage(SC_SELECTED_IMAGES, [encodeSelectedImageBody(img)]));
selectedContextParts.push(
encodeMessage(SC_SELECTED_IMAGES, [encodeSelectedImageBody(img, input.blobStore)])
);
}
}
// The empty selected_context placeholder and mode=1 match cursor-agent's

View File

@@ -0,0 +1,80 @@
import crypto from "node:crypto";
import {
encodeBytes,
encodeMessage,
encodeString,
encodeUInt32Field,
} from "./wire.ts";
const SI_UUID = 2;
const SI_PATH = 3;
const SI_DIMENSION = 4;
const SI_MIME_TYPE = 7;
const SI_BLOB_ID_WITH_DATA = 9;
const SIBD_BLOB_ID = 1;
const SIBD_DATA = 2;
const DIM_WIDTH = 1;
const DIM_HEIGHT = 2;
export type EncodedImage = {
data: Buffer;
mimeType?: string;
width?: number;
height?: number;
uuid: string;
};
export function cursorImageAttachmentPath(uuid: string, mimeType?: string): string {
const normalized = (mimeType || "").toLowerCase();
const ext =
normalized === "image/jpeg" || normalized === "image/jpg"
? "jpg"
: normalized === "image/gif"
? "gif"
: normalized === "image/webp"
? "webp"
: "png";
return `attachment-${uuid}.${ext}`;
}
export function encodeSelectedImageBody(
img: EncodedImage,
blobStore?: Map<string, Buffer>
): Buffer {
const blobId = crypto.createHash("sha256").update(img.data).digest();
if (blobStore) {
blobStore.set(blobId.toString("hex"), img.data);
}
const parts: Buffer[] = [
encodeString(SI_UUID, img.uuid),
encodeString(SI_PATH, cursorImageAttachmentPath(img.uuid, img.mimeType)),
];
if (
typeof img.width === "number" &&
typeof img.height === "number" &&
Number.isFinite(img.width) &&
Number.isFinite(img.height) &&
img.width > 0 &&
img.height > 0
) {
parts.push(
encodeMessage(SI_DIMENSION, [
encodeUInt32Field(DIM_WIDTH, Math.floor(img.width)),
encodeUInt32Field(DIM_HEIGHT, Math.floor(img.height)),
])
);
}
if (img.mimeType) {
parts.push(encodeString(SI_MIME_TYPE, img.mimeType));
}
parts.push(
encodeMessage(SI_BLOB_ID_WITH_DATA, [
encodeBytes(SIBD_BLOB_ID, blobId),
encodeBytes(SIBD_DATA, img.data),
])
);
return Buffer.concat(parts);
}

View File

@@ -2,8 +2,8 @@
* Image resolution + security for Cursor vision input.
*
* Turns OpenAI `image_url` parts (base64 `data:` URIs or remote `http(s)`
* URLs) into decoded bytes ready to inline into a cursor SelectedImage
* (see ../utils/cursorAgentProtobuf.ts::encodeSelectedImageBody).
* URLs) into decoded, JPEG-prepped bytes ready for SelectedImage
* `blobIdWithData` encoding (see cursorAgentProtobuf.ts).
*
* Security (OmniRoute hard rules):
* - SSRF: remote fetches go through the repo's canonical outbound guard
@@ -12,9 +12,9 @@
* cloud-metadata hostnames. Client-supplied image URLs are always held to
* the strict public-only policy (never gated by the private-URL toggle that
* admin-configured provider URLs use).
* - Size cap: each image must decode to <= 1 MiB (matches composer-api).
* Enforced both before base64 decode (cheap pre-check) and while streaming
* a remote body (so a hostile server can't stream gigabytes).
* - Size caps: inbound decode/fetch is bounded (16 MiB) so large clipboard
* PNGs can shrink via JPEG soft-cap prep; the final wire image must be
* <= 1 MiB. Soft target is ~100 KiB JPEG for reliable Cursor hydration.
* - Content type: data URIs and URL responses must be `image/*`.
* - Errors throw `CursorImageError` with a clean, path-free message; the
* executor routes it through the sanitized 400 path (hard rule #12).
@@ -23,6 +23,7 @@
import crypto from "node:crypto";
import dns from "node:dns";
import { isIP } from "node:net";
import sharp from "sharp";
import {
parseAndValidatePublicUrl,
isPrivateHost,
@@ -30,14 +31,47 @@ import {
} from "@/shared/network/outboundUrlGuard";
import type { EncodedImage } from "./cursorAgentProtobuf.ts";
// 1 MiB per image — matches composer-api's MAX_CURSOR_IMAGE_BYTES. Large
// enough for a typical screenshot, small enough to bound request size and
// memory.
/** Final per-image byte cap after prep (composer-api / wire bound). */
export const MAX_CURSOR_IMAGE_BYTES = 1024 * 1024;
// Upper bound on the number of images per request. Each image triggers (at
// most) one remote fetch, so an unbounded count is a DoS vector; 12 is well
// above any realistic vision prompt.
/**
* Inbound decode/fetch bomb ceiling before JPEG prep. Large clipboard PNGs may
* exceed {@link MAX_CURSOR_IMAGE_BYTES} raw but shrink under the wire cap after
* re-encode.
*/
export const MAX_CURSOR_IMAGE_DECODE_BYTES = 16 * 1024 * 1024;
/**
* Soft target for Cursor vision hydration. Prefer JPEG at or under this size.
*/
export const CURSOR_VISION_SOFT_MAX_BYTES = 100 * 1024;
/** Soft target when the client requests `detail: original` or `high`. */
export const CURSOR_VISION_SOFT_MAX_BYTES_HIGH = 256 * 1024;
/** Longest edge after Cursor vision prep. */
export const CURSOR_VISION_MAX_EDGE = 2000;
/** Decode bomb: reject images whose sniffed longest edge exceeds this. */
export const MAX_CURSOR_IMAGE_DECODE_EDGE = 8192;
/** Decode bomb: reject images whose sniffed pixel count exceeds this. */
export const MAX_CURSOR_IMAGE_PIXELS = 25_000_000;
const CURSOR_VISION_JPEG_QUALITIES_DEFAULT = [85, 70, 55, 40] as const;
const CURSOR_VISION_JPEG_QUALITIES_HIGH = [90, 80, 65, 50] as const;
const CURSOR_VISION_SOFT_MIN_EDGE = 256;
const CURSOR_VISION_SOFT_SHRINK = 0.85;
const CURSOR_VISION_PASSTHROUGH_MIME = new Set([
"image/jpeg",
"image/jpg",
"image/png",
"image/gif",
"image/webp",
]);
/** Upper bound on images attached to one Cursor turn. */
export const MAX_CURSOR_IMAGES = 12;
// Wall-clock cap for a single remote image fetch. A malformed env value
@@ -64,6 +98,25 @@ export class CursorImageError extends Error {
}
}
function estimatedBase64DecodedBytes(payload: string): number {
return Math.floor((payload.length * 3) / 4);
}
function isHighDetail(detail: string | undefined): boolean {
const normalized = (detail || "").toLowerCase();
return normalized === "high" || normalized === "original";
}
function softMaxBytesForDetail(detail: string | undefined): number {
return isHighDetail(detail) ? CURSOR_VISION_SOFT_MAX_BYTES_HIGH : CURSOR_VISION_SOFT_MAX_BYTES;
}
function jpegQualitiesForDetail(detail: string | undefined): readonly number[] {
return isHighDetail(detail)
? CURSOR_VISION_JPEG_QUALITIES_HIGH
: CURSOR_VISION_JPEG_QUALITIES_DEFAULT;
}
function decodeDataUrl(url: string): { data: Buffer; mimeType: string } {
// data:[<mediatype>][;base64],<data>
const comma = url.indexOf(",");
@@ -86,16 +139,21 @@ function decodeDataUrl(url: string): { data: Buffer; mimeType: string } {
// Reject on the raw payload length BEFORE the regex/normalize pass, so an
// arbitrarily large data URL can't burn CPU on the whitespace strip. Base64
// expands ~4:3, so 2x the byte cap is a safe upper bound on the encoded text.
if (payload.length > MAX_CURSOR_IMAGE_BYTES * 2) {
throw new CursorImageError("Image input is too large (max 1 MiB). Resize and retry.");
// expands ~4:3, so 2x the decode ceiling is a safe upper bound on the text.
if (payload.length > MAX_CURSOR_IMAGE_DECODE_BYTES * 2) {
throw new CursorImageError("Image input is too large to process safely.");
}
const normalized = payload.replace(/\s/g, "");
// Cheap pre-check: 4 base64 chars -> 3 bytes. Reject obviously oversized
// payloads before allocating the decode buffer.
if (Math.floor((normalized.length * 3) / 4) > MAX_CURSOR_IMAGE_BYTES) {
throw new CursorImageError("Image input is too large (max 1 MiB). Resize and retry.");
if (normalized.length === 0) {
throw new CursorImageError("Image data URL contains invalid base64 data.");
}
// Reject lenient Buffer.from acceptances (wrong alphabet, bad padding).
if (normalized.length % 4 !== 0 || !/^[A-Za-z0-9+/]+={0,2}$/.test(normalized)) {
throw new CursorImageError("Image data URL contains invalid base64 data.");
}
if (estimatedBase64DecodedBytes(normalized) > MAX_CURSOR_IMAGE_DECODE_BYTES) {
throw new CursorImageError("Image input is too large to process safely.");
}
let data: Buffer;
@@ -104,11 +162,16 @@ function decodeDataUrl(url: string): { data: Buffer; mimeType: string } {
} catch {
throw new CursorImageError("Image data URL contains invalid base64 data.");
}
// Buffer.from(base64) silently drops invalid trailing chars; guard against a
// payload that decoded to nothing despite being non-empty.
if (normalized.length > 0 && data.length === 0) {
if (data.length === 0) {
throw new CursorImageError("Image data URL contains invalid base64 data.");
}
// Round-trip guard: Node can silently drop trailing garbage.
if (data.toString("base64").replace(/=+$/, "") !== normalized.replace(/=+$/, "")) {
throw new CursorImageError("Image data URL contains invalid base64 data.");
}
if (data.length > MAX_CURSOR_IMAGE_DECODE_BYTES) {
throw new CursorImageError("Image input is too large to process safely.");
}
return { data, mimeType };
}
@@ -216,10 +279,10 @@ async function fetchImageBytes(url: string): Promise<{ data: Buffer; mimeType: s
// Reject early on an oversized Content-Length, then still cap during read
// (the header is advisory / may be absent).
const declaredLen = Number(response.headers.get("content-length") || "0");
if (Number.isFinite(declaredLen) && declaredLen > MAX_CURSOR_IMAGE_BYTES) {
throw new CursorImageError("Image input is too large (max 1 MiB). Resize and retry.");
if (Number.isFinite(declaredLen) && declaredLen > MAX_CURSOR_IMAGE_DECODE_BYTES) {
throw new CursorImageError("Image input is too large to process safely.");
}
const data = await readCapped(response, MAX_CURSOR_IMAGE_BYTES);
const data = await readCapped(response, MAX_CURSOR_IMAGE_DECODE_BYTES);
return { data, mimeType };
} finally {
clearTimeout(timer);
@@ -249,7 +312,7 @@ async function readCapped(response: Response, cap: number): Promise<Buffer> {
const pushCapped = (chunk: Uint8Array) => {
total += chunk.byteLength;
if (total > cap) {
throw new CursorImageError("Image input is too large (max 1 MiB). Resize and retry.");
throw new CursorImageError("Image input is too large to process safely.");
}
chunks.push(Buffer.from(chunk));
};
@@ -284,22 +347,311 @@ async function readCapped(response: Response, cap: number): Promise<Buffer> {
// Last resort: buffer then cap-check (only exotic non-stream bodies).
const buf = Buffer.from(await response.arrayBuffer());
if (buf.length > cap) {
throw new CursorImageError("Image input is too large (max 1 MiB). Resize and retry.");
throw new CursorImageError("Image input is too large to process safely.");
}
return buf;
}
/** Magic-byte format sniff (independent of declared MIME). */
export function sniffCursorImageFormat(
data: Uint8Array
): "png" | "jpeg" | "gif" | "webp" | undefined {
if (
data.byteLength >= 8 &&
data[0] === 0x89 &&
data[1] === 0x50 &&
data[2] === 0x4e &&
data[3] === 0x47 &&
data[4] === 0x0d &&
data[5] === 0x0a &&
data[6] === 0x1a &&
data[7] === 0x0a
) {
return "png";
}
if (
data.byteLength >= 6 &&
data[0] === 0x47 &&
data[1] === 0x49 &&
data[2] === 0x46 &&
data[3] === 0x38
) {
return "gif";
}
if (data.byteLength >= 4 && data[0] === 0xff && data[1] === 0xd8) return "jpeg";
if (
data.byteLength >= 12 &&
data[0] === 0x52 &&
data[1] === 0x49 &&
data[2] === 0x46 &&
data[3] === 0x46 &&
data[8] === 0x57 &&
data[9] === 0x45 &&
data[10] === 0x42 &&
data[11] === 0x50
) {
return "webp";
}
return undefined;
}
/**
* Sniff PNG/JPEG/GIF/WebP dimensions from raw bytes when the header is present.
* Best-effort only — unknown formats return undefined (dimension is optional).
*/
export function sniffCursorImageDimensions(
data: Uint8Array
): { width: number; height: number } | undefined {
// PNG: signature + IHDR chunk (width/height at bytes 16..23)
if (
data.byteLength >= 24 &&
data[0] === 0x89 &&
data[1] === 0x50 &&
data[2] === 0x4e &&
data[3] === 0x47 &&
data[4] === 0x0d &&
data[5] === 0x0a &&
data[6] === 0x1a &&
data[7] === 0x0a
) {
const width = ((data[16]! << 24) | (data[17]! << 16) | (data[18]! << 8) | data[19]!) >>> 0;
const height = ((data[20]! << 24) | (data[21]! << 16) | (data[22]! << 8) | data[23]!) >>> 0;
if (width > 0 && height > 0) return { width, height };
}
// GIF: "GIF8" + width/height as little-endian u16 at bytes 6..9
if (
data.byteLength >= 10 &&
data[0] === 0x47 &&
data[1] === 0x49 &&
data[2] === 0x46 &&
data[3] === 0x38
) {
const width = data[6]! | (data[7]! << 8);
const height = data[8]! | (data[9]! << 8);
if (width > 0 && height > 0) return { width, height };
}
// WebP: RIFF....WEBP + VP8X / VP8 / VP8L
if (
data.byteLength >= 30 &&
data[0] === 0x52 &&
data[1] === 0x49 &&
data[2] === 0x46 &&
data[3] === 0x46 &&
data[8] === 0x57 &&
data[9] === 0x45 &&
data[10] === 0x42 &&
data[11] === 0x50
) {
const fourcc = String.fromCharCode(data[12]!, data[13]!, data[14]!, data[15]!);
if (fourcc === "VP8X") {
const width = 1 + (data[24]! | (data[25]! << 8) | (data[26]! << 16));
const height = 1 + (data[27]! | (data[28]! << 8) | (data[29]! << 16));
if (width > 0 && height > 0) return { width, height };
} else if (fourcc === "VP8 ") {
if (data[23] === 0x9d && data[24] === 0x01 && data[25] === 0x2a) {
const width = (data[26]! | (data[27]! << 8)) & 0x3fff;
const height = (data[28]! | (data[29]! << 8)) & 0x3fff;
if (width > 0 && height > 0) return { width, height };
}
} else if (fourcc === "VP8L" && data[20] === 0x2f) {
const raw = data[21]! | (data[22]! << 8) | (data[23]! << 16) | (data[24]! << 24);
const width = (raw & 0x3fff) + 1;
const height = ((raw >> 14) & 0x3fff) + 1;
if (width > 0 && height > 0) return { width, height };
}
}
// JPEG: scan for SOF0/SOF2 marker with dimensions
if (data.byteLength >= 4 && data[0] === 0xff && data[1] === 0xd8) {
let offset = 2;
while (offset + 8 < data.byteLength) {
if (data[offset] !== 0xff) break;
const marker = data[offset + 1]!;
// Standalone markers (TEM, RSTn, SOI, EOI) carry no length payload.
if (marker === 0x01 || (marker >= 0xd0 && marker <= 0xd9)) {
offset += 2;
continue;
}
const length = (data[offset + 2]! << 8) | data[offset + 3]!;
if (marker === 0xc0 || marker === 0xc2) {
const height = (data[offset + 5]! << 8) | data[offset + 6]!;
const width = (data[offset + 7]! << 8) | data[offset + 8]!;
if (width > 0 && height > 0) return { width, height };
break;
}
if (length < 2) break;
offset += 2 + length;
}
}
return undefined;
}
type PreparedImage = {
data: Buffer;
mimeType: string;
width?: number;
height?: number;
};
/**
* Re-encode toward a JPEG under the soft vision cap when sharp can decode the
* payload. Fail-closed with CursorImageError on unsupported MIME, decode bombs,
* or undecodable bytes. After the quality ladder, edges shrink iteratively
* until the soft byte cap is met (or the min edge floor is hit).
*/
export async function prepareCursorImageForWire(input: {
data: Buffer;
mimeType: string;
detail?: string;
}): Promise<PreparedImage> {
const mime = input.mimeType.toLowerCase();
const softMax = softMaxBytesForDetail(input.detail);
const qualities = jpegQualitiesForDetail(input.detail);
const lowestQuality = qualities[qualities.length - 1]!;
if (!CURSOR_VISION_PASSTHROUGH_MIME.has(mime)) {
throw new CursorImageError("Image input type is unsupported.");
}
const format = sniffCursorImageFormat(input.data);
const sniffed = sniffCursorImageDimensions(input.data);
if (sniffed) {
const edge = Math.max(sniffed.width, sniffed.height);
const pixels = sniffed.width * sniffed.height;
if (edge > MAX_CURSOR_IMAGE_DECODE_EDGE || pixels > MAX_CURSOR_IMAGE_PIXELS) {
throw new CursorImageError("Image input dimensions are too large.");
}
}
// Soft-cap skip: already soft-capped JPEG that has a real SOF (not SOI-only).
const declaredJpeg = mime === "image/jpeg" || mime === "image/jpg";
const alreadySmallJpeg =
declaredJpeg && format === "jpeg" && sniffed !== undefined && input.data.byteLength <= softMax;
if (alreadySmallJpeg) {
return {
data: input.data,
mimeType: "image/jpeg",
width: sniffed!.width,
height: sniffed!.height,
};
}
try {
// Force a full decode before accepting passthrough / encode.
await sharp(input.data, { failOn: "error" }).resize(1, 1).jpeg({ quality: 1 }).toBuffer();
// Passthrough only when declared MIME matches actual JPEG magic.
if (declaredJpeg && format === "jpeg" && input.data.byteLength <= softMax) {
const dims = sniffed ?? (await sharp(input.data).metadata());
const width = typeof dims.width === "number" ? dims.width : undefined;
const height = typeof dims.height === "number" ? dims.height : undefined;
return {
data: input.data,
mimeType: "image/jpeg",
...(width && height && width > 0 && height > 0 ? { width, height } : {}),
};
}
const meta = await sharp(input.data).metadata();
const width = typeof meta.width === "number" ? meta.width : 0;
const height = typeof meta.height === "number" ? meta.height : 0;
if (width > 0 && height > 0) {
const edge = Math.max(width, height);
if (edge > MAX_CURSOR_IMAGE_DECODE_EDGE || width * height > MAX_CURSOR_IMAGE_PIXELS) {
throw new CursorImageError("Image input dimensions are too large.");
}
}
let targetW = width;
let targetH = height;
if (width > 0 && height > 0 && Math.max(width, height) > CURSOR_VISION_MAX_EDGE) {
const scale = CURSOR_VISION_MAX_EDGE / Math.max(width, height);
targetW = Math.max(1, Math.round(width * scale));
targetH = Math.max(1, Math.round(height * scale));
}
const encodeAt = async (w: number, h: number, quality: number): Promise<Buffer> => {
let pipeline = sharp(input.data, { failOn: "error" });
if (w > 0 && h > 0 && (w !== width || h !== height)) {
pipeline = pipeline.resize(w, h);
}
return pipeline.jpeg({ quality, mozjpeg: true }).toBuffer();
};
let best: Buffer | undefined;
for (const quality of qualities) {
const encoded = await encodeAt(targetW, targetH, quality);
if (!best || encoded.byteLength < best.byteLength) best = encoded;
if (encoded.byteLength <= softMax) {
const outDims = sniffCursorImageDimensions(encoded);
return {
data: encoded,
mimeType: "image/jpeg",
...(outDims ?? (targetW > 0 && targetH > 0 ? { width: targetW, height: targetH } : {})),
};
}
}
while (
best &&
best.byteLength > softMax &&
targetW > 0 &&
targetH > 0 &&
Math.max(targetW, targetH) > CURSOR_VISION_SOFT_MIN_EDGE
) {
const nextW = Math.max(1, Math.round(targetW * CURSOR_VISION_SOFT_SHRINK));
const nextH = Math.max(1, Math.round(targetH * CURSOR_VISION_SOFT_SHRINK));
if (Math.max(nextW, nextH) < CURSOR_VISION_SOFT_MIN_EDGE) {
const scale = CURSOR_VISION_SOFT_MIN_EDGE / Math.max(targetW, targetH);
targetW = Math.max(1, Math.round(targetW * scale));
targetH = Math.max(1, Math.round(targetH * scale));
} else {
targetW = nextW;
targetH = nextH;
}
const encoded = await encodeAt(targetW, targetH, lowestQuality);
if (!best || encoded.byteLength < best.byteLength) best = encoded;
if (encoded.byteLength <= softMax) {
const outDims = sniffCursorImageDimensions(encoded);
return {
data: encoded,
mimeType: "image/jpeg",
...(outDims ?? { width: targetW, height: targetH }),
};
}
if (Math.max(targetW, targetH) <= CURSOR_VISION_SOFT_MIN_EDGE) break;
}
if (best) {
const outDims = sniffCursorImageDimensions(best);
return {
data: best,
mimeType: "image/jpeg",
...(outDims ?? (targetW > 0 && targetH > 0 ? { width: targetW, height: targetH } : {})),
};
}
if (declaredJpeg && format !== "jpeg") {
throw new CursorImageError("Image input is not a valid JPEG.");
}
throw new CursorImageError("Image input could not be prepared for Cursor vision.");
} catch (err) {
if (err instanceof CursorImageError) throw err;
throw new CursorImageError("Image input is undecodable or unsupported.");
}
}
/**
* Resolve OpenAI `image_url` URLs (data: or http(s):) into EncodedImage[]
* ready to inline into a cursor request. Each image gets a stable random uuid.
* Throws CursorImageError (clean message, sanitizable) on any invalid /
* oversized / blocked input.
* ready for SelectedImage blobIdWithData encoding. Each image gets a stable
* random uuid. Throws CursorImageError (clean message, sanitizable) on any
* invalid / oversized / blocked / undecodable input.
*/
export async function resolveCursorImages(imageUrls: string[]): Promise<EncodedImage[]> {
export async function resolveCursorImages(
imageUrls: string[],
options?: { detail?: string }
): Promise<EncodedImage[]> {
if (imageUrls.length > MAX_CURSOR_IMAGES) {
throw new CursorImageError(
`Too many images in one request (max ${MAX_CURSOR_IMAGES}).`
);
throw new CursorImageError(`Too many images in one request (max ${MAX_CURSOR_IMAGES}).`);
}
const out: EncodedImage[] = [];
for (const url of imageUrls) {
@@ -314,10 +666,27 @@ export async function resolveCursorImages(imageUrls: string[]): Promise<EncodedI
if (!data.length) {
throw new CursorImageError("Image input is empty.");
}
if (data.length > MAX_CURSOR_IMAGE_BYTES) {
if (data.length > MAX_CURSOR_IMAGE_DECODE_BYTES) {
throw new CursorImageError("Image input is too large to process safely.");
}
const prepared = await prepareCursorImageForWire({
data,
mimeType,
detail: options?.detail,
});
if (prepared.data.length > MAX_CURSOR_IMAGE_BYTES) {
throw new CursorImageError("Image input is too large (max 1 MiB). Resize and retry.");
}
out.push({ data, mimeType, uuid: crypto.randomUUID() });
out.push({
data: prepared.data,
mimeType: prepared.mimeType,
uuid: crypto.randomUUID(),
...(typeof prepared.width === "number" && typeof prepared.height === "number"
? { width: prepared.width, height: prepared.height }
: {}),
});
}
return out;
}
@@ -327,17 +696,11 @@ export async function resolveCursorImages(imageUrls: string[]): Promise<EncodedI
* Returns the raw url strings (data: or http(s):) in order. Non-image parts
* are ignored. A plain string content has no images.
*/
export function extractImageUrls(
content: unknown
): string[] {
export function extractImageUrls(content: unknown): string[] {
if (!Array.isArray(content)) return [];
const urls: string[] = [];
for (const part of content) {
if (
part &&
typeof part === "object" &&
(part as { type?: unknown }).type === "image_url"
) {
if (part && typeof part === "object" && (part as { type?: unknown }).type === "image_url") {
const imageUrl = (part as { image_url?: unknown }).image_url;
if (typeof imageUrl === "string") {
urls.push(imageUrl);

View File

@@ -3,15 +3,20 @@
* Safe for circular references (WeakSet). Iterative frames only (no recursive call stack).
*
* Budgets:
* - ESTIMATE_SIZE_BYTE_LIMIT (256 KiB): early-exit once counted bytes exceed the limit
* - byteLimit param (default ESTIMATE_SIZE_BYTE_LIMIT, 256 KiB): early-exit
* once counted bytes exceed the limit — pass the caller's own threshold
* explicitly rather than relying on the default, since a caller comparing
* against a bigger configured limit would otherwise never see a size
* above 256 KiB.
* - ESTIMATE_SIZE_NODE_BUDGET: max value visits (containers + primitives/elements)
*
* Arrays are walked by index frame (never pre-push/copy every element reference).
* Plain objects yield own enumerable values incrementally (no Object.keys materialization).
* Node-budget exhaustion returns a value strictly above 256 KiB so callers fail closed.
* Node-budget exhaustion returns a value strictly above the effective byteLimit
* so callers fail closed.
*/
/** Byte early-exit threshold (256 KiB). */
/** Default byte early-exit threshold (256 KiB) when a caller doesn't pass its own. */
export const ESTIMATE_SIZE_BYTE_LIMIT = 262_144;
/**
@@ -74,14 +79,22 @@ function expandContainerFrame(stack: Frame[], frame: Exclude<Frame, ValueFrame>)
stack.push({ t: "v", v: (frame.o as Record<string, unknown>)[next.value] });
}
export function estimateSizeFast(value: unknown): number {
/**
* @param byteLimit - early-exit threshold (default ESTIMATE_SIZE_BYTE_LIMIT,
* 256 KiB). Pass the actual threshold you're comparing against (see
* chatCore/logTruncation.ts::truncateForLog) so raising that threshold
* doesn't silently cap what this function is even capable of reporting —
* the byte check and the node-budget fail-closed fallback both key off this
* value, not the fixed module constant, when a caller supplies one.
*/
export function estimateSizeFast(value: unknown, byteLimit = ESTIMATE_SIZE_BYTE_LIMIT): number {
let bytes = 0;
let visitsLeft = ESTIMATE_SIZE_NODE_BUDGET;
const seen = new WeakSet<object>();
const stack: Frame[] = [{ t: "v", v: value }];
while (stack.length > 0) {
if (visitsLeft <= 0) return ESTIMATE_SIZE_BYTE_LIMIT + 1;
if (visitsLeft <= 0) return byteLimit + 1;
const frame = stack.pop()!;
if (!isValueFrame(frame)) {
@@ -96,7 +109,7 @@ export function estimateSizeFast(value: unknown): number {
const ty = typeof v;
if (ty === "string" || ty === "number" || ty === "boolean") {
bytes = addPrimitiveBytes(bytes, v as string | number | boolean);
if (bytes > ESTIMATE_SIZE_BYTE_LIMIT) return bytes;
if (bytes > byteLimit) return bytes;
continue;
}
if (ty === "object") {

View File

@@ -19,6 +19,8 @@
export const FUNCTIONAL_GATEWAY_MIRROR_SUFFIX = " (via ";
const FUNCTIONAL_GATEWAY_MIRROR = Symbol("functionalGatewayMirror");
export interface FunctionalGatewayMirrorsDeps {
/** Ordered list of passthrough gateway provider ids to consider as mirrors. */
gatewayProviderIds: string[];
@@ -40,9 +42,14 @@ interface GatewayMirrorCatalogEntry {
root?: unknown;
name?: unknown;
display_name?: unknown;
[FUNCTIONAL_GATEWAY_MIRROR]?: true;
[key: string]: unknown;
}
export function isFunctionalGatewayMirror(model: GatewayMirrorCatalogEntry): boolean {
return model?.[FUNCTIONAL_GATEWAY_MIRROR] === true;
}
/**
* Append `<gatewayAlias>/<originalId>` mirror entries for every eligible model.
* Returns the original array reference unchanged when nothing is eligible.
@@ -88,14 +95,14 @@ export function appendFunctionalGatewayMirrors<T extends GatewayMirrorCatalogEnt
// Skip if the id already starts with this gateway alias (would double-prefix).
if (id.startsWith(`${chosenAlias}/`)) continue;
const label =
typeof model.name === "string" && model.name ? model.name : modelId;
const label = typeof model.name === "string" && model.name ? model.name : modelId;
aliases.push({
...model,
id: aliasId,
root: id,
owned_by: chosenProvider,
display_name: `${label}${FUNCTIONAL_GATEWAY_MIRROR_SUFFIX}${chosenProvider})`,
[FUNCTIONAL_GATEWAY_MIRROR]: true,
} as T);
}

View File

@@ -13,6 +13,8 @@
* that proxy to thinking-mode models.
*/
import { requiresReasoningReplay } from "../services/reasoningCache.ts";
const PLACEHOLDER = " ";
type JsonRecord = Record<string, unknown>;
@@ -30,24 +32,6 @@ const THINKING_MODEL_PATTERNS: RegExp[] = [
/\bmimo\b/i, // xiaomi-tokenplan mimo family (e.g. xiaomi-tokenplan/mimo-v2.5-pro)
];
const AUTHENTIC_REASONING_MODEL_PATTERN = /(?:^|\/)kimi-k(?:3|2\.7-code)(?:$|-)/i;
/**
* Native Moonshot K3/K2.7 replay must use the original reasoning content.
* A fabricated placeholder changes preserved-thinking history and is not a
* valid substitute when the client and reasoning cache both lack the field.
*/
export function requiresAuthenticReasoningContent(provider: unknown, model: unknown): boolean {
const normalizedProvider = String(provider ?? "")
.trim()
.toLowerCase();
const normalizedModel = String(model ?? "").trim();
return (
(normalizedProvider === "moonshot" || normalizedProvider === "kimi") &&
AUTHENTIC_REASONING_MODEL_PATTERN.test(normalizedModel)
);
}
export function isThinkingMessageModel(model: string | undefined | null): boolean {
if (!model || typeof model !== "string") return false;
return THINKING_MODEL_PATTERNS.some((re) => re.test(model));
@@ -62,7 +46,11 @@ export function shouldInjectReasoningContentPlaceholder(
.toLowerCase();
return (
(normalizedProvider === "moonshot" || normalizedProvider === "kimi") &&
!requiresAuthenticReasoningContent(normalizedProvider, model) &&
!requiresReasoningReplay({
provider: normalizedProvider,
model: String(model ?? ""),
allowLegacyFallback: false,
}) &&
isThinkingMessageModel(model)
);
}

View File

@@ -62,10 +62,38 @@ export function hasAnyReasoningSignal(value: unknown): boolean {
);
}
const STRIPPABLE_REASONING_FIELDS = [
"reasoning_content",
"reasoning",
"reasoning_text",
"thinking",
"thought",
] as const;
/**
* Strip the internal replay placeholder from a single string reasoning field,
* deleting the field when nothing meaningful remains. Returns true only when a
* present string field was fully stripped to empty (absent/non-string fields
* return false so callers can distinguish "removed" from "never had text").
*/
function stripPlaceholderFromField(target: JsonRecord, field: string): boolean {
const value = target[field];
if (typeof value !== "string") return false;
const stripped = stripInternalReasoningPlaceholder(value);
if (stripped === "") {
delete target[field];
return true;
}
if (stripped !== value) target[field] = stripped;
return false;
}
export function copyOpenAICompatibleReasoningFields(source: JsonRecord, target: JsonRecord) {
if (source.reasoning_content !== undefined) target.reasoning_content = source.reasoning_content;
if (source.reasoning !== undefined) target.reasoning = source.reasoning;
if (source.reasoning_text !== undefined) target.reasoning_text = source.reasoning_text;
if (source.thinking !== undefined) target.thinking = source.thinking;
if (source.thought !== undefined) target.thought = source.thought;
if (Array.isArray(source.reasoning_details)) target.reasoning_details = source.reasoning_details;
if (!getReadableReasoningValue(target)) {
const mirrored = getUnsupportedReasoningValue(source);
@@ -73,15 +101,31 @@ export function copyOpenAICompatibleReasoningFields(source: JsonRecord, target:
}
// ponytail: the internal replay placeholder is request scaffolding, never
// real reasoning — models echo it and it poisons client history + the cache
// (#8081 echo). Strip it from anything we forward to the client.
if (typeof target.reasoning_content === "string") {
const stripped = stripInternalReasoningPlaceholder(target.reasoning_content);
if (stripped === "") delete target.reasoning_content;
else if (stripped !== target.reasoning_content) target.reasoning_content = stripped;
// (#8081 echo). Strip it from anything we forward to the client, including
// non-standard reasoning fields (reasoning_text / thinking / thought) and
// reasoning_details items that non-OpenAI-compatible upstreams (e.g.
// Venice) use (#9765 uncovered path).
for (const field of STRIPPABLE_REASONING_FIELDS) {
stripPlaceholderFromField(target, field);
}
if (typeof target.reasoning === "string") {
const stripped = stripInternalReasoningPlaceholder(target.reasoning);
if (stripped === "") delete target.reasoning;
else if (stripped !== target.reasoning) target.reasoning = stripped;
if (Array.isArray(target.reasoning_details)) {
const cleaned: unknown[] = [];
for (const detail of target.reasoning_details) {
const record = asReasoningRecord(detail);
const next: JsonRecord = { ...record };
// Track whether the item originally carried text/content at all so
// non-text details (e.g. `reasoning.encrypted` carrying only `data`)
// survive untouched.
const hadText = typeof next.text === "string";
const hadContent = typeof next.content === "string";
stripPlaceholderFromField(next, "text");
stripPlaceholderFromField(next, "content");
const textGone = next.text === undefined;
const contentGone = next.content === undefined;
if ((hadText || hadContent) && textGone && contentGone) continue;
cleaned.push(next);
}
if (cleaned.length === 0) delete target.reasoning_details;
else target.reasoning_details = cleaned;
}
}

View File

@@ -1,4 +1,5 @@
import { getPendingById } from "@/lib/usage/usageHistory";
import { getChatLogMaxDepth } from "@/lib/logEnv";
import { sanitizeErrorMessage } from "./error.ts";
type JsonRecord = Record<string, unknown>;
@@ -148,7 +149,7 @@ export function cloneBoundedForLog(value: unknown, depth = 0, key: string | null
if (ArrayBuffer.isView(value)) {
return `[binary ${(value as ArrayBufferView).byteLength} bytes]`;
}
if (depth >= 6) return "[MaxDepth]";
if (depth >= getChatLogMaxDepth()) return "[MaxDepth]";
if (Array.isArray(value)) {
// Idempotence (#7847): an already-bounded array is [marker, ...tail] — MAX_LOG_ARRAY_ITEMS + 1

View File

@@ -29,6 +29,59 @@ export type PipelineStreamErrorHandler = (event: {
statusCode: number;
}) => boolean;
export type ClientDisconnectEvent = { reason: string; duration: number };
/**
* #9653: a client that closes its connection right after reading a fully-completed
* SSE stream can race the stream's own completion bookkeeping — the bytes already
* reached the client, but the transform stream's completion callback (which flips
* `isStreamCompletionRecorded()` to true) hasn't finished bubbling up yet when the
* disconnect handler fires. Persisting immediately in that case records a false
* 499 with zero token usage for a request that actually delivered its full response.
*
* This wraps a disconnect finalizer with a grace period: instead of finalizing
* immediately, poll `isStreamCompletionRecorded()` until it flips true (a real
* completion landed — nothing more to do) or the deadline passes (genuinely gone —
* finalize as a 499 same as before). Pass `gracePeriodMs <= 0` to disable and
* finalize immediately, matching the pre-#9653 behavior.
*/
export function createClientDisconnectGraceHandler({
isStreamCompletionRecorded,
gracePeriodMs,
finalize,
pollIntervalMs = 250,
setTimeoutFn = setTimeout,
}: {
isStreamCompletionRecorded: () => boolean;
gracePeriodMs: number;
finalize: (event: ClientDisconnectEvent) => unknown;
pollIntervalMs?: number;
setTimeoutFn?: (callback: () => void, ms: number) => unknown;
}): (event: ClientDisconnectEvent) => boolean {
return (event) => {
if (isStreamCompletionRecorded()) return true;
if (gracePeriodMs <= 0) {
finalize(event);
return true;
}
const deadline = Date.now() + gracePeriodMs;
const poll = () => {
if (isStreamCompletionRecorded()) return;
if (Date.now() >= deadline) {
finalize(event);
return;
}
setTimeoutFn(poll, pollIntervalMs);
};
setTimeoutFn(poll, pollIntervalMs);
// Claim "handled" immediately so the caller's own immediate-finalize fallback
// doesn't fire while the grace-period poll is still pending.
return true;
};
}
export function finalizeStreamRequestLog({
pendingRequestId,
model,
@@ -107,9 +160,7 @@ export function createStreamFailureFinalizers({
const message = failure.message || "Upstream stream error";
const code = failure.code || failure.type || String(status);
const classification =
failure.code || failure.type
? { code: failure.code, type: failure.type }
: undefined;
failure.code || failure.type ? { code: failure.code, type: failure.type } : undefined;
if (!isFailureCompletionRecorded()) {
const errorBody = buildErrorBody(status, message, undefined, classification);

411
package-lock.json generated
View File

@@ -1,12 +1,12 @@
{
"name": "omniroute",
"version": "3.8.50",
"version": "3.8.49",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "omniroute",
"version": "3.8.50",
"version": "3.8.49",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -30,6 +30,7 @@
"bottleneck": "^2.19.5",
"clsx": "^2.1.1",
"commander": "^15.0.0",
"cron-parser": "^5.6.2",
"csv-stringify": "^6.7.0",
"dompurify": "^3.4.13",
"express": "^5.2.1",
@@ -73,13 +74,14 @@
"recharts": "^3.8.1",
"safe-regex": "^2.1.1",
"selfsigned": "^5.5.0",
"sharp": "^0.35.3",
"smol-toml": "1.7.1",
"socks": "^2.8.7",
"sql.js": "^1.14.1",
"sqlite-vec": "^0.1.9",
"tailwind-merge": "^3.6.0",
"tsx": "^4.23.0",
"undici": "^8.10.0",
"undici": "^8.3.0",
"update-notifier": "^7.3.1",
"uuid": "^14.0.0",
"ws": "^8.18.0",
@@ -132,7 +134,6 @@
"lint-staged": "^17.0.8",
"lockfile-lint": "^5.0.0",
"node-loader": "^2.1.0",
"opencode-ai": "1.18.8",
"playwright-ctrf-json-reporter": "^0.0.29",
"prettier": "^3.8.3",
"promptfoo": "^0.121.18",
@@ -152,7 +153,7 @@
"@atjsh/llmlingua-2": "2.0.3",
"@huggingface/transformers": "3.5.2",
"@tensorflow/tfjs": "4.22.0",
"better-sqlite3": "^13.0.2",
"better-sqlite3": "^13.0.1",
"js-tiktoken": "^1.0.20",
"keytar": "^7.9.0",
"tls-client-node": "^0.2.0",
@@ -3560,7 +3561,6 @@
"resolved": "https://registry.npmjs.org/@img/colour/-/colour-1.1.0.tgz",
"integrity": "sha512-Td76q7j57o/tLVdgS746cYARfSyxk8iEfRxewL9h4OMzYhbW4TAcppl0mT4eyqXddh6L/jwoM75mo7ixa/pCeQ==",
"license": "MIT",
"optional": true,
"engines": {
"node": ">=18"
}
@@ -3693,6 +3693,9 @@
"cpu": [
"arm"
],
"libc": [
"glibc"
],
"license": "LGPL-3.0-or-later",
"optional": true,
"os": [
@@ -3709,6 +3712,9 @@
"cpu": [
"arm64"
],
"libc": [
"glibc"
],
"license": "LGPL-3.0-or-later",
"optional": true,
"os": [
@@ -3725,6 +3731,9 @@
"cpu": [
"ppc64"
],
"libc": [
"glibc"
],
"license": "LGPL-3.0-or-later",
"optional": true,
"os": [
@@ -3741,6 +3750,9 @@
"cpu": [
"riscv64"
],
"libc": [
"glibc"
],
"license": "LGPL-3.0-or-later",
"optional": true,
"os": [
@@ -3757,6 +3769,9 @@
"cpu": [
"s390x"
],
"libc": [
"glibc"
],
"license": "LGPL-3.0-or-later",
"optional": true,
"os": [
@@ -3773,6 +3788,9 @@
"cpu": [
"x64"
],
"libc": [
"glibc"
],
"license": "LGPL-3.0-or-later",
"optional": true,
"os": [
@@ -3789,6 +3807,9 @@
"cpu": [
"arm64"
],
"libc": [
"musl"
],
"license": "LGPL-3.0-or-later",
"optional": true,
"os": [
@@ -3805,6 +3826,9 @@
"cpu": [
"x64"
],
"libc": [
"musl"
],
"license": "LGPL-3.0-or-later",
"optional": true,
"os": [
@@ -3821,6 +3845,9 @@
"cpu": [
"arm"
],
"libc": [
"glibc"
],
"license": "Apache-2.0",
"optional": true,
"os": [
@@ -3843,6 +3870,9 @@
"cpu": [
"arm64"
],
"libc": [
"glibc"
],
"license": "Apache-2.0",
"optional": true,
"os": [
@@ -3865,6 +3895,9 @@
"cpu": [
"ppc64"
],
"libc": [
"glibc"
],
"license": "Apache-2.0",
"optional": true,
"os": [
@@ -3887,6 +3920,9 @@
"cpu": [
"riscv64"
],
"libc": [
"glibc"
],
"license": "Apache-2.0",
"optional": true,
"os": [
@@ -3909,6 +3945,9 @@
"cpu": [
"s390x"
],
"libc": [
"glibc"
],
"license": "Apache-2.0",
"optional": true,
"os": [
@@ -3931,6 +3970,9 @@
"cpu": [
"x64"
],
"libc": [
"glibc"
],
"license": "Apache-2.0",
"optional": true,
"os": [
@@ -3953,6 +3995,9 @@
"cpu": [
"arm64"
],
"libc": [
"musl"
],
"license": "Apache-2.0",
"optional": true,
"os": [
@@ -3975,6 +4020,9 @@
"cpu": [
"x64"
],
"libc": [
"musl"
],
"license": "Apache-2.0",
"optional": true,
"os": [
@@ -5346,6 +5394,9 @@
"cpu": [
"arm64"
],
"libc": [
"glibc"
],
"license": "MIT",
"optional": true,
"os": [
@@ -5362,6 +5413,9 @@
"cpu": [
"arm64"
],
"libc": [
"musl"
],
"license": "MIT",
"optional": true,
"os": [
@@ -5378,6 +5432,9 @@
"cpu": [
"x64"
],
"libc": [
"glibc"
],
"license": "MIT",
"optional": true,
"os": [
@@ -5394,6 +5451,9 @@
"cpu": [
"x64"
],
"libc": [
"musl"
],
"license": "MIT",
"optional": true,
"os": [
@@ -10611,6 +10671,9 @@
"arm64"
],
"dev": true,
"libc": [
"glibc"
],
"license": "MIT",
"optional": true,
"os": [
@@ -10628,6 +10691,9 @@
"arm64"
],
"dev": true,
"libc": [
"musl"
],
"license": "MIT",
"optional": true,
"os": [
@@ -10645,6 +10711,9 @@
"x64"
],
"dev": true,
"libc": [
"glibc"
],
"license": "MIT",
"optional": true,
"os": [
@@ -10662,6 +10731,9 @@
"x64"
],
"dev": true,
"libc": [
"musl"
],
"license": "MIT",
"optional": true,
"os": [
@@ -12708,6 +12780,9 @@
"arm"
],
"dev": true,
"libc": [
"glibc"
],
"optional": true,
"os": [
"linux"
@@ -12721,6 +12796,9 @@
"arm"
],
"dev": true,
"libc": [
"musl"
],
"optional": true,
"os": [
"linux"
@@ -12734,6 +12812,9 @@
"arm64"
],
"dev": true,
"libc": [
"glibc"
],
"optional": true,
"os": [
"linux"
@@ -12747,6 +12828,9 @@
"arm64"
],
"dev": true,
"libc": [
"musl"
],
"optional": true,
"os": [
"linux"
@@ -12760,6 +12844,9 @@
"x64"
],
"dev": true,
"libc": [
"glibc"
],
"optional": true,
"os": [
"linux"
@@ -12773,6 +12860,9 @@
"x64"
],
"dev": true,
"libc": [
"musl"
],
"optional": true,
"os": [
"linux"
@@ -13598,14 +13688,11 @@
}
},
"node_modules/balanced-match": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-4.0.4.tgz",
"integrity": "sha512-BLrgEcRTwX2o6gGxGOCNyMvGSp35YofuYzw9h1IMTRmKqttAZZVU67bdb9Pr2vUHA8+j3i2tJfjO6C6+4myGTA==",
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
"integrity": "sha512-3oSeUO0TMV67hN1AmbXsK4yaqU7tjiHlbxRDZOpH0KW9+CeX4bRAaX0Anxt0tx2MrpRpWwQaPwIlISEJhYU5Pw==",
"dev": true,
"license": "MIT",
"engines": {
"node": "18 || 20 || >=22"
}
"license": "MIT"
},
"node_modules/base64-js": {
"version": "1.5.1",
@@ -13670,9 +13757,10 @@
}
},
"node_modules/better-sqlite3": {
"version": "13.0.2",
"resolved": "https://registry.npmjs.org/better-sqlite3/-/better-sqlite3-13.0.2.tgz",
"integrity": "sha512-jW6oufeDhXZaiX9Lw5A+oerVClx4iFrI6uDj1zu7SqUAjak9vbJvA0NEcKLNxHiQHb6kYCoFzzXYV0YOauhV3g==",
"version": "13.0.1",
"resolved": "https://registry.npmjs.org/better-sqlite3/-/better-sqlite3-13.0.1.tgz",
"integrity": "sha512-LYpmOXdkpQYf4wmlxkdzW01XGlOXNIbjLg45yNkh0FQ4814VbK9PdOFmhZpYbej+EZtR/i3FDdhEG98HqZdgnA==",
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"dependencies": {
@@ -13954,16 +14042,14 @@
}
},
"node_modules/brace-expansion": {
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"version": "1.1.18",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.18.tgz",
"integrity": "sha512-Edep/X9fGqVNmzKBVsDYIOtD+z1tuezV70LBjdCst9Tqu76lsnvRiZ6oTic1n+/BIwX6QDGAO94PN4N2SADvtw==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "20 || >=22"
"balanced-match": "^1.0.0",
"concat-map": "0.0.1"
}
},
"node_modules/braces": {
@@ -15794,6 +15880,18 @@
"node": ">= 6"
}
},
"node_modules/cron-parser": {
"version": "5.7.0",
"resolved": "https://registry.npmjs.org/cron-parser/-/cron-parser-5.7.0.tgz",
"integrity": "sha512-iSpDHpwwW/GhIg4JVODYlWUEpMNSimaHvqOhHpOz1W+Y97z1lL1nf+dpcF17cNwFRpTtKN9devgi1fxflp3Phw==",
"license": "MIT",
"dependencies": {
"luxon": "^3.7.2"
},
"engines": {
"node": ">=18"
}
},
"node_modules/cross-env": {
"version": "10.1.0",
"resolved": "https://registry.npmjs.org/cross-env/-/cross-env-10.1.0.tgz",
@@ -24392,25 +24490,6 @@
"node": ">= 14"
}
},
"node_modules/libxmljs2/node_modules/balanced-match": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
"integrity": "sha512-3oSeUO0TMV67hN1AmbXsK4yaqU7tjiHlbxRDZOpH0KW9+CeX4bRAaX0Anxt0tx2MrpRpWwQaPwIlISEJhYU5Pw==",
"dev": true,
"license": "MIT",
"optional": true
},
"node_modules/libxmljs2/node_modules/brace-expansion": {
"version": "2.1.4",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.4.tgz",
"integrity": "sha512-hGfVzPxthbf3+2yjg/RBs60cB0FhqBS/zvdV/4wn4/BmN0bNMMHPc4V/BbFieqf1TKAGGAHnY4eSjajCl0f2Xg==",
"dev": true,
"license": "MIT",
"optional": true,
"dependencies": {
"balanced-match": "^1.0.0"
}
},
"node_modules/libxmljs2/node_modules/cacache": {
"version": "19.0.1",
"resolved": "https://registry.npmjs.org/cacache/-/cacache-19.0.1.tgz",
@@ -25433,6 +25512,15 @@
"react": "^16.5.1 || ^17.0.0 || ^18.0.0 || ^19.0.0"
}
},
"node_modules/luxon": {
"version": "3.7.2",
"resolved": "https://registry.npmjs.org/luxon/-/luxon-3.7.2.tgz",
"integrity": "sha512-vtEhXh/gNjI9Yg1u4jX/0YVPMvxzHuGgCm6tC5kZyb08yjGWGnqAjGJvcXbqQR2P3MyMEFnRbpcdFS6PBcLqew==",
"license": "MIT",
"engines": {
"node": ">=12"
}
},
"node_modules/magic-string": {
"version": "0.30.21",
"resolved": "https://registry.npmjs.org/magic-string/-/magic-string-0.30.21.tgz",
@@ -26883,24 +26971,6 @@
"node": "*"
}
},
"node_modules/minimatch/node_modules/balanced-match": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
"integrity": "sha512-3oSeUO0TMV67hN1AmbXsK4yaqU7tjiHlbxRDZOpH0KW9+CeX4bRAaX0Anxt0tx2MrpRpWwQaPwIlISEJhYU5Pw==",
"dev": true,
"license": "MIT"
},
"node_modules/minimatch/node_modules/brace-expansion": {
"version": "1.1.18",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.18.tgz",
"integrity": "sha512-Edep/X9fGqVNmzKBVsDYIOtD+z1tuezV70LBjdCst9Tqu76lsnvRiZ6oTic1n+/BIwX6QDGAO94PN4N2SADvtw==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^1.0.0",
"concat-map": "0.0.1"
}
},
"node_modules/minimist": {
"version": "1.2.8",
"resolved": "https://registry.npmjs.org/minimist/-/minimist-1.2.8.tgz",
@@ -28739,205 +28809,6 @@
}
}
},
"node_modules/opencode-ai": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-ai/-/opencode-ai-1.18.8.tgz",
"integrity": "sha512-eZvYK0rIc/NUDQ+s3LsO9gyUU3MswsbNOLZz06iPwVhbg/2jF6bkTaroBgiIdFWKwUn5sj+kSMc4TBYxFkMrNQ==",
"cpu": [
"arm64",
"x64"
],
"dev": true,
"hasInstallScript": true,
"license": "MIT",
"os": [
"darwin",
"linux",
"win32"
],
"bin": {
"opencode": "bin/opencode.exe"
},
"optionalDependencies": {
"opencode-darwin-arm64": "1.18.8",
"opencode-darwin-x64": "1.18.8",
"opencode-darwin-x64-baseline": "1.18.8",
"opencode-linux-arm64": "1.18.8",
"opencode-linux-arm64-musl": "1.18.8",
"opencode-linux-x64": "1.18.8",
"opencode-linux-x64-baseline": "1.18.8",
"opencode-linux-x64-baseline-musl": "1.18.8",
"opencode-linux-x64-musl": "1.18.8",
"opencode-windows-arm64": "1.18.8",
"opencode-windows-x64": "1.18.8",
"opencode-windows-x64-baseline": "1.18.8"
}
},
"node_modules/opencode-darwin-arm64": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-darwin-arm64/-/opencode-darwin-arm64-1.18.8.tgz",
"integrity": "sha512-ZZCIEgTvHxOHk52Aeqhq59t/R0aqs29bPIgu45XE4rkgjmn/XCkTWalCPtyzJHipdcEbq/g0lqsE1OlJV0oNbA==",
"cpu": [
"arm64"
],
"dev": true,
"optional": true,
"os": [
"darwin"
]
},
"node_modules/opencode-darwin-x64": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-darwin-x64/-/opencode-darwin-x64-1.18.8.tgz",
"integrity": "sha512-2EXRMJbRKnFPWI9oDU9tb7jDGmKiPmfjCLtwJMe3EF57h5wfcdEH9sP25bR3Og5NbE2M+PtMcJm0jMeHn2XoLQ==",
"cpu": [
"x64"
],
"dev": true,
"optional": true,
"os": [
"darwin"
]
},
"node_modules/opencode-darwin-x64-baseline": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-darwin-x64-baseline/-/opencode-darwin-x64-baseline-1.18.8.tgz",
"integrity": "sha512-eLXa2tK9LRuZ5e20QG2k4dmWAA5xnLgJ1afRTSD0/ybE6CAeK02i8vFCnFFDaxuBo+gnq+yqO8AkqvN1m64V/Q==",
"cpu": [
"x64"
],
"dev": true,
"optional": true,
"os": [
"darwin"
]
},
"node_modules/opencode-linux-arm64": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-linux-arm64/-/opencode-linux-arm64-1.18.8.tgz",
"integrity": "sha512-7kj3c9JEdryHgK+o8zE/N9KzTOdbiDn6KpY8dl+hM9n5Cnmxezx4IAlgJeC9QxpIx8Omop6CYuZ+17KfrKdKLw==",
"cpu": [
"arm64"
],
"dev": true,
"optional": true,
"os": [
"linux"
]
},
"node_modules/opencode-linux-arm64-musl": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-linux-arm64-musl/-/opencode-linux-arm64-musl-1.18.8.tgz",
"integrity": "sha512-tww5TF/LIOv/GoTNyzGYgqDRhbJrhoMu8R+p5yD/SpnXPg3rcfYREw2wRy9yikyPU9sAQksuIIteTsyGerPjlA==",
"cpu": [
"arm64"
],
"dev": true,
"libc": [
"musl"
],
"optional": true,
"os": [
"linux"
]
},
"node_modules/opencode-linux-x64": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-linux-x64/-/opencode-linux-x64-1.18.8.tgz",
"integrity": "sha512-Sm4fbQ9BdLI6hgN6FYYX8Nql+Sqe/2EKHJu3iWg0UYs93AXN4ROi0rvOmRbMk+ycYgOchb0hL6Ti2opxLx17sg==",
"cpu": [
"x64"
],
"dev": true,
"optional": true,
"os": [
"linux"
]
},
"node_modules/opencode-linux-x64-baseline": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-linux-x64-baseline/-/opencode-linux-x64-baseline-1.18.8.tgz",
"integrity": "sha512-egeEF4tk1rK9flIQjjeSVB9cR/X3zUti0pNAHW6ROJkNkj72z2C2FmjK1hZbfjtteCueMXPLptS23JROHGWL1w==",
"cpu": [
"x64"
],
"dev": true,
"optional": true,
"os": [
"linux"
]
},
"node_modules/opencode-linux-x64-baseline-musl": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-linux-x64-baseline-musl/-/opencode-linux-x64-baseline-musl-1.18.8.tgz",
"integrity": "sha512-S+438BXs48gLeXX/ya4TSNytDy9mliU3sOAf6j9rfFjzGiF/S08LedemSAnHkr0riBtamik1aRPSmTjhQ0dOBg==",
"cpu": [
"x64"
],
"dev": true,
"libc": [
"musl"
],
"optional": true,
"os": [
"linux"
]
},
"node_modules/opencode-linux-x64-musl": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-linux-x64-musl/-/opencode-linux-x64-musl-1.18.8.tgz",
"integrity": "sha512-c+E4Zsp0DYVcuqcDtgxw/4YcFLrVYWdGBR8x4CzpW48ga3RshaH+BlmUiy+GY0yr1x6UR+e2V3w4uzvzm/L9UQ==",
"cpu": [
"x64"
],
"dev": true,
"libc": [
"musl"
],
"optional": true,
"os": [
"linux"
]
},
"node_modules/opencode-windows-arm64": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-windows-arm64/-/opencode-windows-arm64-1.18.8.tgz",
"integrity": "sha512-7NjdtEIiX28kmsKD9jHbFG4bbwBB5T4dAe2UwdnOqCBb2cl+ETV5eO6kbdC/xWxrgOghgZM1Wtw791T5pQPyag==",
"cpu": [
"arm64"
],
"dev": true,
"optional": true,
"os": [
"win32"
]
},
"node_modules/opencode-windows-x64": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-windows-x64/-/opencode-windows-x64-1.18.8.tgz",
"integrity": "sha512-G+NEgEMvu/dEYshH5IaqHVTmsHVuGdORBvVmgphFiknT7q/NXPuoZCMtMIdfNlEFbu54BlzRDdJCR3Mqe98gUw==",
"cpu": [
"x64"
],
"dev": true,
"optional": true,
"os": [
"win32"
]
},
"node_modules/opencode-windows-x64-baseline": {
"version": "1.18.8",
"resolved": "https://registry.npmjs.org/opencode-windows-x64-baseline/-/opencode-windows-x64-baseline-1.18.8.tgz",
"integrity": "sha512-IGbjFyWoSN9rdGUJX7TWkQ1Yl673Q3dDna54b5NtqeRcZ839p+Z47zzM5m883HKAxrdKCC6Z22HYuDcXLV0laA==",
"cpu": [
"x64"
],
"dev": true,
"optional": true,
"os": [
"win32"
]
},
"node_modules/opener": {
"version": "1.5.2",
"resolved": "https://registry.npmjs.org/opener/-/opener-1.5.2.tgz",
@@ -32174,13 +32045,6 @@
"url": "https://github.com/sponsors/isaacs"
}
},
"node_modules/rimraf/node_modules/balanced-match": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
"integrity": "sha512-3oSeUO0TMV67hN1AmbXsK4yaqU7tjiHlbxRDZOpH0KW9+CeX4bRAaX0Anxt0tx2MrpRpWwQaPwIlISEJhYU5Pw==",
"dev": true,
"license": "MIT"
},
"node_modules/rimraf/node_modules/brace-expansion": {
"version": "2.1.4",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-2.1.4.tgz",
@@ -32778,7 +32642,6 @@
"resolved": "https://registry.npmjs.org/sharp/-/sharp-0.35.3.tgz",
"integrity": "sha512-ej0zVHuZGHCiABXcNxeYhpRnPNPAcvbG8RMdBAhDAxLKkCRVSpK3Iyu7qbqw3JMzoj0REeM6f3tJLtVwl0023Q==",
"license": "Apache-2.0",
"optional": true,
"dependencies": {
"@img/colour": "^1.1.0",
"detect-libc": "^2.1.2",
@@ -32828,7 +32691,6 @@
"resolved": "https://registry.npmjs.org/semver/-/semver-7.8.5.tgz",
"integrity": "sha512-Y7/KDsb8LjooZpwaqGyulO6DQlksgCncchHGk+sZIY4SBvUocMBEFH5Ur1fI4dV+Jvl0w6cjvucaIi40puRioA==",
"license": "ISC",
"optional": true,
"bin": {
"semver": "bin/semver.js"
},
@@ -35157,9 +35019,9 @@
"license": "MIT"
},
"node_modules/undici": {
"version": "8.10.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-8.10.0.tgz",
"integrity": "sha512-HvltHd7avK13QIw/oLe4qoOLyoVSoafqJ2jYOrtMRBkbYT31eiBQ8O0ehRKZiEZCMEyLFQNIADpgCWC5fALvYQ==",
"version": "8.9.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-8.9.0.tgz",
"integrity": "sha512-aWZpUj7XoGonMClx4gdDRfgBjqeA+F473aDmROQQbM9n6PRfK/u1q/a0X4wMTgcHfT8H6fpbt98PFuDUwFg2YA==",
"license": "MIT",
"engines": {
"node": ">=22.19.0"
@@ -36950,7 +36812,12 @@
},
"open-sse": {
"name": "@omniroute/open-sse",
"version": "3.8.50"
"version": "3.8.49",
"dependencies": {
"@toon-format/toon": "^4.1.0",
"safe-regex": "^2.1.1",
"smol-toml": "1.7.1"
}
}
}
}

View File

@@ -1,7 +1,7 @@
{
"name": "omniroute",
"version": "3.8.50",
"description": "Unified AI router with 290 providers, RTK+Caveman compression, auto fallback, MCP/A2A, desktop, PWA, and OpenAI-compatible APIs.",
"version": "3.8.49",
"description": "Unified AI router with 160+ providers, RTK+Caveman compression, auto fallback, MCP/A2A, desktop, PWA, and OpenAI-compatible APIs.",
"type": "module",
"bin": {
"omniroute": "bin/omniroute.mjs",
@@ -23,7 +23,6 @@
".env.example",
"scripts/build/postinstall.mjs",
"scripts/build/fixTlsClientNodeBinary.mjs",
"scripts/build/fixPlaywrightAndroid.mjs",
"bin/cli/runtime/",
"scripts/postinstall.mjs",
"scripts/build/postinstallSupport.mjs",
@@ -34,15 +33,11 @@
"scripts/dev/tls-options.mjs",
"scripts/check/check-supported-node-runtime.ts",
"scripts/dev/sync-env.mjs",
"scripts/build/assembleStandalone.mjs",
"scripts/build/backendOnlyPages.mjs",
"scripts/build/build-next-isolated.mjs",
"scripts/build/build-tproxy-native.mjs",
"scripts/build/native-binary-compat.mjs",
"scripts/build/build-next-isolated.mjs",
"scripts/build/runtime-env.mjs",
"README.md",
"LICENSE",
"!**/node_modules/**",
"!**/__tests__/**",
"!**/*.test.ts",
"!**/*.test.tsx",
@@ -115,8 +110,6 @@
"test:unit:ci": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=4 tests/unit/*.test.ts \"tests/unit/{api,auth,authz,build,cli,cli-helper,combo,compression,correctness,cors,db,db-adapters,docs,gamification,guardrails,lib,mcp,memory,runtime,security,services,settings,shared,ui,usage}/**/*.test.ts\" \"tests/unit/**/*.test.mjs\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=4 \"tests/unit/dashboard/**/*.test.ts\" && npm run test:unit:serial",
"test:unit:ci:shard": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=4096 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=4 --test-shard=$TEST_SHARD tests/unit/*.test.ts \"tests/unit/{api,auth,authz,build,cli,cli-helper,combo,compression,correctness,cors,db,db-adapters,docs,gamification,guardrails,lib,mcp,memory,runtime,security,services,settings,shared,ui,usage}/**/*.test.ts\" \"tests/unit/**/*.test.mjs\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=4096 --import tsx --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=4 --test-shard=$TEST_SHARD \"tests/unit/dashboard/**/*.test.ts\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=4096 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=1 --test-shard=$TEST_SHARD \"tests/unit/serial/**/*.test.ts\"",
"test:unit:fast": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-isolation=none tests/unit/*.test.ts \"tests/unit/{api,auth,authz,build,cli,cli-helper,combo,compression,correctness,cors,db,db-adapters,docs,gamification,guardrails,lib,mcp,memory,runtime,security,services,settings,shared,ui,usage}/**/*.test.ts\" \"tests/unit/**/*.test.mjs\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-isolation=none \"tests/unit/dashboard/**/*.test.ts\" && npm run test:unit:serial",
"test:scoped": "bash scripts/quality/test-scoped.sh",
"test:scoped:staged": "bash scripts/quality/test-scoped.sh --staged",
"test:unit:shard": "concurrently --kill-others-on-fail -n s1,s2 \"npm:test:unit:shard:1\" \"npm:test:unit:shard:2\"",
"test:unit:shard:1": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=10 --test-shard=1/2 tests/unit/*.test.ts \"tests/unit/{api,auth,authz,build,cli,cli-helper,combo,compression,correctness,cors,db,db-adapters,docs,gamification,guardrails,lib,mcp,memory,runtime,security,services,settings,shared,ui,usage}/**/*.test.ts\" \"tests/unit/**/*.test.mjs\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=10 --test-shard=1/2 \"tests/unit/dashboard/**/*.test.ts\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=1 --test-shard=1/2 \"tests/unit/serial/**/*.test.ts\"",
"test:unit:shard:2": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=10 --test-shard=2/2 tests/unit/*.test.ts \"tests/unit/{api,auth,authz,build,cli,cli-helper,combo,compression,correctness,cors,db,db-adapters,docs,gamification,guardrails,lib,mcp,memory,runtime,security,services,settings,shared,ui,usage}/**/*.test.ts\" \"tests/unit/**/*.test.mjs\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=10 --test-shard=2/2 \"tests/unit/dashboard/**/*.test.ts\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --max-old-space-size=8192 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=1 --test-shard=2/2 \"tests/unit/serial/**/*.test.ts\"",
@@ -150,7 +143,6 @@
"check:node-runtime": "node --import tsx scripts/check/check-supported-node-runtime.ts",
"check:pack-artifact": "node --import tsx scripts/build/validate-pack-artifact.ts",
"check:pack-boot": "node scripts/check/check-pack-boot.mjs",
"check:install-upgrade": "node scripts/check/check-install-upgrade.mjs",
"check:pack-policy": "node --import tsx scripts/build/validate-pack-artifact.ts --policy-only",
"check:cli-i18n": "node scripts/check/check-cli-i18n.mjs",
"check:openapi-coverage": "node scripts/check/check-openapi-coverage.mjs",
@@ -169,7 +161,6 @@
"check:test-masking": "node scripts/check/check-test-masking.mjs",
"check:test-runner-api": "node scripts/check/check-test-runner-api.mjs",
"check:changelog-integrity": "node scripts/check/check-changelog-integrity.mjs",
"sweep:stale-fragments": "node scripts/release/sweep-stale-fragments.mjs",
"changelog:aggregate": "node scripts/release/aggregate-changelog.mjs",
"check:agent-skills-sync": "node --import tsx/esm scripts/skills/generate-agent-skills.mjs",
"check:build-scope": "node scripts/check/check-build-scope.mjs",
@@ -193,7 +184,6 @@
"check:bundle-size": "node scripts/check/check-bundle-size.mjs",
"check:circular-deps": "node scripts/check/check-circular-deps.mjs",
"check:mutation-ratchet": "node scripts/check/check-mutation-ratchet.mjs",
"check:rtl-ratchet": "node scripts/check/check-rtl-ratchet.mjs",
"check:licenses": "node scripts/check/check-licenses.mjs",
"check:pr-evidence": "node scripts/check/check-pr-evidence.mjs",
"check:vuln-ratchet": "node scripts/check/check-vuln-ratchet.mjs",
@@ -210,11 +200,9 @@
"typecheck:core": "tsc --pretty false -p tsconfig.typecheck-core.json",
"typecheck:noimplicit:core": "tsc --pretty false -p tsconfig.typecheck-noimplicit-core.json",
"check:dashboard-typecheck": "node scripts/check/check-dashboard-typecheck.mjs",
"check:open-sse-typecheck": "node scripts/check/check-open-sse-typecheck.mjs",
"backfill-aggregation": "node --import tsx src/scripts/backfillAggregation.ts",
"env:sync": "node scripts/dev/sync-env.mjs",
"test:integration": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=1 tests/integration/*.test.ts \"tests/integration/combo-matrix/*.test.ts\"",
"test:integration:ci": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=1 --test-shard=$TEST_SHARD tests/integration/*.test.ts \"tests/integration/combo-matrix/*.test.ts\"",
"test:combo:matrix": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=1 \"tests/integration/combo-matrix/*.test.ts\"",
"test:combo:live": "cross-env RUN_COMBO_LIVE=1 DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=1 \"tests/integration/combo-live/*.live.test.ts\"",
"test:combo:live:vps": "node scripts/test/combo-live-vps.mjs",
@@ -225,7 +213,7 @@
"test:e2e": "node scripts/dev/run-playwright-tests.mjs test tests/e2e/*.spec.ts",
"test:protocols:e2e": "node scripts/dev/run-protocol-clients-tests.mjs",
"test:vitest": "vitest run --config vitest.mcp.config.ts",
"test:vitest:ui": "vitest run --config vitest.config.ts tests/unit/ui",
"test:vitest:ui": "vitest run --config vitest.config.ts",
"test:mutation": "stryker run",
"test:ecosystem": "node scripts/dev/run-ecosystem-tests.mjs",
"test:system": "cross-env DISABLE_SQLITE_AUTO_BACKUP=true node --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=1 tests/e2e/system-failover.test.ts",
@@ -244,7 +232,6 @@
"prepare": "husky",
"system-info": "node scripts/dev/system-info.mjs",
"build:cli-api": "node --import tsx/esm scripts/cli/generate-api-commands.mjs",
"postbuild": "node scripts/build/colocate-standalone.mjs",
"release:contributors": "node scripts/release/gen-contributors.mjs",
"release:uncovered": "node scripts/release/list-uncovered-commits.mjs",
"test:coverage:runner": "node --max-old-space-size=8192 --import tsx/esm --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=8 tests/unit/*.test.ts \"tests/unit/{api,auth,authz,build,cli,cli-helper,combo,compression,correctness,cors,db,db-adapters,docs,gamification,guardrails,lib,mcp,memory,runtime,security,services,settings,shared,ui,usage}/**/*.test.ts\" \"tests/unit/**/*.test.mjs\" && cross-env DISABLE_SQLITE_AUTO_BACKUP=true NODE_OPTIONS=--max-old-space-size=8192 c8 --merge-async --output-dir coverage --exclude=tests/** --exclude=**/*.test.* --reporter=text-summary --reporter=html --reporter=json-summary --reporter=lcov --check-coverage --statements 60 --lines 60 --functions 60 --branches 60 node --max-old-space-size=8192 --import tsx --import ./open-sse/utils/setupPolyfill.ts --import ./tests/_setup/isolateDataDir.ts --test --test-force-exit --test-concurrency=8 \"tests/unit/dashboard/**/*.test.ts\" && npm run test:unit:serial",
@@ -268,6 +255,7 @@
"bottleneck": "^2.19.5",
"clsx": "^2.1.1",
"commander": "^15.0.0",
"cron-parser": "^5.6.2",
"csv-stringify": "^6.7.0",
"dompurify": "^3.4.13",
"express": "^5.2.1",
@@ -311,13 +299,14 @@
"recharts": "^3.8.1",
"safe-regex": "^2.1.1",
"selfsigned": "^5.5.0",
"sharp": "^0.35.3",
"smol-toml": "1.7.1",
"socks": "^2.8.7",
"sql.js": "^1.14.1",
"sqlite-vec": "^0.1.9",
"tailwind-merge": "^3.6.0",
"tsx": "^4.23.0",
"undici": "^8.10.0",
"undici": "^8.3.0",
"update-notifier": "^7.3.1",
"uuid": "^14.0.0",
"ws": "^8.18.0",
@@ -330,7 +319,7 @@
"@atjsh/llmlingua-2": "2.0.3",
"@huggingface/transformers": "3.5.2",
"@tensorflow/tfjs": "4.22.0",
"better-sqlite3": "^13.0.2",
"better-sqlite3": "^13.0.1",
"js-tiktoken": "^1.0.20",
"keytar": "^7.9.0",
"tls-client-node": "^0.2.0",
@@ -376,7 +365,6 @@
"lint-staged": "^17.0.8",
"lockfile-lint": "^5.0.0",
"node-loader": "^2.1.0",
"opencode-ai": "1.18.8",
"playwright-ctrf-json-reporter": "^0.0.29",
"prettier": "^3.8.3",
"promptfoo": "^0.121.18",
@@ -408,15 +396,6 @@
"sharp"
]
},
"allowScripts": {
"better-sqlite3": true,
"esbuild": true,
"@swc/core": true,
"@parcel/watcher": true,
"keytar": true,
"protobufjs": true,
"unrs-resolver": true
},
"overrides": {
"fast-xml-parser": "^5.10.1",
"sharp": "^0.35.0",
@@ -447,14 +426,27 @@
"adm-zip": "^0.6.0",
"promptfoo": {
"js-yaml": "^5.2.2",
"@apidevtools/json-schema-ref-parser": {
"js-yaml": "^4.3.1"
},
"undici": "^7.29.0"
},
"socket.io-parser": "^4.2.7",
"tar": "^7.5.21",
"nanoid": "^3.3.17",
"brace-expansion": "^5.0.9",
"minimatch": {
"brace-expansion": "^1.1.18"
},
"libxmljs2": {
"minimatch": {
"brace-expansion": "^2.1.4"
}
},
"rimraf": {
"minimatch": {
"brace-expansion": "^2.1.4"
}
},
"@apidevtools/json-schema-ref-parser": {
"js-yaml": "^4.3.1"
},
"@eslint/eslintrc": {
"js-yaml": "^4.3.1"
},
@@ -464,11 +456,9 @@
"xmlbuilder2": {
"js-yaml": "^4.3.1"
},
"nanoid": "^3.3.17",
"monaco-editor": {
"dompurify": "^3.4.13"
},
"@apidevtools/json-schema-ref-parser": {
"js-yaml": "^4.3.1"
}
}
}

Some files were not shown because too many files have changed in this diff Show More