Commit Graph

8945 Commits

Author SHA1 Message Date
Tiangao
fda9ef78b1 feat(proxylogs): show registry proxy name in proxy log columns (#12814)
* feat(proxylogs): show registry proxy name in proxy log columns

Registry resolution already attaches name to the runtime proxy object; the
name was dropped at the persistence boundary (ProxyInfo had no name field)
and never rendered. Add proxy_name column (base schema + ALTER heal for
existing DBs), persist/hydrate it, render it in the ProxyLogger table and
ProxyLogDetail pane with host:port fallback, and search by name.

Local-only (PMO City): not submitted upstream. Re-apply after upgrades via
patch file (see pmo-city-builds omniroute/Operator/runbooks/upgrade.md).

* test(proxylogs): flush batched writes before asserting persisted row

The v3.8.50 rebase kept upstream's batched proxy-log persistence
(enqueueProxyLogs/flushProxyLogsSync); logProxyEvent no longer writes
synchronously, so the persist+hydrate test closed the DB before the row
was flushed. Flush explicitly first.

* test(proxylogs): drain batched queue in resetStorage to avoid cross-test row bleed

* docs(changelog): add changelog fragment for proxy registry name in proxy logs

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Tiangao (hermes) <montigaud@aikumi.pro>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 13:14:14 -03:00
Abhishek Sharma
95b2e53727 feat(usage): generic billing/quota for openai-compatible connections (#13673)
* feat(usage): generic billing/quota for openai-compatible connections (#13616)

Every other fetcher in services/usage hard-codes one upstream's URL, auth
and response shape, which works because those providers are known. An
openai-compatible connection can point at anything, and its id is minted
per connection -- so it can never be a member of USAGE_SUPPORTED_PROVIDERS
or a case in the dispatcher switch.

So the shape comes from the connection instead. `providerSpecificData.
quotaEndpoint` declares the url, auth mode, optional headers, and a mapping
of dot-paths onto UsageQuota:

    { "url": "...", "auth": "bearer",
      "quotas": { "credits": { "used": "$.data.used_usd",
                               "total": "$.data.limit_usd",
                               "currency": "USD" } } }

Dot/bracket paths (`$.a.b[0].c`) rather than full JSONPath, so the mapping
stays dependency-free and legible in a config field.

Three decisions worth stating:

- **An unresolvable mapping reports nothing, never 0/0.** A quota reading
  0 of 0 renders as fully exhausted, and an operator would act on that. A
  typo'd path must produce no card, not a fake outage.
- **The transport error is not echoed.** The url is operator-supplied and
  can carry a query-string secret; the message says "unreachable" and
  nothing more.
- **The capability is read off the connection, not the id.**
  `supportsProviderQuota` already takes the connection and already has a
  connection-shaped check (moonshot), so the gate goes there. A declared
  url with no `quotas` mapping does NOT count as supported: it can be
  fetched but can never yield a quota, and would leave a permanently empty
  card in Provider Limits.

Verified: 7 new tests; mutations each killed by the right one --
  let an unresolved mapping fall through to 0/0 -> that test alone fails
  echo the transport error                       -> that test alone fails
121 tests pass across this file, usage-families-split, provider-plugin-
manifest, provider-limits* and the quota-visibility suites (the drift
guard from #13134 included). eslint clean on all four files; the three
no-unused-vars errors in usage.ts are byte-identical on the base branch.

* fix(usage): bound the openai-compatible quota fetch with a timeout

An operator-configured quotaEndpoint that never responds would hang
getOpenAiCompatibleUsage()'s fetch() indefinitely, stalling that
connection's Provider Limits sync. Same 15s bound as the other
fetchers in this directory (grokResetCredits.ts's FETCH_TIMEOUT_MS).

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* test(usage): pin the openai-compatible quota fetch timeout

A quota endpoint that accepts the connection and never answers must be
aborted by the fetch signal instead of hanging the Provider Limits sync.
Fails without the AbortSignal.timeout() bound, passes with it.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: abhisheksharma2411 <abhisheksharma2411@users.noreply.github.com>
2026-09-18 13:14:06 -03:00
STAVAN SHAMUVEL WADEKAR
c74cea3d35 feat(mitm): dynamically inject configured models into Antigravity model catalog (#14006)
* feat(mitm): dynamically inject configured models into Antigravity model catalog

- Add /v1internal:fetchAvailableModels to ANTIGRAVITY_TARGET.endpointPatterns in src/mitm/targets/antigravity.ts
- Implement catalog interception and dynamic model merging in AntigravityHandler.intercept() (src/mitm/handlers/antigravity.ts)
- Merge operator's configured combos/models dynamically from the repository into Google Cloud Code's upstream catalog
- Prepend injected models to agentModelSorts recommended group while preserving native models and upstream structure
- Add unit tests covering target endpoint pattern declaration, catalog merging, dynamic combo retrieval, and error propagation in tests/unit/mitm-handler-antigravity.test.ts

Resolves #13959

* test(mitm): isolate DATA_DIR and clean up the combo row in the antigravity catalog test

The DB-backed test ("dynamic catalog pulls configured combos from database
repository") creates a real combo row via src/lib/db/combos.ts, whose module-
level DATA_DIR const resolves once at import time. The PR's own documented
Validation command (`node --import tsx/esm tests/unit/mitm-handler-antigravity.test.ts`)
runs without the `--test` flag, so the existing #10428 eval-probe/test-context
guard in resolveWritableDataDir() never triggers and DATA_DIR falls through to
the real ~/.omniroute home database — writing a permanent test-combo row into
it every run.

Set DATA_DIR to an isolated temp dir at the top of the file (before the
combos.ts import), reset the DB singleton and clean up the temp dir in
test.after(), and wrap the combo creation in try/finally so the created row is
deleted even on assertion failure.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(mitm): prevent model name collision and filter inactive combos in antigravity catalog

* feat(antigravity): integrate native auto groups, fallback to groq, and add bridge proxy

- Inject OmniRoute native auto groups (auto/best-fast, auto/best-coding, auto/best-reasoning, auto/best-free, etc.) into Antigravity IDE & CLI /model selector
- Add bin/antigravity-bridge.mjs with selective proxy routing to isolate native Gemini quota (zero Google token leakage)
- Implement transparent self-healing model remapping to prevent upstream 410 model_shutdown errors on deprecated models
- Update emergencyFallback provider from nvidia to groq/openai/gpt-oss-120b for resilient 0.02s failover

* test(antigravity): add unit test suite for antigravity bridge routing and model self-healing

- Add tests/unit/antigravity-bridge-routing.test.ts covering zero quota leakage for native Gemini models
- Validate OmniRoute auto group routing and display name interception
- Validate retired upstream model self-healing (preventing HTTP 410 crashes)
- Export helper methods from bin/antigravity-bridge.mjs with isMain guard

---------

Co-authored-by: Stavan <stavan794@gmail>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: steve25060 <steve25060@users.noreply.github.com>
2026-09-18 13:13:59 -03:00
dmlanday
d600eba35b fix(cli): preflight the port before serving so a second instance cannot de-register the first (#12485)
* fix(cli): preflight the port before serving so a second instance cannot de-register the first

Starting `omniroute serve` against a port another OmniRoute already owns
produced three identical raw Node stack traces and no explanation:

    Error: listen EADDRINUSE: address already in use 0.0.0.0:20128

The conflict was handed to the child process, so it surfaced only after the
child had been spawned and retried twice on the supervisor's restart budget,
and never named the process holding the port.

The damage was worse than the noise. Both spawns happen after
writePidFile("supervisor") and the failed child's cleanupPidFile("server"), so
a doomed second instance overwrites the pid files of the healthy instance that
owns the port: supervisor/.pid ends up pointing at the dead starter and
server/.pid is deleted, de-registering a server that is up and serving.
Observed live: healthy server 19348 under supervisor 11108, while
supervisor/.pid read 21440 (dead) and server/.pid was gone. `omniroute stop`
still worked, but only by falling through to its killByPort port fallback.

serve now resolves the port owner before spawning anything or touching a pid
file, and exits with a message naming the owning PID and the two ways out
(`omniroute stop`, or `serve --port <other>`). Discovery lives in
findListeningPids() in bin/cli/utils/pid.mjs (netstat on win32, lsof
elsewhere); it mirrors killByPort()'s discovery in stop.mjs, which is worth
consolidating next time that file is touched. A discovery failure reports the
port as free, since a false "busy" would block a legitimate start.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CJf2dxEpiwZqyZujWk57T2

* chore(changelog): link the port-preflight fix to PR 12485

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CJf2dxEpiwZqyZujWk57T2

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: dmlanday <dmlanday@users.noreply.github.com>
2026-09-18 13:13:49 -03:00
Davide Baraldo
1b2349de22 feat(sse): Claude OAuth lower-priority lane + weekly session-limit reset (#13074)
* feat(sse): Claude OAuth lower-priority lane + weekly session-limit reset

Mirror Claude Code's /low-priority and /limit-reset for OmniRoute-managed
Claude subscription accounts (wire contract captured from Claude Code 2.1.263).

Both are opt-in per connection (providerSpecificData.lowPriorityMode /
autoLimitReset, Edit connection -> Claude section, default off) and only act
on the 5-hour usage wall: a 429 carrying
anthropic-ratelimit-unified-status: rejected and, when eligible,
anthropic-ratelimit-unified-slow-offer: treatment. Nothing is sent before
that first wall 429.

- Lower-priority lane: on the wall the executor retries the SAME account
  with `anthropic-usage-limit: slow` and keeps the header on every request
  until anthropic-ratelimit-unified-reset (+60s). The intercepted 429 never
  reaches chatCore, so the connection is not cooled down or rotated away.
  slot_busy (429) / 529 wait slow-retry-after (20s default, 5-600s, +-30%
  jitter) bounded by slow-max-wait (20min default, 1min-6h), then end +
  10min cool-off. weekly_limit / budget_exhausted / off / ineligible, a
  5h-window rollover, or ineligible + overage-in-use end the lane and let
  the response flow to the normal cooldown path.
- Session-limit reset: GET /api/oauth/usage?at_wall=1&skip_spend=1 ->
  juniper_tide block; when arm=reset and available, POST
  /api/organizations/{org}/reset_rate_limits {program: "juniper_tide"} and
  retry at full speed. already_used / not offered memoise next_available_at.
- State is in-memory per connection; the executor owns the abort-aware
  sleep; the pure state machine and the HTTP client are separate modules
  with unit tests; an executor-level test proves the header/retry wiring
  end to end with a mocked upstream.

* fix(sse): make the Claude usage-wall handling race-safe for parallel requests

Two requests on the same Claude OAuth connection can hit the 5-hour wall in
the same instant.

- Lower-priority lane: the executor now tells the decider whether THIS
  request carried `anthropic-usage-limit: slow`. A sibling built while the
  lane was still idle (no header) whose 429 lands after the lane activated
  is re-sent on the lane instead of being misread as a "wall" verdict that
  would end it; its 2xx is not counted as lane telemetry either.
- Session-limit reset: concurrent wall hits share one in-flight status+claim
  round trip (no duplicate POST reset_rate_limits), and for 60s after a
  granted reset stale sibling walls are answered "reset" without touching
  the network, so they retry at full speed instead of re-claiming or
  falling into the slow lane.

Tests cover both races.

* fix(sse): address adversarial review of the Claude usage-wall handling

Three defects found by a 3-lens review of the two previous commits.

1. Lane wait could outlive the request (high). The slot_busy/529 sleep shares
   the request's AbortSignal with chatCore's upstream-start timeout (10 min by
   default), while the lane's own max-wait defaults to 20 min and can reach 6h
   from the server header. A long slot_busy streak was therefore killed
   mid-sleep with a TimeoutError instead of ending gracefully as max_wait with
   its cool-off. The decision now takes a waitCeilingMs — what is left of the
   executor's own timeout, minus a 5s margin — which caps the effective
   max-wait and clamps each individual sleep.

2. A wall 429 surfacing only after a 400-driven intra-attempt retry was missed
   (medium). The context-editing / thinking-budget / effort / auto-learn
   fallbacks all re-fetch and REASSIGN `response`, and the wall check ran
   before them, so such a 429 fell through to the generic path and cooled the
   connection down. The check now runs after those retries, on the final
   response of the attempt.

3. `ineligible` + `overage-in-use: true` ended the lane as plain `ineligible`
   on a 429 (medium) because the status mapping ran first; only the non-429
   tail produced `extra_usage`. Overage takeover now wins on every status.

Also bounds the module-level per-connection maps with the same FIFO policy as
the identity caches in claudeIdentity.ts: the state key falls back to the
access token when a connection id is absent, and OAuth tokens rotate on every
refresh, so the maps could grow for the process lifetime.

Tests cover all three fixes, including an executor-level regression for the
400-then-wall ordering.

* fix(i18n): add the Claude usage-wall toggle strings to pt-BR

`tests/unit/i18n-pt-br.test.ts` (#6695) requires pt-BR.json to carry every key
present in en.json; the four new `providers.claude{LowPriorityMode,AutoLimitReset}*`
keys were only added to en and it, so the gate failed on this branch.

* refactor(sse): keep the usage-wall change inside the frozen quality budgets

The three ratchets this PR tripped were all its own, not inherited:

- file-size (frozen, may only shrink): open-sse/executors/base.ts 1857 > 1751
  and EditConnectionModal.tsx 1653 > 1631.
- complexity / cognitive-complexity (new-code mode): three functions over the
  15 threshold — runClaudeLimitResetAttempt (27), handleClaudeUsageLimitResponse
  (19 / cognitive 23) and observeClaudeLowPriorityResponse (17 / 17).

Extractions, all behavior-preserving:

- New open-sse/executors/claudeUsageLimit.ts owns the executor-side glue (header
  injection, wait accounting, abort-aware sleep, timeout-derived wait ceiling and
  the decision logging) behind a ClaudeUsageLimitGuard, so base.ts keeps a
  three-line call site instead of ~100 lines of mechanics.
- Three long-standing Claude blocks leave base.ts for the modules they belong to:
  mergeCcHeaders + applyStainlessHeaders into config/anthropicHeaders.ts and
  stripClaudeSystemPrefixBlocks into executors/claudeIdentity.ts. base.ts is back
  at its frozen 1750 lines.
- The modal's Claude section becomes ClaudeConnectionFields.tsx (mirroring
  CcCompatibleRequestDefaultsFields) plus a claudeConnectionFields.ts helper that
  de-duplicates the field defaults across the modal's two init sites; the file
  drops to 1622, below its frozen 1631.
- The three over-threshold functions are split into focused helpers
  (observeErrorResponse / observeSuccessResponse, shouldClaimLimitReset,
  resolveLimitResetOffer / runLimitResetClaim / memoiseNotBefore).

Gates now: file-size OK, complexity 0 new violations, cognitive 0 new,
fetch-targets / error-helper / build-scope / deps OK, typecheck clean, ESLint 0,
Prettier clean, 155 unit tests green across the feature and its neighbours.

Still failing and NOT this branch's: pack-policy (unexpected
@omniroute/opencode-plugin-v2 files in the npm artifact) and
mutation-test-coverage (stryker tap.testFiles missing entries for
circuitBreaker.ts and comboStructure.ts) — both reproduce on the untouched base.

---------

Co-authored-by: davidebaraldo <davidebaraldo@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 13:13:41 -03:00
Sean Ford
0fbd8854c5 fix(api): adapt TEI/Infinity request and response shapes on the /v1/rerank node path (#13733)
* feat(api): route /v1/rerank to remote provider nodes behind RERANK_REMOTE_PROVIDER_NODES

POST /v1/rerank only ever dispatched to provider nodes whose base URL hostname
was localhost, 127.0.0.1, or 172.16.0.0/12 — a filter hardcoded in the route.
A rerank node on any other host (a LAN box or Tailscale peer running TEI,
Infinity, vLLM, …) was silently dropped and the request fell through to
"Invalid rerank model", even though the same node served /v1/embeddings
without complaint and had already passed the provider outbound URL policy at
creation time. The memory engine's rerank step calls this route over loopback,
so `rerankProviderModel` could not reach such a node either.

Mirror the audio routes (#3963): loopback nodes stay always-eligible and
unchanged; remote nodes are opt-in via a new `RERANK_REMOTE_PROVIDER_NODES`
feature flag (default off — routing to a remote host changes egress identity)
AND must pass the provider outbound URL policy (`getProviderOutboundGuard()`,
`public-only` deployments never route to private hosts.

- src/shared/network/loopbackNodeHost.ts: one pure definition of the
  loopback host set, replacing three copies (rerank route, audioRegistry,
  localHealthCheck). The shared version also rejects `user@host` URLs, which
  the audio copy did not.
- src/shared/network/providerNodeHost.ts: policy-aware remote-node
  eligibility that mirrors guardProviderNodeBaseUrl() on the creation path.
- src/app/api/v1/_shared/rerankProviderNodes.ts: pure, testable selection
  step + loader, modelled on audioProviderNodes.ts.
- Feature flag definition, FEATURE_FLAGS.md / ENVIRONMENT.md / .env.example
  rows, API_REFERENCE.md and MEMORY.md notes, changelog fragment.
- tests/unit/rerank-remote-provider-nodes.test.ts covers the host
  classification, the three policy modes, the selection step, and the route
  end-to-end (flag off → 400 without contacting the node; flag on → forwarded
  to <base>/v1/rerank with the node credential; flag on + strict policy →
  still excluded). Feature-flag count test bumped to 56.

* chore(changelog): name the #13732 fragment

* fix(api): adapt TEI/Infinity request and response shapes on the /v1/rerank node path

The provider-node branch of POST /v1/rerank already fell back from
<base>/v1/rerank to <base>/rerank on 404 "for Infinity / TEI", but it kept
sending the Cohere body and returned the upstream JSON verbatim. Against
Hugging Face text-embeddings-inference that could never work: TEI requires
the candidate list as `texts` (HTTP 422 otherwise), takes `return_text`,
and answers a bare `[{index, score, text?}]` with no `results` envelope and
`score` instead of `relevance_score`. Thin gateways in front of TEI/Infinity
commonly emit `score` too. Either way the memory engine's applyRerank(),
which reads `results[].relevance_score`, ended up with undefined scores.

Add two pure adapters in src/app/api/v1/_shared/rerankLocalNodeShapes.ts:

- buildLocalRerankRequestBody(): one upstream body carrying both spellings
  (`documents` + `texts`, `return_documents` + `return_text`). TEI's request
  struct is not deny_unknown_fields and the OpenAI-shaped servers (vLLM,
  llama.cpp, Infinity, oMLX) ignore extras, so a single body serves all.
- normalizeLocalRerankResponse(): folds `{results:[…]}`, Voyage-style
  `{data:[…]}`, and TEI's bare array into the Cohere envelope, backfilling
  `relevance_score` from `score`, sorting by score, honouring `top_n`,
  attaching `document.text` when requested, dropping malformed entries, and
  preserving other top-level fields (`model`, `usage`, …).

The route now uses both on the primary and fallback fetch. Cloud registry
providers are untouched (they go through open-sse/handlers/rerank.ts).

tests/unit/rerank-local-node-shapes.test.ts covers the adapters and the
route end-to-end: 404 → /rerank with `texts`, bare TEI array normalized and
top_n-capped; `score`-only gateway → `relevance_score` for the client.

* chore(changelog): name the #13733 fragment

* refactor(api): split the local rerank response normalizer into per-entry helpers

The complexity ratchet (new-code mode) flagged normalizeLocalRerankResponse
at 18/15 on both metrics. Pull the per-entry validation and the document
resolution into toCohereResult() / resolveResultDocument(); behaviour and
tests are unchanged.

---------

Co-authored-by: seanford <seanford@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 13:13:32 -03:00
Marco
55b6ca6573 fix(compression): fall back in-process when the compression worker fa… (#13637)
* fix(compression): fall back in-process when the compression worker fails (#13145)

The worker pool resolved every worker fault with the *uncompressed* body instead
of reporting it. `PendingJob` had no reject path at all, so a thread error, a
worker exit, a dispatch timeout, or an engine error posted back as
`type: "error"` all resolved as `{ compressed: false, stats: null }`.

`applyCompressionAsync` then treated that as a legitimate "nothing to compress"
result and returned it as-is, so the request reached the provider uncompressed
while the response header still announced the selected plan ("stacked") — the
header is emitted before the pipeline runs. Nothing was logged at any level, and
`compression_analytics` stayed empty because rows are only written when a
compressed result is reported. The net effect was compression silently disabled
for every worker-eligible request.

The worker is a throughput optimisation, not a behavioural variant, so a worker
fault must degrade to the in-process pipeline rather than to no compression:

- `PendingJob` gains `reject`; `fail()` delegates to a new `abort()` that clears
  the slot timeout and rejects with a diagnostic cause (thread error, exit code,
  or timeout budget).
- An `error` message from the worker is propagated instead of being swallowed.
- `applyCompressionAsync` catches the rejection and falls through to the
  in-process path, logging the cause. The logger is imported lazily and
  defensively: `compressionWorker.ts` imports this module, so a static import
  would pull the logger into the worker bundle, and a logging failure must never
  be able to break compression itself.

`close()` keeps resolving with the unchanged body — shutdown is not a fault.

The regression test drives a real worker fault via
`OMNI_COMPRESSION_WORKER_TIMEOUT_MS` rather than mocking the module, since this
project's tsx/ESM + node:test setup has no `mock.module()` support. Its options
are fully populated on purpose: `runCompressionAsync` forwards them into
`workerOptions`, and `isStrictlySerializable` rejects an object holding
`undefined` values — which would route the test through the in-process path and
assert nothing. Production requests always carry all of those fields, which is
why the worker path is taken there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LGJQT3E6iJZq4zNGwfjkPs

* fix(compression): keep the timeout path uncompressed, retry only fast worker faults (#13145)

Review follow-up: the in-process fallthrough ran the full pipeline on the main
event loop for *every* worker fault, including a dispatch timeout. A timeout
means the worker already spent its whole budget on that body, so re-running the
same CPU-bound work inline would stall other in-flight requests — strictly worse
than not compressing on a shared gateway.

Faults are now typed by whether recovery is cheap:

- `CompressionWorkerError.retryInProcess` distinguishes fast faults (thread
  error, worker exit, engine throw — no work was done, so the in-process path
  costs what the worker would have) from a dispatch timeout.
- Timeouts keep the original degrade-to-uncompressed behaviour, but are now
  reported. The defect this PR fixes is the silent swallow, not the degrade.

Also strips `reject` from the structured-clone wire job. It is a function, so
leaving it on the object handed to `postMessage` threw `DataCloneError` before
the worker ever saw the job — turning every dispatch into an immediate fault.

Adds the missing `changelog.d/fixes/` fragment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(compression): narrow the worker thread-error type for typecheck:core

@types/node 26 types the Worker "error" event payload as unknown, not
Error, so `error?.message` failed typecheck:core (TS2339). Narrow with
an instanceof check before reading .message.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: marcs7 <marcs7@users.noreply.github.com>
2026-09-18 13:13:25 -03:00
birdleandro-bit
1e8c913ca2 feat(services): add open-wa as a 6th embedded service (#13222)
* feat(services): add open-wa as a 6th embedded service

Adds @open-wa/wa-automate (WhatsApp Web automation via headless
Chromium) following the existing embedded-service pattern, mirroring
Mux's lifecycle-managed-only shape (no Layer 4 executor — this is not
a routing target). Flags/env verified directly against the installed
4.76.0 source rather than trusted from web docs, which mix this
stable v4 line with an unreleased v5 alpha CLI surface.

healthIntervalMs is set to 60s (vs. the usual 5s) for this service:
open-wa's HTTP server does not start listening until the full
WhatsApp handshake resolves, which blocks on a human scanning the
pairing QR code on first pairing. At the default 5s interval the
supervisor's 3-consecutive-failure threshold would declare "error"
~15s into every legitimate start. This is a local, single-service
config change — a proper fix (a startup-grace knob distinct from the
steady-state poll interval) belongs in ServiceSupervisor/HealthChecker
as a follow-up affecting all embedded services.

* fix(db): renumber open-wa seed migration to avoid collision with 163

Migration 163 was already taken by 163_radar_feed_cache_generated_at.sql
on release/v3.8.51 (tip is at 179). Renumbered to 180, the next free slot.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(db): renumber open-wa seed migration to avoid collision with 180

Slot 180 was reused by 180_memory_fts_au_conditional_memory_id.sql
(merged 2026-09-16), so this PR's seed migration moves to the owner's
assigned slot 185.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: birdleandro-bit <birdleandro-bit@users.noreply.github.com>
2026-09-18 13:13:16 -03:00
Beexly
706dc75c13 fix(chatCore): stop executeWithUpstreamStartTimeout leaking its abortPromise listener (hedge-cancelled process exit) (#12406)
* fix(sse): stop mergeAbortSignals from leaking abort listeners

mergeAbortSignals() attached "abort" listeners to its primary/secondary
signals but never removed them once the merged signal settled. Every
executor fetch attempt calls this (fetchWithStartTimeout, once per
URL/retry), so a busy combo request accumulated one live listener per
call on the long-lived combo/client signal. A leaked listener still
fires when that signal is later aborted (e.g. a hedge cancellation
arriving after this merge's own caller already finished), for a merged
output nothing is watching anymore.

Mirrors the already-correct self-cleaning pattern in
open-sse/utils/directResponseStartTimeout.ts's local mergeAbortSignals.

Regression test measures listener growth across repeated merges of the
same long-lived signal: 25 merges leaked exactly 25 listeners pre-fix,
0 post-fix.

(cherry picked from commit 07969655147d3236969b38bfb41280ab4fb52b79)

* fix(server): stop the crash guard re-throwing combo abort reasons

Production crash 2026-08-31 (omniroute.log): on a client disconnect,
handleDisconnect aborted the combo controller and a late abort listener
threw the abort reason on an empty stack:

    Error [AbortError]: hedge-cancelled
        at ... AbortController.abort ... handleDisconnect
    file:///.../src/shared/utils/httpClientAbortGuard.mjs:130  throw err;

isClientAbortError() only knew Node's stream codes and "aborted", so
shouldSwallowUncaught() said false and the guard re-threw, taking the
whole server down.

- Port upstream's AbortError line (name "AbortError" + abort-flavoured
  message) so request_signal_aborted / DOMException aborts are absorbed.
- Add an exact-message match for the combo abort reasons from
  open-sse/services/combo/comboAbortReasons.ts ("hedge-cancelled",
  "combo-per-model-timeout"), name-agnostic because the raw reason is a
  plain Error that only gets name="AbortError" stamped on the way out.
  A losing hedge / stalled target is never a server fault. Inlined so
  this .mjs stays dependency-free for scripts/dev/run-next.mjs.

Tests: port upstream's guard tests, add the exact crash shape, a
child-process replay of the crash (dies pre-fix, survives post-fix), a
genuine-error case that must still crash, and a sync check against
comboAbortReasons.ts. The child-process helper passes a file:// URL, not
a bare path, so the tests run on Windows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 90c9bce8c474b60c37cc63f4421d50feae4c0ad2)

* fix(chatCore): stop executeWithUpstreamStartTimeout leaking its abortPromise listener

Root cause of the 2026-08-31 production exit (Error [AbortError]:
hedge-cancelled), verified by mapping the crash frames in
.build/next/server/chunks/13721.js back to this file:

- The abortPromise abort listener registered on the long-lived client /
  stream signal was never removed in the finally block (only abortListener
  and timeoutAbortListener were), so every executor attempt (and every
  retry) leaked one listener onto that signal.
- Promise.race only subscribes to abortPromise/timeoutPromise once the
  array literal has been evaluated. When execute() threw synchronously the
  race never ran, abortPromise was orphaned, and the next hedge
  cancellation / client disconnect aborted the signal with the string
  reason streamHandler.ts forwards; createAbortError() rebuilt it as an
  AbortError-named Error and rejected a promise nothing awaited. That
  unhandledRejection reached the process crash guard, which re-threw it as
  an uncaughtException and exited with code 7.

Keep a handle to the listener and remove it with the others, and mark the
two race-loser promises as handled so a synchronous throw from execute()
can never orphan them. Race semantics are unchanged (the race still
observes their rejections).

Regression tests: (1) a resolving execute leaves the listener count on
the client signal unchanged; (2) a synchronously throwing execute leaks
no listener and a later abort with the string "hedge-cancelled" produces
no unhandledRejection. Both fail against the previous implementation.

Note: commit 079696551 (mergeAbortSignals cleanup) is correct listener
hygiene but is not on this crash path; this is the fix for the incident.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit e68a50ad854a1945c23e747da4ef15a820cc148d)

* fix(server): absorb raw string abort reasons in the crash guard; document the verified crash path

Follow-ups from the adversarial review of 90c9bce8c:

- open-sse/utils/streamHandler.ts aborts the stream controller with a raw
  string reason (getClientAbortReason / handleDisconnect) and undici
  rejects with signal.reason verbatim, so a cancellation can reach
  process level as a bare string. isClientAbortError() returned false for
  every non-object, which would still have exited the process. Absorb the
  combo abort reasons and the stream-handler disconnect reasons when they
  arrive as strings.
- Correct the mechanism comment: the 2026-08-31 exit was a leaked
  upstreamTimeouts.ts abortPromise listener rejecting a promise nothing
  awaited (unhandledRejection), escalated by this guard, not a listener
  throwing synchronously. The leak is fixed at the source in the previous
  commit; this guard remains the last-resort net.
- Reword the inlining rationale (plain node launcher, no reliance on
  type-stripping for the .ts constants module).
- Tests: the child-process replay now also exercises the
  unhandledRejection route with the exact production error shape and with
  raw string reasons; add unit coverage for string reasons and non-object
  look-alikes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 696fcc8fe1b9b9bc43e6e5f5f5e44b619a86e68a)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: Beexly <Beexly@users.noreply.github.com>
2026-09-18 13:13:06 -03:00
GiauPhan
eb4e3be5dc fix(quota): keep Kiro active while any _freetrial pool has quota (#13088) (#13324)
* fix(quota): keep Kiro active while any _freetrial pool has quota (#13088)

* fix(quota): rename changelog and remove blank line

---------

Co-authored-by: giauphan <giauphan@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 13:12:58 -03:00
Dizzle
0349627c86 fix(opencode): match the upstream free-tier request contract (#14013)
Match the upstream OpenCode free-tier request contract (issue #13935): canonical
ses_/msg_ identity ids, versioned User-Agent, and the measured body requirements
(stream:true + non-empty tools) with a learn-and-reuse tool-name cache, so
no-auth oc/* requests stop being refused with 403 FreeTierError.

Supersedes #13937 (session regex and minimum-version rule kept, credited below).
Complements #14011 (refusal classification) and #13819 (stream_options strip),
both already merged.

Closes #13935

Co-authored-by: AStupidBear <16422976+AStupidBear@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:57:10 -03:00
Prabhjot Singh
2ebafaedce fix(windows): hide supervised server console (#13992)
* fix(windows): hide supervised server console

* fix(windows): port icon.ico fix from #13991 and add regression tests

Adds a source-pattern test asserting the supervised server spawn() passes
windowsHide: true (Hard Rule #8 gap noted in review), and ports the
icon.ico-on-win32 fix from #13991 (credit @prabhtheone) with its own
regression test, so both real fixes ship without #13991's unrelated
comment purge.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:26:26 -03:00
Климентий Горенков
dfcc4baed8 fix(codex): preserve native custom tools in Responses WebSocket requests (#13864)
* fix(codex): preserve native custom tools in Responses WebSocket requests

* test(codex): release per-turn WS leases in the custom-tools passthrough test

The per-account WS lease (release/v3.8.51, added after this branch's fork
point) is non-queued with a default maxConcurrent of 1. This test's bare
prepare() helper never released its lease, and the reused-WebSocket case
re-prepares (acquiring a fresh lease) per turn before releasing the
previous one — both starve the single test connection once merged with
the lease feature. Release after each bare prepare() and raise the test
fixture's maxConcurrent to 2 so a session's sequential turns fit.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:26:17 -03:00
voidstack
e400cf9ac7 fix(db): resolve backup retention from persisted setting on health-check path (#13773)
* fix(db): resolve backup retention from persisted setting on health-check path

#13404 fixed the missing prune call after health-check-repair backups but
only resolved maxFiles/retentionDays from env vars, so the persisted
Storage-page setting (honored for manual/API/auto backups via
getDbBackupMaxFiles/getDbBackupRetentionDays) was silently ignored on this
path. Extract that env->persisted->default precedence into
resolveDbBackupRetention() in backupRetention.ts and share it between
backup.ts and core.ts's createManagedDbBackup().

* docs: add changelog fragment for #13308 persisted-setting follow-up

* fix(db): re-point backup retention fix at managedBackup.ts's prune call

The base drifted since this branch was opened: the health-check-repair backup
path (createManagedDbBackup) moved from core.ts into managedBackup.ts
(writeManagedDbBackup), taking its env-only maxFiles/retentionDays resolution
along with it. This branch's resolveDbBackupRetention() extraction and
backup.ts delegation were already correct and unaffected; only the wiring
that used to live in core.ts needed to move to managedBackup.ts's prune call
so the persisted Storage-page setting is honored on this path too (#13308).

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:26:09 -03:00
小妍儿 ✨
27e0d9b5b7 fix(dashboard): refresh per-connection proxy badges after a proxy save (#13711)
* fix(dashboard): refresh per-connection proxy badges after a proxy save

ProxyConfigModal persists an assignment through
`PUT /api/settings/proxies/assignments`, but the provider page bound its
`onSaved` callback to `fetchProxyConfig()`, which only refetches
`GET /api/settings/proxy` into `proxyConfig`.

The per-account proxy badges read `connProxyMap`, which is filled from a
different endpoint (`GET /api/settings/proxy?resolve=<connectionId>`) by an
effect keyed on `[loading, connections]`. A proxy save changes neither key,
so the effect never re-ran and the saved (or cleared) proxy stayed invisible
until a manual page reload. Account- and combo-level saves refreshed nothing
visible at all; a provider-level save refreshed only the toolbar chip while
the rows that inherit that proxy stayed stale.

Adds `refreshProxyState()`, which re-reads both sources together, and binds
the modal's `onSaved` to it. The callback reads the latest connections from a
ref so it stays referentially stable and does not re-render consumers on every
connections fetch.

Fixes #13710

Also removes the now-unused `no-unused-vars` suppression entry for
useProviderConnections.ts. That entry was already stale on the base branch
(the same eslint invocation reports it on an unmodified tree), and the
pre-commit ratchet refuses to pass while a touched file carries one. Only
that single exact entry was pruned; no baseline was widened.

* docs(changelog): add fragment for #13711

* test(providers): raise proxySaveRefresh test timeout to 30s

The it() case dynamically imports ProviderModalsPanel, the first test in
the repo to pull in that module's ~15 modal components, which alone
consumes most of Vitest's default 5000ms testTimeout and made the test
flake under CI/runner load.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: 千乘妍 (Xiaoyaner) <xiaoyaner0201@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:26:01 -03:00
keii-2596
80ea176022 fix: feat(providers): add kimi/qwen/deepseek/gpt auto-routing families (#13214) (#13709) 2026-09-18 12:25:53 -03:00
Paijo
2374bbf1ee fix(mitm): bound SSE transcript retention and cancel abandoned upstream reads (#13702)
Handler-side collected strings grew without bound before the inspector
clamp; abandoned streams kept the reader alive for the full upstream
lifetime. createBoundedCollector caps retention at 1 MiB while keeping
true responseSize; pipeSSE and server.cjs cancel on downstream close.
Fixes #13395.

Co-authored-by: oyi77 <oyi77@users.noreply.github.com>
2026-09-18 12:25:45 -03:00
Kha Tran
994245a476 fix(translator): prevent schema property name collision in Gemini sanitizer (#13057, #13477) (#13690)
* fix(translator): prevent schema property name collision in Gemini sanitizer (#13057, #13477)

* refactor(translator): reuse SCHEMA_MAP_KEYS in forEachSubschema and add changelog fragment

* test(translator): avoid explicit any in gemini schema collision regression test

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: zcrew0x <zcrew0x@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:25:38 -03:00
dependabot[bot]
44e32a1995 deps: bump electron from 44.0.0 to 44.3.0 in /electron (#13664)
Bumps [electron](https://github.com/electron/electron) from 44.0.0 to 44.3.0.
- [Release notes](https://github.com/electron/electron/releases)
- [Commits](https://github.com/electron/electron/compare/v44.0.0...v44.3.0)

---
updated-dependencies:
- dependency-name: electron
  dependency-version: 44.3.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-18 12:25:30 -03:00
hummern
cc7d19daa9 fix(routing): skip redundant parseAutoPrefix for recognized built-in auto variants (#13647)
* Change hasFree from true to false for pioneer.ai

Pioneer.ai removed the free tier.

Before:

https://web.archive.org/web/20260516140358/https://pioneer.ai/pricing

After:

https://pioneer.ai/pricing

* fix(sse): skip parseAutoPrefix invalid-prefix warning for recognized built-in auto variants

resolveAutoRoutingState() already classifies auto/best-* variants correctly via
classifyAutoModel() before applyAutoPrefix() runs, and the old early-return
preserved that state — so the routing variant was never broken. The real,
observable defect was the spurious 'Invalid auto prefix format' warning logged
on every auto/best-* request, because parseAutoPrefix() only knows the short
aliases (VALID_VARIANTS) and returns valid:false for the best-* built-ins that
AUTO_TEMPLATE_VARIANTS recognizes.

Skip the warning (and the pointless early-return) for any model already present
in AUTO_TEMPLATE_VARIANTS. Add a regression test asserting the warning no
longer fires for auto/best-coding while an genuinely unknown auto/* variant
still warns (proving the log probe detects the message).

* fix(providers): drop out-of-scope Pioneer AI hasFree change from this PR

The Pioneer AI hasFree=false commit (c70e425) leaked into this branch
via a merge and is unrelated/stale vs. the release tip's current
hasFree=true value for this unrelated PR (#13647 is about auto-routing
prefix warnings). Reverting to keep the diff scoped to the actual fix.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Tobias Andersen <turbolego@gmail.com>
Co-authored-by: hummern <hummern@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:25:23 -03:00
Moseyuh333
31ea46934e fix(sse): recognize reasoning_effort in the reactive 400 field-strip retry (#13642)
* fix(sse): recognize reasoning_effort in the reactive 400 field-strip retry

Strict OpenAI-compatible gateways that don't implement the reasoning-effort
knob reject requests with 400 "Unsupported parameter: reasoning_effort".
findOffendingField() did not list it in KNOWN_OFFENDING_FIELDS, so the
generic strip-and-retry in base.ts never fired and the 400 surfaced to the
client — the request died instead of being retried once without the field.

Add "reasoning_effort" to KNOWN_OFFENDING_FIELDS (sibling of the existing
reasoning_budget entry, same FCC/NIM-style recovery) and pin the new match
in provider-field-strips.test.ts.

* chore(changelog): add fix fragment for the reasoning_effort field-strip retry (#13642)
2026-09-18 12:25:15 -03:00
luyuehm
1f8bfe52c5 feat(docker): add self-host compose + 5-minute deploy doc (RIC-739) (#13639)
KISS self-host carrier for the 零月费 + 自托管 product form. One command
brings up the published image + Redis on loopback — no profile choice, no
build step, no multi-tenant anything.

- docker-compose.selfhost.yml: pulls diegosouzapw/omniroute:latest + redis,
  all app ports 127.0.0.1-only by default, Redis not published to host,
  depends_on healthy, healthcheck wired.
- .env.selfhost.example: minimal env (2 EDIT ME lines), no secrets baked in.
- docs/getting-started/SELF_HOST_GUIDE.md: 5-minute deploy, sizing, exposing,
  data/backups, common issues, security checklist, what-it-is-NOT.
- meta.json + DOCKER_GUIDE cross-link.

Graduates to the full docker-compose.yml profiles when the user needs CLI
tools / web-cookie Chromium / sidecars.

Co-authored-by: Ant Rich <ant@richants.com>
2026-09-18 12:25:08 -03:00
Lukas
639d7dac44 docs(resilience): describe the queue-wait and execution deadlines as they are (#13624)
* docs(resilience): describe the queue-wait and execution deadlines as they are

Two docs still describe `requestQueue.maxWaitMs` the way it behaved before
the split, in two different and both incorrect ways:

  RESILIENCE_GUIDE.md  "a legacy persisted name for execution expiration
                        ... bounds limiter-managed execution, not time
                        spent in the local queue ... Queue residence has
                        no time deadline"
  ENVIRONMENT.md       "Max time to wait on a 429 before failing the
                        request"

`rateLimitManager.ts` does the opposite of the first and has nothing to do
with the second. `maxWaitMs` is the queue-wait budget: it covers the slot
wait plus QUEUED residence, and its timer is cleared the moment the job
starts executing (`wrappedFn`). Bottleneck's `expiration` is fed by
`executionMaxWaitMs` (default 600000ms), "never by the queue-wait budget"
in the source's own words, and is raised to the executor's fetch-start
timeout when that is longer.

Issue #13592 is that misreading in practice: the reporter concluded the
deadline "measures execution time after dispatch rather than queue wait"
-- which is what the guide says.

Also corrected while here:

  * The guide's "1-30000ms UI ceiling" is a different setting's bound
    (`comboCooldownWaitSettings`). `requestQueueSettingsSchema.maxWaitMs`
    is `min(1)` with no max, normalised to 1ms-24h.
  * Precedence is documented for the first time: the env var supplies the
    DEFAULT only, a persisted `resilienceSettings.requestQueue` value wins
    over it, and a per-connection `rateLimitOverrides` value wins over
    that. #13592 reports exactly this surprise -- setting
    `RATE_LIMIT_MAX_WAIT_MS` on a deployment that already has a persisted
    value changes nothing.

Documentation only; no behaviour change.

Refs #13592

* docs(changelog): add fragment for the resilience-deadline doc correction
2026-09-18 12:25:00 -03:00
jbovard2016
d574826f37 fix(usage): detach completed request previews (#13623)
* fix(usage): bound completed request retention

Completed request previews used V8 sliced strings that kept multi-megabyte request backing stores alive. Detach and byte-bound cached details, and add a credential-free profiler with cleanup and physical-retention assertions.

* fix(usage): give the JON-562 memory-profile canary realistic timeouts

The 100k-token worker step alone takes ~230s (tsx/esm boot of the full
route/handler module graph plus the real request lifecycle), well past
the driver's hardcoded 180s spawnSync timeout — the resulting SIGKILL
surfaces as `worker.status === null`, indistinguishable from a real
crash. Bump the worker timeout to 300s and the test's own outer/inner
timeouts to match the measured ~230-330s real runtime.

Also make git-branch provenance detached-HEAD safe: `git branch
--show-current` is empty on a detached HEAD (the normal state for a CI
PR checkout, and for this fix worktree itself), which made the canary
throw "git branch is empty" deterministically outside a regular branch
checkout.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:24:52 -03:00
Omkar Prabhu
9b84531e4c fix(cli): pass --legacy-peer-deps to npm install -g in omniroute update (#13579)
* fix(cli): pass --legacy-peer-deps to npm install -g in omniroute update

Problem
-------
Running \
pm install -g omniroute\ prints a wall of ERESOLVE / peer-dependency
warnings because marked-terminal@7.3.0 declares a peer range of marked>=1 <16,
while omniroute ships marked@18. npm's strict peer-resolution mode (the default
since npm 7) flags this mismatch loudly even though the packages work correctly
together at runtime.

Fix
---
Pass --legacy-peer-deps to the npm install -g call that \omniroute update\
issues so that every user who upgrades through the built-in updater gets a
clean, warning-free output. The dry-run log line is updated to match.

Why --legacy-peer-deps is safe here
------------------------------------
The repo already ships .npmrc with legacy-peer-deps=true (added in #11544) so
the published package documents this as its supported install mode. This commit
simply applies the same flag programmatically in the updater so the flag is
always honoured regardless of the caller's local npm config.

Changes
-------
- bin/cli/commands/update.mjs: append --legacy-peer-deps to execSync npm call
  and to the dry-run console.log so output matches the real command
- docs/guides/TROUBLESHOOTING.md: add a supported install snippet and clarify
  that residual deprecation notices come from third-party transitive packages
- tests/unit/cli-update-npm-win32-11335.test.ts: regression test asserting both
  the dry-run string and execSync call carry --legacy-peer-deps
- changelog.d/fixes/: add fragment (number updated after PR is opened)

* chore(changelog): rename fragment to PR #13579
2026-09-18 12:24:43 -03:00
mdigitalbh81
f24c3665df fix(proxy): skip bare TCP health probe for SOCKS5 data plane (#13571)
* fix(proxy): skip bare TCP health probe for SOCKS5 data plane

* docs(changelog): add fragment for SOCKS5 bare TCP probe skip fix

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:24:35 -03:00
Ryan
70eebe9adb fix(arena+analytics): atomic ELO sync (fetch-first) + flatRateAsZero in compression writer (#13446)
* fix(analytics+arena): flatRateAsZero in compression writer; atomic arena sync redesign

* docs(changelog): add fragments for arena ELO sync and compression flat-rate fixes

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: CrashCartCapital <crashcartcapital@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:24:28 -03:00
이창섭
a9c62ba83b fix(playground): improve Compare response scrolling and copy actions (#13317)
* feat(playground): copy individual compare responses

* docs(changelog): add fragment for Compare column copy-to-clipboard

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:24:19 -03:00
easypathuni
4dc73fb36a chore: add Windows helpers to run OmniRoute and Claude Code from a source checkout (#13312)
* chore: add Windows helpers to run OmniRoute and Claude Code from a source checkout

Adds contrib/windows/ with two double-clickable scripts for Windows users who
cloned the repo instead of installing the npm package:

- start-omniroute.bat  -> npm run dev (resolves the repo root from its own path)
- launch-claude.bat    -> node bin/omniroute.mjs launch, forwarding extra args

The README documents a fresh-clone gotcha on Windows with npm >= 11: the
optional better-sqlite3 dependency is silently skipped, the server falls
back to node:sqlite and logs "Module not found: Can't resolve
'better-sqlite3'". Since better-sqlite3@13 ships win32-x64 prebuilds inside
the package, extracting the npm pack into node_modules fixes it without a
compiler. Verified on Windows 11, Node 24.14.0, npm 11.6.1.

* chore(contrib): start Claude Code in the project folder, not the OmniRoute checkout

launch-claude.bat used to cd into the repo root before exec'ing
`omniroute launch`, so Claude Code always started inside the OmniRoute
checkout and loaded this repo's CLAUDE.md/AGENTS.md (~60k chars) into the
user's own coding session. Take the project folder as the first argument
(or prompt for it on double-click) and resolve bin/omniroute.mjs from the
script's own path instead. Remaining args are still forwarded to
`omniroute launch`.
2026-09-18 12:24:11 -03:00
Domenico Massafra
fbe195d69e fix(opencode): preserve catalog display names (#13168)
* fix(opencode): preserve catalog display names

* test(cli): cover OpenCode catalog display-name precedence

Adds the automated unit test the PR body's manual smoke check
(Auto Chat / DeepSeek V4 Pro) was standing in for, covering all four
name-precedence branches: existing custom name, catalog display_name,
native catalog name with owned_by prefix stripped, and the auto/* readable
fallback. Also adds the changelog.d/fixes/ fragment.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: ginettododo <117327638+ginettododo@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:24:03 -03:00
Anh Tran
93be9d544c fix(responses): count input tokens locally for Codex OAuth (#13167)
The ChatGPT subscription backend does not serve
/backend-api/codex/responses/input_tokens for the affected account. Native
requests to that path are intercepted by an OpenAI Cloudflare managed challenge,
while the same path over the bundled Chrome transport returns 404 Not Found.
Forwarding the client preflight can therefore never return a useful count and,
before the companion classifier fix, permanently disabled the healthy Codex
connection on the first challenge.

Add a static /v1/responses/input_tokens route that shadows the generic
Responses passthrough, uses the existing offline o200k_base token counter, and
returns the standard response.input_tokens contract without issuing any upstream
request. Count instructions, structured input, tool definitions and config;
apply a conservative five-percent margin so the failure mode is earlier client
compaction rather than a context-window overflow.

Preserve the API-key and model-policy boundary from the catch-all Responses path.
Tests pin the public schema, prove fetch is never called, cover text,
instructions, structured input, tools, non-text parts, server-held context ids,
invalid JSON, OPTIONS, and the conservative lower bound.

Co-authored-by: anhth2 <anhth2@vng.com.vn>
2026-09-18 12:23:55 -03:00
marioschoenert-code
6614780c36 fix(deps): declare remark-gfm as a direct dependency (#13162)
MarkdownMessage.tsx and other client components import remark-gfm
directly, but it was only present transitively via fumadocs-mdx (a
devDependency). npm's hoisting happens to resolve it, masking the
issue, but pnpm's strict node_modules isolation fails with
"Module not found: Can't resolve 'remark-gfm'" since the package was
never declared as a direct dependency of the app.

Add it explicitly to package.json/package-lock.json and the
supply-chain dependency allowlist.

Co-authored-by: marioschoenert-code <291579285+marioschoenert-code@users.noreply.github.com>
2026-09-18 12:23:48 -03:00
Ozeas Souza
b71f466f1c fix(authz): preserve zed-hosted native-app callback through root middleware redirect (#13140)
* fix(release): let the Electron workflow start again — grant actions:read to the npm leg (#11973)

v3.8.50 shipped with zero desktop assets. The tag push did trigger electron-release.yml
(run 33005490476) but GitHub refused the run at startup:

  Error calling workflow 'npm-publish.yml@5458026'. The nested job 'publish' is
  requesting 'actions: read', but is only allowed 'actions: none'.

npm-publish.yml's `publish` job gained `actions: read` (it downloads the next-build
artefact) and the caller job here never widened its grant — a reusable workflow may not
request more than its caller allows, and the refusal is a startup failure of the WHOLE
run, so the `release` job that attaches the installers, the source archives and the
SBOM never ran either. Nothing about it is visible through the API (no jobs, no
check-runs); only the run page shows the annotation.

- publish-npm: `actions: read` added, with the rule written down (keep the block a
  superset of every job in npm-publish.yml).
- workflow_dispatch: new boolean input `publish_npm` (default true) and the npm leg
  is gated on it, so re-attaching assets to a release whose package already shipped
  does not try to publish the same version twice.
- web-build / build / release checkouts pin `ref: needs.validate.outputs.version`:
  a dispatch builds the tag it names, not the dispatching branch (a tag push resolves
  to the same commit, so nothing changes on the normal path).

actionlint clean; electron-release-desktop-channel-8949, electron-release-efficiency,
build-next-isolated-windows-home-2402, electron-release-latest-yml.repro and
check-workflows suites pass. Next step: dispatch on main with version=v3.8.50 and
publish_npm=false to attach the missing assets.

* fix(ci): stop a stalled Codecov upload from cancelling the Coverage job and the main run (main twin of #11972) (#11978)

Same change as #11972 on release/v3.8.51: the Coverage job had timeout-minutes: 20,
the c8 merge across 8 shards takes ~10 min and the informational Codecov upload hung
for the rest of the budget on two consecutive main runs (33207760653, 33215115341),
ending the job cancelled and turning the run's conclusion cancelled with every
blocking job green. Codecov step: 5-minute ceiling + continue-on-error; job: 30 min.

* fix(release): resync the electron lockfile and let a dispatch build from a repaired ref (#11982)

* fix(release): resync the electron lockfile and let a dispatch build from a repaired ref

The v3.8.50 desktop re-dispatch (run 33238093090) lost its Linux leg at
`npm ci` in electron/: "Missing: electron-builder-squirrel-windows@26.15.3 from
lock file" plus its 12 transitive entries — the optional Windows-installer subtree of
electron-builder had been dropped when the lock was last regenerated, and no CI ran
the desktop legs between then and the tag (v3.8.49 never ran them; v3.8.50 died at
startup, #11973). `npm install --package-lock-only` restores the 13 entries; a clean
`npm ci --ignore-scripts` on the result adds 284 packages with no complaint.

The tag itself carries the broken lock, and the workflow now checks out the tag on
dispatch (#11973), so a dispatch input `build_ref` (default: the version tag) lets the
operator name the repaired line — the v3.8.50 assets will be rebuilt from main, which
is 3.8.50 plus its post-release fixes. Push-triggered runs are unaffected.

actionlint clean; electron-release-desktop-channel-8949, electron-release-efficiency,
electron-release-latest-yml.repro and check-workflows suites pass.

* fix(release): do not regenerate release notes on a re-attach dispatch

`generate_release_notes: true` on an existing release APPENDS GitHub's auto-generated
"What's Changed" block to the curated body — the v3.8.50 re-dispatch (run 33238093090)
added 1,416 chars to the 121 KB notes. Only the tag push should generate notes.

* fix(release): attach the SBOM to the GitHub Release on dispatch publishes too (#12020)

The step was gated on github.event_name == 'release'. v3.8.50's package shipped
through a workflow_dispatch (the staged publish, 11 attempts) and the step was
skipped, so the GitHub Release carried no SBOM — it was attached by hand from the
run's sbom-npm artifact (5.0 MB, 1,886 components). Now it attaches on release or
workflow_dispatch whenever a release for the published tag exists, and says so
when it does not (the workflow artifact remains the durable copy either way).

actionlint and prettier clean; npm-publish-artifact-provenance and
check-workflows-provenance-runner suites pass.

* fix(release): drop the build_ref input — a dispatch builds the ref it is dispatched on (#12032)

Twin of #12022 on main: CodeQL flagged the same input-controlled checkout + npm cache pattern (cache-poisoning/poisonable-step) on main since it's the default branch. Checkouts go back to github.ref; dispatch still works via --ref (documented in the workflow's own on: contract).

Also fixes the packaged-app smoke: it now waits on /api/monitoring/health (which touches the DB) instead of /login (which doesn't), so the smoke can actually distinguish "native driver selected" from "database never opened." electron-smoke-script.test.ts 9/9 (2 new cases).

* fix(ci): accept CVE-2025-68121 in the prebuilt tls-client .so, auto-close base-red issues, guard Scorecard on the default branch (main twin) (#12086)

* fix(ci): accept CVE-2025-68121 in the prebuilt tls-client .so, auto-close base-red issues, guard Scorecard on the default branch

- .trivyignore: CVE-2025-68121 (Go stdlib crypto/tls inside bogdanfinn/tls-client
  v1.15.1, built with go 1.24.1) with justification, expiry and tracker #12084.
  No upstream rebuild exists; the blocking Trivy gate now also names the ignore
  file explicitly.
- nightly-release-green: close the "not green" issue when the validation passes
  again (the workflow only ever opened/commented it, so stale issues outlived
  the fix and stamped new PRs as base-red inherited).
- scorecard: the action only accepts the DEFAULT branch (the active release
  branch, not main) - guard the job on it so pushes to main stop failing.

Refs #12084

(cherry picked from commit 8adf34bada)

* fix(release): never let the tag-push Create Release append auto notes to the curated body

Twin of the release/v3.8.51 commit (see #12085).

Refs #12084

* fix(docker): bump Bun image to 1.4.0 with Turbopack and port the node image's build memory guards (#11719)

Validated in an isolated worktree against main: typecheck:core clean, 15/15 focused tests pass (docker-build-memory-budget, bun-support, resolve-next-build-bundler-flag). Root cause confirmed against the current workflow config (docker-publish.yml triggers on push to both main and release/v*, so this genuinely needed to target main). One out-of-scope change dropped before merging: config/alibaba-free-tier-allowlist.json's validUntil bump (2026-08-27 -> 2027-12-31) was unrelated to the Docker/Bun fix — reverted to the current value, keeping only the Docker/Bun/memory-guard changes this PR is actually about. Thanks for the thorough root-cause writeup and the worker-pool math.

* test(infra): retry recursive temp-dir removal on main (main twin of #11968) (#12246)

* test(infra): retry recursive temp-dir removal on main (main twin of #11968)

`main` has been red since b342c1a361 on the vitest and integration gates:

  ✖ tests/unit/autoCombo/provider-family-combos.test.ts > auto/<family>
  ✖ chat pipeline applies Codex OAuth fingerprint and priority tier inside combos

Both call resetStorage() from beforeEach, which does an fs.rmSync(TEST_DATA_DIR,
{recursive: true, force: true}) with no retry, and intermittently loses the race
with a not-yet-released SQLite handle (ENOTEMPTY).

release/v3.8.51 fixed this in #11968 with a mechanical codemod adding
maxRetries/retryDelay to every recursive rm/rmSync/rmdirSync under tests/, but
that PR landed only on the release branch. Because main only receives work at
the release squash, it stayed broken for the whole cycle — and repo-wide gates
then turn every open PR into main red on checks unrelated to their diff.

This is the --base main twin: re-runs the same codemod that already shipped on
the release branch (scripts/ad-hoc/codemod-rm-maxretries.mjs), so the two
branches converge on identical test-teardown semantics. Test-only; no product
logic is touched.

The remaining three failures reported on #12133 (unit full suite exceeding its
4800s ceiling, package-artifact exceeding 1200s, and the boot-smoke that is
skipped as a consequence) are runner-contention timeouts, not code defects —
validate-release-green.mjs runs those heavy gates concurrently on one shared
hosted runner. There is no fix to port for those.

* chore(scripts): carry the rm-maxretries codemod onto main alongside its output

The codemod that generated the previous commit lives in the repo on
release/v3.8.51 (added by #11968) but was never on main. Bringing it over keeps
the tool next to the change it produced, so the transformation stays
reproducible and auditable from either branch.

* fix(ci): port the release-green ESLint gate fix to main (base-red #12363) (#12618)

Porta para `main` o fix do gate de ESLint que só havia entrado na branch de release — o padrão de PR-companheiro que `_shared/merge-gates.md` §8 prescreve.

As 12 falhas de CI foram discriminadas como o **outro** base-red do main, não deste diff. Todas descendem de um único ponto: `Package Artifact` falha e os 9 shards de E2E mais os 2 Electron Package Smoke consomem esse artefato. A própria issue #12363 lista os dois separadamente:

- ` ESLint: could not parse eslint json` — que é justamente o que este PR conserta;
- ` Package artifact (npm pack policy): gate exceeded its 1200s ceiling` — a raiz da cascata.

O PR toca apenas `scripts/quality/validate-release-green.mjs` e seu teste, então não tem caminho para afetar o build do pacote. Teste portado primeiro e falhando no script atual do main (TDD).

* fix(authz): preserve zed-hosted native-app callback through root middleware redirect

The root middleware intercepts `pathname === "/"` and redirects to
`/dashboard` using `new URL(basePath+"/dashboard", url)`, which drops
the query string entirely.

Zed's native-app sign-in always redirects the browser to the loopback
root — `http://127.0.0.1:<port>/?user_id=...&access_token=...` — ignoring
any path. When the dashboard's own loopback port is reused as
`native_app_port` (see `src/lib/oauth/providers/zed-hosted.ts`'s
`resolveDashboardLoopbackPort`), that redirect lands on `/` of the running
OmniRoute instance. The root page (`src/app/page.tsx`) was already written
to forward `user_id`+`access_token` to `/callback`, but this middleware
runs first and silently discards the payload — making page.tsx's forward
dead code and breaking the entire zed-hosted sign-in flow.

Fix: detect `user_id` + `access_token` in `searchParams` and, when
present, redirect to `/callback${search}` (preserving the query string)
instead of `/dashboard`. Regular root visits (no native callback params)
continue to redirect to `/dashboard` unchanged.

This approach mirrors what `src/app/page.tsx` already does and is
provider-agnostic: any future provider whose native loopback callback lands
on `/` with `user_id`+`access_token` params benefits automatically.

* test(authz): cover zed-hosted native-app callback root redirect (#13140)

Adds automated coverage for the new pathname === "/" branch: with
user_id+access_token both present the redirect now forwards to
/callback preserving the query string; with only one of the two
present, behavior is unchanged (redirect to /dashboard).

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
Co-authored-by: Rouzbeh† <78313022+rqzbeh@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:23:39 -03:00
Bl0ck
62d14a7fc7 feat(audio): expand Fish Audio S2.1 and voice cloning (#13090)
* feat(audio): expand Fish Audio S2.1 and voice cloning

* fix(audio): type Node streaming request init

* test(audio): align Fish Audio CI expectations

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:23:31 -03:00
Tuan Dinh
45fa62a18e feat(build): add build:fast and start:fast to bypass standalone tracing (#13021)
* feat(build): add build:fast and OMNIROUTE_SKIP_STANDALONE to bypass standalone tracing

* feat(build): add start:fast script to run non-standalone builds locally

* docs(changelog): add fragment for #13021
2026-09-18 12:23:23 -03:00
Tuan Dinh
e188da6399 fix(dev): allow Ctrl+C to promptly kill dev server by closing active connections (#13020)
* fix(dev): allow Ctrl+C to promptly kill dev server by closing active connections and adding force-exit timeout

* test(dev): add regression assertions for prompt dev server exit on Ctrl+C

* docs(changelog): add fragment for #13020
2026-09-18 12:23:15 -03:00
Tuan Dinh
4587c7ec3a fix(mitm): add catch-all (*) model mapping fallback for Agent Bridge (#12140) (#13013)
* fix(mitm): add catch-all (*) model mapping fallback for Agent Bridge (#12140)

* docs(changelog): add fragment for #13013
2026-09-18 12:23:07 -03:00
Zach Frederich
80a2bb1230 fix(cli): skip POSIX path lookup on Windows autostart (#12993)
* fix(cli): skip POSIX path lookup on Windows autostart

* docs(changelog): record Windows autostart path fix
2026-09-18 12:22:59 -03:00
Nguyễn Viết Tuấn
d5452d03e7 feat(codex): safely discover compatible models (#12933)
* feat(codex): safely discover compatible models

* docs(changelog): add fragment for #12933

* fix(codex): drop the duplicate GPT-6 Astra registry entries from the merge

release/v3.8.51 had already landed the seven gpt-6-astra* models, and the
merge kept both copies, so the Codex registry listed every Astra id twice.
Keep the release's entries (same ids, capabilities and timeouts) and drop
this branch's copies.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: TheDemonTuan <nguyenviettuanbp@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:22:50 -03:00
ZaimMarzuki
7f1b4a5eb7 feat(dashboard): add sidebar pinned items shortcut section (#12891)
* feat(dashboard): add sidebar pinned items shortcut section

* docs(changelog): add fragment for sidebar pinned items feature

* docs(changelog): update pr number in changelog fragment

* fix(dashboard): scale down sidebar pinned item icons to match text proportions

* fix(dashboard): remove redundant pin icon from PINNED category header

---------

Co-authored-by: ZaimMarzuki <ZaimMarzuki@users.noreply.github.com>
2026-09-18 12:22:42 -03:00
botea
9eb7c2f17e feat(compression): add Hungarian Caveman language pack (#12825)
* feat(compression): add Hungarian Caveman language pack

* docs(changelog): add Hungarian Caveman entry

* chore(changelog): make the fragment a well-formed bullet

changelog.d/ fragments must start with "- " (scripts/release/aggregate-changelog.mjs).

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: botii16 <botii16@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:22:34 -03:00
KeelTrace
95d3b164a5 fix(models): preserve free-model metadata from discovery (#12763)
* fix(models): preserve live free economics in synced discovery

* docs(changelog): add fragment for free-model metadata discovery fix

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:22:26 -03:00
Goni Sulaiman
5c305680ac fix(sse): strip internal markers from the dario and 9router request bodies (#12729) (#12735)
`dario` and `9router` both override transformRequest() without calling the base
implementation, so the internal-marker strip that runs inside
BaseExecutor.transformRequest() never applied to their bodies, and the second
strip before dispatch in BaseExecutor.execute() is not on their path either.
Whatever the routing layer left on the request — including the context-relay /
universal-handoff markers — went upstream verbatim, and strict
OpenAI-compatible gateways reject unknown top-level keys with HTTP 400.

glm and gitlab look like the same class of bypass but are not: glm's
transformRequest() calls super, so the shared strip already covers it, and
gitlab rebuilds its payload field by field, so no extra key can survive to the
wire.

Co-authored-by: Goni Sulaiman <gonisulaimann@users.noreply.github.com>
2026-09-18 12:22:17 -03:00
Juri
db5ae3c33d docs(dependencies): clarify socket.yml is registry-side scan, not CI gate (#12664)
* docs(dependencies): clarify socket.yml is registry-side scan, not CI gate

* docs(dependencies): add changelog fragment for socket.yml scope note

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:22:09 -03:00
wofiporia
6e74739607 fix(pricing): accept sync-written fields on PATCH and surface actionable save errors (#12629)
* fix(pricing): accept sync-written fields on PATCH and surface actionable save errors

* docs(changelog): add fragment for #12629

---------

Co-authored-by: wofiporia <172453170+wofiporia@users.noreply.github.com>
2026-09-18 12:22:00 -03:00
ZaimMarzuki
060f70ed18 feat(dashboard): show exact token counts on hover in usage analytics cards and tables (#12553)
* feat(dashboard): show exact token counts on hover in usage analytics cards and tables

* fix(dashboard): lock tooltip position to prevent top-left slide animation

---------

Co-authored-by: ZaimMarzuki <ZaimMarzuki@users.noreply.github.com>
2026-09-18 12:21:52 -03:00
Gaul Samuyenga
43758c7885 fix(cli): support dashboard API key ids (#12520)
* fix(cli): support dashboard API key ids

* docs(changelog): add fragment for CLI API key route drift fix

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:21:44 -03:00
Mr White
5ba3eee2b0 fix(devin): treat Devin CLI model ids as literal — never strip or synthesize effort suffixes (#12492)
The Devin CLI providers (devin-cli, devin-cli-agentic, devin-desktop; aliases
dv/dva) serve a catalog whose model ids EMBED the reasoning tier:
claude-opus-5-low, claude-opus-5-medium, … and gpt-5-6-sol-max/-low are
distinct upstream models (see registry/devin/catalog.ts).

applyClaudeEffortVariant stripped the trailing -{low,medium,high,xhigh,max}
from any id whose base is a known Claude model, regardless of provider. For
Devin lanes this dispatched a base id that does not exist upstream, e.g.

  dva/claude-opus-5-low  ->  claude-opus-5  ->  400
  'Model is not present in the current Devin catalog: claude-opus-5'

Only accidental double-suffixed ids (dva/claude-opus-5-max-low) survived,
because stripping the outer -low left the real claude-opus-5-max. Symmetrically,
the catalog synthesized -<level> variants on top of tier-embedded ids,
advertising phantom ids (dva/gpt-5-6-sol-max-low, dva/kimi-k3-*) that 400 when
called.

Three gates now treat Devin ids as literal:
- applyClaudeEffortVariant: early return for Devin providers (ids/aliases)
- appendClaudeEffortVariants: no -<level> variants for devin-prefixed ids
- appendSyncedEffortVariants: isSkippedEffortProvider now covers Devin
  providers (they own their suffix mechanism — the tier IS the id)

Validated live on a self-hosted v3.8.51 deployment: dva/claude-opus-5-low,
dva/claude-5-fable-low and the whole tier-embedded catalog now dispatch; the
phantom variant ids disappear from /v1/models. Claude-lane stripping
(claude/cc, e.g. cc/claude-opus-5-high -> claude-opus-5 + reasoning_effort) is
unchanged and covered by existing + new characterization tests.

Co-authored-by: Neuron Mr White <whiteneuron@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:21:36 -03:00
dmlanday
b36c81d3f0 fix(cli): escalate the readiness probe timeout so a slow health response is not a phantom boot failure (#12484)
* fix(cli): escalate the readiness probe timeout so a slow health response is not a phantom boot failure

`omniroute serve` reported "Server did not respond within 60s" over servers that
were up and serving traffic. Every probe of /api/monitoring/health was aborted at
a fixed 2s, and a timed-out probe is classified "hanging", which never counts
toward readiness (#6800). So whenever the first health response takes longer than
2s the poll can never succeed: each abort discards the in-flight request before
the route finishes (its own 1s payload cache is never populated either), and
500ms later the next probe restarts the same work into the same ceiling, for the
whole 60s budget. Reproduced by the new test: against a health route that answers
200 in 3.2s, the old poller ran 12 probes over 30s and reported ready=false every
time.

The per-probe timeout now escalates after each hang (2s, 4s, 8s, 15s), clamped to
the time left in the budget so the caller's total timeout still holds. Only a hang
escalates, so #6800's guarantee is unchanged: a socket that accepts TCP and never
answers still resolves false. waitForServer also reports each probe outcome to an
optional onOutcome callback, and the readiness-timeout diagnostic uses it to say
whether the port was accepting connections, which separates "up and still warming"
from "never bound the port".

Same failure family as #10508, which fixed it by taking a DNS lookup out of the 2s
budget rather than by widening it. The heavy /api/monitoring/health route is what
makes that budget tight in the first place (its own docstring points high-frequency
pollers at /api/health/ping, which is what the Electron readiness poller uses);
switching the CLI probe route is a larger change, left as a follow-up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CJf2dxEpiwZqyZujWk57T2

* chore(changelog): link the readiness-probe fix to PR 12484

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CJf2dxEpiwZqyZujWk57T2

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-09-18 12:21:27 -03:00
killer30001000
d71d0f76e5 fix(usage): render OpenRouter PAYG credit pool with real denominator (#12468)
* fix(usage) handle OpenRouter PAYG credit percentage

OpenRouter PAYG accounts without a per-key limit previously rendered
the credits row as 'total: 0, remainingPercentage: 100, unlimited: true',
treating /credits balance as unlimited even when a real credit pool was
present. Route the credit pool through the credits renderer with the
real denominator: total = totalCredits when positive, used = total -
creditBalance, remaining = creditBalance, remainingPercentage =
round(balance / total * 100), isCredits: true, unlimited: false. Per-key
limit still wins. A non-positive pool surfaces the row but never invents
a 100% bar.

Tests cover: explicit key limit, PAYG account credits without key limit,
key limit taking priority over account credits, and a balance without a
positive denominator.

* fix(usage) render OpenRouter PAYG quota as a metered percentage bar

The frontend parser was routing every OpenRouter 'credits' quota through
buildCreditsQuota(), which sets isCredits: true. QuotaCardExpanded
short-circuits on that flag and shows only the USD balance as a bare
number, so a real PAYG payload (used: 7.33, total: 10, remaining: 2.67,
remainingPercentage: 27) was rendered as '$2.67' instead of the '27% left
/ 7.33 / 10' bar the backend already computed.

Drop isCredits: true for any payload whose total is a positive finite
number - the row then goes through the normal normalizeQuotaEntry() path
with currency preserved as an extra. The balance-only fallback (total 0
or non-finite denominator, used by legacy /credits responses) still uses
buildCreditsQuota() so the row stays renderable, and never invents a
100% percentage.

The frontend test now asserts:
- PAYG positive denominator -> total: 10, remainingPercentage: 27,
  currency: 'USD', isCredits !== true.
- Balance-only payload -> isCredits === true, creditCount === 2.67,
  total: 0, no fabricated 100%.
- NaN denominator -> balance-only fallback.
- Non-credits keys -> unchanged normalizeQuotaEntry() path.
- Mixed payload -> normal quota row + PAYG row, both kept.

* docs(changelog): add OpenRouter PAYG fix fragment

* docs(changelog): remove self credit
2026-09-18 12:21:17 -03:00