mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-09-21 14:22:14 +03:00
80e2ca7c392b345d759dca58e05c2b8944bb73fe
3144 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
80e2ca7c39 |
fix: don't blanket-disable Claude tool-name prefix for non-Anthropic providers (#13856)
* fix: don't blanket-disable Claude tool-name prefix for non-Anthropic providers
The Issue #199/#618 fix that disables the proxy_ tool-name prefix for
requests landing in the general (non-CC-bridge, non-passthrough)
translation branch was scoped to "targetFormat === Claude", but that
condition is true for any provider translating into Claude's wire
format, not just genuine first-party Anthropic traffic. In practice
this let ordinary third-party tool names (e.g. GitHub Copilot's own
client-executed "web_fetch" function tool) pass through unprefixed
and collide with Claude's reserved tool namespace, getting rejected
upstream ("rejected tool(s): web_fetch") for any gh/claude-* model.
#618's actual traffic was real Claude Code (provider "claude") landing
in this fallback branch, so scoping the disable to provider === "claude"
(the same check already used a few lines above in the sibling
isClaudePassthrough branch) keeps that fix intact while restoring
correct prefixing for every other provider that merely targets
Claude's wire format.
Fixes #13835.
* docs: add changelog fragment for #13856
|
||
|
|
c299e6c180 | fix(translator): pair Gemini tool responses per turn to prevent cross-turn ID collision (#13848) | ||
|
|
fa23670ea2 |
fix(codex): forward the caller client version upstream instead of a pinned default (#13708)
* fix(codex): forward the caller client version upstream instead of pinning 0.144.1 The Codex provider reported a hardcoded client version (DEFAULT_CODEX_CLIENT_VERSION) to the ChatGPT backend. Newer gated models reject older clients, e.g.: The 'gpt-6-astra' model requires a newer version of Codex. so the pinned value silently rots on every CLI upgrade, and a user on the latest CLI is still refused. CodexExecutor.buildHeaders also dropped the clientHeaders/model/health arguments that the base class accepts (base.ts:509), so the caller User-Agent never reached the version resolution. - codexClient.ts: add getCodexClientVersionFromHeaders(), which reads the version the caller reports in its User-Agent (codex_cli_rs/<v>, codex_exec/<v>) or a version header, validated against SAFE_HEADER_TOKEN_PATTERN. getCodexUserAgent() now takes an optional version override. - codex.ts: forward clientHeaders/model/health to super.buildHeaders(), and use getCodexClientVersionFromHeaders(clientHeaders) ?? getCodexClientVersion(). Falls back to the existing env override / default when the caller sends nothing. * test(codex): cover getCodexClientVersionFromHeaders and clientHeaders-aware buildHeaders Adds unit coverage for the caller-version forwarding introduced in this PR: real Codex CLI User-Agent parsing, the generic version header, missing/empty headers, non-Codex User-Agent, and CRLF/oversized injection attempts safely returning null/falling back to the default version. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Zeeshan Haque <zeeshan@moonscape.local> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
c8434436af |
feat(antigravity): expose physical send telemetry (#13659)
Co-authored-by: ginettododo <117327638+ginettododo@users.noreply.github.com> |
||
|
|
c43fbb1b3d |
fix(security,resilience): block origin-IP header forwarding and treat 413 as retryable TPM (#13350)
* fix(security): never forward origin-IP headers upstream Operator-set custom upstream headers could carry the client-origin IP (x-forwarded-for, x-real-ip, cf-connecting-ip, forwarded, via, ...) to the upstream provider, disclosing or spoofing it. Extend the FORBIDDEN denylist in upstreamHeaders.ts to cover the whole forwarding/IP set, mirroring the scrubbers already used by the Antigravity (antigravityHeaderScrub.ts) and Cursor CLI (cursorCliProxy.ts) paths, so the protection applies to every provider rather than those two. Covered by tests/unit/upstream-headers-sanitize.test.ts (6 passing). * fix(resilience): treat 413 payload-too-large as retryable TPM rate limit Providers with a tokens-per-minute cap (Groq among them) answer an oversized turn with 413 rather than 429. checkFallbackError() did not list 413 as retryable, so the request failed hard instead of falling back to another account or model. - Add PAYLOAD_TOO_LARGE (413) to HTTP_STATUS and to the retryable set - Return a MODEL_CAPACITY retryable fallback for 413 - Recognise "tokens per minute" / "tpm" as context-overflow patterns --------- Co-authored-by: Themedexperiencesusa <221764849+themedexperiencesusa@users.noreply.github.com> |
||
|
|
7663aadea9 |
feat(models): add Gemini 3.8 Flash tiers to Antigravity and AGY catalogs (#13318)
* feat(models): add Gemini 3.8 Flash tiers to Antigravity and AGY catalogs
* fix(antigravity): handle Gemini 3.8 Flash thought signatures, native tool calls, and output token limits
* fix: remove duplicate Gemini 3.8 model specs
* feat(models): adopt CLI catalog, pricing, and version fallbacks from #12499 (#13318)
* test(models): assert Gemini 3.8 Flash catalog and pricing presence (#13318)
* fix(antigravity): Gemini 3.8 Flash tiers have no shared -tiered endpoint
Live testing against Google's Cloud Code upstream
(daily-cloudcode-pa.googleapis.com/v1internal:streamGenerateContent),
documented in #12499, shows Gemini 3.8 Flash is served directly at
gemini-3.8-flash-high/-medium/-low. Unlike 3.7, there is no
gemini-3.8-flash-tiered endpoint for 3.8.
This branch mapped the bare id and all three tiers to an invented
gemini-3.8-flash-tiered upstream target and declared that model in the
shared Antigravity/AGY catalog. Remove the invented catalog entry, alias
only the bare "gemini-3.8-flash" display id to its default tier
(gemini-3.8-flash-high), and let -high/-medium/-low pass through
verbatim to match what the live endpoint actually serves. Also drops
the now-dead managedModelImport.ts mitm-alias branch that forced the
same invented -tiered target, and updates the catalog/alias tests
accordingly.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* revert: drop CLI catalog/pricing/version-fallback adoption also shipped by #12499
Commit
|
||
|
|
8288a4a012 |
fix(sse): deepseek-web resilience — premature session close + malformed tool-call recovery (#13226)
* fix(sse): deepseek-web collectSSEContent no longer returns a silent partial stub on premature session close
collectSSEContent() (used for the deepseek-web tool-calling / non-stream path)
drained the upstream SSE body and returned whatever content it had once the
reader reported done, with no check that DeepSeek had actually signalled
completion via response/status: "FINISHED".
When the upstream cookie session drops mid-generation (expired session,
anti-bot challenge, network interruption), the HTTP body simply closes
early. That was indistinguishable from a real completion: execute()
returned HTTP 200 with finish_reason "stop" and whatever partial stub text
had arrived so far. Observed in production call logs: a lone "I'll check
that..." / "Vou verificar..." with no continuation, reported as a
successful completion.
collectSSEContent now tracks whether the FINISHED status event was seen. If
the stream ends without it, it throws instead of returning the stub -
execute()'s existing try/catch turns that into a proper 502 that the
client, or a combo's retry/fallback logic, can react to.
Added tests/unit/deepseek-web-premature-close.test.ts covering both the
premature-close error path and the normal FINISHED completion path. Full
deepseek-web unit suite (97 tests) still passes.
* fix(sse): recover malformed deepseek-web tool-call replies and retry when unrecoverable
Two related failure modes on the deepseek-web tool-calling path, both
observed in production call logs from real agentic (VS Code Copilot-style)
usage of the deepseek combo:
1. DeepSeek's web session occasionally leaks malformed/internal formatting
tokens right after an otherwise-complete <tool>{json} body, instead of
a clean </tool> close (observed: a fully valid create_file JSON call
immediately followed by corrupted pseudo-tags). parseLooseJsonObject's
strict JSON.parse rejected the whole block over that trailing garbage,
even though a perfectly valid object sat at the start - so the call was
silently dropped and the raw tagged text was shown to the user instead
of the file being created.
deepseekWebTools.ts: added salvageLeadingJsonObject(), a quote/escape
aware balanced-brace scanner that recovers just the leading JSON object
when the strict parse fails, reusing the same salvage idea already used
elsewhere in this file (findBareJsonCandidates) for bare-JSON detection.
2. When even that salvage cannot recover a call (genuinely truncated JSON,
garbled beyond repair), execute() previously gave up on the first try.
Since this is a scraped, non-deterministic web session rather than a
real API, simply asking again is usually enough to get a clean reply.
deepseek-web.ts: the hasTools branch now detects an unparsed <tool...>
tag surviving in the cleaned content and retries with a brand-new
session, bounded to MAX_TOOL_PARSE_ATTEMPTS (2) - never an unbounded
retry loop, and a reply that parses cleanly on the first try costs no
extra latency.
Builds on the collectSSEContent premature-close fix from the same PR -
that one covers the upstream session dropping mid-stream; this one covers
the session completing but returning malformed tool-call content.
Testing:
- tests/unit/deepseek-web-tools-salvage-leading-json.test.ts (4 tests):
recovery from the exact production-observed corruption pattern, escaped
quotes/nested braces before the garbage, correct non-promotion of
genuinely truncated JSON, and no regression on well-formed blocks.
- tests/unit/deepseek-web-tool-call-retry.test.ts (3 tests): retry
succeeds on a fresh session, retry is bounded (gives up after
MAX_TOOL_PARSE_ATTEMPTS and surfaces the raw content rather than
looping forever), and a clean first reply never triggers a retry.
- Full deepseek-web unit suite: 104/104 passing, no regressions.
- The salvage fix was additionally verified directly against the exact
malformed content captured from a live production call log (not just
the hand-written test fixture).
---------
Co-authored-by: VictorRP7 <187780317+VictorRP7@users.noreply.github.com>
|
||
|
|
db022768df |
fix: FriendliAI 403 credit exhaustion misclassified as auth_error (#13040)
* Add new error signal for depleted credits * chore: add changelog fragment for #13040 FriendliAI credit-exhaustion 403 fix * test(accountFallback): add FriendliAI credit-exhaustion 403 body test (#13040) --------- Co-authored-by: fesshompa <fesshompa@yahoo.com> |
||
|
|
614ff4b60a |
fix(sse): thread errorText into shouldMarkAccountExhaustedFrom429 (#13008)
* fix(sse): thread errorText into shouldMarkAccountExhaustedFrom429
shouldPreserveQuotaSignals() (open-sse/services/quotaResetParsing.ts) gained
an errorText parameter so an explicit quota-exhausted body could override the
apikey-category default, but only one of its two call sites was updated.
checkFallbackError() passes the upstream body; shouldMarkAccountExhaustedFrom429()
still called it with the provider alone.
With errorText undefined the helper's
`Boolean(errorText) && looksLikeQuotaExhausted(errorText)` branch can never be
true, so for every apikey-category provider without per-model quotas the
connection was never marked quota-exhausted -- even when the upstream body
explicitly said a long-window cap was hit.
Thread errorText through the helper and pass it at the src/sse/handlers/chat.ts
call site. The parameter is optional and additive: OAuth-category providers and
callers that pass no body keep their existing behavior, and plain rate limits
("Rate limit exceeded, retry in 20s", "Too Many Requests") still fall through to
the short generic cooldown.
Regression guard: tests/unit/quota-signal-errortext-threading.test.ts
* test(sse): pin the chat.ts call site that forwards errorText
The errorText parameter added in the previous commit is optional, so dropping
it at the only production call site (src/sse/handlers/chat.ts) is neither a
type error nor a test failure -- the four existing cases call the helper
directly and none of them assert the wiring. The half of the patch that makes
it do anything in production could be reverted, or lost in a refactor, with
the whole suite green. That is the same failure mode this branch fixes (a
two-argument helper with a call site silently passing one), one level up.
handleSingleModelChat is not exported, so the call cannot be driven or spied
without widening the production surface. Assert at the source level instead,
following tests/unit/api-key-provider-quota-bypass-scope.test.ts, and parse
the argument list rather than regex-matching formatted text so a Prettier
reflow cannot cause a false failure or a false pass. Also pins that there is
exactly one call site, and what errorStr holds.
Verified by mutation: dropping errorStr at chat.ts now fails 1 of 5.
* chore: number quota signal changelog fragment
Co-Authored-By: Paperclip <noreply@paperclip.ing>
---------
Co-authored-by: TogetherWeOwn <eng@togetherweown.com>
Co-authored-by: togetherweown[bot] <togetherweown[bot]@users.noreply.github.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
a123bd049e |
fix(images): fall back when combo leg returns empty 2xx response (#12982)
* fix(images): fall back when combo leg returns empty 2xx response fetchImageEndpoint() normalized any successful HTTP response to success:true with data.data || [], so an OpenAI-compatible provider returning 200 with an empty/malformed image payload stopped image combos on the first leg and produced an image-less 200. Require at least one usable item (non-empty b64_json or url) before declaring success; empty 2xx becomes a retryable 502 with a sanitized error so executeImageCombo() advances to the next priority leg. Valid responses and direct image-model requests are unchanged. * chore(changelog): add fragment for image-combo empty-2xx fallback (#12982) * test(images): drop misleading #10199 reference from empty-200 fallback test The test file and its production-code comment referenced issue #10199, which belongs to an unrelated, already-merged PR (auto/best-free free-tier filter fix). Rename the test file and update comments to remove the incorrect reference so future readers are not misled. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * test(images): fix response-shape assertions after #12268 combo envelope change release/v3.8.51 already ships #12268, which changed executeImageCombo() to return the handler's payload unchanged ({created, data: [...]}) instead of unwrapping it into a bare array. Update the two assertions that still expected a bare array so the test reflects the combo response shape that is actually live on the target branch; the fallback behavior under test is unaffected. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Tiangao (hermes) <montigaud@aikumi.pro> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
2b7a881c6f |
fix(claude): passthrough tool_use names must match client-declared casing (#12855)
* fix(claude): passthrough tool_use names must match client-declared casing A mapless restoreClaudePassthroughToolUseName upgraded known Claude Code tool names (bash -> Bash) on every Claude-format SSE passthrough. Clients that declared lowercase tool names (pi, OpenCode, ... on claude-format executors like devin-cli-agentic) received a tool_use name they never declared: client-side tool dispatch fails and echoing the history back hard-fails with devin-cli-agentic undeclared_historical_tool (live repro: pi + dva/glm-5-2 on a self-hosted router, 400 on every agentic turn). - restoreClaudePassthroughToolUseName: alias map first (renamed -> original), then normalize to the request's declared tools[] casing, never canonicalize undeclared names. Genuine Claude Code clients keep their #7926 protection (upstream downcase -> declared PascalCase). - devin-agentic serializer: case-insensitive fallback for historical tool_use names + render the declared spelling, so case drift can no longer kill a whole turn. Tests: tests/unit/claude-passthrough-tool-name-mapless-leak.test.ts, tests/unit/devin-agentic-serializer-case-insensitive-history.test.ts * fix(stream): direct ledger lookups in claude passthrough restore restoreClaudeToolName's canonical-upgrade fallback fires even when an alias ledger exists (canonical beats the identity match). The claude passthrough lane always carries a non-empty proxy_ ledger (buildClaudePassthroughToolNameMap), so every lowercase tool_use name was upgraded to Claude Code PascalCase on the SSE path while the JSON path (direct map.get) stayed correct — the live leak behind #12721. * docs(changelog): add fragment for claude passthrough tool_use casing fix Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
88e666e596 |
fix(sse): fix double-escaped tabs in Codex JSON tool call arguments (#12841)
Fixes #12831. When the Codex upstream model produces string values in tool arguments that contain double-escaped tabs (\t inside the JSON string instead of \t), the parser outputs literal backslash-t characters. This breaks editor patches that rely on proper indentation. This commit adds a fixDoubleEscapedTabs sanitization step before parsing to restore them to single tabs. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
c193595db6 |
fix(open-sse): promote reasoning_details text to reasoning_content even when reasoning present (#12665) (#12688)
* fix(open-sse): promote reasoning_details text to reasoning_content even when reasoning present (#12665) OpenRouter thinking models return both a "reasoning" string and a "reasoning_details[].text" array for the same thinking trace. OmniRoute's reasoning promotion gated on "is any readable value present" (which includes the "reasoning" alias), so reasoning_content was never populated and clients like opencode that only read reasoning_content lost all thinking traces. Encrypted-only reasoning_details items are left intact (not flattened). Fixes all three promotion gates: - non-streaming: copyOpenAICompatibleReasoningFields now mirrors reasoning_details[].text into reasoning_content unless reasoning_content itself is present - streaming mirror block: same gate fix on getReadableReasoningValue - streaming passthrough: force re-serialization when sanitize added a reasoning_content the upstream delta did not carry (needsReserialization was false because hasUnsupportedReasoningSignal requires !readable) Tests: non-streaming + streaming unit regressions and an integration E2E that drives the full handleChat path against a mock OpenRouter provider. Also closes the same latent gate in the JSON-to-SSE rehydrator (jsonToSse.ts buildReasoningDelta): a populated reasoning string used to short-circuit the unsupported-alias mirror, so reasoning_details[].text was dropped when synthesizing an SSE stream from a non-streaming JSON body. Adds a #12665 regression test for that path, fixes an over-indented brace in stream.ts (lint), and restores the missing trailing newline in the E2E. * docs(changelog): add fragment for #12688 — fix(open-sse): promote reasoning_details text to reasoning_content even when reasoning present * fix(tests): replace any with typed casts in #12665 regressions to satisfy no-explicit-any gate |
||
|
|
0f5f83c5ed |
fix(resilience): clear the combo LKGP pin only when it names the failed target (#12235)
* fix(resilience): clear the combo LKGP pin only when it names the failed target
Re-applied onto current release/v3.8.51. The branch was 217 commits behind
and dispatchWithCooldownRetry / handleRoundRobinCombo have since moved out
of combo.ts, so this is a re-application, not a rebase: the definition is
still in combo.ts, and the 14 call sites now live in
combo/executeTargetAttempt.ts (4), combo/executeTargetGates.ts (5) and
combo/roundRobinCombo.ts (5). clearStaleLKGP is dependency-injected via
attemptLoopTypes.ts, so that signature takes the new parameter too.
Unchanged in substance. The target-scoped pin is still always cleared; the
combo-level pin is cleared only when it actually names the failed target's
provider, because it records whichever provider last *succeeded* and that
need not be the one failing now. Under `auto` the pin is a scoring input
rather than a hoist, so clearing unconditionally discarded a preference for
a healthy provider every time an unrelated target was skipped. Omitting
`failed` keeps the old behaviour for callers with no target in scope.
4 regression tests pass, including "a pin naming a healthy provider
survives another target being skipped"
mutation: clear the combo pin unconditionally (the pre-fix behaviour)
-> only that test fails, 3 pass
248 tests pass across tests/unit/lkgp*, tests/unit/combo/*
eslint clean on all five changed files (exit 0, no suppressions flag)
* docs(changelog): add fragment for the LKGP pin scope fix
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: abhisheksharma2411 <abhisheksharma2411@users.noreply.github.com>
|
||
|
|
975b29c275 |
feat(core): improve observability for dual-auth fallback execution (#11828)
* feat(core): improve observability for dual-auth fallback execution
* refactor(executors): keep the clinepass auth decision in buildClinepassHeaders
The clinepass case no longer re-decides OAuth vs BYOK off
credentials.authType; it always delegates to buildClinepassHeaders(),
which already keys the decision off credentials.accessToken, and only
the debug log line branches. The OAuth test now uses the real credential
shape ({ accessToken }) instead of an OAuth token stored in apiKey, and
a parity test pins the executor output to buildClinepassHeaders() for
both credential shapes so the two paths cannot drift apart.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
||
|
|
4d1282be31 |
fix(sse): restore maxQueueDepth=0 as unbounded, sanitize refusals at the write, drain the 09-18 base-reds (#14101)
* fix(quality): drain the 09-18 base-reds, part 1 — thinking gate parity, inventory, webpack externals Reproduced on the clean tip |
||
|
|
95b2e53727 |
feat(usage): generic billing/quota for openai-compatible connections (#13673)
* feat(usage): generic billing/quota for openai-compatible connections (#13616) Every other fetcher in services/usage hard-codes one upstream's URL, auth and response shape, which works because those providers are known. An openai-compatible connection can point at anything, and its id is minted per connection -- so it can never be a member of USAGE_SUPPORTED_PROVIDERS or a case in the dispatcher switch. So the shape comes from the connection instead. `providerSpecificData. quotaEndpoint` declares the url, auth mode, optional headers, and a mapping of dot-paths onto UsageQuota: { "url": "...", "auth": "bearer", "quotas": { "credits": { "used": "$.data.used_usd", "total": "$.data.limit_usd", "currency": "USD" } } } Dot/bracket paths (`$.a.b[0].c`) rather than full JSONPath, so the mapping stays dependency-free and legible in a config field. Three decisions worth stating: - **An unresolvable mapping reports nothing, never 0/0.** A quota reading 0 of 0 renders as fully exhausted, and an operator would act on that. A typo'd path must produce no card, not a fake outage. - **The transport error is not echoed.** The url is operator-supplied and can carry a query-string secret; the message says "unreachable" and nothing more. - **The capability is read off the connection, not the id.** `supportsProviderQuota` already takes the connection and already has a connection-shaped check (moonshot), so the gate goes there. A declared url with no `quotas` mapping does NOT count as supported: it can be fetched but can never yield a quota, and would leave a permanently empty card in Provider Limits. Verified: 7 new tests; mutations each killed by the right one -- let an unresolved mapping fall through to 0/0 -> that test alone fails echo the transport error -> that test alone fails 121 tests pass across this file, usage-families-split, provider-plugin- manifest, provider-limits* and the quota-visibility suites (the drift guard from #13134 included). eslint clean on all four files; the three no-unused-vars errors in usage.ts are byte-identical on the base branch. * fix(usage): bound the openai-compatible quota fetch with a timeout An operator-configured quotaEndpoint that never responds would hang getOpenAiCompatibleUsage()'s fetch() indefinitely, stalling that connection's Provider Limits sync. Same 15s bound as the other fetchers in this directory (grokResetCredits.ts's FETCH_TIMEOUT_MS). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * test(usage): pin the openai-compatible quota fetch timeout A quota endpoint that accepts the connection and never answers must be aborted by the fetch signal instead of hanging the Provider Limits sync. Fails without the AbortSignal.timeout() bound, passes with it. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: abhisheksharma2411 <abhisheksharma2411@users.noreply.github.com> |
||
|
|
c74cea3d35 |
feat(mitm): dynamically inject configured models into Antigravity model catalog (#14006)
* feat(mitm): dynamically inject configured models into Antigravity model catalog - Add /v1internal:fetchAvailableModels to ANTIGRAVITY_TARGET.endpointPatterns in src/mitm/targets/antigravity.ts - Implement catalog interception and dynamic model merging in AntigravityHandler.intercept() (src/mitm/handlers/antigravity.ts) - Merge operator's configured combos/models dynamically from the repository into Google Cloud Code's upstream catalog - Prepend injected models to agentModelSorts recommended group while preserving native models and upstream structure - Add unit tests covering target endpoint pattern declaration, catalog merging, dynamic combo retrieval, and error propagation in tests/unit/mitm-handler-antigravity.test.ts Resolves #13959 * test(mitm): isolate DATA_DIR and clean up the combo row in the antigravity catalog test The DB-backed test ("dynamic catalog pulls configured combos from database repository") creates a real combo row via src/lib/db/combos.ts, whose module- level DATA_DIR const resolves once at import time. The PR's own documented Validation command (`node --import tsx/esm tests/unit/mitm-handler-antigravity.test.ts`) runs without the `--test` flag, so the existing #10428 eval-probe/test-context guard in resolveWritableDataDir() never triggers and DATA_DIR falls through to the real ~/.omniroute home database — writing a permanent test-combo row into it every run. Set DATA_DIR to an isolated temp dir at the top of the file (before the combos.ts import), reset the DB singleton and clean up the temp dir in test.after(), and wrap the combo creation in try/finally so the created row is deleted even on assertion failure. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(mitm): prevent model name collision and filter inactive combos in antigravity catalog * feat(antigravity): integrate native auto groups, fallback to groq, and add bridge proxy - Inject OmniRoute native auto groups (auto/best-fast, auto/best-coding, auto/best-reasoning, auto/best-free, etc.) into Antigravity IDE & CLI /model selector - Add bin/antigravity-bridge.mjs with selective proxy routing to isolate native Gemini quota (zero Google token leakage) - Implement transparent self-healing model remapping to prevent upstream 410 model_shutdown errors on deprecated models - Update emergencyFallback provider from nvidia to groq/openai/gpt-oss-120b for resilient 0.02s failover * test(antigravity): add unit test suite for antigravity bridge routing and model self-healing - Add tests/unit/antigravity-bridge-routing.test.ts covering zero quota leakage for native Gemini models - Validate OmniRoute auto group routing and display name interception - Validate retired upstream model self-healing (preventing HTTP 410 crashes) - Export helper methods from bin/antigravity-bridge.mjs with isMain guard --------- Co-authored-by: Stavan <stavan794@gmail> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: steve25060 <steve25060@users.noreply.github.com> |
||
|
|
1b2349de22 |
feat(sse): Claude OAuth lower-priority lane + weekly session-limit reset (#13074)
* feat(sse): Claude OAuth lower-priority lane + weekly session-limit reset
Mirror Claude Code's /low-priority and /limit-reset for OmniRoute-managed
Claude subscription accounts (wire contract captured from Claude Code 2.1.263).
Both are opt-in per connection (providerSpecificData.lowPriorityMode /
autoLimitReset, Edit connection -> Claude section, default off) and only act
on the 5-hour usage wall: a 429 carrying
anthropic-ratelimit-unified-status: rejected and, when eligible,
anthropic-ratelimit-unified-slow-offer: treatment. Nothing is sent before
that first wall 429.
- Lower-priority lane: on the wall the executor retries the SAME account
with `anthropic-usage-limit: slow` and keeps the header on every request
until anthropic-ratelimit-unified-reset (+60s). The intercepted 429 never
reaches chatCore, so the connection is not cooled down or rotated away.
slot_busy (429) / 529 wait slow-retry-after (20s default, 5-600s, +-30%
jitter) bounded by slow-max-wait (20min default, 1min-6h), then end +
10min cool-off. weekly_limit / budget_exhausted / off / ineligible, a
5h-window rollover, or ineligible + overage-in-use end the lane and let
the response flow to the normal cooldown path.
- Session-limit reset: GET /api/oauth/usage?at_wall=1&skip_spend=1 ->
juniper_tide block; when arm=reset and available, POST
/api/organizations/{org}/reset_rate_limits {program: "juniper_tide"} and
retry at full speed. already_used / not offered memoise next_available_at.
- State is in-memory per connection; the executor owns the abort-aware
sleep; the pure state machine and the HTTP client are separate modules
with unit tests; an executor-level test proves the header/retry wiring
end to end with a mocked upstream.
* fix(sse): make the Claude usage-wall handling race-safe for parallel requests
Two requests on the same Claude OAuth connection can hit the 5-hour wall in
the same instant.
- Lower-priority lane: the executor now tells the decider whether THIS
request carried `anthropic-usage-limit: slow`. A sibling built while the
lane was still idle (no header) whose 429 lands after the lane activated
is re-sent on the lane instead of being misread as a "wall" verdict that
would end it; its 2xx is not counted as lane telemetry either.
- Session-limit reset: concurrent wall hits share one in-flight status+claim
round trip (no duplicate POST reset_rate_limits), and for 60s after a
granted reset stale sibling walls are answered "reset" without touching
the network, so they retry at full speed instead of re-claiming or
falling into the slow lane.
Tests cover both races.
* fix(sse): address adversarial review of the Claude usage-wall handling
Three defects found by a 3-lens review of the two previous commits.
1. Lane wait could outlive the request (high). The slot_busy/529 sleep shares
the request's AbortSignal with chatCore's upstream-start timeout (10 min by
default), while the lane's own max-wait defaults to 20 min and can reach 6h
from the server header. A long slot_busy streak was therefore killed
mid-sleep with a TimeoutError instead of ending gracefully as max_wait with
its cool-off. The decision now takes a waitCeilingMs — what is left of the
executor's own timeout, minus a 5s margin — which caps the effective
max-wait and clamps each individual sleep.
2. A wall 429 surfacing only after a 400-driven intra-attempt retry was missed
(medium). The context-editing / thinking-budget / effort / auto-learn
fallbacks all re-fetch and REASSIGN `response`, and the wall check ran
before them, so such a 429 fell through to the generic path and cooled the
connection down. The check now runs after those retries, on the final
response of the attempt.
3. `ineligible` + `overage-in-use: true` ended the lane as plain `ineligible`
on a 429 (medium) because the status mapping ran first; only the non-429
tail produced `extra_usage`. Overage takeover now wins on every status.
Also bounds the module-level per-connection maps with the same FIFO policy as
the identity caches in claudeIdentity.ts: the state key falls back to the
access token when a connection id is absent, and OAuth tokens rotate on every
refresh, so the maps could grow for the process lifetime.
Tests cover all three fixes, including an executor-level regression for the
400-then-wall ordering.
* fix(i18n): add the Claude usage-wall toggle strings to pt-BR
`tests/unit/i18n-pt-br.test.ts` (#6695) requires pt-BR.json to carry every key
present in en.json; the four new `providers.claude{LowPriorityMode,AutoLimitReset}*`
keys were only added to en and it, so the gate failed on this branch.
* refactor(sse): keep the usage-wall change inside the frozen quality budgets
The three ratchets this PR tripped were all its own, not inherited:
- file-size (frozen, may only shrink): open-sse/executors/base.ts 1857 > 1751
and EditConnectionModal.tsx 1653 > 1631.
- complexity / cognitive-complexity (new-code mode): three functions over the
15 threshold — runClaudeLimitResetAttempt (27), handleClaudeUsageLimitResponse
(19 / cognitive 23) and observeClaudeLowPriorityResponse (17 / 17).
Extractions, all behavior-preserving:
- New open-sse/executors/claudeUsageLimit.ts owns the executor-side glue (header
injection, wait accounting, abort-aware sleep, timeout-derived wait ceiling and
the decision logging) behind a ClaudeUsageLimitGuard, so base.ts keeps a
three-line call site instead of ~100 lines of mechanics.
- Three long-standing Claude blocks leave base.ts for the modules they belong to:
mergeCcHeaders + applyStainlessHeaders into config/anthropicHeaders.ts and
stripClaudeSystemPrefixBlocks into executors/claudeIdentity.ts. base.ts is back
at its frozen 1750 lines.
- The modal's Claude section becomes ClaudeConnectionFields.tsx (mirroring
CcCompatibleRequestDefaultsFields) plus a claudeConnectionFields.ts helper that
de-duplicates the field defaults across the modal's two init sites; the file
drops to 1622, below its frozen 1631.
- The three over-threshold functions are split into focused helpers
(observeErrorResponse / observeSuccessResponse, shouldClaimLimitReset,
resolveLimitResetOffer / runLimitResetClaim / memoiseNotBefore).
Gates now: file-size OK, complexity 0 new violations, cognitive 0 new,
fetch-targets / error-helper / build-scope / deps OK, typecheck clean, ESLint 0,
Prettier clean, 155 unit tests green across the feature and its neighbours.
Still failing and NOT this branch's: pack-policy (unexpected
@omniroute/opencode-plugin-v2 files in the npm artifact) and
mutation-test-coverage (stryker tap.testFiles missing entries for
circuitBreaker.ts and comboStructure.ts) — both reproduce on the untouched base.
---------
Co-authored-by: davidebaraldo <davidebaraldo@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
||
|
|
0fbd8854c5 |
fix(api): adapt TEI/Infinity request and response shapes on the /v1/rerank node path (#13733)
* feat(api): route /v1/rerank to remote provider nodes behind RERANK_REMOTE_PROVIDER_NODES POST /v1/rerank only ever dispatched to provider nodes whose base URL hostname was localhost, 127.0.0.1, or 172.16.0.0/12 — a filter hardcoded in the route. A rerank node on any other host (a LAN box or Tailscale peer running TEI, Infinity, vLLM, …) was silently dropped and the request fell through to "Invalid rerank model", even though the same node served /v1/embeddings without complaint and had already passed the provider outbound URL policy at creation time. The memory engine's rerank step calls this route over loopback, so `rerankProviderModel` could not reach such a node either. Mirror the audio routes (#3963): loopback nodes stay always-eligible and unchanged; remote nodes are opt-in via a new `RERANK_REMOTE_PROVIDER_NODES` feature flag (default off — routing to a remote host changes egress identity) AND must pass the provider outbound URL policy (`getProviderOutboundGuard()`, `public-only` deployments never route to private hosts. - src/shared/network/loopbackNodeHost.ts: one pure definition of the loopback host set, replacing three copies (rerank route, audioRegistry, localHealthCheck). The shared version also rejects `user@host` URLs, which the audio copy did not. - src/shared/network/providerNodeHost.ts: policy-aware remote-node eligibility that mirrors guardProviderNodeBaseUrl() on the creation path. - src/app/api/v1/_shared/rerankProviderNodes.ts: pure, testable selection step + loader, modelled on audioProviderNodes.ts. - Feature flag definition, FEATURE_FLAGS.md / ENVIRONMENT.md / .env.example rows, API_REFERENCE.md and MEMORY.md notes, changelog fragment. - tests/unit/rerank-remote-provider-nodes.test.ts covers the host classification, the three policy modes, the selection step, and the route end-to-end (flag off → 400 without contacting the node; flag on → forwarded to <base>/v1/rerank with the node credential; flag on + strict policy → still excluded). Feature-flag count test bumped to 56. * chore(changelog): name the #13732 fragment * fix(api): adapt TEI/Infinity request and response shapes on the /v1/rerank node path The provider-node branch of POST /v1/rerank already fell back from <base>/v1/rerank to <base>/rerank on 404 "for Infinity / TEI", but it kept sending the Cohere body and returned the upstream JSON verbatim. Against Hugging Face text-embeddings-inference that could never work: TEI requires the candidate list as `texts` (HTTP 422 otherwise), takes `return_text`, and answers a bare `[{index, score, text?}]` with no `results` envelope and `score` instead of `relevance_score`. Thin gateways in front of TEI/Infinity commonly emit `score` too. Either way the memory engine's applyRerank(), which reads `results[].relevance_score`, ended up with undefined scores. Add two pure adapters in src/app/api/v1/_shared/rerankLocalNodeShapes.ts: - buildLocalRerankRequestBody(): one upstream body carrying both spellings (`documents` + `texts`, `return_documents` + `return_text`). TEI's request struct is not deny_unknown_fields and the OpenAI-shaped servers (vLLM, llama.cpp, Infinity, oMLX) ignore extras, so a single body serves all. - normalizeLocalRerankResponse(): folds `{results:[…]}`, Voyage-style `{data:[…]}`, and TEI's bare array into the Cohere envelope, backfilling `relevance_score` from `score`, sorting by score, honouring `top_n`, attaching `document.text` when requested, dropping malformed entries, and preserving other top-level fields (`model`, `usage`, …). The route now uses both on the primary and fallback fetch. Cloud registry providers are untouched (they go through open-sse/handlers/rerank.ts). tests/unit/rerank-local-node-shapes.test.ts covers the adapters and the route end-to-end: 404 → /rerank with `texts`, bare TEI array normalized and top_n-capped; `score`-only gateway → `relevance_score` for the client. * chore(changelog): name the #13733 fragment * refactor(api): split the local rerank response normalizer into per-entry helpers The complexity ratchet (new-code mode) flagged normalizeLocalRerankResponse at 18/15 on both metrics. Pull the per-entry validation and the document resolution into toCohereResult() / resolveResultDocument(); behaviour and tests are unchanged. --------- Co-authored-by: seanford <seanford@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
55b6ca6573 |
fix(compression): fall back in-process when the compression worker fa… (#13637)
* fix(compression): fall back in-process when the compression worker fails (#13145) The worker pool resolved every worker fault with the *uncompressed* body instead of reporting it. `PendingJob` had no reject path at all, so a thread error, a worker exit, a dispatch timeout, or an engine error posted back as `type: "error"` all resolved as `{ compressed: false, stats: null }`. `applyCompressionAsync` then treated that as a legitimate "nothing to compress" result and returned it as-is, so the request reached the provider uncompressed while the response header still announced the selected plan ("stacked") — the header is emitted before the pipeline runs. Nothing was logged at any level, and `compression_analytics` stayed empty because rows are only written when a compressed result is reported. The net effect was compression silently disabled for every worker-eligible request. The worker is a throughput optimisation, not a behavioural variant, so a worker fault must degrade to the in-process pipeline rather than to no compression: - `PendingJob` gains `reject`; `fail()` delegates to a new `abort()` that clears the slot timeout and rejects with a diagnostic cause (thread error, exit code, or timeout budget). - An `error` message from the worker is propagated instead of being swallowed. - `applyCompressionAsync` catches the rejection and falls through to the in-process path, logging the cause. The logger is imported lazily and defensively: `compressionWorker.ts` imports this module, so a static import would pull the logger into the worker bundle, and a logging failure must never be able to break compression itself. `close()` keeps resolving with the unchanged body — shutdown is not a fault. The regression test drives a real worker fault via `OMNI_COMPRESSION_WORKER_TIMEOUT_MS` rather than mocking the module, since this project's tsx/ESM + node:test setup has no `mock.module()` support. Its options are fully populated on purpose: `runCompressionAsync` forwards them into `workerOptions`, and `isStrictlySerializable` rejects an object holding `undefined` values — which would route the test through the in-process path and assert nothing. Production requests always carry all of those fields, which is why the worker path is taken there. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LGJQT3E6iJZq4zNGwfjkPs * fix(compression): keep the timeout path uncompressed, retry only fast worker faults (#13145) Review follow-up: the in-process fallthrough ran the full pipeline on the main event loop for *every* worker fault, including a dispatch timeout. A timeout means the worker already spent its whole budget on that body, so re-running the same CPU-bound work inline would stall other in-flight requests — strictly worse than not compressing on a shared gateway. Faults are now typed by whether recovery is cheap: - `CompressionWorkerError.retryInProcess` distinguishes fast faults (thread error, worker exit, engine throw — no work was done, so the in-process path costs what the worker would have) from a dispatch timeout. - Timeouts keep the original degrade-to-uncompressed behaviour, but are now reported. The defect this PR fixes is the silent swallow, not the degrade. Also strips `reject` from the structured-clone wire job. It is a function, so leaving it on the object handed to `postMessage` threw `DataCloneError` before the worker ever saw the job — turning every dispatch into an immediate fault. Adds the missing `changelog.d/fixes/` fragment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(compression): narrow the worker thread-error type for typecheck:core @types/node 26 types the Worker "error" event payload as unknown, not Error, so `error?.message` failed typecheck:core (TS2339). Narrow with an instanceof check before reading .message. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: marcs7 <marcs7@users.noreply.github.com> |
||
|
|
706dc75c13 |
fix(chatCore): stop executeWithUpstreamStartTimeout leaking its abortPromise listener (hedge-cancelled process exit) (#12406)
* fix(sse): stop mergeAbortSignals from leaking abort listeners
mergeAbortSignals() attached "abort" listeners to its primary/secondary
signals but never removed them once the merged signal settled. Every
executor fetch attempt calls this (fetchWithStartTimeout, once per
URL/retry), so a busy combo request accumulated one live listener per
call on the long-lived combo/client signal. A leaked listener still
fires when that signal is later aborted (e.g. a hedge cancellation
arriving after this merge's own caller already finished), for a merged
output nothing is watching anymore.
Mirrors the already-correct self-cleaning pattern in
open-sse/utils/directResponseStartTimeout.ts's local mergeAbortSignals.
Regression test measures listener growth across repeated merges of the
same long-lived signal: 25 merges leaked exactly 25 listeners pre-fix,
0 post-fix.
(cherry picked from commit 07969655147d3236969b38bfb41280ab4fb52b79)
* fix(server): stop the crash guard re-throwing combo abort reasons
Production crash 2026-08-31 (omniroute.log): on a client disconnect,
handleDisconnect aborted the combo controller and a late abort listener
threw the abort reason on an empty stack:
Error [AbortError]: hedge-cancelled
at ... AbortController.abort ... handleDisconnect
file:///.../src/shared/utils/httpClientAbortGuard.mjs:130 throw err;
isClientAbortError() only knew Node's stream codes and "aborted", so
shouldSwallowUncaught() said false and the guard re-threw, taking the
whole server down.
- Port upstream's AbortError line (name "AbortError" + abort-flavoured
message) so request_signal_aborted / DOMException aborts are absorbed.
- Add an exact-message match for the combo abort reasons from
open-sse/services/combo/comboAbortReasons.ts ("hedge-cancelled",
"combo-per-model-timeout"), name-agnostic because the raw reason is a
plain Error that only gets name="AbortError" stamped on the way out.
A losing hedge / stalled target is never a server fault. Inlined so
this .mjs stays dependency-free for scripts/dev/run-next.mjs.
Tests: port upstream's guard tests, add the exact crash shape, a
child-process replay of the crash (dies pre-fix, survives post-fix), a
genuine-error case that must still crash, and a sync check against
comboAbortReasons.ts. The child-process helper passes a file:// URL, not
a bare path, so the tests run on Windows.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 90c9bce8c474b60c37cc63f4421d50feae4c0ad2)
* fix(chatCore): stop executeWithUpstreamStartTimeout leaking its abortPromise listener
Root cause of the 2026-08-31 production exit (Error [AbortError]:
hedge-cancelled), verified by mapping the crash frames in
.build/next/server/chunks/13721.js back to this file:
- The abortPromise abort listener registered on the long-lived client /
stream signal was never removed in the finally block (only abortListener
and timeoutAbortListener were), so every executor attempt (and every
retry) leaked one listener onto that signal.
- Promise.race only subscribes to abortPromise/timeoutPromise once the
array literal has been evaluated. When execute() threw synchronously the
race never ran, abortPromise was orphaned, and the next hedge
cancellation / client disconnect aborted the signal with the string
reason streamHandler.ts forwards; createAbortError() rebuilt it as an
AbortError-named Error and rejected a promise nothing awaited. That
unhandledRejection reached the process crash guard, which re-threw it as
an uncaughtException and exited with code 7.
Keep a handle to the listener and remove it with the others, and mark the
two race-loser promises as handled so a synchronous throw from execute()
can never orphan them. Race semantics are unchanged (the race still
observes their rejections).
Regression tests: (1) a resolving execute leaves the listener count on
the client signal unchanged; (2) a synchronously throwing execute leaks
no listener and a later abort with the string "hedge-cancelled" produces
no unhandledRejection. Both fail against the previous implementation.
Note: commit 079696551 (mergeAbortSignals cleanup) is correct listener
hygiene but is not on this crash path; this is the fix for the incident.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit e68a50ad854a1945c23e747da4ef15a820cc148d)
* fix(server): absorb raw string abort reasons in the crash guard; document the verified crash path
Follow-ups from the adversarial review of 90c9bce8c:
- open-sse/utils/streamHandler.ts aborts the stream controller with a raw
string reason (getClientAbortReason / handleDisconnect) and undici
rejects with signal.reason verbatim, so a cancellation can reach
process level as a bare string. isClientAbortError() returned false for
every non-object, which would still have exited the process. Absorb the
combo abort reasons and the stream-handler disconnect reasons when they
arrive as strings.
- Correct the mechanism comment: the 2026-08-31 exit was a leaked
upstreamTimeouts.ts abortPromise listener rejecting a promise nothing
awaited (unhandledRejection), escalated by this guard, not a listener
throwing synchronously. The leak is fixed at the source in the previous
commit; this guard remains the last-resort net.
- Reword the inlining rationale (plain node launcher, no reliance on
type-stripping for the .ts constants module).
- Tests: the child-process replay now also exercises the
unhandledRejection route with the exact production error shape and with
raw string reasons; add unit coverage for string reasons and non-object
look-alikes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 696fcc8fe1b9b9bc43e6e5f5f5e44b619a86e68a)
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: Beexly <Beexly@users.noreply.github.com>
|
||
|
|
eb4e3be5dc |
fix(quota): keep Kiro active while any _freetrial pool has quota (#13088) (#13324)
* fix(quota): keep Kiro active while any _freetrial pool has quota (#13088) * fix(quota): rename changelog and remove blank line --------- Co-authored-by: giauphan <giauphan@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
0349627c86 |
fix(opencode): match the upstream free-tier request contract (#14013)
Match the upstream OpenCode free-tier request contract (issue #13935): canonical ses_/msg_ identity ids, versioned User-Agent, and the measured body requirements (stream:true + non-empty tools) with a learn-and-reuse tool-name cache, so no-auth oc/* requests stop being refused with 403 FreeTierError. Supersedes #13937 (session regex and minimum-version rule kept, credited below). Complements #14011 (refusal classification) and #13819 (stream_options strip), both already merged. Closes #13935 Co-authored-by: AStupidBear <16422976+AStupidBear@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
80ea176022 | fix: feat(providers): add kimi/qwen/deepseek/gpt auto-routing families (#13214) (#13709) | ||
|
|
994245a476 |
fix(translator): prevent schema property name collision in Gemini sanitizer (#13057, #13477) (#13690)
* fix(translator): prevent schema property name collision in Gemini sanitizer (#13057, #13477) * refactor(translator): reuse SCHEMA_MAP_KEYS in forEachSubschema and add changelog fragment * test(translator): avoid explicit any in gemini schema collision regression test Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: zcrew0x <zcrew0x@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
31ea46934e |
fix(sse): recognize reasoning_effort in the reactive 400 field-strip retry (#13642)
* fix(sse): recognize reasoning_effort in the reactive 400 field-strip retry Strict OpenAI-compatible gateways that don't implement the reasoning-effort knob reject requests with 400 "Unsupported parameter: reasoning_effort". findOffendingField() did not list it in KNOWN_OFFENDING_FIELDS, so the generic strip-and-retry in base.ts never fired and the 400 surfaced to the client — the request died instead of being retried once without the field. Add "reasoning_effort" to KNOWN_OFFENDING_FIELDS (sibling of the existing reasoning_budget entry, same FCC/NIM-style recovery) and pin the new match in provider-field-strips.test.ts. * chore(changelog): add fix fragment for the reasoning_effort field-strip retry (#13642) |
||
|
|
f24c3665df |
fix(proxy): skip bare TCP health probe for SOCKS5 data plane (#13571)
* fix(proxy): skip bare TCP health probe for SOCKS5 data plane * docs(changelog): add fragment for SOCKS5 bare TCP probe skip fix Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
70eebe9adb |
fix(arena+analytics): atomic ELO sync (fetch-first) + flatRateAsZero in compression writer (#13446)
* fix(analytics+arena): flatRateAsZero in compression writer; atomic arena sync redesign * docs(changelog): add fragments for arena ELO sync and compression flat-rate fixes Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: CrashCartCapital <crashcartcapital@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
62d14a7fc7 |
feat(audio): expand Fish Audio S2.1 and voice cloning (#13090)
* feat(audio): expand Fish Audio S2.1 and voice cloning * fix(audio): type Node streaming request init * test(audio): align Fish Audio CI expectations --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
9eb7c2f17e |
feat(compression): add Hungarian Caveman language pack (#12825)
* feat(compression): add Hungarian Caveman language pack * docs(changelog): add Hungarian Caveman entry * chore(changelog): make the fragment a well-formed bullet changelog.d/ fragments must start with "- " (scripts/release/aggregate-changelog.mjs). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: botii16 <botii16@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
5c305680ac |
fix(sse): strip internal markers from the dario and 9router request bodies (#12729) (#12735)
`dario` and `9router` both override transformRequest() without calling the base implementation, so the internal-marker strip that runs inside BaseExecutor.transformRequest() never applied to their bodies, and the second strip before dispatch in BaseExecutor.execute() is not on their path either. Whatever the routing layer left on the request — including the context-relay / universal-handoff markers — went upstream verbatim, and strict OpenAI-compatible gateways reject unknown top-level keys with HTTP 400. glm and gitlab look like the same class of bypass but are not: glm's transformRequest() calls super, so the shared strip already covers it, and gitlab rebuilds its payload field by field, so no extra key can survive to the wire. Co-authored-by: Goni Sulaiman <gonisulaimann@users.noreply.github.com> |
||
|
|
5ba3eee2b0 |
fix(devin): treat Devin CLI model ids as literal — never strip or synthesize effort suffixes (#12492)
The Devin CLI providers (devin-cli, devin-cli-agentic, devin-desktop; aliases
dv/dva) serve a catalog whose model ids EMBED the reasoning tier:
claude-opus-5-low, claude-opus-5-medium, … and gpt-5-6-sol-max/-low are
distinct upstream models (see registry/devin/catalog.ts).
applyClaudeEffortVariant stripped the trailing -{low,medium,high,xhigh,max}
from any id whose base is a known Claude model, regardless of provider. For
Devin lanes this dispatched a base id that does not exist upstream, e.g.
dva/claude-opus-5-low -> claude-opus-5 -> 400
'Model is not present in the current Devin catalog: claude-opus-5'
Only accidental double-suffixed ids (dva/claude-opus-5-max-low) survived,
because stripping the outer -low left the real claude-opus-5-max. Symmetrically,
the catalog synthesized -<level> variants on top of tier-embedded ids,
advertising phantom ids (dva/gpt-5-6-sol-max-low, dva/kimi-k3-*) that 400 when
called.
Three gates now treat Devin ids as literal:
- applyClaudeEffortVariant: early return for Devin providers (ids/aliases)
- appendClaudeEffortVariants: no -<level> variants for devin-prefixed ids
- appendSyncedEffortVariants: isSkippedEffortProvider now covers Devin
providers (they own their suffix mechanism — the tier IS the id)
Validated live on a self-hosted v3.8.51 deployment: dva/claude-opus-5-low,
dva/claude-5-fable-low and the whole tier-embedded catalog now dispatch; the
phantom variant ids disappear from /v1/models. Claude-lane stripping
(claude/cc, e.g. cc/claude-opus-5-high -> claude-opus-5 + reasoning_effort) is
unchanged and covered by existing + new characterization tests.
Co-authored-by: Neuron Mr White <whiteneuron@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
||
|
|
d71d0f76e5 |
fix(usage): render OpenRouter PAYG credit pool with real denominator (#12468)
* fix(usage) handle OpenRouter PAYG credit percentage OpenRouter PAYG accounts without a per-key limit previously rendered the credits row as 'total: 0, remainingPercentage: 100, unlimited: true', treating /credits balance as unlimited even when a real credit pool was present. Route the credit pool through the credits renderer with the real denominator: total = totalCredits when positive, used = total - creditBalance, remaining = creditBalance, remainingPercentage = round(balance / total * 100), isCredits: true, unlimited: false. Per-key limit still wins. A non-positive pool surfaces the row but never invents a 100% bar. Tests cover: explicit key limit, PAYG account credits without key limit, key limit taking priority over account credits, and a balance without a positive denominator. * fix(usage) render OpenRouter PAYG quota as a metered percentage bar The frontend parser was routing every OpenRouter 'credits' quota through buildCreditsQuota(), which sets isCredits: true. QuotaCardExpanded short-circuits on that flag and shows only the USD balance as a bare number, so a real PAYG payload (used: 7.33, total: 10, remaining: 2.67, remainingPercentage: 27) was rendered as '$2.67' instead of the '27% left / 7.33 / 10' bar the backend already computed. Drop isCredits: true for any payload whose total is a positive finite number - the row then goes through the normal normalizeQuotaEntry() path with currency preserved as an extra. The balance-only fallback (total 0 or non-finite denominator, used by legacy /credits responses) still uses buildCreditsQuota() so the row stays renderable, and never invents a 100% percentage. The frontend test now asserts: - PAYG positive denominator -> total: 10, remainingPercentage: 27, currency: 'USD', isCredits !== true. - Balance-only payload -> isCredits === true, creditCount === 2.67, total: 0, no fabricated 100%. - NaN denominator -> balance-only fallback. - Non-credits keys -> unchanged normalizeQuotaEntry() path. - Mixed payload -> normal quota row + PAYG row, both kept. * docs(changelog): add OpenRouter PAYG fix fragment * docs(changelog): remove self credit |
||
|
|
b84059213f |
fix(gemini): preserve response-schema nullability across union flattening (#12310)
* fix(gemini): preserve response-schema nullability across union flattening
cleanJSONSchemaForAntigravity flattens every union spelling of nullable before
the schema reaches Gemini: flattenTypeArrays turns ["string","null"] into
"string" and flattenAnyOfOneOf drops the {"type":"null"} branch. Correct for
tool parameters, wrong for response schemas — a model with nothing to say can
no longer answer null, so it returns the string "null" or fabricates a value,
and either reaches the client as schema-conformant data. Pydantic emits the
anyOf spelling for Optional[str], so the fabricating path is the common one.
A Phase 1b walk now records Gemini's sibling-key spelling, nullable: true, on
any node whose union carries null — before Phase 2 destroys the evidence. The
key is absent from GEMINI_UNSUPPORTED_SCHEMA_KEYS so it survives sanitizing,
and flattenAnyOfOneOf's Object.assign cannot clobber a key the surviving
branch lacks. Opt-in via { preserveNullable: true }, passed only by the
responseSchema call site; the three tool-parameter call sites keep the default.
Closes #12308
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): add fragment for #12308 gemini nullable schema fix
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
||
|
|
d8be3b1a77 |
feat(sse): reserve the Antigravity account for the request's stream lifecycle (re-land of #10011) (#13929)
* feat(sse): reserve the Antigravity account for the request's stream lifecycle Re-land of the account-lease half of #10011 on the current release branch. Its exact-model-scoping half had already shipped in #8050 and its quota half lost to the tip's aggregate-family design (selectAntigravityQuotaWindowNames / antigravityQuotaFamily.ts); none of that is reintroduced here. The lease is a concurrency reservation only and never reads or writes quota state. The Antigravity account selected for a request is reserved for the whole streaming lifecycle of that request, so a concurrent retry — or the credential handoff inside getProviderCredentialsWithQuotaPreflight — cannot re-pick an account already committed to an in-flight upstream stream. The reservation is scoped to (connection, callable upstream model) rather than the whole account, so one account can still serve two different models at once; catalog ids that resolve to the same upstream id (the gemini-3.7-flash tiers, all gemini-3.7-flash-tiered) share one lease. When every eligible account is leased for that model the request returns a structured 503 antigravity_pool_busy with a bounded Retry-After instead of piling onto a busy account. Opt-in behind ANTIGRAVITY_ACCOUNT_LEASE_ENABLED (runtime, default false). With the flag off no reservation is taken, credentials carry no routing descriptor, every release/hold is a no-op on an undefined lease id, and account selection and dispatch behave exactly as before. #10011's original test suite asserted family semantics for a lease that was exact-model scoped and failed deterministically on its own head; the model ids it used (gemini-3.5-flash / gemini-3-flash-agent) no longer exist in the catalog. The contradiction is resolved in favour of one coherent semantic — exact callable upstream model — and the tests assert it against the alias tables as they are on this branch. Co-authored-by: Ardem2025 <openclaw-auto@example.invalid> * fix(sse): widen the Antigravity lease reservation result so auth.ts narrows it The discriminated-union form of reserveAntigravityLeaseForSelection's return type did not narrow under tsconfig.typecheck-api.json, so reading `reserved.lease` after the `reserved.busy` early return raised TS2339 in the API Route Typecheck gate. A single optional-property shape carries the same information and type-checks everywhere. Co-authored-by: Ardem2025 <openclaw-auto@example.invalid> --------- Co-authored-by: Ardem2025 <openclaw-auto@example.invalid> |
||
|
|
04eba4dc05 |
fix(mcp): load audit sqlite via runtime helper (#13223)
* fix(mcp): load audit sqlite via runtime helper * docs(changelog): add fragment for MCP audit sqlite runtime-require fix Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Codex <codex@openai.com> Co-authored-by: chatchawan-simplewish <chatchawan-simplewish@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
04cc8aab67 |
fix(cache): fold the response output contract into the semantic cache signature (#12309)
* fix(cache): fold the response output contract into the semantic cache signature
The signature hashed only {model, messages, temperature, top_p}, so two temp=0
requests with identical messages but different response_format shared a cache
key: the second was served the first's stored body under a 200, violating the
schema it asked for. tools/tool_choice had the same exposure.
generateSignature now takes an optional output contract — response_format,
text.format, tools, tool_choice, collected by outputContractOf() — and folds it
into the digest only when present, so plain-chat signatures (and every cache
entry already written for them) are unchanged. All three call sites pass it;
read/write symmetry is preserved because bodyForCacheWrite snapshots the same
body object the read path hashed (#cache-signature-asymmetry).
Closes #12307
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): add fragment for #12307 semantic-cache output-contract fix
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* fix(cache): populate both constraint spellings in outputContractOf
The merge with #12734 left generateSignature reading the camelCase
constraints (toolChoice/responseFormat) with a snake_case fallback, but
outputContractOf only filled the snake_case keys, so the #12734
"signature is called with tool_choice/tools/response_format from body"
store tests failed on the merged branch. Set both spellings so either
caller shape reads the value it expects.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: amirrezakm <amirrezakm@users.noreply.github.com>
|
||
|
|
aecd50369b |
fix(sse): treat "length" stop_reason as legitimate in detectMalformedNonStream for Claude messages (#12935)
* fix(sse): treat "length" stop_reason as legitimate in detectMalformedNonStream for Claude messages * docs(changelog): add fragment for length stop_reason fix Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: jasminsehic <jasminsehic@users.noreply.github.com> |
||
|
|
6788de8ef9 |
feat(providers): update Openference free models and add Deyin to compatible agents (#13378)
* feat(providers): update Openference free models and add Deyin to compatible agents * docs(providers): regenerate PROVIDER_REFERENCE.md against the current tip Post-merge regeneration so the diff only reflects the Openference free-model addition, not stale eurouter/greenpt/count churn from an out-of-date local generation. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(providers): soften unconfirmed Openference free-forever claim The Openference pricing page (openference.com/pricing) currently lists five paid plans ($15-$120/mo) and no $0 tier in its structured pricing data, so neither the old "3-day trial" note nor a "free forever" claim can be verified against the source. Point readers to the pricing page instead of asserting a specific duration. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: AnhLead <AnhLead@users.noreply.github.com> |
||
|
|
128f06d645 |
fix(translator): support Responses custom tool choice (#13128)
* fix(translator): support Responses custom tool choice * fix(translator): preserve custom tools across response paths * docs(changelog): add fragment for Responses custom tool choice fix Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Pham Tien Duc <phamtienduceng@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: ducphamtien-fonos <ducphamtien-fonos@users.noreply.github.com> |
||
|
|
5a82da7084 |
feat(routing): deterministic routing strategies for self-hosted entry (RIC-740) (#13611)
* feat(routing): self-hosted unified OpenAI-compatible entry (RIC-738) Divert /v1/chat/completions through the self-hosted provider adapters when OMNIROUTE_SELF_HOSTED_PROVIDERS / OMNIROUTE_SELF_HOSTED_PROVIDERS_FILE is set: one OpenAI-compatible contract in, auto-route to the selected provider (x-omniroute-provider header, provider/model prefix, or first provider), standard OpenAI error shape out. Optional OMNIROUTE_SELF_HOSTED_API_KEY guards the entry (D5 reserved); unset = open loopback route. Upstream credentials stay runtime-only and are stripped from echoed responses. Brings in the provider-adapters baseline from sibling branch (RIC-737) that this entry depends on. Includes 21 passing unit tests (provider selection, model-prefix forwarding, header hygiene, auth, error normalization, SSE passthrough, fall-through/misconfig), docs, env example, changelog fragment. * feat(routing): deterministic routing strategies for self-hosted entry (RIC-740) Add the M2 deterministic routing strategy engine (D3 可审计路由) to the self-hosted unified entry: a declarative `strategy:` block expressing five explainable, non-predictive policies — blacklist/whitelist hard filters, cooldown circuit breaker, cost-priority, latency-aware ordering, and an explicit fallback chain. The ordered candidate list is the fallback chain: a failed primary (network or non-2xx) falls through to the next candidate and each failure feeds the breaker. Every response carries an x-omniroute-route-decision header answering "why this model / why not that one". A pinned provider rejected by a hard filter returns 400 (never a silent re-route); no eligible providers returns 503 with the full explainable decision. No ML/predict dependency. Covers the RIC-740 acceptance: 5 strategy types with unit tests + HTTP fault-injection tests (primary down -> fallback works), config matching docs, and no predict/ML deps. Adds docs, .env.example entries, and a changelog fragment. * refactor(routing): reduce complexity-ratchet violations in new self-hosted routing files Extract cost/id validation, pin-blocked resolution, ordering, and env/file source resolution into small helpers so routingStrategies.ts and selfHostedEntry.ts stay under the complexity-ratchets cap. No behavior change — the same 51 routing-strategies/self-hosted-entry/provider-adapters tests pass unmodified. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * docs(routing): document the 5 self-hosted env vars in ENVIRONMENT.md check:env-doc-sync failed because OMNIROUTE_SELF_HOSTED_PROVIDERS(_FILE), OMNIROUTE_SELF_HOSTED_API_KEY and OMNIROUTE_SELF_HOSTED_STRATEGY(_FILE) were present in .env.example but missing from docs/reference/ENVIRONMENT.md. Add them under "6. Tool & Routing Policies". Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Ant Rich <ant@richants.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: luyuehm <luyuehm@users.noreply.github.com> |
||
|
|
3ebea07278 |
fix(sse): replay reasoning for Responses-API targets on plain turns and Anthropic clients (#13031)
* fix(sse): replay reasoning for Responses-API targets on plain turns and Anthropic clients
DeepSeek thinking mode requires the reasoning of every prior assistant turn
to be passed back once the request carries `tools`, including turns that
made no tool call. Since #10540 routed opencode-go/deepseek-v4-* to
`/responses`, the reasoning replay cache had two gaps on Responses-API
targets, and clients that drop `reasoning_content` hit intermittent
`400 The reasoning_text in the thinking mode must be passed back`.
1. Plain (non-tool-call) turns are keyed on a digest of the normalized
OpenAI transcript. Both capture sites used `translatedBody.messages` as
the history, which a Responses body (`input`) does not carry, so the
write-time digest never matched the read side. translateRequest now
reports the pivot transcript it digested via `onReasoningReplayHistory`,
and the streaming / non-streaming capture sites digest that transcript.
2. The Responses replay pass was gated on `sourceFormat === "openai"`, so
Anthropic Messages clients (Claude -> OpenAI -> Responses) got no replay
at all. The pass now runs on the OpenAI pivot for every source format,
right before the Responses conversion discards `messages`.
The reported transcript is a shallow snapshot of the digested fields only
and travels through a callback, not the body, so nothing new reaches the
upstream payload.
* docs(changelog): add fragment for #13031
* fix(sse): guard the Responses capture sites and skip plain-turn writes with no history
Review follow-ups for #13031:
- Add tests/unit/chatcore-reasoning-cache-write-guard-responses.test.ts:
runs the real handleChatCore against a mocked opencode-go/deepseek-v4-flash
Responses upstream (JSON and SSE), then asserts the next turn's upstream
body carries the replayed `reasoning` input item. Removing either capture
site fallback turns both cases red.
- Project the reported transcript down to the digested fields only
(tool_calls keep type/name/arguments, ids are dropped) and document that
`content` is shared by reference.
- Skip the plain-turn cache write when the history is empty: a real request
always has a prior user turn, so an empty history means the transcript
could not be recovered and a one-message digest can never match.
- Changelog wording: the pre-fix write digested only the assistant message.
* test(sse): select the /responses dispatch by URL in the Responses replay guard
Review follow-ups for #13031: the guard picks the upstream body by URL
(`/responses`) and asserts exactly one such dispatch per turn instead of
taking the last fetch, the streaming case asserts the same body shape as the
non-streaming one, and the `historyMessages` doc on
NonStreamingClientTranslateInput names the Responses-shaped fallback.
* docs(routing): name the replay-history hand-off without tripping the hook heuristic
The fabricated-docs gate treats any `onXxx` token in prose as a plugin hook
name and flagged `onReasoningReplayHistory` (a translateRequest option, not a
hook). Point at the option's home file instead.
* chore(quality): freeze chatCore.ts at 6159 for the Responses replay wiring
check:file-size in PR mode caps a frozen file at max(frozen, base). The rebase onto
the v3.8.51 tip (
|
||
|
|
6768b14b54 |
fix(chatcore): block duplicate turn execution with 409 turn_in_progress (#12912)
* fix(chatcore): block duplicate turn execution with 409 turn_in_progress * test(sse): align turn-execution-guard 409 body expectation with buildErrorBody reason field * fix(errors): preserve duplicate turn classification * test(turn-execution-guard): assert ageMs range instead of exact 0 Comparing ageMs to an exact 0 was flaky under real scheduling latency between the two synchronous calls in the test. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * docs(changelog): add fragment for turn execution guard fix (#12912) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * chore(quality): rebaseline chatCore.ts file-size for the turn-execution guard The guard logic lives in the new open-sse/handlers/chatCore/turnExecutionGuard.ts leaf; what grows chatCore.ts is the irreducible call-site wiring at the single execution chokepoint (acquire, the 409 turn_in_progress early return, the release/handoff bookkeeping and the try wrapper that scopes it). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(sse): keep endpointPath outside the turn-guard try so the failure-usage closure can reach it The try/finally that scopes the duplicate-turn guard block-scoped the resolveChatCoreRequestFormat destructuring, but persistFailureUsage is defined above the try and closes over endpointPath — every failure-usage write would have thrown ReferenceError. Moved the destructuring above the guard (it is a pure derivation from the request, so nothing else changes) and narrowed the acquire result with an explicit === false, which the workspace tsconfig (strict: false) needs to see the non-acquired arm's retryCount/ageMs. check:open-sse-typecheck goes from 4 errors to 0; typecheck:core, eslint, prettier and the PR's 4 turn-execution-guard tests stay green. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Jihyun Son <jihyun.son@sk.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: initguru <initguru@users.noreply.github.com> |
||
|
|
6fec29ca2d |
fix(copilot): fallback to copilot-chat on 403 identity denial for standard provider (#13705)
* fix(copilot): fallback to copilot-chat on 403 identity denial for standard provider * fix(copilot): document COPILOT_INTEGRATION_ID, extract identity fallback, add changelog Adds the missing COPILOT_INTEGRATION_ID entry to .env.example (fixes tests/unit/issue-7793-env-doc-sync-repro.test.ts), extracts the GitHub Copilot 403 identity fallback out of open-sse/executors/base.ts into its own module (open-sse/executors/copilotIdentityFallback.ts) to bring the file back under the frozen file-size ratchet, and adds a changelog.d/fixes fragment for the PR. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: tuandinh0801 <tuandinh0801@users.noreply.github.com> |
||
|
|
24706705fa |
fix(streaming): per-provider fetch-start timeout cap override (#11526 follow-up) (#13002)
* fix(streaming): per-provider fetch-start timeout cap override (#11526 follow-up) Buffered gateways (opencode-go / command-code Console Go tiers) legitimately buffer a whole reasoning generation before the first upstream byte, so their streaming requests can exceed the default 110s headers-wait cap. #11526 capped every streaming request at that ceiling, so these long generations died at exactly 'Fetch timeout after 110000ms' (504) before any bytes arrived. Add a per-provider fetchStartTimeoutCapMs registry knob (600s for opencode-go and command-code) and project it into the executor's LegacyProvider so resolveFetchStartTimeout caps only genuinely unbounded providers. * docs(changelog): add fragment for fetch-start cap per-provider override Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: alvinveroy <alvinveroy@users.noreply.github.com> |
||
|
|
01b2467d61 |
feat(sse): add LLM Gateway DevPass quota tracking (#12462)
* feat(sse): add LLM Gateway DevPass quota tracking Surface the LLM Gateway DevPass allowance (GET /v1/key) in OmniRoute's quota telemetry, mirroring the OpenRouter API-key fetcher pattern. - llmgatewayQuotaFetcher.ts: fetch + parse the DevPass /v1/key response (decimal-string USD values), exposing two windows — monthly plan credits and the 7-day premium-model window — with a 45s TTL cache. Pay-as-you-go keys (devPlan "none") and 401/403 fail open (no quota). - Register in chat.ts before registerGenericQuotaFetchers + register the named windows for the dashboard cutoff modal. - usage/llmgateway.ts leaf + usage.ts dispatch case so the Limits page renders the monthly + weekly premium rows. - Add "llmgateway" to USAGE_FETCHER_PROVIDERS, USAGE_SUPPORTED_PROVIDERS, PROVIDER_LIMITS_APIKEY_PROVIDERS, and the dashboard label/order map. - tests: 21 cases covering the parser, auth fail-open, cache TTL, window exhaustion, preflight proceed/block, registration, and the usage leaf. * docs(sse): add changelog fragment + codebase-doc entry for llmgateway quota * refactor(sse): register llmgateway quota via quotaTrackersBatch Move the LLM Gateway fetcher registration out of chat.ts (a frozen file-size-baseline chokepoint) into quotaTrackersBatch.ts, the dedicated side-effect module that exists precisely so new fetchers don't grow chat.ts. The batch import runs at module load, before registerGenericQuotaFetchers(), so the bespoke fetcher still wins over the generic path. Fixes the file-size gate (chat.ts must not grow). --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
f1eabd8885 |
fix(sse): stop direct fetch retry reusing pooled flat response-start budget (#13703) (#14047)
resolveDirectHeadersTimeoutMs() (open-sse/utils/directResponseStartTimeout.ts) now bounds only the pooled dispatcher attempt (attempt 0) with the flat OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS watchdog. The fresh-socket retry (attempt 1) is by construction a brand-new socket with no zombie-socket risk (#10214's rationale only applies to the pooled attempt), so when the caller already attached its own deadline signal it now defers to a generous, configurable backstop (OMNIROUTE_DIRECT_RESPONSE_RETRY_TIMEOUT_MS, default 600s) instead of reusing the identical short flat window — fixing spurious 504s on healthy slow-TTFB reasoning models that need well over 60s total for both attempts. Regression test: tests/unit/proxyfetch-direct-response-start-flat-retry-budget-13703.test.ts (RED against unmodified code: retry cut at 81ms against an 80ms flat budget with ~1920ms of caller deadline unused; GREEN after the fix). New env var documented in .env.example and docs/reference/ENVIRONMENT.md. tlsProfileForProvider's return type in proxyFetch.ts is pulled into a named alias so its signature stays on one line under prettier's canonical formatting -- otherwise prettier's mandatory lint-staged reformat of this frozen file grows it past the check:file-size baseline on every future touch. base-red inherited: #14004 (docs env/docs contract, fixed separately in #14022; chatHelpers file-size drift) |
||
|
|
1b82b2f982 |
fix(combo): stop chars/4 overestimate demoting a verified context override (#13870) (#14046)
filterTargetsByRequestCompatibility ranked combo targets solely on the chars/4 estimateTokens() heuristic. On a repetitive agent-session body the estimate overstates real usage several-fold, so a manually-overridden primary sized correctly for the real request got marked context-incompatible and was reordered behind an unconfirmed catalog "emergency" member with a large but unverified limit_context. Fix: when the reorder branch promotes known-context-compatible targets, split them by whether their pass came from an operator-set model_context_override (trusted) or bare catalog metadata (advisory), and also trust a near-boundary override rejection (required tokens within 5x the override — covering the ~3.7x overestimate the issue measured) over a catalog-only pass. An override target keeps or regains priority over an unconfirmed catalog-only "known compatible" target; two override targets or two catalog-only targets keep resolving purely on their own fit as before. Regression test: tests/unit/combo-13870-chars4-overdrops-override-primary.test.ts (RED before the fix — emergency member promoted to position 0 ahead of the override primary; GREEN after). ⚠️ base-red inherited: #14004 — docs env/docs contract (fixed separately in #14022), chatHelpers file-size drift. Not touched by this branch. |
||
|
|
b0955042dc |
fix(providers): map agentrouter GLM thinking.type adaptive to enabled (#13696) (#14043)
AgentRouter routes GLM models through the generic DefaultExecutor, which has
no GLM-specific handling. When the connection's OpenAI-compatible alternate
format is used, a Claude-style thinking:{type:"adaptive"} field survived
stripUnsupportedParams untouched and reached AgentRouter's upstream GLM
endpoint verbatim, which 400s (thinking.type "adaptive" is not supported by
glm models; must be one of enabled, disabled).
Add a mapThinkingType mechanism to paramSupport.ts's STRIP_RULES (in addition
to the existing drop/clamp mechanisms) and scope a rule to provider
"agentrouter" + model matching /glm-/i that remaps thinking.type from
"adaptive" to "enabled", preserving any other thinking fields (e.g.
budget_tokens) — mirroring the same mapping GlmExecutor already performs for
its own provider.
|