Commit Graph

2732 Commits

Author SHA1 Message Date
backryun
8a35c45029 fix(quality): align UI test fixtures to current component contracts (base-red vitest)
- setup-wizard: provide required serverState prop (component gained it in a merged PR)
- grok-device-oauth-modal: next-intl stub resolves grok flow keys to EN labels
- provider-quota-widget: label now inline (PR #8916 removed AutoRefreshButtonLabel extraction) — test the widget
- use-provider-connections-cursor-refresh + phase1f: match /api/providers?provider=<id> query form; hoist heavy dynamic imports to module scope (timeout flake)
- home-topology: mock next/navigation useRouter (component added node-click navigation)
- cooling/lobe/AutoComboCatalog: raise cold-import describe timeouts to 30-60s
- request-logger-*: align to current detail-view contract
2026-08-11 22:38:23 -03:00
diegosouzapw
200233d301 fix(quality): resolve open-sse type errors + catalog/build regressions (base-red round 3)
Storm-merge splices repaired in the base-fix PR #10131:
- doctor.ts: AppConfig missing brokerSocketPath
- conol-web.ts: Buffer not assignable to BodyInit (Uint8Array)
- tinycms.ts: TinyCmsExecutor.execute return matches BaseExecutor (response/url/transformedBody)
- tinycmsSigner.ts: encodeInto never-narrowing guard + dead wasm URL fallback (Turbopack)
- virtualFactory.ts: options slot for resolutionSnapshot
- bottleneckPatch.ts: insufficient-overlap casts (as unknown as)
- imageCombo.ts: narrow handleImageGeneration union result
- browser-worker.ts: AppConfig + turn.capabilities splice
- conolDiscovery.ts: getProviderOutboundGuard from Policy module
- catalog.ts: drop removed SWR hooks (getCatalogStaleWhileRevalidateMs + accessors), CatalogCachePolicy -> inline settings, resolve 4-arg call
- catalogCache.ts: remove dead inFlight/promise refs
- chat.ts: add isProviderBreakerFailureStatus import
- model-catalog-cache-swr-8728.test.ts: align to #9199 new API (policy injection removed)
2026-08-11 21:09:39 -03:00
diegosouzapw
b1e151f3cc fix(quality): green release/v3.8.50 base-reds round 2 — gateways/conol/deepai corruption, migrations, docs, ratchets, dashboard-typecheck
Base-red fix for issue #9985 after the 2026-08-11 merge storm (99 PRs).

Real defects fixed:
- gateways.ts: close regolo entry (was swallowing naga-ac + chatanywhere from #9421), drop stale duplicate chatanywhere entry (#9594)
- conol-web + deepai registry: correct ../shared import depth + deepai executor:default
- modelSelectModalHelpers: close isProviderModelHidden (#9011)
- driverFactory.test.ts: restore eaten test-closing brace (#9173)
- usageTracking: remove duplicate cache_* props
- modelCapability{Overrides,ResolutionSnapshot,Capabilities}: max_token -> max_output_tokens (#9199 vs #8908) + test align
- videoGeneration: drop duplicate handleFalVideoGeneration import (mediaGeneration/fal canonical, #9982)
- responseSanitizer: cast input_tokens_details before .cached_tokens access
- EditConnectionModal: missing alibaba code fields, hoist validationPsd, providerPageHelpers Badge variant union
- FreeBudgetCard: t() -> labels.noApiKey
- peerRouting + cliRuntime: ProcessEnv typing
- image-combo.test.ts: type any -> unknown
- fal.test.ts: moved to tests/unit/services (collected path) 14 tests green
- remove duplicate 143_job_registry.sql (146 canonical), KNOWN_GAPS fix

Docs/ratchets (owner-authorized rebaselines, annotated):
- CHANGELOG 3.8.50 living section restored + 42 i18n mirrors
- MCP-SERVER.md 104->105 tools + i18n
- ENVIRONMENT.md/.env.example: ADOBE_FIREFLY_CHROME_HEADED + DEBUG_CLAUDE_NONSTREAM
- fabricated-docs allowlist: TELEGRAM proposal env vars
- file-size: 5 grown files + proxyFetch 1207->1220
- dead-code 230->248, codeql 2->9 (drift from merged PRs, not this PR)
- untrack _tasks symlink; agent-skills-sync --apply (config-codex-cli)
2026-08-11 18:23:48 -03:00
Alex Jordan
7ca73697b0 fix(i18n): localize hardcoded web UI copy (#9245)
* fix(i18n): localize hardcoded web UI copy

* test(i18n): cover hardcoded UI regressions

* chore(changelog): add PR 9245 fragment
2026-08-11 10:37:50 -03:00
Xiangzhe
def5334768 feat(api-manager): add provider-level model permissions (#9313)
* feat(api-manager): add provider-level model permissions

Persist canonical provider wildcards alongside exact model grants and
preserve explicit restricted-empty deny-all semantics across API, SQLite,
JSON import, sync, runtime policy, and the dashboard.

Invalidate filtered model catalogs on permission changes and guard against
stale in-flight catalog builders repopulating invalidated cache entries.

* fix(api-manager): show provider and model counts separately in summary

Provider wildcard selections (provider/*) are no longer counted as
individual models in the Selected Models Summary. The header now shows
"N providers · M models" when both are present, or just the non-empty
category when only one type is selected.

* fix(api-manager): separate provider and model permission displays

* fix(api-manager): separate provider wildcard permissions in UI
2026-08-11 10:29:28 -03:00
Aman
3898305df0 fix(rate-limit): separate queue wait from execution timeout (#9164)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-11 10:19:21 -03:00
Bob.Hou
47c819df66 fix(combo): network errors must not trip provider circuit breaker (#9342)
* fix(combo): keep queue/network timeouts out of the provider breaker

A single-model network error (ECONNREFUSED / proxy_unreachable) means we never
reached the provider — the provider may be healthy while only the network path
is broken. OmniRoute's own rate-limit queue timeouts are backpressure we
applied, not an upstream failure. Neither should trip the whole-provider
breaker.

- chatPredicates: the single-model path excludes proxy_unreachable and
  RATE_LIMIT_QUEUE_* from the provider-breaker trip.
- accountFallback.recordProviderFailure: isQueueTimeout short-circuits before
  the breaker ever counts (combo.ts already flags it from errorText).
- chat.ts: the queue/network guard on the allRateLimited _onFailure trip.

Deliberately leaves the combo same-provider dead-proxy leg (#8376) intact:
there a proxy_unreachable on the next same-provider target must still be able
to open the breaker, or a dead proxy burns every attempt until the 503
max-retry limit.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* fix(resilience): dedup same-provider network errors per event

Same-provider combo targets can all fail the same single network event (a VPN
blip) within one request. Without a dedup each target counts once toward the
provider breaker, so one transient blip opens the whole-provider breaker while
the provider is healthy — the antigravity outage this branch originally chased.

recordProviderFailure now keeps a short per-provider window (10s) for
proxy_unreachable failures: the first network error in a window counts, the rest
of that window are the same event and return. A genuinely dead proxy keeps
failing across requests (past the window) and still accumulates to its
threshold, so the #8376 dead-proxy protection is not weakened.

Covered by tests/unit/breaker-network-error-guard.test.ts: same-window errors
dedup to one, cross-window errors still open the breaker.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
2026-08-11 10:17:31 -03:00
Bob.Hou
fb83f43fca fix(rate-limit): patch Bottleneck doExpire capacity leak (#9328) 2026-08-11 10:14:41 -03:00
Chewji
a367bf62f5 fix(providers): scope model-level targetFormat to declaring provider catalog (#9994)
Model-level targetFormat is provider-scoped endpoint semantics: a catalog entry
declares how the DECLARING provider serves the model. getModelTargetFormat()
fell back to getGlobalModel() when the provider's own catalog lacked the model
id, importing another provider's tag into every provider serving that id.

catalog. command-code serves gpt-5.6-luna over its chat-shaped /alpha/generate
endpoint but inherited that tag, so chatCore translated the request to Responses
format (messages -> input). CommandCodeExecutor.buildCommandCodeBody reads
chat-format input.messages -> undefined -> [] -> upstream 502 "Invalid prompt:
messages must not be empty" (call log 1786341194167-774a5b).

Fix: resolve the provider alias (mirroring getProviderModels), only apply the
provider's OWN catalog entry's targetFormat, and skip the global fallback when
the provider has a catalog. Catalog-less providers keep the global fallback
unchanged; ghe-copilot's Responses routing (#8835) is preserved.

Regression test: tests/unit/provider-models-target-format-scoping.test.ts
(red before the fix, green after).
2026-08-11 10:04:24 -03:00
Will Gordon
4795825513 fix(sse): make Claude effort/no-think catalog variants dispatchable on every provider (#9006)
* fix(executors): route Claude-via-Vertex through native rawPredict with real streaming

Claude models on Vertex AI were being sent through the generic OpenAI-
compatible partner endpoint, which 404s/errors for Claude on at least
some projects. Route them through Vertex's native Anthropic Messages
API (publishers/anthropic/.../rawPredict) instead, stripping the
body-level model field rawPredict rejects and injecting the required
anthropic_version field.

rawPredict only ever returns a complete JSON body, never real SSE
framing, so streaming requests now get a genuine Anthropic-format SSE
stream synthesized from that JSON (message_start/content_block_*/
message_delta/message_stop), which the existing claude-to-openai
response translator already knows how to parse.

Also fixes two response-format resolution bugs that silently dropped
a custom model's DB-stored targetFormat override whenever the model
id also existed in the static provider registry (as claude-sonnet-4-6
and claude-opus-4-7 do under vertex): resolveModelOrError had its own
ad-hoc resolution that never consulted the override, and even once
fixed, executeChatWithBreaker discarded the correctly-resolved format
before handleChatCore's own resolution ran a second time.

* docs: add changelog fragment for #8909

* refactor(sse): extract shared Claude effort-model predicate

* fix(sse): strip Claude effort-suffix ids for any provider serving a real Claude model

* fix(sse): keep no-think and CC-discovery catalog variant roots unprefixed

* fix(dashboard): re-qualify no-think playground model ids correctly

* fix(sse): scope Vertex 404s to a per-model lockout via passthroughModels

* docs: add changelog fragment for the Claude catalog/dispatch fix

* fix(sse): align regex naming and changelog formatting

* fix(sse): clarify effort-variant strip comment and add cross-module drift guard

* fix(sse): disambiguate Vertex connection-wide vs per-model 403s

* docs: document Vertex 403 disambiguation in changelog fragment

* fix(sse): correlate reason and resource within the same ErrorInfo detail

* fix(sse): extract Vertex error classifier and rebaseline frozen file sizes

* test: register vertex-passthrough-model-lockout in stryker tap.testFiles

* fix(sse): reconciles rebase-onto-tip drift for 9006

Two categories of inherited base-branch breakage surfaced when
rebasing onto release/v3.8.50's latest tip, both confirmed unrelated
to this PR's own diff:

- check:file-size: base.ts and chat.ts drifted further past their
  frozen caps via already-merged commits (7163081f5 and others) that
  didn't rebaseline after growing them. Documented and bumped in
  file-size-baseline.json.
- chat-helpers.test.ts: two gpt-5.5 routing assertions predate #9275
  (fix(routing): bare model ids route to codex first), which
  deliberately made gpt-5.5 route to codex unconditionally, regardless
  of which other providers are active. Confirmed via #9275's own
  commit message and code comments this is intentional, not a
  regression; verified reproducible on the raw base tip alone, with
  no changes from this PR involved. Updated both assertions and their
  names to match the new, intentional default.

* ci: re-trigger checks after GitHub Actions incident (2026-08-07, resolved)

* ci: re-trigger checks (previous push event was dropped)

* fix(quality): rebaseline combo-routing-engine.test.ts own-comment growth

The ALL_ACCOUNTS_INACTIVE->ALL_TARGETS_SKIPPED fix (a32aed738) added explanatory comments (+7 lines), pushing the file past its frozen 3457 cap. CI's PR-mode check:file-size caught it; local check-file-size.mjs was not re-run after that specific commit.
2026-08-11 10:02:48 -03:00
Bob.Hou
ca6e944bb1 fix(antigravity): propagate switchAuth signal from 429 engine to retry guard (#9351)
When Google returns a 429 with no parseable retry hint, decide429 correctly
classifies it as short_cooldown_switch_auth (switch accounts). But the
executor discarded that decision, keeping only retryMs=60000. The retry
guard then slept 60s against the same URL/account up to 3 times because
60000 <= LONG_RETRY_THRESHOLD_MS (inclusive boundary).

Plumb a switchAuth boolean through tryResolveRetryFromErrorBody so the
retry guard can decline the sleep branch and fall through to URL/account
fallback immediately.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
2026-08-11 10:01:16 -03:00
Xiangzhe
36959c37d3 [v3.8.50] fix(models): keep model catalogs responsive (#9199)
* fix(models): preserve catalog on affinity bookkeeping

Related to #8697.

Focused follow-up to #8728; this does not replace or supersede that contribution.

* docs(changelog): record model catalog affinity fix

* fix(models): keep cold catalog builds responsive

* docs(changelog): record catalog responsiveness fix

* fix(models): snapshot auto candidate capabilities

* fix(models): invalidate capability catalog snapshots

* test(models): register catalog invalidation coverage

* fix(models): bulk-load catalog capability snapshots

Resolve synced capabilities and persisted overrides from one build-local view instead of repeating per-target SQLite reads. Keep ordinary runtime lookups on demand and preserve catalog generation invalidation.

Refs: #9199

* fix(models): snapshot catalog pricing once per build

Production profiling showed per-model models.dev pricing reads and JSON parsing dominated cold catalog builds. Reuse one build-local pricing snapshot during enrichment and yield before publication so queued health checks can run, while preserving fresh reads for ordinary callers.

* docs(changelog): record catalog pricing snapshot
2026-08-11 09:59:45 -03:00
Bob.Hou
2b2d947faf fix(cache): add latency marker + per-key bypass for semantic cache (#8984)
* fix(cache): add latency marker + per-key bypass for semantic cache

Semantic cache silently corrupts latency measurements: a 10s upstream
call served from cache looks like 19ms. Three fixes:

A. Latency marker: cache HIT responses now carry
   X-OmniRoute-Cache-Latency: synthetic so measurement tools can
   distinguish real vs cached latency.

B. Per-key bypass: new apiKeys.cacheDefaultMode ('legacy' | 'bypass')
   lets latency-sensitive clients opt out of cache reads entirely.
   - DB column + migration (134)
   - rowParser parseCacheDefaultMode
   - API create default + PATCH update
   - checkSemanticCache returns null on bypass

C. Type safety: ApiKeyRow/ApiKeyView/params updated, superRefine
   guard includes cacheDefaultMode.

Cache write path intentionally unchanged: apiKeyId is already in the
cache signature (semanticCache.ts:140), so per-key isolation prevents
cross-key pollution.

Changed test files:
- tests/unit/chatcore-semantic-cache.test.ts (3 new tests)

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* docs: document semantic cache latency impact + bypass configuration

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
2026-08-11 09:55:07 -03:00
JK TAN
16b67f5f68 fix(dashboard): make quota providers expandable (#9025) 2026-08-11 09:54:40 -03:00
Jan Leon
a99c795a67 Add native ChatGPT Web provider for Codex clients (#8949)
* Bypass proxy compaction for native Codex context

* Add native ChatGPT Web provider pipeline

* Add managed browser and tunnel deployment

* Add ChatGPT Web setup and doctor UI

* Document and test ChatGPT Web integration

* fix(security): register chatgpt-web-codex-doctor in LOCAL_ONLY_API_PATTERNS

The diagnostic route under /api/providers/{id}/chatgpt-web-codex-doctor
was not registered in the spawn-capable route guard. Adding it for
parity with the existing /login pattern.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* fix(providers): route chatgpt-web-codex admin routes through a service boundary

The provider CRUD/doctor routes imported chatgpt-web-codex helpers
(finalizeValidatedChatGptWebCodexSecrets, encode/decodeChatGptWebCodexSecrets,
getChatGptWebCodexDoctorStatus) directly from open-sse/executors/**, which
no-restricted-imports (EXECUTOR_IMPORT_RESTRICTION) forbids for src/app/**
files — executor implementations must stay behind an open-sse handler or
service boundary.

Add open-sse/services/chatgptWebCodexAdmin.ts as a thin re-export boundary
(mirroring the existing tokenRefresh.ts re-export pattern) and import from
there instead. No behavior change.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-08-11 09:53:39 -03:00
engmarcosjr
1d33025c70 fix(translator): restore TitleCase tool names on the Claude to Gemini path (#9993)
Gemini lowercases tool names in functionCall responses, so the request
translator must publish a lowercase alias (read -> Read) for
gemini-to-claude to restore the casing Claude Code registered.

claude-to-gemini.ts filtered identity entries (Read -> Read) out of
_toolNameMap, so no alias reached the response translator and
normalizeToolName() - whose REVERSE_MAP is keyed by TitleCase - left the
lowercase name untouched, surfacing as 'No such tool available: read'.

Reuse buildChangedToolNameMap(), which #9568 already introduced for the
openai-to-gemini path.

Closes #9713

Co-authored-by: Marcos Jr <engenheiromarcosjr@gmail.com>
2026-08-11 09:52:40 -03:00
Jan Leon
d774ccecac fix(routing): account for active OAuth sessions (#8940) 2026-08-11 09:51:35 -03:00
Jonathan Bailey
84e83e2f19 feat(providers): add DeepSeek V4 thinking effort aliases (#9485)
* feat(providers): add DeepSeek V4 thinking effort aliases

* docs(changelog): add DeepSeek effort alias entry

* fix(catalog): scope effort-tier fallback to declared models and harden resolver

Addresses reviewer findings on #9485:

- CRITICAL #1: catalog no longer synthesizes unresolvable effort aliases for
  static reasoning models without declared tiers (cheaperinference, cline, etc.)
- CRITICAL #2: tiered static models survive synced-coverage suppression so
  normal installs with synced DeepSeek base models still expose aliases
- WARNING #3: registry suffix resolution short-circuits when the raw id matches
  a direct custom or synced model, preserving custom apiFormat/targetFormat
- WARNING #4: empty synced effort array no longer erases the registry fallback
- WARNING #5: isFlash check is robust to suffixed/prefixed model ids
- Added regression tests for blast radius, custom-model shadowing, none-path,
  and suffixed isFlash

* fix(combos): expose static registry effort tiers in Combo Builder (#9485)

Static provider registry models (e.g. DeepSeek V4 Flash/Pro) declare
supportedThinkingEfforts, but buildModelOptions() only ran
appendSyncedEffortVariants() over DB-synced rows. Synced metadata for a
DeepSeek connection can omit supportedThinkingEfforts, so the catalog/
Playground surfaced the declared aliases while the Combo Builder picker
showed only the bare base ids.

Feed builtInModels with declared effort tiers through the same
appendSyncedEffortVariants() utility used for synced rows, inheriting the
base entry's contextLength/outputTokenLimit/supportedEndpoints/
supportsThinking and preserving its source. DeepSeek is not skipped by
shouldExposeSyncedEffortVariants(), so Flash (none/low/high/max) and Pro
(none/high/max) aliases now appear in the Combo Builder for any connection
whose synced rows omit effort metadata.

Regression test seeds a DeepSeek connection with effort-less synced rows
and asserts the exact alias sets, source preservation, and metadata
inheritance.
2026-08-11 09:50:55 -03:00
Shixi Li
a9d6fd3d9a feat(api): add per-key prompt compression bypass (#10001)
* feat(api): add per-key compression bypass

* docs(changelog): note per-key compression bypass

* chore(db): renumber API key compression migration

* fix(compression): preserve hard kill during adaptive planning

* chore(db): refresh migration gap allowlist
2026-08-11 09:20:41 -03:00
Arul Kumaran
aa4e72097a fix(bun): make server child and outbound fetch Bun-safe (#9761)
* chore(changelog): v3.8.49 reconciliation — 200 missing bullets + 22 restored credits

Phase 0a of /generate-release. Measured commit<->CHANGELOG coverage over the real
cycle range (2c62333b0..HEAD, 933 non-merge commits) instead of the last tag: 180
merged PRs had no bullet at all (they landed without a changelog.d fragment) and a
further 19 were invisible because the merge-train landed them under a generic
'Train 1D: merge via --admin' subject that carries no PR reference.

- +200 bullets, all with PR back-reference and author attribution (1179 -> 1379)
- 🙌 Contributors 156 -> 178; credits @terrafirmbot-source for #7904, which shipped
  through the conflict-resolved #8685 without any attribution
- closed-PR credit audit over the 32 human PRs closed unmerged this cycle: 12 had
  already landed under the author's own follow-up PR and were verified credited
- rollup bullet for the direct release-branch maintenance (merge-train landings,
  ratchet re-pins, base-red sweeps) that carries no PR of its own
- [3.8.49] header dated 2026-07-28 (was TBD) in the root file and the 42 i18n mirrors

Coverage after: 0 commits uncovered.

* chore(quality): v3.8.49 pre-flight — clear 4 base-reds, absorb cycle drift

Pre-flight sweep (Phase 0). Test suites ran on the dedicated 32-core box so the
self-inflicted load of `node --test` could not fabricate timing flakes.

Base-reds fixed (all real, all from merged cycle PRs that did not update their
characterization tests):

- providers-constants-split / quota-plan-registry / provider-translate-path GOLDEN:
  #8861 added the Xiaomi MiMo Token Plan provider, so APIKEY_PROVIDERS is 195 (was
  194), knownProviders() is 12 (was 11) and the translate-path snapshot gains one
  purely additive entry. Counts aligned to the shipped catalog, never relaxed.
- agent-skills-content: skills/config-codex-cli/ was added by #8709 with a custom
  block, so the custom-block set is 13, not 12.
- chatcore-compression-integration: #8595/#8560 deliberately decoupled REACTIVE
  context compaction from the `enabled` master switch, so a body above 70% of the
  window is pruned even with compression off. The test was sized above that
  threshold, which made it assert against intended behavior; it now stays below it
  and keeps testing the invariant it was written for (resolveBasePlan short-circuits
  to "off" before reading comboOverrides).

Static gates:

- 3 shellcheck directives were malformed (`# shellcheck disable=SC2086 — text`; the
  em-dash makes shellcheck reject the whole directive as SC1125) in ci.yml and
  nightly-release-green.yml — the comment now sits on its own line.
- gitleaks: 2 new generic-api-key false positives allowlisted with justification —
  a localStorage key for the sponsor banner (#8723) and the PUBLIC Adobe Firefly
  web x-api-key, whose only literals are in JSDoc (the runtime reads it through
  resolvePublicCred, per Hard Rule #11). secretFindings back to 0.
- zizmor 176 -> 189 and bundleSize 6762 -> 7666 rebaselined with the measurement and
  the reason; both are ordinary cycle drift absorbed at release.

Environment-dependent failures classified out, not silenced: the two tproxy tests
assert the native addon is unavailable/unprivileged and therefore fail when the
suite runs as root on the build box (they pass as a normal user), and the
consoleInterceptor rate-limit test is a 4s-timing flake under load (6/6 isolated).

* test(codex): align the Responses HTTP e2e to the #8507 input-item contract

Fifth and last base-red of the v3.8.49 pre-flight. #8507 (#8083) deliberately sets
`status: "completed"` on Responses input items so strict upstream validators accept
them; codex-chat-reasoning-http-e2e still asserted the pre-#8507 shape, so it failed
against intended behavior. Expectation updated with the reason inline — the assertion
is not relaxed, it now pins the current contract.

The test was never reached in the first pre-flight sweep (the run was interrupted
during the integration phase, and this file sorts after the one that failed).

* docs(release): v3.8.49 feature-documentation sync

Phase 1 step 6b. Swept the cycle's 284 New Features bullets against the existing
docs before writing anything: nearly every large theme (Kimi, xAI OAuth, session
affinity, bun:sqlite, Firecrawl, Opus 5, omniglyph, GCF v3.2, homologation suite)
was already covered. Six real gaps were left undocumented by the PRs that shipped
them, each verified in source before being written up:

- CredentialMaskerGuardrail (#7683) is registered in guardrails/registry.ts but the
  GUARDRAILS table listed only 3 of the 4 guardrails
- the cacheAffinity scoring factor and the cache-optimized combo strategy (#8008):
  the docs still said 12 factors / 18 strategies, the code has 13 / 19
- the optional dashboard OIDC login gate (#6973) — /api/auth/oidc/{login,callback}
  had no mention in AUTHZ_GUIDE
- GET /api/usage/cache-health (#8827) and GET /api/usage/model-latency-stats (#6873)
  were missing from the API reference

README "What's New" gains one bullet (routing transparency) and merges two others
rather than growing a second changelog. PROVIDER_REFERENCE regenerated with the
generator (Firecrawl reclassified to Search, Xiaomi MiMo added by #8861).

check:docs-all green: 134 docs, 813 internal links, no fabricated API/env/CLI
references. Known pre-existing drift left alone and reported: stale nominal counts
in ARCHITECTURE/CODEBASE_DOCUMENTATION (soft), the 9-factor mentions scattered in
AUTO-COMBO, and the auto-combo diagram SVG (the renderer needs a browser this
environment does not have — the .mmd source is updated and the .md says so).

* chore(release): v3.8.49 — clear the release-PR CI in one pass

Every finding from the first full ci.yml run on the release PR, fixed or justified
together so a single re-push clears the board.

Lint / check:route-validation:t06 — three routes read request.json() with no visible
Zod validation. The two proxy-subscriptions routes validated with a hand-rolled
parsePayload(); they now use real Zod schemas (src/lib/proxySubscription/schema.ts)
reproducing the same acceptance rules, error strings and status codes. chat/completions
is the proxy's hottest path and parses the body ONCE on purpose (#4380 OOM crash-loop),
so it now safeParses the ALREADY-PARSED object against a deliberately permissive
structural schema — proven not to change behavior: absent model and model:null still
pass through, role "developer" still reaches 200, a ~300 KB payload is accepted, and
the body is still read exactly once. 25 new tests.

i18n UI value drift — 13 English strings rewritten during the cycle left stale
translations in up to 41 locales (317 pairs). Eleven are genuine rewrites and now carry
the pipeline's __MISSING__:<english> marker so the runtime serves corrected English until
translation catches up; vi forbids that marker by test, so it got a real translation.

PR Test Policy — 33 files flagged. Each was verified against the SOURCE, not the diff:
26 assert reductions are legitimate (mostly the #7866 Qwen OAuth provider removal and the
#8013 Antigravity refactor deleting the surface under test) and are allowlisted with the
PR and the evidence; 5 deleted files have verified replacements. One was NOT legitimate:
#7528's GraphQL->WebSocket migration dropped four muse-spark continuation scenarios whose
logic is still live — connection isolation, cache eviction after a failed turn (the commit
itself says "was missing"), parallel-chat cache collision, and the empty-content guard.
All four are restored against the new transport and each was verified to fail when the
corresponding production mechanism is broken.

Quality Ratchet / openapiCoverage — 36.6% against a baseline of 38: the cycle added routes
faster than the spec. Eight real endpoints are now documented from their route.ts
(usage cache-health and model-latency-stats, the two OIDC endpoints, and the five
proxy-subscriptions paths), bringing it to 38.1%.

Quality Gates (Extended) / zizmor — the runner measures 190 where the devbox measures 189
on the same commit, a delta already recorded in this baseline's history. Baselined to the
runner's number.

Also: the driverFactory better-sqlite3 guard moved from a mid-body t.skip() to a declared
{ skip: <condition> } test option. Same behavior for the optional native dependency, but
the skip now shows up in the report and is distinguishable from a test.skip() that silences
a test outright. Verified under both runners: 15/15 on Node, 14/14 on Bun.

SonarCloud Code Analysis stays red and is not a blocker: sonar.qualitygate.wait=false since
#7038 makes the job informative, the built-in gate cannot be swapped on the FREE plan, and
main has no branch protection.

* chore(quality): close the last two release-PR reds

test-masking — I had missed one of the 34 flagged files: my first pass grepped only
paths under tests/, so open-sse/services/__tests__/tierResolver.test.ts was invisible.
Same #7866 cause as the other eight qwen-driven reductions: the "classifies Qwen as
free" case and qwen's entry in the batch list went with the removed provider, and the
batch indices dropped from 10 to 9 (61→59). Allowlisted with that evidence.

dast-smoke — all four Schemathesis findings are on the two OIDC endpoints documented
in the previous commit, and none is a defect. /api/auth/oidc/* is a BROWSER redirect
flow: it answers 302 to the IdP and 302 back to /login?oidc_error=... on every failure,
which Schemathesis reads as "accepted a schema-violating request", and it answers 400
when OIDC is not configured, which it reads as "rejected a schema-compliant request".
Keeping the endpoints in the spec is right — operators need them, and they are what
brought openapi coverage back over the baseline — so the flow is excluded from the fuzz
instead, with the reason inline in the workflow. The rest of /api/auth and /api/keys
stays in scope.

* test(db): reword the driverFactory skip comment so the gate stops counting it

The anti-test-masking gate greps text, not code: my explanation of WHY the
better-sqlite3 guard moved out of the test body spelled the runner API out
literally, and those two mentions inside a comment were counted as two new skip
markers — the exact signal the previous commit set out to clear. Same explanation,
phrased without the call syntax.

Verified with the gate's own exported helpers against the merge-base: 0 modified-file
violations, 0 deletion violations. Test still 15/15.

* fix(dashboard): unbreak the vitest:ui gate — 2 real production bugs + the i18n test seam

The Vitest job is a BLOCKING gate that had not run to completion once in this whole
release: rounds 1-3 cancelled it via cancel-in-progress on each successive fix push,
so its red was indistinguishable from green. Round 4 finally ran it and the suite was
broken cycle-wide.

Root cause of the suite: #7935 instrumented ~180 shared/dashboard components with
next-intl's useTranslations/useLocale without updating the tests that mount them, so
every one of them threw "context from NextIntlClientProvider was not found". Fixed at
the shared seam (tests/_setup/vitestUiPolyfills.ts) rather than per file: a translator
built from the REAL en.json via next-intl's own createTranslator, memoized per
namespace — the naive version returns a fresh function each call and any component
whose useCallback/useEffect depends on t spins forever, which reads as a hang, not a
failure. A local mock still wins over the default. 22 files fixed by the seam alone,
15 realigned to the real strings; no assert removed or weakened.

Two production bugs the suite was hiding, both pre-existing and both with a failing
regression test already in the tree:

- RequestLoggerDetail crashed on a structured error object. #7920 gave the component
  formatErrorForDisplay for exactly this case, then #8213's combo-503 / cooldown
  checks went to the raw field and called .toLowerCase() on it. Both paths now use
  the helper.
- The logs detail modal reopened on first close again. #6830 fixed that by reading the
  deep-link id ONCE; the #8354 page rewrite regressed it by reading the live
  searchParams every render, so the prop flips mid-session and re-fires the child's
  deep-link effect exactly as the modal closes. Frozen at mount again.

Also tightens i18nUiCoverage 75.5 -> 99, which the ratchet demanded under
--require-tighten: the metric genuinely improved as the async translation workflow
paid off the debt that the v3.8.39/.44/.47 rebaselines had been recording. The
collector subtracts placeholders, so this release's 317 __MISSING__ markers are
already netted out of the 99.

Two UI files still fail locally under 20-worker concurrency (combos-page-smoke,
evals-tab-smoke) — cold-import flakes that pass isolated and with a larger timeout.

* test(e2e): repair the four shards the first green Build finally exercised

test-e2e has `needs: [build]`, and the release PR's Build died on every round
until now — so the 9-shard matrix produced ZERO signal for this whole cycle
while ~200 PRs merged. The first successful Build surfaced four independent
breakages, each traced to the commit that caused it:

- providers-management (#7361): the single-connection delete moved from
  window.confirm() to a ConfirmModal, so page.once("dialog") never fired and
  the DELETE was never sent (deleteCalls stayed 0). Click the modal instead.
- providers-bailian-coding-plan (#7882): the free-text Base URL field was
  deliberately replaced by a region step whose choice resolves the endpoint
  (global-sg -> coding-intl.dashscope, china-beijing -> coding.dashscope).
  Both cases rewritten against the region step; the invalid-URL case is
  unreachable from this modal now, so it covers the CN choice instead.
- group-b-activity-feed: the stack-trace guard ran against page.content(),
  which embeds the serialized i18n payload — zenmux's "endpoint at
  /api/v1/chat/completions" is prose, not a leak. Assert on rendered
  innerText and require the :line:col every real stack frame carries.
- navigation (#8292): APP_ROUTE_PATTERN accepted only /login and /dashboard,
  but the new prefetch spec is the sole caller passing /home, so waitForURL
  never resolved and the retry loop burned the full 180s timeout.

E2E is green on main (9/9 on 07-22 and 07-23), so all four are cycle
regressions, not pre-existing debt. Tests only — no production code touched.

* fix(dashboard): stop the /home quick-start cards from prefetching too

#8292 fixed half the RSC prefetch storm: it added prefetch={false} to the
sidebar's navigation and logo links, but /home — the landing route, and the
one its own e2e guard visits — renders five more internal Links in the
quick-start cards. First paint still fired 12 speculative RSC requests for
/dashboard/{analytics,logs,providers,api-manager} and /docs.

That PR shipped the test that would have caught this, but the test never got
to its assertion: gotoDashboardRoute("/home") hung because APP_ROUTE_PATTERN
accepted only /login and /dashboard, so the retry loop burned the whole 180s
timeout with no assertion error. With that helper repaired in the previous
commit, navigation.spec.ts finally ran and reported the 12 requests.

Validated both ways, per Hard Rule #18:
- tests/unit/sidebar-prefetch-policy-8281.test.ts extended to /home — red on
  the parent commit (5 internal Links, 5 without prefetch={false}), green here.
- the e2e assertion expect(speculativeRequests).toEqual([]) is the end-to-end
  guard; it is what surfaced the defect in the first place.

* refactor(dashboard): shrink HomePageClient back under the size gate

The prefetch fix in the parent commit tripped check:file-size — the frozen
budget for this file is 1377 lines and a naive fix measured 1391, because
`href` + `prefetch={false}` + `className` no longer fits Prettier's 100-column
budget, so three one-line <Link> elements each expanded to five.

Followed the gate's own first suggestion (extract/DRY) before touching the
baseline: the quick-start links repeated the same className literal four
times, and the docs link carried a 180-char one inline. Hoisting both into
INLINE_LINK / DOCS_LINK collapses five wrapped <Link> blocks back to a single
line each and removes the duplication — 1391 -> 1381.

The remaining +4 over the frozen budget is the five prefetch attributes
themselves, which cannot be expressed in fewer lines. Rebaselined to 1381
with the rationale recorded in file-size-baseline.json under
_rebaseline_2026_07_29_8281_home_quickstart_prefetch.

tests/unit/sidebar-prefetch-policy-8281.test.ts still passes (2/2): it matches
whole <Link ...> blocks, so it is indifferent to the wrapping and only checks
that every internal link opts out of prefetch.

* fix(bun): use native fetch for direct outbound requests

* test(bun): cover native direct fetch path

* fix(bun): preload polyfill for next build workers

* fix(bun): expose AsyncLocalStorage globally

* fix(bun): filter non-page Fumadocs metadata

* fix(bun): defer docs-only route dependencies

* chore(skills): sync generated OmniRoute agent skill docs

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-11 09:16:23 -03:00
NOXX - Commiter
f5ce51a9ff feat(providers): add Conol (conol.ai) web session provider (#8974)
* feat(providers): add Conol web support

* fix(conol): preserve sessions and image turns

* fix(conol): pin session model and effort via /model endpoint

Conol ignores agentModel/agentEffort on POST /api/sessions, so every
session silently ran on the downgraded account default (the create
response reports modelDowngraded: true / effectiveModel).

Sessions are now created empty and configured out-of-band against
POST /api/sessions/{id}/model before the first turn is submitted, in the
order the web client uses: modelPreset, then agentModel, then agentEffort.
The ordering is load-bearing because the model call resets agentEffort to
null server-side.

Effort now defaults to xhigh when the caller does not pin one via the
-<effort> model suffix, and is clamped onto the ladder each model actually
advertises, so xhigh degrades to high on claude-sonnet-5 and is skipped
entirely for models without an effort ladder such as openrouter/fusion.

Model and effort are also dropped from the session binding key so switching
models re-pins the existing session instead of stranding it and losing the
conversation history. Re-pinning only happens on an actual change, so
steady-state follow-ups cost no extra round trips.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-11 09:08:23 -03:00
Ababil
55c2b35eb7 fix(codebuddy-cn): replace agent system prompts to bypass Tencent content filter (#9723)
* fix(codebuddy-cn): replace agent system prompts to bypass Tencent content filter

Tencent's content filter flags CLI agent system prompts (e.g. 'You are
Claude Code, Anthropic's official CLI...') as prompt injection / sensitive
content and rejects the entire request with error:

  抱歉,系统检测到您当前输入的信息存在敏感内容,我无法响应您的请求

This patch adds detection and replacement logic to the CodeBuddyCnExecutor:

- Regex-based identity marker detection (Claude Code, Cursor, Windsurf,
  Cline, Aider, Copilot, Cody, etc.) + length catch-all (>2000 chars)
- Handles both top-level 'system' field (Anthropic format) and messages
  array with role:'system' (OpenAI format)
- Preserves original content shape (string vs typed content blocks)
- Strips oversized tool descriptions (>64KB) that can also trigger the filter
- Replaces with neutral prompt, leaving legitimate user prompts untouched

Based on approach from rafilajhh/9router commit 7f7d7ce.

* test(codebuddy-cn): add regression coverage for system prompt replacement

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-08-11 09:06:47 -03:00
NOXX - Commiter
d6543d71ae fix(usage): reject impossible provider token counts (#8927) 2026-08-11 09:00:26 -03:00
AmirHossein Rezaei
c7be8a4870 fix(providers): drop dead Cloudflare Workers AI free catalog IDs (#8717) (#8804)
Four of the original six free-catalog model IDs return 400/403/410 from
Workers AI. Remove them from freeModelCatalog + cloudflare-ai registry,
keep the live replacements from #8763, and move the 30M monthlyTokens
budget onto @cf/meta/llama-3.3-70b-instruct-fp8-fast.

Co-authored-by: MumuTW <42820974+MumuTW@users.noreply.github.com>
2026-08-11 08:58:34 -03:00
Ke Jin
549739a9b7 fix(kimi): apply K3 effort policy to aliases (#10005) 2026-08-11 08:55:41 -03:00
Aman
1423259fec fix(opencode): fallback unsupported DeepSeek json schema output (#9992) 2026-08-11 08:52:00 -03:00
backryun
2b5253da79 fix(types): narrow Claude stream deltas (#9990) 2026-08-11 08:50:10 -03:00
backryun
33d58ed420 fix(types): validate Fal video result URLs (#9989) 2026-08-11 08:48:22 -03:00
backryun
861011c02b fix(types): preserve GHE Copilot executor configuration (#9988) 2026-08-11 08:46:32 -03:00
backryun
425396d11d fix(types): narrow chatCore local contracts (#9987) 2026-08-11 08:44:27 -03:00
backryun
390efaafaf fix(types): narrow chat dispatch contracts (#9986) 2026-08-11 08:42:38 -03:00
backryun
16c566d146 fix(copilot-web): restore browser authentication (#9984) 2026-08-11 08:40:48 -03:00
Xiangzhe
3e2b166869 feat(combo): add quota-only priority fallback (#9983)
Add a per-target priority option that advances only after trusted quota exhaustion while preserving retry, nested Combo, quality, and Global Fallback semantics.
2026-08-11 08:38:55 -03:00
rinseaid
82115193c2 fix(media): support Gemini Omni Flash video (#9982)
* feat(media): add provider-neutral video and music generation

* fix(db): clean audit tables by created timestamp

* fix(media): support Fal-hosted Grok video

* fix(media): route Fal video references to Grok

* fix(media): support Gemini Omni Flash video

* fix(media): use Gemini Omni Flash Fal endpoint

---------

Co-authored-by: rinseaid <rinseaid@rinseaid.net>
2026-08-11 08:37:07 -03:00
Vasily Larin
16bf95fe33 fix(executors): strip redundant oneOf matching sibling enum (#9828)
* fix(executors): strip redundant oneOf matching sibling enum

The Codex private Responses endpoint intermittently returns a 502 upstream_empty_response for tool parameters that combine oneOf:[{const,...}] with a sibling enum containing the same value set.

When the const and enum sets match exactly, oneOf adds no constraint beyond enum. Add stripRedundantOneOfConstEnum to normalizeCodexTools to remove only this semantically redundant form.

The schema-aware recursive walker requires non-empty, unique string const branches containing annotations only, string enum values, and an exact set match. It preserves bare oneOf[const], narrowing or non-matching sets, type-discriminated oneOf, empty oneOf, non-string values, and anyOf/allOf.

Run the normalization after stripUnsupportedRegexPatterns and before assigning tool.parameters. Add focused regression coverage for matching, non-matching, nested, immutable, and Chat-to-Responses cases.

* docs(changelog): update PR number in changelog fragment
2026-08-11 08:35:16 -03:00
Chewji
029a43359c fix(command-code): include tool call arguments (#9821)
* fix(command-code): normalize malformed tool call arguments and fix test assertion handling

* fix(command-code): resolve toolName from assistant calls and update version header to 1.15.1

* refactor(command-code): consolidate pre-pass message tool metadata extraction and add unknown fallback test

* fix(command-code): fallback unnamed tool calls to unknown to satisfy upstream name validation

* fix(db): rename 139_job_registry -> 143 to avoid collision with 139_ccr_blocks

release/v3.8.50 owns version 139 (ccr_blocks, #9061). The #9631 job
registry cherry-pick (5e5919dcc) landed its migration as 139_job_registry,
recreating the version collision that fix 21a3cb32f had already resolved
on the standalone branch. The migration runner throws on startup, which
makes getDbInstance() fail and every route return 500.

Bump the job registry migration to 143 (next free slot; 140 is taken by
connection_runtime_state) so the runner stops throwing. The SQL is
idempotent (CREATE TABLE IF NOT EXISTS + INSERT OR IGNORE), so DBs that
never applied it just pick it up on next boot; no DB can have recorded
version 139 as job_registry because the collision always threw before
any migration ran.

* fix(command-code): emit arguments on tool-result parts to satisfy /alpha/generate schema

* fix(command-code): rename tool names colliding with upstream built-ins to satisfy /alpha/generate result normalization

The upstream server normalizes tool-call/tool-result parts against its own
built-in registry for matching names. A tool named `tool_search` collides
with a server-side built-in, so the result is rejected mid-stream with
`input[N] missing required field 'arguments'` (verified live: renaming the
pair makes the identical request pass; the server pairs each result with the
nearest preceding tool-call, so any result following such a call is affected).

Rename colliding names consistently on the wire (definitions + calls +
results) via a request-scoped toolNameMap, then un-rename on the response
path so the client still sees its original tool names.
2026-08-11 08:31:09 -03:00
backryun
641153f60a provider(agnes):refresh model catalog (#9998) 2026-08-11 07:52:22 -03:00
Diego Rodrigues de Sa e Souza
c61cdc30aa fix(providers): switch minimax from claude to openai format so images work (#9463)
* fix(providers): switch minimax from claude to openai format so images work

The Anthropic-compatible /anthropic/v1/messages endpoint rejects image
input with 403. MiniMax's OpenAI-compatible /v1/chat/completions endpoint
supports image_url natively for MiniMax-M3.

- minimax + minimax-cn: format claude→openai, baseUrl→/v1/chat/completions
- Remove Anthropic-Version header + ?beta=true suffix (not needed for openai)
- Remove minimax/minimax-cn from ?beta=true executor case
- Update cache-control tests (openai format uses different caching path)
- Fix reasoning-split test names (no longer claude format)

TDD: 2 registry tests assert format=openai (red→green).
Refs: Hermes Agent #15715, MiniMax OpenAI-compatible API docs.

* fix(sse): re-align stream-readiness-policy tests with minimax's openai format

PR #9463 switched minimax/minimax-cn from claude to openai format so images
work. The stream-readiness bump for Claude-format replicas is keyed off the
registry's format field (single source of truth), so minimax legitimately
falls out of that group now. Swap the "Claude-format replica" test fixtures
to agentrouter (still format: "claude") and add explicit coverage that
minimax no longer gets the claude_format_heavy_reasoning bump.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-11 07:25:28 -03:00
Diego Rodrigues de Sa e Souza
31fce91fbd feat(providers): add Naga.ac and ChatAnywhere aggregator providers (#6674) (#9421)
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-11 07:23:31 -03:00
backryun
668ed4813f refactor(providers): remove retired GitHub Models (#9023) 2026-08-11 07:19:32 -03:00
jhordanjw123
b4fc835d25 [v3.8.50] feat(providers): add support for TinyCMS Web (#8736)
* feat(providers): add support for TinyCMS Web including WASM-based cryptographic signing and Proof-of-Work emulation

* feat(providers): add unit tests, ESLint suppressions, and fix hardcoded userid for TinyCMS Web

- Add unit tests for WASM init, UUID validation, challenge flow (15 tests)
- Add WASM source comment explaining binary origin
- Replace hardcoded userid with dynamic provider-specific data
- Add ESLint suppressions for no-explicit-any in WASM bridge code
- Add explanatory comments for DOM shim (runtime WASM-bindgen, not test mocks)

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* refactor(providers): extract TinyCMS DOM shims into an explicit setup function

tinycmsSigner.ts installed its window/document/HTMLCanvasElement/
CanvasRenderingContext2D shims for the wasm-bindgen glue as a module-load
side effect. That meant merely importing the module (even transitively,
e.g. through the provider registry from an unrelated test) mutated
global state for the rest of the test process.

Extract the shim installation into setupDomMocks(), which returns a
restore callback:
- initTinyCmsWasm() calls it once before instantiating the WASM module
  (production path — unchanged behavior, still automatic).
- tests/unit/provider-tinycms-web.test.ts now calls it explicitly in a
  `before` hook and restores the previous globals in `after`, so the
  shims never leak into other test files.

As a side effect, replacing five separate `as any` casts with a single
typed `global as Record<string, any>` handle drops the file's
no-explicit-any count from 5 to 1; eslint-suppressions.json updated to
match.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* docs(providers): regenerate PROVIDER_REFERENCE.md for tinycms-web

Mechanical `npm run gen:provider-reference` run after merging release/
v3.8.50 into this branch — the generated table was stale for both the
new tinycms-web entry this PR adds and the release's own cheaperinference
addition. Total providers 290 -> 292, Web Cookie Providers 31 -> 32.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-08-11 04:45:34 -03:00
LeonG606
565b3e0356 fix(sse): preserve client cache boundaries when hoisting system roles (#9457)
Hoisting a mid-conversation `system`/`developer` message into the top-level
`system` field carried its `cache_control` marker along. Anthropic assembles the
cache prefix as tools -> system -> messages, so the marker ended the cached
prefix at the system block and left the accumulated conversation without a
breakpoint: that turn was billed as fresh input and the next one rebuilt the
cache.

`relocateHoistedCacheBoundary` moves the marker to the nearest preceding block
that can carry a breakpoint, skipping thinking blocks, empty text and anything
the upstream normalisation discards or empties out. If that block already
carries the client's own marker, both are kept - unless the hoisted one, now
ahead of the target in `system[]`, would put a 5m breakpoint before a 1h one,
which Anthropic rejects; it is dropped in that case. Either way the breakpoint
count never grows.

normalizeClaudeUpstreamMessages rewrites tool_result and inlined file/document
blocks into plain text after the hoist, which silently discarded any marker on
them - including a relocated one. The replacement block now inherits it.

Both hoisting implementations share the helper; a fix touching only
claudeSystemRole.ts would leave extractSystemMessagesToBody broken, and the
native Claude path reaches the former through normalizeClaudeUpstreamMessages.
Capability-gated hoisting for strict providers (#7293) is unaffected.

Fixes #9436

Co-authored-by: LeonG606 <leongudat01@gmail.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
2026-08-11 04:35:49 -03:00
Kittisak Tangsiri
b13c3cd202 fix(translator): normalize streamed optional tool arguments (#9423)
* fix: preserve Codex cache usage for Claude suggestions

Co-Authored-By: Claude <noreply@anthropic.com>

* fix: normalize streamed optional tool arguments

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Kittisak Tangsiri <kittisak@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-11 04:35:30 -03:00
Dizzle
0cbdc95023 feat(sse): server-side template expansion for combo system prompts (#5501) (#9414)
* feat(sse): server-side template expansion for combo system prompts (#5501)

* fix(quality-gates): register combo-system-prompt-templates-5501 test in stryker tap.testFiles

check:mutation-test-coverage --strict flagged tests/unit/combo-system-prompt-templates-5501.test.ts
as covering src/shared/utils/circuitBreaker.ts without being listed in stryker.conf.json
tap.testFiles, so its mutant kills wouldn't count.

Co-authored-by: maxmad64bis <maxmad64bis@users.noreply.github.com>

---------

Co-authored-by: Max <maxmad64@gmail.com>
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: maxmad64bis <maxmad64bis@users.noreply.github.com>
2026-08-11 04:35:20 -03:00
Chewji
e58c5ee060 fix(command-code): preserve literal max effort for command-code provider (#9257)
* fix(command-code): preserve literal max effort for command-code provider

* test(command-code): type the new sanitizeReasoningEffortForProvider assertions

The 3 new command-code reasoning-effort test cases cast the function's
unknown return value with `as any`, which pushes the file's frozen
no-explicit-any suppression count (48) to 51 and trips the "No new
ESLint warnings" gate. Use a minimal EffortCarrierResult shape instead
of any, matching the fields the assertions actually read.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* test(v1-models): type the API key lookup in the #9320 auth-leak regression test

The release-tip test file added by #9320 used `(k: any)` in an Array.find
callback, which is not covered by config/quality/eslint-suppressions.json
(the file was added after the suppressions snapshot was frozen). That
leaves the "No new ESLint warnings" gate red for any branch that merges
this exact release/v3.8.50 tip, unrelated to this PR's own diff. Fixing
it here with a minimal derived type (Awaited<ReturnType<typeof
getApiKeys>>[number]) unblocks the gate without touching the frozen
suppressions baseline.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-08-11 04:34:46 -03:00
Bob.Hou
f10dca4318 fix(combo): recover provider circuit breaker from HALF_OPEN on success (#9207)
The combo success path called recordProviderSuccess (cooldown-only)
without notifying the circuit breaker. When a provider breaker entered
HALF_OPEN after repeated failures, successful probe requests never
transitioned it back to CLOSED -- the breaker stayed stuck indefinitely.

Production evidence: agy breaker HALF_OPEN with 699 requests at 98%
success rate, never recovering.

Root cause: combo.ts calls recordProviderSuccess from
providerCooldownTracker.ts (resets cooldown failureCount only) but
never calls breaker._onSuccess(). The failure path in accountFallback.ts
calls breaker._onFailure(), creating an asymmetry.

Fix: add recordProviderSuccess to accountFallback.ts as the symmetric
counterpart of recordProviderFailure. Uses getProviderBreaker (not
configureProviderBreaker) to avoid overwriting the breaker's resetTimeout
with default profile values. Calls breaker._onSuccess() for all non-OPEN
states (CLOSED/DEGRADED/HALF_OPEN), matching execute()'s behavior.
2026-08-11 04:34:23 -03:00
Will Gordon
57744aeb14 feat(cursor): proactively renews Cursor sessions and fixes manual refresh (#9173)
* refactor(cursor): extracts token extraction into shared lib

Moves tryIdeAuth/tryAgentAuth and supporting helpers out of the
auto-import route into src/lib/cursor/tokenExtractor.ts, and adds
an agent-cli-state.json fallback candidate path to tryAgentAuth
(alongside the existing auth.json candidate) so the extraction
logic can be reused by the upcoming renewal orchestrator.

* feat(cursor): adds cursor-agent-backed token renewal orchestrator

Builds the renewal orchestrator in src/lib/cursor/renewal.ts: a
bounded, unattended-safe --list-models nudge, a side-effect-free
status availability check, an in-flight spawn lock keyed by
command, and renewCursorConnection() which nudges cursor-agent
then independently re-scrapes the IDE and cursor-agent credential
sources to detect whichever refreshed. Extends cursorAgent.ts's
binary resolution and spawn helper with fixed-paths-only mode and
a SIGKILL follow-up for background use. Adds a generic keyed-mutex
utility (src/shared/utils/keyedMutex.ts) for serializing a
connection's renew-then-persist cycle, and forwards a busy-timeout
through driverFactory's node:sqlite fallback path.

* feat(cursor): proactively renews Cursor sessions in the sweep

Adds src/lib/tokenHealthCheckCursor.ts, sweep-side glue that calls
the renewal orchestrator and persists the result, wired into
tokenHealthCheck.ts's checkConnection() via a new Cursor-specific
branch placed ahead of the generic no-refresh-token fallthrough.
Carves out a non-terminal exception for a Cursor connection that
already landed at testStatus "expired" via the request-time 401
path, excluding permanently-dead account_deactivated connections.
Extends buildRefreshFailureUpdate() with an overrides param so
Cursor's failure path can use a distinct, non-terminal errorCode
instead of the generic refresh_failed/expired taxonomy.

* feat(cursor): adds local-only manual refresh route

Adds POST /api/providers/[id]/refresh-cursor, a dedicated
loopback-only route that calls the renewal orchestrator on demand
for a single Cursor connection, bounded by a 30s per-connection
cooldown. Classifies the new route in LOCAL_ONLY_API_PATTERNS and
closes the manage-scope-bypass gap for dynamic-segment spawn-capable
routes under /api/providers/ via a new SPAWN_CAPABLE_PATTERNS /
SPAWN_CAPABLE_PATTERN_ANCESTORS mechanism, which also retroactively
covers the pre-existing /login route. The existing shared
/api/providers/[id]/refresh route is untouched and stays
remote-reachable for every other provider.

* feat(cursor): surfaces a dismissible cursor-agent nudge

Adds GET /api/providers/cursor/agent-availability, a credential-free
LOCAL_ONLY route returning only { cursorAgentAvailable: boolean },
backed by a 5-minute cached wrapper around the renewal orchestrator's
existing availability check. Surfaces a dismissible dashboard banner
on the Cursor provider page suggesting cursor-agent installation
when it isn't detected, following the existing dismissible-banner
convention. Also fixes a pre-existing bracket character in a
routeGuard.ts comment that was silently truncating
check-openapi-security-tiers.mjs's view of LOCAL_ONLY_API_PREFIXES.

* fix(cursor): wires manual refresh button to the new route

Branches handleRefreshToken to call the dedicated Cursor refresh
route instead of the generic /refresh route, which silently 502s
for Cursor connections today since they carry no refresh token.
Every other provider's refresh behavior is unaffected. Adds the
cursorSessionUnchanged i18n key and syncs it (plus a pre-existing,
unrelated 28-key backlog) across all 42 locale files.

* fix(cursor): addresses Phase 4/4.5 review findings

Restores the legacy stdout/stderr auth-pattern fallback in
checkCursorAgentAvailability() that the plan's Task 2 Step 4
required but the implementation had dropped. Threads an optional
deps parameter through checkCursorConnectionIfNeeded() so its
error branch is reachable in tests, and switches both it and the
manual-refresh route to exhaustive switch statements over the
renewal result. Adds a short-lived host-keyed dedup cache around
tryIdeAuth() so multiple due Cursor connections sharing a host
don't each open the same state.vscdb file in one sweep tick.
Adds opportunistic eviction to the manual-refresh cooldown map,
an outer try/catch to the availability route for defense-in-depth
consistency with the plan's other routes, and corrects a stale
JSDoc claim about the /login route's auth check. Documents the
now-empirically-confirmed agent-cli-state.json schema mismatch
found while validating against a real cursor-agent install.

* docs(cursor): adds changelog fragments for the renewal plan

Adds one fragment per user-facing outcome per changelog.d/README.md's
convention for a PR that both fixes and adds. PR number placeholder
to be filled in once the PR is opened.

* fix(i18n): translates the new Cursor keys into Vietnamese

The i18n:sync-ui run in an earlier commit left __MISSING__
sentinels for the 4 new Cursor keys in every locale, but
Vietnamese has a dedicated completeness test requiring zero
internal missing markers. Provides real translations for
cursorSessionUnchanged, cursorAgentNudgeTitle,
cursorAgentNudgeBody, and cursorAgentNudgeDismiss.

* fix(cursor): addresses quality-gate Layer 1.5 findings

Restores a comment that misrepresented execFile's actual argv shape
after an earlier bracket-removal fix, this time avoiding literal
closing-bracket characters entirely so the openapi checker's naive
array parser can't be broken by either version. Bounds the sweep-
and manual-route-triggered tryIdeAuth() busy-timeout to 250ms
(down from the interactive auto-import path's 2000ms), since both
share the main event loop with all other in-flight requests and
should fail fast on a WAL-lock collision rather than block the
whole instance for up to ~4s. Has the manual refresh route bypass
the sweep's IDE-auth dedup cache so a click always sees a fresh
read, consistent with this plan's existing "manual actions never
see stale cached data" convention. Documents the previously-missing
agent-availability route in ROUTE_GUARD_TIERS.md's spawn-capable
table.

* fix(cursor): adds SIGKILL follow-up to the status-check spawn

Matches the nudge spawn's existing SIGTERM+SIGKILL pattern so an
unresponsive cursor-agent status check can't leak a lingering
process if it ignores SIGTERM.

* docs(cursor): fills in the PR number for changelog fragments

Renames the 3 changelog.d fragments to their PR-numbered filenames and replaces the (#PR) placeholder with #9173, now that the PR exists.

* fix(cursor): corrects changelog fragments to reference PR #9173

The prior commit only staged the git mv rename — a git add invocation with a stale (pre-rename) pathspec aborted before the actual (#PR) -> (#9173) content edit was staged, so the rename landed without the fix it was meant to carry. This captures the actual content change.

* docs(cursor): regenerates the agent-skills catalog for the new route

check:agent-skills-sync (CI's Merge integrity gate) requires SKILL.md files to stay in sync with the live route catalog. Adding /api/providers/cursor/agent-availability in an earlier commit needed a regen this branch never ran.

* chore(quality): rebaselines file-size caps grown by agentrouter merges

Two already-merged agentrouter commits (564c204ef, ec150a006) on release/v3.8.50 grew open-sse/executors/base.ts, open-sse/handlers/chatCore.ts, and tests/unit/chatcore-translation-paths.test.ts past their frozen caps before this PR branched — unrelated to the Cursor renewal changes here. No PR branch is left to fix the growth in-place, so the caps are bumped to the current real sizes, following the existing release-green rebaseline precedent in this file.

* fix(sse): imports getModel helpers from db/models, not localDb

A recently-merged agentrouter commit added a @/lib/localDb import in chatCore.ts, violating the no-restricted-imports rule (Hard Rule #2 — never barrel-import from localDb.ts). Points the import at the owning module, src/lib/db/models.ts, where both functions are actually defined, and prunes the now-stale suppression entry.

* fix(sse): scopes CC-relay anthropic-beta to its own requestDefaults

Two already-merged agentrouter commits widened usesClaudeCodeProtocol()'s native-Claude system-transform block (billing header + selectBetaFlags-derived anthropic-beta) to also run for generic CC-compatible relay connections, not just real claude traffic and agentrouter's own wire-image mimicry. selectBetaFlags() has no visibility into a relay's own providerSpecificData.requestDefaults, so its header replacement silently wiped out an earlier context-1m append and force-included redact-thinking regardless of the relay's own opt-in. Restores both for plain CC-compatible relays only; real claude/agentrouter traffic is unaffected.

Also bumps four stale hardcoded Codex/Claude Code CLI version-string test assertions (0.144.1->0.146.0, 2.1.219->2.1.220) that drifted when the same two commits bumped the version constants without updating their tests, and rebaselines base.ts's frozen file-size cap for this fix's own +35 lines.

* fix(sse): preserves bare CC-relay native treatment and context-1m

The previous commit's fix was too broad in one direction: excluding ALL CC-compatible relays from the native-Claude header block broke two pre-existing tests (cc-compatible-provider.test.ts, v3.6.6) that rely on that treatment for a 'vanilla' relay with no providerSpecificData.requestDefaults configured.

Refines the gate to this whole native-Claude header-replacement block: replace headers for real claude traffic, agentrouter's wire-image mimicry, OR a CC-relay with no requestDefaults at all — only a relay with EXPLICIT requestDefaults (context1m/redactThinking/summarizeThinking) gets to keep buildHeaders()'s own correctly-computed header set. A redact-thinking-beta strip (unconditional, a no-op when native treatment didn't apply) covers the one remaining gap: selectBetaFlags() force-includes it for a bare relay's opaque client, which a bare relay never explicitly opted into.

Verified against all three previously-conflicting pre-existing tests simultaneously: executor-default-base.test.ts's '1M beta' test, both cc-compatible-provider.test.ts SSE-forcing tests, and provider-request-failure-pipeline.test.ts's 'keeps request beta headers' test (the last of which was already broken by the raw agentrouter merge, confirmed via direct comparison against that exact commit).

* fix(sse): fills in remaining stale CLI version literals

The same two agentrouter commits bumped Codex/Claude Code CLI version constants (0.144.1->0.146.0, 2.1.219->2.1.220) without updating every hardcoded test assertion. This round covers the ones the previous version-string commit missed: the anthropic-cache-fingerprint billing-version constant, a cc-bridge-transforms body assertion, the UI-mirror parity test's own snapshot plus its RoutingTab.tsx source of truth, an integration test's User-Agent assertion (inconsistent with its own dynamic Version assertion two lines up), and the translate-path golden snapshot. Also updates a stale doc comment referencing the old literal by value instead of by constant name.

* fix(cursor): imports from db/ modules, not the localDb barrel

Both files violated Hard Rule #2 (never barrel-import from localDb.ts) — a genuine lint error that had gone uncaught locally. refresh-cursor/route.ts imported getCachedProviderConnectionById from @/lib/localDb instead of its owning module, @/lib/db/readCache. tokenHealthCheckCursor.ts copied the same pattern from its sibling tokenHealthCheckCopilot.ts (an existing, already-suppressed violation) for updateProviderConnection; imports it from @/lib/db/providers instead, with no circular-import fallout (verified via the existing token-health-check-cursor and refresh-cursor-route test suites).

* fix(db): removes stale raw-SQL allowlist entry for cursor route

The cursor auto-import route no longer contains raw SQL — that query
now lives in src/lib/cursor/tokenExtractor.ts, outside the
route/handler scope check-db-rules scans. The allowlist entry was
stale, tripping the stale-enforcement gate.

* fix(test): registers cursor test files in stryker tap.testFiles

Three unit test files covering mutation-tested modules
(route-guard-cursor-agent-availability, route-guard-cursor-refresh,
cursor-renewal) were missing from stryker.conf.json's tap.testFiles,
tripping the mutation-test-coverage gate's drift detection.

* chore(ci): retriggers checks (stuck GH Actions runner on shard 2/4)

* fix(sse): restores CC-relay context1m/redact-thinking test coverage

Rebasing onto release/v3.8.50's new tip (35405be60, an unrelated
agentrouter protocol-inference commit) silently flipped two assertions
this branch's own earlier fix (687fbda62) depends on, in the same test
files that commit touched for other reasons:

- executor-default-base.test.ts: calls[0] (a bare CC-relay with no
  requestDefaults) expected redact-thinking-beta absent; flipped to
  present. calls[1] (context1m+redactThinking requestDefaults) expected
  the context-1m beta preserved; flipped to absent.
- provider-request-failure-pipeline.test.ts: expected Accept:
  text/event-stream and the context-1m beta present for a relay with
  explicit requestDefaults; flipped to application/json and absent.

35405be60 did not touch open-sse/executors/base.ts at all, so these
were test-only edits made without visibility into the still-unmerged
CC-relay header-preservation fix on this branch — they quietly matched
the assertions back to the pre-fix (buggy) behavior instead. Restores
the original, validated expectations; all three interdependent test
files (executor-default-base, cc-compatible-provider,
provider-request-failure-pipeline) verified passing together again.

* ci: re-trigger checks after GitHub Actions incident (2026-08-07, resolved)

* ci: re-trigger checks (previous push event was dropped)

* fix(quality): restore dropped vi.json cursor-renewal keys + rebaseline test growth

vi.json was missing 4 keys (cursorSessionUnchanged, cursorAgentNudgeTitle/Body/Dismiss) that this PR's own pre-merge branch had translated -- the original merge's 'git checkout --theirs' resolution for the 7 conflicted locale files discarded them since upstream's vi.json has no cursor-token-renewal feature. Restored from pre-merge tip a38003e30. Also rebaselines combo-routing-engine.test.ts (3457->3464) for the comment growth from the ALL_ACCOUNTS_INACTIVE fix, caught by CI's PR-mode check:file-size.

* chore(tests): drop explanatory comments on ALL_TARGETS_SKIPPED assertions

Kept the assertion value fix (ALL_ACCOUNTS_INACTIVE -> ALL_TARGETS_SKIPPED); the comments were unnecessary. Reverts the file-size baseline bump these comments caused (combo-routing-engine.test.ts back to its original 3457).
2026-08-11 04:31:24 -03:00
Gioxa
ea085a1517 test(mcp): guard Node 24 bundled MCP startup (#9162) 2026-08-11 04:31:06 -03:00
Aman
58fa99406a fix(translator): honor Chat targets for Responses clients (#9161)
Honor explicit Chat targets for Responses-shaped clients while preserving native Responses providers and selecting token fields from the outbound protocol.

Includes focused regression coverage and the required changelog fragment.
2026-08-11 04:30:56 -03:00
3g0r1ch
2e799b33a7 fix: skills & memory — tool-name encoding, schema normalization, warm-cache, combo id, Ponytail catalog (#9058)
* feat(skills): add Ponytail minimalism skill as external catalog entry

- Add 'external' SkillCategory + SkillArea
- Register ponytail (MIT, DietrichGebert/ponytail) in CURATED_SKILLS
- Generator: external skills carry content in custom block, no api/cli body
- Generate skills/ponytail/SKILL.md with original content preserved
- Update catalog test counts 45 -> 46

* fix(skills+memory): builtin handler fallback in executor, skip vector upsert for deleted memories

- skills: Next.js compiles SkillExecutor into multiple chunks (own singleton
  each); route chunk lacked builtin handlers registered at startup via
  instrumentation. execute() now falls back to builtinSkills registry, so
  POST /api/skills/executions works for file_read/web_fetch/etc.
- memory: scheduleVectorUpsert is fire-and-forget and embeddings are slow;
  health-check verify (create->delete test memory) left queued upserts
  failing with 'memory not found' every 30s. Check existence before embedding
  and skip quietly.

* fix(skills): encode tool names with @ and . for providers rejecting them

Skill tools were advertised as 'name@version' (e.g. test-fr2@1.0.0), but
DeepSeek/Groq/OpenAI reject function names not matching ^[a-zA-Z0-9_-]+$.
Names already valid are left untouched; invalid ones are reversibly encoded
as omr_skill_<base64url> and decoded in interception before registry lookup.

* fix(combos): include DB id column in combo records for dashboard links

getCombos() selected only data/sort_order/context_cache_protection, so
combos whose JSON blob lacked an id field returned id: undefined. The
dashboard then linked to /dashboard/combos/undefined and Combo Control
Center failed with 'Combo not found'. Merge the id column into parsed
rows (authoritative, only when the blob has no id).

* fix(skills): normalize flat skill schemas to object schema for Gemini/Claude

Stored skill schemas are flat property maps ({ text: { type: string } }),
which OpenAI-compatible providers tolerate but Gemini
(function_declarations[].parameters) rejects with 'Unknown name ... Cannot
find field'. Wrap bare maps into { type: 'object', properties: {...} } for
all three tool formats.

* fix(skills): warm registry cache before skill injection in chat path

injectSkills() lists the in-memory skillRegistry, which is empty after a
cold start until something calls loadFromDatabase(). The interception path
already warms the cache (#2815); the injection path did not, so skills
were silently skipped (no_enabled_skills) for the first requests after
restart. Warm the cache for the chat owner before injection.

---------

Co-authored-by: Egor <egorich-print@users.noreply.github.com>
2026-08-11 04:30:38 -03:00