Commit Graph

1924 Commits

Author SHA1 Message Date
diegosouzapw
e52842b926 Merge branch 'compression-core' of github.com:Egorich-print/OmniRoute into fix/combine-9115 2026-08-04 08:33:32 -03:00
Diego Rodrigues de Sa e Souza
7163081f5e fix(agentrouter): retry on 400 content-blocked + burst guard (#9323)
The agentrouter.org upstream WAF returns 400 content-blocked
intermittently when:
  1. messages[].content contains a blocked keyword (Lorem ipsum, the
     phrase 'language model' alone, 'virtual assistant', etc.); or
  2. Requests from the same IP/key arrive in a burst, after which the
     WAF's per-IP suspicion bucket starts blocking content that would
     normally pass. The bucket relaxes after ~5-10s of idle.

Apply three mitigations:

1. Burst guard (open-sse/services/wafRateLimit.ts)
   Per-bucket (provider+url) gate that enforces a 500ms minimum gap
   between outbound requests to agentrouter. Configurable via
   configureWafRateLimit(). Tested in tests/unit/wafRateLimit.test.ts.

2. Reactive retry (BaseExecutor.WAF_RETRY_CONFIG in base.ts)
   New WAF_RETRY_CONFIG with maxAttempts=2, delayMs=1500,
   backoffMultiplier=2. When the upstream returns 400 with a body that
   matches /content[_-]blocked/i, retry the same URL with exponential
   backoff (1.5s, 3.0s) before falling through to the 429/401/fallback
   chain. Tested in tests/unit/base-executor-waf-retry.test.ts.

3. Documentation (docs/security/AGENTROUTER_WAF.md)
   Blocklist of always-blocked and almost-always-blocked patterns,
   behavior under load, guidance for prompts/tool output, and pointers
   to the relevant code paths in OmniRoute.

These are belt-and-suspenders: the burst guard prevents the WAF from
activating on normal traffic, and the reactive retry recovers when it
does anyway. Together they should eliminate the intermittent
400 content-blocked that Claude Code sees when running through
agentrouter via OmniRoute.

Refs #9275 follow-up. Test: 'WAF retry config shape' and 'WAF retry
differs from generic' guard the WAF_RETRY_CONFIG contract so future
refactors don't accidentally collapse the two retry paths.

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-03 18:22:14 -03:00
Diego Rodrigues de Sa e Souza
a72e1656eb fix(routing): bare model ids route to codex first; validate synced candidates (#9275)
* fix(routing): bare model ids route to codex first; validate synced candidates

Two bare-model-routing bugs surfaced in the field when an OmniRoute
deployment had a codex subscription whose cookie quota was exhausted
(retry-after 429047s / ~5 days) AND an active kiro connection whose
upstream sync briefly advertised 'claude-opus-5' before kiro vendored
it into the static registry.

  1. Bare 'gpt-5.6-sol' (and friends) routed to the codex provider even
     when the user had explicitly configured 'agentrouter' as their
     provider (via model_provider in codex CLI). With codex in cooldown,
     every bare request 429'd. Fix: extend CODEX_NATIVE_UNPREFIXED_MODELS
     to include the full gpt-5.6-sol tier set + gpt-5.5 + the related
     codex-native ids. The Codex CLI default is now actually honored;
     users can still prefix 'agentrouter/gpt-5.6-sol' to opt into a
     specific provider.

  2. Bare 'claude-opus-5' silently routed to 'kiro' when kiro's synced
     /v1/models catalog had that id (likely from a transient upstream
     quirk). kiro's static registry never cataloged claude-opus-5, so
     the upstream call 404'd. Fix: validate activeSyncedProviders against
     MODEL_TO_PROVIDERS before merging them into the candidate list.
     Auto-discovery still wins when the model id has no static entry
     (brand-new models from upstream keep working).

Bonus: when handleNoCredentials returns a 404 'No active credentials for
provider: X' error, surface the top-3 candidate aliases (e.g.
'anthropic/claude-opus-5, claude/claude-opus-5, agentrouter/claude-opus-5')
so the operator can pick a working prefix instead of staring at a wall.

Tests (all pass, 25 regression tests preserved):
  - tests/unit/fix-bare-model-precedence.test.ts (7 tests)
  - tests/unit/fix-synced-model-validation.test.ts (3 tests)
  - tests/unit/fix-error-message-candidates.test.ts (3 tests)
  - tests/unit/fix-bare-routing-fallback.test.ts (7 tests)

* fix(tests): replace lorem ipsum with neutral text to avoid agentrouter WAF

The agentrouter.org WAF blocks requests containing 'lorem ipsum' in
messages[].content. When Claude Code reads test files via the Read tool,
the content appears in tool_result blocks which can trigger the filter.

Replace 'lorem ipsum dolor sit amet' with 'example content for testing
purposes' in compression harness test to avoid false positives.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-03 18:09:31 -03:00
Diego Rodrigues de Sa e Souza
92e8960f77 feat(models): functional gateway mirrors + fix synced-substitution (#9217)
* fix(models): preserve static registry models not covered by synced discovery

* feat(models): add functional gateway mirror synthesizer

* feat(models): add functional gateway mirror gate predicate

* feat(models): add functional gateway mirror DB gate

* feat(models): wire functional gateway mirrors into /v1/models

* refactor(models): extract synced-coverage helper to pure leaf (file-size gate)

* fix(db): re-export functional gateway mirrors gate from localDb (db-rules)

* fix(i18n): translate functional gateway mirror flag for Vietnamese (locale completeness)

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-02 20:35:58 -03:00
Diego Rodrigues de Sa e Souza
35405be602 fix(agentrouter): infer protocol from client endpoint
fix(agentrouter): infer protocol from client endpoint

- /v1/responses resolves AgentRouter as openai-responses
- /v1/chat/completions resolves AgentRouter as openai
- /v1/messages resolves AgentRouter as claude
- Per-request protocol and credential cloning (no SQLite mutation)
- Codex 0.146.0 and Claude Code 2.1.220 identity alignment
- response.completed.usage.total_tokens normalization for strict Codex clients

Closes #9224
2026-08-02 15:10:54 -03:00
diegosouzapw
d53f9bd813 fix(responses): avoid usage normalization short-circuit 2026-08-02 10:05:03 -03:00
Diego Rodrigues de Sa e Souza
7b2e4b4837 fix(responses): normalize terminal usage for Codex (#9192)
* fix(responses): normalize terminal usage for Codex

* refactor(responses): reduce stream gate growth

* refactor(responses): keep stream within size ratchet

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-02 02:29:21 -03:00
diegosouzapw
ec150a0069 fix(agentrouter): honor alternate protocol in chat pipeline 2026-08-01 13:38:25 -03:00
diegosouzapw
564c204efe fix(agentrouter): support Claude and Codex protocols 2026-08-01 12:18:09 -03:00
Diego Rodrigues de Sa e Souza
0ef50886ef feat(g1): rewrite combo-strategy check to runtime-import approach (#9131)
G1 (v3.8.51): section (2) of check-known-symbols no longer regex-scans
strategy === "..." literals from combo source. The handled set now comes
from a runtime-imported dispatch registry (open-sse/services/combo/
strategyDispatch.ts) that imports the real ordering functions and enumerates
which strategies they implement. This keeps the canonical-not-handled gate
correct under the upcoming R0.3 registry dispatch, which removes the
strategy === branches the regex relied on.

- Adds HANDLED_COMBO_STRATEGIES registry (all 20 canonical strategies) + binds
  the real dispatch leaves (applyStrategyOrdering, resolveAutoStrategyOrder,
  tryFusionDispatch, tryPipelineDispatch, resolveComboTargetPipeline).
- main() imports the registry instead of reading/sourcing combo files.
- extractHandledStrategies + diffComboStrategies stay exported (pure, tested).
- New TDD test proves the runtime enumeration covers canonical exactly.

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-01 11:08:49 -03:00
Egor
3d444c2209 fix(memory): enable agent memory save/update via MCP tools + builtin stream guard
- memoryTools: apiKeyId now optional; falls back to caller principal
  (HTTP auth headers / OMNIROUTE_API_KEY env) so agents can store memory
  without knowing their key id
- memorySkillsInjection: server-side memory_* builtins only injected for
  non-stream requests (stream clients execute tools client-side via MCP)
- memoryBuiltins: memory_save/update/search/delete builtin tools with
  per-provider schemas + interception dispatch
- retrieval: fix toFts5MatchQuery import (ReferenceError on FTS5 path)
- tests: MCP auto-owner fallback, cross-principal isolation, stream guard
2026-08-01 10:56:07 +03:00
Egor
1c3f6dcc90 fix(skills): warm registry cache before skill injection in chat path
injectSkills() lists the in-memory skillRegistry, which is empty after a
cold start until something calls loadFromDatabase(). The interception path
already warms the cache (#2815); the injection path did not, so skills
were silently skipped (no_enabled_skills) for the first requests after
restart. Warm the cache for the chat owner before injection.
2026-07-31 14:30:05 +03:00
Diego Rodrigues de Sa e Souza
371c10ea5f feat(sse): Cheaper Inference provider — chat + native Responses + images, sponsor rail 2nd (#9043)
Registers Cheaper Inference (api.cheaperinference.com) as an OSS-sponsor gateway provider.

- Canonical provider `cheaperinference` (alias `cinf`) + routing registry with 39 measured text models
- Dedicated executor: forces `store:false` on the native /v1/responses endpoint (the shared strip in
  chatCore.ts deletes `store` for every provider != openai, so without this every Responses request
  400'd) and resolves chat-vs-responses URL from the per-model targetFormat
- 3 image models (grok-imagine, nano-banana-pro, nano-banana-2), prefix-only: the two nano-banana ids
  already belong to adobe-firefly, which keeps the bare-id routing
- Resale pricing measured from GET /v1/models (30% off list); sponsor rail Kimi 1st / Cheaper
  Inference 2nd via an explicit rank map; supporter badge in 43 locales; README row

No quota card: the gateway exposes no balance API (/v1/wallet and /v1/balance both 404).

Validated live end-to-end through OmniRoute: chat, native Responses, streaming and image generation
all 200 with real content; the Firefly collision guard verified at runtime.
2026-07-31 07:53:41 -03:00
Egor
2baf2f4820 feat(nvidia): forward quota headers + add quota check API
- responseHeaders: promote NVIDIA NIM quota/usage headers (x-nvcf-*,
  x-quota-*, x-ratelimit-*) to priority 2 so they survive the 768-byte
  upstream header forwarding budget (were previously dropped as priority 3)
- new GET /api/v1/quotas: lists every provider connection with saturation
  (0..1), remaining percent, and data source using existing saturation
  signals — lets operators check key budgets via API instead of live monitoring
2026-07-31 09:32:10 +03:00
Egor
4826385f19 feat(resilience): enable retry for all providers + attempt header + memory by default
- chatCore: maxAttempts 1→2 for all providers (previously only model-scope and codex got retries)
- combo.ts: add x-omniroute-attempt response header showing how many attempts were made
- memory/settings: enable memory injection by default (maxTokens 1000) — auto-extraction already wired in chatCore
2026-07-31 09:13:29 +03:00
Egor
0515e69a3f Merge branch 'pr/8949' into feat/personal-build
# Conflicts:
#	tests/unit/providers-constants-split.test.ts
2026-07-31 08:40:42 +03:00
Egor
00059f7bdd Merge branch 'pr/8914' into feat/personal-build 2026-07-31 08:39:49 +03:00
Egor
b4f0d97f10 Merge branch 'pr/9006' into feat/personal-build 2026-07-31 08:39:36 +03:00
Egor
bf83381794 Merge branch 'pr/9014' into feat/personal-build 2026-07-31 08:39:30 +03:00
Egor
71a8f6f98e Merge branch 'pr/9015' into feat/personal-build
# Conflicts:
#	open-sse/translator/response/gemini-to-claude.ts
2026-07-31 08:39:24 +03:00
Will Gordon
a48f256f51 fix(sse): clarify effort-variant strip comment and add cross-module drift guard 2026-07-30 17:33:48 -04:00
Prudhvivuda
b2ce078c92 fix(sse): preserve Claude Code tool-name casing via Gemini/Antigravity (#9008)
Stop blindly lowercasing PascalCase tool_use names on the Gemini→Claude path so Claude Code no longer rejects Read/WebSearch as missing tools.
2026-07-30 16:32:11 -04:00
Prudhvivuda
8a3888e510 fix(sse): preserve Gemini thought_signature on Claude Desktop tool turns
Claude→Gemini direct translators dropped thoughtSignature, so Gemini 3
tool follow-ups returned 400. Store on the response path, re-attach (or
context-fallback) on the request path, and thread signatureNamespace.

Closes #8979
2026-07-30 16:31:44 -04:00
Prudhvivuda
368d1c0e87 fix(sse): route Poe API-key traffic through DefaultExecutor (#8969)
Stop aliasing canonical `poe` to PoeWebExecutor so API-key requests hit
api.poe.com Chat Completions / Responses / Claude-only Messages instead of
the web GraphQL path that returned HTTP 405.
2026-07-30 16:31:41 -04:00
Will Gordon
c2c622ad82 fix(sse): align regex naming and changelog formatting 2026-07-30 16:07:34 -04:00
Will Gordon
cf2055ce3e fix(sse): scope Vertex 404s to a per-model lockout via passthroughModels 2026-07-30 15:10:47 -04:00
Will Gordon
2a1c946aa6 fix(sse): keep no-think and CC-discovery catalog variant roots unprefixed 2026-07-30 15:00:26 -04:00
Will Gordon
0d2678360f fix(sse): strip Claude effort-suffix ids for any provider serving a real Claude model 2026-07-30 14:51:53 -04:00
Will Gordon
a8fb526e9c refactor(sse): extract shared Claude effort-model predicate 2026-07-30 14:46:45 -04:00
Will Gordon
0e66f7e566 Merge remote-tracking branch 'upstream/release/v3.8.50' into fix/vertex-claude-catalog-dispatch 2026-07-30 14:15:46 -04:00
Diego Rodrigues de Sa e Souza
2c243cf1fc feat(sse): deprecate the gemini-cli upstream provider with a real migration path (#8980)
* feat(sse): deprecate the gemini-cli upstream provider with a real migration path

Stored `gemini-cli` connections were being kept alive for nothing. Measured before
touching anything:

  routable?    absent from PROVIDERS, from REGISTRY, from OAUTH_PROVIDERS, and no
               executor references it → the connection can NEVER serve a request
  refreshing?  yes, and successfully — it redeemed against PROVIDERS.gemini's client
               (681255809395-oo8ft2o…), the same public Gemini CLI / Code Assist OAuth
               client

So the scheduler made periodic upstream calls to Google to keep a credential fresh
that had nowhere to go. That is the waste this removes.

This is a deprecation, not a deletion, and the difference is deliberate. The path was
not dead code: #8232 added it after a user report (the UI advertises automatic OAuth
rotation and these rows never rotated), and #8275 narrowed it to exactly the legacy
refresh. Simply dropping it from `supportsTokenRefresh` would have produced a SILENT
skip — `Skipping … (refresh unsupported)` — leaving the row at "active" forever, doing
nothing. Worse than before.

Instead:

  DEPRECATED_PROVIDERS + isDeprecatedProvider/getDeprecationNotice in tokenRefresh
      one place naming the provider and where to migrate. A test asserts the migration
      target is itself routable, so the notice can never point somewhere useless.

  _getAccessTokenInternal returns the ESTABLISHED unrecoverable envelope
      { error: "unrecoverable_refresh_error", code: "provider_deprecated", migrateTo }
      Reusing `error` means isUnrecoverableRefreshError and the manual-refresh route
      already stop retrying — no new contract for callers to learn. The distinct `code`
      is what makes it legible. A bare `null` would read as transient and retry forever.

  tokenHealthCheck marks the connection terminal with the reason
      Placed after the existing terminal-status guard, which makes it idempotent for
      free: once "expired", later sweeps skip the row, so it writes once instead of
      rewriting the same reason every cycle.

  the manual-refresh route stops lying
      It said "Refresh token expired. Please re-authenticate this account." — false
      here: the token is fine, the provider is gone. Re-authenticating would loop
      against something that no longer exists. It now reports the deprecation and the
      migration target.

`gemini` uses the same OAuth client, so re-adding the account there is a working path,
not advice to start over.

Deliberately NOT touched:

  Category A — the gemini-cli CLIENT identity (#7034): clientIdentityProfiles.ts,
      clientApi.ts, googApiKeyAuth.ts. Same string, opposite direction — requests
      ARRIVING from the Gemini CLI, where OmniRoute is the server. Deleting these is the
      failure this change must never cause, so a test now asserts the profile survives.
      Audited: `git diff --name-only` touches none of those files.

  errorClassifier.ts's isCloudCodeProvider list still names gemini-cli. It is a
      defensive 403→PROJECT_ROUTE_ERROR list shared with cloudcode/cloud-code; the entry
      is unreachable for a non-routable provider, and editing a shared classification
      path for a dead string is risk without upside.

Tests — 42 across the six files that mention the identifier, all green:

    gemini-cli-legacy-refresh.test.ts        5   (3 assertions REWRITTEN, see below)
    gemini-cli-deprecation.test.ts           5   (new)
    client-identity-profiles.test.ts         9   (category A, untouched)
    service-token-refresh.test.ts           14
    errorclassifier-antigravity-403.test.ts  4
    gemini-cli-ansi-sanitization.test.ts     5   (category C, untouched)

The three rewritten assertions in the legacy file are alignment, not weakening, and the
gate is right to ask: each is now STRONGER. "refresh succeeds against Google's token
endpoint" became "zero upstream calls happen at all"; "a 400 surfaces invalid_grant"
became "the envelope is unchanged but the code says provider_deprecated" plus a control
asserting `gemini` still reports invalid_grant, proving the real path was not blunted.
The file's header keeps the whole #8232#8275 → deprecation arc, because each step is
why the next made sense. Count unchanged; no test deleted, so no allowlist entry needed.

* docs(changelog): fragment for #8980

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-07-30 09:36:54 -03:00
Diego Rodrigues de Sa e Souza
888ce1a73b chore(sse): drop the iflow entry from the token-refresh TTL map (#8966)
* chore(sse): drop the iflow entry from the token-refresh TTL map

The `iflow` provider was removed from the product, but its 24-hour refresh-lead
entry outlived it in REFRESH_LEAD_MS. Surfaced during v3.8.49 homologation on the
production VPS, where startup logs carry:

    [CREDENTIALS] Warning: unknown provider "iflow" in credentials file, skipping.

Measured across src/, open-sse/, tests/ and docs/ — the identifier had exactly two
occurrences repo-wide: the map entry and one test assertion. Nothing dispatches on
it, so `getRefreshLeadMs("iflow")` now falls through to TOKEN_EXPIRY_BUFFER_MS like
any other unknown provider.

The test assertion was not deleted, it was MOVED: from "returns explicit lead time
for known providers" to "falls back to TOKEN_EXPIRY_BUFFER_MS for unknown
providers". That is alignment to the new behavior and strictly more coverage than
before — a silent reintroduction of the entry now turns the fallback case red
instead of passing unnoticed. Flagged explicitly because the test-masking gate
rightly treats a removed assertion as suspicious.

Also removes the now-redundant "Non-rotating providers" section header: every
remaining entry under it is Google-backed and the following comment already says
"permanent (non-rotating)".

    node --import tsx/esm --test tests/unit/service-token-refresh.test.ts
    # 14 pass, 0 fail

* docs(changelog): fragment for #8966

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-07-30 09:07:54 -03:00
Jan Leon
8941eafcb9 Add native ChatGPT Web provider pipeline 2026-07-30 06:55:41 +02:00
Jan Leon
b02c586cb4 Bypass proxy compaction for native Codex context 2026-07-30 05:03:35 +02:00
Lucas Israel
347bfe257c fix: validate live Claude Devin bridge 2026-07-29 10:02:02 -03:00
Lucas Israel
71af74bee4 fix: fail closed on incompatible Devin ACP behavior 2026-07-29 10:02:02 -03:00
Lucas Israel
3899495f60 fix: close Devin bridge live runtime gaps 2026-07-29 10:02:02 -03:00
Lucas Israel
aa5856376e feat: harden Devin ACP bridge contracts 2026-07-29 10:02:01 -03:00
Lucas Israel
a52d5fb31f fix: fail closed around Devin ACP execution 2026-07-29 10:02:01 -03:00
Lucas Israel
44ba571521 feat: add initial Devin agentic provider 2026-07-29 10:02:01 -03:00
Will Gordon
1e629721b9 fix(executors): route Claude-via-Vertex through native rawPredict with real streaming
Claude models on Vertex AI were being sent through the generic OpenAI-
compatible partner endpoint, which 404s/errors for Claude on at least
some projects. Route them through Vertex's native Anthropic Messages
API (publishers/anthropic/.../rawPredict) instead, stripping the
body-level model field rawPredict rejects and injecting the required
anthropic_version field.

rawPredict only ever returns a complete JSON body, never real SSE
framing, so streaming requests now get a genuine Anthropic-format SSE
stream synthesized from that JSON (message_start/content_block_*/
message_delta/message_stop), which the existing claude-to-openai
response translator already knows how to parse.

Also fixes two response-format resolution bugs that silently dropped
a custom model's DB-stored targetFormat override whenever the model
id also existed in the static provider registry (as claude-sonnet-4-6
and claude-opus-4-7 do under vertex): resolveModelOrError had its own
ad-hoc resolution that never consulted the override, and even once
fixed, executeChatWithBreaker discarded the correctly-resolved format
before handleChatCore's own resolution ran a second time.
2026-07-28 19:18:44 -04:00
diegosouzapw
ed2db6cb19 chore(release): open v3.8.50 development cycle
Cut from the v3.8.49 freeze tip (parallel-cycle model). Bumps package.json x3,
openapi.yaml, lockfile, adds the living [3.8.50] CHANGELOG section with the three
canonical headings (New Features / Bug Fixes / Maintenance) so aggregate-changelog.mjs
cannot mis-target an older published section, and syncs the 42 i18n mirrors.
2026-07-28 15:45:27 -03:00
NOXX - Commiter
fff11cbb57 fix(adobe-firefly): default gpt-image detailLevel to maximal (5) (#8863)
* fix(adobe-firefly): default gpt-image detailLevel to maximal (5)

GPT Image 2 quality is dominated by generationSettings.detailLevel (1-5).
The SPA often defaults to 3 (medium); missing/auto quality previously mapped
to 3 as well. Default now to 5 (high/max) so API clients and Media without
an explicit quality still get maximal detail. Explicit low/medium still honored.

* chore(quality): rebaseline adobeFireflyClient + changelog fragment

adobeFireflyClient.ts 2317->2322 (+5) — this PR's own growth at the existing
payload-build site. Covered by tests/unit/adobe-firefly.test.ts.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-07-28 14:37:22 -03:00
Diego Rodrigues de Sa e Souza
85b9c1754e feat: Xiaomi MiMo Token Plan provider + per-connection API protocol selector (#8861)
* feat(sse): add alternateFormats registry field and resolver

* feat(sse): honor per-connection targetFormat in getTargetFormat

Registry-driven format lookup now resolves an alternate protocol
declared for the provider when the connection's providerSpecificData
carries a matching targetFormat, falling back to the entry's default
format otherwise.

* feat(sse): resolve base URL from selected alternate format

resolveBaseUrl now falls back to the connection's selected alternate
protocol (providerSpecificData.targetFormat) before the provider's
default base URL, while a manual providerSpecificData.baseUrl override
still wins over both.

* feat(sse): apply alternate format auth header and extra headers

DefaultExecutor's registry authHeader lookup and BaseExecutor's shared
header preamble now both honor a connection's selected alternate
protocol: the alternate's authHeader wins over the registry default,
and its extra headers (e.g. Anthropic-Version) are merged in.

* refactor(sse): extract resolveAlternate helper into BaseExecutor

Centralizes the getRegistryEntry() + resolveAlternateFormat() pair
that resolveBaseUrl, buildHeadersPreamble, and DefaultExecutor's
authHeader lookup each duplicated, so a future call-site can't diverge
from the shared precedence. Also translates the PT-BR comments added
in the previous three commits to match the surrounding English. Pure
refactor — no behavior change.

* feat(sse): declare Anthropic-compatible variant for xiaomi-mimo

The provider publishes the same catalog over /anthropic/v1/messages on the same
host. Selecting it also required bypassing the per-provider URL normalizers in
DefaultExecutor.buildUrl(): normalizeXiaomiMimoChatUrl() appends /chat/completions
unconditionally, which mangled the alternate's already-complete endpoint into
.../anthropic/v1/messages/chat/completions.

* feat(sse): add xiaomi-mimo-token-plan provider with monthly quota

Token Plan is a separate product: tp- keys authenticate only on the regional
token-plan-sgp host and return 401 on api.xiaomimimo.com, where the existing
xiaomi-mimo provider points. Same pattern as qwen-cloud-token-plan.

Registers the monthly token allowance (no balance API upstream) and declares
the Anthropic-compatible variant on the token-plan host.

* feat(dashboard): add API protocol selector to connection modal

Providers that declare alternateFormats in the registry now expose an opt-in
protocol dropdown on the connection modal. The choice persists to
providerSpecificData.targetFormat as an explicit null when set back to the
default, since the PUT route merges { ...existing, ...incoming } and an omitted
key would keep the previous override.

* fix(i18n,quality): vi parity for the protocol selector + own-growth rebaselines

The three new provider keys landed only in en/pt-BR, so the vi parity test failed
(tests/unit/i18n-vi-completeness.test.ts asserts key parity AND no __MISSING__
markers — running i18n:sync-ui would have satisfied the first and broken the
second). Added translated values instead. Scoped to vi: it and pt-BR are the only
locales with a parity test.

Rebaselines are this PR's own growth: EditConnectionModal.tsx 1283->1316 (the
selector field) and open-sse/executors/base.ts 1540->1562 (alternate-format
resolution at the existing buildUrl/headers chokepoint).

Adds the changelog fragment.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-07-28 12:56:12 -03:00
Bob.Hou
ec0edef499 fix(token-refresh): discover projectId during token refresh (#8860)
* fix(token-refresh): discover projectId during token refresh

The token refresh path (tokenRefresh.ts) did not discover projectId
for antigravity/agy accounts. Dashboard and health check refresh use
this path, not the executor path.

Add ensureAntigravityProjectAssigned call after refreshGoogleToken
for antigravity/agy providers when projectId is empty.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* chore(quality): rebaseline the token-refresh test file + changelog fragment

tests/unit/token-refresh-service.test.ts 1311->1378 (+67) — the four cases
covering projectId discovery on the tokenRefresh.ts path. Growth is the tests
this PR adds, nothing else.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-07-28 12:40:08 -03:00
Will Gordon
f1fcdbfa6e fix(executors): route current Claude generations through Vertex partner endpoint (#8852)
* fix(executors): route current Claude generations through Vertex partner endpoint

PARTNER_MODELS pinned three Claude 3.x prefixes (claude-3-5-sonnet,
claude-3-opus, claude-3-haiku). Every newer Claude generation on Vertex
(claude-sonnet-4-6, claude-haiku-4-5, etc.) fell through to the
Google-publisher branch instead, producing an invalid
publishers/google/models/claude-... path.

Replace the pinned prefixes with a single generic "claude-" prefix:
any Claude model on Vertex is always an Anthropic partner model, never
a Google one, so this can't go stale again the way pinned version
strings did.

Fixes #1985

* docs: add changelog fragment for #8852
2026-07-28 11:49:52 -03:00
Bob.Hou
38cee62d42 fix(antigravity): discover projectId during token refresh (#8842)
* fix(antigravity): discover projectId during token refresh

The initial OAuth exchange can fail to populate projectId via
loadCodeAssist (network timeout, account not yet onboarded). The
runtime transformRequest path already recovers via
ensureAntigravityProjectAssigned, but refreshCredentials did not --
after a token refresh the per-token memoization cache is invalidated
(new access token = new cache key), so every subsequent request
triggers a fresh loadCodeAssist round-trip that may fail again.

Add a best-effort ensureAntigravityProjectAssigned call in
refreshCredentials when projectId is empty, mirroring the pattern in
transformRequest. Persist the discovered id so it survives the next
refresh or restart.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* chore(quality): rebaseline antigravity.ts and test file-size

antigravity.ts grew from 1493 to 1528 lines (+35) with projectId
discovery in refreshCredentials. executor-antigravity.test.ts is a new
test file at 1098 lines (above cap 1000) with 4 new test cases.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* fix(antigravity): match ExecutorLog arity in the refresh discovery log

`ExecutorLog.info` (open-sse/executors/base.ts) is `(tag, message) => void` —
two parameters. The discovery log passed a third metadata object, which failed
typecheck:core with TS2554 on antigravity.ts:777. Bind the message to a local
and pass two arguments, matching the sibling warn on the catch branch. Kept to
two lines so the file stays at its frozen size; ran Prettier, which also wrapped
the pre-existing over-width `const msg` line below.

Also adds the changelog fragment for the fix.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-07-28 10:25:03 -03:00
hppsc1215
85f1d78d11 fix(ghe-copilot): route OpenAI-native models via Responses API (#8835)
* fix(ghe-copilot): route OpenAI-native models via Responses API

- Add targetFormat: 'openai-responses' to gpt-5.4-mini, gpt-5.3-codex, gpt-5-mini,
  mai-code-1-flash, and oswe-vscode-prime in ghe-copilot registry.
- Register gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna with openai-responses targetFormat.
- Update GheCopilotExecutor.buildUrl() to route openai-responses and codex models to
  <gheUrl>/responses while keeping Claude and Gemini on /chat/completions.
- Add unit tests verifying targetFormat parity and buildUrl endpoint routing.

* docs(changelog): add fragment for #8835

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Alex <sefias_methue@hotmail.de>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-07-28 10:24:45 -03:00
Tito Fauzan Putra
0809c73430 fix(providers): correct Codex GPT-5.6 context window (#8838)
* fix(providers): correct Codex GPT-5.6 context window (#7702)

* docs(changelog): add fragment for #8838

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* test(providers): align the remaining GPT-5.6 context-window assertions

Three suites assert the pinned GPT_5_6_CODEX_CAPABILITIES contract through the
VS Code and provider-models routes, and still expected 372000. They only surface
in a full run, so the focused loop on this PR stayed green while `npm run test:unit`
failed with five `272000 !== 372000`.

The two conservative-merge cases in provider-models-route-codex keep testing what
they tested: live 999999 still exceeds the pinned value (pinned wins) and live
100000 is still below it (live wins). Only the pinned number and the comments
naming it move.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-07-28 10:24:36 -03:00
rinseaid
d30f484089 fix(reasoning): sanitize streamed K3 think tags (#8821)
* fix(reasoning): sanitize streamed K3 think tags

* refactor(stream): move think-tag helpers into thinkTagParser leaf module

open-sse/utils/stream.ts is frozen at 2887 lines with zero headroom, and the
kimi-coding-apikey think-tag handling added here pushed it to 2951. The parsing
half of that work has no SSE dependency, so it moves to thinkTagParser.ts (an
unfrozen leaf): the open-tag lookahead predicate and the end-of-stream flush
delta assembly. stream.ts keeps only the SSE envelope - enqueue, payload
collection, logging.

Getting back under the frozen ceiling needed more than just the PR's own
added lines, since even the fully self-gating helpers still cost a handful of
call-site lines stream.ts has zero room for. Along the way this also
deduplicates a synthetic chat-completion-chunk literal that was copy-pasted
three times in stream.ts (textual tool-call flush, think-tag flush, terminal
finish_reason synthesis) into one buildSyntheticChatChunk() in
streamHelpers.ts - pure DRY, no behavior change.

Behavior is unchanged: same tests, same counts.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: rinseaid <rinseaid@rinseaid.net>
Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-07-28 01:27:20 -03:00