Commit Graph

5987 Commits

Author SHA1 Message Date
Diego Rodrigues de Sa e Souza
eaea0347ac fix(executor): guard claude/anthropic buildHeaders against empty credentials and extend dual-Bearer parity for third-party baseUrls (#8653) 2026-08-04 21:36:04 -03:00
Diego Rodrigues de Sa e Souza
37edd74f2d fix(proxy-health): include credentials in proxy health check URLs (#8853) 2026-08-04 21:36:00 -03:00
Diego Rodrigues de Sa e Souza
0b70a14a3b fix(auth): setting first dashboard login password no longer fails with HTTP 400 PASSWORD_REQUIRED (#8950) 2026-08-04 21:35:29 -03:00
Diego Rodrigues de Sa e Souza
28a1f4d1b6 fix(deps): bump transitive deps for 20 Dependabot CVE alerts
Bumps ip-address, hono, fast-uri, socket.io-parser, undici (v6+v7), protobufjs, tar via targeted package.json overrides. Closes 20 Dependabot alerts (2026-08-04). npm audit → 0 vulnerabilities.
2026-08-04 19:08:34 -03:00
diegosouzapw
ed2c4dbab3 fix(deps): bump transitive deps for 20 Dependabot CVE alerts
Bumps ip-address, hono, fast-uri, socket.io-parser, undici (v6+v7),
protobufjs, and tar via targeted package.json overrides.

All patches are lockfile-only (no code change, range already covers).
Verified: npm audit → 0 vulnerabilities.
Note: brace-expansion NOT in overrides (separate major lines need
different patches; each resolved within its parent range).

Co-authored-by: wgordon17 <22222756+wgordon17@users.noreply.github.com>
2026-08-04 18:51:49 -03:00
Diego Rodrigues de Sa e Souza
0965b041fa chore(ci): stop dependabot from grouping ioredis majors with routine bumps (#9425)
* chore(ci): stop dependabot from grouping ioredis majors with routine bumps

ioredis is loaded through a dynamic import in the distributed quota store, so a
breaking major passes build, typecheck and both test suites and only surfaces at
runtime for operators running Redis-backed quota. #9310 grouped ioredis 5.10.1 to
6.0.0 with 9 unrelated production bumps; majors get their own PR from now on.

* docs(changelog): add fragment for #9425

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-04 18:07:16 -03:00
Nick Sullivan
7b8055c7f8 fix(resilience): count STREAM_EARLY_EOF as a provider failure in combo routing (#9251)
* fix(resilience): count STREAM_EARLY_EOF as a provider failure in combo routing

A STREAM_EARLY_EOF is an upstream that accepted the request (HTTP 200), opened
the SSE stream, then closed it without emitting a single non-ping event. The
combo path classified it together with STREAM_READINESS_TIMEOUT through
isStreamReadinessFailureErrorBody(), and the readiness exemption in
shouldRecordProviderBreakerFailure meant the whole-provider circuit breaker
never saw it.

During a provider-wide outage that makes the breaker blind. Over a 7-day window
on our router we recorded 311 of these events, 302 of them on one model, 265
inside the upstream's published incident window — and the provider breaker sat
at CLOSED / failure_count=0 the entire time. Every request kept being dispatched
to the failing provider instead of shedding to the next combo target.

The two codes are different signals. The readiness probe is a pre-flight
liveness check on a connection we have not committed to, so failing it means
"this connection looks stale". An early EOF means the provider took the request
and then failed to serve it. The single-model path already treats it that way:
shouldTripProviderBreakerForResult has no readiness exemption, so a 502 early
EOF trips the breaker there. This makes the combo path consistent.

isStreamReadinessFailureErrorBody keeps matching both codes, because the
transient-retry and round-robin semaphore-cooldown paths in combo.ts do want
identical treatment for both. Only the breaker needs to tell them apart, so the
distinction is added as a narrow predicate and an optional argument rather than
by changing the shared classifier. Omitting the new argument reproduces the
previous behaviour exactly.

Follows the additive-override pattern established by the isProxyUnreachable
work, and leaves the existing exclusions for client aborts and plain 429s
untouched.

* test: register stream-early-eof-breaker in stryker tap.testFiles

The mutation test-coverage gate (check:mutation-test-coverage --strict)
detects unit tests that cover a mutated module but are missing from
stryker.conf.json tap.testFiles, so their mutant kills would not count.

comboPredicates.ts is one of the mutated modules, and the new
stream-early-eof-breaker.test.ts covers it, so the gate correctly flagged
the omission. 8376-econnrefused-breaker.test.ts -- the test this one is
modeled on -- is already registered; this just brings the new file in line.

No production code change.

---------

Co-authored-by: Nick Sullivan <nick@technick.ai>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
2026-08-04 18:07:09 -03:00
Xiangzhe
3440c118e0 feat(usage): show Grok Build billing limits (#9205)
* feat(usage): show Grok Build billing limits

* test(usage): keep Grok quota reset fixture in the future

* fix(i18n): add Grok billing labels to pt-BR

* fix(i18n): add Grok billing labels to Vietnamese
2026-08-04 18:07:02 -03:00
nguyenha935
712910612b fix(db): bundle and verify the sql.js fallback (#9044)
Co-authored-by: nguyenha935 <208228297+nguyenha935@users.noreply.github.com>
2026-08-04 18:06:55 -03:00
Dizzle
0538eec05e test(dashboard): drop stale next-intl mock breaking ProviderDetailPageClient smoke (#9150)
The local vi.mock("next-intl") predates the #7935 global polyfill and returns a
useTranslations without .rich, crashing t.rich() in ProviderParamFilterSection:199.
The global polyfill (backed by the real createTranslator) now covers this file;
assertions only check DOM/fetch, never translated text.

Co-authored-by: Max <maxmad64@gmail.com>
2026-08-04 18:06:46 -03:00
Dizzle
0ca25d61f4 fix(dashboard): apply provider Auto Sync per connection and fan out the master toggle (#9149)
* fix(dashboard): add per-connection autoSync toggle handler

* fix(dashboard): render per-connection autoSync toggle in ConnectionRow

* fix(dashboard): wire canAutoSync into ConnectionsListPanel

* fix(dashboard): wire per-connection autoSync toggle into provider page

* fix(dashboard): make master autoSync toggle all-on with fan-out

* docs(dashboard): add changelog fragment for per-connection autoSync

* fix(dashboard): correct disable toast and assert fan-out classification

* test(dashboard): pin fan-out classification branches symmetrically

* docs(dashboard): fill changelog fragment with PR number

* fix(dashboard): port autoSync i18n keys to vi and pt-BR locales

* fix(dashboard): localize autoSync keys across all 43 locales

---------

Co-authored-by: Max <maxmad64@gmail.com>
2026-08-04 18:06:35 -03:00
Bob.Hou
8027c60726 test(sse): expect the trailing period in the no-credentials message (#9392)
#9275 started appending a candidate-alias hint to the zero-active-credentials
error and terminated the provider name with a period, so the two sentences read
as one message. The two vscode tokenized-route tests still assert the old
unterminated string and now fail on every pull request opened against this
branch.

The Quality Gates workflow only runs on pull_request to release/**, never on
push, so the branch itself never re-runs these shards and the drift stayed
invisible after the merge.

Assert what the handler actually produces. Keeping the comparison exact rather
than loosening it to a prefix match is deliberate -- the exact form is what
caught the drift.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
2026-08-04 17:08:14 -03:00
Diego Rodrigues de Sa e Souza
4a3dcf6b0b fix(routing): only let Codex-native bare ids preempt a provider when codex is active (#9447)
* fix(routing): only let Codex-native bare ids preempt a provider when codex is active

#9275 widened CODEX_NATIVE_UNPREFIXED_MODELS from a single id to gpt-5.5 plus the
gpt-5.6-sol/terra/luna tiers, so bare Codex CLI ids would reach the ChatGPT
subscription instead of fanning out to whichever provider won the inference race.
The early return it added never consulted the active-provider set, which made the
codex-only guard 30 lines below unreachable for every id in the set:

  if (CODEX_NATIVE_UNPREFIXED_MODELS.has(modelId)) return { provider: "codex", ... }

An OpenAI-only install therefore had bare gpt-5.5 routed to codex and failed with
'no active credentials for provider: codex' on a model OpenAI serves, and an install
whose codex connection was merely inactive failed identically. This also silently
reverted #5887's compatibility boundary.

The preference now only PREEMPTS another provider when a codex connection is active.
Ids that no other provider catalogs (codex-auto-review) still resolve to codex with no
connection at all — there is nothing to preempt and 'no codex credentials' is the
honest error. With codex active the preference still beats OpenAI, which is the point
of #9275, and an explicit openai/ prefix overrides it either way.

Tests: the three assertions that encode the intended #9275 change now expect codex
(plus a new one pinning the explicit-prefix override); the rest were already correct
and pass again untouched. Adds a regression test for the OpenAI-only case.

* docs(changelog): correct fragment id to #9447

* test(routing): seed an active codex connection in the bare-precedence guards

The two files #9275 added assert that bare gpt-5.5 / gpt-5.6-sol reach codex, but
they ran against an empty database — so they also pinned 'codex wins with no codex
connection at all', which is the regression #9447 removes. That put them in direct
contradiction with plan3-p0 / chat-helpers / codex-gpt55-routing-5887, which assert
openai for the very same input: no implementation could satisfy both, which is why
the release could not go green.

Seeding an active codex connection keeps the contract these files were written to
guard (codex beats openai for a Codex-native bare id) while dropping the accidental
'even with no codex configured' half. Cases that need no connection are left as they
were: the tier-only ids and codex-auto-review have no alternative provider to preempt,
and the explicit-prefix overrides are unaffected.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-04 17:08:08 -03:00
Diego Rodrigues de Sa e Souza
16ed707148 feat(providers): filter detail connections server-side (#9247)
* feat(providers): filter detail connections server-side

Filter provider detail requests at the database boundary while preserving
the full per-provider connection set needed by search, pagination, and bulk
actions. Alias-backed provider pages keep their existing aggregate behavior.

Co-authored-by: RobertsXML <RobertsXML@proton.me>
Inspired-by: https://github.com/decolua/9router/pull/2998

* chore(changelog): fragment for #9247

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: RobertsXML <RobertsXML@proton.me>
2026-08-04 14:34:22 -03:00
Dizzle
8b97ef99aa fix(db): persist account egress IP into proxy_logs (#9291)
* fix(db): persist account egress IP into proxy_logs

The account egress IP (outbound IP the upstream saw, resolved via proxyEgress.ts
echo-IP probe with 5-min cache) was computed and surfaced in the proxy_logs
console and ring buffer, but never persisted: proxy_logs.egress_ip did not
exist, so the value was lost on restart and real traffic could not be
attributed to the node/IP active at that instant.

- migration 134 adds proxy_logs.egress_ip (nullable, backward-compatible)
- schemaColumns.ensureProxyLogsColumns() idempotent reconciler
- proxyLogger self-heals the schema in loadFromDb(), persists egress_ip on
  INSERT, and matches it in search
Follows the session_tag (#8249) migration + schemaColumns reconciler pattern;
base SCHEMA_SQL untouched.

* docs(changelog): add 9291 fragment for proxy_logs egress_ip

---------

Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
2026-08-04 14:34:17 -03:00
Dizzle
e50f2329dc fix(lib): memoize catalog pricing/capability lookups to fix cold /v1/models freeze (#8697) (#8987)
Root cause: a cold GET /v1/models catalog rebuild froze the entire server 41-54s.
node --prof profiling found a systemic missing-memoization pattern — a per-model
function rescanning a static or synced data structure with Object.entries()/
Object.keys() (or hitting SQLite) on every call instead of once per rebuild. Fixed
6 instances of the same pattern, found by iteratively re-profiling the full catalog
sweep after each fix (plus a whitebox review pass) until no further hotspot of this
shape remained:

1. getModelsDevPricing() (modelsDevSync.ts) — re-ran a synchronous SQLite query and
   re-JSON.parse'd ~180 blobs on every call (up to ~6091x instead of once per
   request). Memoized via the existing modelCatalogCacheVersion invalidation signal
   (same pattern as getCachedRawProviderConnections/getCachedProviderNodes in
   db/readCache.ts). Dominant cost of the original 41-54s freeze.

2. findInsensitive() (modelMetadataRegistry.ts, resolveCatalogPricing) — rebuilt a
   full Object.entries() scan on every case-insensitive lookup miss, twice per
   model. Replaced with a lowercase-key index built once per distinct pricing
   object and cached by identity (WeakMap). Warns once at index-build time on a
   case-insensitive key collision instead of silently discarding the second value.

3. getSyncedCapability() (modelsDevSync.ts) — ran a per-model SQLite SELECT on cold
   cache instead of self-warming the whole-table cache; no caller in the
   /v1/models build path ever primed it, so a cold rebuild ran one SQLite
   round-trip per model per call site. Now self-warms via the existing bulk
   getSyncedCapabilities() on first miss. Measured as the dominant remaining cost
   after fixes 1-2 (~70% of a full catalog sweep).

4. getCanonicalModelSpecId() (shared/constants/modelSpecs.ts) — up to 3 separate
   linear scans over the static MODEL_SPECS table per call (exact ci, alias ci,
   prefix). Replaced with a lazy, lowercase-key index built once (MODEL_SPECS never
   changes at runtime); prefix-match iteration order preserved exactly so
   resolution outcomes are unchanged.

5. getStaticSpecCanonicalModelId() (modelCapabilities.ts) — duplicated the same
   exact+alias scan as (4) in a second, separate rescan. Now reuses the shared
   index via a new exported helper (findModelSpecIdByExactOrAlias) instead of
   maintaining a second cache over the same static table.
   reverseModelsDevProviders() (modelCapabilities.ts) — rescanned
   Object.entries(MODELS_DEV_PROVIDER_MAP) (also static) on every call; memoized
   by provider key. Result is frozen (readonly) since it is now shared across
   calls instead of freshly allocated each time.

6. resolveModelAlias() (shared/constants/modelSpecs.ts) — rescanned
   Object.entries(MODEL_SPECS) unconditionally once per model (verified 1:1 call
   ratio, no short-circuit). Case-sensitive exact match (Array.includes(), no
   .toLowerCase()) — uses a dedicated exact-match index, deliberately not the
   case-insensitive alias index from fix 4/5 (would silently broaden matches).

Measured on a 1940-pair real-catalog sample (static PROVIDER_MODELS registry):
cold sweep 828ms -> 356ms after fixes 3-5 on top of 1-2, extrapolating to roughly
1s on the real ~6091-model catalog, down from the original 41-54s freeze.

Complementary to the stale-serve fix in #8801 (upstream) — neither alone
eliminates the freeze.

Tests: call-count regression guards for every fix (DB prepare / Object.entries /
Object.keys call counts staying constant instead of scaling with iteration count),
plus correctness coverage for case-insensitive/case-sensitive resolution. All
pre-existing consumer suites re-verified passing (96 tests total across 19 files).

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-04 14:34:10 -03:00
Xiangzhe
455906c181 fix(reasoning): forward Ollama Cloud thinking (#9290) 2026-08-04 14:34:02 -03:00
Bob.Hou
b6bcc491bc fix(token-refresh): exempt transient errors from exponential backoff (#9242)
* fix(token-refresh): exempt transient errors from exponential backoff

A refresh that failed on a network timeout was treated exactly like one
that failed on a revoked token: the streak incremented and the circuit
backed off exponentially, up to four hours. A brief upstream blip could
therefore park a healthy account for the rest of the day.

Transient failures now take a flat two-minute retry window instead of
advancing the streak. Classification checks structured signals first
(err.name for AbortError/TimeoutError, then err.code and err.cause.code)
and only falls back to matching the message text, so it does not depend
on upstream wording. Everything else keeps the existing exponential path.

Two properties worth preserving on sight:

  - A transient failure never shortens a longer permanent backoff. The
    new window is only adopted when the existing one is not already
    further out.
  - testStatus is preserved on both paths, so a connection whose access
    token is still valid keeps serving requests while its refresh
    retries.

Only a successful refresh clears the circuit. A successful request does
not, because requests do not refresh tokens.

* chore(quality): rebaseline file-size for tokenHealthCheck.ts

src/lib/tokenHealthCheck.ts lands at 1021 lines, above the 1000 cap. The
file consolidates token-refresh health checking that was previously split
across auth.ts and tokenRefresh.ts, and the refresh circuit state machine
does not divide cleanly, so splitting it to satisfy the cap would cost
more than it buys.

Scoped to this file only. Baseline entries for files this branch does not
touch are left at their upstream values.
2026-08-04 14:33:54 -03:00
NOXX - Commiter
45d375aa0b fix(api): defer media body size limits to providers (#8843)
Image and video payloads vary by provider and base64 encoding adds substantial overhead. Exempt media routes from OmniRoute's global request-body cap so provider-specific validation determines whether a request is too large. Keep finite body limits for non-media routes and cover both header and streamed-body admission paths.
2026-08-04 14:33:48 -03:00
Diego Rodrigues de Sa e Souza
f11d883f22 fix(cli-tools): enable Apply for compatible providers (#9250)
* fix(cli-tools): resolve models for compatible providers

Keep the CLI tools Apply flow usable when a dynamic OpenAI-compatible or
Anthropic-compatible connection has no static catalog entry. Resolve its
public prefix, connection default model, and prefix-backed catalog entries
before gating the cards.

Co-authored-by: lazysaltyfish <7127935+lazysaltyfish@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2995

* chore(changelog): fragment for #9250

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: lazysaltyfish <7127935+lazysaltyfish@users.noreply.github.com>
2026-08-04 10:06:37 -03:00
Diego Rodrigues de Sa e Souza
2cb77bbca7 fix(translator): harden Claude format detection for model validation (#9253)
* fix(translator): harden Claude format detection for model validation

Co-authored-by: Ervareza Naurian <rianskp644@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/2949

* chore(changelog): fragment for #9253

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: Ervareza Naurian <rianskp644@gmail.com>
2026-08-04 10:06:31 -03:00
Shixi Li
a8216c92fe fix(sse): preserve error-only stream diagnostics (#9022)
* fix(sse): preserve error-only stream diagnostics

* test(ci): register stream readiness mutation coverage

* chore(changelog): finalize PR 9022 fragment
2026-08-04 10:06:25 -03:00
ikelvingo
c790b57af8 fix(translator): pass output_config.effort=max through verbatim (#9053)
The claude->openai translator was unconditionally rewriting max to xhigh, which broke any OpenAI-shape upstream that accepts max literally (e.g. ollama-cloud, opencode-go deepseek, moonshot k3, native Claude). Provider-aware effort policy is owned by sanitizeReasoningEffortForProvider in the executor; the translator should only do form conversion.

Regression guard: tests/unit/base-executor-sanitize-effort.test.ts end-to-end case (claude -> ollama-cloud preserves max).
2026-08-04 10:06:18 -03:00
dependabot[bot]
9acf79f04f chore(deps): bump github/codeql-action from 4 to 4.37.3 (#9082)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4 to 4.37.3.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4...v4.37.3)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-04 10:06:10 -03:00
dependabot[bot]
09665ab455 chore(deps): bump docker/login-action from 4 to 4.5.2 (#9081)
Bumps [docker/login-action](https://github.com/docker/login-action) from 4 to 4.5.2.
- [Release notes](https://github.com/docker/login-action/releases)
- [Commits](https://github.com/docker/login-action/compare/v4...v4.5.2)

---
updated-dependencies:
- dependency-name: docker/login-action
  dependency-version: 4.5.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-04 10:06:03 -03:00
Bob.Hou
9ee6435f0e fix(classify): recognize Modal 'usage limit reached' as quota exhausted (#9079)
Modal-hosted OpenAI-compatible endpoints (self-hosted Kimi K3 via
Modal free tier) return HTTP 429 with body {"error":"usage limit
reached"} when the account's credit is exhausted. Previously no
QUOTA_PATTERNS regex matched this bare-string error shape, so the 429
fell through to rate_limit (60s short cooldown). Combined with combo
round-robin's per-conversation session stickiness (#3825), this kept
re-targeting the same exhausted connection every turn instead of
locking it out and failing over to an account with remaining credit.

Add a substring pattern matching the JSON key/value pair
"error":"usage limit reached" with tolerance for trailing
punctuation and whitespace. Only the exact "error" key matches;
different keys or qualified transient messages like "Per-minute usage
limit reached" stay classified as rate_limit.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
2026-08-04 10:05:56 -03:00
Bob.Hou
edd9b0d664 fix(combos): include id column in getCombos query (#8905)
getCombos() SELECT was missing the id column, so returned combo objects
had their id come only from the JSON data blob. If the data blob lacked
an id field, callers (including the Dashboard) saw null — making the
combo appear to have no primary key and impossible to delete.

Add id to the SELECT so the database column value is always available.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
2026-08-04 10:05:48 -03:00
Jade Guo
ba353aa3d6 docs(db): specify MySQL conformance semantics (#8947)
* docs(db): specify MySQL conformance semantics

* docs(db): deepen MySQL conformance specification

* docs(db): close MySQL conformance gaps
2026-08-04 10:05:41 -03:00
Bob.Hou
224bc0a5a5 docs(guides): add Antigravity (Google One AI) onboarding guide (#8904)
Signed-off-by: Minxi Hou <houminxi@gmail.com>
2026-08-04 10:05:35 -03:00
MumuTW
2e5854906d docs: slim AGENTS.md (#8839) 2026-08-04 10:05:29 -03:00
Will Gordon
2e4268003a fix(ci): merge-queue tolerance for Build (advisory), drops paid-tier batching (#9233)
* fix(ci): restores dast-smoke queue tolerance, drops batching

PR #7329 (an unrelated cliproxy feature PR) silently reverted two prior
Mergify fixes when it touched .mergify.yml from a stale branch:

- #7225's tolerance for the advisory dast-smoke check, which hangs
  recurrently on GitHub-hosted runners (issue #7226) and had been
  dequeuing every queue attempt it touched.
- #7220's removal of batch_size/batch_max_wait_time, which is a paid
  Mergify tier feature this repo's free plan does not have (the queue
  command fails outright with it set).

Restores both fixes verbatim. No PR has used the queue label since
#7329 landed two weeks ago, so this had gone unnoticed.

* fix(ci): restores the auto-enqueue merge_protections_settings block

The first pass of this fix missed a second piece #7329 clobbered in
the same diff hunk: merge_protections_settings.auto_merge_conditions,
the actual mechanism that puts a queue-labeled PR into the queue (the
older rules-based autoqueue path it replaced is EOL). Without it, the
queue label was a no-op even after restoring the check-failure
tolerance and dropping batching.

.mergify.yml now matches commit 9875ccf4e (the last known-good state
before #7329) byte-for-byte, confirmed via sha256.

* fix(ci): retargets queue tolerance from dast-smoke to Build (advisory)

Evidence review found the prior fix's dast-smoke exception is stale:
dast-smoke has failed only twice ever, none since 2026-07-13 (0/30 in
the last ~3.3h across many PRs). Meanwhile Build (advisory), added to
quality.yml 2026-07-27, has a 100% failure rate on every sampled PR
since — confirmed via job logs to be the same class of runner hang
(dies mid "Creating an optimized production build", never a real
compile error), just in a check dast-smoke's tolerance never covered.

Retargets the merge_conditions exception accordingly so the queue can
actually tolerate the failure mode it faces today, instead of one
that's been dormant for weeks.
2026-08-04 08:57:11 -03:00
Diego Rodrigues de Sa e Souza
7163081f5e fix(agentrouter): retry on 400 content-blocked + burst guard (#9323)
The agentrouter.org upstream WAF returns 400 content-blocked
intermittently when:
  1. messages[].content contains a blocked keyword (Lorem ipsum, the
     phrase 'language model' alone, 'virtual assistant', etc.); or
  2. Requests from the same IP/key arrive in a burst, after which the
     WAF's per-IP suspicion bucket starts blocking content that would
     normally pass. The bucket relaxes after ~5-10s of idle.

Apply three mitigations:

1. Burst guard (open-sse/services/wafRateLimit.ts)
   Per-bucket (provider+url) gate that enforces a 500ms minimum gap
   between outbound requests to agentrouter. Configurable via
   configureWafRateLimit(). Tested in tests/unit/wafRateLimit.test.ts.

2. Reactive retry (BaseExecutor.WAF_RETRY_CONFIG in base.ts)
   New WAF_RETRY_CONFIG with maxAttempts=2, delayMs=1500,
   backoffMultiplier=2. When the upstream returns 400 with a body that
   matches /content[_-]blocked/i, retry the same URL with exponential
   backoff (1.5s, 3.0s) before falling through to the 429/401/fallback
   chain. Tested in tests/unit/base-executor-waf-retry.test.ts.

3. Documentation (docs/security/AGENTROUTER_WAF.md)
   Blocklist of always-blocked and almost-always-blocked patterns,
   behavior under load, guidance for prompts/tool output, and pointers
   to the relevant code paths in OmniRoute.

These are belt-and-suspenders: the burst guard prevents the WAF from
activating on normal traffic, and the reactive retry recovers when it
does anyway. Together they should eliminate the intermittent
400 content-blocked that Claude Code sees when running through
agentrouter via OmniRoute.

Refs #9275 follow-up. Test: 'WAF retry config shape' and 'WAF retry
differs from generic' guard the WAF_RETRY_CONFIG contract so future
refactors don't accidentally collapse the two retry paths.

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-03 18:22:14 -03:00
Diego Rodrigues de Sa e Souza
a72e1656eb fix(routing): bare model ids route to codex first; validate synced candidates (#9275)
* fix(routing): bare model ids route to codex first; validate synced candidates

Two bare-model-routing bugs surfaced in the field when an OmniRoute
deployment had a codex subscription whose cookie quota was exhausted
(retry-after 429047s / ~5 days) AND an active kiro connection whose
upstream sync briefly advertised 'claude-opus-5' before kiro vendored
it into the static registry.

  1. Bare 'gpt-5.6-sol' (and friends) routed to the codex provider even
     when the user had explicitly configured 'agentrouter' as their
     provider (via model_provider in codex CLI). With codex in cooldown,
     every bare request 429'd. Fix: extend CODEX_NATIVE_UNPREFIXED_MODELS
     to include the full gpt-5.6-sol tier set + gpt-5.5 + the related
     codex-native ids. The Codex CLI default is now actually honored;
     users can still prefix 'agentrouter/gpt-5.6-sol' to opt into a
     specific provider.

  2. Bare 'claude-opus-5' silently routed to 'kiro' when kiro's synced
     /v1/models catalog had that id (likely from a transient upstream
     quirk). kiro's static registry never cataloged claude-opus-5, so
     the upstream call 404'd. Fix: validate activeSyncedProviders against
     MODEL_TO_PROVIDERS before merging them into the candidate list.
     Auto-discovery still wins when the model id has no static entry
     (brand-new models from upstream keep working).

Bonus: when handleNoCredentials returns a 404 'No active credentials for
provider: X' error, surface the top-3 candidate aliases (e.g.
'anthropic/claude-opus-5, claude/claude-opus-5, agentrouter/claude-opus-5')
so the operator can pick a working prefix instead of staring at a wall.

Tests (all pass, 25 regression tests preserved):
  - tests/unit/fix-bare-model-precedence.test.ts (7 tests)
  - tests/unit/fix-synced-model-validation.test.ts (3 tests)
  - tests/unit/fix-error-message-candidates.test.ts (3 tests)
  - tests/unit/fix-bare-routing-fallback.test.ts (7 tests)

* fix(tests): replace lorem ipsum with neutral text to avoid agentrouter WAF

The agentrouter.org WAF blocks requests containing 'lorem ipsum' in
messages[].content. When Claude Code reads test files via the Read tool,
the content appears in tool_result blocks which can trigger the filter.

Replace 'lorem ipsum dolor sit amet' with 'example content for testing
purposes' in compression harness test to avoid false positives.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-03 18:09:31 -03:00
Diego Rodrigues de Sa e Souza
84b1e5e12f docs: use hard links, not a symlink, for a worktree's node_modules (#9059)
The worktree-isolation recipe in `CLAUDE.md` told agents to symlink `node_modules` from the main checkout. That silently breaks the dev server.

Turbopack refuses a symlink that resolves outside the project root, so `npm run dev` dies with a FATAL panic while typecheck, lint and both test runners keep passing — the message names "filesystem root", not the worktree, so it reads like a Next/build problem.

`cp -al` gives the same benefit (no per-worktree npm install) without the defect: ~5s for the 4.4 GB tree and near-zero extra disk. Verified on this very worktree: same inode, link count 2.

Mirrored into the two translated CLAUDE.md copies that carry the command (zh-CN translated, pl left in English to match its surrounding section).
2026-08-02 20:39:02 -03:00
Diego Rodrigues de Sa e Souza
92e8960f77 feat(models): functional gateway mirrors + fix synced-substitution (#9217)
* fix(models): preserve static registry models not covered by synced discovery

* feat(models): add functional gateway mirror synthesizer

* feat(models): add functional gateway mirror gate predicate

* feat(models): add functional gateway mirror DB gate

* feat(models): wire functional gateway mirrors into /v1/models

* refactor(models): extract synced-coverage helper to pure leaf (file-size gate)

* fix(db): re-export functional gateway mirrors gate from localDb (db-rules)

* fix(i18n): translate functional gateway mirror flag for Vietnamese (locale completeness)

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-02 20:35:58 -03:00
Diego Rodrigues de Sa e Souza
743ccc895a docs(marketing): track Cheaper Inference link clicks via ?utm_source=omniroute (#9258)
Adds the `?utm_source=omniroute` tracking parameter to every public-facing URL where Cheaper Inference is clickable.

- README: the two `<a href>` targets in the Open Source Friends table row (logo + "Get an API key" CTA)
- `gateways.ts`: the `website` field and the `apiHint` text

Not changed on purpose: `api.cheaperinference.com/*` endpoints (technical, not clicks), JSDoc mentions (descriptive text), and the `<sub>cheaperinference.com</sub>` label under the logo (plain text, not a link).
2026-08-02 20:27:53 -03:00
Diego Rodrigues de Sa e Souza
35405be602 fix(agentrouter): infer protocol from client endpoint
fix(agentrouter): infer protocol from client endpoint

- /v1/responses resolves AgentRouter as openai-responses
- /v1/chat/completions resolves AgentRouter as openai
- /v1/messages resolves AgentRouter as claude
- Per-request protocol and credential cloning (no SQLite mutation)
- Codex 0.146.0 and Claude Code 2.1.220 identity alignment
- response.completed.usage.total_tokens normalization for strict Codex clients

Closes #9224
2026-08-02 15:10:54 -03:00
Diego Rodrigues de Sa e Souza
cfb44dab91 Merge pull request #9213 from diegosouzapw/fix/responses-usage-short-circuit
fix(responses): avoid Codex usage normalization short-circuit
2026-08-02 10:06:12 -03:00
diegosouzapw
d53f9bd813 fix(responses): avoid usage normalization short-circuit 2026-08-02 10:05:03 -03:00
diegosouzapw
e16865e394 refactor: remove unused dynamic.ts and source.config.mjs files from .source directory 2026-08-02 08:40:36 -03:00
diegosouzapw
8846f08c73 feat(gitignore): add .source/dynamic.ts to ignore list 2026-08-02 08:40:36 -03:00
Diego Rodrigues de Sa e Souza
b532894a3e docs(readme): Affiliates Promo section (AgentRouter coupon) (#9194)
* docs(readme): add Affiliates Promo section (coupon for AgentRouter)

New collapsible <details> block under '🤝 Supported by our Open Source Friends',
labeld 'Affiliates Promo' to keep it visually and editorially separate from
actual sponsors. Sized at ~half (icon 32px, sub-tag text) so it does not
compete with the partner block above.

First entry: AgentRouter — $100 signup credit (per FREE_TIERS.md), free server
with higher latency, first-class support since v3.8.50. Models surfaced:
claude-opus-4-8, claude-opus-5, gpt-5.6-sol — with a live-list link to
agentrouter.org/v1/models so users can verify.

Clear caveat: 'Affiliate link — OmniRoute has no sponsorship or partnership
with this provider.' — and a footer inviting more coupons via issue.

* docs(readme): leave Affiliates Promo expanded by default

The block only has a single entry today; collapsing it would hide the
AgentRouter coupon from a casual skim. Add 'open' to <details> so the
content is visible on first load. Users can still collapse it manually.

* docs(readme): drop 'see live list at agentrouter.org/v1/models' sentence

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-02 02:32:26 -03:00
Diego Rodrigues de Sa e Souza
7b2e4b4837 fix(responses): normalize terminal usage for Codex (#9192)
* fix(responses): normalize terminal usage for Codex

* refactor(responses): reduce stream gate growth

* refactor(responses): keep stream within size ratchet

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-02 02:29:21 -03:00
diegosouzapw
fc35dc248f feat(gitignore): add .source and .playwright-cli to ignore list 2026-08-01 19:39:09 -03:00
diegosouzapw
ec150a0069 fix(agentrouter): honor alternate protocol in chat pipeline 2026-08-01 13:38:25 -03:00
diegosouzapw
564c204efe fix(agentrouter): support Claude and Codex protocols 2026-08-01 12:18:09 -03:00
Diego Rodrigues de Sa e Souza
b38f3a4c02 feat(test:scoped): add TIA-based local test runner (#8084 D1) (#9143)
- npm run test:scoped: runs only tests impacted by your changes
- npm run test:scoped:staged: for staged changes (pre-commit)
- Uses select-impacted-tests.mjs with impact map when available
- Falls back to heuristic (changed test files) when no map
- Hub file changes suggest full suite
- 7 unit tests for the selection logic

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-01 11:32:04 -03:00
Diego Rodrigues de Sa e Souza
0ef50886ef feat(g1): rewrite combo-strategy check to runtime-import approach (#9131)
G1 (v3.8.51): section (2) of check-known-symbols no longer regex-scans
strategy === "..." literals from combo source. The handled set now comes
from a runtime-imported dispatch registry (open-sse/services/combo/
strategyDispatch.ts) that imports the real ordering functions and enumerates
which strategies they implement. This keeps the canonical-not-handled gate
correct under the upcoming R0.3 registry dispatch, which removes the
strategy === branches the regex relied on.

- Adds HANDLED_COMBO_STRATEGIES registry (all 20 canonical strategies) + binds
  the real dispatch leaves (applyStrategyOrdering, resolveAutoStrategyOrder,
  tryFusionDispatch, tryPipelineDispatch, resolveComboTargetPipeline).
- main() imports the registry instead of reading/sourcing combo files.
- extractHandledStrategies + diffComboStrategies stay exported (pure, tested).
- New TDD test proves the runtime enumeration covers canonical exactly.

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-01 11:08:49 -03:00
Diego Rodrigues de Sa e Souza
8fac6bcd48 feat(.50): completa itens restantes — G13, G14, gap34, docs, R0.2 (#9126)
* feat(ci): G0 — quality rail (PR→release/**) ganha ratchets+segurança do trilho A

O refactor de god-files do trilho 3.8.50→3.9.0 acontece em PRs→release/**, e esse
trilho pulava o motor de ratchet, o CodeQL ratchet e todos os scanners de segurança
— exatamente onde a rede era necessária (5 das 13 causas da reconciliação de 07-24
eram regressões reais shipadas por CI verde por-PR).

Modo enxuto, jobs EXISTENTES (a .51 consolida lanes; nenhum job novo):

- lint-guard: quality:collect + ratchet --allow-missing + require-tighten +
  check:codeql-ratchet. O job já escreve .artifacts/eslint-results.json, então o
  motor entra a custo ZERO de ESLint (um inventário, dois consumidores). Coverage
  ausente degrada gracioso (--allow-missing); autoridade de coverage segue no
  trilho A. + permissions security-events:read para o CodeQL ratchet.
- fast-gates: check:cycles, check:lockfile, duplication, dead-code, type-coverage,
  compression-budget + install endurecido dos scanners (gh release download,
  zizmor PINADO 1.25.2 = mesmo auditor do ci.yml) + secrets/vuln/workflows/
  openapi-breaking com --ratchet (self-skip sem binário; só regressão medida bloqueia).
- Fora de propósito: bundle-size (self-skip sem build → configuração morta) e o
  run de coverage (fast-unit já roda a suíte cheia).

Runners intocados: guard tests/unit/vps-runner-variable-scope.test.ts verde;
teste novo tests/unit/quality-rail-gate-membership.test.ts pina a MEMBERSHIP dos
gates no trilho B (red antes da edição, green depois).

Validação no tip puro (nenhum base-red fabricado para a fila de PRs abertos):
cycles OK · lockfile OK · duplication 4.26% (base 5.72%) · dead-code 226 (base
227) · type-coverage 94.13% (base 92.17%) · compression OK · secrets 0 (base 0) ·
vuln 5 (base 10) · codeql 0 (base 0) · oasdiff 0 (base 0) · zizmor 178 (base 190)
· actionlint exit 0 no arquivo editado · quality-ratchet 56 métricas OK +
require-tighten OK com --allow-missing.

Refs #8084

* feat(.50): G13 golden-set, G14 import boundaries, gap34 deterministic, docs sync, R0.2 dead hooks

Integra os itens restantes da 3.8.50:

- G13: golden-set determinístico para combo.ts e chatCore.ts via seams públicas
- G14: no-restricted-imports para localDb barrel fora de src/lib/db/ e executors em src/app/
- Gap34: teste determinístico de timeout DuckDuckGo sem rede real
- Docs: golden path de contribuição + sincronização de números canônicos
- R0.2: remoção dos 7 hooks mortos do BUILTIN_EVENTS + UI marketplace ajustada

* fix(r0.2): remove marketplace tab remnants from plugins page — fixes dashboard typecheck regression

* chore(r0.2): remove pluginWorker.ts, signing.ts, sandbox.ts — zero importers confirmed

* fix(docs): remove OMNIROUTE_PLUGINS_ALLOW_EXEC reference — env var removed with pluginWorker.ts in R0.2

* fix(env): remove dead OMNIROUTE_PLUGINS_ALLOW_EXEC from .env.example — consumer removed in R0.2

* fix(test): update sidebar-visibility assertion for R0.2 marketplace removal

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-01 10:44:23 -03:00
Diego Rodrigues de Sa e Souza
45c91e22c2 feat(ci): G0 — quality rail (PR→release/**) ganha ratchets+segurança do trilho A (#9108)
O refactor de god-files do trilho 3.8.50→3.9.0 acontece em PRs→release/**, e esse
trilho pulava o motor de ratchet, o CodeQL ratchet e todos os scanners de segurança
— exatamente onde a rede era necessária (5 das 13 causas da reconciliação de 07-24
eram regressões reais shipadas por CI verde por-PR).

Modo enxuto, jobs EXISTENTES (a .51 consolida lanes; nenhum job novo):

- lint-guard: quality:collect + ratchet --allow-missing + require-tighten +
  check:codeql-ratchet. O job já escreve .artifacts/eslint-results.json, então o
  motor entra a custo ZERO de ESLint (um inventário, dois consumidores). Coverage
  ausente degrada gracioso (--allow-missing); autoridade de coverage segue no
  trilho A. + permissions security-events:read para o CodeQL ratchet.
- fast-gates: check:cycles, check:lockfile, duplication, dead-code, type-coverage,
  compression-budget + install endurecido dos scanners (gh release download,
  zizmor PINADO 1.25.2 = mesmo auditor do ci.yml) + secrets/vuln/workflows/
  openapi-breaking com --ratchet (self-skip sem binário; só regressão medida bloqueia).
- Fora de propósito: bundle-size (self-skip sem build → configuração morta) e o
  run de coverage (fast-unit já roda a suíte cheia).

Runners intocados: guard tests/unit/vps-runner-variable-scope.test.ts verde;
teste novo tests/unit/quality-rail-gate-membership.test.ts pina a MEMBERSHIP dos
gates no trilho B (red antes da edição, green depois).

Validação no tip puro (nenhum base-red fabricado para a fila de PRs abertos):
cycles OK · lockfile OK · duplication 4.26% (base 5.72%) · dead-code 226 (base
227) · type-coverage 94.13% (base 92.17%) · compression OK · secrets 0 (base 0) ·
vuln 5 (base 10) · codeql 0 (base 0) · oasdiff 0 (base 0) · zizmor 178 (base 190)
· actionlint exit 0 no arquivo editado · quality-ratchet 56 métricas OK +
require-tighten OK com --allow-missing.

Refs #8084

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-01 09:55:02 -03:00