Commit Graph

7915 Commits

Author SHA1 Message Date
Diego Rodrigues de Sa e Souza
c8dc982eaa fix(ci): drop the stale ESLint cache restore-keys fallback from ci.yml (#11600) (#11996)
Fixes the blocking Lint job's own ci.yml cache: PR #11963 removed the stale restore-keys fallback from quality.yml but left ci.yml's two "Restore ESLint file cache" steps carrying the same prefix-match fallback that lets a cache from a different lint config report stale per-file verdicts. Byte-level parity with #11963's already-merged fix.

Deliberately half of #11600 — the other half (run-eslint-json.mjs) is covered by PR #11983 from a parallel session, so the two don't collide on the same file.
2026-08-29 15:27:39 -03:00
Diego Rodrigues de Sa e Souza
9ec4d39a74 fix(ci): webpack for docker-publish even on omni-build (#12050)
Turbopack had 31 GB on omniroute-113-6 and still panicked
(TurbopackInternalError: there must be a path to a root, run
33253576569). The same tree's arm64 webpack build on hosted ARM
succeeded. Dockerfile already documents webpack as the Docker
escape hatch. Keep amd64 on the one omni-build slot (#12048).
2026-08-29 15:18:33 -03:00
Diego Rodrigues de Sa e Souza
a9aee94a00 docs(ops): the .113 heavy-build ceiling is one runner, not two (#12048)
* docs(ops): the .113 heavy-build ceiling is one runner, not two

Two concurrent next-builds (15.4 GB + 17.2 GB RSS) OOM-killed one on 2026-08-29 17:26 UTC;
systemd booked the kill on the other runner's unit and its job died with the same
"shutdown signal" text a hosted-runner OOM shows. omni-build now lives on
omniroute-113-5 only; 113-6 keeps omni-release. The janitor ceiling counts every
listener on the box (4 OmniRoute + OmniHeuris + OmniMind = 6). The second heavy slot
returns when the Proxmox VM gets more RAM; the exact command is in the doc.

* docs(ops): apply the single-heavy-slot text (previous commit only carried formatting)
2026-08-29 14:41:20 -03:00
Diego Rodrigues de Sa e Souza
47f7e5a306 fix(release): the packaged-app smoke verifies the database opened, not a driver line the primary path never prints (twin of #12032) (#12047)
* fix(release): the packaged-app smoke verifies the database opened, not a driver line the primary path never prints (release/v3.8.51 twin of #12032)

Same change as #12032 on main: the packaged app opens SQLite during the smoke but
its primary open path prints no "[DB] Driver: …" line (only the recovery path and
the sql.js fallback do), so the #7592 assertion failed every Linux release leg. The
guard rejects the sql.js fallback line, accepts a native driver line, and otherwise
accepts demonstrable database activity; after readiness the smoke requests
/api/monitoring/health and waits for that activity outside the readiness loop.
electron-smoke-script suite 10/10.

* fix(release): reapply the smoke rework on top of release/v3.8.51's own copy of the script

The previous commit copied main's file wholesale and dropped this branch's
ensureSmokeEnvDirs(currentPlatform) fix and its tests; this reapplies only the
DB-open evidence change as a patch. electron-smoke-script suite green.
2026-08-29 14:07:58 -03:00
Diego Rodrigues de Sa e Souza
38e2616464 fix(ci): stop hosted docker-publish OOM and unpaint Build (advisory) (#12021)
* fix(ci): stop hosted docker-publish OOM and unpaint Build (advisory)

docker-publish was firing 8 concurrent hosted builds on every merge
storm; each died ResourceExhausted in npm run build (#11976). One
publish per ref, webpack instead of Turbopack so native RSS stays
inside the V8 heap we can cap. Build (advisory) is skipped: continue-on-error
still reports FAILURE and was painting every fork PR red.

Closes #11976

* fix(ci): run docker-publish amd64 on omni-build and share the heavy lane

The .113 box is 31 GB / 32 cores — enough for one next-build. Hosted
ubuntu-24.04 is ~7 GB and ResourceExhausted every publish (#11976).
amd64 now targets [self-hosted, omni-build] (Turbopack) when
USE_VPS_RUNNER is on, joins the existing heavy-build-main group so it
queues beside ci.yml Build instead of becoming a third heavy, and
falls back to hosted + webpack if the VPS is off. arm64 stays on
ubuntu-24.04-arm with webpack (no ARM box).

* test(ci): align the advisory-build contract with the hosted-OOM skip

if: ${{ false }} tripped zizmor obfuscation (194→195). Bare if: false
skips the job without a new finding. The #7307 test now pins the skip
and keeps the job body as the restore recipe.
2026-08-29 09:52:00 -03:00
Diego Rodrigues de Sa e Souza
c4bd8b8ec4 fix(release): electron lockfile resync, build_ref, curated notes and SBOM on dispatch (twin of #11982 + #12020) (#12022)
* fix(release): resync the electron lockfile, build a dispatch from a repaired ref, keep curated notes, attach the SBOM on dispatch (release/v3.8.51 twin of #11982 + #12020)

Same four changes as #11982 and #12020 on main, applied to this branch's own copies:

- electron/package-lock.json regenerated (271 -> 284 entries): the optional
  electron-builder-squirrel-windows subtree was missing and `npm ci` refused the lock
  (EUSAGE) on the Linux and macOS legs; a clean `npm ci --ignore-scripts` on the
  result exits 0.
- electron-release.yml: `build_ref` dispatch input (default: the version tag) and
  `generate_release_notes` only on the tag push (a re-attach dispatch appended
  GitHub's auto notes to the curated body on v3.8.50).
- npm-publish.yml: the SBOM attaches to the GitHub Release on workflow_dispatch
  publishes too, whenever a release for the tag exists.

actionlint and prettier clean; electron-release-desktop-channel-8949,
electron-release-efficiency, electron-release-latest-yml.repro, check-workflows
and npm-publish-artifact-provenance suites pass.

* fix(release): validate build_ref in the validate job before any checkout uses it

CodeQL (actions/cache-poisoning/poisonable-step, high) on release/v3.8.51 — the
default branch: a raw dispatch input checked out next to setup-node's npm cache is a
cache-poisoning vector. The input now goes through the validate job's regex
allowlist (main or release/vX.Y.Z, empty = the version tag) and every build job
checks out needs.validate.outputs.build_ref, never the input itself.

* fix(release): drop the build_ref input — a dispatch builds the ref it is dispatched on

CodeQL (actions/cache-poisoning/poisonable-step) tracks the input through the
validate job's output regardless of the regex allowlist: an input-controlled
checkout next to setup-node's npm cache on the default branch is a cache-poisoning
vector. The ref is not an input any more; the checkouts use github.ref, so
`gh workflow run electron-release.yml --ref v3.8.50 -f version=v3.8.50` rebuilds
the tag and `--ref main` builds the repaired line. The tag-push path is unchanged.
2026-08-29 09:28:03 -03:00
Diego Rodrigues de Sa e Souza
e6de61f0c2 fix(sse): stop the auto-combo candidates inspector from dropping blocked rows (#9133) (#11994)
* fix(sse): stop the auto-combo candidates inspector from dropping blocked rows (#9133)

prepareVirtualAutoComboInputs applied filterResilienceBlockedCandidates
before the #7819 read-only candidate inspector ever saw the pool, so a
model-locked or cooled-down candidate silently disappeared from
/auto-combo/*/candidates instead of showing up as reachable:false with a
reason (modelLocked/connectionCooldown/breakerState were dead fields by
construction). Add an opt-in `skip` parameter so the inspector builds its
own unfiltered pool; routing (createVirtualAutoCombo/createBuiltinAutoCombo
called without a prepared override) is unchanged. Also aligns
isModelLocked's model argument to the bare model id, matching every lock
writer and the routing-side filter, instead of the "provider/model" string.

Regression test: tests/unit/auto-combo-candidates-locked-model-visible.test.ts
(red before the fix — locked account's row silently missing; green after).

* chore(quality): register the #9133 regression test in stryker tap.testFiles

tests/unit/auto-combo-candidates-locked-model-visible.test.ts covers
open-sse/services/accountFallback.ts (via isModelLocked) but wasn't listed,
so its mutant kills wouldn't count toward mutation coverage.

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-29 08:10:40 -03:00
Diego Rodrigues de Sa e Souza
02ba573730 fix(providers): scope Antigravity mitmAlias tier ids to the safe static alias (#11824) (#11988)
Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-29 08:10:37 -03:00
Diego Rodrigues de Sa e Souza
bd04bb9cc6 fix(sse): set X-OmniRoute-Selected-Connection-Id on successful combo dispatches (#11810) (#11986)
Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-29 08:10:33 -03:00
Diego Rodrigues de Sa e Souza
322b218f06 fix(cli): drop the never-produced dist/index.cjs requirement from prepublish's opencode-plugin skip check (#11787) (#11990)
* fix(cli): drop the never-produced dist/index.cjs requirement from prepublish's opencode-plugin skip check (#11787)

* test(build): resolve tsup/npm portably in the #11787 regression test instead of a hardcoded .bin path

The old test assumed @omniroute/opencode-plugin/node_modules/.bin/tsup
already existed. A fresh checkout (CI's npm ci never installs this
standalone package's own deps) has no such node_modules at all, so the
test failed with MODULE_NOT_FOUND in CI while passing locally on a devbox
that had installed it before. Mirror scripts/build/prepublish.ts's own
install-then-resolveLocalBinEntry approach.

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-29 08:10:29 -03:00
Diego Rodrigues de Sa e Souza
71093eda77 fix(codex): keep parallel_tool_calls:false on translated Responses Lite path (#11707) (#11984)
enforceCodexResponsesLiteParallelToolCalls() forces parallel_tool_calls:false
at the top of CodexExecutor.execute(), but transformRequest() early-returns
the body before its RESPONSES_API_ALLOWLIST field filter only when
_nativeCodexPassthrough is set. Any request that reaches the codex
executor via the translated (non-native-passthrough) path never gets that
flag, so the allowlist filter silently deleted parallel_tool_calls right
before the fetch body was sent, reproducing the reported upstream
rejection ('X-OpenAI-Internal-Codex-Responses-Lite requires
parallel_tool_calls to be false') for every model.

Add parallel_tool_calls to RESPONSES_API_ALLOWLIST so the value survives
the translated path too. Update the sibling #2608 allowlist test that
previously asserted parallel_tool_calls gets stripped like other Chat
Completions-only fields -- it is a legitimate Responses API field that
must now survive.

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-29 08:10:26 -03:00
Diego Rodrigues de Sa e Souza
fb9cbe9566 fix(ci): pass --pass-on-unpruned-suppressions in run-eslint-json.mjs (#11600) (#11983)
Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-29 08:10:21 -03:00
Diego Rodrigues de Sa e Souza
674cc5feb1 fix(db): rate-limit Arena ELO fetch-failure warnings on repeated timeouts (#11500) (#11989)
* fix(db): rate-limit Arena ELO fetch-failure warnings on repeated timeouts (#11500)

* test(quality): split the #11500 fetch-failure-dedup tests into their own file

tests/unit/arena-elo-sync.test.ts crossed the 1000-line new-test-file cap
(file-size gate, PR mode). The two new tests don't need the file's DB
fixture (fetchArenaLeaderboards() never touches the DB), so they move to a
self-contained sibling file instead of growing the frozen suite.

---------

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-29 08:10:17 -03:00
Diego Rodrigues de Sa e Souza
d2ad71cf56 fix(kie): route flux/kontext to its dedicated endpoint, not the Market createTask flow (#11296) (#11985)
flux/kontext is catalogued with isMarket: true, so handleKieImageGeneration
routed it through KIE's unified Market createTask endpoint with
model: "flux/kontext". KIE does not expose Flux Kontext through the Market
catalog at all -- it lives under a dedicated API tree
(POST /api/v1/flux/kontext/generate, poll GET /api/v1/flux/kontext/record-info,
models flux-kontext-pro/flux-kontext-max) -- so the Market endpoint rejected it
with "model name not supported", matching the reporter's exact error text.

Special-case flux/kontext ahead of the isMarket branch so it hits the
dedicated endpoint/payload shape instead of being treated as a Market entry.
z-image/4.0-*/4.5-* remains intentionally untouched (still blocked on
reporter/live confirmation per the existing in-code comment).

Co-authored-by: Markus Hartung <mail@hartmark.se>
2026-08-29 08:10:13 -03:00
Tobias Andersen
c705147de2 docs(i18n): finish Freepik → Magnific rebrand in locale strings and README (#11772)
Finishes the Freepik → Magnific rebrand from #10594 across 40 locale files and 3 README feature-list bullets (README.md, docs/i18n/it, docs/i18n/tr) — legacy `freepik` alias intentionally left in code/tests/redirects for backward compatibility, and historical CHANGELOG entries left untouched as documented history.

The README bullet had base-drifted since the PR branched (release tip's "What's New" changelog snippet had already dropped two providers mentioned nowhere else in the codebase, unrelated to this PR's scope) — resolved by keeping the tip's current bullet shape and applying only the Freepik→Magnific rename on top, in both the combined-worktree validation and the pushed branch.

Validated: all 40 edited locale JSON files parse; re-verified after resync onto the updated tip (post #11762/#11774/#11781).
2026-08-29 05:30:13 -03:00
Tobias Andersen
d8879371ea fix(combo): lock GitHub models rejected as "not supported" for future requests (#11781)
Follow-up to #11762/#11774, same bug class in combo's own model-lockout wiring: GitHub rejects several models (gpt-5.4, gpt-5.3-codex, etc.) with a 400 that's permanently unavailable for this account's Copilot integration, but nothing recorded a cross-request lockout — combo's #5249 in-request advance guard is correct but doesn't persist, so the same doomed model gets retried from scratch on every new request, indefinitely.

Fix: on a model-scoped 400 (`isModelScoped400`), call `lockModelIfPerModelQuota(provider, connectionId, rawModel, "model_capacity", 1h)`. GitHub already has per-model-quota enabled, so only the rejected model locks — siblings keep working. `isModelLocked()` is already checked pre-dispatch, so no other wiring needed.

Validated: 3/3 new tests + fixed a pre-existing test-isolation gap in combo-model-scoped-400-advance.test.ts (shared model name across sub-tests without clearing lockout state). Thanks!
2026-08-29 05:21:50 -03:00
Tobias Andersen
91f9a01fda fix(resilience): stop hammering permanently-moved endpoints and billing-suspended accounts (#11774)
Follow-up to #11762, same bug class hitting freeaiapikey (410 permanently-moved endpoint) and fireworks (412 billing-suspension) — both fell through checkFallbackError's generic transient-cooldown branch and got retried every ~1 minute for a full day.

Fix: `ENDPOINT_PERMANENTLY_MOVED_PATTERNS`/`isEndpointPermanentlyMoved()` → 24h lockout; `ACCOUNT_SUSPENDED_BILLING_PATTERNS`/`isAccountSuspendedForBilling()` → treated as credits-exhausted (1h cooldown), independent of status code so it also catches Fireworks' 412.

#11762 landed first and touched the same file — rebased/re-merged onto the updated tip (additive, no logic changes) and re-validated: 13/13 tests pass. Thanks for tracing this with real production logs again!
2026-08-29 05:19:15 -03:00
Tobias Andersen
87b3bdf85e fix(resilience): lock permanently retired models instead of short backoff (Gemini ban prevention) (#11762)
Root-caused via a real Gemini-ban incident log: deprecated-model 404/410s (e.g. gemini-2.5-flash "no longer available to new users") fell through checkFallbackError's generic transient-cooldown branch, so combo/auto-routing kept re-selecting a permanently dead model every cooldown window forever — the hammering that got the account flagged as abusive.

Fix: `MODEL_PERMANENTLY_UNAVAILABLE_PATTERNS` + `isModelPermanentlyUnavailable()` classify these as a 24h lockout instead, surfaced via `quotaResetHintMs` so combo's per-request model-lockout honors it in full.

Validated: 6/6 new tests + 133/133 existing accountFallback/error-classification tests, no regressions. Thanks for tracing this end-to-end with real production logs!
2026-08-29 05:09:37 -03:00
Diego Rodrigues de Sa e Souza
3b752f9d4c chore(quality): type the 55 no-explicit-any sites frozen under #11924 (#11975)
Production (open-sse/utils/socksConnectorWithFamily.ts, 4 sites): every cast was
redundant — undici's buildConnector.BuildOptions already has `timeout?: number | null`,
socks' SocksClientOptions has `timeout?: number`, and Agent.Options' `connect` /
`connectTimeout` narrow to the connector's parameter types on their own. Behaviour
unchanged; check:open-sse-typecheck stays at the frozen 5.

Tests (51 sites): the socks-timeout mocks now carry the real types — the patched
SocksClient.createConnection is typed as the static it replaces, the fake
buildConnector returns buildConnector.connector, the proxy is a SocksProxy, the
dynamic import is typed as the module it loads; the e2e suite passes a SocksProxy and
Agent.Options and no longer casts undici's fetch init (its RequestInit already has
`dispatcher`); the isFree suites narrow getCustomModels()' JSON to a declared row
shape, feed deliberately-wrong values through `unknown`, and stop casting for
zod's safeParse, which takes unknown.

The six files' suppression entries are removed: 1238 → 1232 files, 5487 → 5432
suppressed. ESLint without the suppressions file reports 0 problems on all six;
with it, no stale entry is left. The five suites pass (4, 2, 5, 4, 4).
2026-08-29 03:06:50 -03:00
Diego Rodrigues de Sa e Souza
757cc3bb9d fix(release): let the Electron workflow start again — grant actions:read to the npm leg (release/v3.8.51 twin of #11973) (#11974)
Same three changes as #11973 on main, applied to this branch's newer copy of the
workflow so the v3.8.51 tag does not repeat v3.8.50's zero-asset release:
publish-npm grants actions:read (the called publish job requests it — a caller that
grants less is refused at startup and the release job dies with it), a publish_npm
dispatch input gates the npm leg, and web-build/build/release check out the tag
named by the dispatch. actionlint clean; the five workflow-pinning suites pass.
2026-08-29 02:42:13 -03:00
Diego Rodrigues de Sa e Souza
60bbf0f8f0 fix(ci): stop a stalled Codecov upload from cancelling the Coverage job and the main run (#11972)
The job has timeout-minutes: 20; the c8 merge across 8 shards takes ~10 min and the
Codecov upload (declared informational) then hung for the rest of the budget on two
consecutive main runs (33207760653, 33215115341) — GitHub cancels the step, the job
ends cancelled, and the run's conclusion turns cancelled although every blocking job
was green. The upload step now has its own 5-minute ceiling and continue-on-error;
the job budget is 30 min. check-workflows suite 32/32; zizmor ratchet unchanged.
2026-08-29 02:10:25 -03:00
Diego Rodrigues de Sa e Souza
3d4f3e4960 test(infra): retry recursive temp-dir removal instead of failing a shard on ENOTEMPTY (#11966) (#11968)
* test(infra): retry recursive temp-dir removal instead of failing a shard on ENOTEMPTY (#11966)

Two shards on release/v3.8.51 went red in one day with the same signature —
"ENOTEMPTY, Directory not empty: /tmp/omniroute-<test>-XXXXXX" — from
combo-same-provider-cascade (Unit Tests fast-path 4/4, on a PR that touches only
.github/) and auth-policy-embeddings-webfetch-7785 (the 20k-test TIA step). Both pass
alone and on re-run: the cleanup races something still writing into the directory
(SQLite WAL/-shm checkpoint, a worker, the backup) and under a loaded hosted runner
the window opens. 1154 test files do their own cleanup with
fs.rmSync(dir, { recursive: true, force: true }); 57 already asked for retries.

One-shot codemod (scripts/ad-hoc/codemod-rm-maxretries.mjs, kept for the record):
every rm / rmSync / rmdirSync option object with `recursive: true` and no
`maxRetries` gains `maxRetries: 5, retryDelay: 100` — Node itself then retries
ENOTEMPTY/EBUSY/EPERM for up to ~0.5 s before giving up. 2243 call sites in 1292
files under tests/, the shared tests/_setup/isolateDataDir.ts exit hook included.
Only the option object changes: no call site, assertion or import is touched.

Validation: prettier and ESLint (with the frozen suppressions) clean on all 1292
files; a random 20-file sample runs green (quota-redis-store hangs identically on
the untouched tree — it needs a Redis on localhost, an environment matter). The
four unit shards on this PR are the full run.

* fix(quality): let check-forgotten-sibling-tests read a 1,000-file diff

The gate shells out to `git diff` through execFileSync with Node's default 1 MB
maxBuffer; the 1,292-file codemod in this PR is the first diff large enough to
overflow it, and the gate died with `spawnSync git ENOBUFS` before comparing
anything. 64 MB is far above any real PR and costs nothing when unused.
2026-08-29 01:17:40 -03:00
Diego Rodrigues de Sa e Souza
751710616a fix(ci): run the build-bearing nightly jobs on the box's light pool (#11965) (#11967)
Four nightly jobs run a backend-only `next build` on ubuntu-latest (7 GB):
Schemathesis, promptfoo injection guard, garak probes and the axe a11y suite
(self-building webServer). On release/v3.8.51 three of them died with the hosted
VM shutdown signature and nobody saw it — nightlies have no audience — and the
fourth passes by a margin of minutes. They now target [self-hosted, omni-light]
(hosted fallback when USE_VPS_RUNNER is off), a new two-listener label on the .113
box for jobs that need ~6 GB, not the 14-16 GB of a full build; they run once a day
in the 04:00-06:00 UTC window, when the box is idle.

Fleet reshaped the same day and documented in docs/ops/RUNNER_BOX.md: 4 active
OmniRoute listeners (omniroute-113-5/-6 omni-build, omniroute-113/-2 omni-light),
omniroute-113-3/-4/-7/-8 disabled (systemctl enable --now brings one back), janitor
ceiling MAX_ACTIVE_RUNNERS=4. The remaining headroom limit is the VM's 31 GB of RAM
(2 heavy + 2 light ≈ 42 GB peak, inside the 16 GB swap); more RAM on the Proxmox VM
is the lever that turns the label ceilings into 3 heavy + 2 light.

check:workflows --ratchet unchanged (194/194); check-workflows and
backend-only-smoke-workflows suites pass; docs-sync PASS.
2026-08-28 23:26:52 -03:00
Diego Rodrigues de Sa e Souza
d7cdfcad43 fix(ci): take the two hosted-runner builds off the PR rail (#11946, option 3) (#11962)
The hosted 7 GB runner cannot build release/v3.8.51 in any profile: `Build App`
(build.yml, push on every branch, full `build:release`) died in 19 of the last 30
runs — the branch tip included — with "The runner has received a shutdown signal"
~8 min into `next build`, swapfile and all; the advisory quality.yml build failed on
8/8 recent fork PRs with the same recipe; and `DAST smoke (PR)`'s backend-only build
died ~7 min in before the server even started, hidden as a permanently red
continue-on-error check. Together they painted every PR into release/** red with
zero signal and, on build.yml, produced an artefact nothing downloads.

- build.yml: workflow_dispatch only. The bundle is validated where a build fits —
  ci.yml `Build` on the self-hosted omni-build pool after every merge to main, and
  nightly-release-green.yml on the same pool for release/**.
- dast-smoke.yml: pull_request into main only (plus workflow_dispatch to smoke a
  release branch by hand); main's tree still builds on the hosted runner in ~5.5 min.
- quality.yml: the fork-only rationale of `Build (advisory)` updated to say why
  own-origin PRs no longer get a hosted build either. Behaviour unchanged.

check:workflows --ratchet: 194 zizmor findings, baseline 194. check-workflows and
backend-only-smoke-workflows suites pass. Trade-off stated in the PR: own-origin PRs
into release/** lose a pre-merge build that was not succeeding anyway; the nightly
rail files a base-red issue within a day if a merge breaks the build.
2026-08-28 22:52:15 -03:00
Diego Rodrigues de Sa e Souza
034314262e docs(agents): sync-back landings are fast-forward, never squash (#11964)
Records the v3.8.50 → v3.8.51 precedent in the single source of truth: a
main → release/vX+1 sync PR lands by fast-forward push so main stays an ancestor
of the release branch (squash re-conflicts the next sync-back on every file main
touched — 551 conflicts this cycle before the two-step merge), plus the two
post-landing checks (ancestry assert; ratchet files carried main's freezes).
The full procedure lives in the generate-release Phase 5 skill.
2026-08-28 22:44:06 -03:00
Diego Rodrigues de Sa e Souza
3d125647c8 chore(ci): cap the unit shards at 30 min and stop restoring stale ESLint caches (#11963)
- quality.yml fast-unit: timeout-minutes: 30. A shard finishes in ~10 min; without
  a ceiling a hung test process holds the PR for GitHub's 6 h default. On
  2026-08-28 shard 1/4 sat 64 min without a line of output — twice at the same spot,
  a timing race that vanished on the third run — while the other three shards were
  long green. A fast red plus a re-run beats a silent multi-hour hold.
- quality.yml lint-guard + the earlier ESLint cache block: drop the
  `restore-keys: eslint-<os>-` fallback (#11600, P-II.1 of the v3.8.50 postmortem).
  The key already hashes the lint config, the suppressions file and the lockfile;
  the fallback restored a cache built under a DIFFERENT configuration and its stale
  per-file verdicts are how 215 pre-existing errors stayed invisible for a cycle.
  Exact key or a cold full lint — never a partial cache from another configuration.

check:workflows --ratchet unchanged (194/194); check-workflows suite 32/32.
2026-08-28 22:43:11 -03:00
Diego Rodrigues de Sa e Souza
d0f69e4c70 chore(quality): re-freeze the ESLint suppressions on release/v3.8.51 from a clean-room run (#11955)
* chore(quality): re-freeze the ESLint suppressions on release/v3.8.51 from a clean-room run

`No new ESLint warnings` failed on every PR against release/v3.8.51 with exit 2:
"There are suppressions left that do not occur anymore". Measured in a depth-1 clone
with `npm ci` from the branch's own lockfile and the job's exact command
(`npm run lint:json -- --max-warnings 0`): 56 errors — 55 `no-explicit-any` in six
files that landed while the base was red (#11843 isFree tests: 22; b7102140d5 socks
connect timeout: 33) plus one `no-unused-vars` — and stale entries for files that no
longer violate. The devbox figure previously quoted in #11924 (280, with 224
react-hooks/*) does not reproduce on the lockfile install and is withdrawn.

- config/quality/eslint-suppressions.json: `--prune-suppressions` (two stale file
  entries removed) and the 55 pre-existing `any` frozen at their exact counts — the
  file is a ratchet, counts only go down; the debt stays tracked in #11924.
- open-sse/services/adobeFireflyCatalog.ts: remove `GPT_SIZE_MAP`, a constant the
  f3d9279b44 split left behind with no reader (the real violation, fixed not frozen).

Verification in the clean room after both changes, same command as CI: exit 0,
0 errors, 0 warnings (1238 files / 5487 suppressions).

* chore(quality): tighten openapiCoverage.pct to the measured 39 (require-tighten)

With ESLint back to 0/0 on this PR, the job's next step (check-quality-ratchet
--require-tighten) started failing: openapiCoverage.pct improved from 38.4 to 39
(delta 0.6 > slack 0.5) and the baseline must be tightened in the same PR. 39 is
the value CI collect-metrics measured on run 33213844112 and a clean-room checkout
of 777d9d1629 reproduces it; the cycle's new routes landed documented in
docs/openapi.yaml. Only this metric moves; annotation follows the file's convention.
2026-08-28 22:07:40 -03:00
diegosouzapw
777d9d1629 test(translator): fix the relative imports of the relocated deferred-finish test
#11940 moved tests/unit/translator/openai-to-claude-trailing-usage.test.ts one level
up so a collector would run it, but kept the ../../../ import path from the old
directory, so the file failed to load and painted Unit Tests fast-path (3/4) red on
every PR since a94fe23e89. The path now matches its new location (5/5 pass).
2026-08-28 19:23:37 -03:00
diegosouzapw
9661611e31 Merge remote-tracking branch 'origin/main' into chore/sync-main-into-3851-20260828c 2026-08-28 19:03:24 -03:00
Diego Rodrigues de Sa e Souza
24c0643a94 test(check): escape the runs-on fixture with JSON.stringify, not a quote-only replace (#11942)
CodeQL js/incomplete-sanitization (alert #888 on #11929): the hand-rolled replace
only escaped double quotes, so a backslash in the fixture would have produced a
malformed YAML scalar. JSON.stringify covers every escape the double-quoted YAML
scalar needs. Test-only change (7/7 pass).
2026-08-28 19:01:54 -03:00
Diego Rodrigues de Sa e Souza
a94fe23e89 fix(release): drain the twelve reds every PR against release/v3.8.51 was born with (#11940)
* fix(release): drain the twelve reds every PR against release/v3.8.51 was born with

Measured on the cycle tip: fifteen unit files were red on every PR. Two came
from the v3.8.50 sync-back (fixed in #11929); the other thirteen predate it and
are the branch's own drift. This sweep clears all of them but the ESLint debt
(#11924), each with the smallest change that keeps the guard honest:

- .env.example + ENVIRONMENT.md: NEXT_PUBLIC_SW_BUILD_ID / OMNIROUTE_SW_BUILD_ID /
  SOURCE_VERSION (#11779 service-worker cache busting) documented — the env/docs
  contract gate was failing on every PR.
- stryker.conf.json: the six tests the mutation gate found covering mutated modules
  (four retirement runtime-block suites, combo connection-aware expansion, tunnel
  error sanitization) registered in tap.testFiles.
- dependency-allowlist: eslint-plugin-react-hooks 7.0.1 approved; its findings are
  tracked in #11924.
- i18n: the six combo.sort.* strings (d5dfcfff58) translated for vi (strict parity)
  and pt-BR.
- docs/providers/CHATGPT_WEB.md: the retirement test is migration-168, not 163.
- g4f gateways: authHint now says member key, which the discontinued-providers
  guard asserts.
- tests realigned to the catalog the branch actually ships: qwen-web (#11713) and
  chatgpt-web (#11720) are retired, so web-session-contract and
  token-health-check-webcookie use perplexity-web, grok-web and chatgpt-web-codex.
- db-core-init: the two minimal legacy fixtures gained the columns migrations 164-168
  UPDATE (error_code, last_error*, test_status) — they exist on every real legacy DB
  (base CREATE TABLE); the fixtures simply never declared them.
- no-js-extension guard: a .js specifier whose target is a genuine JavaScript file
  (open-sse/lib/deepseek-pow-hash.js, shared with a worker) is not the #10674
  defect; the test now skips targets that exist as .js.

All twelve files pass locally; docs-sync, docs-counts, env-doc-sync, the tap
drift gate and the fabricated-docs gates are green on the tree.

* test(release): move the deferred-finish translator test into a collected path

tests/unit/translator/ is not one of the unit collectors (package.json test:unit,
merge-train.sh, build-test-impact-map, check-test-discovery), so the suite that
dd35750e5f added there never ran — check:test-discovery flagged it as a new orphan
on every PR. Relocated next to its sibling openai-to-claude-trailing-usage-11817
under tests/unit/, where the root glob collects it (5/5 pass).

* fix(dashboard): type the four sort-method sites #11812 left red on the dashboard typecheck ratchet

d5dfcfff58 added the combo model sort and raised combos/page.tsx from 23 to 27
scoped TypeScript errors (TS2339 +1, TS2345 +2, TS2322 +1), which fails
check:dashboard-typecheck on every PR against release/v3.8.51:

- initialSortMethod: sanitizeComboRuntimeConfig() is untyped, so config.modelSort is
  unknown; narrow it before reading .method (normalizeSortMethod takes unknown anyway).
- handleAddModels: the batch path passes ComboBuilderDraftModelStep[] to the ComboStep[]
  sort helpers without the cast handleSortChange already uses; mirror it.
- ComboSortSelect expects a translate-with-fallback (k, f) => string, but received
  next-intl's Translator whose second argument is a values object. Pass the page's
  getI18nOrFallback adapter instead of the raw translator — that is also what makes
  the `has()` check and the fallback text actually work at runtime.

Baseline untouched (no widening). Scoped tsc: 0 new/regressed errors.
2026-08-28 18:18:33 -03:00
Diego Rodrigues de Sa e Souza
8dfdd95187 test(release): align five suites with the contracts #11933, #11919 and #11876 shipped on release/v3.8.51 (#11944)
Eleven PRs landed on release/v3.8.51 while the branch carried fifteen base reds, and
nine more red tests hid among them. None is a defect in the shipped code; each test
still encoded the contract that the merged PR deliberately replaced:

- openai-to-claude finish deferral (dd35750e5f, #11933): a finish chunk that carries no
  usage is now held until the end-of-stream flush that production performs
  (open-sse/utils/stream.ts flush -> translateResponse(..., null, state)). The drivers in
  stream-markdown-token-boundary, translator-tool-call-shim and
  gemini-malformed-function-call-finish-reason-2462 fed the finish chunk and asserted
  the terminal events immediately; they now mirror the flush. Assertions unchanged.
- authoritative live catalog (3d2832b836, #11919 fixes #11829): a synced catalog replaces
  the static registry, so model-lifecycle-integration no longer expects the static-only
  gpt-5.6-sol row to survive a sync. The #8627 contract the file guards (stale chat rows
  suppressed, typed media retained) is untouched.
- provider asset provenance (#11876): the unit shards check out with depth 1. The fixture
  pinned a historical commit as auditedCommit (absent on a shallow clone), the
  "binds auditedCommit" case relied on the repository root commit (the grafted HEAD on
  a shallow clone, which matches the physical snapshot), and the real-manifest case
  needs the audited commit fetched. The fixture now audits HEAD, the mismatch case
  builds a dangling empty-tree commit (no ref written), and the real-manifest case
  skips only on a shallow checkout that lacks the commit - the gate itself keeps
  running on both fetch-depth-0 rails, which the next test asserts.

All five files pass locally (30, 11, 38, 3 and 18 tests); lint with the frozen
suppressions is clean.
2026-08-28 18:10:56 -03:00
Diego Rodrigues de Sa e Souza
226538fa27 feat(ci): publish to npm through Trusted Publishing (OIDC) by default (#11931)
* feat(ci): publish to npm through Trusted Publishing (OIDC) by default

npm rejects provenance from self-hosted runners and is retiring tokens that
bypass 2FA; v3.8.49 answered with staged publishing (WS1.3) so a leaked token
could never publish alone — at the price of a manual `npm stage approve` per
release. Trusted Publishing gives the same guarantee with no token at all: the
github-hosted stage-npm job exchanges GitHub's id-token for a credential scoped
to that run, provenance included, and the flow is automatic again as it was up
to v3.8.48.

publish_mode gains `auto` (the default, also the path for the release event);
`staged` now runs only when asked for; `direct` stays as the emergency token
fallback. Until the owner registers the Trusted Publisher on npmjs.com
(diegosouzapw/OmniRoute, workflow npm-publish.yml) the automatic step fails
with ENEEDAUTH and either other mode can be dispatched — documented in
docs/ops/RELEASE_CHECKLIST.md.

* docs(release): date the checklist for the Trusted Publishing change and drop the env-var claim

check-deprecated-versions flags a touched doc whose header still says
2026-06-28 / v3.8.40; the fabricated-docs gate read the backticked NPM_TOKEN as
an environment variable the code never reads (it is a repository secret).
2026-08-28 18:02:51 -03:00
diegosouzapw
fb7445eaa3 test(check): escape the runs-on fixture with JSON.stringify, not a quote-only replace
CodeQL js/incomplete-sanitization (#888): the hand-rolled replace only escaped
double quotes, so a backslash in the fixture would have produced a malformed YAML
scalar. JSON.stringify covers every escape the double-quoted YAML scalar needs.
2026-08-28 17:29:04 -03:00
diegosouzapw
9968e1ce6e Merge remote-tracking branch 'origin/release/v3.8.51' into chore/sync-main-into-3851-20260828b 2026-08-28 17:27:31 -03:00
Diego Rodrigues de Sa e Souza
f907b5ea8e fix(ci): cap heavy builds at two runners with the omni-build label (#11932)
The .113 box (31 GB) holds one next-build (14–16 GB RSS) comfortably and two
at the edge; on 2026-08-28 the kernel killed main's build twice while PR
builds ran beside it. Labels are the runner-side cap: only omniroute-113-5
and omniroute-113-6 carry omni-build (added through the runners API, no
re-registration), and every job that runs a next build — ci.yml build,
npm-publish.yml publish, both nightly-release-green validations — now asks
for that label. A third heavy job queues on GitHub instead of racing for
memory. The six other runners keep omni-release and no longer take builds.
Pairs with the heavy-build-* concurrency lanes (#11901); documented in
docs/ops/RUNNER_BOX.md.
2026-08-28 17:19:45 -03:00
Diego Rodrigues de Sa e Souza
5b38ec717d fix(ci): keep the next-build artefact on disk, not on the runner's tmpfs (#11896)
* fix(ci): keep the next-build artefact on disk, not on the runner's tmpfs

On the .113 pool /tmp is a 12 GB tmpfs — it is RAM. The 1.3 GB next-build
artefact was parked there four times over: the Build job tar'd it to
/tmp/e2e-build.tar.gz (6 min), three E2E jobs downloaded it to /tmp/ and
extracted from there, and npm-publish.yml pulled it with gh run download into
/tmp/next-build. Measured on the v3.8.50 publish runs: that download step took
27 min (9th attempt) and 32 min (10th) — 42% of a 76-minute job — while the
very same bytes upload from disk in 2 min and the box pulls from GitHub at
7.3 MB/s (1.3 GB ≈ 3 min). Network was never the bottleneck; a tmpfs at 75%
under memory pressure was.

Every site now uses $RUNNER_TEMP / ${{ runner.temp }}: per-runner, on disk
(_work/_temp under the runner dir on the pool, /home/runner/work/_temp on
hosted images), and cleaned by the runner between jobs.

It also removes a latent race: e2e-build.tar.gz is a FIXED name under a /tmp
shared by every runner on the box, so two E2E shards on different runners could
overwrite each other's download mid-extraction. RUNNER_TEMP is per runner.

The supply-chain guard in tests/unit/npm-publish-artifact-provenance.test.ts
pins the candidate-run selection and the --name, not the directory; it stays
green. check:workflows --ratchet: zizmor unchanged at the baseline.

* fix(ci): download the next-build artefact to a workspace-relative dir (pwsh has no $RUNNER_TEMP)

The Electron Package Smoke matrix runs on windows-latest, whose default shell
is pwsh: $RUNNER_TEMP is empty there (pwsh spells it $env:RUNNER_TEMP), so the
first cut's tar -xzf "$RUNNER_TEMP/e2e-build.tar.gz" tried to open
'/e2e-build.tar.gz' and failed. A path relative to the workspace works in bash
and pwsh alike, and hosted workspaces are ephemeral. The producer (Build, Linux,
bash) and npm-publish keep $RUNNER_TEMP.
2026-08-28 17:08:58 -03:00
diegosouzapw
5ade9e0851 fix(sync): repair the two regressions the v3.8.50 sync-back left on release/v3.8.51
Fifteen unit files were red on this branch's PRs; running them on the pre-sync
tip (d5dfcfff58) and on the synced one showed thirteen already failed before
the sync — the cycle's own drift — and exactly two regressed:

- open-sse/services/tokenExtractionConfig.ts: git kept BOTH sides' identical
  volcengine-console config (23 entries instead of 22). The duplicate is gone.
- src/lib/usage/providerLimits.ts: the sync took release/v3.8.50's cooldown
  release helper, which is looser than this branch's #11277 contract (it frees
  an extra_usage block when the policy is off and a window with no reset
  evidence). tests/unit/provider-limits-recovery.test.ts pins the contract;
  the pre-sync call site is restored and the unused helper and its imports
  dropped. 20/20 again, siblings unchanged.
2026-08-28 16:41:42 -03:00
Diego Rodrigues de Sa e Souza
33763f06cc chore(changelog): add missing fragments for #11919/#11918/#11916 (#11938)
Adds the 3 missing changelog fragments.
2026-08-28 16:26:32 -03:00
Bob.Hou
6b259812a7 fix(sse): preserve store parameter semantics for openai-compatible responses (#11826) (#11916)
stripStore() now forces store=false for stateless OpenAI-compatible Responses-API targets unless the connection explicitly opts in via providerSpecificData.openaiStoreEnabled, instead of only handling the openai/agentrouter cases — a client-supplied store value previously passed through untouched to backends that don't actually persist responses server-side. Closes #11826. Thanks!
2026-08-28 16:25:19 -03:00
Bob.Hou
dc75a02ca7 fix(models): expose custom node models in canonical prefix mode (#11832) (#11918)
Custom provider-node models (synced, custom, and alias-backed) now appear under their configured prefix in the unified catalog when the operator's model-id prefix mode is canonical, instead of being dropped whenever alias-inclusion was otherwise disabled. Closes #11832. Thanks!
2026-08-28 16:25:05 -03:00
Bob.Hou
3d2832b836 fix(models): suppress static registry models when live catalog is synced (#11829) (#11919)
Suppresses stale static registry models (including effort-tier variants) for any provider whose active connection has an authoritative live synced catalog, not just providers using exclusive-synced-listing — closing a gap where a connection with providerUsesAuthoritativeLiveCatalog kept serving both the live-synced models and the stale static rows side by side. Closes #11829. 4/4 focused tests passing. Thanks!
2026-08-28 16:24:53 -03:00
Diego Rodrigues de Sa e Souza
cea1baa797 fix(ui): guard remaining ProviderIcon lookups against prototype collisions (#11920 port) (#11935)
Ports the 3 still-needed guards from #11920 that #11880 didn't cover. 90/90 + 4/4 focused tests passing.
2026-08-28 16:19:07 -03:00
Diego Rodrigues de Sa e Souza
dd35750e5f fix(sse): defer OpenAI-to-Claude finish emission until real usage arrives (#11915 follow-up on #11883) (#11933)
Merges #11883's already-merged usage-harvesting extraction with #11915's finish-deferral mechanism, verified to fix a real remaining bug: the client-visible message_delta carried stale/zero usage when finish_reason arrived before the trailing usage chunk. 86/86 tests passing across 16 translator regression files.
2026-08-28 16:11:38 -03:00
Diego Rodrigues de Sa e Souza
c661e1c811 port(playground): specific step warnings from #11882, keep #11862's string-step handling (#11930)
Ports the specific-warning improvement from #11882 (combo-ref/provider-wildcard steps get their own message instead of a generic count) onto #11862's already-merged crash fix. 4/4 focused tests passing.
2026-08-28 15:51:37 -03:00
diegosouzapw
529e4415c5 chore(release): sync main into release/v3.8.51 — the five post-release pipeline fixes
Brings e4683cd22d (#11867 Alibaba allowlist time bomb), 09de69edc7 (#11891
config expiry detector), e71be03398 (#11893 runner janitor), 9dc8eab70e
(#11895 provenance × self-hosted lint) and f564b64f7d (#11901 heavy-build
lanes). main is already an ancestor of this branch (v3.8.50 sync-back), so the
merge is exactly these five commits.

# Conflicts:
#	tests/unit/alibaba-free-tier-allowlist.test.ts
2026-08-28 15:49:53 -03:00
NoxzRCW
f08f35d6f0 fix(providers): pass xAI reasoning_effort xhigh through to grok-4.6+ (#11879)
normalizeXaiReasoningEffort() folded xhigh onto high before the request reached xAI, so anyone picking xhigh on grok-4.6 silently got high instead. xhigh is a real xAI tier (grok-4.6+); xAI already degrades it itself on unsupported models, so forwarding verbatim is safe everywhere. Closes #11816. Measured against live grok-4.6: reasoning_tokens 830 (high) vs 1052 (xhigh) — previously indistinguishable. Thanks!
2026-08-28 15:49:42 -03:00
NoxzRCW
d846692c30 fix(dashboard): guard provider icon lookups against prototype collisions (#11880)
getLobeProviderIcon() indexed two plain-object maps with no own-property check — a provider id that lowercases to an Object.prototype member (e.g. constructor) resolved through the prototype chain and threw on the follow-up .color/.mono lookup, surfacing as the misleading 'Failed to load providers, check your connection' error boundary card with a healthy server and clean logs. Thanks for the precise root-cause trace!
2026-08-28 15:49:32 -03:00
NoxzRCW
c5ebbb733c fix(skills): expand shorthand property types in injected tool schemas (#11881)
Every request through a strictly-validating provider (reproduced on opencode-go/glm-5.3-flash) failed with a 400: normalizeInputSchema() wrapped a skill's shorthand property map without expanding string values, so every injected omr_skill_* tool carried an invalid JSON Schema. Closes #11856. Thanks for the root-cause!
2026-08-28 15:49:22 -03:00
NoxzRCW
b8c7ee599d fix(translator): keep upstream usage from trailing empty-choices chunks (#11883)
openaiToClaudeResponse() returned early on !chunk.choices?.[0], dropping the trailing usage-only chunk many OpenAI-compatible upstreams send when stream_options.include_usage is set (confirmed on Fireworks kimi-k3) — state.usage stayed undefined and billing fell back to an uncached token estimate. 154/154 focused assertions across the fix + regression suite. Thanks for tracking down the billing impact!
2026-08-28 15:49:14 -03:00