mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-09-20 05:42:19 +03:00
815dedd3c9bbc2781ae4ef07f880869fbf7beb43
8802 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
815dedd3c9 |
fix(resilience): extend process crash guard to combo hedge cancels and upstream fetch failures (#13636) (#14064)
Closes a real process-killer: the direct-response start timeout could fire after the fetch promise had already settled, and aborting at that point delivered the abort reason to a promise nobody was awaiting — Node promotes that to an `unhandledRejection` → `uncaughtException` and the process dies (#12861). The timer is now a no-op once the attempt has settled, and the same guard is extended to combo hedge cancels and upstream fetch failures. Validated as a combined board first (this PR merged with the 11 siblings of the same batch on the release tip): eslint on every changed file with the suppressions file, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, i18n new-key coverage, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 275 passing / 0 failing focused node:test cases across the 28 test files the batch touches. Then re-validated alone on the fresh tip before this merge: conflicts re-resolved, file sizes rebaselined for this PR's own growth, eslint and this PR's focused tests re-run. Thanks @HouMinXi! Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com> |
||
|
|
9b92477b40 |
fix(providers): declare Agnes official thinking effort tiers (#13655) (#14063)
Declares the accepted thinking-effort tiers per Agnes chat model (2.0/2.5: none/low/medium/high/max; 3.0 adds minimal/xhigh), so the generic declared-tier clamp maps `xhigh`/`off` onto values the upstream accepts instead of forwarding them verbatim and collecting a 400. Validated as a combined board first (this PR merged with the 11 siblings of the same batch on the release tip): eslint on every changed file with the suppressions file, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, i18n new-key coverage, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 275 passing / 0 failing focused node:test cases across the 28 test files the batch touches. Then re-validated alone on the fresh tip before this merge: conflicts re-resolved, file sizes rebaselined for this PR's own growth, eslint and this PR's focused tests re-run. Thanks @HouMinXi! Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com> |
||
|
|
9956f13b35 |
fix(db): never TRUNCATE-checkpoint a live WAL (SIGBUS under traffic) (#14005)
* fix(db): never TRUNCATE-checkpoint a live WAL A live TRUNCATE checkpoint rewrites the shared wal-index while other processes hold it mapped; dereferencing the stale mapping SIGBUSes the process. Two production crashes six hours apart, coredump stack in better-sqlite3 native memcpy (issue #13973). Remove the periodic TRUNCATE scheduler. Runtime checkpoints are PASSIVE-only, which move pages without changing the wal-index geometry, while TRUNCATE stays on the shutdown path where reclaiming the file is safe. The 256MB size guard now warns instead of escalating to a live TRUNCATE, busy PASSIVE ticks feed the persisted busy telemetry that the TRUNCATE tick used to carry, and a positive OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS logs a one-time deprecation warning. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(db): RESTART the WAL when it exceeds the size guard The 256MB size guard only warned, so a live WAL could keep growing until the next restart. wal_checkpoint(RESTART) starts a new WAL file without rewriting the mapped wal-index, which is what SIGBUS'd the process when we used TRUNCATE under traffic. Related to #13973. Signed-off-by: Minxi Hou <houminxi@gmail.com> * docs: drop a fake TRUNCATE env name from the WAL guard row Backticks around TRUNCATE made the env/docs checker treat it as a variable. The VACUUM rows next to it were never part of this change and are not in the base docs. Signed-off-by: Minxi Hou <houminxi@gmail.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com> |
||
|
|
7cc454d931 |
fix(compression): repair the unit-test base-reds left by the 09-17 merge wave (#14082)
Every PR against release/v3.8.51 is born red on all four Unit Tests fast-path
shards. Reproduced on the pure tip (
|
||
|
|
24fc202d9f |
fix(ci): repair the API Route Typecheck base-red blocking every PR (#14079)
The gate fails on the pure release/v3.8.51 tip (3 files above the frozen baseline), so every PR against the release is born red on it. None of the three is a PR defect — they are drift from merged work: - src/lib/usage/glmResetCards.ts: entered the gate's scope when #12754 added the route that imports it. runWithProxyContext is an untyped async helper (Promise<any>), so runWithConnectionFetch<T> could not return T. Every call site passes an async callback and awaits it — declare that contract: fn: () => Promise<T> → Promise<T>. - src/sse/handlers/chat.ts: handleSingleModelChat had no return annotation, so runWithTransientBackendRetry<T extends ResponseLike> fell back to the constraint and the value could no longer feed withSessionHeader(Response). Annotated Promise<Response> (every return path builds a Response). - src/app/api/internal/codex-responses-ws/route.ts: the bridge helpers return either { error: Response } or a payload, but as unannotated object-literal unions TypeScript synthesised error?: undefined on the success member, and every "error" in x guard stopped narrowing (TS2339 x5 on the destructure). Explicit return types keep the discriminant real; the ApiKeyMetadata alias now points at the policy shape (the wider one both sources are assignable to), which clears the TS2740 self-mismatch; logger.warn → log.warn (the module logger factory has no .warn). check-api-typecheck.mjs: OK — 283 errors, baseline ratcheted DOWN from 294 (codex-responses-ws 7→4 TS2339, TS2740 1→0; combos/test and keys/[id] 1→0). typecheck:core clean; 43/43 unit tests in the touched areas green. |
||
|
|
1b484c57a0 |
chore(quality): rebaseline the last ceiling the 09-17 merge wave moved (#14016)
tests/unit/chatcore-translation-paths.test.ts 3447 -> 3449, from #13173 (Fable mid-conversation cache prefixes — the new assertions for that case). This is the only one left. The other two this PR originally carried (chatHelpers.ts and chatCore.ts) were absorbed by the rebaselines the merging PRs brought with them, so the branch was rebuilt on the current tip rather than shipping stale numbers. Measured clean: check:file-size and test-file-size both pass with no violations. |
||
|
|
94aa978c2a |
fix(ci): document OMNIROUTE_STRIP_SYSTEM_PREAMBLE — the env/docs base red blocking every PR (#14022)
* fix(ci): document OMNIROUTE_STRIP_SYSTEM_PREAMBLE (env/docs contract base-red) * fix(ci): allowlist COMBO_LOOP_SAFETY_TIMEOUT_MS as a doc-only source constant The env/docs contract gate had a SECOND violation on the release tip, added after this branch was cut: #13857's comboTimeoutMs narrative in ENVIRONMENT.md cites COMBO_LOOP_SAFETY_TIMEOUT_MS, which is a source constant (open-sse/services/combo/comboPredicates.ts:35 — `10 * 60 * 1000`), not an operator-facing env var. The doc regex captured the SHOUTY_NAME and reported it as documented-but-missing-from-.env.example. DOC_ONLY_ALLOWLIST already exists for exactly this class (see CLI_COMPAT_OMITTED_PROVIDER_IDS, LOCAL_ONLY_API_PREFIXES, VACUUM). Gate now reports all three directions in sync. |
||
|
|
b45e0a4b59 |
fix(i18n): review the 64 non-pt-BR dashboard catalogs for translation quality (#14078)
* fix(i18n): review-locale reviews every leaf of a catalog that did not exist at --since * fix(i18n): review the 64 non-pt-BR dashboard catalogs for translation quality Runs scripts/i18n/review-locale.mjs over every locale except pt-BR (done in #13885): 75,263 corrections applied against the English source, 1,677 of them reverted because the "correction" replaced a real translation with the plain English term (the real-translation ratio gate counts those as untranslated). zh-CN/zh-TW provider term normalised after the run. review-locale.mjs hardening found by the run: per-batch retries with backoff (a skipped batch is listed, not fatal), catalog checkpoint every 25 batches, and setDeep resolving leaf keys that contain a dot. |
||
|
|
1603c86e06 |
test: realign two stale assertions with product behavior (#13313) (#13315)
* test: expect the jina alias prefix in the custom-model catalog case
The custom-model assertions expected ids prefixed `jina-ai/`, but the catalog
prefixes model ids with the provider alias, which is `jina`. The synced-model
test directly above asserts `jina/` and passes, so the two cases contradicted
each other within the same file.
The expectation predates the alias: the test was written in v3.7.9 (2026-05-04)
and `alias: "jina"` was added in v3.8.36 (2026-06-25).
Aligns the two assertions with the sibling test and with the product. The file
now passes 44/44 (was 43 with 1 failure).
Confirmed the assertions still bite: renaming the alias to `jina-XX` fails
exactly these two cases.
* test: pin the pt-BR pack in the two language-pack fixtures
Both tests assert the Portuguese output-style string but configured
languageConfig with enabled:false and defaultLanguage:"en".
resolveOutputStyleLanguage returns "en" on its first line when enabled is not
true, so the English pack was injected and the assertion could never hold.
autoDetect:true would not have helped either: the user turns in these fixtures
are English, so the detector resolves back to "en". The pack under test has to
be pinned, hence autoDetect:false with an explicit defaultLanguage.
This restores coverage rather than just turning the suite green. Mutating the
pt-BR pack string in outputMode.ts now fails exactly these two tests; with the
old fixture the file reported 9 pass / 2 fail whether the pack was intact or
mutated, so it detected nothing. File is 11/11 (was 9 + 2 failures).
The third languageConfig fixture in this file belongs to an rtk test that makes
no language assertion and is left untouched.
* test: keep the canonical jina-ai prefix in the catalog case
Reverts
|
||
|
|
176d632a2d |
feat(api): add per-key allowAutoCombos to gate the built-in auto/* combos (#13670)
* feat(api): add per-key allowAutoCombos to gate the built-in auto/* combos
`auto/*` combos currently bypass per-key authorization entirely. They are
virtual — synthesised in the catalog, never stored as combo rows — so
`resolveRequestedComboName()` returns null for them and
`isComboAllowedForKey()` fails open:
const comboName = await resolveRequestedComboName(modelStr);
if (!comboName) return { allowed: true, comboName: null };
`validateModelAccess()` then sets `requestedComboName = modelStr` for any
`auto/` id and returns before `isModelAllowedForKey()` runs, so
`allowedModels` and `blockedModels` are skipped for those ids too.
The effect is that `allowedCombos` does not constrain `auto/*`: a key
scoped to a single cheap lane can still send `auto/best-coding` and reach
every model on the gateway. `blockedModels: ["auto/*"]` only unadvertises
the ids — it cannot deny them.
Add an explicit per-key flag instead of tightening the fail-open, which
would silently revoke `auto/*` from every key whose `allowedCombos` lacks
an entry for it. `allow_auto_combos` is NOT NULL DEFAULT 1 and the row
parser treats anything but an explicit falsy value as allowed, so every
existing key keeps working and opting out is deliberate.
When set to false:
- `validateModelAccess()` rejects `auto/*` for that key;
- the catalog skips the `auto/*` synthesis loop for it, reusing the
existing `hideAuto` break so the key is not offered ids it cannot use.
Settable via PATCH /api/keys/[id]. The create path and the dashboard
toggle are deliberately left for a follow-up: the API Manager control
needs UI strings across all message catalogs, which does not belong in
the same change as the policy fix.
* feat(dashboard): add the Auto Combos toggle to API key permissions
Exposes the `allowAutoCombos` flag in the API Manager permissions modal so
the per-key gate can be managed from the dashboard rather than only over
the API.
The control mirrors the prompt-compression toggle: a small dedicated
component, a `role="switch"` button, and labels from the `settings`
message namespace.
Defaults to ON. State reads `apiKey?.allowAutoCombos !== false` — using
`!== false` rather than `=== true` so a key that predates the column, or
one that has never been configured, renders as enabled and matches the
`NOT NULL DEFAULT 1` column.
The field is threaded through all three positional lists (the save
handler signature, the modal prop type and the onSave call) plus the
PATCH payload, so no later argument shifts position.
UI strings are added to en.json and to vi.json. Vietnamese is translated
rather than left as a sync placeholder because
tests/unit/i18n-vi-completeness.test.ts asserts key parity with English
and bans `__MISSING__` markers in that locale. The remaining locales fall
back to English at runtime; `i18n:check-ui-coverage` still passes well
clear of its threshold. They are deliberately not mass-synced here: a
full `i18n:sync-ui` run also replicates ~844 unrelated pre-existing gaps
across all 50 catalogs, which does not belong in this change.
* feat(api): advertise the combo description in /v1/models
A combo's description is stored on its record and returned by
GET /api/combos, but the catalog row never carried it, so no client could
show it.
Claude Code's gateway model discovery reads exactly `id`, `display_name`
and `description` from each entry in the /v1/models `data` array and
renders the description in the /model picker — an entry without one reads
"From gateway" instead. Other OpenAI-compatible clients surface it too.
Emit it only when the combo actually has one, so rows for combos without
a description are byte-identical to before. The value is typeof-narrowed
and trimmed because ComboRecord is Record<string, unknown>, and
`comboMetadata` still spreads last so context and capability metadata
keep precedence.
`display_name` is deliberately not sent: a combo's id is already its
human-chosen name, and the field is only consulted when it differs from
the id.
Ref: https://code.claude.com/docs/en/llm-gateway-protocol.md#model-discovery
* fix(api): list a key's allowed combos in /v1/models
`allowedCombos` gates combos; `modelAccessMode`, `allowedModels` and
`blockedModels` gate provider models. The catalog consulted only the
latter, so a key with `modelAccessMode: "restricted"` and an empty
`allowedModels` received an empty catalog — zero rows — while every combo
in its `allowedCombos` dispatched normally. The catalog contradicted the
key.
Observed on a live gateway: a key with 24 entries in `allowedCombos` and
`restricted` + `allowedModels: []` returned {"object":"list","data":[]},
yet `claude-orchestrate` answered 200 on that same key.
Gate combo rows on `allowedCombos` instead of hiding them. Listing a
combo the key can already dispatch grants no new access, so this is a
consistency fix rather than a relaxation, and it needs no opt-in: the
rule is simply that a key's catalog shows what that key can use.
auto/* rows are exempt. They fail open at dispatch — they resolve to no
stored combo — and their synthesis is already gated by allowAutoCombos,
so gating them here would make the catalog stricter than dispatch.
The decision lives in a new exported helper, isComboNameAllowedForKey(),
which wraps the existing matchesComboAccessRule. An absent list means no
combo restriction, matching validateComboAccess, which skips the check
when allowedCombos is not an array; an empty list allows nothing.
Also advertise `display_name` on combo rows from an operator-set
`displayName` field. Claude Code uses it as the picker entry's name when
it differs from the id, which lets a combo carry a discovery-compatible
id and still read cleanly. It is never derived from the combo name — an
unset field advertises nothing.
* fix(api): accept displayName on the combo schemas
The previous commit advertises `display_name` in /v1/models from a
combo's `displayName`, but neither createComboSchema nor
updateComboSchema declared the field, so Zod stripped it from every
request body and the value could never be set. The endpoint would have
answered 200 and written nothing — the feature was unreachable.
This is the same silent no-op that made `blockedModels` unsettable on
API keys: a field plumbed through the route and the store, missing only
its schema declaration.
Declare it on both schemas and count it in updateComboSchema's "no valid
fields" guard, so a body carrying only `displayName` is a valid update
rather than being rejected as empty. Nullable on update so a label can be
cleared.
* feat(api): add per-key catalogScope to scope what /v1/models advertises
A key had no way to say which kinds of thing its catalog should list. It
always advertised whatever the key's model and combo policies permitted,
mixed together. A client that builds its model picker from /v1/models —
Claude Code's gateway discovery, for one — then sees provider models
alongside the curated combos it was meant to offer.
Add a three-way per-key setting: "all" (default), "combos", "models".
This is a listing preference, not an access control: narrowing it never
changes what the key may dispatch, which the model policy and
allowedCombos continue to decide. That is why it is an explicit setting
rather than implied behaviour — unlike gating combo rows on
allowedCombos, which was a correctness fix and needed no opt-in.
Defaults to "all" everywhere: the column, the parser, the metadata and
the UI state, so every existing key is unchanged. The parser widens to
"all" on an unrecognised value rather than narrowing, so a bad value can
never silently hide rows an operator expects to see.
The dashboard control is a segmented radio group beside the Auto Combos
toggle. UI strings are added to en.json and vi.json; the remaining
locales fall back to English, and vi is translated rather than left as a
sync placeholder because tests/unit/i18n-vi-completeness.test.ts asserts
key parity and bans markers there.
* fix(api): invalidate the model catalog on key visibility changes
updateApiKeyPermissions already advances the unified /v1/models catalog
generation for the fields that change what a key may dispatch, but the two
fields this branch introduces -- allowAutoCombos and catalogScope -- were
missing from that predicate. Both change what the catalog advertises, so a
PATCH toggling either one left the request-shaped catalog cache serving the
previous listing until its TTL expired, and the dashboard's API-key screen
could show a catalog that disagreed with the key it had just written.
Add the two fields to the existing predicate -- no new cache machinery. The
call still runs only after a successful write, so a no-op or failed update
does not invalidate, and unrelated metadata edits (isActive, rate limits)
still leave the catalog cached.
Observed on a live deployment before the fix: PATCH catalogScope="combos"
returned 200 and the column read back "combos", yet GET /v1/models kept
returning the previous mixed rows until a process restart, after which the
same key correctly returned combo-only rows.
* docs(changelog): add fragment for per-key allowAutoCombos and catalogScope
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* chore(quality): rebaseline the two ceilings this PR's own growth moved
src/app/api/v1/models/catalog.ts 2075 -> 2117 and src/lib/db/apiKeys.ts
1625 -> 1659. Measured on the clean tip first: catalog.ts sits at 2074 (under
its 2075 ceiling) and apiKeys.ts at 1620 (under 1625), so none of this is
inherited — it is the feature itself. Gating the built-in auto/* combos per key
means the permission field has to be read, validated and carried all the way to
the catalog filter, and each of those is an explicit call site rather than
something extractable without hiding the gate.
Covered by the PR's 25 tests. The other violations in this tree (chatHelpers.ts,
chatCore.ts, chatcore-translation-paths.test.ts) are inherited base-reds and were
left untouched.
---------
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
||
|
|
b9dd80c8c2 |
fix(chat): preserve suffix reasoning intent across model attempts (#13720)
* chat/suffix-effort: keep reasoning intent tied to each model attempt Carry resolved suffix effort through dispatch without treating a derived value as explicit client input. Prepare reasoning defaults and dependent parameter constraints for each handler attempt so a replacement model does not inherit the original model's suffix. Keep explicit reasoning choices in context-aware request hashes to avoid sharing concurrent responses across different effort settings. Preserve the legacy hash interface and tenant namespace. Exercise retries, credential refresh, tool follow-ups, replacement models and overlapping requests with local HTTP and targeted regression tests. Signed-off-by: Minxi Hou <houminxi@gmail.com> * chat/upstream-body: separate normalization from async payload preparation Keep synchronous per-attempt normalization together so payload preparation stays within the function size and complexity limits without changing its ordering or explicit reasoning semantics. Condense redundant provider selection comments to retain the formatted file within its size ceiling. Signed-off-by: Minxi Hou <houminxi@gmail.com> * changelog: record the suffix-effort propagation fix Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(quality): rebaseline file-size cap for chatHelpers.ts growth The release tip independently grew src/sse/handlers/chatHelpers.ts from 1164 to 1213 lines while the frozen cap sat at 1214; this PR's own +3 lines (threading resolvedThinkingEffort through resolveModelOrError and executeChatWithBreaker) push the merged result to 1217, past the cap. Owner-approved exception for this file only, with the measured growth breakdown recorded in the baseline entry. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
7d69b02a29 |
fix(db): skip integrity scans during health polling (#13149)
* fix(db): skip integrity scans during health polling * test(db): update health error fixture for scan-free polling * docs(changelog): add fragment for health poll integrity skip * fix(db): keep the #13149 dashboard skip inside the #13717 managed health check Merge fallout only: runManagedDbHealthCheck moved behind the health coordinator on the release tip, so the per-call skipIntegrityCheck now travels through it. A waived integrity scan is part of the job identity, so it is never replayed from the 60s diagnosis cache to a caller that asked for the full scan. Co-authored-by: cryptiklemur <cryptiklemur@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: cryptiklemur <cryptiklemur@users.noreply.github.com> |
||
|
|
5f9e153971 |
fix(mcp): fall back when better-sqlite3 export is not callable (#13903)
* mcp/audit: fall back when better-sqlite3 export is not callable
Dashboard MCP status polls reopen a failed native sqlite load every 30s
because a minified TypeError ("a is not a function") was not treated as
a native load failure and a failed open was not cached. Classify that
shape, fall back to node:sqlite, cache the miss, and refuse to ship a
Docker image without better_sqlite3.node.
Signed-off-by: Minxi Hou <houminxi@gmail.com>
* mcp/audit: force native better-sqlite3 compile in Docker
better-sqlite3 13 ships a linux prebuild. Bare `node-gyp rebuild`
then only TOUCHes stamp files and never writes
build/Release/better_sqlite3.node, so the new test -f gate fails the
image build. Pass --force_build=1, matching the package's own
build-release script.
Signed-off-by: Minxi Hou <houminxi@gmail.com>
* db/core: keep native-load classification under the file-size cap
The audit fallback added two TypeError fingerprints in core.ts and
crossed the frozen 1788-line cap. Move the classifier into
sqliteLoadError.ts and re-export it so existing importers stay stable.
Signed-off-by: Minxi Hou <houminxi@gmail.com>
* build/bootstrap: keep the encrypted-credentials probe narrow
The native-load classifier was copied into scripts/build/bootstrap-env.mjs
alongside the runtime one, but the two files consume its verdict in opposite
directions. In src/lib/db/sqliteLoadError.ts a true verdict means "the driver
is unusable, cascade to node:sqlite", so treating a non-callable export as a
load failure is what we want. In the bootstrap the verdict feeds
hasEncryptedCredentials, where true means "no encrypted credentials found" and
clears the way to generate a fresh STORAGE_ENCRYPTION_KEY.
With the TypeError patterns in the bootstrap copy, a binding that loads but
exports something non-callable over a database full of enc:v1: rows reads as an
empty database, and the operator silently loses access to every stored
credential. Drop those two patterns from the bootstrap copy only, and note in
both files why the pair is deliberately not identical.
A corrupt binding still fails loudly there, now with the database path, the
underlying message, and a rebuild hint, so the narrower classifier does not
cost any diagnosability.
Signed-off-by: Minxi Hou <houminxi@gmail.com>
* fix(mcp): keep audit logging recoverable when the database is created later
getDb() cached a null for the "storage.sqlite does not exist yet" branch, and
closeAuditDb() returns before clearing a falsy cache — so an MCP server started
before the app created the database stayed without audit logging for the whole
process lifetime. Only a genuine driver-load failure is cached now; the
not-found branch retries, which is how it recovers when the file appears.
Covered by a new test that fails without the change.
Also replace the fabricated minified TypeError text ("a is not a function")
thrown by the loader with "better-sqlite3 export is not a function": the
operator sees a diagnosable message and isNativeSqliteLoadError() still
classifies it (it matches on "is not a function").
---------
Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
||
|
|
4be37d149b |
feat(dashboard): expose comboTimeoutMs next to Target timeout (#13857)
* feat(dashboard): expose comboTimeoutMs next to Target timeout The runtime already applied config.comboTimeoutMs as the whole-combo wall-clock budget (0 = 10-minute hang-stop). Schema treated it as an unknown passthrough key and the dashboard only painted Target timeout, so operators could not raise the 15-step failover ceiling from the UI. Declare comboTimeoutMs on comboRuntimeConfigSchema, mount both knobs in the combo editor Advanced panel and Combo defaults, and keep comboTimeoutMs longer than targetTimeoutMs so failover still has time. Signed-off-by: Minxi Hou <houminxi@gmail.com> * docs(changelog): attach #13857 to comboTimeoutMs fragment Signed-off-by: Minxi Hou <houminxi@gmail.com> * test(dashboard): name comboTimeoutMs store as milliseconds The input is seconds; the stored config field is milliseconds. The old title said "in seconds" while asserting 1_200_000. Signed-off-by: Minxi Hou <houminxi@gmail.com> * chore(i18n,changelog): translate #13857 keys into all locales and drop the CHANGELOG hunk --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
ec780ef2dc |
feat(compression): make Lite tool-result truncation length configurable (#13915)
* feat(compression): make Lite tool-result truncation length configurable Lite truncated tool results at a hardcoded 2000 characters. Coding-agent payloads (file reads, crash dumps) lost the middle of the content with no supported way to raise the cap. Honor lite.maxToolLength from settings, then OMNIROUTE_LITE_MAX_TOOL_LENGTH, then 2000. Existing installs keep the old length. Related to #13178. Signed-off-by: Minxi Hou <houminxi@gmail.com> #13178 stays open. * compression/lite: keep a stored cap when a step or toggle write is incomplete An out-of-range step maxToolLength was still a number, so it hid a valid global cap and fell through to env. A toggle-only settings PUT replaced the whole lite row and dropped the stored cap. Save treated an out-of-range number like a cleared field. Reject the bad Save, merge omitted caps, and use null to clear. Related to #13178. Signed-off-by: Minxi Hou <houminxi@gmail.com> * compression/lite: stop dashboard copy from hard-coding a 2000-char cap The page overlays schema descriptions from i18n. Updating only LITE_SCHEMA left operators seeing "over 2,000 characters" after the cap became configurable. Also assert the Save error string, not the Save button. Related to #13178. Signed-off-by: Minxi Hou <houminxi@gmail.com> * chore(changelog): move #13915 entry to a changelog.d fragment --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
e7872e57c3 |
fix(db): reclaim freed pages incrementally instead of a blocking VACUUM in the cleanup scheduler (#12821) (#12830)
* fix(db): reclaim freed pages incrementally instead of a blocking VACUUM in the cleanup scheduler (#12821) startCleanupScheduler() ran a synchronous whole-database VACUUM on the event loop whenever a cleanup pass deleted at least one row - 30 s after every start and every 6 h. With node:sqlite that blocks every route (/healthz included) for the duration: 7 min 55 s on a 540 MB storage.sqlite to reclaim six rows. It also bypassed vacuumScheduler, the app-level owner of full VACUUMs and the operator's scheduledVacuum / vacuumHour settings. cleanup.ts no longer issues a full VACUUM. After each pass reclaimFreedPages() branches on PRAGMA auto_vacuum: - INCREMENTAL: drain the freelist with PRAGMA incremental_vacuum(N) in ~1 MiB batches (N from page_size), pausing between batches for as long as the last one took (<=250 ms), PASSIVE checkpoint every 64 batches and a TRUNCATE checkpoint at the end so the main file shrinks in WAL mode; hard caps of 2048 batches / 30 s per pass, the remainder waits for the next pass. - FULL: nothing to do, SQLite reclaims on commit. - NONE: incremental_vacuum is a no-op, so record a request via the new vacuumScheduler.requestFullVacuum(); the rebuild runs in the configured window (or via the Storage page button). scheduledVacuum=never is honored. vacuumScheduler persists fullVacuumRequestedAt / fullVacuumRequestReason, clears them on the next successful runNow(), and hydrates from key_value before an early request so it cannot clobber a persisted lastRunAt. Loop robustness: db.exec() rather than pragma() (bun:sqlite's all() steps a zero-column pragma once), SQLITE_BUSY/LOCKED and a handle closed under the pass stop it quietly, other errors stop it with partial progress logged. Also drops the duplicate cleanupProxyLogs() call in the scheduled pass - runAutoCleanup() already covers proxy_logs. Tests: new tests/unit/db/cleanup-reclaim-freed-pages.test.ts (INCREMENTAL drain/pause/checkpoint, page_size-derived batch, caps, FULL no-op, NONE defers and leaves page_count untouched, runScheduledCleanupPass() path); vacuum-scheduler.test.ts covers requestFullVacuum persistence, restart survival and clearing; cleanup-column-fix.test.mjs now asserts incremental_vacuum and the absence of a full VACUUM statement. * chore(changelog): name the #12821 fragment after its PR (#12830) * fix(db): extract reclaimFreedPages into its own module and fix full-suite regressions Split the #12821 incremental-vacuum reclamation logic out of cleanup.ts into src/lib/db/reclaimFreedPages.ts (re-exported for callers/tests) so cleanup.ts stays under the file-size cap after the #13011 reconciliation merge grew it past the 1200-line threshold. Also fixes two full-suite failures surfaced by running the cleanup/vacuumScheduler/db-health suite post-merge (not just this PR's own 3 test files, per the plan-file's mandatory item): - tests/unit/cleanup-column-fix.test.mjs scanned cleanup.ts's raw source for the PRAGMA incremental_vacuum invariant, which now lives in the extracted module — updated to scan both files. - tests/unit/db/cleanup-reclaim-freed-pages.test.ts asserted the freelist count is byte-for-byte unchanged when auto_vacuum=NONE. The tip's runAutoCleanup() now also runs cleanupCompressionRunTelemetry(), which lazily creates its table on first use (ensureCompressionRunTelemetryTable) — a legitimate one-time page cost from a freshly migrated DB, unrelated to reclaimFreedPages()'s own behavior. Loosened the assertion to a small tolerance while keeping the page_count assertion that actually guards against a full rebuild. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * chore(db): drop the reclaimable-bytes VACUUM gate test superseded by incremental reclaim tests/unit/vacuum-reclaimable-threshold.test.ts pinned cleanup.ts's vacuumAfterCleanup()/getReclaimableBytes()/getVacuumMinReclaimableBytes() (#13079). This branch removes the inline post-cleanup full VACUUM entirely in favour of reclaimFreedPages() (#12821), which reads the same freelist_count / page_size signal and defers a full VACUUM to the vacuum scheduler when auto_vacuum=NONE. With those three exports gone the file cannot compile, and the behaviour it guarded no longer exists. --------- Co-authored-by: insoln <is@careerum.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
241e63bfea |
feat(usage): redeem GLM Coding Plan Reset Cards from Provider Limits (#12754)
* feat(usage): redeem GLM Coding Plan Reset Cards from Provider Limits z.ai sells Reset Cards that clear an exhausted GLM coding-plan window (5-hour or weekly) ahead of its natural rollover, but OmniRoute only ever read the passive nextResetTime, so redeeming one meant leaving the dashboard. Add the wire layer for z.ai's two reset endpoints (/api/biz/customer-package-reset/list and /use), which authenticate with the same Bearer API key as /api/monitor/usage/quota/limit and report failures inside an HTTP-200 envelope, so callers must inspect success/code rather than the status line. The banked count rides along with the quota poll - only for keys that actually report a resettable window, and strictly best-effort so a card-less account or a transient failure still renders its quotas. The existing reset-credit card, picker and confirmation flow, until now gated to Codex, now also drive glm/glm-cn/glmt/zai through the new /api/usage/glm-reset-card route, reusing z.ai's requestId as the idempotency key so a retry cannot burn two cards. * test(usage): cover GLM reset-card edge cases * test(dashboard): require GLM reset-card copy * fix(usage): harden GLM reset-card redemption * fix(usage): treat missing GLM key as empty * fix(usage): fence GLM reset-card operations and coalesce lease-window duplicates - Acquire a synthetic 60s exclusive-connection lease around each list/use wire operation; release in finally so a competing lease can acquire immediately after success or failure. - Coalesce same-key duplicates that arrive after lease acquisition by checking the in-flight attempt before loading the connection. - Run the post-commit quota refresh outside the lease (redemption is already committed; the refresh is auxiliary and failure-tolerant). - Do not discard a retained ambiguous attempt on a lease-conflict 409. - Harden transport error mapping: static messages for proxy transport failures, keep explicit direct routing for unproxied connections through list, use, and refresh. * fix(i18n): sync GLM reset-card keys to pt-BR and vi locales --------- Co-authored-by: insoln <is@careerum.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
4d9c4d3d8f |
fix(codex): fail the stream when the websocket closes before a terminal event (#12737)
* fix(codex): fail the stream when the websocket closes before a terminal event * docs(changelog): add fragment for Codex websocket premature-close fix * chore(quality): rebaseline file-size baselines for #12737 test and executor growth * fix(codex): log websocket failures and harden premature-close tests Review follow-up: failController now logs the failure (code + message) via nextInput.log, and onclose surfaces the WS close code/reason in the log line (the public payload stays sanitized through the allowlist). Adds regression tests for the onerror-before-onclose sequence and a close with zero prior events. * chore(quality): re-measure the #12737 codex.ts file-size rebaseline after the release merge The PR's own annotation was written against base 1505 and the frozen cap was already at 1528 on release/v3.8.51 (which grew the file independently). With this PR's +13 lines the merged file measures 1530, so the frozen entry moves 1528->1530 (+2 of genuine PR growth; the remaining 11 lines fit the headroom the tip already had). No other frozen entry touched. --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
ac0a63117e |
fix(providers): read the vLLM context window from max_model_len (#12897)
* fix(providers): read the vLLM context window from max_model_len normalizeDiscoveredModels resolved the window from inputTokenLimit, context_length, contextLength and top_provider.context_length. vLLM reports it as max_model_len and nothing else, so a synced vLLM model carried no inputTokenLimit and the resolver fell back to the 128K default - half the window on a 250K deployment. The native vllm provider and the OpenAI/Anthropic-compatible custom providers all pass raw records through this function, so one chain entry covers the three connection shapes. Closes #12858 * docs(changelog): fragment for #12897 --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
9688032451 |
test(dashboard): reactivate request logger coverage (#13843)
* test(dashboard): reactivate request logger coverage * test(dashboard): assert the request-logger modal by its own label The drift this PR unblocks is real — the next-intl mock lost its `useLocale` export in #7935, so the suite died on `No "useLocale" export is defined`. Restoring it is the actual fix. The dialog assertions were also loosened to `[role="dialog"]`, and that part was not needed: 20 components under src/ render that role, so the selector stops proving this particular modal is the one on screen. Measured — keeping only the `useLocale` fix and restoring a precise selector still passes 6/6. Now asserting `[aria-label="ariaLabel"]`, which is what RequestLoggerDetail.tsx:469 renders (the mock returns the key rather than the translation). --------- Co-authored-by: Paco Cartones <pacocartones@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
1ea87603c0 |
fix(qoder): unwrap split SSE error envelopes (#13838)
Co-authored-by: Paco Cartones <pacocartones@users.noreply.github.com> |
||
|
|
28ce4cacb2 |
fix(streaming): track progress across chunk boundaries (#13839)
Co-authored-by: Paco Cartones <pacocartones@users.noreply.github.com> |
||
|
|
1deb77a00d |
test(dashboard): reactivate discovery page coverage (#13842)
Co-authored-by: Paco Cartones <pacocartones@users.noreply.github.com> |
||
|
|
6f2eafd138 |
test(dashboard): reactivate API endpoints coverage (#13841)
Co-authored-by: Paco Cartones <pacocartones@users.noreply.github.com> |
||
|
|
4be690736c |
test(dashboard): reactivate webhook wizard coverage (#13844)
Co-authored-by: Paco Cartones <pacocartones@users.noreply.github.com> |
||
|
|
44d1760c3a |
fix(nlpcloud): restore chatbot endpoint coverage (#13845)
Co-authored-by: Paco Cartones <pacocartones@users.noreply.github.com> |
||
|
|
c40a5f0432 |
fix(build): externalize @modelcontextprotocol/sdk to heal MCP initialize 500 on standalone builds (#13859)
* build/mcp: externalize @modelcontextprotocol/sdk in standalone server bundle The SDK client graph contains a module-level class-extends-Client cycle against the top-level-await Client module. Webpack's TLA runtime evaluates that circular subgraph out of order when it is inlined into route chunks, throwing "Cannot access 'l' before initialization" during module evaluation. Every request to /api/mcp/stream then answers HTTP 500 on initialize and the failed module is evicted and re-evaluated per request, which floods the logs with the same ReferenceError. Node's native ESM loader resolves the same circular graph through live bindings, so keep the SDK external to the server bundle like the other packages that break only when bundled. Signed-off-by: Minxi Hou <houminxi@gmail.com> * changelog: record @modelcontextprotocol/sdk externalize fix (#13859) Signed-off-by: Minxi Hou <houminxi@gmail.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> |
||
|
|
2387d051c1 |
fix(sse): strip Codex temperature on native Responses passthrough (#12585)
* fix(sse): strip Codex temperature on native Responses passthrough Codex /responses rejects sampling params with FastAPI 400 Unsupported parameter: temperature. Native passthrough returned before the Responses allowlist, so client temperature reached upstream on combo traffic. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(sse): extract Codex passthrough param strip under file-size cap Keep temperature/top_p (and #3317 client-only fields) stripped before native Codex /responses passthrough returns. Move the call to open-sse/executors/codex/stripPassthroughRejectedParams.ts so executors/codex.ts stays under its frozen 1505-line cap. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
fc6b240328 |
fix(oauth): classify an embedded invalid_grant in a refresh error body (#13466)
* fix(oauth): classify an embedded invalid_grant in a refresh error body
Cline answers a dead refresh_token with
400 {"data":"","error":"failed to refresh token: invalid_grant","success":false}
The code is the tail of a sentence — neither a bare code nor an
"error":"<code>" field pair — so extractOAuthErrorCode returned null.
A null classification means refreshClineToken emits no unrecoverable
sentinel, so a permanently consumed refresh_token is handled as a
TRANSIENT failure. tokenHealthCheck therefore never reaches its
unrecoverable branch, and never runs the credentialsChangedSinceSweep
race guard, the "please re-authenticate this account" message, or the
dead-token clear for rotating providers. The connection instead stays
active with errorCode "refresh_failed", retries the same consumed token
3x per sweep behind an exponential backoff, and 401s every request
routed to it indefinitely with no actionable operator signal.
Scan for a known unrecoverable code embedded in the error value as a
last resort, after the exact-match and nested-JSON paths, delimited on
both sides so server_error, xinvalid_grant, my_invalid_grant_flag and a
502 HTML page all still classify as null.
Also add cline to ROTATION_LOCK_GROUP: refreshClineToken reads a new
refreshToken out of every response body and a measured refresh rotated a
connection's stored token, so sibling connections must not refresh
concurrently. cline was already listed in tokenHealthCheck's
ROTATING_REFRESH_PROVIDERS but missing from the serializer.
* docs(changelog): add fragment for cline refresh token error classification
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
||
|
|
ec60915d91 |
fix(api): accept blockedModels in the key permissions schema (#13666)
* fix(api): accept blockedModels in the key permissions schema `PATCH /api/keys/[id]` already destructures `blockedModels`, forwards it into the update payload, and `updateApiKeyPermissions()` writes it to the `blocked_models` column. Only the first link was missing: `updateKeyPermissionsSchema` never declared the field, so Zod stripped it from the parsed body and the destructured value was always `undefined`. The request answered 200 and wrote nothing. The API Manager permissions modal sends `blockedModels` on every save (ApiManagerPageClient.tsx), so the Claude-Code family-blocking control silently did nothing and an existing deny-list could not be cleared. `blockedModels` is the deny-list half of the model policy — read by `isModelAllowedForKey()` before the allow-list and winning over it — so that half was only reachable by editing the database by hand. Declare the field mirroring `allowedModels` (trimmed, non-empty, max 1000) and count it in the "No valid fields to update" guard so a body carrying only `blockedModels` is a valid update. Left out of `createKeySchema` deliberately: the create route does not read `blockedModels`, so declaring it there would be dead weight. * docs(changelog): add fragment for blockedModels key schema fix Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |
||
|
|
c97f61b2ac |
fix(sse): prevent Anthropic 400s for Claude-native handoffs (#12668)
* fix(sse): prevent Anthropic 400s for Claude-native handoffs * docs(changelog): add fragment for Claude-native handoff 400 fix * refactor(sse): satisfy file-size and complexity ratchets Keep the Claude wire-body guard while staying under the frozen per-file line baselines and the complexity ratchets measured against release/v3.8.51. Extract the final constraint coordinator, split system-message normalization into focused helpers, and isolate handoff response parsing. Reflow the universal-handoff explanation to absorb the added source-format argument without growing the frozen file. |
||
|
|
dd70dbdaa0 |
fix(sse): Anthropic OAuth 403 "Request not allowed" is a per-request refusal — cooldown with backoff instead of an instant ban (#12859) (#12864)
* fix(sse): Anthropic OAuth 403 "Request not allowed" is a per-request refusal, not a ban
A single upstream 403 on the `claude` OAuth connection was classified
FORBIDDEN and written as the terminal `banned` connection state
(chatCore -> writeTerminalStatus). From then on every request to that
provider was short-circuited with "All 1 connection(s) banned by
upstream - please reconnect in the dashboard" without touching Anthropic,
until an operator reconnected.
Anthropic's OAuth surface answers a small fraction of otherwise-valid
requests with 403 {"type":"permission_error","message":"Request not
allowed"}. On the reporting install the same token returned 200 forty
seconds before the 403 and again right after the connection was
re-enabled; a revoked or expired token is a 401 authentication_error, not
this. It is a refusal of one request, not of the credential.
Classify it as the new non-terminal PROVIDER_ERROR_TYPES.REQUEST_REJECTED
(scoped to provider `claude` and the "Request not allowed" body) and list
that type in authTerminalStatus.isNonTerminalProviderError, mirroring the
Cloudflare FINGERPRINT_REJECTION precedent. The combo layer still falls
through to the next target for the failing request; the connection stays
active for the next one. Any other claude 403 keeps its previous
classification.
Tests: error-classifier.test.ts covers the Anthropic body, the
gateway-flattened "[403]: Request not allowed" message, the same body from
a non-Anthropic provider (still FORBIDDEN), other claude 403s (unchanged),
and the helpers; anthropic-request-not-allowed-not-a-ban.test.ts pins
resolveTerminalConnectionStatus() -> null for the new type even with a
`permanent` fallback verdict, and `banned` for a generic claude 403.
* fix(sse): cooldown with backoff and streak escalation for REQUEST_REJECTED (#12859)
Not "ignore the 403" either: if Anthropic ever made "Request not allowed"
systematic, re-sending every request into it would be the wrong thing to
do to an OAuth account. chatCore now handles REQUEST_REJECTED explicitly:
- exclude the connection via setConnectionRateLimitUntil for a growing
cooldown (5 -> 15 -> 45 min) so a sporadic refusal costs minutes, not a
reconnect, and a systematic one cannot become a stream of 403s;
- escalate to the terminal `banned` state only for 3 refusals within a
60-minute window (services/requestRejectedStreak.ts, in-memory per
connection; a restart forgets the streak, erring towards more cooldowns
rather than an operator-undone ban), with a last_error that says so;
- probe-origin failures record but never cool down or ban (#9817).
The existing "request not allowed" text rule (5 s) is unaffected:
markAccountUnavailable skips a connection that already has a future
rateLimitedUntil, so the minute-scale cooldown written here wins.
Tests: request-rejected-streak.test.ts pins the window/threshold/backoff
arithmetic; anthropic-request-not-allowed-cooldown-escalation.test.ts drives
the real chat route against a mocked 403 upstream on a `claude` OAuth
connection: 300 s cooldown, then 900 s, then banned on the third refusal;
a different claude 403 body still bans on the first response.
* chore(changelog): name the #12859 fragment after its PR (#12864)
* refactor(sse): move the REQUEST_REJECTED branch into a chatCore leaf; register its tests for mutation coverage
chatCore.ts is frozen at 5984 lines by the file-size ratchet; the branch
body now lives in open-sse/handlers/chatCore/requestRejectedFailure.ts
(chatCore: 5974 -> 5983). stryker.conf.json tap.testFiles gains the two new
DB-backed tests so their mutant kills count (check:mutation-test-coverage).
* fix(sse): count refusal episodes, reset on success, keep the dashboard honest (#12859 review)
Review findings on the first cut of the REQUEST_REJECTED handling:
- A burst of in-flight requests that all got the 403 within seconds
produced streak 1, 2, 3 and a ban from one upstream event. The streak
now counts cooldown *episodes*: a refusal that lands while the
connection is already excluded is the same event and is not counted.
- Nothing reset the streak on a healthy response, so sporadic refusals
on a busy install could still accumulate to a ban. chatHelpers'
onRequestSuccess now clears it (only a real success does - the recovery
tick's clearAccountError is an elapsed cooldown, not a success).
Clearing the cooldown by hand in the dashboard clears it too.
- The third rung of the ladder was unreachable (the third refusal
escalates): the ladder is now 5 -> 15 min, sourced from COOLDOWN_MS next
to the existing 5 s "request not allowed" rule, with a note on why that
rule is superseded for claude. The 60-min window becomes a 24 h
staleness bound - "consecutive" is defined by successes, not by time.
- Probe-origin refusals no longer touch the streak (#9817).
- The cooldown is written like every other connection-level cooldown:
ISO rateLimitedUntil + testStatus "unavailable" (+ lastErrorAt), so the
dashboard shows the countdown and the recovery tick restores "active".
- One refusal is re-seeded from the persisted row after a restart so a
crash loop cannot reset the count on every boot.
Docs: RESILIENCE_GUIDE terminal states + CODEBASE_DOCUMENTATION resilience
row mention the streak module. Tests cover the burst, the success reset,
the seed, and the ISO/unavailable shape end-to-end through the chat route.
* chore(sse): drop unrelated Prettier churn in auth.ts / providers route
* style(api): keep providers route Prettier-clean
* refactor(sse): share the "exclude connection for a cooldown" leaf between GEO_BLOCKED, GCP_PROJECT_REQUIRED and the new branch
The release tip moved chatCore.ts to its frozen 5984 lines, so the
REQUEST_REJECTED branch cannot add a single net line. The GEO_BLOCKED and
GCP_PROJECT_REQUIRED branches were the same eight statements with different
constants and log wording; both now call
open-sse/handlers/chatCore/connectionCooldown.ts::excludeConnectionForCooldown
(behaviour, probe guard and log lines preserved verbatim). chatCore.ts ends
9 lines below the base it branched from.
* chore(chatCore): tighten the cooldown comments to keep the file under its size ceiling
After merging release/v3.8.51, chatCore.ts sat at 6150 lines against a
frozen ceiling of 6146. Condense the explanatory comments this PR added
to the GEO_BLOCKED and GCP_PROJECT_REQUIRED branches; no code change.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: insoln <is@careerum.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
||
|
|
47159ed56b |
fix(combo): answer 503 + Retry-After, not 404, when a weighted pool is only cooling down (#12956)
* fix(combo): answer 503 + Retry-After, not 404, when a weighted pool is only cooling down The weighted strategy filters targets before dispatch (open circuit breaker, provider cooldown, model lockout, availability probe) and drops them silently. When that emptied the pool the host returned the 404 "Combo has no executable targets" with the "switch combo / reconnect the missing providers" recovery hint — for a pool that was configured, connected and merely cooling down. Claude Code renders a 404 from /v1/messages as "this model may not exist". - targetResolution.ts: the eligibility predicate now reports which gate excluded a target and, for the resilience gates, the remaining time; the exclusions of fully-excluded steps travel out of resolveWeightedSelection. When the weighted pool ends empty and at least one exclusion is a resilience timer, the pipeline returns an early 503 and logs the reasons at warn level. - pinRecovery.ts: buildAllTargetsCoolingDownResponse() — 503 `all_targets_cooling_down`, Retry-After = earliest exclusion to lapse, every excluded target in diagnostics.excluded, `wait` recovery hint with retry_after_seconds; formatPreDispatchExclusions() for the log line. - error.ts: `all_targets_cooling_down` joins the public error identifiers. - A pool emptied only by the availability probe keeps the 404. - docs: RESILIENCE_GUIDE debugging entry; changelog fragment. * chore(changelog): name the fragment after PR #12956 and link issue #12954 * refactor(combo): keep weighted exhaustion below complexity ratchet --------- Co-authored-by: insoln <is@careerum.com> |
||
|
|
bc7f68fb91 |
fix(resilience): lock the exact model, not the quota family, on 5xx model-lockout failures (#12957)
* fix(resilience): lock the exact model, not the quota family, on 5xx model-lockout failures A 5xx model-lockout failure — a transport error (terminated, EHOSTUNREACH, connect timeout), an upstream server error, or OmniRoute's own synthesized 502 from quality validation — is evidence about one model endpoint at that moment, not about the account's quota family. recordModelLockoutFailure() wrote it under the quota-family key regardless, so for codex (whose family key is the whole `codex` scope, i.e. every gpt-5* model) one empty stream on gpt-5.6-luna removed gpt-5.6-sol and gpt-5.6-terra from routing too, for 2–30 min with exponential escalation, while the quota was untouched. - exactModelLock.ts: resolveLockoutScope(status, explicit) — 429/403/402 (and 404, already narrowed by getModelLockKey) keep the family key; any other status uses the exact provider/connection/model key. An explicit `scope` option still wins. - recordModelLockoutFailure() resolves the scope once for key + lock fn. - decayModelFailureCount() now walks every key shape (family, not_found, exact) so success-decay reaches exact-scope locks; null model stays a no-op. - getAllModelLockouts() parses the `exact:` marker out of the key so the Model Cooldowns card lists the bare model and can clear it by that name. - docs: RESILIENCE_GUIDE §3 key-scope-by-status; changelog fragment. * chore(changelog): name the fragment after PR #12957 and link issue #12955 --------- Co-authored-by: insoln <is@careerum.com> |
||
|
|
21d756d7f0 |
fix(combos): accept isHidden in updateComboSchema (#12898)
* fix(combos): accept isHidden in updateComboSchema A combo's visibility is stored on the record and honoured by the builder option list and the dashboard grid, but updateComboSchema never listed isHidden. The PUT handler spreads the validated body, so zod stripped the field: a visibility-only update was rejected as "No valid fields to update", and a mixed update succeeded while dropping the visibility change. Closes #12836 * docs(changelog): fragment for #12898 |
||
|
|
8074e3d596 |
fix(resilience): honor declared effort vocabulary in reasoning rule gate (#12686)
* fix(resilience): honor declared effort vocabulary in reasoning rule gate The reasoning-routing rule capabilityFor() hardcoded a gpt-5.6-(sol|terra|luna) whitelist for forced max/ultra, rejecting every other thinking-capable model even when the model's resolved capabilities declare the requested tier (synced supportedThinkingEfforts or an operator Model Overrides reasoning_efforts override). This 400'd direct calls with "Reasoning effort 'max' is not supported by the configured target" for models like Merge Gateway zai/glm-5.3-flash, which natively accepts low|high|max. The gate now treats a declared vocabulary containing the requested tier as authoritative, mirroring the dispatch-time sanitizer (open-sse/executors/base/reasoningEffort.ts) which already forwards declared tiers verbatim. Undeclared models keep the legacy gpt-5.6 regex verdicts and the unknown passthrough. * fix(resilience): gate forced max against the static registry the sanitizer clamps with Adversarial review finding: the gate read supportedThinkingEfforts from getResolvedModelCapabilities, which prefers the DB override over the registry. For a registered model with a narrow registry vocabulary and a widening operator override, the gate passed forced max but the dispatch-time sanitizer (executors/base/reasoningEffort.ts) clamps against the STATIC registry and would silently downgrade max to the registry ceiling — converting a loud 400 into a silent wrong-effort request. Order of precedence in the gate now: 1. static registry vocabulary (authoritative — matches sanitizer clamping) 2. declared/overridden vocabulary for unregistered providers (#8057 path) 3. legacy gpt-5.6 regex, then unknown/unsupported verdicts Also pins the test fixture to a synthetic model id so a future models.dev sync row cannot flip the unknown-precondition assertion. * fix(resilience): gate registry lookup mirrors the dispatch sanitizer exactly Review findings on the forced max/ultra gate: - resolve the registry through getProviderModels (id->alias namespace) and match entry aliases, mirroring reasoningEffort.ts — a raw provider id or alias-spelled model no longer skips the registry branch and diverges from dispatch clamping - treat an empty declared vocabulary as no declaration (falls through), matching the sanitizer's declaredRanked.length>0 guard — before, a model declaring [] was gated to unsupported while dispatch forwarded verbatim - an operator-declared vocabulary that excludes the forced tier is terminal; the legacy gpt-5.6 regex can no longer resurrect a tier the override narrowed away - rewrite the registry-outranks-override test: create the matching rule so the decision is non-null, assert unconditionally, pin gpt-5.6 narrowing, alias namespace parity, and use the deterministic xai/grok-4.6 fixture * docs(changelog): clarify override scope for registry-declared models * test: drop placeholder issue reference from test names * chore(changelog): name fragment after PR #12686 |
||
|
|
821d02ba13 |
fix(providers): parse per-vendor-route reasoning.effort_values in discovery (#12730)
OpenAI-compatible model discovery does not recognize per-vendor-route reasoning vocabularies declared under vendors.<vendor>.capabilities.reasoning in GET /v1/models (Merge Gateway's documented catalog shape), so synced models carry no supportedThinkingEfforts/defaultThinkingEffort and operator effort data resets on every model sync; models whose upstream accepts a native max tier cannot be used with forced-max reasoning rules. Parse the shape into the existing supportedThinkingEfforts pipeline, intersected across vendor routes: the same canonical model declares different vocabularies per route and unpinned requests self-narrow to a route honoring the requested level, so a synced tier must be honored on every route the model can land on. Routes without effort_values declare no effort control and are excluded; disjoint vocabularies produce an authoritative empty list (no fall-through to generic tier shapes). detectDefaultThinkingEffort falls back to the intersection's highest tier ranked by the canonical effort order — only when the vendors shape is the record's winning vocabulary source, never escaping a flat or nested declared list. Detection is shape-gated, not provider-gated; Zod-validated (Hard Rule #7) with malformed vendor and tier entries dropped individually (discarding a whole route would widen the intersection, fail-open). Precedence: flat field > reasoning.supported_efforts / metadata (#7694) > vendor-route intersection > capabilities.effort_tiers (#9160) / supported_reasoning_levels / thinking.levels (#8347). |
||
|
|
7e0c9f526a |
feat(sse): allow disabling conversation tracking (#13150)
* feat(sse): allow disabling conversation tracking * docs: document OMNIROUTE_DISABLE_CONVERSATION_TRACKING |
||
|
|
30451e63af |
fix(sse): preserve Fable mid-conversation cache prefixes (#13173)
* fix(sse): preserve Fable mid-conversation cache prefixes * docs(changelog): add fragment for fable cache prefix fix |
||
|
|
46730700f1 |
feat(usage): show separate Fable weekly limits (#13266)
* feat(usage): show separate Fable weekly limits * docs(changelog): add fragment for fable weekly usage |
||
|
|
2fa6ef0bdd |
security(runtime): harden TLS provenance, lifecycle, and public error boundaries (#11742)
* security(deps): pin and verify tls-client native artifacts * docs(changelog): link tls-client provenance PR * security(runtime): harden TLS and public error boundaries * security(runtime): resolve CodeQL error-boundary findings * security(lmarena): close public stream error boundary * fix(lmarena): normalize public error statuses * chore(quality): rebaseline chatCore.ts for the surviving log-boundary hardening open-sse/handlers/chatCore.ts 6219 -> 6287. This is the one part of #11742 that survived the rebase: sanitizeErrorMessage on the plugin onError hook, on the semaphore-timeout path and on failureMessage before it reaches console.log and the call log, sanitizeUpstreamDetails on the malformed-response log, and getSafeErrorMetadata + try/catch where hostile (Proxy) metadata could throw. That is the LOG boundary, which is broader than Hard Rule #12 (responses). The rest of the PR was dropped as already landed on the tip. |
||
|
|
57729db54d |
chore(quality): rebaseline the four ceilings the 2026-09-17 merge wave moved (#14002)
Measured on the clean tip (
|
||
|
|
83fa4328f3 |
feat(providers): add xKiro (#12648)
* test(catalog): pin the 2026-09-02 free-tier re-audit facts for gemini, ollama-cloud, groq, nara and mistral * feat(providers): add xKiro (5M tokens/day free plan, 39 pinned free models) * fix(catalog): re-audit gemini, ollama-cloud, groq, nara and mistral against official pages * docs(providers): xKiro in the provider reference, counts and free-tier headline (~1.66B) * fix(catalog): restore the console-verified Mistral 1B pool and harden its regression test * docs(free-tiers): move headline to the re-audited ~1.50B and refresh pool counts * chore(free-tiers): retire stale Groq free-tier text and preset model; fix catalog header * docs(providers): align the remaining visible provider/executor counts with the catalog * docs(providers): align remaining free-tier count chips and metadata * docs(free-tiers): state the evidence-comment rule honestly and retire the last "14.4K RPD" Groq texts * docs(free-tiers): retire the stale Gemini onboarding quota text * docs(free-tier): refresh catalog-entry counts to 442 after base sync * docs(providers): re-sync provider and free-tier counts after merging release/v3.8.51 * docs(providers): re-sync residual counts after the base merge * docs(free-tiers): restore README spacing lost in the merge and re-sync the guide counts * docs(free-tiers): re-sync numbers after merging release/v3.8.51 (Cerebras reclassified upstream) * fix(docs): keep the NaraRouter plans endpoint out of the API-path checker; rebaseline gateways.ts (+3) * chore(quality): rebaseline gateways.ts file-size cap for the xKiro entry (+20) --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> |
||
|
|
b6975537c1 |
fix(providers): remove the chipotle/pepper provider (#13131) (#13913)
* fix(providers): remove the chipotle/pepper provider (#13131) amelia.chipotle.com (the reverse-engineered Amelia chat-widget backend chipotle/pepper-1 talked to) now returns 404 on every route, including root, from its Azure Application Gateway — confirmed live 2026-09-15. This regressed from a WS handshake timeout (#4037, June 2026) to a fully decommissioned host, so the upstream protocol cannot be fixed. Owner decided to retire the provider entirely (Option B), following the phind/kluster quiet-removal precedent: no REMOVED_PROVIDERS.md entry (reserved for operator takedowns), just a one-line note under FREE_TIERS.md "Removed / no free tier". Removed every surface: executor, registry entry, executors/index.ts and providers/index.ts wiring, noauth provider catalog entry, ProviderIcon generic-fallback set, the autoCombo exclusion-list comment, the chipotle_error code from the sanitizer allowlist, PROVIDER_REFERENCE.md (regenerated), and every doc/test reference. Regression test: tests/unit/issue-13131-chipotle-provider-removed.test.ts asserts the provider is fully gone from the executor registry, the provider REGISTRY and the noauth catalog, and that the executor module no longer resolves — not a live-network repro (flaky/third-party). Several existing tests used "chipotle" only as a generic noAuth-provider example (proxy scoping, error classification, onboarding, fallback text) with no chipotle-specific behavior under test; those were re-pointed at another still-existing noAuth provider (cloudflare-playground / duckduckgo-web) rather than weakened. * test(providers): document the agnes-cn/chipotle count coincidence (#13131) provider-node-reserved-prefix.test.ts's REGISTRY id+alias walk was already red on the base tip (414 vs. expected 412) from agnes-cn (#13399, +id/+alias). Removing chipotle's REGISTRY id/alias in this PR nets it back to 412, making the test pass again without a numeric edit — record why in a comment so it doesn't read as an untracked coincidence later. |
||
|
|
872376bdc1 |
fix(chat): reject null/non-object entries in messages[] (#12643) (#13755)
* fix(chat): reject null/non-object entries in messages[] (#12643) A messages array containing null (or any non-object entry, e.g. [null] or [42]) passed every existing entry guard in chat.ts (#5110/#6402/#6407/#6412) and reached downstream translators/session helpers that read `.role` / `.content` directly off each entry (openai-to-claude.ts, sessionManager.ts, contextManager.ts's fixToolPairs), crashing with a raw TypeError and surfacing as an HTTP 500 instead of a clean 400. The route's Zod schema is intentionally wide (z.array(z.unknown())), so this shape check belongs in the handler's guard chain. Adds one more entry-shape guard clause to the same chokepoint, rejecting the request with a clear 400 before any routing or upstream call. Regression test: tests/unit/chat-messages-entry-objects-12643.test.ts PR #12644 (@soroush5) proposed this exact fix but was closed without merging on 2026-09-12; this re-implements it fresh against the current tip using the same guard shape and error message. Originally-proposed-by: @soroush5 in #12644 Co-authored-by: soroush5 <mrsoroushahmadi@gmail.com> * chore(quality): refix the chat.ts ceiling for the merged tree This branch rebaselined src/sse/handlers/chat.ts against an older tip. After merging the current release tip the combined file is 2520 lines, so the 2500 ceiling no longer covers it. The tip alone is already at 2509 — above the 2500 this PR had frozen — so most of the gap is inherited, not introduced here. This PR's own contribution is the +10 of the messages-entry guard itself. Ceiling refixed at the value the gate reports for the merged tree. --------- Co-authored-by: soroush5 <mrsoroushahmadi@gmail.com> |
||
|
|
d6f720bceb |
feat(i18n): new-key gate rejects __MISSING__ markers; skills translate new keys in parallel (#13996)
On 2026-09-16 eight feature PRs added 61 keys to src/i18n/messages/en.json and stamped `__MISSING__:<en>` into all 65 locales instead of translating. check-new-key-coverage accepted the marker as "the key reached the locale", so nothing blocked the PRs, and the blocking real-translation ratio gate then failed on the release tip for everybody (pt-BR 3.2 % > 2.5 % + 0.5). - scripts/i18n/check-new-key-coverage.mjs: a leaf whose value starts with `__MISSING__:` is judged exactly like an absent leaf; the FAIL message names the marker as the cause and prints the per-locale sync-ui-keys command and the parallel runner. Header/JSDoc updated. - tests/unit/i18n-new-key-coverage.test.ts: "a new key that only carries a __MISSING__ marker is flagged" (was the inverse case, which encoded the old contract); the other six cases unchanged and green. - scripts/i18n/translate-new-keys.sh (+ `npm run i18n:translate-new-keys`): committed, detached-safe runner — flock queue, N workers (default 5), 3 attempts per locale of `sync-ui-keys.mjs --translate-markers --batch-size=40`, per-locale logs/.exit + batch.log/batch.status/batch.rc/batch.pid under _artifacts/i18n-new-keys/, non-zero exit while any locale still carries a marker, refuses to start (exit 2, names the five OMNIROUTE_TRANSLATION_* vars) when the backend env is absent. Reads only the OMNIROUTE_TRANSLATION_* lines of the repo .env; kills nothing, matches nothing by name. - docs: QUALITY_GATES.md (gate table + check-new-key-coverage section) and I18N.md (gate table + "Translating the keys a branch adds" subsection). The implementation/port/merge skills reference the new shared snippet `.agents/skills/_shared/i18n-translate-new-keys.md` (skills repo, separate). |
||
|
|
ceafa55824 |
fix(providers): select and verify the requested gemini-web model/mode before answering (#13381) (#13919)
Root cause: GeminiWebExecutor.execute() opened the identical fixed https://gemini.google.com/app URL and ran the identical Playwright interaction sequence for every advertised gweb/<model> id. `model` was read only AFTER the response was captured, purely to stamp the OpenAI-shaped response — never to influence what was actually clicked/typed, so two different advertised models produced byte-identical automation and the response `model` field was a caller-supplied label, not an observed fact. Fix (owner decision, Option B): a new model -> Gemini UI mode map (open-sse/executors/gemini-web/modeSelection.ts) drives an in-browser selection step before anything is typed — try the mode control, read back the active-mode indicator, and only proceed on a confirmed match. An unconfirmed model, or a requested Extended Thinking control that cannot be confirmed (#13381 follow-up comment), fails closed with 400 unsupported_control_for_provider instead of silently running the account default under the requested label. The selectors involved are UNVALIDATED (no live Gemini account from this checkout) — see the PR's "Selector set is UNVALIDATED" section and the required live smoke. Regression test: tests/unit/issue-13381-gemini-web-model-selection.test.ts |
||
|
|
d8ad12f22d |
docs(agents): advance the documented Bun pin to 1.4.2 (#13946)
#13661 bumped the exact `bun` devDependency (and `@types/bun`) from 1.4.0 to 1.4.2, and #12977 moved the `oven/bun` image to 1.4.2-slim. AGENTS.md still documented 1.4.0, so the guide disagreed with package.json for every agent reading it as authority. Version string only. The surrounding policy is deliberately unchanged: Bun stays confined to the allow-listed gate/generator scripts plus the `test:bun:db` smoke suite, and Node remains the only supported runtime. Operator approved editing this protected surface for this change. |
||
|
|
b7192b72e2 |
fix(thinking): parse/scrub DSML tool-call markers and recognize adaptive thinking (#12905)
* fix(thinking): recognize adaptive thinking + parse/scrub DSML tool-call markers
Two defects combined to break DeepSeek-V4-Flash turns and raise 502
empty_response on Claude Code autocompact.
Defect 1 — DSML tool-call markers leaked as visible content:
DeepSeek-V4-Flash occasionally emits tool calls in a non-standard DSML
text format using full-width pipes instead of the OpenAI tool_calls JSON.
Two shapes appear in production call logs:
- complete block: <|DSML|:Read><path>...</path></|DSML|:Read>
- stray closers (truncated call): </|DSML|parameter></|DSML|invoke>
</|DSML|tool_calls>, sometimes trailing a system-prompt echo
The openai-compatible path never parsed these, so the markers leaked to
the client as visible content and the turn ended incomplete.
Fix: add open-sse/utils/dsmlToolCalls.ts — parseDsmlToolCalls() converts
complete DSML blocks into OpenAI tool_calls and strips stray closing
markers from content (streaming-safe via a holdback for partial openers).
Wire it into the response translator before extractXmlInvokeBlocks so
DSML and XML invoke tool calls share the same pending queue.
Defect 2 — adaptive thinking silently suppressed:
A prior inline === 'enabled' check on body.thinking.type silently
suppressed adaptive (the intent Claude Code actually sends), so
reasoning was dropped. The model then emitted DSML tool-call markers
as plain text, producing an incomplete stop finish. Fix: use
hasActiveClaudeThinking() (which recognizes enabled AND adaptive) to
set requestedThinking, thread it through stream.ts and translator
state, and gate thinking block emission on state.requestedThinking
so upstream reasoning_content only relays when the client opted in.
Tests: 29/29 (6 dsml-tool-calls, 5 thinking-active-claude-adapter,
3 translator-resp-dsml-integration, 15 translator-resp-openai-to-claude
incl. requestedThinking suppression regression). typecheck:core clean.
* fix(sse): strip echoed system-prompt preamble + preserve large analysis/summary blocks
DeepSeek-V4 and similar models echo the OMNIROUTE_SYSTEM_INSTRUCTION_APPEND
directive (appended to the system tail by claude-to-openai.ts) and whole chunks
of the system prompt (<analysis>/<system-reminder>/<summary> blocks, prose
reproductions of the superpowers skill section) verbatim at the START of their
reply — the 'system message leak' persisting after the request-side fix.
Add two streaming-safe preamble strippers in directivePreambleStripper.ts:
- createDirectivePreambleStripper(directive): drops a leading reproduction of
the exact configured directive across arbitrary SSE chunk boundaries.
- createSystemPreambleStripper(): removes <analysis>/<system-reminder>/
<summary> echo blocks and known prose heads (Phase B) from the very start
of a stream, only while the stream is still a preamble.
Wire both into openai-to-claude.ts content-delta path: chain the exact-directive
stripper then the system-echo stripper before DSML/XML-invoke parsing, so a
leading system echo is dropped before it reaches the client.
Preserve large blocks (>= SYSTEM_ECHO_THRESHOLD=1000 chars) and blocks with no
trailing content — these are the model's real response (e.g. a Claude Code
autocompact summary), not a short system-echo. Stops the autocompact
empty-response regression where a whole-summary <analysis> block was stripped
to empty (3a8515).
Regression: origin's markdown-boundary feature (bufferedPrefix /
splitMarkdownBoundary, commit
|
||
|
|
20f3900889 |
fix(claude-web): add charset=utf-8 to Content-Type headers to fix Arabic/Persian UTF-8 mojibake (#13416) (#13419)
* fix(claude-web): add charset=utf-8 to Content-Type headers to fix Arabic/Persian UTF-8 mojibake Fixes #13416 The Claude Web endpoint and the outer SSE streaming pipeline were returning Content-Type headers without an explicit charset parameter: - stream.ts responseHeaders() returned 'application/json' and 'text/event-stream' without charset - responseHeaders.ts buildStreamingResponseHeaders() returned 'text/event-stream' without charset While RFC 8259 defaults JSON to UTF-8 and the SSE spec defaults text/event-stream to UTF-8, some HTTP clients (notably VS Code Chat on Windows) fall back to ISO-8859-1/Latin-1 when no charset is declared, causing multi-byte UTF-8 characters to appear as mojibake. For example, the Persian word for hello (سلام, UTF-8 bytes D8 B3 D9 84 D8 A7 D9 85) was decoded as Latin-1, producing the garbled output 'سلام'. Fix: - claude-web/stream.ts: append '; charset=utf-8' to the Content-Type header in the responseHeaders() helper, with a guard to avoid double appending if the caller already includes a charset - chatCore/responseHeaders.ts: hardcode 'text/event-stream; charset=utf-8' in buildStreamingResponseHeaders() Tests: - 11 new regression tests in claude-web-utf8-mojibake-13416.test.ts covering Persian, Arabic, mixed-script, emoji, and chunk-boundary-split scenarios across both streaming and buffered response paths - All 7 existing claude-web-stream tests pass - All 13 response header tests pass - All 8 adaptive-admission-lifecycle tests pass * test(sse): confirm streaming charset header and add byte-level UTF-8 repro (#13416) Independently verified the mojibake root cause before trusting the charset fix: OmniRoute's claude-web decoder already reconstructs a Persian/Arabic multi-byte UTF-8 sequence split across a chunk boundary correctly because it decodes with TextDecoder({ stream: true }); a naive per-chunk decode (without stream state) is what actually produces the U+FFFD garbling. Updated the 2 exact Content-Type assertions in chat-pipeline.test.ts to match the new "text/event-stream; charset=utf-8" header, which is already the convention used by the other streaming executors (uc.ts, maxai.ts, codex-app-server.ts, etc). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Koosha Pari <koosha@phenotype.ai> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> |