* fix(resilience): honor declared effort vocabulary in reasoning rule gate
The reasoning-routing rule capabilityFor() hardcoded a gpt-5.6-(sol|terra|luna)
whitelist for forced max/ultra, rejecting every other thinking-capable model
even when the model's resolved capabilities declare the requested tier (synced
supportedThinkingEfforts or an operator Model Overrides reasoning_efforts
override). This 400'd direct calls with "Reasoning effort 'max' is not
supported by the configured target" for models like Merge Gateway
zai/glm-5.3-flash, which natively accepts low|high|max.
The gate now treats a declared vocabulary containing the requested tier as
authoritative, mirroring the dispatch-time sanitizer
(open-sse/executors/base/reasoningEffort.ts) which already forwards declared
tiers verbatim. Undeclared models keep the legacy gpt-5.6 regex verdicts and
the unknown passthrough.
* fix(resilience): gate forced max against the static registry the sanitizer clamps with
Adversarial review finding: the gate read supportedThinkingEfforts from
getResolvedModelCapabilities, which prefers the DB override over the registry.
For a registered model with a narrow registry vocabulary and a widening
operator override, the gate passed forced max but the dispatch-time sanitizer
(executors/base/reasoningEffort.ts) clamps against the STATIC registry and
would silently downgrade max to the registry ceiling — converting a loud 400
into a silent wrong-effort request.
Order of precedence in the gate now:
1. static registry vocabulary (authoritative — matches sanitizer clamping)
2. declared/overridden vocabulary for unregistered providers (#8057 path)
3. legacy gpt-5.6 regex, then unknown/unsupported verdicts
Also pins the test fixture to a synthetic model id so a future models.dev
sync row cannot flip the unknown-precondition assertion.
* fix(resilience): gate registry lookup mirrors the dispatch sanitizer exactly
Review findings on the forced max/ultra gate:
- resolve the registry through getProviderModels (id->alias namespace) and
match entry aliases, mirroring reasoningEffort.ts — a raw provider id or
alias-spelled model no longer skips the registry branch and diverges from
dispatch clamping
- treat an empty declared vocabulary as no declaration (falls through),
matching the sanitizer's declaredRanked.length>0 guard — before, a model
declaring [] was gated to unsupported while dispatch forwarded verbatim
- an operator-declared vocabulary that excludes the forced tier is terminal;
the legacy gpt-5.6 regex can no longer resurrect a tier the override
narrowed away
- rewrite the registry-outranks-override test: create the matching rule so
the decision is non-null, assert unconditionally, pin gpt-5.6 narrowing,
alias namespace parity, and use the deterministic xai/grok-4.6 fixture
* docs(changelog): clarify override scope for registry-declared models
* test: drop placeholder issue reference from test names
* chore(changelog): name fragment after PR #12686
Merged. The failure mode was concrete — a matched reasoning rule dropped on native Responses/Anthropic paths, model-suffix/account defaults, or fallback preparation, and `_omnirouteReasoningRule` leaking upstream as `Unsupported parameter` — and the fix is carried in the request-local credential context through dispatch, refreshed credentials and fallbacks, with forced effort winning over defaults and client-forged markers dropped at ingress. The 11-case integration suite exercises the real routing/translation modules.
Validated as a combined board first (this PR merged with the 4 siblings of the JxnLexn wave on the release tip): eslint with the frozen suppressions, typecheck:core, check:open-sse-typecheck, complexity, cognitive-complexity, changelog-integrity, docs-sync, migration-numbering, provider-consistency and a duplicate-identifier audit all green, plus 77 passing / 0 failing focused node:test cases across the test files the wave touches. The wave's i18n fill (new keys carried to all 66 locales), free-tier doc counts and file-size rebaseline land in one follow-up PR right after the wave, as with #13904.
Thank you — and for keeping this a runtime-only change with the editor and service-tier work in their own PRs.
* feat(routing): add reasoning-based model and effort routing
* refactor(routing): modularize reasoning and auto-routing pipeline
* fix(routing): remove redundant DB re-export and prevent SQL scan false positives
* fix(routing): resolve reasoning routing review blockers
* fix(i18n): keep release ranking fallbacks outside reasoning
* fix(db): renumber reasoning-routing migration past release tip (124→125)
124_generic_session_affinity_ttl.sql (#7274) has since landed on
release/v3.8.49 at version 124, colliding with this PR's own
124_reasoning_routing_rules.sql. Renumbers to 125 (the next free slot
past the current release tip) and updates the one filename reference
in docs/routing/REASONING_ROUTING.md.
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* chore(db): renumber reasoning-routing migration 125→126 (slot taken by #7360)
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* refactor(api): compact temp-path decls in exportAll GET (complexity-ratchet lines budget)
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
* refactor(api): single-statement auth guard in exportAll GET (function under 80-line cap)
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>