Integrated into release/v3.8.29 — reconciled with #4188 (chatCore split part A): both byte-identical extractions stacked; baseline chatCore.ts -> 5060.
Integrated into release/v3.8.29. Thanks @megamen32! Reconciled on your branch: the branch had drifted behind release and would have reverted the centralized vision detection (#4072), the no-thinking gateway variants (#4145) and #4164's auto/* loop — those are kept, and your genuine improvement is applied on top: each auto/* /v1/models entry now carries advertised context/output limits + capabilities (createBuiltinAutoCombo, with a fallback to the minimal entry on resolve failure). Added a regression test (Rule #18) and bumped the catalog.ts file-size baseline 1440→1463 with justification.
Integrated into release/v3.8.29. combo god-file split part 3 (QG v2 Fase 9 T5 D8, reduced): auto-strategy scoring/intent/tag-routing/candidate-pool/quota-soft registry extracted byte-identically to open-sse/services/combo/autoStrategy.ts (434 LOC, <cap); buildAutoCandidates kept in combo.ts to avoid a combo⇄autoStrategy cycle; 4 public symbols re-exported. Validated: lint clean, file-size OK (combo.ts 3819→3432), check:cycles OK, autoStrategy has no combo-barrel import, 16/16 re-export tests.
Integrated into release/v3.8.29. chatCore god-file split Fase 2 part A (QG v2 Fase 9 T5 C2-C3-C5): 3 byte-identical leaf modules (passthroughHelpers/responseHeaders/telemetryHelpers) under open-sse/handlers/chatCore/; chatCore.ts re-exports the 5 previously-public symbols so existing importers resolve. Validated: lint clean, file-size OK (chatCore 5445→5265), 108/109 symbol-importer tests (the 1 fail = known executor-200ms timing flake under load; 65/65 clean on isolated re-run).
Follow-up to #4091/#4153: the generic capture re-attach in
createPreparedRequestLogger().body() was only regression-tested on the
native-Claude OAuth cloak. Add a dedicated test driving the real
cloakAntigravityToolPayload (_ide suffix scheme) through the
JSON.parse(JSON.stringify()) capture round-trip, asserting the reverse
map survives, stays non-enumerable, and that all-native traffic yields no
spurious map. Proven RED with the #4153 re-attach removed. No production
change.
Closes#4181
Integrated into release/v3.8.29 — add moonshotai/kimi-k2.7-code to the kmca (kimi-coding-apikey) catalog (KIMI_CODING_SHARED.models), requested in discussion #3737. On resync: the PR's qwen-web PROVIDER_MODELS_CONFIG addition (issue #3931 bug #3) was already shipped by #4172 — dropped the duplicate, kept release's entry (which has the Array.isArray guard); kept #4183's KIMI_K27_MODELS spread + added the moonshotai-prefixed entry (union). Also reconciled the file-size baseline for openai-to-gemini.ts (#4180 merged without its +20 bump). Validated: 13/13 tests (kmca catalog + qwen-web parse + kimi registration) + file-size green.
Integrated into release/v3.8.29 — correct contextLength for theoldllm models. Adds an entry-level defaultContextLength (200000) + per-model overrides (GPT-5.4 400K, Gemini 3 Flash/Pro 1M, Claude/DeepSeek 200K) so the models report their real context windows instead of null. Thanks @herjarsa! On review we added a regression test (the resolved window was null before this fix), satisfying the PR Test Policy gate. 4/4 tests + eslint + prettier clean.
Integrated into release/v3.8.29 — default includeThoughts for modern Gemini (2.5+) in the openai->gemini translator, so the model's reasoning is marked thought:true and routed to reasoning_content instead of leaking into visible content (#4170). Thanks @dhaern! On review we regenerated the gemini golden snapshots (basic-chat/with-system) and committed the new modern-gemini-default-thinking snapshot. 61/61 translator + golden tests pass; eslint clean. Respects explicit client thinking config (only applies when none is set) and excludes gemini-1.x / non-thinking 2.0.
Integrated into release/v3.8.29 — register Kimi K2.7 Code (kimi-k2.7-code + -highspeed). New KIMI_K27_MODELS in providers/shared.ts (262144 ctx, vision+reasoning, temperature/top_p stripped = fixed sampling), wired into kimi-coding, kimi-coding-apikey, moonshot and kimi. Validated: 6/6 unit tests (all 4 providers advertise it + context + reasoning + param stripping) + file-size; PR CI CLEAN. OAuth coding endpoint validated live on the test VPS.
When an upstream Gemini SSE stream emitted partial content followed by a JSON
error object ({"error":{"code":503,...,"status":"UNAVAILABLE"}}) instead of a
candidates payload, the gemini→openai translator dropped it: the no-candidate
branch only handled promptFeedback and returned null otherwise, so the stream
ended with finish_reason "stop" and the client got HTTP 200 with a truncated
body — masking the failure and skipping combo fallback.
geminiToOpenAIResponse now detects an error object (optionally wrapped in
response), records it as state.upstreamError preserving the real status (503/
UNAVAILABLE, or 429 for RESOURCE_EXHAUSTED), and lets stream.ts error the stream
out via the existing onFailure/buildErrorBody/controller.error path — the same
mechanism the openai-responses translator already uses.
Closes#4177
Co-authored-by: hartmark <hartmark@users.noreply.github.com>
Integrated into release/v3.8.29 — TLS-terminating capture for TPROXY (decrypt 2/N). New src/mitm/tproxy/tlsCapture.ts: TLS-terminates the intercepted client with a per-SNI leaf from the dynamic CA (#4173), feeds plaintext to an internal http.Server, captures to the Traffic Inspector as source 'tproxy' (sanitizeHeaders + maskSecret), and forwards re-encrypted via an injected seam (realForward uses connectMarked for anti-loop, #4169). Secure-by-default: rejectUnauthorized defaults true; errors via sanitizeErrorMessage; bounded timeouts + closeAllConnections. CaptureSource enum + Zod widened with 'tproxy'. Validated: 17/17 unit tests (real local TLS round-trip + secret masking + 502/error path), targeted typecheck of changed files clean, file-size + test-discovery green.
Integrated into release/v3.8.29 — import-only-free-models connection option + free/paid list filters and free-first sort. Thanks @felipesartori! On review we relocated the 3 React-component tests from tests/unit/ to tests/unit/ui/ so a runner actually collects them (check:test-discovery had flagged them as orphans), bumped the file-size baseline for the two connection modals (cohesive toggle, mirroring #3879/#2997), and synced the branch with the current release. 12/12 vitest + 16/16 node tests green.
Integrated into release/v3.8.29 — raise Node heap for local next build (extends #4104 Docker OOM fix to the native path). Validated: 10/10 tests, typecheck:core, file-size, test-discovery, eslint green.
qwen-web (cookie provider) had no PROVIDER_MODELS_CONFIG entry, so its model-
discovery page returned an empty/stale local catalog — the OAuth fallback at the
top of the route only fires for provider===qwen, so qwen-web fell through to the
no-config branch.
Added a qwen-web entry that fetches the public https://chat.qwen.ai/api/v2/models
endpoint (no auth header configured/sent) and parses the
{ data: { data: [{ id, name, owned_by }] } } shape, with a flatter { data: [] }
fallback.
This is Problem #3 of #3931 (diagnosed by @thezukiru). Problem #1 (validator
bare-token false-positive) shipped earlier in the merged PR #3958; Problem #2
(empty stream from Qwen WAF bot-detection on the streaming endpoint) is a
separate upstream/stealth concern and stays open.
TDD: tests/unit/qwen-web-models-discovery-3931.test.ts mocks the upstream and
asserts source==='api' + the live ids (and the flatter shape), RED 0/2 ->
GREEN 2/2. Rebaselined route.ts 2512->2531.
Co-authored-by: thezukiru <thezukiru@users.noreply.github.com>
OmniRoute ships a zero-setup auto/* catalog (auto/best-coding, auto/pro-
reasoning, …, 16 AUTO_TEMPLATE_VARIANTS) that the dashboard advertises and that
resolve on demand via createBuiltinAutoCombo, but the /v1/models listing only
emitted persisted DB combos + provider models. Clients that build their model
picker from /v1/models (e.g. Hermes Agent) never saw any auto/* option.
The catalog now emits every AUTO_TEMPLATE_VARIANTS id (owned_by: combo) at the
top of the list, deduped against persisted combos so a user combo can't shadow
or double a built-in id.
TDD: tests/unit/models-catalog-auto-combos-4164.test.ts asserts every auto/*
variant is advertised, they sit at the top, and ids are unique (RED 2/3 ->
GREEN 3/3). Aligned the existing 'context_length for individual chat models'
test to exclude owned_by:combo entries (combos/routers resolve dynamically and
have no fixed context_length — they are not individual chat models).
Under concurrent load, requests exceeding the per-connection rate-limit queue
budget (resilienceSettings.requestQueue.maxWaitMs) are dropped by Bottleneck with
its raw 'This job timed out after <maxWaitMs> ms.' message. That string is
indistinguishable from an upstream gateway timeout, so the 502 body and call-log
last_error masquerade as a provider outage across unrelated providers (TI:0|TO:0)
— the #4165 reporter spent ~3h misdiagnosing local queue saturation as upstream
failures.
withRateLimit now rewrites that specific Bottleneck error into a clear,
OmniRoute-owned message: names the knob (requestQueue.maxWaitMs, Settings →
Resilience), disclaims an upstream timeout, preserves the original as cause, and
tags code=RATE_LIMIT_QUEUE_TIMEOUT. Behavior unchanged — the job is still dropped
so combo falls back.
TDD: tests/unit/rate-limit-queue-timeout-message-4165.test.ts drives a real
Bottleneck expiration (tiny maxWaitMs + a slow job) and asserts the surfaced
error carries the code + clear message and no longer leaks 'This job timed out'
(RED 1/2 -> GREEN 2/2). Rebaselined rateLimitManager.ts 1022->1035 at the
existing catch chokepoint.
models.dev catalogs Pixtral 12B under the short id `pixtral-12b` (attachment:
true, image modality), while requests use the Mistral API alias
`pixtral-12b-latest`. getSyncedCapabilityForResolved tried only exact / raw /
static-spec-canonical ids, all of which miss the short form, so vision fell
through to the #4071 model-id heuristic and `attachment` stayed null.
Add a last-resort fallback that retries the synced lookup with a trailing
`-latest` stripped, so synced metadata wins for these aliases. Models whose
`-latest` id is stored verbatim (e.g. pixtral-large-latest) keep resolving
directly. TDD: tests/unit/model-capabilities-mistral-vision-sync-4073.test.ts
(RED 2/4 -> GREEN 4/4) seeds a synced fixture and asserts the verdict comes via
the synced path (attachment === true/false), not the heuristic.
The models.dev sync remains manual-only (no scheduled refresh) — documented as a
separate follow-up.
Follow-up to #4054. The Request Logger still froze auto-refresh on some hosts
(reported on 3.8.28 Docker, works on 3.8.24). #4054 made the INITIAL visibility
fail-open, but the pause is event-driven: a host that fires a one-shot
visibilitychange->hidden and then keeps reporting 'hidden' (or recovers without
firing the event again) left visibleRef stuck false, so the interval ticked but
never polled - only the manual Refresh button worked.
The poll tick now also re-checks the LIVE document.visibilityState, and a window
'focus' listener re-arms polling (a focused window is a reliable signal the page
is actively viewed). A genuinely backgrounded browser tab still pauses (reports
'hidden' and never receives focus), preserving the #3109 optimization.
TDD: two new jsdom regression tests (RED on current release, GREEN after);
the existing 'pause on real background' test stays green.
Restore MCP/third-party tool names on the native Claude path (#4091): re-attach the non-enumerable _toolNameMap dropped by the request-capture JSON round-trip, so the response-side un-cloak restores MCP/snake_case names. Integrado em release/v3.8.29.
Case-by-case follow-up on the 2026-06-17 refresh deferred items. Each dead tier was
re-verified against its official source on 2026-06-18 before flipping (qwen-web lesson):
- aimlapi -> hasFree:false. docs.aimlapi.com/faq/free-tier states "The Free Tier is
currently paused"; now pay-as-you-go (min $20 top-up).
- gitlawb + gitlawb-gmi -> hasFree:false. Free MiMo revoked 2026-05-24 (GitHub
Gitlawb/openclaude#1345 "OPENGATEWAY REVOKED THE ONLY 1 FREE MODEL"); Nemotron promo
ended 2026-06; Opengateway is now a pay-as-you-go credit gateway. Removed 40 stale
gitlawb-gmi catalog records.
- yi -> hasFree:false. Yi-Light retired; platform.01.ai is pay-as-you-go (Yi-Lightning
paid); open weights are download-only. Removed 1 stale catalog record.
- iflytek / sparkdesk -> kept hasFree:true with a ToS-caution freeNote: Spark Lite is
free but the ToS prohibits programmatic extraction / relay to third parties (iflytek
also requires Chinese real-name auth).
- monsterapi -> freeNote tightened to "one-time signup trial" (recurring plan = 0
credits/mo); the budget catalog already classified it one-time-initial.
The dead providers stay registered (paid use still works); only the free-tier
advertisement is removed. FREE_TIERS.md "Removed" note + budget table reconciled.
Tests: gitlawb-provider guard updated (hasFree now false for both gitlawb variants);
discontinued-providers-2026 extended to cover the 4 new flips + an iflytek ToS-caution
assertion. Re-verified live 2026-06-18 (Hard Rule #18).
Mid-stream continuation for truncated streams (FCC port, Task 4.4). Opt-in (STREAM_RECOVERY_MIDSTREAM_ENABLED, default OFF); plain-text OpenAI-compat only, nunca com tool-call; OFF byte-idêntico ao #4131. Integrado em release/v3.8.29.
Per-provider sliding-window rate-limit fallback (FCC port, Fase 8.2): opt-in (empty map = no-op), burst-free sliding window as a floor for header-less providers; Bottleneck still applies. Integrado em release/v3.8.29 (conflito de file-size-baseline resolvido por união com #4145).
No-thinking gateway model IDs (FCC port, Fase 8.1): synthetic claude-3-omniroute-no-thinking/<provider>/<model> catalog id that resolves to the real model with reasoning suppressed. Integrado em release/v3.8.29.
Owner found the 46px graph-paper cells a touch large for the dashboard
layout; tightening --grid-size to 32px (a ~30% reduction). Mirrors the
matching change on the marketing site so the two stay identical.
Guard test + design.md updated. Build verified: --grid-size compiles to 32px.
Follow-ups from the 2026-06-17 free-tier refresh (#4089) that did NOT land in the
original merge — orphan commit ccc13c27c was pushed after the PR had already merged,
so these three follow-ups never shipped. Re-applied onto release/v3.8.29.
- providers.ts: hasFree flipped to false for confirmed-dead free tiers
(phind x2, kluster, glhf, chutes) so onboarding stops advertising a free tier
that no longer exists. Deliberately NOT flipped: gitlawb (dedicated test guards
hasFree:true), aimlapi/theoldllm (medium-confidence research) — flipping risks a
qwen-web-style false removal.
- New test tests/unit/discontinued-providers-2026.test.ts guards both the flips and
the intentional keeps.
- Headline rounded ~1.5B -> ~1.6B in README (hero/badge/intro/bullet), FREE_TIERS
honest-headline sentence, and FREE-TIERS-GUIDE total. The TL;DR table + live card +
SVG keep the precise computed ~1.54B.
- FREE_TIERS per-provider table regenerated from the per-model catalog (pool-deduped,
2026-06-17): the old inflated rate-limit rows (sparkdesk 2.59B / tencent 2.07B /
siliconflow 1.73B) are gone — uncapped providers now show 'uncapped*'.
Grid wallpaper opacity tuned up so it's actually visible on the dense
dashboard: light 0.045->0.07, dark 0.035->0.06 (the site's 0.045/0.035
read as invisible behind the cards/chrome that cover most of the viewport).
Focus-ring reconcile (design.md C6): the form controls (Input, Select,
Textarea, Toggle, Checkbox) now focus on the accent (violet) ring to match
the global --focus-ring, instead of the red ring-primary that collided with
the red error state. Error rings stay red.
Guard test updated (grid opacity + new C6 assertion). Build verified:
--grid-line compiles to #00000012 / #ffffff0f and focus:ring-accent/30 +
focus:border-accent/50 utilities are generated.
Unified visual identity (design phases 1-4): grid, primitives, tables, form controls. Integrado em release/v3.8.29: clsx + tailwind-merge adicionados à dependency-allowlist (já declarados no package.json, consumidos por src/shared/utils/cn.ts) para o check:deps anti-slopsquatting passar.
Free DuckDuckGo web search as last-resort provider (FCC port, Fase 6). Integrado em release/v3.8.29: conflito de file-size-baseline.json resolvido (união dos rebaselines duckduckgo_free + stream_recovery) e teste de contagem de providers atualizado 12->13.