Validated in merge-train 2026-07-18 @ 9084b408b: 12198/12201 pass; single red = earlyStreamKeepalive timer test, confirmed load-flake (6/6 green isolated on release tip AND the merged tree; no boarded PR touches keepalive)
Validated in merge-train 2026-07-18 @ 9084b408b: 12198/12201 pass; single red = earlyStreamKeepalive timer test, confirmed load-flake (6/6 green isolated on release tip AND the merged tree; no boarded PR touches keepalive)
Validated in merge-train 2026-07-18 @ 9084b408b: 12198/12201 pass; single red = earlyStreamKeepalive timer test, confirmed load-flake (6/6 green isolated on release tip AND the merged tree; no boarded PR touches keepalive)
Validated in merge-train 2026-07-18 @ 9084b408b: 12198/12201 pass; single red = earlyStreamKeepalive timer test, confirmed load-flake (6/6 green isolated on release tip AND the merged tree; no boarded PR touches keepalive)
* docs(readme): standardize all tables to full content width
Add a 1px transparent spacer.svg and per-table header spacers so every
markdown table renders at the same ~890px full content width on GitHub
instead of collapsing to its own content width. No table text changed.
* chore(changelog): fragment for #7666
* docs(readme): replace free-tier budget mockup with animated SMIL card
Single detailed card (1200x872, 10s loop, SMIL only — plays inside GitHub's
img sandbox): ~1.6B/mo hero + honest-math panel (struck-through ~10B, 15
providers ToS-flagged), animated budget bar of the 21 countable free pools,
full per-model grid (Mistral Large 3 1.00B -> Auto 25K), ~616M first-month
signup-credit chips, permanently-free no-cap providers + $10 OpenRouter
top-up, and a live used/remaining footer.
The generated mockup docs/screenshots/free-tier-budget-card.svg stays in
place — it is produced by scripts/research/gen-budget-card-svg.mjs and still
referenced by the i18n READMEs (zh-CN/zh-TW); only the root README embed
changes. Registered in the hand-authored table in docs/diagrams/README.md.
* chore(changelog): fragment for #7665
* docs(readme): animate CLI command list + compression flow as SMIL SVGs
Two more README ASCII/text blocks become hand-authored animated SVGs
(SMIL only, GitHub <img>-sandbox safe, DESIGN_SYSTEM.md palette),
following the tier-cascade / pool / combo pattern:
- cli-terminal.svg — compact terminal window (640x500) cycling three
real CLI screens (providers list / combo list / health) with
character-by-character typing, output formats copied from the actual
bin/cli printers (headings, column layout, status colors, circuit
breaker block), plus a scrolling ticker carrying the full 30-subcommand
list the image replaces (also preserved in the img alt).
- compression-pipeline.svg — the 'Client -> 10 engines -> Provider'
flow line as an animated funnel: 10,000 tok in, ~1,080 tok out, token
dots evaporating engine by engine behind the cells, RTK -> Caveman
default stack highlighted, a code token passing through untouched
(always preserved byte-perfect) and the stacked savings math badge.
Registered both in docs/diagrams/README.md (hand-authored table).
* docs(changelog): add fragment for #7637 (CLI terminal + compression SVGs)
* docs(readme): enlarge CLI terminal diagram (full-width, 1200x700)
Per review: the mini 640x500 terminal read too small. Rebuild it as a
full-width widescreen terminal (viewBox 1200x700, embedded at width=100%)
with larger type, wider aligned columns, 6 provider rows and 4 combo rows
so each screen fills the frame. Same 3 real CLI screens, same SMIL, same
DESIGN_SYSTEM.md palette, same command-ticker footer.
* docs(readme): animate pool + combo blocks as SMIL SVG diagrams
Replace the two remaining ASCII blocks in the README with hand-authored
animated SVGs (16s loops, SMIL only — play inside GitHub's <img> sandbox,
DESIGN_SYSTEM.md palette), following the tier-cascade.svg pattern:
- pool-fair-share.svg — key pool "team-codex" fair-share quota: weights
50/30/20, generous mode lending idle shares, 50% threshold crossing,
strict mode holding each key to its cap (verbatim README copy).
- combo-always-on.svg — combo "always-on" priority strategy: 4 fallback
layers with coral hand-off on failure and an uptime bar that never
drops (zero downtime).
Both blocks keep their full flow text in the img alt. Registered in
docs/diagrams/README.md (hand-authored table).
* docs(changelog): add fragment for #7626 (pool + combo SVG diagrams)
* docs: sync provider count to 259 (docs-counts strict gate)
The auto-generated catalog (docs/reference/PROVIDER_REFERENCE.md) is at
259 providers; README.md, AGENTS.md and CLAUDE.md still said 253 —
tripping the strict Provider-count check in check:docs-counts for every
PR targeting the release branch (surfaced red on #7615's Docs Gates
fast-path run). Updates the 8 provider-count mentions across the three
files (marketing badges and AES-256-GCM strings untouched).
* docs(changelog): add fragment for #7616 (provider-count sync)
* docs(readme): replace tier-cascade ASCII diagram with animated SMIL SVG
The 4-tier auto-fallback block in the README becomes a self-contained
animated SVG (docs/diagrams/tier-cascade.svg, 16 KB): a 16s loop in 4
acts where requests flow from the IDE through the smart router into the
active tier, and each quota-out/budget-hit transition hands the traffic
down to the next tier, ending on the always-on free tier. SMIL only — no
JS, no external fonts — so it animates inside GitHub's camo/<img>
sandbox. Content is verbatim from the previous ASCII art; the full flow
is preserved in the img alt text. docs/diagrams/README.md gains a
hand-authored-diagrams section documenting it.
* docs(changelog): add fragment for #7615 (animated tier-cascade SVG)
* docs(readme): align tier-cascade SVG palette with DESIGN_SYSTEM.md
Retrofit to the canonical tokens (docs/architecture/DESIGN_SYSTEM.md §3.1):
dark bg #0b0e14 + the 32px graph-paper grid wallpaper (the product/site
signature), surface #161b22, borders rgba(255,255,255,.08), radius 14,
text-muted #a1a1aa. Brand semantics fixed: the router hub glyph + glow now
use primary #e54d5e (matching the favicon hub mark) and the title carries
the --grad-brand gradient (primary → accent-3); exhaustion states
(quota out / budget hit flashes, spent-tier status dots, hand-off dots)
move from brand coral to the semantic error token #ef4444; topology paths
use accent #6366f1 with accent-2 #8b5cf6 request dots; success stays
#22c55e. Re-validated (0 warnings) and re-verified frame-by-frame.
* feat(providers): add Segmind image+video provider (#6656)
Segmind exposes 200+ hosted image/video models under a single
`POST https://api.segmind.com/v1/{model}` REST shape: x-api-key auth,
JSON request body, raw media bytes response (no JSON envelope).
- New IMAGE_PROVIDERS + VIDEO_PROVIDERS registry entries (format:
"segmind") with a curated starter model list (Flux, SDXL, SD3.5,
Kandinsky for image; Wan, Hunyuan, LTX, Kling for video).
- New connection-metadata entry in specialty-media.ts; segmind added
to IMAGE_ONLY_PROVIDER_IDS and VIDEO_PROVIDER_IDS.
- Dedicated handlers (imageGeneration/providers/segmind.ts,
videoGeneration/providers/segmind.ts) built on a shared REST client
(utils/segmindClient.ts) that centralizes the fetch/error/log path
so both stay under the complexity/max-lines ratchets.
- Extracted the pre-existing Alibaba DashScope video handler out of
the frozen videoGeneration.ts into videoGeneration/providers/
dashscope.ts (no behavior change) to make room for the new Segmind
dispatch branch under the frozen file-size baseline.
- Error responses route through sanitizeErrorMessage() (Hard Rule
#12) — verified by dedicated no-leak tests.
- Regenerated docs/reference/PROVIDER_REFERENCE.md (250 -> 251
providers) and synced the plain-text provider counts in README.md/
AGENTS.md/CLAUDE.md (anchors/badges left untouched).
Tests: tests/unit/segmind-image-video-provider-6656.test.ts (11
cases — registry shape, connection metadata, IMAGE_ONLY/VIDEO_
PROVIDER_IDS membership, mocked-fetch request mapping for both
image and video, and sanitized-error-path assertions for both
upstream error bodies and network exceptions). No live Segmind key
required; response shape (raw media bytes, x-api-key auth) is
sourced from https://docs.segmind.com/ and corroborated against
https://www.segmind.com/models/flux-schnell/api,
https://www.segmind.com/models/sdxl1.0-txt2img/api, and
https://www.segmind.com/models/wan2.1-t2v/api.
Gates run clean: check-file-size, check:complexity-ratchets
(2055/889, both under baseline), typecheck:core,
typecheck:noimplicit:core (no new errors), lint (targeted files),
check:cycles, check:docs-counts (STRICT provider-count drift
resolved), check:docs-sync, check:any-budget:t11,
check:tracked-artifacts, check:provider-consistency,
check:known-symbols.
* test(providers): align APIKEY_PROVIDERS count 167→168 for the new segmind provider (#6656)
Adding segmind to specialty-media.ts grows APIKEY_PROVIDERS by one;
providers-constants-split.test.ts hardcodes the family-partition total.
Legitimate count alignment, not a weakened assertion — all 4 partition/
dedup checks still enforced.
* feat(sse): add Microsoft Designer as image provider (#6672)
Adds `microsoft-designer-web` — an unofficial, reverse-engineered
Bearer-token web-session image provider, modeled on the existing
`chatgpt-web`/`copilot-m365-web` "-web" provider category.
- Registers the provider in WEB_COOKIE_PROVIDERS (src/shared/constants/
providers/web-cookie.ts) and IMAGE_PROVIDERS (open-sse/config/
imageRegistry.ts, new "designer-web" format).
- New handler open-sse/handlers/imageGeneration/providers/designerWeb.ts
implements the submit-then-poll DallE.ashx flow (Bearer access_token +
ClientId/SessionId/UserId headers -> form POST -> poll for
image_urls_thumbnail), wired into handleImageGeneration()'s dispatch.
- The upstream ClientId header is a fixed, publicly-shared value (not a
secret) — routed through resolvePublicCred() per Hard Rule #11, never
as a string literal.
- Registers the token-based credential requirement in
webSessionCredentials.ts so the provider-connect UI asks for the
right field; connection validation falls back to the existing generic
web-cookie session-ping validator (no dedicated validator needed).
- Extracted the KIE image-model catalog into a co-located
open-sse/config/providers/registry/kie/models.ts module (mirrors the
existing lmarena/directModels.ts pattern) to keep imageRegistry.ts
under the file-size cap while adding the new provider entry.
- Regenerated docs/reference/PROVIDER_REFERENCE.md (250 -> 251
providers) and updated the plain-text counts in README.md, AGENTS.md,
CLAUDE.md.
Tests (tests/unit/microsoft-designer-web-6672.test.ts, 16 cases):
registry-entry shape assertions, the resolvePublicCred() shape
assertion (Hard Rule #11), and the pure header/form-body/response-
parsing helpers plus the handler's submit/poll/error/timeout paths
against a mocked fetch — no live Designer session required.
Reverse-engineered from the g4f MicrosoftDesigner.py provider reference
(researched during #6672 triage); the exact upstream response shape has
not been validated against a live Designer session, so the poll-loop
implementation follows the documented g4f contract as closely as
possible without a live capture.
* fix(providers): satisfy web-cookie executor contract + document designer-web env vars (#6672)
Registers `deepinfra` in the video-gen registry, reusing the DeepInfra
native /v1/inference/{model} endpoint already proven for reranking in
this codebase (same host, Bearer auth, non-OpenAI response shape).
Confirmed synchronous against DeepInfra's own docs (POST {prompt} ->
{video_url, seed, request_id, inference_status}), so no polling loop
is needed. Reuses the already-registered `deepinfra` API-key provider
credential (chat) — no new credential/OAuth flow.
To keep the frozen videoGeneration.ts file-size ratchet from growing,
the new deepinfra-video adapter lives in its own co-located module
(open-sse/handlers/videoGeneration/deepinfraHandler.ts, following the
existing googleFlowHandler.ts pattern), and the pre-existing Leonardo
handler was extracted into videoGeneration/leonardoHandler.ts (pure
code move, no behavior change) to make room.
Adds Novita AI to the video-generation subsystem (VIDEO_PROVIDERS), alongside
its existing text/chat gateway registration. Novita's async video APIs are
per-model (POST /v3/async/<model-slug>, e.g. wan-t2v, kling-v1.6-t2v) sharing
one poll endpoint (GET /v3/async/task-result?task_id=...) — confirmed against
Novita's published API reference. Seeds Wan 2.1 T2V and Kling V1.6 T2V models;
reuses the stored novita provider Bearer apiKey (no separate credential flow).
To stay under the frozen videoGeneration.ts file-size cap, extracted the
existing Alibaba/DashScope handler into a co-located sibling module
(videoGeneration/dashscopeHandler.ts) alongside the new Novita handler
(videoGeneration/novitaHandler.ts) and its pure helpers (videoGeneration/novita.ts).
Also tags novita in VIDEO_PROVIDER_IDS (src/shared/constants/providers.ts) so
it surfaces as a video-capable provider in PROVIDER_REFERENCE.md and A2A
provider-discovery, and regenerates the provider reference doc.
Tests: tests/unit/video-novita-6658.test.ts (18 cases) covering registry
shape, pure helpers (URL building, param normalization, task-id/result
parsing), and full handler wiring (submit->poll->mp4, missing credentials,
missing task_id, task FAILED, task timeout).
* feat(providers): add Freepik (Magnific Mystic) image generation provider (#6654)
Adds an official, API-key-based Freepik image-gen provider using the Mystic
endpoint (POST /v1/ai/mystic -> async task_id -> GET /v1/ai/mystic/{id}
polling), modeled on the existing leonardo.ts generationId adapter pattern.
- open-sse/config/providers/registry/freepik/index.ts: new registry module
(kept separate to avoid pushing the frozen imageRegistry.ts over the
file-size cap) with the 6 real Mystic style models (realism, fluid, zen,
flexible, super_real, editorial_portraits) — not the "Flux/Imagen3" list
from the original feature request, which independent research showed was
stale.
- open-sse/handlers/imageGeneration/providers/freepik.ts: submit+poll
adapter; all error paths route through sanitizeErrorMessage() (Hard Rule
#12), configurable poll interval/timeout via body.poll_interval_ms /
poll_timeout_ms for fast, deterministic tests.
- Registered in providers.ts (IMAGE_ONLY_PROVIDER_IDS) and
apikey/specialty-media.ts (catalog metadata), with the corrected free-tier
note (one-time ~€5 credit, not a recurring "100/month" allotment).
Drops the "100 free credits/month" and "Flux/Imagen3 selectable models"
claims from the original issue - verification showed the free tier is a
one-time ~€5 API credit and Imagen 3 only underlies the `fluid` style, not a
separately selectable model. Domain: api.freepik.com is still live as of
this writing despite Freepik's April-2026 API-docs rebrand to Magnific
(docs.freepik.com -> docs.magnific.com); noted inline for future
re-verification.
Closes#6654
* test: align APIKEY_PROVIDERS count to 171 after freepik + release merge (#7597)
* feat(providers): add Gladia as an async speech-to-text provider (#6657)
Adds Gladia's async pre-recorded transcription API (upload → POST
/v2/pre-recorded → poll result_url) following the existing
AssemblyAI/Kie.ai async-STT pattern:
- New `gladia` entry in AUDIO_TRANSCRIPTION_PROVIDERS
(open-sse/config/audioRegistry.ts), authenticated via the
`x-gladia-key` custom header.
- New `handleGladiaTranscription()` handler
(open-sse/handlers/audioTranscription.ts) wired into the
format dispatch table.
- New `x-gladia-key` case in `buildAuthHeaders()`
(open-sse/config/registryUtils.ts).
- Registered `gladia` in AUDIO_ONLY_PROVIDERS
(src/shared/constants/providers/audio.ts) so it appears in the
auto-generated provider catalog; regenerated
docs/reference/PROVIDER_REFERENCE.md (250 -> 251 providers) and
synced the plain-text provider counts in README.md, AGENTS.md,
and CLAUDE.md.
Real-time/streaming transcription is explicitly out of scope for
this change — OmniRoute has no WebSocket audio-ingestion layer
today; only the async/pre-recorded path (which covers every other
async STT provider already wired in) is implemented.
Tests: 5 new node:test cases in
tests/unit/audio-transcription-handler.test.ts covering the
upload→submit→poll happy path, a terminal Gladia error, and a
missing result_url guard, plus a buildAuthHeaders case in
tests/unit/registry-utils.test.ts for the new x-gladia-key header.
* chore(providers): sync provider counts to 253 + fix base-red APIKEY partition count 168→169 (#6657)
Rebasing gladia onto the advanced release surfaced two count drifts the
RUN_ALL suite trips on: (1) docs provider count is now 253 (multiple
providers merged since this branch was cut); (2) providers-constants-split
already expects 168 but the release has 169 APIKEY entries — a pre-existing
base-red from an earlier provider merge that didn't update the test. Gladia
is STT (adds no APIKEY entry), so 169 is the correct value; aligning it here
also un-reds the release. All 4 partition/dedup checks still enforced.
* feat(providers): add FreeTheAi as an OpenAI-compatible gateway provider (#6670)
FreeTheAi is a free-tier, Discord-signup gateway aggregator — same shape
as hackclub/chutes: OpenAI-compatible chat/completions + /v1/models
discovery, no custom executor/translator needed.
- Registry entry: open-sse/config/providers/registry/freetheai/index.ts
(format: openai, executor: default, apikey/bearer auth, passthroughModels)
- Provider metadata: src/shared/constants/providers/apikey/gateways.ts
- Listed in AGGREGATOR_PROVIDER_IDS (src/shared/constants/providers.ts)
- Unit test verifying registry entry, getExecutor() resolution, aggregator
classification, and provider metadata (tests/unit/provider-registry-freetheai.test.ts)
* chore(providers): sync counts (APIKEY 170, providers 253) after rebase onto advanced release (#6670)
The release advanced heavily since this branch was cut; realign the
family-partition count to the true post-rebase value (170) and the doc
provider totals to 253. freetheai adds exactly one gateway; the rest of
the delta is pre-existing release drift. All partition/dedup checks enforced.
* feat(sse): add EdgeTTS audio-tts provider (#6668)
Registers Microsoft Edge "Read Aloud" as a new no-API-key AUDIO_SPEECH_PROVIDERS
entry — the first WebSocket-transport TTS provider in the registry. Reverse-
engineered/unofficial endpoint, same class of integration already accepted for
other "-web" style providers (chatgpt-web.ts, copilot-web.ts).
- open-sse/executors/edgeTts.ts: pure Sec-MS-GEC token construction (SHA-256
over a public trusted-client-token + rounded Windows file-time ticks, ported
from rany2/edge-tts drm.py), WS message framing (speech.config/ssml),
binary-chunk demuxing, SSML building/escaping, and the WS synth call itself
(injectable WebSocket ctor for tests, lazy `import("ws")` in production so
it never enters esbuild's top-level CJS bundle graph). Per-client-IP
sliding-window throttle (SlidingWindowLimiter) since there's no per-user key
— one abusive deployment could otherwise get the shared trusted token
rate-limited for everyone.
- open-sse/utils/publicCreds.ts: embeds the trusted-client-token via
resolvePublicCred() (Hard Rule #11) — it's a constant hardcoded in every
Edge build and every open-source edge-tts port, not a per-user secret.
- Extracted open-sse/utils/audioResponse.ts (shared response helpers) and
open-sse/executors/awsPollyTts.ts (AWS Polly handler) out of
open-sse/handlers/audioSpeech.ts to stay under its frozen file-size ratchet
baseline while making room for the new branch — no behavior change to
either extracted piece.
- src/app/api/v1/audio/speech/route.ts: thread the caller's IP through to the
handler for the new throttle.
Tests: tests/unit/edgetts-provider.test.ts (23 cases) — Sec-MS-GEC determinism
and cross-check against a hand-derived reference vector, message framing,
binary demux, SSML escaping/injection-safety, registry lookup, publicCreds
shape, and the error path via an injected fake WebSocket (upstream failure ->
sanitized 502, no stack/path leak; Hard Rule #12), plus the per-IP rate limit.
No live upstream is required or used — the reverse-engineered protocol can't
be validated against real credentials, but every pure/testable seam is
covered per the TDD path in the bug/feature validation gate.
* test(mutation): register edgetts-provider.test.ts in stryker tap.testFiles (#6668)
The new provider's unit test covers a mutated module, so the strict
mutation-test-coverage gate requires it in stryker.conf.json's
tap.testFiles. Single-line addition (kept the file's existing formatting).
Notion AI has no public inference API (see closed request #3272), so this
adds it as a new entry in the established web-cookie provider category
(chatgpt-web, claude-web, grok-web, ...): cookie-based auth via the
token_v2 session cookie posted to Notion's undocumented internal
POST /api/v3/runInferenceTranscript endpoint, translating its NDJSON
transcript-patch stream into OpenAI-compatible chat completions.
- NotionWebExecutor (open-sse/executors/notion-web.ts): resolves the
token_v2 cookie (+ optional space_id/notion_browser_id), builds a
Notion transcript from the chat messages, parses the NDJSON response
(cumulative-snapshot semantics, mirroring gemini-web.ts's handling of
#7163), and returns a chat.completion or pseudo-streamed SSE response.
All error paths route through makeExecutorErrorResult (sanitized).
- RegistryEntry under open-sse/config/providers/registry/notion-web/,
registered in providers/index.ts REGISTRY and executors/index.ts
(alias "nw").
- WEB_COOKIE_PROVIDERS entry (src/shared/constants/providers/web-cookie.ts)
with subscriptionRisk + webCookie risk notice, clearly labeled
"(Unofficial/Experimental)".
- Cookie-probe validator (validateNotionWebProvider) against Notion's
getSpaces endpoint, and a webSessionCredentials.ts UI entry for the
"Add session cookie" flow.
- Regenerated docs/reference/PROVIDER_REFERENCE.md and the
provider/translate-path golden snapshot (purely additive diffs); synced
the "251 providers" count across README/AGENTS/CLAUDE.md
(check:docs-counts STRICT gate).
Tests: tests/unit/executor-notion-web.test.ts (22 cases — registry
consistency, mocked-upstream request/response translation, NDJSON
snapshot parsing, cookie resolution, sanitized error paths) plus the
existing executor-web-cookie-sweep, provider-alias-uniqueness,
check-provider-consistency, web-session-credentials, and
provider-translate-path-golden suites all pass with notion-web included.
Adds felo-web, a free no-signup no-API-key chat/search-agent aggregator
(felo.ai), following the same architectural pattern as the existing
duckduckgo-web/blackbox-web "-web" scrape family:
- POST /api-proxy/main/search/threads opens a search thread and returns a
stream_key.
- GET /api/message/v1/stream/{stream_key} streams Felo's bespoke
data:{...}-line SSE, translated into OpenAI-compatible chunks.
- 5 models (felo-chat/search/scholar/social/document) map to Felo's
chat/google/scholar/social/document search categories.
Registered in providers.ts (noauth.ts, no-auth like duckduckgo-web),
providerRegistry.ts, and executors/index.ts. Free-tier catalog entries
added with tos: "avoid" (reverse-engineered endpoint, no published API —
same ToS posture as the other -web scrape providers).
No live network access was available in this environment to smoke-test
against the real felo.ai endpoint, so validation is TDD via mocked fetch
(tests/unit/felo-web-executor.test.ts): thread-creation payload shape,
SSE parsing (answer-snapshot diffing + final_contexts drop), streaming
and non-streaming response translation, and error/timeout paths that
route through sanitizeErrorMessage() per the error-sanitization rule.
Registers Rev AI as a 13th async-job STT provider, mirroring the
AssemblyAI/Kie.ai upload -> submit -> poll pattern already used by the
audio transcription handler:
- audioRegistry.ts: new "rev-ai" entry (bearer auth, async: true,
format: "rev-ai") with machine/low_cost/fusion transcriber models.
- audioTranscription.ts: handleRevAiTranscription() submits the job
with the media file inline in the multipart body (field "media"),
polls GET /jobs/{id} until "transcribed"/"failed", then fetches the
plain-text transcript. buildMultipartBody() gained an optional
fileFieldName param (default "file") so Rev AI's "media" field name
doesn't require a bespoke multipart builder. Errors route through
the existing upstreamErrorResponse()/errorResponse() helpers.
- providers/audio.ts: catalog entry (id/alias/name/icon/color/website)
for the dashboard connection UI.
- validation/audioMiscProviders.ts + validation.ts: validateRevAiProvider
wired into the provider "Test Connection" dispatcher.
Streaming (WebSocket) STT is scoped as a follow-up per the analyzed
plan — no existing precedent to extend, needs its own design pass.
Closes#6655.
The Discord invite (discord.gg/EkzRkpzKYt) was returning "Invalid Invite";
replace it with the new permanent invite across README + docs + zh-CN/zh-TW
i18n mirrors, and refresh the WhatsApp Brasil group link.
Reported-by: WhatsApp community (support mesh)
Rebuilt onto release/v3.8.49 feature-only: the branch's original file-size
decomposition of CostOverviewTab collided with the release's own component
extraction (#7272 TopListCard). Kept just the 180D/365D range delta —
CostRange/COST_RANGE_VALUES, the RANGE_OPTIONS selector, the getRangeStartIso
handlers in the analytics and requests-by-provider-date usage routes, and the
range180d/range365d i18n labels across all locales.
Inspired-by: 9router#2361
Codex's rolling "session" quota window only starts counting down once a
request lands inside it, so an idle connection's window keeps sliding
forward and the first real request after a long idle period pays for the
whole warm-up latency. This adds a strictly opt-in, per-connection
scheduler (default OFF) that watches an enabled connection's reported
resetAt and, once it slides forward, fires one tiny non-billed-model
request through the real Codex executor to keep the window warm.
- New in-process scheduler (src/lib/services/quotaAutoPing.ts), fully
dependency-injected (settings, DB, credential refresh, usage fetch,
executor, circuit breaker) and clock-injectable for deterministic tests.
Reimplemented in TS from the shipped 9router
src/shared/services/quotaAutoPing.js (Codex half only).
- Migration 123: last_ping_at / last_pinged_reset_key on
provider_connections, so the scheduler never re-pings the same reset
window twice.
- Settings: codexAutoPing.connections map, default {} (nobody opted in),
validated by the shared Zod schema.
- Respects the existing resilience layers: skips a connection whose
provider circuit breaker is open or whose rateLimitedUntil cooldown is
active, and applies its own 15-minute failure cooldown after a failed
ping.
- UI: new per-connection toggle (Settings -> AI -> Codex Quota Auto-Ping)
with an explicit "consumes real quota" tooltip, wired to the settings
PATCH route; i18n keys added to all 43 locales (EN authored, others
filled from the EN fallback pending real translation).
- 15 new deterministic unit tests covering enable/disable, first-reset
observation (cache-only, no ping), reset-slide ping, stable-reset
no-op, min-ping-interval, same-resetKey dedupe, session/weekly quota
exhaustion, non-OAuth skip, circuit-breaker-open skip, cooldown skip,
failure-cooldown skip, failed-ping bookkeeping, the real executor call
shape, and credential-refresh failure handling.
Antigravity's 2-bucket auto-ping is out of scope for this PR (no upstream
reference exists for its reset shape) and is tracked as a follow-up.
Ref #6977 (backend + settings + minimal UI toggle; Antigravity follow-up
tracked separately — not closing the issue from this PR)
* fix(providers): cap grok-cli tools at 200 for cli-chat-proxy
xAI's cli-chat-proxy enforces a hard limit of 200 tools per request and
returns a 400 above that ceiling. A client fanning a large MCP toolset
through Grok Build/Composer (e.g. Claude Code with many registered
tools) can exceed it. transformRequest() now caps the tools array
defensively before forwarding, and the grok-cli registry entries are
annotated supportsReasoning:false to document the existing (already
unconditional) reasoning_effort/reasoning strip for these two models.
Co-authored-by: Joseph Yaksich <294273268+gitcommit90@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2534
* chore(changelog): fragment for #6986
---------
Co-authored-by: Joseph Yaksich <294273268+gitcommit90@users.noreply.github.com>
Newly-added provider connections default testStatus to null until an operator explicitly runs a connection test. The combo builder's active-providers filter only kept testStatus === active/success, so a freshly-added custom provider was excluded from activeProviders — ModelSelectModal's loadCustomProviderModels() effect never fired for it, and its models never populated the combo model picker.
Extracted the eligibility check into isEligibleActiveConnection (src/lib/combos/builderDraft.ts), treating a never-tested connection the same as a known-good one (consistent with deriveConnectionStatus in builderOptions.ts, which only flags error on an explicit error/fail testStatus).
Reported-by: fajarbossit (https://github.com/decolua/9router/issues/2057)
gemini-cli (and any @google/genai-based client) sends its credential
exclusively via x-goog-api-key and it is not client-configurable to use
Authorization/x-api-key instead. Add it as an unconditional fallback,
after Authorization: Bearer and x-api-key, before the path-scoped URL
token, in both the real enforcement gate
(src/server/authz/policies/clientApi.ts::extractBearer()) and the
general extractor (src/sse/services/auth.ts::extractApiKey()).
The header-read/trim logic is extracted into a new leaf module
(src/sse/services/googApiKeyAuth.ts) shared by both call sites, so the
frozen auth.ts file only takes the minimal chokepoint wiring
(config/quality/file-size-baseline.json rebaselined 2458->2461 with
justification, matching this repo's established extraction pattern).
Closes#7034
* fix(openai): strip reasoning_effort when GPT-5.x tools present (port from 9router#2540)
Raw api.openai.com Chat Completions rejects GPT-5.x reasoning models that carry both function tools and an active reasoning_effort with HTTP 400 ("Function tools with reasoning_effort are not supported ... Please use /v1/responses instead"). The existing forceResponsesUpstream guard only reroutes openai-compatible-* connections carrying MCP/tool_search tool shapes; the plain openai provider had no equivalent guard, so gpt-5.x models used with a coding client (function tools + any explicit reasoning effort) still hit the upstream 400. Add stripGpt5ReasoningWhenTools() (gpt5SamplingGuard.ts), wired into chatCore.ts alongside the existing sampling guard, to drop reasoning_effort/reasoning when function tools are present and reasoning is active, letting the request succeed on /v1/chat/completions.
Reported-by: Tech Solution (@techsolutionmta) (https://github.com/decolua/9router/issues/2540)
* fix(openai): scope reasoning-strip guard to /chat/completions only
stripGpt5ReasoningWhenTools gated on provider+model-name alone, so once
#7242 routes the public GPT-5.6 family to /v1/responses (targetFormat
"openai-responses", which natively supports tools + reasoning), the two
PRs would compose into the worst of both worlds: routed to the endpoint
that supports reasoning, but reasoning stripped anyway. Pass the
request's already-resolved targetFormat into the guard and skip the
strip whenever it is not going out over /chat/completions, so the
guard tracks the actual upstream surface instead of a model-name list
that would need updating for every future GPT-5.x family.
Reported-by: Tech Solution (@techsolutionmta) (https://github.com/decolua/9router/issues/2540)
* fix(providers): honor configured proxy on Grok Build egress
The grok-cli executor reaches Grok Build over raw `https.request()` (forced
IPv4, to dodge Cloudflare blocking on the direct path) rather than the
process-wide patched `fetch()` that every other executor uses. `https.request()`
never consults the proxy AsyncLocalStorage context, so the proxy the caller
already pinned upstream in chatHelpers.ts (`runWithProxyContext`) was silently
ignored on BOTH grok-cli paths: chat inference (`nativePost`) and OAuth token
refresh (`nativeHttpsPost`, POST https://auth.x.ai/oauth2/token).
User-visible effect: an operator who assigns a proxy to a Grok Build connection
(or provider/global scope) still egresses on the host's real IP — an IP leak
that defeats account-isolation/anonymity setups, and breaks Grok Build entirely
for operators who must egress through a proxy.
Fix is delta-only: `resolveGrokRequestDispatch()` reads the already-resolved
proxy via the shared `resolveProxyForRequest()` and returns either an
HttpsProxyAgent bound to it, or — when no proxy is configured — the existing
forced-IPv4 direct options, unchanged. Only HTTP/HTTPS CONNECT proxies are
supported on this path; an explicitly configured proxy of another kind (SOCKS5)
fails closed rather than silently leaking direct, matching the fail-closed
convention for OAuth/account proxies (#3051). The proxy URL is never logged, so
proxy credentials cannot leak into logs.
Regression test: tests/unit/grok-cli-proxy-selection.test.ts (RED before the fix
— `resolveGrokRequestDispatch` did not exist and both request builders hardcoded
`family: 4` with no agent; GREEN after).
Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2343
* chore(changelog): fragment for #7244
---------
Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>
* feat(sse): add native xAI Grok Imagine video generation provider
OmniRoute's /v1/videos surface already supported 10 provider formats
(vertex-veo, google-flow, comfyui, sdwebui-video, kie-video, runwayml,
haiper-video, veoaifree-web, leonardo-video, dashscope-video), but xAI
had no native entry — Grok Imagine was only reachable indirectly through
the kie proxy market (kie's "grok-imagine/text-to-video" models), which
requires a separate kie.ai account and bills through kie.
This registers xai as a first-class video provider that talks to
api.x.ai/v1/videos directly, reusing the stored xai Bearer apiKey that
the existing image-generation "xai" entry in imageRegistry.ts already
uses — no new credential flow. The new xai-video handler format mirrors
the DashScope create+poll shape, adapted to xAI's request_id / status
("pending" | "processing" | "done" | "failed") job model.
User-visible effect: `xai/grok-imagine-video` works on
POST /v1/videos/generations against a user's own xAI key.
Co-authored-by: ann <daohuyentfqn2l@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/2593
* chore(changelog): fragment for #7238
* fix(sse): extract xAI Grok Imagine video handler to fix file-size ratchet
videoGeneration.ts grew to 1407 lines (frozen cap 1265) after adding the
Grok Imagine handler. Extract handleXaiVideoGeneration into a co-located
module (open-sse/handlers/videoGeneration/xaiGrokImagineHandler.ts),
following the googleFlowHandler.ts precedent — same pattern already used
for the Google Flow video handler. File now sits at 1261 lines, under cap.
No behavior change; existing tests (video-xai-grok-imagine.test.ts) cover
the handler through the public handleVideoGeneration() entry point and
pass unmodified.
* refactor(sse): decompose xAI Grok Imagine handler to fix complexity ratchets
The file-size red was masking two ratchet regressions (the gate aborts on
the first failure): complexity 2058 > 2056 and cognitive 891 > 890. Both
came from the PR's own handleXaiVideoGeneration — a single 107-line
function with complexity 37 / cognitive 24, tripping `complexity`,
`max-lines-per-function` (2 complexity-ratchet violations) and
`sonarjs/cognitive-complexity` (1 cognitive violation).
Decompose it into four cohesive units instead of rebaselining:
- resolveXaiVideoOptions() — timeouts/credential/endpoints/prompt
- buildXaiVideoPayload() — OmniRoute body -> xAI create payload
- createXaiVideoJob() — create-job POST -> request_id | error
- pollXaiVideoJob() — poll loop -> terminal outcome
- buildXaiVideoResponse() — outcome -> OpenAI-like response
Both ratchets now sit exactly at baseline (complexity 2056, cognitive 890)
and file-size stays under cap. pollXaiVideoJob reads Date.now() only in the
loop condition, so the caller keeps its timeout budget semantics. No
behavior change; the 7 existing tests pass unmodified.
---------
Co-authored-by: ann <daohuyentfqn2l@gmail.com>
* feat(providers): add Mixedbread AI as embeddings provider (#6660)
Registers Mixedbread AI (https://api.mixedbread.com) in the
EMBEDDING_PROVIDERS registry alongside the other bearer-auth embedding
providers (Voyage AI, Jina AI, Nomic, ...): OpenAI-compatible
/v1/embeddings endpoint, exposing mxbai-embed-large-v1 and
mxbai-embed-2d-large-v1 (both 1024d, Matryoshka). Adds a matching
provider metadata entry (icon/color/authHint/free-tier note) modeled
on the nomic block, regenerates docs/reference/PROVIDER_REFERENCE.md,
and syncs the 250->251 provider-count mentions in README/AGENTS/CLAUDE
required by the strict docs-counts gate.
No executor/translator changes needed — the embeddings handler is a
generic pass-through with no provider-specific branching.
* test(providers): align APIKEY_PROVIDERS count 167→168 for the new 6660 provider (#6660)
Adding the mixedbread embeddings provider to specialty-media.ts grows APIKEY_PROVIDERS
by one; providers-constants-split.test.ts hardcodes the family-partition
total. Legitimate count alignment (the code genuinely added a provider),
not a weakened assertion — all 4 partition/dedup checks still enforced.
* fix(sse): stop dropping tool_search and stop leaking OpenAI-only params in Responses->Chat translation (#7532, #7533)
#7532: `openai-responses.ts` unconditionally dropped `tool_search` when
downgrading a Responses-shaped request to Chat Completions, hiding the tool
from the model and breaking Codex's deferred/lazy tool-discovery protocol for
any provider that gets downgraded (e.g. built-in providers like opencode-go).
tool_search carries `execution: "client"` — the client resolves the call
locally regardless of wire shape — so it is now mapped to a normal Chat
function tool, mirroring the existing local_shell -> shell pattern in the
same file, instead of being silently discarded.
#7533: the same translator unconditionally copied two GPT-5/OpenAI-only
fields (`verbosity`, `prompt_cache_key`) into the translated Chat body
regardless of destination provider. A strict-protocol non-OpenAI upstream
(NVIDIA confirmed by the reporter) 400s on unrecognized top-level parameters.
Both fields are now gated on `credentials.provider === "openai"`, stripped
otherwise; the existing OpenAI-destined behavior (needed for #517's
prompt-caching fix) is preserved byte-identical via a dedicated sanity test.
Regression tests: tests/unit/tool-search-filtered-responses-to-chat-7532.test.ts,
tests/unit/verbosity-prompt-cache-key-provider-gate-7533.test.ts. Two existing
tests that encoded the old buggy contract (unconditional tool_search drop /
unconditional field leak with no credentials) were aligned to the corrected
contract: tests/unit/translator-openai-responses-req.test.ts,
tests/unit/openai-responses-verbosity.test.ts.
Gates run green: file-size, complexity, cognitive-complexity, typecheck:core,
lint (scoped to changed files), and the full touched-area unit test suite
(329 tests, 0 failures).
* fix(sse): keep prompt_cache_key/verbosity for the codex destination (#7533)
The #7533 provider gate allowlisted only "openai", but /v1/responses routes
EVERY request through this downgrade (handleResponsesCore ->
convertResponsesApiFormat) regardless of provider, and codex is an
OpenAI-operated upstream (chatgpt.com/backend-api/codex). Gating it out
stripped prompt_cache_key for Codex and silently re-broke the prompt-cache
affinity #517 exists to protect — with no test covering it.
Allowlist is now {openai, codex} and carries two #517 regression guards.
Non-OpenAI upstreams (NVIDIA) still get both fields stripped, per #7533.
Non-stream Codex (ChatGPT account) chat 502'd with "Response body is already
used". On the wreq-js TLS-fingerprint transport the Response is backed by a
native body handle, and merely accessing response.body disturbs it so a later
.text() throws. The Codex non-stream upstream response has an empty content-type,
so peekCodexSseTransientError early-returns — but its guard evaluated
!response.body (touching .body) before the content-type check, consuming the
body; chatCore's readNonStreamingResponseBody then re-read it and 502'd.
Streaming was unaffected. Reorder the guard to check content-type first.
Validated live on the VPS (192.168.0.15): codex/gpt-5.5 and codex/gpt-5.6-terra
non-stream now return 200; streaming still works. Regression test drives the real
peek with a destructive-.body mock.
CodeQL js/stack-trace-exposure flags ANY error-derived value returned in
the mock route bridge's 500 path, not just error.stack — swapping .stack
for error.message (in #7354, alert #736) left sibling alert #737 open on
the same line. Replace the body with a static string; the test only
asserts status===200, so the 500 body is never inspected. Clears the last
open CodeQL alert repo-wide, unblocking the Quality Ratchet on every PR.
* docs(troubleshooting): document Avast/AVG README.md false positive (#5946)
Avast/AVG quarantine the packaged README.md with MD:HttpRequest-inf[Susp] --
a heuristic false positive on the ~15 http://localhost:20128 examples the file
carries (README ships via package.json -> files, landing at
node_modules/omniroute/README.md).
Adds a Troubleshooting section explaining the detection is benign, how to stop
the notifications (AV exclusion), how to report the false positive upstream,
and why we do not mangle the localhost examples to dodge one vendor heuristic.
Documentation only -- no functional change.
Reported-by: DemonNCoding
* docs(changelog): add fragment for #7295
* fix(sse): silence noisy proxy-failure log on caller-initiated abort
The pinned-proxy dispatch path in proxyFetch.ts logged every failure —
including a plain caller abort or the caller's own AbortSignal timeout
firing — as "[ProxyFetch] Proxy request failed ... fail-closed". A
client cancelling its own request is not a proxy transport failure and
shouldn't be misreported as one in ops logs/alerting; it still
propagates to the caller unchanged (fail-closed behavior is untouched).
Co-authored-by: TuyulSpam <287281626+TuyulSpam@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2589
* chore(changelog): fragment for #7266
---------
Co-authored-by: TuyulSpam <287281626+TuyulSpam@users.noreply.github.com>
* feat(db): include xp_audit_log in automatic retention/prune (#6801)
* fix(i18n): mirror retentionXpAuditLog into pt-BR.json (#6801)
pt.json and pt-BR.json are distinct files; the key landed only in pt.json,
so the i18n pt-BR integrity test (no drift, #6695) went red.
Fusion's panel-model extraction in combo.ts only recognized plain string
or {model: string} entries in combo.models; a {kind:"combo-ref", comboName}
step (a first-class, Zod-validated combo-step shape the dashboard already
lets you add to a fusion panel) had neither field, so it was silently
filtered out — no error, no warning, and an opaque 400 if it was the only
panel member.
A combo-ref panel member is now dispatched as one black-box panel voice
(a recursive handleComboChat call into the referenced combo, reusing the
same executeComboRefUnit + cycle/depth guards every other combo-ref-
consuming strategy already uses), not a fan-out of the referenced combo's
own targets.
New module open-sse/services/combo/fusionPanel.ts keeps the frozen
combo.ts god-file's growth minimal (extraction/dispatch-wrapper logic
lives there; the fusion branch itself only wires it in).
Add a per-connection cache capability override (supportsPromptCaching,
cacheControlPassthrough) stored in provider_specific_data.cache, consulted
first by providerSupportsCaching() / providerHonorsOpenAIFormatCacheControl()
before falling back to the hardcoded CACHING_PROVIDERS name sets. Unblocks
prompt_cache_key injection, the compression cache-aware guard, and
cache_control passthrough for openai-compatible-chat-<uuid>-style custom
connections that can never match the hardcoded provider-name sets. Default
(no override) is byte-identical to current behavior.