Files
OmniRoute/open-sse/handlers/imageGeneration.ts
Diego Rodrigues de Sa e Souza 1bda6c15dc Release v3.8.44 (#5925)
* fix(install): add pnpm-workspace.yaml allowBuilds + pnpm.json for pnpm 11+

pnpm 11 introduced ERR_PNPM_IGNORED_BUILDS for native addon packages.
Without explicit allowBuilds approval, these packages silently skip build scripts
and OmniRoute fails to start with missing native modules.

Changes:
- pnpm-workspace.yaml: Set allowBuilds=true for all 13 native addon packages
  (@parcel/watcher, @swc/core, better-sqlite3, core-js, esbuild, keytar, koffi,
  libxmljs2, onnxruntime-node, protobufjs, sharp, tls-client-node, unrs-resolver)
- pnpm.json: Migrate onlyBuiltDependencies from package.json (deprecated field)
  to the new pnpm.json config file per pnpm 11 spec.

Tested on: pnpm 11.9.0, Node 24, Windows 11.

Fixes: pnpm install ERR_PNPM_IGNORED_BUILDS on fresh clone with pnpm 11.

* chore(release): open v3.8.44 development cycle

* test(security): parse Kimi Web URL host instead of substring match (CodeQL #689) (#5928)

Alert js/incomplete-url-substring-sanitization: the Kimi Web executor
test asserted result.url.includes("www.kimi.com"), which a hostile host
like www.kimi.com.evil.net would also satisfy. Parse the URL and assert
on the exact hostname (new URL(result.url).hostname === "www.kimi.com"),
which is both a stronger check and clears the CodeQL warning.

* refactor(translator): extract thinking-budget fitting from openai-to-claude (#5932)

Extract the thinking-budget fitting cluster (fitThinkingToMaxTokens +
private safeCapMaxOutputTokens + MIN_* constants) verbatim into the pure
leaf openai-to-claude/thinkingBudget.ts. Host re-exports fitThinkingToMaxTokens
so external importers keep working and imports it back for internal use.

Host 822 -> 738 LOC (under the 800 cap). No behavior change: byte-identical
bodies, public export set unchanged. Adds a split-guard test; all consumer
tests stay green (translator-openai-to-claude, strip-empty, minimax-m3, passthrough).

* chore(release): pipeline hardening — test-masking pre-flight gate + contributors/uncovered helpers (#5926)

* chore(ci): add test-masking PR-context gate to release-green pre-flight

Reproduce check:test-masking (vs origin/main) inside validate-release-green so
non-allowlisted net-assert reductions surface in the local pre-flight instead of
in a ~40-min CI layer on the release PR. run() now merges a per-gate opts.env so
GITHUB_BASE_REF reaches the child. HARD gate; skipped under --quick.

Context: v3.8.43 release cost 3 CI round-trips for PR-context gates (test-masking,
file-size, pr-evidence) that check:release-green did not reproduce locally.

* chore(release): add contributors generator + uncovered-commit reconciliation helpers

- scripts/release/gen-contributors.mjs: reproducible `### 🙌 Contributors` table for a
  CHANGELOG version (parenthetical-group parser → accurate per-PR attribution, noise-handle
  denylist). v3.8.43 shipped without the section (a real miss) because it was hand-built.
  npm run release:contributors <version> [--inject].
- scripts/release/list-uncovered-commits.mjs: lists commits since the last tag with no
  CHANGELOG bullet (v3.8.43 had 123/176 uncovered at reconciliation start). Advisory,
  maintainer-side. npm run release:uncovered.
- 20 unit tests (parenthetical attribution, noise exclusion, idempotent injection, coverage window).

* chore(quality): absorb web-cookie-providers-new file-size drift from #5928 (base-red on release/v3.8.44)

* refactor(translator): split openai-responses request translator into pure leaves (#5940)

Extract the shared pure primitives and the chat->Responses direction out of the
894-line openai-responses.ts request translator:
- openai-responses/helpers.ts: pure primitives (toRecord/toString/clampCallId/
  normalizeVerbosity/etc + markers/regexes/JsonRecord), zero host imports
- openai-responses/toResponses.ts: openaiToOpenAIResponsesRequest (chat->Responses),
  imports the helpers leaf

Host keeps openaiResponsesToOpenAIRequest (Responses->chat, imported by production)
plus both register() directions, and re-exports openaiToOpenAIResponsesRequest so
external importers (tests) keep working.

Host 894 -> 529 LOC (under the 800 cap). Verbatim bodies (multiset check: leaf A 54/54,
leaf B 294 lines, fn1 intact), public export set unchanged, leaves never import the host
(no cycle). Adds a split-guard test; all consumer tests stay green (responses-translation-fixes
37, verbosity 4, reasoning-effort 4, orphaned-tool-filter 8, empty-tool-name-loop 8,
headroom-responses-format 3).

* chore(ci): pr-evidence FAIL output tells you to push (body edit does not re-run the gate) (#5944)

ci.yml ignores the 'edited' event, so adding the Evidence block to the PR body after a
push does not re-run check:pr-evidence — you need another commit. The FAIL report now
says so, at the exact place someone sees the red check. + 5 unit tests (classification +
hint-on-fail / no-hint-on-pass). Decided against a separate edited-triggered workflow:
pr-evidence is not a required check (no ruleset gates it; release PRs merge UNSTABLE, not
BLOCKED), so the gap is cosmetic and the generate-release skill already puts Evidence in
the body before the first push.

* fix(providers): Perplexity Web emits real tool_calls in streaming mode (mirror chatgpt-web toolMode) (#5927) (#5937)

Perplexity Web (Pro/Max) only converted <tool>{...}</tool> text into
OpenAI tool_calls for non-streaming requests (hasTools && !stream).
Streaming requests -- the default for agentic coding clients -- got
the raw <tool> text as plain delta.content and never emitted a
tool_calls SSE delta, so clients could not execute tools.

Reuses the provider-agnostic buildToolModeResponse()/
toolCompletionToSseStream() helpers already shipped for chatgpt-web
(#5240): when tools are requested, buffer the full completion and
convert it into either a JSON completion or a terminal SSE replay
carrying delta.tool_calls + finish_reason: tool_calls, regardless of
the caller's stream flag. Extended buildToolModeResponse()'s idSeed
to be caller-supplied (default 'cgpt', perplexity-web passes 'pplx')
so tool_call ids stay provider-specific without duplicating the
helper. Non-tool streaming is unchanged (still lives token-by-token
via buildStreamingResponse).

* fix(discovery): resolve duplicate /v1 paths and redirect aborts (#5904)

Integrated into release/v3.8.44. Thanks @hamsa0x7 for diagnosing the doubled /v1 discovery path and the REDIRECT_BLOCKED probe-loop abort (#5899). De-scoped to the discovery fix (the #5903 session-affinity work is handled by #5943) and added Rule #18 regression guards.

* docs(changelog): record #5926 + #5944 (release-pipeline hardening) under v3.8.44 Maintenance (#5952)

* docs(claude): add Hard Rule #22 — cross-session safety (git stash + in-flight PRs) (#5955)

Integrated into release/v3.8.44 — Hard Rule #22 (cross-session safety).

* refactor(translator): extract pure helpers from response/openai-responses (#5949)

Extract the 5 stateless helpers (normalizeToolName, stripEmptyOptionalToolArgs,
normalizeOutputIndex, normalizeUpstreamFailure, extractResponsesReasoningSummaryText)
verbatim into the pure leaf openai-responses/pureHelpers.ts (no stream state, no host
import). Host imports them back and re-exports normalizeUpstreamFailure for external
importers (tests).

Host 1091 -> 1001 LOC. The stateful streaming core stays in the host (out of scope).
Byte-identical bodies (multiset 73/73), no cycle. Adds a split-guard; consumer tests
stay green (responses-translation-fixes 37, combo-param-validation-fallback-4519 5).

* docs(compression): document upstream sync policy for RTK/Caveman engines (#5830) (#5948)

Integrated into release/v3.8.44 — docs-only upstream sync policy for RTK/Caveman engines (closes #5830). All 7 checks green.

* fix(sse): strip ANSI/VT100 codes from gemini-cli stream frames (#5934)

Integrated into release/v3.8.44 — ReDoS-safe ANSI/VT100 strip for gemini-cli stream frames (port of upstream #2273, thanks @anki1kr). PR test green (5/5), file-size gate OK.

* fix(translator): strict Anthropic content-block compliance in antigravity→openai request (#5935)

Integrated into release/v3.8.44 — strict Anthropic content-block compliance in antigravity→openai (port upstream #2296). PR test green (9/9). UNSTABLE red is the pre-existing environmental setup-claude base-red (opencode-plugin dist not built in fast-path), not a regression from this PR.

* fix(mcp): auto-recover stale streamable HTTP sessions on initialize (#5957)

Integrated into release/v3.8.44 — MCP stale streamable-HTTP session auto-recovery (thanks @Chewji9875).

* fix(providers): validate v0 Platform API keys via chats endpoint (#5954)

Integrated into release/v3.8.44 — v0-vercel Platform API key validation (thanks @vittoroliveira-dev).

* fix(api): relax provider-scoped chat completion validation (#5907)

Integrated into release/v3.8.44 — relaxed provider-scoped chat validation + regression test (thanks @nickwizard).

* fix(providers): strip /v1 unconditionally to avoid /v1/v1/models fetch error (#5899) (#5920)

Integrated into release/v3.8.44 — unconditional /v1 strip in both models-discovery paths + regression test (thanks @anki1kr).

* fix(resilience): per-window is_exhausted + honor quota-exhaustion preflight for priority combos (#5923) (#5941)

Integrated into release/v3.8.44.

* fix(resilience): honor active codex session affinity over per-request reset-aware re-scoring (#5903) (#5943)

Integrated into release/v3.8.44.

* fix(thinking): only inject redacted_thinking replay block when tool_use present and thinking enabled (#5945) (#5953)

Integrated into release/v3.8.44.

* feat(providers): add ClinePass API-key provider (#5942)

Integrated into release/v3.8.44 — ClinePass API-key (BYOK) provider (port upstream 9router#2304, co-authored @adentdk). Validated locally: 16 clinepass tests green; fixed the APIKEY count 158→159 + translate-path golden snapshot (clinepass is a genuine new provider). Remaining UNSTABLE red is the pre-existing environmental setup-claude base-red (opencode-plugin dist not built in fast-path). Supersedes stub #5541.

* feat(api): add /v1/ocr endpoint (Mistral OCR) + Mistral moderation (#5950)

Integrated into release/v3.8.44 — /v1/ocr endpoint (Mistral OCR) + Mistral moderation (port upstream 9router#2064, co-authored @waguriagentic). Validated locally: 14 ocr-route tests + moderation/servicekind/endpoint-category suites green (CORS→Zod→handler + no-stack-leak assertion). Reds are inherited DRIFT only: cognitive-complexity ratchet (none from OCR files — pre-existing cycle drift, rebaselined at release) + environmental setup-claude base-red.

* fix(codex): convert chat json schema to responses text format (#5933)

Integrated into release/v3.8.44 — converts Chat Completions json_schema response_format → Responses API text.format on the Codex path, and preserves existing text.format through verbosity normalization. Base redirected main→release; the openai-responses.ts split that landed this cycle was reconciled by re-applying the delta onto openai-responses/toResponses.ts. Validated locally: 48 translator-openai-responses-req + 8 codex-verbosity tests green.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(providers): add Claude Sonnet 5 support across the model pipeline (#5833)

Integrated into release/v3.8.44 — wires claude-sonnet-5 end-to-end (registries, modelSpecs, pricing ×3, cost, Sonnet-family fallback, 1M-ctx, static models). Reconciled the add/add overlap with the already-merged #5796 (kept the PR's superset test with the family-fallback assertion). Validated locally: kiro-sonnet-5 + catalog + pricing/modelSpecs/fallback suites all green. Thanks @ggiak!

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(relay): gate bifrost auto routing by provider manifest (#5870)

Integrated into release/v3.8.44 — gates Bifrost auto-routing by the provider plugin manifest (only manifest-eligible providers reach the sidecar; ineligible/unknown fall back to the TS path with explicit reasons). Superset of #5869 (carries the full manifest + registry + docs). Resolved an integration-test conflict in favor of the release (which already subsumes this PR's readiness/removeDirWithRetry improvements). Validated locally: 4 provider-plugin-manifest + 11 relay-routing-backend tests green. Thanks @KooshaPari!

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* refactor(translator): extract pure message helpers from openai-to-kiro (852→751) (#5947)

* refactor(translator): extract pure message helpers from openai-to-kiro

Extract the pure tool/message helpers (parseToolInput, normalizeKiroToolSchema,
serializeToolResultContent) verbatim into the leaf openai-to-kiro/messageHelpers.ts.
The host imports them back for convertMessages. They were module-private, so the
public export set is unchanged (no re-export needed).

Host 852 -> 751 LOC. Byte-identical bodies (multiset 99/99), leaf has zero imports
(no cycle). Adds a split-guard; consumer tests stay green (translator-openai-to-kiro 33,
translator-ai-sdk-image-parts 3).

* chore: re-trigger CI (stuck runner on 2/2 shard)

* refactor(executors): extract pure prompt + composer helpers from cursor (#5960)

Extract two pure clusters from the cursor executor into sibling leaves:
- cursor/prompt.ts: isRecordLike + toolChoiceDirectiveLine + buildCursorOutputConstraints
- cursor/composer.ts: composer thinking-as-content decoding (isComposerModel,
  visibleComposerContentFromThinking, composerReasoningRemainder + markers)

Host imports both back for internal use and re-exports the 3 composer helpers for
external importers (tests). Host 1576 -> 1451 LOC. Byte-identical bodies (verbatim
multiset prompt 65/65, composer 32/32), leaves have zero imports (no cycle). Adds a
split-guard; consumer tests stay green (cursor-composer-thinking, cursor-streaming,
cursor-agent-tool-calls, translator-openai-to-cursor, cursor-agent-system-prompt).

* refactor(executors): extract pure SSE-collect parsing from antigravity (#5962)

Extract the pure SSE-payload -> collected-stream parser (AntigravityCollectedStream,
stripZeroWidth, parseAntigravityTextualToolCall, addAntigravityTextualToolCall,
processAntigravitySSEPayload/Text, flushAntigravitySSEText) verbatim into the leaf
antigravity/sseCollect.ts. Host imports the helpers it uses and re-exports
processAntigravitySSEPayload for external importers (tests).

Host 1812 -> 1671 LOC. Byte-identical bodies (verbatim multiset 135/135), leaf does
not import the host (no cycle). Credit/quota state, auth, and HTTP dispatch untouched.
Adds a split-guard; consumer tests stay green (executor-agy 8, executor-antigravity 26,
antigravity-sse-collect-socket-release, copilot-agent-antigravity-parity 6).

* refactor(executors): extract pure model maps + resolvers from chatgpt-web (#5967)

Extract the static model maps (MODEL_MAP, MODEL_FORCED_EFFORT, THINKING_CAPABLE_SLUGS)
and the pure thinking-effort resolvers (isThinkingCapableModel, normalizeThinkingEffort,
resolveThinkingEffort, ResolvedChatGptModel, resolveChatGptModel) verbatim into the pure
leaf chatgpt-web/models.ts. Host imports the two resolvers it uses back.

Host 3205 -> 3076 LOC. Byte-identical bodies (verbatim multiset 120/120), leaf has zero
imports (no cycle). Auth/PoW/session/HTTP dispatch and all module caches untouched.
Adds a split-guard; consumer tests stay green (chatgpt-web 86, chatgpt-web-tools-5240 4,
chatgpt-web-sha3-boringssl-5531 5).

* refactor(executors): decompose grok-web into pure tool/markup leaves (#5994)

Extract the pure OpenAI<->Grok tool-translation, native-tool mapping, markup cleanup,
and NDJSON stream types out of the 1872-line grok-web executor into 4 sibling leaves:
- grok-web/types.ts: GrokStreamResponse/GrokStreamEvent (stream types)
- grok-web/tool-bridge.ts: OpenAI<->Grok tool translation + registry + classifiers
- grok-web/native-tools.ts: native-tool selection/scoring + native->OpenAI mapping
- grok-web/text-cleanup.ts: Grok markup stripping + GrokMarkupFilter

Layered, acyclic: types <- tool-bridge <- native-tools; text-cleanup <- types; host
imports the leaves. All symbols module-private (no host re-export). Host 1872 -> 887 LOC.
Byte-identical bodies (verbatim per-leaf), no cycle, all new leaves <= 800 cap
(tool-bridge split at line 753 to stay under). Auth/cookie/TLS/HTTP dispatch untouched.
Adds a split-guard; consumer tests stay green (grok-web 62, grok-cli-oauth 15,
grok-cli-strip-params 2).

* refactor(executors): extract pure quota parsing from codex (#5999)

Extract the pure Codex quota-snapshot parsing + reset/cooldown scheduling
(CodexQuotaSnapshot, parseCodexQuotaHeaders, getCodexResetTime,
getCodexDualWindowCooldownMs) verbatim into the leaf codex/quota.ts. Host re-exports
the 4 symbols so handlers/chatCore/codexQuota.ts + tests keep resolving.

Host 1539 -> 1427 LOC. Byte-identical bodies (verbatim 98/98), leaf has zero imports
(only Date, no cycle). WS transport, auth, HTTP dispatch untouched. Adds a split-guard;
consumer tests stay green (executor-codex 40, codex-quota-fetcher 7, chatcore-codex-quota 5).

* refactor(executors): extract pure stream formatters from deepseek-web (#6000)

Extract the pure content/citation formatters (isThinkingModel, isSearchModel,
cleanDeepSeekToken, formatStreamContent, DeepSeekSearchResult, appendSearchCitations)
verbatim into the leaf deepseek-web/stream-format.ts. Host imports the 5 it uses back
into transformSSE/collectSSEContent (cleanDeepSeekToken stays internal to the leaf).

Host 1147 -> 1108 LOC. Byte-identical bodies (verbatim 34/34), leaf has zero imports
(no cycle), all module-private (no re-export). PoW/auth/token-cache/HTTP dispatch
untouched. Adds a split-guard; consumer tests stay green (deepseek-web 35,
deepseek-web-rolling-window-2942 5, deepseek-web-tools-execute 3).

* refactor(api): add validatedJsonBody helper (salvage #5075) (#5931)

Fuses JSON body parsing + Zod validation into a single call that returns
either type-narrowed data or a ready-to-return 400 NextResponse with the
standard error envelope. Salvaged as the Tier 1 portable helper from the
closed refactor PR #5075; the bulk route migration is intentionally not
ported. Adds a focused 6-case regression test.

Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>

* feat(qoder): drive PAT auth via qodercli, add dashboard quota, fix connection display (#5816)

Integrated into release/v3.8.44 — Qoder PAT auth via qodercli binary + dashboard quota + dual-auth connection fix. Thanks @AgentKiller45 (co-author @judy459)!

Validated locally (release-green on its own merits): lint 0, typecheck:core 0, 104 qoder/usage/UI tests green, file-size gate OK (owner-approved qoderCli.ts baseline-freeze 666→989), env-doc-sync fixed (documented QODER_CLI_CONFIG_DIR).

The 2 remaining CI reds are INHERITED base-reds, not caused by this PR: (1) LEDGER-4 minimax-m3 supportsVision (minimax-m3 base + cline-pass/minimax-m3 from the already-merged #5942); (2) mutation-test-coverage missing 3 tests in stryker.conf (#5903/#5942/#5923). Both cleaned up separately.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(providers): minimax-m3 supportsVision (LEDGER-4) + stryker tap.testFiles drift (#6012)

Release-green cleanup — clears LEDGER-4 minimax-m3 supportsVision + stryker tap.testFiles drift base-reds. Validated locally.

* fix(registry): flag cline-pass/minimax-m3 as multimodal (supportsVision) (#6003)

The cline-pass provider's minimax-m3 entry was missing supportsVision, breaking the
LEDGER-4 registry-consistency test (all minimax-m3 entries must set supportsVision to
match lite.ts — minimax-m3 is multimodal). Every other minimax-m3 registry entry
(trae, bazaarlink, cline, ollama-cloud, ...) already sets it. This was a base-red on
release/v3.8.44 inherited by every open PR.

Validated by the existing failing-then-passing guard tests/unit/review-reviews-v3814-fixes.test.ts
(LEDGER-4).

* refactor(executors): extract pure payload construction from claude-web (#6006)

Extract the pure Claude-web payload types + transforms + default tools/style
(ClaudeWebRequestPayload, ClaudeWebStreamChunk, DEFAULT_CLAUDE_MODEL,
generateMessageUUIDs, getDefaultTools, getDefaultPersonalizedStyle, transformToClaude,
transformFromClaude) verbatim into the leaf claude-web/payload.ts. Host imports the 3
it uses back (ClaudeWebRequestPayload type + the two transforms).

Host 1056 -> 835 LOC. Byte-identical bodies (verbatim 149/149), leaf imports only
randomUUID (no host import, no cycle), all module-private (no re-export). Cookie/auth/
Turnstile/TLS/HTTP dispatch untouched. Adds a split-guard; consumer tests stay green
(claude-web 13, claude-web-auto-refresh 6).

* refactor(executors): extract pure upstream-header helpers from base (#6008)

Extract the pure upstream-header helpers (mergeUpstreamExtraHeaders, getCustomUserAgent,
setUserAgentHeader, applyConfiguredUserAgent, isOpenAICompatibleEndpoint,
stripStainlessHeadersForOpenAICompat) verbatim into the leaf base/headers.ts. base.ts is
imported by ~18 executors, so the host re-exports all 6 to keep those import paths intact;
it also imports the 4 it uses internally in the BaseExecutor class. The trivial JsonRecord
type alias is redefined locally in the leaf to avoid a base<->leaf cycle.

Host 1539 -> 1451 LOC. Byte-identical bodies (verbatim 78/78), leaf does not import the
host (no cycle). typecheck:core validates all base importers still resolve via the
re-export. Adds a split-guard; consumer tests stay green (executor-base-utils 22,
executor-default-base 49, executor-strip-stainless-openai-compat 6, plus executor sanity
via typecheck).

* refactor(executors): extract pure wire protocol from perplexity-web (#6014)

Extract the pure Perplexity wire protocol (consts, SSE stream types, SSE parsing,
OpenAI<->Perplexity message translation, request/query builders, content extraction,
sseChunk) verbatim into the leaf perplexity-web/protocol.ts. Host imports back the 10
symbols it uses; everything module-private (no re-export). Session cache, TLS fetch,
auth, and the executor class stay in the host.

Host 1028 -> 534 LOC. Byte-identical bodies (verbatim), leaf imports only randomUUID
(no host import, no cycle). Adds a split-guard; consumer tests stay green
(perplexity-web 26, streaming-tools-5927 2, tls-client 6, key-validation-models 2).

* refactor(executors): extract pure URL normalizers from default (#6015)

Extract the pure per-provider chat-URL normalizers (normalizeBailianMessagesUrl,
normalizeDataRobotChatUrl, normalizeAzureAiChatUrl, normalizeWatsonxChatUrl,
normalizeOciChatUrl, normalizeSapChatUrl, normalizeXiaomiMimoChatUrl,
normalizeOpenAIChatUrl, getOpenRouterConnectionPreset) verbatim into the leaf
default/urlNormalizers.ts. Host imports them back into buildUrl/transformRequest; the
now-dead build*ChatUrl/normalizeBaseUrl imports move to the leaf. All module-private
(no re-export).

Host 864 -> 815 LOC (shrunk below its frozen baseline). Byte-identical bodies (verbatim
45/45), leaf does not import the host (no cycle). buildHeaders/execute/auth untouched.
Adds a split-guard; consumer tests stay green (executor-default-base 49,
anthropic-compatible-bearer 3, strip-client-metadata 3).

* feat(webfetch): support self-hosted FireCrawl instances (#5793)

Integrated into release/v3.8.44 — self-hosted FireCrawl support (FIRECRAWL_BASE_URL/FIRECRAWL_TIMEOUT_MS). Re-cut clean onto the release tip (branch was fossilized from a pre-v3.8.40 snapshot). Validated: 4 firecrawl tests green, env-doc-sync + docs-sync pass. UNSTABLE red is the inherited environmental setup-claude base-red.

* feat(xai): register XaiExecutor with reasoning-effort suffix parsing (#5800)

Integrated into release/v3.8.44 — XaiExecutor with reasoning-effort suffix parsing. Re-cut clean onto the release tip (branch was fossilized). Validated: 6 xai-executor tests green, provider-consistency OK, typecheck:core 0 errors, env-doc-sync in sync. UNSTABLE red is the inherited environmental setup-claude base-red.

* feat(discovery): Phase 2 — reporter, /api/discovery/* routes (strict loopback-only) + dashboard UI (#5939)

* feat(discovery): Phase 2 reporter — discoveryResults DB module + service wiring

Adds src/lib/db/discoveryResults.ts (CRUD over the discovery_results table
from migration 074) and wires the opt-in discovery service to persist and read
findings through it: persistDiscoveryResult / getDiscoveryResults /
getDiscoveryResultById / markVerified / deleteDiscoveryResult, with
(provider, method, endpoint) upsert de-duplication. Re-exported from localDb.

The service stays opt-in / default-off. The /api/discovery/* routes and the
dashboard UI tab are intentionally deferred to Phase 2b — they need the
local-only enforcement model (Hard Rules #15/#17 territory) decided first.

TDD: tests/unit/db/discovery-results.test.ts (8 cases, DB + service delegation),
isolated DATA_DIR with resetDbInstance cleanup.

* feat(discovery): Phase 2b — /api/discovery/* routes (strict loopback-only)

Adds the discovery HTTP surface on top of the reporter DB module:
  GET    /api/discovery/results            list findings (optional ?providerId)
  GET    /api/discovery/results/:id        one finding (404 if absent)
  DELETE /api/discovery/results/:id        delete a finding
  POST   /api/discovery/scan               scan a provider + persist findings
  POST   /api/discovery/verify/:id         mark a finding verified

Authorization: strict loopback-only. "/api/discovery/" is added to
LOCAL_ONLY_API_PREFIXES so the central authz pipeline (proxy.ts →
runAuthzPipeline → managementPolicy) rejects non-loopback callers with a 403
LOCAL_ONLY before any handler runs. It is deliberately NOT in
LOCAL_ONLY_MANAGE_SCOPE_BYPASS_PREFIXES — no remote manage-scope bypass —
because POST /scan issues outbound probes to provider endpoints (SSRF-adjacent)
and must never be tunnel-reachable. Handlers also call requireManagementAuth
(defense in depth) and return sanitized errors via createErrorResponse.

Tests:
- tests/unit/authz/discovery-routes-local-only.test.ts (8) — security guard:
  isLocalOnlyPath true + not manage-scope-bypassable for all four paths.
- tests/unit/api/discovery-routes.test.ts (6) — handler integration over an
  isolated DATA_DIR: list/filter, by-id 200/404/400, scan persist + 400 on
  empty/malformed body, verify 200/404, delete 200/404, no stack-trace leak.

* feat(discovery): Phase 2c — dashboard UI tab (Tools → Discovery)

Adds the /dashboard/discovery page (DiscoveryPageClient) that consumes the
Phase 2b /api/discovery/* routes: scan a provider, list findings, verify or
delete them. Registered in the sidebar under the Tools group (icon
travel_explore) and given a "discovery" i18n namespace + sidebar keys in
en.json (other locales fall back to en via next-intl until synced — the
locale files are in a pre-existing coverage deficit unrelated to this change).

Registers the UI test path in vitest.config.ts (advisory ui suite).

Tests: src/app/(dashboard)/dashboard/discovery/__tests__/DiscoveryPageClient.test.tsx
(3 cases: loads+renders results, empty state, fetches /api/discovery/results on
mount; stable useTranslations mock to avoid the fetch-loop). NOTE: the ui vitest
suite cannot run in this workspace — @testing-library/dom (a @testing-library/
react peer dep) is absent from node_modules, which fails ALL existing ui tests
equally; the test runs in CI. Component verified locally via typecheck + lint.

* test(discovery): register discovery-routes-local-only in stryker tap.testFiles

The mutation-test-coverage gate (--strict) flags any unit test covering a
mutated module that isn't listed in stryker.conf.json tap.testFiles. This PR's
tests/unit/authz/discovery-routes-local-only.test.ts covers src/server/authz/
routeGuard.ts (a mutated module, which this PR edits by adding the
/api/discovery/ local-only prefix), so it must be registered for its mutant
kills to count. No behavior change.

* refactor(discovery): split DiscoveryPageClient to satisfy max-lines-per-function

The complexity ratchet (max-lines-per-function: 80) flagged the single
184-line DiscoveryPageClient function (+1 over baseline). Extract the data
layer into two hooks (useDiscoveryResults for list/loading/feedback,
useDiscoveryActions for scan/verify/delete), a shared callApi helper, and two
presentational sub-components (DiscoveryScanForm, DiscoveryResultCard). Every
function is now under the 80-line ceiling; complexity gate back to baseline
1995. No behavior change — same exported component, same endpoints, same props.

* test(sidebar): include discovery in omni-proxy item-order snapshot

Adding the Discovery item to the Tools group (this PR's sidebar entry) extends
the ordered omni-proxy section list. Update the exact-match deepEqual snapshot
in sidebar-visibility.test.ts to include "discovery" in its position (after
traffic-inspector). The assertion stays exact — this reflects the intentional
new item, it does not weaken the check.

* docs(changelog): restore release bullets eaten by merge auto-resolve; re-add discovery bullet additively

* chore(quality): bump testFrozen for translator-openai-responses-req.test.ts (1097 -> 1172)

Base-red inherited from #5933, which grew the test file to 1171 lines
(Hard Rule #18 regression tests) without adjusting the frozen cap. The
release tip itself fails check:file-size; this unblocks every PR into
release/v3.8.44. File untouched by this PR.

* chore(quality): restore stryker tap.testFiles entries eaten by merge auto-resolve

The merge of origin/release/v3.8.44 silently dropped the 3 entries added
on the release side (#5903, clinepass, #5923). Took the release version
verbatim and re-added only this PR's entry (discovery-routes-local-only)
in alphabetical order. check:mutation-test-coverage green locally.

* chore(quality): reconcile inherited v3.8.44 merge-burst drift + include discovery in tools-group order test

- complexity 1995->2003 and cognitive 856->859: both measure IDENTICAL on
  the pristine release tip (3a3d618fe) and this PR's merged HEAD — the PR
  is complexity-net-zero; drift is from the 2026-07-02 merge burst
  (notes added to both baselines, same family as prior reconciliations).
- sidebar-tools-group.test.ts: append 'discovery' to the expected
  TOOLS_GROUP order — the intentional new sidebar item this PR adds
  (same expected-value update already made in sidebar-visibility.test.ts).

* feat(providers): custom icon URL for compatible provider nodes (#5815)

Integrated into release/v3.8.44 — custom icon URL for compatible provider nodes (DB migration 113 + nodes.ts + Zod schema + API routes + catalog + ProviderIcon UI). Re-cut onto the release tip (branch was fossilized ~13 real files); reconciled icon_url into the release's evolved nodes.ts/routes via 3-way. Validated: 14 backend + 5 frontend(vitest) + 24 page-utils tests green, typecheck:core 0, provider-consistency OK, file-size/env-doc-sync pass. UNSTABLE red is the inherited environmental setup-claude base-red.

* feat(api): add /v1/audio/translations endpoint (#5809)

Integrated into release/v3.8.44 — /v1/audio/translations endpoint (Whisper-style audio translation) + audioTranslation handler + translation providers in audioRegistry. Re-cut clean onto the release tip (branch was fossilized). Validated: 8 route tests (incl. no-stack-leak), typecheck:core 0, route-guard-membership OK, docs gates pass. UNSTABLE red is the inherited environmental setup-claude base-red.

* feat(dashboard): wildcard-CORS runtime warning + CORS security doc (#5602) (#5759)

Integrated into release/v3.8.44 — wildcard-CORS runtime warning banner + docs/security/CORS.md security guide (#5602). Re-cut clean onto the release tip (branch was fossilized). Validated: 20+9 backend + 2 banner(vitest) tests green, typecheck:core 0, docs-sync/symbols/fabricated/doc-links pass. UNSTABLE red is the inherited environmental setup-claude base-red.

* refactor(executors): extract pure JSONL stream translation from huggingchat (#6016)

Extract the pure JSONL->OpenAI-SSE translation (sseChunk, parseJsonlLine,
streamJsonlToOpenAi, readJsonlResponse) verbatim into the leaf huggingchat/jsonlStream.ts.
They consume a passed-in ReadableStream (no fetch/network/state). Host imports back the
two it uses; all module-private (no re-export).

Host 812 -> 594 LOC. Byte-identical bodies (verbatim), leaf has zero imports (no cycle).
Cookie/auth/multipart/execute untouched. Adds a split-guard; consumer tests stay green
(executor-huggingchat 6, huggingchat-model-catalog 3).

* refactor(executors): extract pure Meta AI response parser from muse-spark-web (#6017)

Extract the pure Meta AI SSE/JSON response parsing + content/reasoning/error extraction
(parseMetaSseFrames, readMetaJsonPayloads, collect*/extract*/classify* helpers,
parseMetaAiResponseText, isRecord, the reasoning/renderer key arrays, MetaSseFrame/
ParsedMetaAiResponse types) verbatim into the leaf muse-spark-web/response-parser.ts.
Host imports back the 3 it uses; all module-private (no re-export).

Host 1301 -> 925 LOC. Byte-identical bodies (verbatim), leaf has zero imports (no cycle).
Conversation cache, cookie/auth, fetch, executor class untouched. Adds a split-guard;
consumer tests stay green (muse-spark-cookie-copy-5449 2, muse-spark-web-continuation 6).

* refactor(executors): extract pure EventStream framing from kiro (#6018)

Extract the pure AWS EventStream binary framing (ByteQueue, CRC32 table + crc32,
TEXT_ENCODER/TEXT_DECODER, KIRO_VERIFY_FULL_CRC, parseEventFrame, EventFrame type)
verbatim into the self-contained leaf kiro/eventstream.ts (local JsonRecord alias to avoid
a cycle). Host imports back the 3 it uses (ByteQueue, TEXT_ENCODER, parseEventFrame).

Host 943 -> 758 LOC. Byte-identical bodies (verbatim 145/145), leaf has zero host imports
(no cycle). Auth/token-refresh/streaming-state/executor class untouched; the test-imported
flushBufferedToolArgs/resolveKiroRegion/kiroRuntimeHost stay exported on the host. Adds a
split-guard; consumer tests stay green (executor-kiro 9, kiro-tool-args-streaming 7,
kiro-iam-region 10).

* refactor(executors): extract challenge solver from duckduckgo-web (#6020)

Extract the DuckDuckGo anti-abuse challenge solver + FE signals (CHALLENGE_STUBS,
countHtmlElements, buildHtmlLookup, sha256Base64, solveDuckDuckGoChallenge,
makeDuckDuckGoFeSignals) verbatim into the leaf duckduckgo-web/challenge.ts. The vm
sandbox + 5s timeout (SECURITY note) are preserved. Host imports back the two it uses.

Host 924 -> 788 LOC. Byte-identical bodies (verbatim 132/132), leaf does not import the
host (no cycle). The now-dead createHash/parse5 host imports are removed; vm stays (still
used in host). Auth/cookie/warm/seed/executor untouched. Adds a split-guard; consumer
tests stay green (duckduckgo-web-executor 15, duckduckgo-domain-4037 8).

* test(cli): deflake setup-claude.test.ts — silence console to stop stdout/report interleaving (#5959) (#6019)

Integrated into release/v3.8.44. Deflakes tests/unit/cli/setup-claude.test.ts (#5959) — verified in CI: setup-claude now passes in Unit Tests fast-path (2/2).

Merged with --admin over two PRE-EXISTING base-reds proven independent of this test-only change (this PR only touches setup-claude.test.ts + CHANGELOG):
- Fast Quality Gates → check:test-discovery: tests/unit/executors/{firecrawl-fetch,xai-executor}.test.ts are orphaned on release/v3.8.44 (added by #5793/#5800); the shard glob 'tests/unit/{api,...,ui}/**' omits 'executors'. Both blobs exist on the pristine base.
- Unit Tests fast-path (2/2): tests/unit/settings-i18n-keys.test.ts → 'direct translation calls have English messages' fails on the pristine base too (unrelated i18n base-red).

* fix(cli): stabilize setup-claude.test.ts flake — inject dry-run log sink (#6021)

* fix(cli): stabilize setup-claude.test.ts flake — inject dry-run log sink (#5959)

Root cause (isolated empirically, 5/10 fail on the pristine base): the
dry-run path of syncClaudeProfilesFromModels console.log's a multi-byte
box-drawing heading ("── [dry-run] … ──"). Under the node:test runner
that write lands on the test child's stdout and corrupts the runner's
V8-serialized event stream ~50% of the time ("Unable to deserialize
cloned data due to invalid or unsupported version"), killing the file at
the first logging test. ASCII-only logging never reproduced it (0/20);
the unicode heading alone reproduced it (10/20).

Fix: syncClaudeProfilesFromModels accepts an injectable log sink
(opts.log, CLI default unchanged: console.log). The dry-run test injects
a collector — keeping unicode off the child's stdout — and gains
assertions on the dry-run report (path + parsed settings content), which
FAIL on the old code (log ignored) and PASS on the new one.

Validation: 0/30 failures post-fix vs 5/10 pre-fix on the same tree.

Baselines: complexity 2003->2006 and cognitive 859->860 are inherited
post-3a3d618fe release drift — measured identical on the pristine base
with and without this change (notes added in both files).

* test(ci): collect the orphaned tests/unit/executors/ directory (base-red unblock)

#5800 created tests/unit/executors/ outside every unit-runner brace glob,
so its 2 test files (firecrawl-fetch, xai-executor) never ran anywhere and
check:test-discovery flags them as NEW orphans on the pristine base,
red-flagging every PR into release/v3.8.44. Added 'executors' to the
runner globs in package.json (7 scripts), ci.yml unit shards, quality.yml
TIA glob, build-test-impact-map.mjs, and the test-discovery gate's
COLLECTORS (the gate enforces those stay in sync). Both files pass when
actually collected (10/10); cli+executors under suite flags: 99/99.

* chore(quality): complexity baseline 2006 -> 2007 (CI-observed value)

The GitHub fast-gates runner measures 2007 where local measures 2006 —
the same local-vs-CI off-by-one documented in the 2026-06-26 note. Pin
the CI-observed value so the gate is deterministic where it runs.

* fix(i18n): add the 6 missing en.json keys flagged by settings-i18n-keys (base-red unblock)

providers.iconUrlLabel/iconUrlHint (referenced by AddCompatibleProviderModal
and EditCompatibleNodeModal) and settings.authz.cors.wildcard.title/desc
(the #5602 CORS_ALLOW_ALL banner in AuthzSection) shipped without their
en.json messages — 'direct translation calls have English messages' fails
on the pristine release tip, red-flagging every PR. git log -S proves the
keys never existed (not a merge-eat). Scanner test: 10/10 green.

* refactor(executors): extract reasoning-effort (base) + tool-normalization (codex) leaves (#6030)

Two pure-leaf follow-ups closing the Block H tail:

- base/reasoningEffort.ts: provider-aware reasoning_effort sanitation
  (MISTRAL/GITHUB reject patterns, supportsMaxEffortForProvider,
  sanitizeReasoningEffortForProvider). Deps are config/services only
  (PROVIDER_CLAUDE, isClaudeCodeCompatible, supportsClaudeMaxEffort/supportsXHighEffort)
  so the leaf never imports the host — no cycle. base.ts re-exports
  sanitizeReasoningEffortForProvider for its external importers (mimoThinking + tests).
  base.ts 1466 -> 1312 LOC.

- codex/tools.ts: Responses-API tool normalization (CODEX_HOSTED_TOOL_TYPES hosted-tool
  passthrough, isCodexFreePlan gating, normalizeCodexTools). Self-contained
  (console.debug only). codex.ts re-exports isCodexFreePlan + normalizeCodexTools for
  external importers (tests + provider services). codex.ts 1430 -> 1268 LOC.

Byte-identical bodies (verbatim: base 100/100, codex 126/126); both leaves have zero host
imports. Adds two split-guards asserting the leaf owns the symbol and both import paths
resolve to the same function. Consumer tests stay green (base-executor-sanitize-effort 34,
executor-codex 40, mimoThinking 9, codex-free-plan-image-generation 3, issue-fixes 6).

* test(ci): move orphaned executor tests to top-level so a runner collects them (#6027)

Integrated into release/v3.8.44 — collect orphaned executor tests (check:test-discovery base-red).

* test(cli): deflake cli-setup-opencode.test.ts — silence console (#5959-class landmine) (#6033)

The command under test prints CLI progress with multi-byte glyphs
(printSuccess "✔" in the happy paths, printError "✖" in the dist-missing
path that test 4 exercises) via console.log. Under the node:test runner
those child-stdout writes interleave with the V8-serialized report frames
and can corrupt the stream — the exact #5959 mechanism proven for
setup-claude.test.ts; this file's ✖ line was already visible entangled in
red CI runs. No test here asserts on stdout, so silence console.log/info/
warn for the file (same pattern as #6019/#6021, restored in after()).

Validation: pre-fix the ✖/✔ lines reach stdout every run (grep-able);
post-fix stdout is clean, 4/4 tests green, 0/20 failures across 20 runs.

* feat(agy): support Google Cloud project ID settings (#5905)

* feat(agy): support Antigravity project ID settings

* refactor(agy): collapse Antigravity family project gate

---------

Co-authored-by: Nikolay Alafuzov <alafuzov_nn@rusklimat.ru>

* feat(proxy): add Webshare proxy pool import and sync (#5993)

* feat(proxy): add Webshare proxy pool import and sync

Adds Webshare (https://proxy.webshare.io) as a fourth source in the
free-proxy provider framework alongside 1proxy, Proxifly, and IPLocate.
WebshareProvider paginates the account's `/api/v2/proxy/list/` endpoint
(Authorization: Token <key>), upserts proxies into the shared
`free_proxies` table via the existing db/freeProxies.ts helpers, and
tombstones proxies the account no longer lists (recycled/retired IDs)
while never touching rows already promoted into the live proxy pool.

Unlike the other sources, Webshare is a paid per-account list, so it is
gated on FREE_PROXY_WEBSHARE_API_KEY rather than a plain on/off flag.
No DB migration needed — reuses the existing free_proxies table and
proxy_registry-on-promote path.

Co-authored-by: ricatix <d.enistraju155@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1176

* chore(changelog): restore release entries + add webshare bullet

---------

Co-authored-by: ricatix <d.enistraju155@gmail.com>

* feat(api-keys): add per-key device/connection tracking (#5998)

* feat(api-keys): add per-key device/connection tracking

Tracks distinct client devices (SHA-256 fingerprint of IP + User-Agent)
seen with each API key, with a 30-minute TTL and per-key/global caps. The
tracker is in-memory only (module-scoped Map, same pattern as
sessionManager.ts — no global.* singleton) and never stores the raw IP:
it is masked before being written.

Hooked into open-sse/handlers/chatCore.ts (the real chat entry) rather
than the legacy src/sse/handlers path. New GET /api/keys/[id]/devices
management route exposes masked device details for a key, and the
API Keys dashboard tab gets a "Devices" count badge alongside the
existing Sessions badge.

This is a new granularity distinct from the existing maxSessions cap
(src/lib/db/apiKeys.ts), which limits concurrent sticky-routing sessions
rather than tracking device identity.

Co-authored-by: Muhammad Mugni Hadi <mugni@rukita.co>
Inspired-by: https://github.com/decolua/9router/pull/931

* chore(changelog): restore release entries + add api-keys device-tracking bullet

---------

Co-authored-by: Muhammad Mugni Hadi <mugni@rukita.co>

* fix(providers): only apply openai-family model inference fallback when no cataloged provider serves the id (#5852) (#5938)

resolveModelByProviderInference() in open-sse/services/model.ts had an
unconditional /^gpt-/i heuristic that hijacked any model id starting with
gpt-/o1/o3 into provider openai, even when the id is cataloged under other
providers. This broke bare (non-combo) requests for open-weight models like
gpt-oss-120b (served by fireworks/cerebras/scaleway/byteplus/sambanova/
heroku), which don't exist on openai's catalog, producing a 404 with no
fallback.

Gate the heuristic on providers.length === 0 so it only fires for genuinely
uncataloged openai-family ids, letting cataloged ids fall through to the
existing single-candidate / ambiguous-candidate resolution paths.

Regression guard: tests/unit/gptoss-provider-inference-5852.test.ts

* fix(cc-compatible): send SSE accept for streamed requests (#5958)

Integrated into release/v3.8.44 — SSE Accept header for streamed cc-compatible requests (thanks @rdself).

* fix: deepseek-web reliability — auto-refresh on 401/403, refresh v2.0.0 client headers, fix token-kind bulk import (#5988)

Integrated into release/v3.8.44 — deepseek-web auto-refresh + v2.0.0 headers + token-kind bulk import (thanks @backryun).

* feat(providers): support Vercel AI Gateway embeddings and images (#5968)

* feat(providers): support Vercel AI Gateway embeddings and images

Extends the existing vercel-ai-gateway (alias vag) provider — currently
chat-only — with embeddings and image generation support, since the
gateway's OpenAI-compatible /v1 API also exposes /embeddings and
/images/generations. Adds entries to EMBEDDING_PROVIDERS
(embeddingRegistry.ts) and IMAGE_PROVIDERS (imageRegistry.ts) modeled
on the existing openai entries.

Out of scope for this PR (tracked as follow-ups): the /v1/credits
usage reader, retry:{429:2} tuning, and claude->reasoning_effort
mapping.

Co-authored-by: Ngô Tấn Tài <tantai@newnol.io.vn>
Inspired-by: https://github.com/decolua/9router/pull/1704

* chore(changelog): restore release entries + add vercel-gateway media bullet

---------

Co-authored-by: Ngô Tấn Tài <tantai@newnol.io.vn>

* feat(cli-tools): add Crush CLI tool to the dashboard (#5970)

* feat(cli-tools): add Crush CLI tool to the dashboard

Add a `crush` entry to the dashboard CLI-Tools catalog and a new
`/api/cli-tools/crush-settings` route (GET/POST/DELETE), cloned from the
`pi` tool's route as a template. OmniRoute already ships a `crush` CLI
command path (bin/cli/commands/setup-crush.mjs) but the dashboard catalog
had no matching entry.

The new route writes the real Crush config shape (providers.omniroute as
an openai-compat provider block) to the same canonical config path
(~/.config/crush/crush.json) that setup-crush.mjs's resolveCrushTarget()
already writes to, so the dashboard and the CLI command agree on one
location. Adds CLI_TOOL_RUNTIME_CONFIG.crush for detection/status, and
bumps EXPECTED_CODE_COUNT (18 -> 19) plus the catalog-count/schema tests
that enumerate the full tool list.

Co-authored-by: dopaemon <polarisdp@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1233

* chore(changelog): restore release entries + add crush cli bullet

---------

Co-authored-by: dopaemon <polarisdp@gmail.com>

* feat(dashboard): suggest HuggingFace Hub media models (#5990)

* feat(dashboard): suggest HuggingFace Hub media models

MVP scope:
- imageRegistry.ts: add an image kind entry for the huggingface provider
  (HF Inference API text-to-image), with a dedicated "huggingface-image"
  format since the endpoint returns raw image bytes rather than JSON.
- New handler open-sse/handlers/imageGeneration/providers/huggingface.ts,
  wired into imageGeneration.ts's format dispatch.
- New pure helper module open-sse/services/hfModelSuggestions.ts: maps a
  dashboard media kind to an HF Hub pipeline_tag and sorts/limits raw HF
  Hub search results (unit-tested directly).
- New route GET /api/v1/providers/suggested-models proxies the public HF
  Hub models search API server-side (Zod-validated query, buildErrorBody
  on every error path, no HF token exposed client-side — this project has
  no server-side HF search token config, so it calls unauthenticated).
- UI: ImageExampleCard now fetches suggested HF Hub models for the
  huggingface provider and merges them into the model picker as a
  selectable chip row, alongside the existing static provider models list.
- i18n: adds media.suggestedModels to en.json only.

Co-authored-by: yicone <yicone@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1633

* chore(changelog): restore release entries + add hf-hub media suggest bullet

---------

Co-authored-by: yicone <yicone@gmail.com>

* feat(dashboard): collapse and sort provider quota rows by remaining (#5977)

* feat(dashboard): collapse and sort provider quota rows by remaining

Sort the expanded quota list by remaining percentage (highest first)
and collapse it to the first 3 rows by default, with a "Show N more" /
"Show less" toggle when a connection reports more than 3 quotas. This
keeps the most at-risk quotas out of view below a long list of
healthy ones.

Extracts the sort/slice logic into pure helpers
(sortQuotasByRemaining, getVisibleQuotas) exported from
QuotaCardExpanded.tsx and unit-tests them directly.

Co-authored-by: CườngNH <j2.cuong@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1919

* chore(changelog): restore release entries + add quota collapse/sort bullet

---------

Co-authored-by: CườngNH <j2.cuong@gmail.com>

* feat(providers): refresh The Old LLM (Free) model catalog (#5181)

* feat(dashboard): add tool-source diagnostics settings toggle (#5978)

* feat(dashboard): add tool-source diagnostics settings toggle

Adds a Settings > Advanced card (cloned from DebugModeCard) that lets
operators flip the existing `logToolSources` flag from the UI instead
of editing the DB row directly. The backend gate (chatCore.ts) and DB
default were already present but had no toggle. Also adds
`logToolSources` to the /api/settings Zod PATCH schema (it is `.strict()`,
so the key was previously rejected) and en-only i18n strings.

Co-authored-by: DuyPrX <93126969+DuyPrX@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/1825

* chore(changelog): restore release entries + add tool-source toggle bullet

---------

Co-authored-by: DuyPrX <93126969+DuyPrX@users.noreply.github.com>

* feat(oauth): import Codex connection from a raw ChatGPT access token (#5995)

* feat(oauth): import Codex connection from a raw ChatGPT access token

OmniRoute's only Codex import path (/api/oauth/codex/import) required both
access_token and refresh_token, leaving no import path for a user who only
has a bare ChatGPT website access token (no refresh token).

- src/lib/db/providers.ts: createProviderConnection gains an explicit
  authType "access_token" branch — intentionally never deduped (no stable
  long-lived identity to match on) — and derives the connection name from
  email/name the same way "oauth" does.
- src/lib/oauth/services/codexImport.ts: export extractCodexAccountInfo so
  the new import path reuses the existing JWT decode instead of duplicating
  one.
- New route POST /api/oauth/codex/import-token (Zod-validated body
  { accessToken, name? }); errors routed through buildErrorBody /
  sanitizeErrorMessage. The executor's refreshCredentials() already
  degrades safely to null when there is no refresh token, forcing re-auth
  on expiry instead of a refresh exchange.
- OAuthModal.tsx: the callback-URL manual-paste path for codex now detects
  an eyJ-prefixed pasted token and posts it to the new endpoint, mirroring
  the existing grok-cli raw-token paste pattern.

Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/1290

* chore(changelog): restore release entries + add codex token-import bullet

---------

Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>

* fix(resilience): parse Retry-After from 429 JSON body for cooldown (#5974)

Integrated into release/v3.8.44 — parse Retry-After from 429 JSON body for cooldown (incl. #6013 retry-after-json extraction by @KooshaPari).

* fix(embeddings): forward connection-level proxy to embedding requests (#5975)

Integrated into release/v3.8.44 — forward connection-level proxy to embedding requests.

* fix(api): guard shared API client against non-JSON error responses (#5973)

Integrated into release/v3.8.44 — guard shared API client against non-JSON error responses.

* feat(dashboard): surface Codex banked reset credits per account (#5199)

* feat(providers): add NVIDIA NIM image generation (#5971)

* feat(providers): add NVIDIA NIM image generation

NVIDIA already exists as a chat provider (integrate.api.nvidia.com,
OpenAI-compatible) but image generation is served on a different host
(ai.api.nvidia.com/v1/genai/<model>) with a native NIM body shape, so it
gets a dedicated `nvidia-nim` image format and handler rather than reusing
the OpenAI image path.

Adds the 4 FLUX models (flux.1-dev, flux.1-schnell, flux.1-kontext-dev,
flux.2-klein-4b) to IMAGE_PROVIDERS, plus handleNvidiaNimImageGeneration()
which shapes the per-model NIM request body (flux.1-dev's mode/cfg_scale
and 768-1344px/64px-increment dimension validation, flux.1-kontext-dev's
required input image + aspect_ratio, schnell/klein's optional array-form
edit image) and normalizes the NIM response (artifacts[]/images[]/data[]/
single-value shapes) into the OpenAI `{created, data}` shape.

Co-authored-by: eng2007 <aleksey.semenov@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1195

* chore(changelog): restore release entries + add nvidia-nim image bullet

---------

Co-authored-by: eng2007 <aleksey.semenov@gmail.com>

* feat(providers): add Augment (Auggie CLI) local provider (#5972)

* feat(providers): add Augment (Auggie CLI) local provider

Adds a new local, no-auth provider that spawns the user's local `auggie`
CLI (`auggie --print --quiet --model <m> --`) and pipes a flattened prompt
via stdin, wrapping stdout as an OpenAI-compatible SSE stream or a single
chat.completion JSON body depending on the request's `stream` flag.

Auth is delegated entirely to `auggie login` outside OmniRoute — the
connection is registered `noAuth: true` and `refreshCredentials()` is a
no-op, matching the existing `NOAUTH_PROVIDERS` credential-less flow
(synthetic connection, no DB row required). An optional connection row is
still admitted via `FREE_APIKEY_PROVIDER_IDS` for display/priority
tracking, consistent with `opencode`. The dashboard "Test Connection"
flow spawns `auggie --version` to confirm the CLI is installed and
runnable, since there is no API key to validate upstream.

Security hardening (spawn is an untrusted-input sink):
- Command injection: spawn no longer passes `shell: true` on Windows. The
  binary is resolved to a concrete path/name and argv is handed straight to
  the OS loader, so no cmd.exe metacharacter interpretation is possible.
- Argument injection (flag smuggling): `model` is validated against the
  registry allowlist (`auggieProvider.models`) before any spawn — a model
  that is unknown or starts with "-" is rejected with a sanitized error and
  the subprocess is never started. A trailing `--` marks end-of-options in
  the argv as belt-and-suspenders.

Co-authored-by: chamdanilukman <16629923+chamdanilukman@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/1200

* test(golden): regenerate translate-path for auggie provider

---------

Co-authored-by: chamdanilukman <16629923+chamdanilukman@users.noreply.github.com>

* feat(providers): add ModelScope OpenAI-compatible provider (#5965)

* feat(providers): add ModelScope OpenAI-compatible provider

Ports ModelScope (Alibaba 魔搭) as a new API-key, OpenAI-compatible
provider — upstream 9router PR #1764. The upstream PR hardcoded
`https://api-inference.modelscope.ai/...` (`.ai` TLD); verified against
ModelScope's own API-Inference docs and third-party integration guides
that the real production domain is `api-inference.modelscope.cn`
(`.cn` TLD) and shipped that instead. Also drops the PR's static
5-model snapshot in favor of `passthroughModels: true` with an empty
seed list + `modelsUrl`, since ModelScope's open-model catalog moves
fast.

Updates the providers-constants-split characterization test's hardcoded
APIKEY_PROVIDERS count (159 -> 160) to match the new entry.

Co-authored-by: Umar Javed <114807145+tn5052@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/1764

* chore(changelog): restore release entries + add modelscope bullet

* test(golden): regenerate translate-path for modelscope provider

---------

Co-authored-by: Umar Javed <114807145+tn5052@users.noreply.github.com>

* feat(providers): add Qiniu OpenAI-compatible provider (#5966)

* feat(providers): add Qiniu OpenAI-compatible provider

Wires Qiniu (七牛云) AI inference gateway as a BYOK API-key provider.
Qiniu proxies many upstream models (DeepSeek V3/V4, Claude, Kimi and
more) behind a single key, so it ships with an empty static seed and
relies on passthroughModels + the live /v1/models catalog instead of a
single stale hardcoded model id.

- metadata: src/shared/constants/providers/apikey/gateways.ts
- registry entry: open-sse/config/providers/registry/qiniu/index.ts
  (format openai, executor default, bearer auth, baseUrl
  https://api.qnaigc.com/v1/chat/completions, modelsUrl
  https://api.qnaigc.com/v1/models)
- added to NAMED_OPENAI_STYLE_PROVIDERS so model import serves the live
  catalog and falls back to the (empty) local catalog on error, same
  pattern as the existing dgrid/zenmux/orcarouter gateways
- tests: tests/unit/qiniu-provider.test.ts (metadata, registry
  resolution, passthrough validation, live /v1/models fetch + fallback)

Co-authored-by: JiangZhuo <jiangzhuo@qiniu.com>
Inspired-by: https://github.com/decolua/9router/pull/911

* chore(changelog): restore release entries + add qiniu bullet

* test(golden): regenerate translate-path for qiniu provider

* test(providers): bump APIKEY count 160→161 for qiniu

---------

Co-authored-by: JiangZhuo <jiangzhuo@qiniu.com>

* feat(providers): add b.ai OpenAI-compatible provider (#5969)

* feat(providers): add b.ai OpenAI-compatible provider

Adds bai as a new OpenAI-compatible BYOK provider, distinct from the
existing thebai/theb.ai provider, using passthrough model discovery
(no hardcoded model list, live catalog served from api.b.ai/v1/models).

Co-authored-by: Delynn Assistant <zhen@dkzhen.org>
Inspired-by: https://github.com/decolua/9router/pull/963

* test(golden): regenerate translate-path for b.ai provider

* test(providers): bump APIKEY count 161→162 for b.ai

---------

Co-authored-by: Delynn Assistant <zhen@dkzhen.org>

* feat(providers): add Nube.sh OpenAI-compatible provider (#5936)

* feat(providers): add Nube.sh OpenAI-compatible provider

Nube.sh is a live BYOK OpenAI-compatible gateway (LiteLLM proxy) at
https://ai.nube.sh/api/v1, Bearer/API-key auth. Registered as an apikey
inference-host with an OpenAI-format, default-executor registry entry.

Its live model catalog is only reachable with a valid key
(/api/v1/models returns 401 unauthenticated), so no model IDs are
hardcoded — the entry uses passthroughModels + modelsUrl for live
enumeration instead of shipping unverifiable IDs.

Co-authored-by: whale9820 <87256750+whale9820@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2294

* test(golden): regenerate translate-path for nube provider

* test(providers): bump APIKEY count 162→163 for nube

---------

Co-authored-by: whale9820 <87256750+whale9820@users.noreply.github.com>

* feat(providers): add Charm Hyper OpenAI-compatible provider (#5961)

* feat(providers): add Charm Hyper OpenAI-compatible provider

Registers Charm Hyper (hyper.charm.land) as a new API-key gateway
provider: OpenAI-compatible chat completions format, bearer auth,
free tier (100 monthly Hypercredits). Models are resolved via
passthrough (modelsUrl + live /v1/models import) instead of a
hardcoded upstream model list, since the specific model catalog is
not publicly documented.

Co-authored-by: whale <admin@dyntech.cc>
Inspired-by: https://github.com/decolua/9router/pull/2006

* test(golden): regenerate translate-path for charm-hyper provider

* test(providers): bump APIKEY count 163→164 for charm-hyper

---------

Co-authored-by: whale <admin@dyntech.cc>

* feat(providers): add SumoPod and X5Lab OpenAI-compatible providers (#5963)

* feat(providers): add SumoPod and X5Lab OpenAI-compatible providers

Both are OpenAI-compatible BYOK aggregator gateways, wired via the
default executor with bearer API-key auth. Neither ships a hardcoded
model list — both use passthroughModels with an empty seed list and a
live /v1/models fetcher, so the catalog always reflects what each
gateway actually serves instead of speculative model IDs.

- SumoPod: https://ai.sumopod.com/v1/chat/completions (sk- keys)
- X5Lab: https://api.x5lab.dev/v1/chat/completions (x5- keys)

Regression guard: tests/unit/sumopod-x5lab-provider.test.ts.

Co-authored-by: Rigel Ramadhani Waloni <rigel8911@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1288

* chore(changelog): restore release entries + add sumopod/x5lab bullet

* test(golden): regenerate translate-path for sumopod + x5lab providers

* test(providers): bump APIKEY count 164→166 for sumopod + x5lab

---------

Co-authored-by: Rigel Ramadhani Waloni <rigel8911@gmail.com>

* feat(server): support reverse-proxy basePath deployment (#5992)

* feat(server): support reverse-proxy basePath deployment

Adds OMNIROUTE_BASE_PATH (opt-in, empty by default) to next.config.mjs
using Next.js's native basePath support so a deployment behind a
reverse-proxy subpath (e.g. https://host/omniroute/) works without
manual header stripping. Next.js strips the configured prefix from
nextUrl.pathname before route classification, so classifyRoute() and
isLocalOnlyPath() keep matching un-prefixed paths.

The two hardcoded auth redirect targets in
src/server/authz/pipeline.ts (root "/" -> "/dashboard" and
unauthenticated dashboard -> "/login") now prefix with
request.nextUrl.basePath so they stay inside the deployed subpath.
Default empty basePath is a no-op for existing root-path deployments.

Co-authored-by: zocomputer <help@zocomputer.com>
Inspired-by: https://github.com/decolua/9router/pull/1810

* docs(env): document OMNIROUTE_BASE_PATH in .env.example + ENVIRONMENT.md; restore changelog

* docs(env): document AUGGIE_BIN + CLI_AUGGIE_BIN (base-red from #5972 auggie)

---------

Co-authored-by: zocomputer <help@zocomputer.com>

* refactor(combo): extract buildTargetTimeoutRunner from handleComboChat (#6036)

Bloco J (hot-path decomposition), Task 1. Extract the per-target-timeout dispatch wrapper
(handleComboChat's handleSingleModelWithTimeout closure) verbatim into the leaf
combo/targetTimeoutRunner.ts as a factory buildTargetTimeoutRunner({handleSingleModel,
comboTargetTimeoutMs, log}). The per-model abort still comes from target.modelAbortSignal,
so the outer request signal is intentionally not a dependency. Host call-sites unchanged.

combo.ts shrinks ~60 LOC; leaf is 91 LOC (<800). Body byte-identical (verbatim), no cycle.
This is the first slice toward extracting the shared attempt-loop/success/error handlers
(Tasks 3-4) that de-duplicate handleComboChat and handleRoundRobinCombo. Adds a dedicated
test (5) so the failover path can be mutated independently. Consumer tests stay green
(combo-strategy-fallbacks 24, combo-499-abort 5, empty-content-failover 3, body-400-stop 1,
priority-quota-exhaustion 2, rr-streaming-lock 1, rr-session-stickiness 2).

Plan: _tasks/superpowers/plans/2026-07-03-blocoJ-combo-hotpath-decomposition.md

* feat(cli-tools): add CodeWhale CLI tool (#5996)

CodeWhale (https://github.com/Hmbown/CodeWhale) is the actively-maintained
successor to DeepSeek TUI — same author, renamed project. Added as a dual
entry alongside the existing "deepseek-tui" catalog entry (rather than a
hard rename) so users who still run the old DeepSeek TUI binary keep a
working dashboard card, while new users are steered to "codewhale".

New /api/cli-tools/codewhale-settings route writes the primary
~/.codewhale/config.toml and keeps an existing legacy
~/.deepseek/config.toml in sync (read fallback + best-effort write sync),
mirroring deepseek-tui-settings/route.ts. CLI_TOOLS and cliRuntime catalogs
updated; catalog cardinality tests/constants bumped accordingly (18→19
visible code tools, 28→29 total).


Inspired-by: https://github.com/decolua/9router/pull/1761

Co-authored-by: aristorinjuang <aristorinjuang@gmail.com>

* feat(i18n): auto-detect browser language on first visit (#5979)

* feat(i18n): auto-detect browser language on first visit

Adds a pure detectBrowserLocale() matcher (exact match, zh-HK/zh-MO
folded to zh-TW, language-prefix match, else null) plus a client-only
LocaleAutoDetect component mounted once in the root layout. On first
visit (no locale cookie set), it reads navigator.languages, computes a
match against the supported locales, and persists it via the same
cookie/localStorage writer LanguageSelector already used for manual
selection (now extracted to shared/lib/persistLocale.ts) before
refreshing the router.

Co-authored-by: anmingwei <anmingwei@dobest.com>
Inspired-by: https://github.com/decolua/9router/pull/1324

* chore(changelog): restore release entries + add browser-lang-detect bullet

---------

Co-authored-by: anmingwei <anmingwei@dobest.com>

* fix(dashboard): render Update-now API errors as text, not the raw envelope object (#5991) (#6028)

Integrated into release/v3.8.44 — fix(dashboard) render Update-now API errors as text, not the raw envelope object (#5991).

Merged with --admin: the fix is a one-line frontend change funneling the error body through the already-tested extractApiErrorMessage() helper, guarded by tests/unit/ui/home-update-error-render-5991.test.ts (3/3 pass, 3/3 fail on pre-fix source). The release branch is under a heavy parallel-merge storm (tip advanced ~6× mid-CI), so the branch is synced to the latest tip and landed atomically to avoid perpetual CONFLICTING; unit-shard reds seen earlier were pre-existing base-reds/flakes unrelated to this source-scan-only change.

* feat(api): expose provider plugin manifest (#6001)

* feat(api): expose provider plugin manifest

* test(translator): split responses chat request coverage

* test(mutation): register provider coverage tests

* feat(api): expose provider plugin manifest

* fix(ci): fail closed for prerelease latest promotion

* chore(ci): reconcile provider manifest complexity gate

* feat(api): expose provider plugin manifest

* test(translator): split responses chat request coverage

* test(mutation): register provider coverage tests

* fix(ci): fail closed for prerelease latest promotion

* chore: rebase onto release tip; drop out-of-scope translator test split + promote-script tweak

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* docs(changelog): add provider plugin manifest entry

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* chore(stryker): register account-fallback-retry-after-json test (base-red)

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: kooshapari <kooshapari@users.noreply.github.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@outlook.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(providers): add CN sign-up geo-restriction notices for SenseNova & StepFun (#5462)

* feat(sidecar): advertise provider manifest url (#6007)

* feat(sidecar): advertise provider manifest url via X-OmniRoute-Provider-Manifest-Url header

Re-cut onto release tip: manifest-url feature only (dropped stale-base noise).

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* docs(changelog): add sidecar manifest-url entry

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* chore(complexity): rebaseline 2009->2015 (inherited release-tip drift; feature adds 0)

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(autoCombo): latency/speed-optimized routing mode + omniroute_pick_fastest_model MCP tool (#6011)

* feat(autoCombo): latency/speed-optimized routing mode + omniroute_pick_fastest_model MCP tool

* test(translator): split responses chat request coverage

* refactor(mcp): extract fastest-model tool modules

* fix(i18n): cover provider icon and cors labels

* test(mutation): register latency coverage files

* test(ci): collect executor unit tests

* refactor(ci): reduce latency path complexity

* fix(mcp): include models catalog module

* feat(autoCombo): latency/speed-optimized routing + omniroute_pick_fastest_model MCP tool

Re-cut onto release tip: keep speed-routing + MCP tool + supporting catalog split;
drop out-of-scope translator split, en.json/ci.yml/package.json orphans, and unrelated
proxyFetch/responsesStreamHelpers/tokenLimitCounter refactors.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: kooshapari <kooshapari@users.noreply.github.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@outlook.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* docs(changelog): restore #5181/#5199/#5462 feature bullets eaten by merge

* feat(usage): on-demand period-scoped usage-data reset (re-cut onto release tip) (#5831)

* chore(quality): rebaseline eslintWarnings 4199->4256 + cognitiveComplexity 860->861 (v3.8.44 cycle drift)

Inherited v3.8.44 cycle drift measured on release tip 72ee80649 by the release-green
pre-flight during the /review-prs fix-batch round. The Quality Ratchet does NOT run on
PR->release fast-gates, so eslint warnings + cognitive complexity accrue unmeasured
across the cycle. Cyclomatic complexity is already green (2012 < baseline 2015) and
needs no bump. Each value carries a dated justification note; no production code touched.

* feat(claude-code): opt-in auto-permission classifier compat mode (re-cut onto release tip) (#5810)

* feat(providers): client-identity header profiles for compatible nodes (re-cut) + forbid cookie in custom headers (#5812)

* docs(openapi): document 9 newly-added routes to restore coverage ratchet (v3.8.44)

Documents the routes added this cycle that dropped openapiCoverage 36.9%->36.2%
below the ratchet baseline: 2 public v1 endpoints (/v1/ocr Mistral-OCR-compatible,
/v1/audio/translations Whisper-compatible) with full request/response specs, plus 7
dashboard/CLI-local routes marked x-internal:true (suggested-models, provider-plugin-
manifest, keys/{id}/devices, settings/purge-usage-history, oauth/codex/import-token,
cli-tools crush-settings + codewhale-settings). Coverage 36.2%->37.8% (207/547),
above baseline 36.9. check:openapi-routes/security-tiers/fabricated-docs all pass.

* refactor(sse): decompose handleComboChat auto-strategy region (Block J Task 2 — parseAutoConfig + resolveAutoStrategyOrder) (#6049)

* refactor(sse): extract pure parseAutoConfig leaf from handleComboChat

Block J Task 2 (safe slice): the auto-strategy config-resolution block in
handleComboChat is a pure function of (combo, eligibleTargets) with no side
effects, no early returns and no mutation. Extract it verbatim into
open-sse/services/combo/autoConfig.ts::parseAutoConfig so the god-function
shrinks and the derivation is independently unit-testable.

Behavior is byte-identical (verbatim-audited); combo.ts 3309->3280 LOC.
Adds tests/unit/combo-auto-config-split.test.ts (5 cases) pinning the
strategy-precedence, candidate-pool, weights and fallback derivations.

* refactor(sse): extract resolveAutoStrategyOrder leaf from handleComboChat

Block J Task 2 (coupled slice): the ~215-line `if (strategy === "auto")`
branch of handleComboChat is extracted into
open-sse/services/combo/resolveAutoStrategy.ts::resolveAutoStrategyOrder.

The branch is a control-flow region (mutates orderedTargets +
autoUsedExplicitRouter, early-returns 429, side-effect _registerExecutionCandidates),
so it is not a pure byte-identical move: the two `return unavailableResponse(...)`
exits become `{ earlyResponse }` and the mutated locals are returned instead of
closed over. Every other logic line is verbatim (semantic diff = only those
wrappers + the deeper getLKGP import path). `buildAutoCandidates` lives in
combo.ts, so it is injected via deps to keep the leaf acyclic (same DI pattern as
buildTargetTimeoutRunner) — which also makes the branch independently testable.

combo.ts 3280->3065 LOC. typecheck:core + check:cycles clean; dead host imports
removed. 60/60 consumer tests (router-strategies / auto-combo-engine /
combo-strategy-fallbacks / scoring-clamp / candidate-expansion / hidden-models)
cover the routable path end-to-end; new tests/unit/combo-resolve-auto-strategy-split.test.ts
pins the DI contract + the early-429 and default-ordering exits.

* test(sse): point quota-bypass source scan at resolveAutoStrategy leaf

The 'auto combo disables hard provider quota cutoffs when relay requests bypass'
source scan asserted combo.ts contains the bypass logic
(relayOptions?.bypassProviderQuotaPolicy === true + quotaPreflight enabled:false).
That block was extracted verbatim into combo/resolveAutoStrategy.ts (Block J
Task 2), so the scan now reads the leaf. Behavior unchanged.

* fix(ci): release-green base-reds — #5695 test regex + file-size rebaseline (#6093)

- tests/unit/ui/quick-start-api-keys-link-5695.test.ts: tolerate Prettier
  splitting <Link href=...> across lines (\s+) so the step1Desc regex matches
  the multi-line /dashboard/api-manager Link instead of skipping to step2's
  single-line /dashboard/providers Link. Code is correct; the test was brittle.
- config/quality/file-size-baseline.json: rebaseline 5 files that grew via
  already-merged PRs on the release tip (ApiManagerPageClient 3017->3058,
  OAuthModal 969->989, cliRuntime 1090->1100, webProvidersA 805->809,
  deepseek-web.test 1081->1092). Dated note added; shrink tracked in #3501.

* fix(translator): wrap Kiro system prompt in <system-reminder> (port from 9router#2306) (#6053)

Kiro/CodeWhisperer has no system role, so system messages were normalized to a
user turn with no wrapper — the full Claude Code system prompt then appeared as
raw user text, polluting the model context. Wrap system-origin content in
<system-reminder> tags before merging it into the Kiro user message. Real user
turns are unaffected. Existing history-merge tests aligned to the wrapped value.

Reported-by: VitzS7 (https://github.com/decolua/9router/issues/2306)

* fix(translator): strip multipleOf from antigravity/gemini tool schemas (port from 9router#2309) (#6052)

`multipleOf` is not part of the Gemini/antigravity OpenAPI 3.0 schema subset, so
leaving it in function_declaration parameters triggered a hard upstream 400
("Unknown name multipleOf"). Add it to GEMINI_UNSUPPORTED_SCHEMA_KEYS so it is
stripped at every schema level; minimum/maximum stay (Gemini accepts them).

Reported-by: abil0321 (https://github.com/decolua/9router/issues/2309)

* fix(kimi-web, qwen-web): align model catalog with live /models + map scenario per model (#5915)

* fix(kimi-web): align catalog with live models

Update the kimi-web catalog and request scenario selection to match
www.kimi.com's live GetAvailableModels response.

* fix(qwen-web): stop aliasing qwen3-coder-plus

Keep qwen3-coder-plus as its own model because it is present in the
live Qwen web models catalog.

* feat(minimax): extract M3 <think> to reasoning_content on OpenAI-format tiers (#6050)

MiniMax M3 is registered with format:"openai" on 8 provider tiers (trae,
huggingchat, bazaarlink, ollama-cloud, opencode, cline, opencode-zen,
codebuddy-cn), where its raw <think>...</think> tags leaked directly into
`content` instead of surfacing as a separate `reasoning_content` field.

OmniRoute already has the extraction primitive
(extractThinkingFromContent in responseSanitizer/reasoning.ts); it was just
gated to deepseek-r1/r1-distill/qwq. Extend the allowlist
(isTextualReasoningTagNativeRoute) with a minimax-m3-only pattern, excluding
the two direct minimax/minimax-cn tiers, which stay on Anthropic's Messages
format (targetFormat: "claude") and already surface reasoning natively.


Inspired-by: https://github.com/decolua/9router/pull/2231

Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
Co-authored-by: zmf963 <19422469+zmf963@users.noreply.github.com>

* fix: unwrap Cline response envelope (#6046)

Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>

* refactor(sse): extract applyStrategyOrdering leaf from handleComboChat (Block J Task 3) (#6063)

* refactor(sse): extract applyStrategyOrdering leaf from handleComboChat

Block J Task 3: the ~177-line else-if chain covering every non-auto combo
strategy (lkgp / strict-random / random / fill-first / p2c / least-used /
cost-optimized / reset-aware / reset-window / context-optimized / headroom /
quota-share) is extracted into
open-sse/services/combo/applyStrategyOrdering.ts::applyStrategyOrdering.

Each branch only reorders orderedTargets (no early returns, no other mutable
state), so the extraction is a clean verbatim move returning the reordered list;
the host replaces the chain with `else { orderedTargets = await
applyStrategyOrdering(strategy, orderedTargets, deps); }`. Semantic diff vs the
original chain = only the leading `if` (was `} else if`), the trailing return
and the deeper getLKGP import path — no logic line changed. None of the 13 strategy
helpers live in combo.ts, so no DI/cycle (unlike the auto branch).

combo.ts 3065->2883 LOC (3309->2883 across Task 2+3). typecheck:core + check:cycles
clean; 9 dead host imports removed (targetSorters block emptied). 47/47 consumer
tests (router-strategies / combo-strategy-fallbacks / rr-session-stickiness /
tag-routing) cover the DB-backed branches end-to-end; new
tests/unit/combo-apply-strategy-ordering-split.test.ts pins random / fill-first /
unknown exits.

* test(sse): point #2359 modelStr-guard scans at applyStrategyOrdering leaf

The LKGP fallback + non-auto strategy ordering (the two target.modelStr string-
method call sites) were extracted verbatim from combo.ts into the
applyStrategyOrdering leaf (Block J Task 3). The #2359 source scans now read the
leaf that owns those usages; the guard and the no-unguarded-usage assertions are
unchanged in intent.

* chore(ci): scan combo strategy leaves in check:known-symbols

Block J decomposed the combo dispatch: the `strategy === "..."` branches for
the 12 non-auto strategies moved to combo/applyStrategyOrdering.ts and the auto
branch to combo/resolveAutoStrategy.ts. The known-symbols gate previously scanned
only combo.ts, so it would report those strategies as canonicalNotHandled. Scan
all three dispatch files. Verified: 18/18 canonical strategies via dispatch.

* fix(combo): fallback to sibling model on 500 for per-model-quota providers (#5976)

* fix(combo): fallback to sibling model on 500 for per-model-quota providers

Two issues prevented combo fallback when gemini/gemma-4-31b-it returned 500:

1. targetExhaustion: connection-level exhaustion marked the shared gemini
   connection as exhausted, skipping the sibling model (gemma-4-26b-a4b-it).
   Skip markConnectionLevelExhaustion for per-model-quota providers (gemini,
   github, passthrough, compatible) since a model-level 500 does not mean
   the connection is bad.

2. combo retry loop: the auth layer records a model lockout on 500, but the
   retry loop did not check isModelLocked before retrying — it retried the
   same locked model instead of falling back. Add isModelLocked guard before
   the transient-retry decision.

* fix tests timeout

* fix: clear quota fallback CI gates

* quality-gate: extract test SSE stream helpers

* drop scope creep

* fix(combo): retry sibling models only on 500 errors

* fix(combo): reconcile onto release/v3.8.44 — keep targetExhaustion 500 fix, drop slow integration test

Reconciled by maintainer onto the current release tip:
- kept the core fix (targetExhaustion.ts model-500 guard for per-model-quota
  providers + the isModelLocked retry early-return in combo.ts) and its unit test
- dropped tests/integration/combo-concurrent-failure-recovery.test.ts +
  _sseTestHelpers.ts: they use Math.random()-based delays and 30s timeouts, run
  >3min and are flake-prone in the test:integration CI job; the unit test
  (tests/unit/combo/combo-target-exhaustion.test.ts, 21 cases) fully covers the fix
- CHANGELOG entry added

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: Koosha Pari <kooshapari@gmail.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@outlook.com>
Co-authored-by: hartmark <hartmark@users.noreply.github.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(xai): surface Grok usage on quota dashboard via local usageHistory aggregation (#5806)

xAI has no public per-account quota API (the billing console requires a
session cookie, not an API key). Add getXaiUsage(connectionId), mirroring
the existing Xiaomi MiMo self-track pattern: sum tokens routed to the
connection from usage_history via getMonthlyProviderTokensForConnection
and surface them as a cumulative, uncapped quota (unlimited: true,
remaining: 100 — xAI has no fixed monthly cap). Register 'xai' in
USAGE_FETCHER_PROVIDERS and wire a switch case in getUsageForProvider.


Inspired-by: https://github.com/decolua/9router/pull/2150

Co-authored-by: ron <devestacion@gmail.com>

* feat(services): add Mux managed embedded service (#6034)

Adds Mux (coder/mux — local agent-orchestration daemon) as a fourth-tier
embedded service built on the existing ServiceSupervisor framework, the
same shape as 9Router and CLIProxyAPI:

- Installer (src/lib/services/installers/mux.ts): npm install/update via
  runNpm (array args + env-based prefix, no shell interpolation), modeled
  on ninerouter.ts. Mux ships an npm package (`mux`) with a documented
  headless `mux server --host <host> --port <port>` mode, so no
  git-clone+build path was needed.
- Registered in bootstrap.ts (SERVICES[] + buildSpawnArgsFactory).
- DB seed migration 113 (version_manager row, not_installed/auto_start=0).
- 7 API endpoints under /api/services/mux/ (install/start/stop/restart/
  update/status/auto-start) plus the shared [name]/logs SSE endpoint,
  mirroring the cliproxy route shape and delegating errors through
  createErrorResponse().
- Dashboard tab (MuxServiceTab) reusing ServiceStatusCard,
  ServiceLifecycleButtons, AutoStartToggle, ServiceLogsPanel.
- Docs: EMBEDDED-SERVICES.md (service table, architecture diagram, API
  reference, key-injection section), openapi.yaml, ENVIRONMENT.md,
  .env.example.

Security:
- Every /api/services/mux/* route is covered by the existing
  LOCAL_ONLY_API_PREFIXES "/api/services/" prefix (Hard Rule #17);
  added an explicit isLocalOnlyPath regression test for all 8 routes.
- Mux binds to 127.0.0.1 explicitly (never 0.0.0.0) as defense-in-depth,
  since it orchestrates AI agents that can execute host commands.
- The bearer token is generated the same way as 9Router's key
  (getOrCreateApiKey) and injected via MUX_SERVER_AUTH_TOKEN (mux's
  documented env form) rather than a CLI flag, so it never appears in
  `ps`/process listings.
- No shell interpolation anywhere in the installer (Hard Rule #13): all
  npm/spawn args are static arrays; the install prefix and auth token
  travel via the env option.


Inspired-by: https://github.com/decolua/9router/pull/1802

Co-authored-by: Ansh7473 <Ansh7473@users.noreply.github.com>

* feat(services): promote Bifrost to embedded/supervised service (#5670) (#5817)

Promotes Bifrost (@maximhq/bifrost — Go AI-gateway) from an env-only relay
sidecar to a first-class embedded/supervised service, matching the existing
cliproxy/9router model. Implements item #2 of #5670; the broader RouterBackend
contract (items #1, #3-#5) stays out of scope.

- Installer (npm-style, ninerouter model): install/update/getInstalledVersion/
  getLatestVersion (1h cache)/resolveSpawnArgs (Go single-dash flags, pinned
  BIFROST_TRANSPORT_VERSION), needsApiKey=false
- Bootstrap SERVICES entry (healthPath /v1/models) + spawn-args factory branch
- Migration 113 seeds the version_manager row (not_installed, port 8080,
  auto_update=1, provider_expose=1)
- 7 lifecycle API routes under /api/services/bifrost/ (verbatim from cliproxy,
  errors sanitized) — loopback-only via existing LOCAL_ONLY_API_PREFIXES
- Shared [name]/logs branch for bifrost
- Dashboard tab + registration in the services page shell
- Relay auto-wiring: getBifrostRoutingConfig defaults BIFROST_BASE_URL to the
  supervised port when the instance is running; explicit env still wins; the
  env-only relay path (/v1/relay/.../bifrost) stays unchanged (compat layer)
- Docs (EMBEDDED-SERVICES, openapi) + unit tests (installer/route-guard/routing,
  19 tests) + RUN_SERVICES_INT-gated integration lifecycle

Note: the actual Go-binary install/start/health path requires a documented VPS
live-test before merge (Hard Rule #18 / spec section 7); the gated integration
harness is the vehicle for that run.

* fix(ci): document BIFROST_PORT to clear env-doc-sync base-red

The Bifrost embedded-service merge referenced process.env.BIFROST_PORT
(src/lib/services/bootstrap.ts, default 8080) without adding it to
.env.example / ENVIRONMENT.md, so check:env-doc-sync failed on the release
tip and reddened Fast Quality Gates for every open PR->release. Docs-only.

* fix(providers): emulate OpenAI tool_calls in GitLab Duo executor (#6051) (#6111)

Co-authored-by: felssxs <felssxs@users.noreply.github.com>

* fix(providers): strip orphan tool_result on Antigravity MITM path (#6026) (#6115)

* fix(registry): update grok-cli model context lengths (#5913)

grok-build 128k→256k, grok-composer-2.5-fast 128k→200k to match actual Grok CLI /context capacities so context-aware routing stops filtering these models out. Registry-only.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(proxy): batch delete, auto-test, health scheduler + transitive alias fix (#5918)

Proxy-registry batch management (batch-delete, auto-test, background health scheduler) + fix resolveProviderAlias to follow the alias chain transitively (oc -> opencode -> opencode-zen). Probe target now operator-configurable via PROXY_HEALTH_TEST_URL. Scope-creep files from the original branch dropped.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(minimax): extract M3 reasoning_content on OpenAI-format tiers (#6073)

MiniMax M3 leaks raw <think>...</think> into content on 8 OpenAI-format provider tiers; extract it into reasoning_content, leaving the direct minimax/minimax-cn (Claude-format) tiers untouched. Replacement for the stale #5804 branch.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(ci): harden provider translate-path golden across CI runners (#6076)

Normalize OS/arch-derived request headers (X-Stainless-Os/Arch, (OS;arch) UAs, and Antigravity's os.platform()-derived platform substring) in the golden so the test is runner-independent. Fixes the Mac-literal Antigravity UA that would have failed on Linux CI. Supersedes stale #6002.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* test(embeddings): pin seeded connection to direct egress in route-edge-coverage (#5975 collateral)

#5975 made the embeddings service honor the connection-level proxy. The pre-existing
route-edge-coverage embeddings edge-case tests seed an openai connection while the
settings-proxy suite has left a provider-level proxy (provider.local:8080) in the shared
DATA_DIR that resetStorage() does not clear — inert before #5975, but now the leaked
proxy fast-fails the embedding upstream with PROXY_UNREACHABLE.

These tests do not exercise proxying, so seedOpenAIConnection now pins the connection to
proxyEnabled:false, making resolveProxyForConnection return a direct egress regardless of
leaked global proxyConfig. No assertions weakened; 16/16 in the file pass. Regression
surfaced by the concurrency=1 full-suite run; passes on #5975's parent, red after it.

* fix(config): externalize ws for copilot-m365-web executor (#6130, closes #6062)

Re-lands the #6098 ws-externalization fix onto release/v3.8.44 (it had merged to main by mistake and was reverted). Externalize ws/bufferutil/utf-8-validate so the copilot-m365-web WebSocket masking path works at runtime.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(providers): update Perplexity Web models (#6106)

Refresh the Perplexity Web model catalog + mode/model_preference mappings to the current live set. Regression guard: perplexity-web.test.ts.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(providers): update Gemini Web cookies and models (#6095)

Refresh Gemini Web cookie handling + model catalog. Regression guard: gemini-web.test.ts.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(models): normalize GLM-5.2 provider context (#6091)

Hosted GLM-5.2 provider aliases now respect their declared context caps instead of inheriting the native 1M; native/bare + verified OpenCode/ZenMux routes stay at 1M. Regression guards added.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(combo): prefer known context capacity over unknown (#6088)

When a combo filters a target for exceeding a known context limit, prefer remaining known-compatible targets over unknown-metadata ones. Regression guard: combo-context-window-filter.test.ts.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix: keep Claude tool results adjacent (#6035)

Reattach OpenAI tool_result adjacent to tool_use before Claude send (#6026). Integrated into release/v3.8.44.

* fix(security): persist IP filter config + enforce it in the authz pipeline (#6131) (#6132)

Integrated into release/v3.8.44 — IP filter persistence + authz-pipeline enforcement (closes #6131). HARD-neutro: validate-release-green on the merge shows the same 3 pre-existing base-reds as the release baseline (test-masking cycle-wide, unit red-herring, integration batch-E2E env); #6131's own tests + ip-filter/pipeline suites all green.

* fix(codex): use access_token.exp instead of id_token.exp for import expiresAt (#6075) (#6084)

Prefer access_token.exp over id_token.exp for Codex auth import (#6075). Integrated into release/v3.8.44.

* fix(compression): send patch-only to PUT /api/settings/compression in CompressionHub (#6039) (#6077)

Send patch-only to PUT /api/settings/compression in CompressionHub (#6039). Integrated into release/v3.8.44.

* fix: reqId ReferenceError in safety-net redirect, dead code, filename typo (#6097)

Fix reqId ReferenceError in safety-net combo redirect + dead-code + DESING→DESIGN rename. Integrated into release/v3.8.44.

* fix(combo): expand fingerprint-based providers into per-fingerprint combo targets (#6082)

Expand fingerprint-based providers into per-fingerprint combo targets. Integrated into release/v3.8.44.

* fix(auth): persist quota preflight account lockouts (#6090)

Persist quota preflight account lockouts until reset window. Integrated into release/v3.8.44.

* fix(combos): expand OpenCode/MiMo fingerprint accounts in combo builder (#6087) (#6092)

Expand OpenCode/MiMo fingerprint accounts in combo builder (#6087). Integrated into release/v3.8.44.

* chore(quality): rebaseline v3.8.44 release-green drift (eslint/cognitive/cyclomatic/file-size)

Measured on release tip 32e4c906e during the #6131/#5975 release-green pass:
eslintWarnings 4256->4270 (+14), cognitiveComplexity 861->867 (+6), cyclomatic
count 2015->2026 (+11), and testFrozen caps for models-catalog-route (1507->1600),
perplexity-web (959->999), route-edge-coverage (1234->1241, my #5975 comment +7).
Inherited cycle drift (the Quality Ratchet does not run on PR->release fast-gates);
compression 'bun not found' is a local-env false and codeql is within baseline, so
neither is rebaselined. No production code touched.

* fix(accountFallback): persist per-account 429 cascade + classify 'Monthly usage limit. Resets in N days.' (#6061)

Persist per-account 429 cascade + classify 'Monthly usage limit. Resets in N days'. Integrated into release/v3.8.44.

* feat(build): backend-only fast build (skip the dashboard frontend) (#6119)

Backend-only fast build (skip dashboard frontend). Integrated into release/v3.8.44.

* fix(provider-limits): clear transient rate-limit state when quota recovers (#6128)

Clear transient rate-limit state when quota recovers. Integrated into release/v3.8.44.

* docs: Normalize mixed-language documentation content (#6105)

Normalize mixed-language documentation to English. Integrated into release/v3.8.44.

* chore docs

* i18n(zh-CN): translate CHANGELOG entries and section headings (#6043)

Adopt zh-CN as a translated locale: translate CHANGELOG + supporting docs. Integrated into release/v3.8.44.

* chore(quality): rebaseline residual eslint + file-size drift (v3.8.44)

Residual drift on release tip 716041223 (moving target): eslintWarnings 4270->4279
(+9 as the branch advanced past the prior rebaseline) and testFrozen/frozen file-size
caps for providerLimits.ts (955->982), accountFallback.ts (1790->1864) and
sse-auth.test.ts (1553->1600). All inherited from parallel-session merges (e.g. #6128);
the two production god-files ideally warrant decomposition rather than a bump (tracked
as debt). No production code touched.

* fix(repo): remove Windows case-conflicting DESIGN duplicate (#6140)

Remove stale root DESIGN.md (Windows case-conflict with design.md). Integrated into release/v3.8.44.

* fix(provider-limits): close TOCTOU race in quota recovery clear (I2) (#6139)

Close TOCTOU race in quota recovery clear via CAS primitive (I2 from #6128). Integrated into release/v3.8.44.

* fix(glm): suppress </think> close marker leak in GLM Anthropic transport (#6133)

Suppress </think> close-marker leak in GLM Anthropic transport. Integrated into release/v3.8.44.

* fix(cli): give setup-claude a fallback profile generator like setup-codex (#6138)

Give setup-claude a fallback profile generator like setup-codex. Integrated into release/v3.8.44.

* fix(onboarding): route provider-details link by node id, not provider slug (#6145) (#6145)

Route onboarding provider-details link by node id (#6145). Integrated into release/v3.8.44.

* fix(translator): strip Responses-only truncation field before Chat Completions forwarding (#6109)

Strip Responses-only truncation field before Chat Completions forwarding (#2311). Integrated into release/v3.8.44.

* fix(mitm): guard against concurrent MITM server starts (#6107)

Guard against concurrent MITM server starts (#2316). Integrated into release/v3.8.44.

* feat(models): add claude-sonnet-5 to Antigravity catalog (#6103)

Add claude-sonnet-5 to Antigravity catalog. Integrated into release/v3.8.44.

* fix(providers): strip thinking param for minimax-m2.7 on NVIDIA NIM (#6102)

Strip unsupported thinking param for minimax-m2.7 on NVIDIA NIM. Integrated into release/v3.8.44.

* feat(providers): add Kenari OpenAI-compatible gateway (#6104)

Add Kenari OpenAI-compatible gateway (BYOK). Integrated into release/v3.8.44.

* feat(sse): per-request Auto-Combo controls (X-OmniRoute-Mode / X-OmniRoute-Budget) — closes #6023 #6024 #6025 (#6057)

Per-request Auto-Combo controls (X-OmniRoute-Mode / X-OmniRoute-Budget). Integrated into release/v3.8.44.

* feat(resilience): throttle concurrent upstream quota fetches — closes #6009 (#6058)

Throttle concurrent upstream quota fetches (#6009). Integrated into release/v3.8.44.

* fix(oauth): graceful 400 for keychain-import-only providers (zed) (#6041) (#6054)

Graceful 400 for keychain-import-only providers on OAuth route (zed, #6041). Integrated into release/v3.8.44.

* fix(dashboard): resolve broken Card import breaking next build (base-red from #6061) (#6155)

* fix(dashboard): resolve broken Card import breaking next build (base-red from #6061)

CoolingConnectionsPanel imported `Card` from `@/components/ui/card`, a path
that does not exist in this repo (there is no shadcn-style `src/components/ui/`).
The PR->release fast-gates do not run `next build`, so the broken import slipped
in and `next build` failed with:

  Module not found: Can't resolve '@/components/ui/card'

Fix: the <Card> here was only a styled container, so replace it with a <div>
carrying the equivalent Tailwind classes (border/bg/padding + rounded-card
shadow-sm). Also normalize the file from CRLF to LF (it shipped with CRLF).

Adds a vitest/jsdom regression test (tests/unit/ui/CoolingConnectionsPanel.test.tsx)
that fails-without-fix (Vite: 'Failed to resolve import @/components/ui/card')
and passes with it, plus renders/empty-state coverage. Rule #18.

* fix(dashboard): stop client CoolingConnectionsPanel dragging server DB barrel into browser bundle

Second base-red from #6061, surfaced once the broken Card import was fixed:

  ./node_modules/ioredis/built/connectors/StandaloneConnector.js
  Module not found: Can't resolve 'net'
  Import trace: ioredis <- rateLimiter.ts <- apiKeys.ts <- @/lib/localDb
                <- CoolingConnectionsPanel.tsx (a "use client" component)

The client panel imported `formatResetCountdown` from `@/lib/localDb` — the
server-side DB re-export barrel — which transitively pulls better-sqlite3/ioredis
(node:net) into the browser bundle. That violates the CLAUDE.md rule 'never
barrel-import from localDb'.

`formatResetCountdown` is a pure date-formatting function, so move its
implementation to the client-safe `@/shared/utils/formatting` (alongside
formatTime/formatDuration) and re-export it from db/providers/rateLimit.ts for the
existing server callers + barrel. The panel now imports it directly from the
shared util — no server code in the client bundle.

Tests (Rule #18):
- tests/unit/format-reset-countdown.test.ts (node:test, blocking test:unit) —
  pure-function coverage: null/past/invalid, s, m+s, h+m, ISO string.
- tests/unit/ui/CoolingConnectionsPanel.test.tsx mock updated to the new module.

* fix(release): v3.8.44 Phase-0 pre-flight — base-red sweep + ratchet absorption

- fix(models): stop resolveProviderAlias at registered provider ids so oc/
  reaches the no-auth opencode provider again (#2901 contract, regressed by
  #5918's transitive chain; transitivity kept across alias-only hops)
- fix(auggie): handle async EPIPE 'error' events on child stdin so a
  fast-exiting CLI surfaces a sanitized error instead of crashing (both
  spawn sites); deflakes auggie-executor tests
- test: align provider family count 166->167 (Kenari #6104), regenerate
  translate-path golden on Linux (+kenari), opencode quota scope
  provider->connection (#6061)
- quality(test-masking): add _deletedWithReplacement allowlist support to
  check-test-masking.mjs (deletion exempt ONLY when the declared replacement
  test exists in HEAD; 5 new gate unit tests) + reduction allowlist entries
  for the verified #5958/#6088/#5816 migrations + targetExhaustion->
  combo-target-exhaustion replacement (#5976, 21 cases/52 asserts vs 13/37)
- quality(file-size): absorb v3.8.44 cycle drift (oauth route 960,
  providerLimits 998, chat 1662, auth 2426) with justification; #6158 will
  restore the oauth-route freeze
- changelog: bullets for the above + the #6155 cooling-panel build fix

* chore(release): v3.8.44 — 2026-07-04

Release reconciliation + close (generate-release Phases 0a/1):
- CHANGELOG [3.8.44]: 21 PR refs added to existing bullets, 62 new bullets
  (incl. restoration of ~10 bullets erased by the stale-branch merge in
  1f6ec5bc8), 3 Maintenance rollups, #6061/#6130 credit fixes, 🙌 Contributors
  table (35 external contributors) — coverage 144/153 cycle commits by #ref
- 42 docs/i18n CHANGELOG mirrors synced (EN content; i18n workflow translates)
- README: What's New refreshed for v3.8.44 highlights
- build scope: exclude electron/node_modules + electron/dist-electron + .build
  from tsconfig (local build-output leak poisoned next build with 8GB OOM —
  same class as the 2026-06-25 incident; scope 14765→5207, gate green)
- quality: cyclomatic baseline 2026→2028 (+2 inherited end-of-cycle drift;
  verified the release-captain code fixes add 0 new violations)

* fix(release): v3.8.44 one-pass release-PR CI sweep

- fix(dashboard): /dashboard/system/proxy 500'd on EVERY render — #5918 put
  useProxyBatchOperations(load) before the const load declaration (TDZ
  ReferenceError, digest 539380095). Hook block moved after load; SSR
  renderToString regression test added (the exact crash mode).
- fix(server): TRACE/TRACK/CONNECT crashed Next's middleware adapter
  (undici cannot represent them) into a raw 500 on every route — the raw
  HTTP method guard now answers 405 + Allow up-front (dast-smoke
  Schemathesis finding on /api/keys/{id}/devices); guard test added.
- fix(api): restore Zod validation on the provider-scoped chat route via a
  .passthrough() schema preserving #5907's relaxed semantics (t06 gate).
- docs(openapi): /api/keys/{id}/devices 401 now refs the management error
  envelope (Schemathesis schema-conformance).
- quality: rebaseline i18nUiCoverage 77.5->76.8 (+~1352 new en.json UI keys
  from the cycle await the async translation workflow; v3.8.39 precedent).
- CodeQL: dismissed 2 incomplete-url-substring FPs on unit-test asserts
  (v3.8.35 precedent) with Hard Rule #14 justifications.
- changelog: bullets for the above + 42 i18n mirrors re-synced

* fix(release): round-2 CI findings — LocaleAutoDetect refresh gating + ratchet tighten

- fix(i18n): LocaleAutoDetect (#5979) refreshed the router on EVERY cookie-less
  first visit, even when the detected locale matched the server-rendered
  <html lang> — re-navigating mid-interaction (flaky e2e 'execution context
  destroyed' + visible flash for new visitors). Refresh now only fires when
  the locale actually differs; regression test added.
- quality: tighten openapiCoverage.pct 36.9->39.3 (require-tighten gate on the
  release PR; value measured by the CI Quality Ratchet on 00c55afcb)
- quality(file-size): shrink the ProxyRegistryManager TDZ note to fit the
  1117-line freeze (prettier reflow added a line at commit time)
- changelog bullet + 42 i18n mirrors re-synced

* test(release): collect the #6082 fingerprint-expansion ghost test

check:test-discovery (Lint job, layered behind the round-1 t06 fix) flagged
tests/e2e/fingerprint-expansion.test.ts as a NEW orphan — it is a node:test
server-boot test that no runner collected, so it had never run. Moved to
tests/integration/ (the collector for this shape), fixed the helper import,
and verified it actually passes (3/3 on first-ever run). CHANGELOG ref updated.

---------

Co-authored-by: Chirag Singhal <76880977+chirag127@users.noreply.github.com>
Co-authored-by: Hamsa_M <116961508+hamsa0x7@users.noreply.github.com>
Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com>
Co-authored-by: Vittor Guilherme Borges de Oliveira <vittoroliveira.dev@gmail.com>
Co-authored-by: nickwizard <35692452+nickwizard@users.noreply.github.com>
Co-authored-by: Ankit <177378174+anki1kr@users.noreply.github.com>
Co-authored-by: Fadhil Yusuf <33994304+yusufrahadika@users.noreply.github.com>
Co-authored-by: Giorgos Giakoumettis <giorgos@yiakoumettis.gr>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>
Co-authored-by: AgentKiller45 <jamalzzj45@gmail.com>
Co-authored-by: Nikolay Alafuzov <alafuzov_nn@rusklimat.ru>
Co-authored-by: ricatix <d.enistraju155@gmail.com>
Co-authored-by: Muhammad Mugni Hadi <mugni@rukita.co>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: backryun <bakryun0718@proton.me>
Co-authored-by: Ngô Tấn Tài <tantai@newnol.io.vn>
Co-authored-by: dopaemon <polarisdp@gmail.com>
Co-authored-by: yicone <yicone@gmail.com>
Co-authored-by: CườngNH <j2.cuong@gmail.com>
Co-authored-by: DuyPrX <93126969+DuyPrX@users.noreply.github.com>
Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>
Co-authored-by: eng2007 <aleksey.semenov@gmail.com>
Co-authored-by: chamdanilukman <16629923+chamdanilukman@users.noreply.github.com>
Co-authored-by: Umar Javed <114807145+tn5052@users.noreply.github.com>
Co-authored-by: JiangZhuo <jiangzhuo@qiniu.com>
Co-authored-by: Delynn Assistant <zhen@dkzhen.org>
Co-authored-by: whale9820 <87256750+whale9820@users.noreply.github.com>
Co-authored-by: whale <admin@dyntech.cc>
Co-authored-by: Rigel Ramadhani Waloni <rigel8911@gmail.com>
Co-authored-by: zocomputer <help@zocomputer.com>
Co-authored-by: aristorinjuang <aristorinjuang@gmail.com>
Co-authored-by: anmingwei <anmingwei@dobest.com>
Co-authored-by: janeza2 <49841619+janeza2@users.noreply.github.com>
Co-authored-by: zmf963 <19422469+zmf963@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Koosha Pari <kooshapari@gmail.com>
Co-authored-by: hartmark <hartmark@users.noreply.github.com>
Co-authored-by: ron <devestacion@gmail.com>
Co-authored-by: Ansh7473 <Ansh7473@users.noreply.github.com>
Co-authored-by: felssxs <felssxs@users.noreply.github.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Arthur Bodera <abodera@gmail.com>
Co-authored-by: Semianchuk Vitalii <fix20152@gmail.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: NOXX - Commiter <artur1992123@mail.ru>
Co-authored-by: Milan Soni <123074437+Iammilansoni@users.noreply.github.com>
Co-authored-by: Devin <studyzy@gmail.com>
Co-authored-by: Raxxoor <manker_lol@hotmail.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
2026-07-04 13:00:30 -03:00

2880 lines
85 KiB
TypeScript

import { randomUUID } from "crypto";
/**
* Image Generation Handler
*
* Handles POST /v1/images/generations requests.
* Proxies to upstream image generation providers using OpenAI-compatible format.
*
* Request format (OpenAI-compatible):
* {
* "model": "openai/gpt-image-2",
* "prompt": "a beautiful sunset over mountains",
* "n": 1,
* "size": "1024x1024",
* "quality": "standard", // optional: "standard" | "hd"
* "response_format": "url" // optional: "url" | "b64_json"
* }
*/
import { getImageProvider, parseImageModel } from "../config/imageRegistry.ts";
import { HTTP_STATUS } from "../config/constants.ts";
import { applyAntigravityClientProfileHeaders } from "../services/antigravityClientProfile.ts";
import { getAntigravityEnvelopeUserAgent } from "../services/antigravityIdentity.ts";
import { kieExecutor } from "../executors/kie.ts";
import { mapImageSize } from "../translator/image/sizeMapper.ts";
import { getCodexClientVersion, getCodexUserAgent } from "../config/codexClient.ts";
import { ChatGptWebExecutor } from "../executors/chatgpt-web.ts";
import { getChatGptImage, findChatGptImageBySha256 } from "../services/chatgptImageCache.ts";
import { createHash } from "node:crypto";
import { saveCallLog } from "@/lib/usageDb";
import { sleep } from "../utils/sleep.ts";
import {
getKieErrorMessage,
getKieErrorStatus,
isJsonObject,
parseKieResultJson,
} from "../utils/kieTask.ts";
import {
submitComfyWorkflow,
pollComfyResult,
fetchComfyOutput,
extractComfyOutputFiles,
} from "../utils/comfyuiClient.ts";
import { fetchRemoteImage } from "@/shared/network/remoteImageFetch";
import { FetchTimeoutError, fetchWithTimeout, getConfiguredTimeout } from "@/shared/utils/fetchTimeout";
import { sanitizeErrorMessage, sanitizeUpstreamDetails } from "../utils/error.ts";
// --- Per-provider handlers (extracted to co-located files in PR-#4582-batch) ---
// Imported locally so internal callers (handleImageGeneration / handleImageEdit)
// resolve to a real binding. extractMarkdownImageUrls + CHATGPT_WEB_IMAGE_ID_RE
// are still used by handleImageEdit below, so they are imported (not re-defined).
import { handleSDWebUIImageGeneration } from "./imageGeneration/providers/sdWebUI.ts";
import { handleHyperbolicImageGeneration } from "./imageGeneration/providers/hyperbolic.ts";
import { handleHuggingFaceImageGeneration } from "./imageGeneration/providers/huggingface.ts";
import { handleComfyUIImageGeneration } from "./imageGeneration/providers/comfyUI.ts";
import { handleImagen3ImageGeneration } from "./imageGeneration/providers/imagen3.ts";
import { handleIdeogramImageGeneration } from "./imageGeneration/providers/ideogram.ts";
import { handleHaiperImageGeneration } from "./imageGeneration/providers/haiper.ts";
import { handleLeonardoImageGeneration } from "./imageGeneration/providers/leonardo.ts";
import {
handleChatGptWebImageGeneration,
extractMarkdownImageUrls,
CHATGPT_WEB_IMAGE_ID_RE,
} from "./imageGeneration/providers/chatgptWeb.ts";
import { handleNvidiaNimImageGeneration } from "./imageGeneration/providers/nvidiaNim.ts";
interface KieImageOptions {
model: string;
provider: string;
providerConfig: {
baseUrl: string;
statusUrl?: string;
};
body: Record<string, unknown> & {
prompt?: unknown;
size?: unknown;
n?: unknown;
timeout_ms?: unknown;
poll_interval_ms?: unknown;
};
credentials?: {
apiKey?: string;
accessToken?: string;
} | null;
log?: {
info: (scope: string, message: string) => void;
error: (scope: string, message: string) => void;
} | null;
}
const OPENAI_IMAGE_TO_IMAGE_MODELS = new Set([
"black-forest-labs/FLUX.2-max",
"black-forest-labs/FLUX.2-pro",
"black-forest-labs/FLUX.2-flex",
"black-forest-labs/FLUX.2-dev",
"openai/gpt-image-1.5",
"Wan-AI/Wan2.6-image",
"Qwen/Qwen-Image-2.0-Pro",
"Qwen/Qwen-Image-2.0",
"google/flash-image-3.1",
"google/gemini-3-pro-image",
"flux-kontext-max",
"flux-kontext",
"flux-kontext-pro",
"qwen-image",
]);
const IMAGE_ASPECT_RATIO_PATTERN = /^\d+:\d+$/;
/**
* Resolve the upstream images endpoint for a custom (OpenAI-compatible) image
* provider node (#3205).
*
* Custom provider nodes store their base URL the same way the chat path does:
* in `credentials.providerSpecificData.baseUrl` (e.g. `https://example.com/v1`),
* NOT as a top-level `credentials.baseUrl`. Older callers may still pass a
* top-level `baseUrl`, so we honor that as a secondary source. When neither is
* present we fall back to `fallback` (the built-in Gemini OpenAI endpoint).
*
* Resolution order: providerSpecificData.baseUrl → credentials.baseUrl → fallback.
*
* A node base URL like `https://example.com/v1` is normalized and the
* OpenAI-compatible `/images/generations` path appended (mirroring
* `buildOpenAICompatibleUrl` in services/provider.ts). A node URL that already
* ends in `/images/generations` is returned as-is (no double-append). The
* `fallback` value is assumed to already be a complete URL and is returned
* verbatim.
*/
export function resolveImageBaseUrl(
credentials:
| { baseUrl?: unknown; providerSpecificData?: { baseUrl?: unknown } | null }
| null
| undefined,
fallback: string,
endpoint: "generations" | "edits" = "generations"
): string {
const psd = credentials?.providerSpecificData;
const psdBaseUrl =
psd && typeof psd === "object" && typeof psd.baseUrl === "string" && psd.baseUrl.trim()
? psd.baseUrl.trim()
: null;
const topLevelBaseUrl =
typeof credentials?.baseUrl === "string" && credentials.baseUrl.trim()
? credentials.baseUrl.trim()
: null;
const nodeBaseUrl = psdBaseUrl || topLevelBaseUrl;
if (!nodeBaseUrl) return fallback;
// A single configured node serves both image routes: honor a base URL that already
// points at the requested OpenAI image path, and rewrite one that points at the other
// image endpoint (e.g. `.../images/generations` requested for edits) (#3214/#3215).
const suffix = `/images/${endpoint}`;
// Trim trailing slashes without a backtracking-prone regex (`/\/+$/` is a
// polynomial-ReDoS pattern on long runs of "/" — CodeQL js/polynomial-redos).
let normalized = nodeBaseUrl;
while (normalized.endsWith("/")) normalized = normalized.slice(0, -1);
if (normalized.endsWith(suffix)) return normalized;
const stripped = normalized.replace(/\/images\/(?:generations|edits)$/, "");
return `${stripped}${suffix}`;
}
function normalizeImageAspectRatio(value: unknown, fallbackSize: unknown): string {
if (typeof value === "string") {
const trimmedValue = value.trim();
if (IMAGE_ASPECT_RATIO_PATTERN.test(trimmedValue)) return trimmedValue;
}
return mapImageSize(typeof fallbackSize === "string" ? fallbackSize : null);
}
function parseJsonOrNull(value: string): unknown | null {
try {
return JSON.parse(value);
} catch {
return null;
}
}
function sanitizeImageProviderError(errorText: string): unknown {
const parsed = parseJsonOrNull(errorText);
if (parsed !== null) {
return sanitizeUpstreamDetails(parsed) || sanitizeErrorMessage(errorText);
}
return sanitizeErrorMessage(errorText);
}
const BFL_MODEL_ENDPOINTS = {
"flux-2-max": "/v1/flux-2-max",
"flux-2-pro": "/v1/flux-2-pro",
"flux-2-flex": "/v1/flux-2-flex",
"flux-2-klein-9b": "/v1/flux-2-klein-9b",
"flux-2-klein-4b": "/v1/flux-2-klein-4b",
"flux-kontext-pro": "/v1/flux-kontext-pro",
"flux-kontext-max": "/v1/flux-kontext-max",
"flux-pro-1.1": "/v1/flux-pro-1.1",
"flux-pro-1.1-ultra": "/v1/flux-pro-1.1-ultra",
"flux-dev": "/v1/flux-dev",
"flux-pro": "/v1/flux-pro",
};
const BFL_EDIT_MODELS = new Set([
"flux-2-max",
"flux-2-pro",
"flux-2-flex",
"flux-kontext-pro",
"flux-kontext-max",
]);
const BFL_FAILURE_STATUSES = new Set(["Error", "Failed", "Content Moderated", "Request Moderated"]);
function formatImageProviderError(err) {
const sanitized = sanitizeErrorMessage(err);
const message = (sanitized || "").replace(/^Error:\s*/i, "").trim();
return message ? `Image provider error: ${message}` : "Image provider error";
}
const STABILITY_GENERATION_ENDPOINTS = {
"sd3.5-large": "/v2beta/stable-image/generate/sd3",
"sd3.5-large-turbo": "/v2beta/stable-image/generate/sd3",
"sd3.5-medium": "/v2beta/stable-image/generate/sd3",
"sd3.5-flash": "/v2beta/stable-image/generate/sd3",
"stable-image-ultra": "/v2beta/stable-image/generate/ultra",
"stable-image-core": "/v2beta/stable-image/generate/core",
};
const STABILITY_EDIT_ENDPOINTS = {
inpaint: "/v2beta/stable-image/edit/inpaint",
outpaint: "/v2beta/stable-image/edit/outpaint",
erase: "/v2beta/stable-image/edit/erase",
"search-and-replace": "/v2beta/stable-image/edit/search-and-replace",
"search-and-recolor": "/v2beta/stable-image/edit/search-and-recolor",
"remove-background": "/v2beta/stable-image/edit/remove-background",
"replace-background-and-relight": "/v2beta/stable-image/edit/replace-background-and-relight",
fast: "/v2beta/stable-image/upscale/fast",
conservative: "/v2beta/stable-image/upscale/conservative",
creative: "/v2beta/stable-image/upscale/creative",
sketch: "/v2beta/stable-image/control/sketch",
structure: "/v2beta/stable-image/control/structure",
style: "/v2beta/stable-image/control/style",
"style-transfer": "/v2beta/stable-image/control/style-transfer",
};
const STABILITY_CONTROL_MODELS = new Set(["sketch", "structure", "style", "style-transfer"]);
function appendOptionalFormValue(formData, key, value) {
if (value === undefined || value === null || value === "") return;
formData.append(key, String(value));
}
function appendImageFormValue(formData, key, source, filename) {
formData.append(
key,
new Blob([source.buffer], {
type: source.contentType || "application/octet-stream",
}),
filename
);
}
const FAL_PRESET_SIZES = {
"1024x1024": "square_hd",
"512x512": "square",
"1792x1024": "landscape_16_9",
"1024x1792": "portrait_16_9",
"1024x768": "landscape_4_3",
"768x1024": "portrait_4_3",
"1536x1024": "landscape_3_2",
"1024x1536": "portrait_3_2",
"576x1024": "portrait_16_9",
"1024x576": "landscape_16_9",
};
/**
* Handle image generation request
* @param {object} options
* @param {object} options.body - Request body
* @param {object} options.credentials - Provider credentials { apiKey, accessToken }
* @param {object} options.log - Logger
* @param {string} [options.resolvedProvider] - Pre-resolved provider ID (from route layer custom model resolution)
*/
export async function handleImageGeneration({
body,
credentials,
log,
resolvedProvider = null,
signal = null,
clientHeaders = null,
}) {
let provider, model;
if (resolvedProvider) {
// Provider was already resolved by the route layer (custom model from DB)
// Extract model name from the full "provider/model" string
provider = resolvedProvider;
const modelStr = body.model || "";
model = modelStr.startsWith(provider + "/") ? modelStr.slice(provider.length + 1) : modelStr;
} else {
// Standard path: resolve from built-in image registry
const parsed = parseImageModel(body.model);
provider = parsed.provider;
model = parsed.model;
}
if (!provider) {
return {
success: false,
status: 400,
error: `Invalid image model: ${body.model}. Use format: provider/model`,
};
}
const providerConfig = getImageProvider(provider);
// For custom models without a built-in provider config, use OpenAI-compatible handler
// with a synthetic config based on the provider's credentials
if (!providerConfig) {
if (!resolvedProvider) {
return {
success: false,
status: 400,
error: `Unknown image provider: ${provider}`,
};
}
// Custom model: use OpenAI-compatible format with provider's base URL
// The credentials were already resolved by the route layer
if (log) {
log.info("IMAGE", `Custom model ${provider}/${model} — using OpenAI-compatible handler`);
}
const syntheticConfig = {
id: provider,
// #3205: custom OpenAI-compatible nodes store their base URL in
// credentials.providerSpecificData.baseUrl (same as the chat path —
// see executors/default.ts:buildUrl / services/provider.ts:buildProviderUrl).
// Previously only the (always-absent) top-level credentials.baseUrl was
// read, so every custom image node fell back to the Gemini endpoint and
// returned "Please pass a valid API key".
baseUrl: resolveImageBaseUrl(
credentials,
`https://generativelanguage.googleapis.com/v1beta/openai/images/generations`
),
authType: "apikey",
authHeader: "bearer",
format: "openai",
};
return handleOpenAIImageGeneration({
model,
provider,
providerConfig: syntheticConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "gemini-image") {
return handleGeminiImageGeneration({ model, providerConfig, body, credentials, log });
}
if (providerConfig.format === "imagen3") {
return handleImagen3ImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "hyperbolic") {
return handleHyperbolicImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "huggingface-image") {
return handleHuggingFaceImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "fal-ai") {
return handleFalAIImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "stability-ai") {
return handleStabilityAIImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "black-forest-labs") {
return handleBlackForestLabsImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "recraft") {
return handleRecraftImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "topaz") {
return handleTopazImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "chatgpt-web") {
return handleChatGptWebImageGeneration({
model,
provider,
body,
credentials,
log,
signal,
clientHeaders,
});
}
if (providerConfig.format === "nanobanana") {
return handleNanoBananaImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "kie-image") {
return handleKieImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "sdwebui") {
return handleSDWebUIImageGeneration({ model, provider, providerConfig, body, log });
}
if (providerConfig.format === "comfyui") {
return handleComfyUIImageGeneration({ model, provider, providerConfig, body, log });
}
if (providerConfig.format === "codex-responses") {
return handleCodexImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "haiper-image") {
return handleHaiperImageGeneration({ model, provider, providerConfig, body, credentials, log });
}
if (providerConfig.format === "leonardo-image") {
return handleLeonardoImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "ideogram-image") {
return handleIdeogramImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
if (providerConfig.format === "nvidia-nim") {
return handleNvidiaNimImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
});
}
return handleOpenAIImageGeneration({ model, provider, providerConfig, body, credentials, log });
}
function normalizeKieImageResult(recordData: unknown): string[] {
const record = isJsonObject(recordData) ? recordData : {};
const data = isJsonObject(record.data) ? record.data : {};
const response = isJsonObject(data.response) ? data.response : {};
const resultJson = parseKieResultJson(recordData);
const urls = new Set<string>();
const add = (val: unknown) => {
if (typeof val === "string" && val.startsWith("http")) urls.add(val);
if (Array.isArray(val)) {
val.forEach((v) => {
if (typeof v === "string" && v.startsWith("http")) urls.add(v);
});
}
};
// Check resultJson (common in Market API)
add(resultJson?.resultUrls);
add(resultJson?.imageUrls);
add(resultJson?.resultUrl);
add(resultJson?.imageUrl);
// Check data.response (common in 4o-image API)
add(response.resultUrls);
add(response.resultUrl);
// Check direct data fields
add(data.resultImageUrls);
add(data.resultImageUrl);
add(data.url);
return Array.from(urls);
}
async function handleKieImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}: KieImageOptions) {
const startTime = Date.now();
const token = credentials?.apiKey || credentials?.accessToken;
const timeoutMs = normalizePositiveNumber(body.timeout_ms, 300000);
const pollIntervalMs = normalizePositiveNumber(body.poll_interval_ms, 2500);
const prompt = typeof body.prompt === "string" ? body.prompt : String(body.prompt ?? "");
const size = typeof body.size === "string" ? body.size : undefined;
if (!token) {
return saveImageErrorResult({
provider,
model,
status: 401,
startTime,
error: "KIE API key is required",
});
}
// Check if model is a Market model (unified API)
const fullRegistry = getImageProvider(provider);
const modelEntry = fullRegistry?.models?.find((m) => m.id === model);
const isMarket = modelEntry?.isMarket || model.includes("/");
const { imageUrl } = extractImageInputs(body);
let baseUrl = "";
let payload: Record<string, unknown> = {};
if (isMarket) {
// Unified Market API endpoint
baseUrl = `${providerConfig.baseUrl.replace(/\/$/, "")}/api/v1/jobs/createTask`;
const input: Record<string, unknown> = {
prompt,
aspect_ratio: mapImageSize(size, "1:1"),
};
if (imageUrl) {
input.image_url = imageUrl;
}
payload = {
model,
input,
};
} else {
// Legacy/Direct endpoint
const modelPath = model.replace("-t2i", "").replace("-i2i", "");
baseUrl = providerConfig.baseUrl.includes(model)
? providerConfig.baseUrl
: `https://api.kie.ai/api/v1/${modelPath}/generate`;
payload = {
prompt,
size: mapImageSize(size, "1:1"),
nVariants: body.n || 1,
};
}
if (log) {
const promptPreview = String(body.prompt ?? "").slice(0, 60);
log.info(
"IMAGE",
`${provider}/${model} (${isMarket ? "market" : "direct"}) | prompt: "${promptPreview}..."`
);
}
try {
const endpoint = isMarket ? "/api/v1/jobs/createTask" : new URL(baseUrl).pathname;
const createBaseUrl = isMarket ? providerConfig.baseUrl : baseUrl.replace(endpoint, "");
const createData = await kieExecutor.createTask({
baseUrl: createBaseUrl,
token,
payload,
endpoint,
});
const taskId = createData?.data?.taskId || createData?.taskId;
if (!taskId) {
const errorMessage =
createData?.msg ||
createData?.message ||
createData?.error ||
"KIE image generation did not return taskId";
if (log) {
log.error("IMAGE", `KIE createTask failed: ${JSON.stringify(createData)}`);
}
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: errorMessage,
requestBody: payload,
});
}
// Use statusUrl from providerConfig if available, fallback to dynamic derivation
const statusUrl = isMarket
? `${providerConfig.baseUrl.replace(/\/$/, "")}/api/v1/jobs/recordInfo`
: providerConfig.statusUrl && !providerConfig.statusUrl.includes("jobs/recordInfo")
? providerConfig.statusUrl
: baseUrl.replace(/\/generate$/, "/record-info");
const { data: recordData, state } = await kieExecutor.pollTask({
statusUrl,
taskId: String(taskId),
token,
timeoutMs,
pollIntervalMs,
});
if (state === "success") {
if (log) {
log.info("IMAGE", `KIE poll success for task ${taskId}`);
}
const urls = normalizeKieImageResult(recordData);
const images = urls.map((url: string) => ({ url, revised_prompt: prompt }));
return saveImageSuccessResult({
provider,
model,
startTime,
requestBody: payload,
responseBody: { images_count: images.length },
images,
});
}
const record = isJsonObject(recordData) ? recordData : {};
const recordDataBody = isJsonObject(record.data) ? record.data : {};
const errorMessage =
recordDataBody.errorMessage ||
recordDataBody.failMsg ||
record.msg ||
"KIE image task failed";
if (log) {
log.error("IMAGE", `KIE poll failed for task ${taskId}: ${JSON.stringify(recordData)}`);
}
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: String(errorMessage),
requestBody: payload,
});
} catch (err: unknown) {
return saveImageErrorResult({
provider,
model,
status: getKieErrorStatus(err, 502),
startTime,
error: `Image provider error: ${getKieErrorMessage(err, "KIE image generation failed")}`,
});
}
}
/**
* Handle Gemini-format image generation (Antigravity / Nano Banana)
* Uses Gemini's generateContent API with responseModalities: ["TEXT", "IMAGE"]
*/
async function handleGeminiImageGeneration({ model, providerConfig, body, credentials, log }) {
const startTime = Date.now();
const url = providerConfig.baseUrl;
const provider = "antigravity";
const credentialRecord = credentials || {};
const token = credentialRecord.accessToken || credentialRecord.apiKey;
const providerSpecificData = credentialRecord.providerSpecificData;
const providerSpecificProjectId =
providerSpecificData && typeof providerSpecificData === "object"
? (providerSpecificData as Record<string, unknown>).projectId
: null;
const credentialProjectId =
typeof credentialRecord.projectId === "string" ? credentialRecord.projectId.trim() : "";
const providerProjectId =
typeof providerSpecificProjectId === "string" ? providerSpecificProjectId.trim() : "";
const projectId = credentialProjectId || providerProjectId || null;
const candidateCount =
typeof body.n === "number" && Number.isFinite(body.n) && body.n > 0 ? Math.floor(body.n) : 1;
const promptText = typeof body.prompt === "string" ? body.prompt : String(body.prompt ?? "");
// Summarized request for call log
const logRequestBody = {
model: body.model,
prompt: promptText.slice(0, 200),
size: body.size || "default",
n: candidateCount,
};
if (!projectId || typeof projectId !== "string") {
return saveImageErrorResult({
provider,
model,
status: 400,
startTime,
error:
"Missing Google projectId for Antigravity account. Please reconnect OAuth in Providers so OmniRoute can fetch your Cloud Code project.",
requestBody: logRequestBody,
});
}
const antigravityBody = {
project: projectId,
requestId: `image_gen/${Date.now()}/${randomUUID()}/0`,
request: {
contents: [
{
role: "user",
parts: [{ text: promptText }],
},
],
generationConfig: {
candidateCount,
imageConfig: {
aspectRatio: normalizeImageAspectRatio(body.aspect_ratio, body.size),
},
},
},
model,
userAgent: getAntigravityEnvelopeUserAgent(credentialRecord),
requestType: "image_gen",
};
const headers = {
"Content-Type": "application/json",
Authorization: `Bearer ${token}`,
};
applyAntigravityClientProfileHeaders(headers, credentialRecord, antigravityBody);
delete headers["x-goog-user-project"];
if (log) {
const promptPreview = promptText.slice(0, 60);
log.info(
"IMAGE",
`antigravity/${model} (gemini) | prompt: "${promptPreview}..." | format: gemini-image`
);
}
try {
const response = await fetch(url, {
method: "POST",
headers,
body: JSON.stringify(antigravityBody),
});
if (!response.ok) {
const errorText = await response.text();
const safeError = sanitizeImageProviderError(errorText);
const safeErrorLog =
typeof safeError === "string" ? safeError : JSON.stringify(safeError ?? {});
if (log) {
log.error("IMAGE", `antigravity error ${response.status}: ${safeErrorLog.slice(0, 200)}`);
}
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: response.status,
model: `antigravity/${model}`,
provider,
duration: Date.now() - startTime,
error: safeErrorLog.slice(0, 500),
requestBody: logRequestBody,
}).catch(() => {});
return { success: false, status: response.status, error: safeError };
}
const data = await response.json();
const responseBody = data.response || data;
// Extract image data from Antigravity's wrapped Gemini response.
const images = [];
const candidates = responseBody.candidates || [];
for (const candidate of candidates) {
const parts = candidate.content?.parts || [];
for (const part of parts) {
if (part.inlineData) {
images.push({
b64_json: part.inlineData.data,
revised_prompt: parts.find((p) => p.text)?.text || promptText,
});
}
}
}
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 200,
model: `antigravity/${model}`,
provider,
duration: Date.now() - startTime,
tokens: { prompt_tokens: 0, completion_tokens: 0 },
requestBody: logRequestBody,
responseBody: { images_count: images.length },
}).catch(() => {});
return {
success: true,
data: {
created: Math.floor(Date.now() / 1000),
data: images,
},
};
} catch (err) {
if (log) {
log.error("IMAGE", `antigravity fetch error: ${err.message}`);
}
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 502,
model: `antigravity/${model}`,
provider,
duration: Date.now() - startTime,
error: err.message,
requestBody: logRequestBody,
}).catch(() => {});
return {
success: false,
status: 502,
error: `Image provider error: ${sanitizeErrorMessage((err as Error).message || err)}`,
};
}
}
/**
* Handle OpenAI-compatible image generation (standard providers + Nebius fallback)
*/
async function handleOpenAIImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}) {
const startTime = Date.now();
// Summarized request for call log
const logRequestBody = {
model: body.model,
prompt:
typeof body.prompt === "string"
? body.prompt.slice(0, 200)
: String(body.prompt ?? "").slice(0, 200),
size: body.size || "default",
n: body.n || 1,
quality: body.quality || undefined,
};
// Build upstream request (OpenAI-compatible format)
const upstreamBody: Record<string, unknown> = {
model: model,
prompt: body.prompt,
};
// Pass optional parameters
if (body.n !== undefined) upstreamBody.n = body.n;
if (body.size !== undefined) upstreamBody.size = body.size;
if (body.quality !== undefined) upstreamBody.quality = body.quality;
if (body.response_format !== undefined) upstreamBody.response_format = body.response_format;
if (body.style !== undefined) upstreamBody.style = body.style;
const { imageUrl } = extractImageInputs(body);
if (imageUrl && OPENAI_IMAGE_TO_IMAGE_MODELS.has(model)) {
upstreamBody.image_url = imageUrl;
}
// Build headers
const headers = {
"Content-Type": "application/json",
};
const token = credentials.apiKey || credentials.accessToken;
if (providerConfig.authHeader === "bearer") {
headers["Authorization"] = `Bearer ${token}`;
} else if (providerConfig.authHeader === "x-api-key") {
headers["x-api-key"] = token;
}
if (log) {
const promptPreview =
typeof body.prompt === "string"
? body.prompt.slice(0, 60)
: String(body.prompt ?? "").slice(0, 60);
log.info(
"IMAGE",
`${provider}/${model} | prompt: "${promptPreview}..." | size: ${body.size || "default"}`
);
}
const requestBody = JSON.stringify(upstreamBody);
// Try primary URL
let result = await fetchImageEndpoint(
providerConfig.baseUrl,
headers,
requestBody,
provider,
log
);
// Fallback for providers with fallbackUrl (e.g., Nebius)
if (
!result.success &&
providerConfig.fallbackUrl &&
[404, 410, 502, 503].includes(result.status)
) {
if (log) {
log.info("IMAGE", `${provider}: primary URL failed (${result.status}), trying fallback...`);
}
result = await fetchImageEndpoint(
providerConfig.fallbackUrl,
headers,
requestBody,
provider,
log
);
}
// Save call log after result is determined
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: result.status || (result.success ? 200 : 502),
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
tokens: { prompt_tokens: 0, completion_tokens: 0 },
error: result.success
? null
: typeof result.error === "string"
? result.error.slice(0, 500)
: null,
requestBody: logRequestBody,
responseBody: result.success ? { images_count: result.data?.data?.length || 0 } : null,
}).catch(() => {});
return result;
}
/**
* OpenAI-compatible image *edit* forwarder for custom providers (#3214 / #3215).
*
* Mirrors `handleOpenAIImageGeneration` but posts multipart/form-data to the node's
* `/images/edits` endpoint and returns the upstream OpenAI-compatible response. Kept
* separate from the chatgpt-web edit flow, which continues a saved conversation node
* rather than forwarding a stateless edit. The fetch helper leaves Content-Type unset so
* `fetch` derives the multipart boundary from the FormData body.
*/
export async function handleOpenAIImageEdit({
model,
provider,
credentials,
prompt,
imageBytes,
imageMime,
size,
responseFormat,
n = 1,
log,
}: {
model: string;
provider: string;
credentials:
| {
apiKey?: string;
accessToken?: string;
baseUrl?: unknown;
providerSpecificData?: { baseUrl?: unknown } | null;
}
| null
| undefined;
prompt: string;
imageBytes: Buffer;
imageMime?: string | null;
size?: string | null;
responseFormat?: string | null;
n?: number;
log?: { info: (tag: string, message: string) => void } | null;
}) {
const startTime = Date.now();
const url = resolveImageBaseUrl(
credentials,
`https://generativelanguage.googleapis.com/v1beta/openai/images/edits`,
"edits"
);
// Build the multipart body as a Buffer with an explicit boundary instead of a global
// `FormData`. In production `globalThis.fetch` is patched with node_modules/undici's fetch,
// whose `FormData` class differs from `globalThis.FormData` — passing a native FormData
// makes undici serialize it as the string "[object FormData]" (text/plain), dropping every
// field (including `model`, which reaches the upstream empty). A Buffer body is accepted
// verbatim by any fetch implementation. (#3273)
const boundary = `----OmniRouteImageEdit${randomUUID().replace(/-/g, "")}`;
const CRLF = "\r\n";
const partBuffers: Buffer[] = [];
const appendField = (name: string, value: string) => {
partBuffers.push(
Buffer.from(
`--${boundary}${CRLF}Content-Disposition: form-data; name="${name}"${CRLF}${CRLF}${value}${CRLF}`
)
);
};
appendField("model", model);
appendField("prompt", prompt);
if (size) appendField("size", size);
if (responseFormat) appendField("response_format", responseFormat);
appendField("n", String(n || 1));
partBuffers.push(
Buffer.from(
`--${boundary}${CRLF}Content-Disposition: form-data; name="image"; filename="image.png"${CRLF}` +
`Content-Type: ${imageMime || "image/png"}${CRLF}${CRLF}`
)
);
partBuffers.push(imageBytes);
partBuffers.push(Buffer.from(`${CRLF}--${boundary}--${CRLF}`));
const multipartBody = Buffer.concat(partBuffers);
const headers: Record<string, string> = {
"Content-Type": `multipart/form-data; boundary=${boundary}`,
};
const token = credentials?.apiKey || credentials?.accessToken;
if (token) headers["Authorization"] = `Bearer ${token}`;
if (log) {
log.info("IMAGE", `${provider}/${model} (edit) | prompt: "${prompt.slice(0, 60)}..." -> ${url}`);
}
const result = await fetchImageEndpoint(
url,
headers,
multipartBody as unknown as BodyInit,
provider,
log
);
saveCallLog({
method: "POST",
path: "/v1/images/edits",
status: result.status || (result.success ? 200 : 502),
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
tokens: { prompt_tokens: 0, completion_tokens: 0 },
error: result.success
? null
: typeof result.error === "string"
? result.error.slice(0, 500)
: null,
requestBody: { model, prompt: prompt.slice(0, 200), size: size || "default", n: n || 1 },
responseBody: result.success ? { images_count: result.data?.data?.length || 0 } : null,
}).catch(() => {});
return result;
}
export async function handleImageEdit({
provider,
model,
body,
imageBytes,
credentials,
log,
signal = null,
clientHeaders = null,
}: {
provider: string;
model: string;
body: Record<string, any>;
imageBytes: Buffer;
imageMime?: string; // accepted for symmetry with route layer; not used
credentials: any;
log: any;
signal?: AbortSignal | null;
clientHeaders?: Record<string, string> | null;
}) {
const startTime = Date.now();
const prompt = typeof body.prompt === "string" ? body.prompt.trim() : "";
if (!prompt) {
return saveImageErrorResult({
provider,
model,
status: 400,
startTime,
error: "Prompt is required for image edit",
});
}
if (!credentials?.apiKey) {
return saveImageErrorResult({
provider,
model,
status: 401,
startTime,
error: "ChatGPT Web credentials missing session cookie",
});
}
const imageHash = createHash("sha256").update(imageBytes).digest("hex");
const cached = findChatGptImageBySha256(imageHash);
const wantsBase64 = body.response_format === "b64_json";
const requestBody = {
model,
prompt: prompt.slice(0, 500),
size: body.size || undefined,
image_hash: imageHash.slice(0, 16),
image_bytes: imageBytes.length,
cached_match: Boolean(cached?.entry.context),
};
if (!cached?.entry.context) {
// chatgpt-web's image_gen tool can only edit an image when we continue
// the original conversation node. If we never generated this image (or
// its 30-minute TTL elapsed), there's no node to continue. Return a
// clear, actionable error — much better than silently spawning an
// unrelated image and confusing the user.
log?.warn?.(
"IMAGE",
`chatgpt-web edit: no cached match for sha256=${imageHash.slice(0, 16)} (bytes=${imageBytes.length}); returning 400`
);
return saveImageErrorResult({
provider,
model,
status: 400,
startTime,
error:
"chatgpt-web image edit only works for images recently generated through this OmniRoute instance " +
"(cache window: 30 minutes). Re-generate the image and try the edit immediately, or disable image-edit " +
"in your client to use plain chat-completion edit prompts instead.",
requestBody,
});
}
// Build a synthetic chat thread that surfaces the cached image URL on
// the assistant turn. The executor's parseOpenAIMessages picks up the
// URL, findCachedImageContext resolves it to {conversationId,
// parentMessageId}, and looksLikeImageEditRequest fires on the user
// prompt — together producing a continuation request that actually
// edits the saved image.
//
// The synthetic user prompt is anchored with both an edit verb AND an
// image-gen verb so the executor's heuristics fire regardless of what
// wording the caller used ("now make it brighter", "tweak this", ...):
// - looksLikeImageEditRequest: matches "edit" + "image" within 120 chars
// - looksLikeImageGenRequest: matches "generate" + "image" within 40 chars
// Either match alone would set forImageGen, but covering both is cheap
// insurance for prompts that don't fit common phrasings.
const messages: Array<{ role: string; content: string }> = [
{
role: "assistant",
// The base URL is irrelevant — only the path is parsed by
// CACHED_IMAGE_URL_RE in the executor's findCachedImageContext.
content: `![image](http://internal/v1/chatgpt-web/image/${cached.id})`,
},
{
role: "user",
content: `Edit the image and generate the new image: ${prompt}`,
},
];
const executor = new ChatGptWebExecutor();
const result = await executor.execute({
model,
body: { messages },
stream: false,
credentials,
signal,
log,
clientHeaders,
});
const responseText = await result.response.text();
if (result.response.status >= 400) {
return saveImageErrorResult({
provider,
model,
status: result.response.status,
startTime,
error: responseText,
requestBody,
});
}
let content = "";
try {
const json = JSON.parse(responseText);
content = String(json?.choices?.[0]?.message?.content || "");
} catch {
content = responseText;
}
const urls = extractMarkdownImageUrls(content);
if (urls.length === 0) {
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: `ChatGPT Web edit completed without returning image markdown: ${content.slice(0, 300)}`,
requestBody,
});
}
const images: Array<{ url?: string; b64_json?: string }> = [];
for (const url of urls) {
if (!wantsBase64) {
images.push({ url });
continue;
}
const id = url.match(CHATGPT_WEB_IMAGE_ID_RE)?.[1];
const cachedNew = id ? getChatGptImage(id) : null;
if (!cachedNew) {
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: "ChatGPT Web image bytes expired before b64_json conversion",
requestBody,
});
}
images.push({ b64_json: cachedNew.bytes.toString("base64") });
}
return saveImageSuccessResult({
provider,
model,
startTime,
requestBody,
responseBody: { images_count: images.length, edit_match: Boolean(cached?.entry.context) },
images,
});
}
async function handleFalAIImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}) {
const startTime = Date.now();
const token = credentials.apiKey || credentials.accessToken;
const { imageUrl, imageUrls } = extractImageInputs(body);
const upstreamBody: Record<string, unknown> = {
prompt: body.prompt,
sync_mode: body.sync_mode ?? true,
};
if (body.n !== undefined) upstreamBody.num_images = Number(body.n) || 1;
if (body.negative_prompt) upstreamBody.negative_prompt = body.negative_prompt;
if (body.seed !== undefined) upstreamBody.seed = body.seed;
if (body.style) upstreamBody.style = normalizeRecraftStyle(body.style);
const outputFormat = normalizeRequestedImageFormat(body, "png");
if (outputFormat) upstreamBody.output_format = outputFormat;
if (model.includes("flux-pro/v1.1") && !model.includes("ultra")) {
upstreamBody.image_size = mapFalImageSize(body.size, "landscape_4_3");
} else if (
model.includes("bytedance/") ||
model.includes("stable-diffusion") ||
model.includes("ideogram") ||
model.includes("recraft/v3")
) {
upstreamBody.image_size = mapFalImageSize(body.size, "square_hd");
} else {
upstreamBody.aspect_ratio = body.aspect_ratio || mapFalAspectRatio(body.size, "1:1");
}
if (body.quality === "hd" && model.includes("ultra")) {
upstreamBody.raw = true;
}
if (imageUrl && model.includes("flux-pro/v1.1-ultra")) {
upstreamBody.image_url = imageUrl;
}
if (imageUrls.length > 0 && model.includes("ideogram")) {
upstreamBody.image_urls = imageUrls;
}
if (log) {
const promptPreview = String(body.prompt ?? "").slice(0, 60);
log.info("IMAGE", `${provider}/${model} (fal-ai) | prompt: "${promptPreview}..."`);
}
try {
const response = await fetch(`${providerConfig.baseUrl.replace(/\/$/, "")}/${model}`, {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Key ${token}`,
},
body: JSON.stringify(upstreamBody),
});
if (!response.ok) {
const errorText = await response.text();
if (log)
log.error("IMAGE", `${provider} error ${response.status}: ${errorText.slice(0, 200)}`);
return saveImageErrorResult({
provider,
model,
status: response.status,
startTime,
error: errorText,
requestBody: upstreamBody,
});
}
const payload = await response.json();
const images = await normalizeProviderImagePayload(payload, body, log);
return saveImageSuccessResult({
provider,
model,
startTime,
requestBody: upstreamBody,
responseBody: { images_count: images.length },
created: payload.created,
images,
});
} catch (err) {
if (log) log.error("IMAGE", `${provider} fetch error: ${err.message}`);
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: `Image provider error: ${sanitizeErrorMessage((err as Error).message || err)}`,
});
}
}
async function handleStabilityAIImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}) {
const startTime = Date.now();
const token = credentials.apiKey || credentials.accessToken;
const endpoint = STABILITY_GENERATION_ENDPOINTS[model] || STABILITY_EDIT_ENDPOINTS[model];
if (!endpoint) {
return {
success: false,
status: 400,
error: `Unsupported Stability AI image model: ${model}`,
};
}
const { imageUrl, maskUrl } = extractImageInputs(body);
const upstreamBody: Record<string, unknown> = {
output_format:
model === "remove-background"
? normalizeRequestedImageFormat(body, "png", ["png", "webp"])
: normalizeRequestedImageFormat(body, "png"),
};
const formData = new FormData();
appendOptionalFormValue(formData, "output_format", upstreamBody.output_format);
if (body.prompt) {
upstreamBody.prompt = body.prompt;
appendOptionalFormValue(formData, "prompt", body.prompt);
}
if (body.negative_prompt) {
upstreamBody.negative_prompt = body.negative_prompt;
appendOptionalFormValue(formData, "negative_prompt", body.negative_prompt);
}
if (body.seed !== undefined) {
upstreamBody.seed = body.seed;
appendOptionalFormValue(formData, "seed", body.seed);
}
try {
if (STABILITY_GENERATION_ENDPOINTS[model]) {
if (model.startsWith("sd3.5")) {
upstreamBody.model = model;
appendOptionalFormValue(formData, "model", model);
}
if (imageUrl) {
const imageSource = await resolveImageSource(imageUrl);
upstreamBody.mode = "image-to-image";
appendOptionalFormValue(formData, "mode", "image-to-image");
upstreamBody.image = imageSource.base64;
appendImageFormValue(formData, "image", imageSource, "image");
if (body.strength !== undefined) {
upstreamBody.strength = body.strength;
appendOptionalFormValue(formData, "strength", body.strength);
}
} else {
upstreamBody.mode = "text-to-image";
appendOptionalFormValue(formData, "mode", "text-to-image");
}
if (!model.startsWith("sd3.5") || !imageUrl) {
const aspectRatio = body.aspect_ratio || mapImageSize(body.size);
upstreamBody.aspect_ratio = aspectRatio;
appendOptionalFormValue(formData, "aspect_ratio", aspectRatio);
}
if (body.style_preset) {
upstreamBody.style_preset = body.style_preset;
appendOptionalFormValue(formData, "style_preset", body.style_preset);
}
} else {
if (imageUrl) {
const imageSource = await resolveImageSource(imageUrl);
upstreamBody.image = imageSource.base64;
appendImageFormValue(formData, "image", imageSource, "image");
}
if (maskUrl && shouldIncludeStabilityMask(model)) {
const maskSource = await resolveImageSource(maskUrl);
upstreamBody.mask = maskSource.base64;
appendImageFormValue(formData, "mask", maskSource, "mask");
}
if (body.search_prompt) {
upstreamBody.search_prompt = body.search_prompt;
appendOptionalFormValue(formData, "search_prompt", body.search_prompt);
}
if (body.grow_mask !== undefined) {
upstreamBody.grow_mask = body.grow_mask;
appendOptionalFormValue(formData, "grow_mask", body.grow_mask);
}
if (body.control_strength !== undefined) {
upstreamBody.control_strength = body.control_strength;
appendOptionalFormValue(formData, "control_strength", body.control_strength);
}
if (body.creativity !== undefined) {
upstreamBody.creativity = body.creativity;
appendOptionalFormValue(formData, "creativity", body.creativity);
}
if (body.left !== undefined) {
upstreamBody.left = body.left;
appendOptionalFormValue(formData, "left", body.left);
}
if (body.right !== undefined) {
upstreamBody.right = body.right;
appendOptionalFormValue(formData, "right", body.right);
}
if (body.up !== undefined) {
upstreamBody.up = body.up;
appendOptionalFormValue(formData, "up", body.up);
}
if (body.down !== undefined) {
upstreamBody.down = body.down;
appendOptionalFormValue(formData, "down", body.down);
}
if (body.style_preset) {
upstreamBody.style_preset = body.style_preset;
appendOptionalFormValue(formData, "style_preset", body.style_preset);
}
if (STABILITY_CONTROL_MODELS.has(model) && !upstreamBody.prompt) {
upstreamBody.prompt = body.prompt || "";
appendOptionalFormValue(formData, "prompt", body.prompt || "");
}
}
if (log) {
const promptPreview = String(body.prompt ?? "").slice(0, 60);
log.info("IMAGE", `${provider}/${model} (stability-ai) | prompt: "${promptPreview}..."`);
}
const response = await fetch(`${providerConfig.baseUrl.replace(/\/$/, "")}${endpoint}`, {
method: "POST",
headers: {
Accept: "application/json",
Authorization: `Bearer ${token}`,
},
body: formData,
});
if (!response.ok) {
const errorText = await response.text();
if (log)
log.error("IMAGE", `${provider} error ${response.status}: ${errorText.slice(0, 200)}`);
return saveImageErrorResult({
provider,
model,
status: response.status,
startTime,
error: errorText,
requestBody: upstreamBody,
});
}
const contentType = response.headers.get("content-type") || "";
let payload;
if (contentType.includes("application/json")) {
payload = await response.json();
} else {
const buffer = Buffer.from(await response.arrayBuffer());
payload = { image: buffer.toString("base64") };
}
const images = await normalizeProviderImagePayload(payload, body, log);
return saveImageSuccessResult({
provider,
model,
startTime,
requestBody: upstreamBody,
responseBody: { images_count: images.length },
created: payload.created,
images,
});
} catch (err) {
if (log) log.error("IMAGE", `${provider} fetch error: ${err.message}`);
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: `Image provider error: ${sanitizeErrorMessage((err as Error).message || err)}`,
});
}
}
async function handleBlackForestLabsImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}) {
const startTime = Date.now();
const token = credentials.apiKey || credentials.accessToken;
const endpoint = BFL_MODEL_ENDPOINTS[model];
if (!endpoint) {
return {
success: false,
status: 400,
error: `Unsupported Black Forest Labs image model: ${model}`,
};
}
const { imageUrl, maskUrl } = extractImageInputs(body);
const upstreamBody: Record<string, unknown> = {
prompt: body.prompt,
output_format: normalizeRequestedImageFormat(body, "png"),
};
try {
if (BFL_EDIT_MODELS.has(model) && imageUrl) {
upstreamBody.input_image = (await resolveImageSource(imageUrl)).base64;
} else if (imageUrl && isHttpUrl(imageUrl)) {
upstreamBody.image_url = imageUrl;
}
if (maskUrl && (model === "flux-pro-1.0-fill" || model === "flux-kontext-pro")) {
upstreamBody.mask = (await resolveImageSource(maskUrl)).base64;
}
if (model === "flux-kontext-pro" || model === "flux-kontext-max") {
upstreamBody.aspect_ratio = body.aspect_ratio || mapImageSize(body.size);
} else if (typeof body.size === "string" && body.size.includes("x")) {
const { width, height } = parseSizeToDimensions(body.size, 1024);
upstreamBody.width = width;
upstreamBody.height = height;
}
if (body.seed !== undefined) upstreamBody.seed = body.seed;
if (body.n !== undefined && model.includes("ultra"))
upstreamBody.num_images = Number(body.n) || 1;
if (body.quality === "hd" && model.includes("ultra")) upstreamBody.raw = true;
if (body.left !== undefined) upstreamBody.left = body.left;
if (body.right !== undefined) upstreamBody.right = body.right;
if (body.top !== undefined) upstreamBody.top = body.top;
if (body.bottom !== undefined) upstreamBody.bottom = body.bottom;
if (body.steps !== undefined) upstreamBody.steps = body.steps;
if (body.guidance !== undefined) upstreamBody.guidance = body.guidance;
if (body.grow_mask !== undefined) upstreamBody.grow_mask = body.grow_mask;
if (body.safety_tolerance !== undefined) upstreamBody.safety_tolerance = body.safety_tolerance;
if (log) {
const promptPreview = String(body.prompt ?? "").slice(0, 60);
log.info("IMAGE", `${provider}/${model} (black-forest-labs) | prompt: "${promptPreview}..."`);
}
const response = await fetch(`${providerConfig.baseUrl.replace(/\/$/, "")}${endpoint}`, {
method: "POST",
headers: {
"Content-Type": "application/json",
Accept: "application/json",
"x-key": token,
},
body: JSON.stringify(upstreamBody),
});
if (!response.ok) {
const errorText = await response.text();
if (log)
log.error("IMAGE", `${provider} error ${response.status}: ${errorText.slice(0, 200)}`);
return saveImageErrorResult({
provider,
model,
status: response.status,
startTime,
error: errorText,
requestBody: upstreamBody,
});
}
const initialPayload = await response.json();
const finalPayload = initialPayload.polling_url
? await pollBlackForestLabsResult({
pollingUrl: initialPayload.polling_url,
token,
body,
log,
})
: initialPayload;
const images = await normalizeProviderImagePayload(finalPayload, body, log);
return saveImageSuccessResult({
provider,
model,
startTime,
requestBody: upstreamBody,
responseBody: { images_count: images.length },
created: finalPayload.created,
images,
});
} catch (err) {
if (log) log.error("IMAGE", `${provider} fetch error: ${err.message}`);
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: `Image provider error: ${sanitizeErrorMessage((err as Error).message || err)}`,
});
}
}
async function handleRecraftImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}) {
const startTime = Date.now();
const token = credentials.apiKey || credentials.accessToken;
const upstreamBody: Record<string, unknown> = {
model,
prompt: body.prompt,
};
if (body.n !== undefined) upstreamBody.n = body.n;
if (body.size !== undefined) upstreamBody.size = body.size;
if (body.response_format !== undefined) upstreamBody.response_format = body.response_format;
if (body.style !== undefined) upstreamBody.style = body.style;
if (log) {
const promptPreview = String(body.prompt ?? "").slice(0, 60);
log.info("IMAGE", `${provider}/${model} (recraft) | prompt: "${promptPreview}..."`);
}
try {
const response = await fetch(
`${providerConfig.baseUrl.replace(/\/$/, "")}/v1/images/generations`,
{
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${token}`,
},
body: JSON.stringify(upstreamBody),
}
);
if (!response.ok) {
const errorText = await response.text();
if (log)
log.error("IMAGE", `${provider} error ${response.status}: ${errorText.slice(0, 200)}`);
return saveImageErrorResult({
provider,
model,
status: response.status,
startTime,
error: errorText,
requestBody: upstreamBody,
});
}
const payload = await response.json();
const images = await normalizeProviderImagePayload(payload, body, log);
return saveImageSuccessResult({
provider,
model,
startTime,
requestBody: upstreamBody,
responseBody: { images_count: images.length },
created: payload.created,
images,
});
} catch (err) {
if (log) log.error("IMAGE", `${provider} fetch error: ${err.message}`);
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: `Image provider error: ${sanitizeErrorMessage((err as Error).message || err)}`,
});
}
}
async function handleTopazImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}) {
const startTime = Date.now();
const token = credentials.apiKey || credentials.accessToken;
const { imageUrl } = extractImageInputs(body);
if (!imageUrl) {
return {
success: false,
status: 400,
error: `Topaz model ${model} requires an input image`,
};
}
try {
const imageSource = await resolveImageSource(imageUrl);
const formData = new FormData();
const blob = new Blob([imageSource.buffer], { type: imageSource.contentType || "image/png" });
formData.append("image", blob, "image.png");
if (typeof body.size === "string" && body.size.includes("x")) {
const { width, height } = parseSizeToDimensions(body.size, 1024);
formData.append("output_width", String(width));
formData.append("output_height", String(height));
}
if (log) {
const promptPreview = String(body.prompt ?? "enhance image").slice(0, 60);
log.info("IMAGE", `${provider}/${model} (topaz) | prompt: "${promptPreview}..."`);
}
const response = await fetch(`${providerConfig.baseUrl.replace(/\/$/, "")}/image/v1/enhance`, {
method: "POST",
headers: {
Accept: "image/jpeg",
"X-API-Key": token,
},
body: formData,
});
if (!response.ok) {
const errorText = await response.text();
if (log)
log.error("IMAGE", `${provider} error ${response.status}: ${errorText.slice(0, 200)}`);
return saveImageErrorResult({
provider,
model,
status: response.status,
startTime,
error: errorText,
});
}
const contentType = response.headers.get("content-type") || "image/jpeg";
const buffer = Buffer.from(await response.arrayBuffer());
const base64 = buffer.toString("base64");
const wantsBase64 = body.response_format === "b64_json";
const images = [
wantsBase64
? { b64_json: base64, revised_prompt: body.prompt }
: { url: `data:${contentType};base64,${base64}`, revised_prompt: body.prompt },
];
return saveImageSuccessResult({
provider,
model,
startTime,
responseBody: { images_count: images.length },
images,
});
} catch (err) {
if (log) log.error("IMAGE", `${provider} fetch error: ${err.message}`);
return saveImageErrorResult({
provider,
model,
status: 502,
startTime,
error: `Image provider error: ${sanitizeErrorMessage((err as Error).message || err)}`,
});
}
}
async function pollBlackForestLabsResult({ pollingUrl, token, body, log }) {
const timeoutMs = normalizePositiveNumber(body.timeout_ms, 300000);
const pollIntervalMs = normalizePositiveNumber(body.poll_interval_ms, 1500);
const deadline = Date.now() + timeoutMs;
while (Date.now() < deadline) {
const response = await fetch(pollingUrl, {
method: "GET",
headers: {
"x-key": token,
},
});
if (!response.ok) {
const errorText = await response.text();
throw new Error(`BFL polling failed (${response.status}): ${errorText}`);
}
const payload = await response.json();
const status = payload?.status;
if (status === "Ready") {
return payload;
}
if (BFL_FAILURE_STATUSES.has(status)) {
throw new Error(`BFL image generation failed: ${status}`);
}
if (log) {
log.info("IMAGE", `black-forest-labs polling status: ${String(status || "Pending")}`);
}
await sleep(pollIntervalMs);
}
throw new Error(`BFL polling timed out after ${timeoutMs}ms`);
}
function extractImageInputs(body) {
const imageUrls = [];
const seen = new Set();
const pushCandidate = (candidate) => {
if (typeof candidate !== "string") return;
const trimmed = candidate.trim();
if (!trimmed || seen.has(trimmed)) return;
seen.add(trimmed);
imageUrls.push(trimmed);
};
pushCandidate(body?.image_url);
pushCandidate(body?.image);
if (Array.isArray(body?.imageUrls)) {
for (const candidate of body.imageUrls) pushCandidate(candidate);
}
if (Array.isArray(body?.image_urls)) {
for (const candidate of body.image_urls) pushCandidate(candidate);
}
if (Array.isArray(body?.messages)) {
for (const msg of body.messages) {
if (!Array.isArray(msg?.content)) continue;
for (const part of msg.content) {
if (part?.type === "image_url") {
pushCandidate(part?.image_url?.url);
}
}
}
}
return {
imageUrl: imageUrls[0] || null,
imageUrls,
maskUrl:
typeof body?.mask_url === "string"
? body.mask_url
: typeof body?.mask === "string"
? body.mask
: null,
};
}
async function resolveImageSource(source) {
if (typeof source !== "string" || source.trim().length === 0) {
throw new Error("Invalid image source");
}
const trimmed = source.trim();
const dataUriMatch = /^data:([^;]+);base64,(.+)$/i.exec(trimmed);
if (dataUriMatch) {
const [, contentType, base64] = dataUriMatch;
return {
buffer: Buffer.from(base64, "base64"),
base64,
contentType,
};
}
if (isHttpUrl(trimmed)) {
const remoteImage = await fetchRemoteImage(trimmed);
return {
buffer: remoteImage.buffer,
base64: remoteImage.buffer.toString("base64"),
contentType: remoteImage.contentType,
};
}
return {
buffer: Buffer.from(trimmed, "base64"),
base64: trimmed,
contentType: "application/octet-stream",
};
}
function parseSizeToDimensions(size, fallback = 1024) {
if (typeof size !== "string" || !size.includes("x")) {
return { width: fallback, height: fallback };
}
const [widthRaw, heightRaw] = size.split("x");
const width = Number(widthRaw);
const height = Number(heightRaw);
return {
width: Number.isFinite(width) && width > 0 ? width : fallback,
height: Number.isFinite(height) && height > 0 ? height : fallback,
};
}
function normalizeRequestedImageFormat(
body,
fallback = "png",
allowedFormats = ["jpeg", "png", "webp"]
) {
const formatCandidate =
typeof body?.output_format === "string"
? body.output_format.toLowerCase()
: typeof body?.response_format === "string" &&
!["url", "b64_json"].includes(body.response_format.toLowerCase())
? body.response_format.toLowerCase()
: fallback;
if (allowedFormats.includes(formatCandidate)) {
return formatCandidate;
}
return fallback;
}
function mapFalImageSize(size, fallback = "square_hd") {
if (typeof size !== "string") return fallback;
if (FAL_PRESET_SIZES[size]) return FAL_PRESET_SIZES[size];
if (size.includes("x")) {
const { width, height } = parseSizeToDimensions(size, 1024);
return { width, height };
}
return fallback;
}
function mapFalAspectRatio(size, fallback = "1:1") {
if (!size) return fallback;
return mapImageSize(size);
}
function normalizeRecraftStyle(style) {
if (style === "vivid") return "digital_illustration";
if (style === "natural") return "realistic_image";
return style;
}
function shouldIncludeStabilityMask(model) {
return new Set([
"inpaint",
"erase",
"search-and-replace",
"search-and-recolor",
"replace-background-and-relight",
]).has(model);
}
async function normalizeProviderImagePayload(payload, body, log) {
const candidates = [];
const pushCandidate = (value) => {
if (value === undefined || value === null) return;
candidates.push(value);
};
if (Array.isArray(payload?.data)) {
for (const item of payload.data) pushCandidate(item);
}
if (Array.isArray(payload?.images)) {
for (const item of payload.images) pushCandidate(item);
}
if (payload?.image) pushCandidate({ b64_json: payload.image });
if (payload?.url) pushCandidate({ url: payload.url });
if (payload?.sample) pushCandidate({ url: payload.sample });
if (payload?.result?.sample) pushCandidate({ url: payload.result.sample });
if (Array.isArray(payload?.result?.images)) {
for (const item of payload.result.images) pushCandidate(item);
}
const normalized = [];
for (const candidate of candidates) {
const item = await normalizeProviderImageCandidate(candidate, body);
if (item) normalized.push(item);
}
if (normalized.length === 0 && log) {
log.warn(
"IMAGE",
`Provider returned no recognizable image payload: ${JSON.stringify(payload).slice(0, 240)}`
);
}
return normalized;
}
async function normalizeProviderImageCandidate(candidate, body) {
const wantsBase64 = body?.response_format === "b64_json";
let url = null;
let b64 = null;
if (typeof candidate === "string") {
const dataUriMatch = /^data:[^;]+;base64,(.+)$/i.exec(candidate);
if (dataUriMatch) {
b64 = dataUriMatch[1];
} else if (isHttpUrl(candidate)) {
url = candidate;
} else {
b64 = candidate;
}
} else if (candidate && typeof candidate === "object") {
url =
firstString(candidate.url, candidate.image_url, candidate.sample, candidate.file_url) || null;
b64 =
firstString(candidate.b64_json, candidate.image, candidate.base64, candidate.data) || null;
}
if (wantsBase64 && !b64 && url) {
b64 = (await resolveImageSource(url)).base64;
}
if (url && !wantsBase64) {
return { url, revised_prompt: body?.prompt };
}
if (b64) {
return { b64_json: b64, revised_prompt: body?.prompt };
}
if (url) {
return { url, revised_prompt: body?.prompt };
}
return null;
}
function firstString(...values) {
for (const value of values) {
if (typeof value === "string" && value.length > 0) return value;
}
return null;
}
function isHttpUrl(value) {
return typeof value === "string" && /^https?:\/\//i.test(value);
}
/**
* Codex image generation — translate GPT-Image-style /v1/images/generations
* request into a /v1/responses call with the `image_generation` hosted tool,
* parse the SSE stream, and return the base64 PNG in OpenAI image response shape.
*
* Requires ChatGPT OAuth credentials (Codex provider connection). The hosted
* image_generation tool is only served upstream under ChatGPT auth; API-key
* users will receive a 400 from OpenAI.
*/
export function extractImageGenerationCalls(
sseText: string
): Array<{ b64: string; revisedPrompt: string | null }> {
const results: Array<{ b64: string; revisedPrompt: string | null }> = [];
const lines = String(sseText || "").split("\n");
for (const line of lines) {
const trimmed = line.trim();
if (!trimmed.startsWith("data:")) continue;
const payload = trimmed.slice(5).trim();
if (!payload || payload === "[DONE]") continue;
let evt: Record<string, unknown>;
try {
evt = JSON.parse(payload) as Record<string, unknown>;
} catch {
continue;
}
if (evt?.type !== "response.output_item.done") continue;
const item = evt.item as Record<string, unknown> | undefined;
if (!item || item.type !== "image_generation_call") continue;
const result = typeof item.result === "string" ? item.result : "";
if (!result) continue;
const revisedPrompt = typeof item.revised_prompt === "string" ? item.revised_prompt : null;
results.push({ b64: result, revisedPrompt });
}
return results;
}
// The image_generation hosted tool accepts { "auto" | "low" | "medium" | "high" }
// for `quality`. Legacy image clients often send "standard" / "hd". Map those values
// so OpenWebUI's quality dropdown doesn't silently get rejected upstream.
function mapLegacyImageQualityToImageTool(value: string): string {
const normalized = value.toLowerCase();
if (normalized === "standard") return "medium";
if (normalized === "hd") return "high";
return normalized;
}
async function handleCodexImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}) {
const startTime = Date.now();
const prompt = typeof body.prompt === "string" ? body.prompt : "";
if (!prompt.trim()) {
return saveImageErrorResult({
provider,
model,
status: 400,
startTime,
error: "Prompt is required for Codex image generation",
});
}
const requestedCount =
Number.isInteger(body.n) && (body.n as number) > 0 ? (body.n as number) : 1;
if (log && requestedCount > 1) {
log.warn(
"IMAGE",
`Codex hosted image_generation returns one image per call; requested n=${requestedCount} will fan out in parallel`
);
}
const token = credentials?.accessToken || credentials?.apiKey;
if (!token) {
return saveImageErrorResult({
provider,
model,
status: 401,
startTime,
error: "Codex credentials missing accessToken — reconnect the Codex provider",
});
}
const workspaceId =
credentials?.providerSpecificData &&
typeof credentials.providerSpecificData === "object" &&
!Array.isArray(credentials.providerSpecificData)
? (credentials.providerSpecificData as Record<string, unknown>).workspaceId
: undefined;
// Forward size/quality from the GPT-Image-style body into the hosted tool so
// OpenWebUI's size/quality selectors actually take effect. Everything else
// (model, n, background, moderation, output_compression) is left to the
// Codex backend's defaults — today that's `gpt-image-2`.
const toolConfig: Record<string, unknown> = { type: "image_generation", output_format: "png" };
if (typeof body.size === "string" && body.size.trim()) {
toolConfig.size = body.size.trim();
}
if (typeof body.quality === "string" && body.quality.trim()) {
toolConfig.quality = mapLegacyImageQualityToImageTool(body.quality.trim());
}
const upstreamBody: Record<string, unknown> = {
model,
instructions:
"You must call the image_generation tool exactly once to fulfill the user's request. Do not add narration.",
input: [
{
role: "user",
content: [{ type: "input_text", text: prompt }],
},
],
tools: [toolConfig],
stream: true,
store: false,
};
const headers: Record<string, string> = {
"Content-Type": "application/json",
Accept: "text/event-stream",
Authorization: `Bearer ${token}`,
Version: getCodexClientVersion(),
"User-Agent": getCodexUserAgent(),
originator: "codex_cli_rs",
};
if (typeof workspaceId === "string" && workspaceId) {
headers["chatgpt-account-id"] = workspaceId;
headers["session_id"] = workspaceId;
}
if (log) {
log.info(
"IMAGE",
`${provider}/${model} (codex-responses) | prompt: "${prompt.slice(0, 60)}..."`
);
}
const fetchOneImage = async () => {
let response: Response;
try {
response = await fetch(providerConfig.baseUrl, {
method: "POST",
headers,
body: JSON.stringify(upstreamBody),
});
} catch (err) {
if (log) log.error("IMAGE", `${provider} fetch error: ${(err as Error).message}`);
return {
ok: false as const,
error: {
provider,
model,
status: 502,
startTime,
error: `Image provider error: ${(err as Error).message}`,
requestBody: upstreamBody,
},
};
}
if (!response.ok) {
const errorText = await response.text();
if (log)
log.error("IMAGE", `${provider} error ${response.status}: ${errorText.slice(0, 200)}`);
return {
ok: false as const,
error: {
provider,
model,
status: response.status,
startTime,
error: errorText,
requestBody: upstreamBody,
},
};
}
const rawSSE = await response.text();
const items = extractImageGenerationCalls(rawSSE);
if (items.length === 0) {
return {
ok: false as const,
error: {
provider,
model,
status: 502,
startTime,
error:
"Codex completed without producing an image_generation_call — the model may have declined the tool",
requestBody: upstreamBody,
},
};
}
return { ok: true as const, items };
};
const imageResults = await Promise.all(
Array.from({ length: requestedCount }, () => fetchOneImage())
);
const collected: Array<{ b64_json: string; revised_prompt?: string }> = [];
for (const imageResult of imageResults) {
if (!imageResult.ok) return saveImageErrorResult(imageResult.error);
for (const item of imageResult.items) {
collected.push({
b64_json: item.b64,
...(item.revisedPrompt ? { revised_prompt: item.revisedPrompt } : {}),
});
}
}
const wantsUrl = body.response_format !== "b64_json";
const data = wantsUrl
? collected.map((item) => ({
url: `data:image/png;base64,${item.b64_json}`,
...(item.revised_prompt ? { revised_prompt: item.revised_prompt } : {}),
}))
: collected;
return saveImageSuccessResult({
provider,
model,
startTime,
requestBody: upstreamBody,
responseBody: { images_count: data.length },
images: data,
});
}
export function saveImageSuccessResult({
provider,
model,
startTime,
requestBody = null,
responseBody = null,
created = null,
images,
}) {
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 200,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
requestBody,
responseBody,
}).catch(() => {});
return {
success: true,
data: {
created: created || Math.floor(Date.now() / 1000),
data: images,
},
};
}
export function saveImageErrorResult({ provider, model, status, startTime, error, requestBody = null }) {
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
error: typeof error === "string" ? error.slice(0, 500) : String(error).slice(0, 500),
requestBody,
}).catch(() => {});
return {
success: false,
status,
error,
};
}
/**
* Fetch a single image endpoint and normalize response
*/
async function fetchImageEndpoint(url, headers, body, provider, log) {
try {
let response;
try {
response = await fetchWithTimeout(url, {
method: "POST",
headers,
body,
timeoutMs: getConfiguredTimeout(),
});
} catch (err: unknown) {
const isAbortError =
typeof err === "object" &&
err !== null &&
"name" in err &&
(err as { name?: unknown }).name === "AbortError";
if (err instanceof FetchTimeoutError || isAbortError) {
const message = err instanceof Error ? err.message : String(err);
if (log) {
log.error("IMAGE", `${provider} fetch error: ${message}`);
}
return {
success: false,
status: 504,
error: `Image provider error: ${sanitizeErrorMessage(message || err)}`,
};
}
throw err;
}
if (!response.ok) {
const errorText = await response.text();
if (log) {
log.error("IMAGE", `${provider} error ${response.status}: ${errorText.slice(0, 200)}`);
}
return {
success: false,
status: response.status,
error: errorText,
};
}
const data = await response.json();
// Normalize response to OpenAI format
return {
success: true,
data: {
created: data.created || Math.floor(Date.now() / 1000),
data: data.data || [],
},
};
} catch (err: unknown) {
const message = err instanceof Error ? err.message : String(err);
if (log) {
log.error("IMAGE", `${provider} fetch error: ${message}`);
}
return {
success: false,
status: 502,
error: `Image provider error: ${sanitizeErrorMessage(message || err)}`,
};
}
}
/**
* Handle Hyperbolic image generation
* Uses { model_name, prompt, height, width } and returns { images: [{ image: base64 }] }
*/
async function handleNanoBananaImageGeneration({
model,
provider,
providerConfig,
body,
credentials,
log,
}) {
const startTime = Date.now();
const token = credentials.apiKey || credentials.accessToken;
// Route to pro URL for "nanobanana-pro" model
const isPro = model === "nanobanana-pro";
const submitUrl = isPro && providerConfig.proUrl ? providerConfig.proUrl : providerConfig.baseUrl;
const statusUrl = providerConfig.statusUrl;
const aspectRatio =
typeof body.aspectRatio === "string"
? body.aspectRatio
: typeof body.aspect_ratio === "string"
? body.aspect_ratio
: mapImageSize(body.size);
let resolution =
typeof body.resolution === "string"
? body.resolution
: inferResolutionFromSize(body.size) || "1K";
if (body.quality === "hd" && resolution === "1K") {
resolution = "2K";
}
const upstreamBody = isPro
? {
prompt: body.prompt,
resolution,
aspectRatio,
...(Array.isArray(body.imageUrls) ? { imageUrls: body.imageUrls } : {}),
}
: {
prompt: body.prompt,
type:
Array.isArray(body.imageUrls) && body.imageUrls.length > 0
? "IMAGETOIAMGE"
: "TEXTTOIAMGE",
numImages: Number.isFinite(body.n) ? Math.max(1, Number(body.n)) : 1,
image_size: aspectRatio,
...(Array.isArray(body.imageUrls) ? { imageUrls: body.imageUrls } : {}),
};
if (log) {
const promptPreview = String(body.prompt ?? "").slice(0, 60);
log.info(
"IMAGE",
`${provider}/${model} (nanobanana ${isPro ? "pro" : "flash"}) | prompt: "${promptPreview}..."`
);
}
try {
const submitResp = await fetch(submitUrl, {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${token}`,
},
body: JSON.stringify(upstreamBody),
});
if (!submitResp.ok) {
const errorText = await submitResp.text();
if (log) {
log.error(
"IMAGE",
`${provider} submit error ${submitResp.status}: ${errorText.slice(0, 200)}`
);
}
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: submitResp.status,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
error: errorText.slice(0, 500),
}).catch(() => {});
return { success: false, status: submitResp.status, error: errorText };
}
const submitData = await submitResp.json();
// Backward compatibility: handle providers returning image payload synchronously
const hasSyncPayload =
Boolean(submitData?.image) ||
Array.isArray(submitData?.images) ||
Array.isArray(submitData?.data) ||
Boolean(submitData?.data?.[0]?.url) ||
Boolean(submitData?.data?.[0]?.b64_json);
if (hasSyncPayload) {
const syncResult = normalizeNanoBananaSyncPayload(submitData, body.prompt);
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 200,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
responseBody: { images_count: syncResult.data?.length || 0, mode: "sync" },
}).catch(() => {});
return {
success: true,
data: { created: Math.floor(Date.now() / 1000), data: syncResult.data },
};
}
const taskId = submitData?.data?.taskId || submitData?.taskId;
if (!taskId) {
const errorText = `NanoBanana submit did not return taskId: ${JSON.stringify(submitData).slice(0, 400)}`;
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 502,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
error: errorText,
}).catch(() => {});
return { success: false, status: 502, error: errorText };
}
if (!statusUrl) {
const errorText = "NanoBanana statusUrl is not configured";
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 500,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
error: errorText,
}).catch(() => {});
return { success: false, status: 500, error: errorText };
}
const timeoutMs = normalizePositiveNumber(
body.timeout_ms,
normalizePositiveNumber(process.env.NANOBANANA_POLL_TIMEOUT_MS, 120000)
);
const pollIntervalMs = normalizePositiveNumber(
body.poll_interval_ms,
normalizePositiveNumber(process.env.NANOBANANA_POLL_INTERVAL_MS, 2500)
);
let lastTaskData = null;
const deadline = Date.now() + timeoutMs;
while (Date.now() < deadline) {
const pollResp = await fetch(`${statusUrl}?taskId=${encodeURIComponent(taskId)}`, {
method: "GET",
headers: { Authorization: `Bearer ${token}` },
});
if (!pollResp.ok) {
const errorText = await pollResp.text();
if (log) {
log.error(
"IMAGE",
`${provider} poll error ${pollResp.status}: ${errorText.slice(0, 200)}`
);
}
return { success: false, status: pollResp.status, error: errorText };
}
const pollData = await pollResp.json();
const taskData = pollData?.data || pollData;
lastTaskData = taskData;
const successFlag = Number(taskData?.successFlag);
if (successFlag === 1) {
const normalized = await normalizeNanoBananaTaskResult(taskData, body, log);
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 200,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
responseBody: { images_count: normalized.length, mode: "async", taskId },
}).catch(() => {});
return {
success: true,
data: {
created: Math.floor(Date.now() / 1000),
data: normalized,
},
};
}
if (successFlag === 2 || successFlag === 3) {
const errorText =
taskData?.errorMessage || `NanoBanana task failed (successFlag=${String(successFlag)})`;
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 502,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
error: errorText.slice(0, 500),
responseBody: { taskId, successFlag, errorCode: taskData?.errorCode ?? null },
}).catch(() => {});
return { success: false, status: 502, error: errorText };
}
await sleep(pollIntervalMs);
}
const timeoutError = `NanoBanana task timeout after ${timeoutMs}ms (taskId=${taskId}, successFlag=${String(lastTaskData?.successFlag ?? "unknown")})`;
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 504,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
error: timeoutError,
responseBody: { taskId, lastSuccessFlag: lastTaskData?.successFlag ?? null },
}).catch(() => {});
return { success: false, status: 504, error: timeoutError };
} catch (err) {
if (log) log.error("IMAGE", `${provider} fetch error: ${err.message}`);
saveCallLog({
method: "POST",
path: "/v1/images/generations",
status: 502,
model: `${provider}/${model}`,
provider,
duration: Date.now() - startTime,
error: err.message,
}).catch(() => {});
return {
success: false,
status: 502,
error: `Image provider error: ${sanitizeErrorMessage((err as Error).message || err)}`,
};
}
}
function normalizeNanoBananaSyncPayload(data, prompt) {
const images = [];
if (data.image) {
images.push({ b64_json: data.image, revised_prompt: prompt });
} else if (Array.isArray(data.images)) {
for (const img of data.images) {
images.push({
b64_json: typeof img === "string" ? img : img?.image || img?.data,
revised_prompt: prompt,
});
}
} else if (Array.isArray(data.data)) {
for (const img of data.data) {
if (!img) continue;
images.push(img);
}
}
return { data: images.filter(Boolean) };
}
async function normalizeNanoBananaTaskResult(taskData, body, log) {
const response = taskData?.response || {};
const urlCandidates = [
response?.resultImageUrl,
response?.originImageUrl,
taskData?.resultImageUrl,
taskData?.originImageUrl,
].filter((v) => typeof v === "string" && v.length > 0);
if (Array.isArray(response?.resultImageUrls)) {
for (const u of response.resultImageUrls) {
if (typeof u === "string" && u.length > 0) urlCandidates.push(u);
}
}
const b64Candidates = [
response?.resultImageBase64,
response?.resultImage,
taskData?.resultImageBase64,
taskData?.resultImage,
].filter((v) => typeof v === "string" && v.length > 0);
if (Array.isArray(response?.resultImageBase64List)) {
for (const b64 of response.resultImageBase64List) {
if (typeof b64 === "string" && b64.length > 0) b64Candidates.push(b64);
}
}
const wantsBase64 = body.response_format === "b64_json";
if (wantsBase64) {
if (b64Candidates.length > 0) {
return b64Candidates.map((b64) => ({ b64_json: b64, revised_prompt: body.prompt }));
}
if (urlCandidates.length > 0) {
const firstUrl = urlCandidates[0];
const remoteImage = await fetchRemoteImage(firstUrl);
const base64 = remoteImage.buffer.toString("base64");
return [{ b64_json: base64, revised_prompt: body.prompt }];
}
}
if (urlCandidates.length > 0) {
return urlCandidates.map((url) => ({ url, revised_prompt: body.prompt }));
}
if (b64Candidates.length > 0) {
return b64Candidates.map((b64) => ({ b64_json: b64, revised_prompt: body.prompt }));
}
if (log) {
log.warn(
"IMAGE",
`NanoBanana task completed without image payload: ${JSON.stringify(taskData).slice(0, 240)}`
);
}
return [];
}
function inferResolutionFromSize(size) {
if (typeof size !== "string") return null;
const [wRaw, hRaw] = size.split("x");
const width = Number(wRaw);
const height = Number(hRaw);
if (!Number.isFinite(width) || !Number.isFinite(height) || width <= 0 || height <= 0) return null;
const longestSide = Math.max(width, height);
if (longestSide <= 1024) return "1K";
if (longestSide <= 2048) return "2K";
return "4K";
}
function normalizePositiveNumber(value, fallback) {
const n = Number(value);
if (!Number.isFinite(n) || n <= 0) return fallback;
return Math.floor(n);
}
/**
* Handle SD WebUI image generation (local, no auth)
* POST {baseUrl} with { prompt, negative_prompt, width, height, steps }
* Response: { images: ["base64..."] }
*/