Files
OmniRoute/docs/i18n/id/README.md
Diego Rodrigues de Sa e Souza 1bda6c15dc Release v3.8.44 (#5925)
* fix(install): add pnpm-workspace.yaml allowBuilds + pnpm.json for pnpm 11+

pnpm 11 introduced ERR_PNPM_IGNORED_BUILDS for native addon packages.
Without explicit allowBuilds approval, these packages silently skip build scripts
and OmniRoute fails to start with missing native modules.

Changes:
- pnpm-workspace.yaml: Set allowBuilds=true for all 13 native addon packages
  (@parcel/watcher, @swc/core, better-sqlite3, core-js, esbuild, keytar, koffi,
  libxmljs2, onnxruntime-node, protobufjs, sharp, tls-client-node, unrs-resolver)
- pnpm.json: Migrate onlyBuiltDependencies from package.json (deprecated field)
  to the new pnpm.json config file per pnpm 11 spec.

Tested on: pnpm 11.9.0, Node 24, Windows 11.

Fixes: pnpm install ERR_PNPM_IGNORED_BUILDS on fresh clone with pnpm 11.

* chore(release): open v3.8.44 development cycle

* test(security): parse Kimi Web URL host instead of substring match (CodeQL #689) (#5928)

Alert js/incomplete-url-substring-sanitization: the Kimi Web executor
test asserted result.url.includes("www.kimi.com"), which a hostile host
like www.kimi.com.evil.net would also satisfy. Parse the URL and assert
on the exact hostname (new URL(result.url).hostname === "www.kimi.com"),
which is both a stronger check and clears the CodeQL warning.

* refactor(translator): extract thinking-budget fitting from openai-to-claude (#5932)

Extract the thinking-budget fitting cluster (fitThinkingToMaxTokens +
private safeCapMaxOutputTokens + MIN_* constants) verbatim into the pure
leaf openai-to-claude/thinkingBudget.ts. Host re-exports fitThinkingToMaxTokens
so external importers keep working and imports it back for internal use.

Host 822 -> 738 LOC (under the 800 cap). No behavior change: byte-identical
bodies, public export set unchanged. Adds a split-guard test; all consumer
tests stay green (translator-openai-to-claude, strip-empty, minimax-m3, passthrough).

* chore(release): pipeline hardening — test-masking pre-flight gate + contributors/uncovered helpers (#5926)

* chore(ci): add test-masking PR-context gate to release-green pre-flight

Reproduce check:test-masking (vs origin/main) inside validate-release-green so
non-allowlisted net-assert reductions surface in the local pre-flight instead of
in a ~40-min CI layer on the release PR. run() now merges a per-gate opts.env so
GITHUB_BASE_REF reaches the child. HARD gate; skipped under --quick.

Context: v3.8.43 release cost 3 CI round-trips for PR-context gates (test-masking,
file-size, pr-evidence) that check:release-green did not reproduce locally.

* chore(release): add contributors generator + uncovered-commit reconciliation helpers

- scripts/release/gen-contributors.mjs: reproducible `### 🙌 Contributors` table for a
  CHANGELOG version (parenthetical-group parser → accurate per-PR attribution, noise-handle
  denylist). v3.8.43 shipped without the section (a real miss) because it was hand-built.
  npm run release:contributors <version> [--inject].
- scripts/release/list-uncovered-commits.mjs: lists commits since the last tag with no
  CHANGELOG bullet (v3.8.43 had 123/176 uncovered at reconciliation start). Advisory,
  maintainer-side. npm run release:uncovered.
- 20 unit tests (parenthetical attribution, noise exclusion, idempotent injection, coverage window).

* chore(quality): absorb web-cookie-providers-new file-size drift from #5928 (base-red on release/v3.8.44)

* refactor(translator): split openai-responses request translator into pure leaves (#5940)

Extract the shared pure primitives and the chat->Responses direction out of the
894-line openai-responses.ts request translator:
- openai-responses/helpers.ts: pure primitives (toRecord/toString/clampCallId/
  normalizeVerbosity/etc + markers/regexes/JsonRecord), zero host imports
- openai-responses/toResponses.ts: openaiToOpenAIResponsesRequest (chat->Responses),
  imports the helpers leaf

Host keeps openaiResponsesToOpenAIRequest (Responses->chat, imported by production)
plus both register() directions, and re-exports openaiToOpenAIResponsesRequest so
external importers (tests) keep working.

Host 894 -> 529 LOC (under the 800 cap). Verbatim bodies (multiset check: leaf A 54/54,
leaf B 294 lines, fn1 intact), public export set unchanged, leaves never import the host
(no cycle). Adds a split-guard test; all consumer tests stay green (responses-translation-fixes
37, verbosity 4, reasoning-effort 4, orphaned-tool-filter 8, empty-tool-name-loop 8,
headroom-responses-format 3).

* chore(ci): pr-evidence FAIL output tells you to push (body edit does not re-run the gate) (#5944)

ci.yml ignores the 'edited' event, so adding the Evidence block to the PR body after a
push does not re-run check:pr-evidence — you need another commit. The FAIL report now
says so, at the exact place someone sees the red check. + 5 unit tests (classification +
hint-on-fail / no-hint-on-pass). Decided against a separate edited-triggered workflow:
pr-evidence is not a required check (no ruleset gates it; release PRs merge UNSTABLE, not
BLOCKED), so the gap is cosmetic and the generate-release skill already puts Evidence in
the body before the first push.

* fix(providers): Perplexity Web emits real tool_calls in streaming mode (mirror chatgpt-web toolMode) (#5927) (#5937)

Perplexity Web (Pro/Max) only converted <tool>{...}</tool> text into
OpenAI tool_calls for non-streaming requests (hasTools && !stream).
Streaming requests -- the default for agentic coding clients -- got
the raw <tool> text as plain delta.content and never emitted a
tool_calls SSE delta, so clients could not execute tools.

Reuses the provider-agnostic buildToolModeResponse()/
toolCompletionToSseStream() helpers already shipped for chatgpt-web
(#5240): when tools are requested, buffer the full completion and
convert it into either a JSON completion or a terminal SSE replay
carrying delta.tool_calls + finish_reason: tool_calls, regardless of
the caller's stream flag. Extended buildToolModeResponse()'s idSeed
to be caller-supplied (default 'cgpt', perplexity-web passes 'pplx')
so tool_call ids stay provider-specific without duplicating the
helper. Non-tool streaming is unchanged (still lives token-by-token
via buildStreamingResponse).

* fix(discovery): resolve duplicate /v1 paths and redirect aborts (#5904)

Integrated into release/v3.8.44. Thanks @hamsa0x7 for diagnosing the doubled /v1 discovery path and the REDIRECT_BLOCKED probe-loop abort (#5899). De-scoped to the discovery fix (the #5903 session-affinity work is handled by #5943) and added Rule #18 regression guards.

* docs(changelog): record #5926 + #5944 (release-pipeline hardening) under v3.8.44 Maintenance (#5952)

* docs(claude): add Hard Rule #22 — cross-session safety (git stash + in-flight PRs) (#5955)

Integrated into release/v3.8.44 — Hard Rule #22 (cross-session safety).

* refactor(translator): extract pure helpers from response/openai-responses (#5949)

Extract the 5 stateless helpers (normalizeToolName, stripEmptyOptionalToolArgs,
normalizeOutputIndex, normalizeUpstreamFailure, extractResponsesReasoningSummaryText)
verbatim into the pure leaf openai-responses/pureHelpers.ts (no stream state, no host
import). Host imports them back and re-exports normalizeUpstreamFailure for external
importers (tests).

Host 1091 -> 1001 LOC. The stateful streaming core stays in the host (out of scope).
Byte-identical bodies (multiset 73/73), no cycle. Adds a split-guard; consumer tests
stay green (responses-translation-fixes 37, combo-param-validation-fallback-4519 5).

* docs(compression): document upstream sync policy for RTK/Caveman engines (#5830) (#5948)

Integrated into release/v3.8.44 — docs-only upstream sync policy for RTK/Caveman engines (closes #5830). All 7 checks green.

* fix(sse): strip ANSI/VT100 codes from gemini-cli stream frames (#5934)

Integrated into release/v3.8.44 — ReDoS-safe ANSI/VT100 strip for gemini-cli stream frames (port of upstream #2273, thanks @anki1kr). PR test green (5/5), file-size gate OK.

* fix(translator): strict Anthropic content-block compliance in antigravity→openai request (#5935)

Integrated into release/v3.8.44 — strict Anthropic content-block compliance in antigravity→openai (port upstream #2296). PR test green (9/9). UNSTABLE red is the pre-existing environmental setup-claude base-red (opencode-plugin dist not built in fast-path), not a regression from this PR.

* fix(mcp): auto-recover stale streamable HTTP sessions on initialize (#5957)

Integrated into release/v3.8.44 — MCP stale streamable-HTTP session auto-recovery (thanks @Chewji9875).

* fix(providers): validate v0 Platform API keys via chats endpoint (#5954)

Integrated into release/v3.8.44 — v0-vercel Platform API key validation (thanks @vittoroliveira-dev).

* fix(api): relax provider-scoped chat completion validation (#5907)

Integrated into release/v3.8.44 — relaxed provider-scoped chat validation + regression test (thanks @nickwizard).

* fix(providers): strip /v1 unconditionally to avoid /v1/v1/models fetch error (#5899) (#5920)

Integrated into release/v3.8.44 — unconditional /v1 strip in both models-discovery paths + regression test (thanks @anki1kr).

* fix(resilience): per-window is_exhausted + honor quota-exhaustion preflight for priority combos (#5923) (#5941)

Integrated into release/v3.8.44.

* fix(resilience): honor active codex session affinity over per-request reset-aware re-scoring (#5903) (#5943)

Integrated into release/v3.8.44.

* fix(thinking): only inject redacted_thinking replay block when tool_use present and thinking enabled (#5945) (#5953)

Integrated into release/v3.8.44.

* feat(providers): add ClinePass API-key provider (#5942)

Integrated into release/v3.8.44 — ClinePass API-key (BYOK) provider (port upstream 9router#2304, co-authored @adentdk). Validated locally: 16 clinepass tests green; fixed the APIKEY count 158→159 + translate-path golden snapshot (clinepass is a genuine new provider). Remaining UNSTABLE red is the pre-existing environmental setup-claude base-red (opencode-plugin dist not built in fast-path). Supersedes stub #5541.

* feat(api): add /v1/ocr endpoint (Mistral OCR) + Mistral moderation (#5950)

Integrated into release/v3.8.44 — /v1/ocr endpoint (Mistral OCR) + Mistral moderation (port upstream 9router#2064, co-authored @waguriagentic). Validated locally: 14 ocr-route tests + moderation/servicekind/endpoint-category suites green (CORS→Zod→handler + no-stack-leak assertion). Reds are inherited DRIFT only: cognitive-complexity ratchet (none from OCR files — pre-existing cycle drift, rebaselined at release) + environmental setup-claude base-red.

* fix(codex): convert chat json schema to responses text format (#5933)

Integrated into release/v3.8.44 — converts Chat Completions json_schema response_format → Responses API text.format on the Codex path, and preserves existing text.format through verbosity normalization. Base redirected main→release; the openai-responses.ts split that landed this cycle was reconciled by re-applying the delta onto openai-responses/toResponses.ts. Validated locally: 48 translator-openai-responses-req + 8 codex-verbosity tests green.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(providers): add Claude Sonnet 5 support across the model pipeline (#5833)

Integrated into release/v3.8.44 — wires claude-sonnet-5 end-to-end (registries, modelSpecs, pricing ×3, cost, Sonnet-family fallback, 1M-ctx, static models). Reconciled the add/add overlap with the already-merged #5796 (kept the PR's superset test with the family-fallback assertion). Validated locally: kiro-sonnet-5 + catalog + pricing/modelSpecs/fallback suites all green. Thanks @ggiak!

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(relay): gate bifrost auto routing by provider manifest (#5870)

Integrated into release/v3.8.44 — gates Bifrost auto-routing by the provider plugin manifest (only manifest-eligible providers reach the sidecar; ineligible/unknown fall back to the TS path with explicit reasons). Superset of #5869 (carries the full manifest + registry + docs). Resolved an integration-test conflict in favor of the release (which already subsumes this PR's readiness/removeDirWithRetry improvements). Validated locally: 4 provider-plugin-manifest + 11 relay-routing-backend tests green. Thanks @KooshaPari!

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* refactor(translator): extract pure message helpers from openai-to-kiro (852→751) (#5947)

* refactor(translator): extract pure message helpers from openai-to-kiro

Extract the pure tool/message helpers (parseToolInput, normalizeKiroToolSchema,
serializeToolResultContent) verbatim into the leaf openai-to-kiro/messageHelpers.ts.
The host imports them back for convertMessages. They were module-private, so the
public export set is unchanged (no re-export needed).

Host 852 -> 751 LOC. Byte-identical bodies (multiset 99/99), leaf has zero imports
(no cycle). Adds a split-guard; consumer tests stay green (translator-openai-to-kiro 33,
translator-ai-sdk-image-parts 3).

* chore: re-trigger CI (stuck runner on 2/2 shard)

* refactor(executors): extract pure prompt + composer helpers from cursor (#5960)

Extract two pure clusters from the cursor executor into sibling leaves:
- cursor/prompt.ts: isRecordLike + toolChoiceDirectiveLine + buildCursorOutputConstraints
- cursor/composer.ts: composer thinking-as-content decoding (isComposerModel,
  visibleComposerContentFromThinking, composerReasoningRemainder + markers)

Host imports both back for internal use and re-exports the 3 composer helpers for
external importers (tests). Host 1576 -> 1451 LOC. Byte-identical bodies (verbatim
multiset prompt 65/65, composer 32/32), leaves have zero imports (no cycle). Adds a
split-guard; consumer tests stay green (cursor-composer-thinking, cursor-streaming,
cursor-agent-tool-calls, translator-openai-to-cursor, cursor-agent-system-prompt).

* refactor(executors): extract pure SSE-collect parsing from antigravity (#5962)

Extract the pure SSE-payload -> collected-stream parser (AntigravityCollectedStream,
stripZeroWidth, parseAntigravityTextualToolCall, addAntigravityTextualToolCall,
processAntigravitySSEPayload/Text, flushAntigravitySSEText) verbatim into the leaf
antigravity/sseCollect.ts. Host imports the helpers it uses and re-exports
processAntigravitySSEPayload for external importers (tests).

Host 1812 -> 1671 LOC. Byte-identical bodies (verbatim multiset 135/135), leaf does
not import the host (no cycle). Credit/quota state, auth, and HTTP dispatch untouched.
Adds a split-guard; consumer tests stay green (executor-agy 8, executor-antigravity 26,
antigravity-sse-collect-socket-release, copilot-agent-antigravity-parity 6).

* refactor(executors): extract pure model maps + resolvers from chatgpt-web (#5967)

Extract the static model maps (MODEL_MAP, MODEL_FORCED_EFFORT, THINKING_CAPABLE_SLUGS)
and the pure thinking-effort resolvers (isThinkingCapableModel, normalizeThinkingEffort,
resolveThinkingEffort, ResolvedChatGptModel, resolveChatGptModel) verbatim into the pure
leaf chatgpt-web/models.ts. Host imports the two resolvers it uses back.

Host 3205 -> 3076 LOC. Byte-identical bodies (verbatim multiset 120/120), leaf has zero
imports (no cycle). Auth/PoW/session/HTTP dispatch and all module caches untouched.
Adds a split-guard; consumer tests stay green (chatgpt-web 86, chatgpt-web-tools-5240 4,
chatgpt-web-sha3-boringssl-5531 5).

* refactor(executors): decompose grok-web into pure tool/markup leaves (#5994)

Extract the pure OpenAI<->Grok tool-translation, native-tool mapping, markup cleanup,
and NDJSON stream types out of the 1872-line grok-web executor into 4 sibling leaves:
- grok-web/types.ts: GrokStreamResponse/GrokStreamEvent (stream types)
- grok-web/tool-bridge.ts: OpenAI<->Grok tool translation + registry + classifiers
- grok-web/native-tools.ts: native-tool selection/scoring + native->OpenAI mapping
- grok-web/text-cleanup.ts: Grok markup stripping + GrokMarkupFilter

Layered, acyclic: types <- tool-bridge <- native-tools; text-cleanup <- types; host
imports the leaves. All symbols module-private (no host re-export). Host 1872 -> 887 LOC.
Byte-identical bodies (verbatim per-leaf), no cycle, all new leaves <= 800 cap
(tool-bridge split at line 753 to stay under). Auth/cookie/TLS/HTTP dispatch untouched.
Adds a split-guard; consumer tests stay green (grok-web 62, grok-cli-oauth 15,
grok-cli-strip-params 2).

* refactor(executors): extract pure quota parsing from codex (#5999)

Extract the pure Codex quota-snapshot parsing + reset/cooldown scheduling
(CodexQuotaSnapshot, parseCodexQuotaHeaders, getCodexResetTime,
getCodexDualWindowCooldownMs) verbatim into the leaf codex/quota.ts. Host re-exports
the 4 symbols so handlers/chatCore/codexQuota.ts + tests keep resolving.

Host 1539 -> 1427 LOC. Byte-identical bodies (verbatim 98/98), leaf has zero imports
(only Date, no cycle). WS transport, auth, HTTP dispatch untouched. Adds a split-guard;
consumer tests stay green (executor-codex 40, codex-quota-fetcher 7, chatcore-codex-quota 5).

* refactor(executors): extract pure stream formatters from deepseek-web (#6000)

Extract the pure content/citation formatters (isThinkingModel, isSearchModel,
cleanDeepSeekToken, formatStreamContent, DeepSeekSearchResult, appendSearchCitations)
verbatim into the leaf deepseek-web/stream-format.ts. Host imports the 5 it uses back
into transformSSE/collectSSEContent (cleanDeepSeekToken stays internal to the leaf).

Host 1147 -> 1108 LOC. Byte-identical bodies (verbatim 34/34), leaf has zero imports
(no cycle), all module-private (no re-export). PoW/auth/token-cache/HTTP dispatch
untouched. Adds a split-guard; consumer tests stay green (deepseek-web 35,
deepseek-web-rolling-window-2942 5, deepseek-web-tools-execute 3).

* refactor(api): add validatedJsonBody helper (salvage #5075) (#5931)

Fuses JSON body parsing + Zod validation into a single call that returns
either type-narrowed data or a ready-to-return 400 NextResponse with the
standard error envelope. Salvaged as the Tier 1 portable helper from the
closed refactor PR #5075; the bulk route migration is intentionally not
ported. Adds a focused 6-case regression test.

Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>

* feat(qoder): drive PAT auth via qodercli, add dashboard quota, fix connection display (#5816)

Integrated into release/v3.8.44 — Qoder PAT auth via qodercli binary + dashboard quota + dual-auth connection fix. Thanks @AgentKiller45 (co-author @judy459)!

Validated locally (release-green on its own merits): lint 0, typecheck:core 0, 104 qoder/usage/UI tests green, file-size gate OK (owner-approved qoderCli.ts baseline-freeze 666→989), env-doc-sync fixed (documented QODER_CLI_CONFIG_DIR).

The 2 remaining CI reds are INHERITED base-reds, not caused by this PR: (1) LEDGER-4 minimax-m3 supportsVision (minimax-m3 base + cline-pass/minimax-m3 from the already-merged #5942); (2) mutation-test-coverage missing 3 tests in stryker.conf (#5903/#5942/#5923). Both cleaned up separately.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(providers): minimax-m3 supportsVision (LEDGER-4) + stryker tap.testFiles drift (#6012)

Release-green cleanup — clears LEDGER-4 minimax-m3 supportsVision + stryker tap.testFiles drift base-reds. Validated locally.

* fix(registry): flag cline-pass/minimax-m3 as multimodal (supportsVision) (#6003)

The cline-pass provider's minimax-m3 entry was missing supportsVision, breaking the
LEDGER-4 registry-consistency test (all minimax-m3 entries must set supportsVision to
match lite.ts — minimax-m3 is multimodal). Every other minimax-m3 registry entry
(trae, bazaarlink, cline, ollama-cloud, ...) already sets it. This was a base-red on
release/v3.8.44 inherited by every open PR.

Validated by the existing failing-then-passing guard tests/unit/review-reviews-v3814-fixes.test.ts
(LEDGER-4).

* refactor(executors): extract pure payload construction from claude-web (#6006)

Extract the pure Claude-web payload types + transforms + default tools/style
(ClaudeWebRequestPayload, ClaudeWebStreamChunk, DEFAULT_CLAUDE_MODEL,
generateMessageUUIDs, getDefaultTools, getDefaultPersonalizedStyle, transformToClaude,
transformFromClaude) verbatim into the leaf claude-web/payload.ts. Host imports the 3
it uses back (ClaudeWebRequestPayload type + the two transforms).

Host 1056 -> 835 LOC. Byte-identical bodies (verbatim 149/149), leaf imports only
randomUUID (no host import, no cycle), all module-private (no re-export). Cookie/auth/
Turnstile/TLS/HTTP dispatch untouched. Adds a split-guard; consumer tests stay green
(claude-web 13, claude-web-auto-refresh 6).

* refactor(executors): extract pure upstream-header helpers from base (#6008)

Extract the pure upstream-header helpers (mergeUpstreamExtraHeaders, getCustomUserAgent,
setUserAgentHeader, applyConfiguredUserAgent, isOpenAICompatibleEndpoint,
stripStainlessHeadersForOpenAICompat) verbatim into the leaf base/headers.ts. base.ts is
imported by ~18 executors, so the host re-exports all 6 to keep those import paths intact;
it also imports the 4 it uses internally in the BaseExecutor class. The trivial JsonRecord
type alias is redefined locally in the leaf to avoid a base<->leaf cycle.

Host 1539 -> 1451 LOC. Byte-identical bodies (verbatim 78/78), leaf does not import the
host (no cycle). typecheck:core validates all base importers still resolve via the
re-export. Adds a split-guard; consumer tests stay green (executor-base-utils 22,
executor-default-base 49, executor-strip-stainless-openai-compat 6, plus executor sanity
via typecheck).

* refactor(executors): extract pure wire protocol from perplexity-web (#6014)

Extract the pure Perplexity wire protocol (consts, SSE stream types, SSE parsing,
OpenAI<->Perplexity message translation, request/query builders, content extraction,
sseChunk) verbatim into the leaf perplexity-web/protocol.ts. Host imports back the 10
symbols it uses; everything module-private (no re-export). Session cache, TLS fetch,
auth, and the executor class stay in the host.

Host 1028 -> 534 LOC. Byte-identical bodies (verbatim), leaf imports only randomUUID
(no host import, no cycle). Adds a split-guard; consumer tests stay green
(perplexity-web 26, streaming-tools-5927 2, tls-client 6, key-validation-models 2).

* refactor(executors): extract pure URL normalizers from default (#6015)

Extract the pure per-provider chat-URL normalizers (normalizeBailianMessagesUrl,
normalizeDataRobotChatUrl, normalizeAzureAiChatUrl, normalizeWatsonxChatUrl,
normalizeOciChatUrl, normalizeSapChatUrl, normalizeXiaomiMimoChatUrl,
normalizeOpenAIChatUrl, getOpenRouterConnectionPreset) verbatim into the leaf
default/urlNormalizers.ts. Host imports them back into buildUrl/transformRequest; the
now-dead build*ChatUrl/normalizeBaseUrl imports move to the leaf. All module-private
(no re-export).

Host 864 -> 815 LOC (shrunk below its frozen baseline). Byte-identical bodies (verbatim
45/45), leaf does not import the host (no cycle). buildHeaders/execute/auth untouched.
Adds a split-guard; consumer tests stay green (executor-default-base 49,
anthropic-compatible-bearer 3, strip-client-metadata 3).

* feat(webfetch): support self-hosted FireCrawl instances (#5793)

Integrated into release/v3.8.44 — self-hosted FireCrawl support (FIRECRAWL_BASE_URL/FIRECRAWL_TIMEOUT_MS). Re-cut clean onto the release tip (branch was fossilized from a pre-v3.8.40 snapshot). Validated: 4 firecrawl tests green, env-doc-sync + docs-sync pass. UNSTABLE red is the inherited environmental setup-claude base-red.

* feat(xai): register XaiExecutor with reasoning-effort suffix parsing (#5800)

Integrated into release/v3.8.44 — XaiExecutor with reasoning-effort suffix parsing. Re-cut clean onto the release tip (branch was fossilized). Validated: 6 xai-executor tests green, provider-consistency OK, typecheck:core 0 errors, env-doc-sync in sync. UNSTABLE red is the inherited environmental setup-claude base-red.

* feat(discovery): Phase 2 — reporter, /api/discovery/* routes (strict loopback-only) + dashboard UI (#5939)

* feat(discovery): Phase 2 reporter — discoveryResults DB module + service wiring

Adds src/lib/db/discoveryResults.ts (CRUD over the discovery_results table
from migration 074) and wires the opt-in discovery service to persist and read
findings through it: persistDiscoveryResult / getDiscoveryResults /
getDiscoveryResultById / markVerified / deleteDiscoveryResult, with
(provider, method, endpoint) upsert de-duplication. Re-exported from localDb.

The service stays opt-in / default-off. The /api/discovery/* routes and the
dashboard UI tab are intentionally deferred to Phase 2b — they need the
local-only enforcement model (Hard Rules #15/#17 territory) decided first.

TDD: tests/unit/db/discovery-results.test.ts (8 cases, DB + service delegation),
isolated DATA_DIR with resetDbInstance cleanup.

* feat(discovery): Phase 2b — /api/discovery/* routes (strict loopback-only)

Adds the discovery HTTP surface on top of the reporter DB module:
  GET    /api/discovery/results            list findings (optional ?providerId)
  GET    /api/discovery/results/:id        one finding (404 if absent)
  DELETE /api/discovery/results/:id        delete a finding
  POST   /api/discovery/scan               scan a provider + persist findings
  POST   /api/discovery/verify/:id         mark a finding verified

Authorization: strict loopback-only. "/api/discovery/" is added to
LOCAL_ONLY_API_PREFIXES so the central authz pipeline (proxy.ts →
runAuthzPipeline → managementPolicy) rejects non-loopback callers with a 403
LOCAL_ONLY before any handler runs. It is deliberately NOT in
LOCAL_ONLY_MANAGE_SCOPE_BYPASS_PREFIXES — no remote manage-scope bypass —
because POST /scan issues outbound probes to provider endpoints (SSRF-adjacent)
and must never be tunnel-reachable. Handlers also call requireManagementAuth
(defense in depth) and return sanitized errors via createErrorResponse.

Tests:
- tests/unit/authz/discovery-routes-local-only.test.ts (8) — security guard:
  isLocalOnlyPath true + not manage-scope-bypassable for all four paths.
- tests/unit/api/discovery-routes.test.ts (6) — handler integration over an
  isolated DATA_DIR: list/filter, by-id 200/404/400, scan persist + 400 on
  empty/malformed body, verify 200/404, delete 200/404, no stack-trace leak.

* feat(discovery): Phase 2c — dashboard UI tab (Tools → Discovery)

Adds the /dashboard/discovery page (DiscoveryPageClient) that consumes the
Phase 2b /api/discovery/* routes: scan a provider, list findings, verify or
delete them. Registered in the sidebar under the Tools group (icon
travel_explore) and given a "discovery" i18n namespace + sidebar keys in
en.json (other locales fall back to en via next-intl until synced — the
locale files are in a pre-existing coverage deficit unrelated to this change).

Registers the UI test path in vitest.config.ts (advisory ui suite).

Tests: src/app/(dashboard)/dashboard/discovery/__tests__/DiscoveryPageClient.test.tsx
(3 cases: loads+renders results, empty state, fetches /api/discovery/results on
mount; stable useTranslations mock to avoid the fetch-loop). NOTE: the ui vitest
suite cannot run in this workspace — @testing-library/dom (a @testing-library/
react peer dep) is absent from node_modules, which fails ALL existing ui tests
equally; the test runs in CI. Component verified locally via typecheck + lint.

* test(discovery): register discovery-routes-local-only in stryker tap.testFiles

The mutation-test-coverage gate (--strict) flags any unit test covering a
mutated module that isn't listed in stryker.conf.json tap.testFiles. This PR's
tests/unit/authz/discovery-routes-local-only.test.ts covers src/server/authz/
routeGuard.ts (a mutated module, which this PR edits by adding the
/api/discovery/ local-only prefix), so it must be registered for its mutant
kills to count. No behavior change.

* refactor(discovery): split DiscoveryPageClient to satisfy max-lines-per-function

The complexity ratchet (max-lines-per-function: 80) flagged the single
184-line DiscoveryPageClient function (+1 over baseline). Extract the data
layer into two hooks (useDiscoveryResults for list/loading/feedback,
useDiscoveryActions for scan/verify/delete), a shared callApi helper, and two
presentational sub-components (DiscoveryScanForm, DiscoveryResultCard). Every
function is now under the 80-line ceiling; complexity gate back to baseline
1995. No behavior change — same exported component, same endpoints, same props.

* test(sidebar): include discovery in omni-proxy item-order snapshot

Adding the Discovery item to the Tools group (this PR's sidebar entry) extends
the ordered omni-proxy section list. Update the exact-match deepEqual snapshot
in sidebar-visibility.test.ts to include "discovery" in its position (after
traffic-inspector). The assertion stays exact — this reflects the intentional
new item, it does not weaken the check.

* docs(changelog): restore release bullets eaten by merge auto-resolve; re-add discovery bullet additively

* chore(quality): bump testFrozen for translator-openai-responses-req.test.ts (1097 -> 1172)

Base-red inherited from #5933, which grew the test file to 1171 lines
(Hard Rule #18 regression tests) without adjusting the frozen cap. The
release tip itself fails check:file-size; this unblocks every PR into
release/v3.8.44. File untouched by this PR.

* chore(quality): restore stryker tap.testFiles entries eaten by merge auto-resolve

The merge of origin/release/v3.8.44 silently dropped the 3 entries added
on the release side (#5903, clinepass, #5923). Took the release version
verbatim and re-added only this PR's entry (discovery-routes-local-only)
in alphabetical order. check:mutation-test-coverage green locally.

* chore(quality): reconcile inherited v3.8.44 merge-burst drift + include discovery in tools-group order test

- complexity 1995->2003 and cognitive 856->859: both measure IDENTICAL on
  the pristine release tip (3a3d618fe) and this PR's merged HEAD — the PR
  is complexity-net-zero; drift is from the 2026-07-02 merge burst
  (notes added to both baselines, same family as prior reconciliations).
- sidebar-tools-group.test.ts: append 'discovery' to the expected
  TOOLS_GROUP order — the intentional new sidebar item this PR adds
  (same expected-value update already made in sidebar-visibility.test.ts).

* feat(providers): custom icon URL for compatible provider nodes (#5815)

Integrated into release/v3.8.44 — custom icon URL for compatible provider nodes (DB migration 113 + nodes.ts + Zod schema + API routes + catalog + ProviderIcon UI). Re-cut onto the release tip (branch was fossilized ~13 real files); reconciled icon_url into the release's evolved nodes.ts/routes via 3-way. Validated: 14 backend + 5 frontend(vitest) + 24 page-utils tests green, typecheck:core 0, provider-consistency OK, file-size/env-doc-sync pass. UNSTABLE red is the inherited environmental setup-claude base-red.

* feat(api): add /v1/audio/translations endpoint (#5809)

Integrated into release/v3.8.44 — /v1/audio/translations endpoint (Whisper-style audio translation) + audioTranslation handler + translation providers in audioRegistry. Re-cut clean onto the release tip (branch was fossilized). Validated: 8 route tests (incl. no-stack-leak), typecheck:core 0, route-guard-membership OK, docs gates pass. UNSTABLE red is the inherited environmental setup-claude base-red.

* feat(dashboard): wildcard-CORS runtime warning + CORS security doc (#5602) (#5759)

Integrated into release/v3.8.44 — wildcard-CORS runtime warning banner + docs/security/CORS.md security guide (#5602). Re-cut clean onto the release tip (branch was fossilized). Validated: 20+9 backend + 2 banner(vitest) tests green, typecheck:core 0, docs-sync/symbols/fabricated/doc-links pass. UNSTABLE red is the inherited environmental setup-claude base-red.

* refactor(executors): extract pure JSONL stream translation from huggingchat (#6016)

Extract the pure JSONL->OpenAI-SSE translation (sseChunk, parseJsonlLine,
streamJsonlToOpenAi, readJsonlResponse) verbatim into the leaf huggingchat/jsonlStream.ts.
They consume a passed-in ReadableStream (no fetch/network/state). Host imports back the
two it uses; all module-private (no re-export).

Host 812 -> 594 LOC. Byte-identical bodies (verbatim), leaf has zero imports (no cycle).
Cookie/auth/multipart/execute untouched. Adds a split-guard; consumer tests stay green
(executor-huggingchat 6, huggingchat-model-catalog 3).

* refactor(executors): extract pure Meta AI response parser from muse-spark-web (#6017)

Extract the pure Meta AI SSE/JSON response parsing + content/reasoning/error extraction
(parseMetaSseFrames, readMetaJsonPayloads, collect*/extract*/classify* helpers,
parseMetaAiResponseText, isRecord, the reasoning/renderer key arrays, MetaSseFrame/
ParsedMetaAiResponse types) verbatim into the leaf muse-spark-web/response-parser.ts.
Host imports back the 3 it uses; all module-private (no re-export).

Host 1301 -> 925 LOC. Byte-identical bodies (verbatim), leaf has zero imports (no cycle).
Conversation cache, cookie/auth, fetch, executor class untouched. Adds a split-guard;
consumer tests stay green (muse-spark-cookie-copy-5449 2, muse-spark-web-continuation 6).

* refactor(executors): extract pure EventStream framing from kiro (#6018)

Extract the pure AWS EventStream binary framing (ByteQueue, CRC32 table + crc32,
TEXT_ENCODER/TEXT_DECODER, KIRO_VERIFY_FULL_CRC, parseEventFrame, EventFrame type)
verbatim into the self-contained leaf kiro/eventstream.ts (local JsonRecord alias to avoid
a cycle). Host imports back the 3 it uses (ByteQueue, TEXT_ENCODER, parseEventFrame).

Host 943 -> 758 LOC. Byte-identical bodies (verbatim 145/145), leaf has zero host imports
(no cycle). Auth/token-refresh/streaming-state/executor class untouched; the test-imported
flushBufferedToolArgs/resolveKiroRegion/kiroRuntimeHost stay exported on the host. Adds a
split-guard; consumer tests stay green (executor-kiro 9, kiro-tool-args-streaming 7,
kiro-iam-region 10).

* refactor(executors): extract challenge solver from duckduckgo-web (#6020)

Extract the DuckDuckGo anti-abuse challenge solver + FE signals (CHALLENGE_STUBS,
countHtmlElements, buildHtmlLookup, sha256Base64, solveDuckDuckGoChallenge,
makeDuckDuckGoFeSignals) verbatim into the leaf duckduckgo-web/challenge.ts. The vm
sandbox + 5s timeout (SECURITY note) are preserved. Host imports back the two it uses.

Host 924 -> 788 LOC. Byte-identical bodies (verbatim 132/132), leaf does not import the
host (no cycle). The now-dead createHash/parse5 host imports are removed; vm stays (still
used in host). Auth/cookie/warm/seed/executor untouched. Adds a split-guard; consumer
tests stay green (duckduckgo-web-executor 15, duckduckgo-domain-4037 8).

* test(cli): deflake setup-claude.test.ts — silence console to stop stdout/report interleaving (#5959) (#6019)

Integrated into release/v3.8.44. Deflakes tests/unit/cli/setup-claude.test.ts (#5959) — verified in CI: setup-claude now passes in Unit Tests fast-path (2/2).

Merged with --admin over two PRE-EXISTING base-reds proven independent of this test-only change (this PR only touches setup-claude.test.ts + CHANGELOG):
- Fast Quality Gates → check:test-discovery: tests/unit/executors/{firecrawl-fetch,xai-executor}.test.ts are orphaned on release/v3.8.44 (added by #5793/#5800); the shard glob 'tests/unit/{api,...,ui}/**' omits 'executors'. Both blobs exist on the pristine base.
- Unit Tests fast-path (2/2): tests/unit/settings-i18n-keys.test.ts → 'direct translation calls have English messages' fails on the pristine base too (unrelated i18n base-red).

* fix(cli): stabilize setup-claude.test.ts flake — inject dry-run log sink (#6021)

* fix(cli): stabilize setup-claude.test.ts flake — inject dry-run log sink (#5959)

Root cause (isolated empirically, 5/10 fail on the pristine base): the
dry-run path of syncClaudeProfilesFromModels console.log's a multi-byte
box-drawing heading ("── [dry-run] … ──"). Under the node:test runner
that write lands on the test child's stdout and corrupts the runner's
V8-serialized event stream ~50% of the time ("Unable to deserialize
cloned data due to invalid or unsupported version"), killing the file at
the first logging test. ASCII-only logging never reproduced it (0/20);
the unicode heading alone reproduced it (10/20).

Fix: syncClaudeProfilesFromModels accepts an injectable log sink
(opts.log, CLI default unchanged: console.log). The dry-run test injects
a collector — keeping unicode off the child's stdout — and gains
assertions on the dry-run report (path + parsed settings content), which
FAIL on the old code (log ignored) and PASS on the new one.

Validation: 0/30 failures post-fix vs 5/10 pre-fix on the same tree.

Baselines: complexity 2003->2006 and cognitive 859->860 are inherited
post-3a3d618fe release drift — measured identical on the pristine base
with and without this change (notes added in both files).

* test(ci): collect the orphaned tests/unit/executors/ directory (base-red unblock)

#5800 created tests/unit/executors/ outside every unit-runner brace glob,
so its 2 test files (firecrawl-fetch, xai-executor) never ran anywhere and
check:test-discovery flags them as NEW orphans on the pristine base,
red-flagging every PR into release/v3.8.44. Added 'executors' to the
runner globs in package.json (7 scripts), ci.yml unit shards, quality.yml
TIA glob, build-test-impact-map.mjs, and the test-discovery gate's
COLLECTORS (the gate enforces those stay in sync). Both files pass when
actually collected (10/10); cli+executors under suite flags: 99/99.

* chore(quality): complexity baseline 2006 -> 2007 (CI-observed value)

The GitHub fast-gates runner measures 2007 where local measures 2006 —
the same local-vs-CI off-by-one documented in the 2026-06-26 note. Pin
the CI-observed value so the gate is deterministic where it runs.

* fix(i18n): add the 6 missing en.json keys flagged by settings-i18n-keys (base-red unblock)

providers.iconUrlLabel/iconUrlHint (referenced by AddCompatibleProviderModal
and EditCompatibleNodeModal) and settings.authz.cors.wildcard.title/desc
(the #5602 CORS_ALLOW_ALL banner in AuthzSection) shipped without their
en.json messages — 'direct translation calls have English messages' fails
on the pristine release tip, red-flagging every PR. git log -S proves the
keys never existed (not a merge-eat). Scanner test: 10/10 green.

* refactor(executors): extract reasoning-effort (base) + tool-normalization (codex) leaves (#6030)

Two pure-leaf follow-ups closing the Block H tail:

- base/reasoningEffort.ts: provider-aware reasoning_effort sanitation
  (MISTRAL/GITHUB reject patterns, supportsMaxEffortForProvider,
  sanitizeReasoningEffortForProvider). Deps are config/services only
  (PROVIDER_CLAUDE, isClaudeCodeCompatible, supportsClaudeMaxEffort/supportsXHighEffort)
  so the leaf never imports the host — no cycle. base.ts re-exports
  sanitizeReasoningEffortForProvider for its external importers (mimoThinking + tests).
  base.ts 1466 -> 1312 LOC.

- codex/tools.ts: Responses-API tool normalization (CODEX_HOSTED_TOOL_TYPES hosted-tool
  passthrough, isCodexFreePlan gating, normalizeCodexTools). Self-contained
  (console.debug only). codex.ts re-exports isCodexFreePlan + normalizeCodexTools for
  external importers (tests + provider services). codex.ts 1430 -> 1268 LOC.

Byte-identical bodies (verbatim: base 100/100, codex 126/126); both leaves have zero host
imports. Adds two split-guards asserting the leaf owns the symbol and both import paths
resolve to the same function. Consumer tests stay green (base-executor-sanitize-effort 34,
executor-codex 40, mimoThinking 9, codex-free-plan-image-generation 3, issue-fixes 6).

* test(ci): move orphaned executor tests to top-level so a runner collects them (#6027)

Integrated into release/v3.8.44 — collect orphaned executor tests (check:test-discovery base-red).

* test(cli): deflake cli-setup-opencode.test.ts — silence console (#5959-class landmine) (#6033)

The command under test prints CLI progress with multi-byte glyphs
(printSuccess "✔" in the happy paths, printError "✖" in the dist-missing
path that test 4 exercises) via console.log. Under the node:test runner
those child-stdout writes interleave with the V8-serialized report frames
and can corrupt the stream — the exact #5959 mechanism proven for
setup-claude.test.ts; this file's ✖ line was already visible entangled in
red CI runs. No test here asserts on stdout, so silence console.log/info/
warn for the file (same pattern as #6019/#6021, restored in after()).

Validation: pre-fix the ✖/✔ lines reach stdout every run (grep-able);
post-fix stdout is clean, 4/4 tests green, 0/20 failures across 20 runs.

* feat(agy): support Google Cloud project ID settings (#5905)

* feat(agy): support Antigravity project ID settings

* refactor(agy): collapse Antigravity family project gate

---------

Co-authored-by: Nikolay Alafuzov <alafuzov_nn@rusklimat.ru>

* feat(proxy): add Webshare proxy pool import and sync (#5993)

* feat(proxy): add Webshare proxy pool import and sync

Adds Webshare (https://proxy.webshare.io) as a fourth source in the
free-proxy provider framework alongside 1proxy, Proxifly, and IPLocate.
WebshareProvider paginates the account's `/api/v2/proxy/list/` endpoint
(Authorization: Token <key>), upserts proxies into the shared
`free_proxies` table via the existing db/freeProxies.ts helpers, and
tombstones proxies the account no longer lists (recycled/retired IDs)
while never touching rows already promoted into the live proxy pool.

Unlike the other sources, Webshare is a paid per-account list, so it is
gated on FREE_PROXY_WEBSHARE_API_KEY rather than a plain on/off flag.
No DB migration needed — reuses the existing free_proxies table and
proxy_registry-on-promote path.

Co-authored-by: ricatix <d.enistraju155@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1176

* chore(changelog): restore release entries + add webshare bullet

---------

Co-authored-by: ricatix <d.enistraju155@gmail.com>

* feat(api-keys): add per-key device/connection tracking (#5998)

* feat(api-keys): add per-key device/connection tracking

Tracks distinct client devices (SHA-256 fingerprint of IP + User-Agent)
seen with each API key, with a 30-minute TTL and per-key/global caps. The
tracker is in-memory only (module-scoped Map, same pattern as
sessionManager.ts — no global.* singleton) and never stores the raw IP:
it is masked before being written.

Hooked into open-sse/handlers/chatCore.ts (the real chat entry) rather
than the legacy src/sse/handlers path. New GET /api/keys/[id]/devices
management route exposes masked device details for a key, and the
API Keys dashboard tab gets a "Devices" count badge alongside the
existing Sessions badge.

This is a new granularity distinct from the existing maxSessions cap
(src/lib/db/apiKeys.ts), which limits concurrent sticky-routing sessions
rather than tracking device identity.

Co-authored-by: Muhammad Mugni Hadi <mugni@rukita.co>
Inspired-by: https://github.com/decolua/9router/pull/931

* chore(changelog): restore release entries + add api-keys device-tracking bullet

---------

Co-authored-by: Muhammad Mugni Hadi <mugni@rukita.co>

* fix(providers): only apply openai-family model inference fallback when no cataloged provider serves the id (#5852) (#5938)

resolveModelByProviderInference() in open-sse/services/model.ts had an
unconditional /^gpt-/i heuristic that hijacked any model id starting with
gpt-/o1/o3 into provider openai, even when the id is cataloged under other
providers. This broke bare (non-combo) requests for open-weight models like
gpt-oss-120b (served by fireworks/cerebras/scaleway/byteplus/sambanova/
heroku), which don't exist on openai's catalog, producing a 404 with no
fallback.

Gate the heuristic on providers.length === 0 so it only fires for genuinely
uncataloged openai-family ids, letting cataloged ids fall through to the
existing single-candidate / ambiguous-candidate resolution paths.

Regression guard: tests/unit/gptoss-provider-inference-5852.test.ts

* fix(cc-compatible): send SSE accept for streamed requests (#5958)

Integrated into release/v3.8.44 — SSE Accept header for streamed cc-compatible requests (thanks @rdself).

* fix: deepseek-web reliability — auto-refresh on 401/403, refresh v2.0.0 client headers, fix token-kind bulk import (#5988)

Integrated into release/v3.8.44 — deepseek-web auto-refresh + v2.0.0 headers + token-kind bulk import (thanks @backryun).

* feat(providers): support Vercel AI Gateway embeddings and images (#5968)

* feat(providers): support Vercel AI Gateway embeddings and images

Extends the existing vercel-ai-gateway (alias vag) provider — currently
chat-only — with embeddings and image generation support, since the
gateway's OpenAI-compatible /v1 API also exposes /embeddings and
/images/generations. Adds entries to EMBEDDING_PROVIDERS
(embeddingRegistry.ts) and IMAGE_PROVIDERS (imageRegistry.ts) modeled
on the existing openai entries.

Out of scope for this PR (tracked as follow-ups): the /v1/credits
usage reader, retry:{429:2} tuning, and claude->reasoning_effort
mapping.

Co-authored-by: Ngô Tấn Tài <tantai@newnol.io.vn>
Inspired-by: https://github.com/decolua/9router/pull/1704

* chore(changelog): restore release entries + add vercel-gateway media bullet

---------

Co-authored-by: Ngô Tấn Tài <tantai@newnol.io.vn>

* feat(cli-tools): add Crush CLI tool to the dashboard (#5970)

* feat(cli-tools): add Crush CLI tool to the dashboard

Add a `crush` entry to the dashboard CLI-Tools catalog and a new
`/api/cli-tools/crush-settings` route (GET/POST/DELETE), cloned from the
`pi` tool's route as a template. OmniRoute already ships a `crush` CLI
command path (bin/cli/commands/setup-crush.mjs) but the dashboard catalog
had no matching entry.

The new route writes the real Crush config shape (providers.omniroute as
an openai-compat provider block) to the same canonical config path
(~/.config/crush/crush.json) that setup-crush.mjs's resolveCrushTarget()
already writes to, so the dashboard and the CLI command agree on one
location. Adds CLI_TOOL_RUNTIME_CONFIG.crush for detection/status, and
bumps EXPECTED_CODE_COUNT (18 -> 19) plus the catalog-count/schema tests
that enumerate the full tool list.

Co-authored-by: dopaemon <polarisdp@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1233

* chore(changelog): restore release entries + add crush cli bullet

---------

Co-authored-by: dopaemon <polarisdp@gmail.com>

* feat(dashboard): suggest HuggingFace Hub media models (#5990)

* feat(dashboard): suggest HuggingFace Hub media models

MVP scope:
- imageRegistry.ts: add an image kind entry for the huggingface provider
  (HF Inference API text-to-image), with a dedicated "huggingface-image"
  format since the endpoint returns raw image bytes rather than JSON.
- New handler open-sse/handlers/imageGeneration/providers/huggingface.ts,
  wired into imageGeneration.ts's format dispatch.
- New pure helper module open-sse/services/hfModelSuggestions.ts: maps a
  dashboard media kind to an HF Hub pipeline_tag and sorts/limits raw HF
  Hub search results (unit-tested directly).
- New route GET /api/v1/providers/suggested-models proxies the public HF
  Hub models search API server-side (Zod-validated query, buildErrorBody
  on every error path, no HF token exposed client-side — this project has
  no server-side HF search token config, so it calls unauthenticated).
- UI: ImageExampleCard now fetches suggested HF Hub models for the
  huggingface provider and merges them into the model picker as a
  selectable chip row, alongside the existing static provider models list.
- i18n: adds media.suggestedModels to en.json only.

Co-authored-by: yicone <yicone@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1633

* chore(changelog): restore release entries + add hf-hub media suggest bullet

---------

Co-authored-by: yicone <yicone@gmail.com>

* feat(dashboard): collapse and sort provider quota rows by remaining (#5977)

* feat(dashboard): collapse and sort provider quota rows by remaining

Sort the expanded quota list by remaining percentage (highest first)
and collapse it to the first 3 rows by default, with a "Show N more" /
"Show less" toggle when a connection reports more than 3 quotas. This
keeps the most at-risk quotas out of view below a long list of
healthy ones.

Extracts the sort/slice logic into pure helpers
(sortQuotasByRemaining, getVisibleQuotas) exported from
QuotaCardExpanded.tsx and unit-tests them directly.

Co-authored-by: CườngNH <j2.cuong@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1919

* chore(changelog): restore release entries + add quota collapse/sort bullet

---------

Co-authored-by: CườngNH <j2.cuong@gmail.com>

* feat(providers): refresh The Old LLM (Free) model catalog (#5181)

* feat(dashboard): add tool-source diagnostics settings toggle (#5978)

* feat(dashboard): add tool-source diagnostics settings toggle

Adds a Settings > Advanced card (cloned from DebugModeCard) that lets
operators flip the existing `logToolSources` flag from the UI instead
of editing the DB row directly. The backend gate (chatCore.ts) and DB
default were already present but had no toggle. Also adds
`logToolSources` to the /api/settings Zod PATCH schema (it is `.strict()`,
so the key was previously rejected) and en-only i18n strings.

Co-authored-by: DuyPrX <93126969+DuyPrX@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/1825

* chore(changelog): restore release entries + add tool-source toggle bullet

---------

Co-authored-by: DuyPrX <93126969+DuyPrX@users.noreply.github.com>

* feat(oauth): import Codex connection from a raw ChatGPT access token (#5995)

* feat(oauth): import Codex connection from a raw ChatGPT access token

OmniRoute's only Codex import path (/api/oauth/codex/import) required both
access_token and refresh_token, leaving no import path for a user who only
has a bare ChatGPT website access token (no refresh token).

- src/lib/db/providers.ts: createProviderConnection gains an explicit
  authType "access_token" branch — intentionally never deduped (no stable
  long-lived identity to match on) — and derives the connection name from
  email/name the same way "oauth" does.
- src/lib/oauth/services/codexImport.ts: export extractCodexAccountInfo so
  the new import path reuses the existing JWT decode instead of duplicating
  one.
- New route POST /api/oauth/codex/import-token (Zod-validated body
  { accessToken, name? }); errors routed through buildErrorBody /
  sanitizeErrorMessage. The executor's refreshCredentials() already
  degrades safely to null when there is no refresh token, forcing re-auth
  on expiry instead of a refresh exchange.
- OAuthModal.tsx: the callback-URL manual-paste path for codex now detects
  an eyJ-prefixed pasted token and posts it to the new endpoint, mirroring
  the existing grok-cli raw-token paste pattern.

Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/1290

* chore(changelog): restore release entries + add codex token-import bullet

---------

Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>

* fix(resilience): parse Retry-After from 429 JSON body for cooldown (#5974)

Integrated into release/v3.8.44 — parse Retry-After from 429 JSON body for cooldown (incl. #6013 retry-after-json extraction by @KooshaPari).

* fix(embeddings): forward connection-level proxy to embedding requests (#5975)

Integrated into release/v3.8.44 — forward connection-level proxy to embedding requests.

* fix(api): guard shared API client against non-JSON error responses (#5973)

Integrated into release/v3.8.44 — guard shared API client against non-JSON error responses.

* feat(dashboard): surface Codex banked reset credits per account (#5199)

* feat(providers): add NVIDIA NIM image generation (#5971)

* feat(providers): add NVIDIA NIM image generation

NVIDIA already exists as a chat provider (integrate.api.nvidia.com,
OpenAI-compatible) but image generation is served on a different host
(ai.api.nvidia.com/v1/genai/<model>) with a native NIM body shape, so it
gets a dedicated `nvidia-nim` image format and handler rather than reusing
the OpenAI image path.

Adds the 4 FLUX models (flux.1-dev, flux.1-schnell, flux.1-kontext-dev,
flux.2-klein-4b) to IMAGE_PROVIDERS, plus handleNvidiaNimImageGeneration()
which shapes the per-model NIM request body (flux.1-dev's mode/cfg_scale
and 768-1344px/64px-increment dimension validation, flux.1-kontext-dev's
required input image + aspect_ratio, schnell/klein's optional array-form
edit image) and normalizes the NIM response (artifacts[]/images[]/data[]/
single-value shapes) into the OpenAI `{created, data}` shape.

Co-authored-by: eng2007 <aleksey.semenov@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1195

* chore(changelog): restore release entries + add nvidia-nim image bullet

---------

Co-authored-by: eng2007 <aleksey.semenov@gmail.com>

* feat(providers): add Augment (Auggie CLI) local provider (#5972)

* feat(providers): add Augment (Auggie CLI) local provider

Adds a new local, no-auth provider that spawns the user's local `auggie`
CLI (`auggie --print --quiet --model <m> --`) and pipes a flattened prompt
via stdin, wrapping stdout as an OpenAI-compatible SSE stream or a single
chat.completion JSON body depending on the request's `stream` flag.

Auth is delegated entirely to `auggie login` outside OmniRoute — the
connection is registered `noAuth: true` and `refreshCredentials()` is a
no-op, matching the existing `NOAUTH_PROVIDERS` credential-less flow
(synthetic connection, no DB row required). An optional connection row is
still admitted via `FREE_APIKEY_PROVIDER_IDS` for display/priority
tracking, consistent with `opencode`. The dashboard "Test Connection"
flow spawns `auggie --version` to confirm the CLI is installed and
runnable, since there is no API key to validate upstream.

Security hardening (spawn is an untrusted-input sink):
- Command injection: spawn no longer passes `shell: true` on Windows. The
  binary is resolved to a concrete path/name and argv is handed straight to
  the OS loader, so no cmd.exe metacharacter interpretation is possible.
- Argument injection (flag smuggling): `model` is validated against the
  registry allowlist (`auggieProvider.models`) before any spawn — a model
  that is unknown or starts with "-" is rejected with a sanitized error and
  the subprocess is never started. A trailing `--` marks end-of-options in
  the argv as belt-and-suspenders.

Co-authored-by: chamdanilukman <16629923+chamdanilukman@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/1200

* test(golden): regenerate translate-path for auggie provider

---------

Co-authored-by: chamdanilukman <16629923+chamdanilukman@users.noreply.github.com>

* feat(providers): add ModelScope OpenAI-compatible provider (#5965)

* feat(providers): add ModelScope OpenAI-compatible provider

Ports ModelScope (Alibaba 魔搭) as a new API-key, OpenAI-compatible
provider — upstream 9router PR #1764. The upstream PR hardcoded
`https://api-inference.modelscope.ai/...` (`.ai` TLD); verified against
ModelScope's own API-Inference docs and third-party integration guides
that the real production domain is `api-inference.modelscope.cn`
(`.cn` TLD) and shipped that instead. Also drops the PR's static
5-model snapshot in favor of `passthroughModels: true` with an empty
seed list + `modelsUrl`, since ModelScope's open-model catalog moves
fast.

Updates the providers-constants-split characterization test's hardcoded
APIKEY_PROVIDERS count (159 -> 160) to match the new entry.

Co-authored-by: Umar Javed <114807145+tn5052@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/1764

* chore(changelog): restore release entries + add modelscope bullet

* test(golden): regenerate translate-path for modelscope provider

---------

Co-authored-by: Umar Javed <114807145+tn5052@users.noreply.github.com>

* feat(providers): add Qiniu OpenAI-compatible provider (#5966)

* feat(providers): add Qiniu OpenAI-compatible provider

Wires Qiniu (七牛云) AI inference gateway as a BYOK API-key provider.
Qiniu proxies many upstream models (DeepSeek V3/V4, Claude, Kimi and
more) behind a single key, so it ships with an empty static seed and
relies on passthroughModels + the live /v1/models catalog instead of a
single stale hardcoded model id.

- metadata: src/shared/constants/providers/apikey/gateways.ts
- registry entry: open-sse/config/providers/registry/qiniu/index.ts
  (format openai, executor default, bearer auth, baseUrl
  https://api.qnaigc.com/v1/chat/completions, modelsUrl
  https://api.qnaigc.com/v1/models)
- added to NAMED_OPENAI_STYLE_PROVIDERS so model import serves the live
  catalog and falls back to the (empty) local catalog on error, same
  pattern as the existing dgrid/zenmux/orcarouter gateways
- tests: tests/unit/qiniu-provider.test.ts (metadata, registry
  resolution, passthrough validation, live /v1/models fetch + fallback)

Co-authored-by: JiangZhuo <jiangzhuo@qiniu.com>
Inspired-by: https://github.com/decolua/9router/pull/911

* chore(changelog): restore release entries + add qiniu bullet

* test(golden): regenerate translate-path for qiniu provider

* test(providers): bump APIKEY count 160→161 for qiniu

---------

Co-authored-by: JiangZhuo <jiangzhuo@qiniu.com>

* feat(providers): add b.ai OpenAI-compatible provider (#5969)

* feat(providers): add b.ai OpenAI-compatible provider

Adds bai as a new OpenAI-compatible BYOK provider, distinct from the
existing thebai/theb.ai provider, using passthrough model discovery
(no hardcoded model list, live catalog served from api.b.ai/v1/models).

Co-authored-by: Delynn Assistant <zhen@dkzhen.org>
Inspired-by: https://github.com/decolua/9router/pull/963

* test(golden): regenerate translate-path for b.ai provider

* test(providers): bump APIKEY count 161→162 for b.ai

---------

Co-authored-by: Delynn Assistant <zhen@dkzhen.org>

* feat(providers): add Nube.sh OpenAI-compatible provider (#5936)

* feat(providers): add Nube.sh OpenAI-compatible provider

Nube.sh is a live BYOK OpenAI-compatible gateway (LiteLLM proxy) at
https://ai.nube.sh/api/v1, Bearer/API-key auth. Registered as an apikey
inference-host with an OpenAI-format, default-executor registry entry.

Its live model catalog is only reachable with a valid key
(/api/v1/models returns 401 unauthenticated), so no model IDs are
hardcoded — the entry uses passthroughModels + modelsUrl for live
enumeration instead of shipping unverifiable IDs.

Co-authored-by: whale9820 <87256750+whale9820@users.noreply.github.com>
Inspired-by: https://github.com/decolua/9router/pull/2294

* test(golden): regenerate translate-path for nube provider

* test(providers): bump APIKEY count 162→163 for nube

---------

Co-authored-by: whale9820 <87256750+whale9820@users.noreply.github.com>

* feat(providers): add Charm Hyper OpenAI-compatible provider (#5961)

* feat(providers): add Charm Hyper OpenAI-compatible provider

Registers Charm Hyper (hyper.charm.land) as a new API-key gateway
provider: OpenAI-compatible chat completions format, bearer auth,
free tier (100 monthly Hypercredits). Models are resolved via
passthrough (modelsUrl + live /v1/models import) instead of a
hardcoded upstream model list, since the specific model catalog is
not publicly documented.

Co-authored-by: whale <admin@dyntech.cc>
Inspired-by: https://github.com/decolua/9router/pull/2006

* test(golden): regenerate translate-path for charm-hyper provider

* test(providers): bump APIKEY count 163→164 for charm-hyper

---------

Co-authored-by: whale <admin@dyntech.cc>

* feat(providers): add SumoPod and X5Lab OpenAI-compatible providers (#5963)

* feat(providers): add SumoPod and X5Lab OpenAI-compatible providers

Both are OpenAI-compatible BYOK aggregator gateways, wired via the
default executor with bearer API-key auth. Neither ships a hardcoded
model list — both use passthroughModels with an empty seed list and a
live /v1/models fetcher, so the catalog always reflects what each
gateway actually serves instead of speculative model IDs.

- SumoPod: https://ai.sumopod.com/v1/chat/completions (sk- keys)
- X5Lab: https://api.x5lab.dev/v1/chat/completions (x5- keys)

Regression guard: tests/unit/sumopod-x5lab-provider.test.ts.

Co-authored-by: Rigel Ramadhani Waloni <rigel8911@gmail.com>
Inspired-by: https://github.com/decolua/9router/pull/1288

* chore(changelog): restore release entries + add sumopod/x5lab bullet

* test(golden): regenerate translate-path for sumopod + x5lab providers

* test(providers): bump APIKEY count 164→166 for sumopod + x5lab

---------

Co-authored-by: Rigel Ramadhani Waloni <rigel8911@gmail.com>

* feat(server): support reverse-proxy basePath deployment (#5992)

* feat(server): support reverse-proxy basePath deployment

Adds OMNIROUTE_BASE_PATH (opt-in, empty by default) to next.config.mjs
using Next.js's native basePath support so a deployment behind a
reverse-proxy subpath (e.g. https://host/omniroute/) works without
manual header stripping. Next.js strips the configured prefix from
nextUrl.pathname before route classification, so classifyRoute() and
isLocalOnlyPath() keep matching un-prefixed paths.

The two hardcoded auth redirect targets in
src/server/authz/pipeline.ts (root "/" -> "/dashboard" and
unauthenticated dashboard -> "/login") now prefix with
request.nextUrl.basePath so they stay inside the deployed subpath.
Default empty basePath is a no-op for existing root-path deployments.

Co-authored-by: zocomputer <help@zocomputer.com>
Inspired-by: https://github.com/decolua/9router/pull/1810

* docs(env): document OMNIROUTE_BASE_PATH in .env.example + ENVIRONMENT.md; restore changelog

* docs(env): document AUGGIE_BIN + CLI_AUGGIE_BIN (base-red from #5972 auggie)

---------

Co-authored-by: zocomputer <help@zocomputer.com>

* refactor(combo): extract buildTargetTimeoutRunner from handleComboChat (#6036)

Bloco J (hot-path decomposition), Task 1. Extract the per-target-timeout dispatch wrapper
(handleComboChat's handleSingleModelWithTimeout closure) verbatim into the leaf
combo/targetTimeoutRunner.ts as a factory buildTargetTimeoutRunner({handleSingleModel,
comboTargetTimeoutMs, log}). The per-model abort still comes from target.modelAbortSignal,
so the outer request signal is intentionally not a dependency. Host call-sites unchanged.

combo.ts shrinks ~60 LOC; leaf is 91 LOC (<800). Body byte-identical (verbatim), no cycle.
This is the first slice toward extracting the shared attempt-loop/success/error handlers
(Tasks 3-4) that de-duplicate handleComboChat and handleRoundRobinCombo. Adds a dedicated
test (5) so the failover path can be mutated independently. Consumer tests stay green
(combo-strategy-fallbacks 24, combo-499-abort 5, empty-content-failover 3, body-400-stop 1,
priority-quota-exhaustion 2, rr-streaming-lock 1, rr-session-stickiness 2).

Plan: _tasks/superpowers/plans/2026-07-03-blocoJ-combo-hotpath-decomposition.md

* feat(cli-tools): add CodeWhale CLI tool (#5996)

CodeWhale (https://github.com/Hmbown/CodeWhale) is the actively-maintained
successor to DeepSeek TUI — same author, renamed project. Added as a dual
entry alongside the existing "deepseek-tui" catalog entry (rather than a
hard rename) so users who still run the old DeepSeek TUI binary keep a
working dashboard card, while new users are steered to "codewhale".

New /api/cli-tools/codewhale-settings route writes the primary
~/.codewhale/config.toml and keeps an existing legacy
~/.deepseek/config.toml in sync (read fallback + best-effort write sync),
mirroring deepseek-tui-settings/route.ts. CLI_TOOLS and cliRuntime catalogs
updated; catalog cardinality tests/constants bumped accordingly (18→19
visible code tools, 28→29 total).


Inspired-by: https://github.com/decolua/9router/pull/1761

Co-authored-by: aristorinjuang <aristorinjuang@gmail.com>

* feat(i18n): auto-detect browser language on first visit (#5979)

* feat(i18n): auto-detect browser language on first visit

Adds a pure detectBrowserLocale() matcher (exact match, zh-HK/zh-MO
folded to zh-TW, language-prefix match, else null) plus a client-only
LocaleAutoDetect component mounted once in the root layout. On first
visit (no locale cookie set), it reads navigator.languages, computes a
match against the supported locales, and persists it via the same
cookie/localStorage writer LanguageSelector already used for manual
selection (now extracted to shared/lib/persistLocale.ts) before
refreshing the router.

Co-authored-by: anmingwei <anmingwei@dobest.com>
Inspired-by: https://github.com/decolua/9router/pull/1324

* chore(changelog): restore release entries + add browser-lang-detect bullet

---------

Co-authored-by: anmingwei <anmingwei@dobest.com>

* fix(dashboard): render Update-now API errors as text, not the raw envelope object (#5991) (#6028)

Integrated into release/v3.8.44 — fix(dashboard) render Update-now API errors as text, not the raw envelope object (#5991).

Merged with --admin: the fix is a one-line frontend change funneling the error body through the already-tested extractApiErrorMessage() helper, guarded by tests/unit/ui/home-update-error-render-5991.test.ts (3/3 pass, 3/3 fail on pre-fix source). The release branch is under a heavy parallel-merge storm (tip advanced ~6× mid-CI), so the branch is synced to the latest tip and landed atomically to avoid perpetual CONFLICTING; unit-shard reds seen earlier were pre-existing base-reds/flakes unrelated to this source-scan-only change.

* feat(api): expose provider plugin manifest (#6001)

* feat(api): expose provider plugin manifest

* test(translator): split responses chat request coverage

* test(mutation): register provider coverage tests

* feat(api): expose provider plugin manifest

* fix(ci): fail closed for prerelease latest promotion

* chore(ci): reconcile provider manifest complexity gate

* feat(api): expose provider plugin manifest

* test(translator): split responses chat request coverage

* test(mutation): register provider coverage tests

* fix(ci): fail closed for prerelease latest promotion

* chore: rebase onto release tip; drop out-of-scope translator test split + promote-script tweak

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* docs(changelog): add provider plugin manifest entry

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* chore(stryker): register account-fallback-retry-after-json test (base-red)

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: kooshapari <kooshapari@users.noreply.github.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@outlook.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(providers): add CN sign-up geo-restriction notices for SenseNova & StepFun (#5462)

* feat(sidecar): advertise provider manifest url (#6007)

* feat(sidecar): advertise provider manifest url via X-OmniRoute-Provider-Manifest-Url header

Re-cut onto release tip: manifest-url feature only (dropped stale-base noise).

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* docs(changelog): add sidecar manifest-url entry

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* chore(complexity): rebaseline 2009->2015 (inherited release-tip drift; feature adds 0)

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(autoCombo): latency/speed-optimized routing mode + omniroute_pick_fastest_model MCP tool (#6011)

* feat(autoCombo): latency/speed-optimized routing mode + omniroute_pick_fastest_model MCP tool

* test(translator): split responses chat request coverage

* refactor(mcp): extract fastest-model tool modules

* fix(i18n): cover provider icon and cors labels

* test(mutation): register latency coverage files

* test(ci): collect executor unit tests

* refactor(ci): reduce latency path complexity

* fix(mcp): include models catalog module

* feat(autoCombo): latency/speed-optimized routing + omniroute_pick_fastest_model MCP tool

Re-cut onto release tip: keep speed-routing + MCP tool + supporting catalog split;
drop out-of-scope translator split, en.json/ci.yml/package.json orphans, and unrelated
proxyFetch/responsesStreamHelpers/tokenLimitCounter refactors.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: kooshapari <kooshapari@users.noreply.github.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@outlook.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* docs(changelog): restore #5181/#5199/#5462 feature bullets eaten by merge

* feat(usage): on-demand period-scoped usage-data reset (re-cut onto release tip) (#5831)

* chore(quality): rebaseline eslintWarnings 4199->4256 + cognitiveComplexity 860->861 (v3.8.44 cycle drift)

Inherited v3.8.44 cycle drift measured on release tip 72ee80649 by the release-green
pre-flight during the /review-prs fix-batch round. The Quality Ratchet does NOT run on
PR->release fast-gates, so eslint warnings + cognitive complexity accrue unmeasured
across the cycle. Cyclomatic complexity is already green (2012 < baseline 2015) and
needs no bump. Each value carries a dated justification note; no production code touched.

* feat(claude-code): opt-in auto-permission classifier compat mode (re-cut onto release tip) (#5810)

* feat(providers): client-identity header profiles for compatible nodes (re-cut) + forbid cookie in custom headers (#5812)

* docs(openapi): document 9 newly-added routes to restore coverage ratchet (v3.8.44)

Documents the routes added this cycle that dropped openapiCoverage 36.9%->36.2%
below the ratchet baseline: 2 public v1 endpoints (/v1/ocr Mistral-OCR-compatible,
/v1/audio/translations Whisper-compatible) with full request/response specs, plus 7
dashboard/CLI-local routes marked x-internal:true (suggested-models, provider-plugin-
manifest, keys/{id}/devices, settings/purge-usage-history, oauth/codex/import-token,
cli-tools crush-settings + codewhale-settings). Coverage 36.2%->37.8% (207/547),
above baseline 36.9. check:openapi-routes/security-tiers/fabricated-docs all pass.

* refactor(sse): decompose handleComboChat auto-strategy region (Block J Task 2 — parseAutoConfig + resolveAutoStrategyOrder) (#6049)

* refactor(sse): extract pure parseAutoConfig leaf from handleComboChat

Block J Task 2 (safe slice): the auto-strategy config-resolution block in
handleComboChat is a pure function of (combo, eligibleTargets) with no side
effects, no early returns and no mutation. Extract it verbatim into
open-sse/services/combo/autoConfig.ts::parseAutoConfig so the god-function
shrinks and the derivation is independently unit-testable.

Behavior is byte-identical (verbatim-audited); combo.ts 3309->3280 LOC.
Adds tests/unit/combo-auto-config-split.test.ts (5 cases) pinning the
strategy-precedence, candidate-pool, weights and fallback derivations.

* refactor(sse): extract resolveAutoStrategyOrder leaf from handleComboChat

Block J Task 2 (coupled slice): the ~215-line `if (strategy === "auto")`
branch of handleComboChat is extracted into
open-sse/services/combo/resolveAutoStrategy.ts::resolveAutoStrategyOrder.

The branch is a control-flow region (mutates orderedTargets +
autoUsedExplicitRouter, early-returns 429, side-effect _registerExecutionCandidates),
so it is not a pure byte-identical move: the two `return unavailableResponse(...)`
exits become `{ earlyResponse }` and the mutated locals are returned instead of
closed over. Every other logic line is verbatim (semantic diff = only those
wrappers + the deeper getLKGP import path). `buildAutoCandidates` lives in
combo.ts, so it is injected via deps to keep the leaf acyclic (same DI pattern as
buildTargetTimeoutRunner) — which also makes the branch independently testable.

combo.ts 3280->3065 LOC. typecheck:core + check:cycles clean; dead host imports
removed. 60/60 consumer tests (router-strategies / auto-combo-engine /
combo-strategy-fallbacks / scoring-clamp / candidate-expansion / hidden-models)
cover the routable path end-to-end; new tests/unit/combo-resolve-auto-strategy-split.test.ts
pins the DI contract + the early-429 and default-ordering exits.

* test(sse): point quota-bypass source scan at resolveAutoStrategy leaf

The 'auto combo disables hard provider quota cutoffs when relay requests bypass'
source scan asserted combo.ts contains the bypass logic
(relayOptions?.bypassProviderQuotaPolicy === true + quotaPreflight enabled:false).
That block was extracted verbatim into combo/resolveAutoStrategy.ts (Block J
Task 2), so the scan now reads the leaf. Behavior unchanged.

* fix(ci): release-green base-reds — #5695 test regex + file-size rebaseline (#6093)

- tests/unit/ui/quick-start-api-keys-link-5695.test.ts: tolerate Prettier
  splitting <Link href=...> across lines (\s+) so the step1Desc regex matches
  the multi-line /dashboard/api-manager Link instead of skipping to step2's
  single-line /dashboard/providers Link. Code is correct; the test was brittle.
- config/quality/file-size-baseline.json: rebaseline 5 files that grew via
  already-merged PRs on the release tip (ApiManagerPageClient 3017->3058,
  OAuthModal 969->989, cliRuntime 1090->1100, webProvidersA 805->809,
  deepseek-web.test 1081->1092). Dated note added; shrink tracked in #3501.

* fix(translator): wrap Kiro system prompt in <system-reminder> (port from 9router#2306) (#6053)

Kiro/CodeWhisperer has no system role, so system messages were normalized to a
user turn with no wrapper — the full Claude Code system prompt then appeared as
raw user text, polluting the model context. Wrap system-origin content in
<system-reminder> tags before merging it into the Kiro user message. Real user
turns are unaffected. Existing history-merge tests aligned to the wrapped value.

Reported-by: VitzS7 (https://github.com/decolua/9router/issues/2306)

* fix(translator): strip multipleOf from antigravity/gemini tool schemas (port from 9router#2309) (#6052)

`multipleOf` is not part of the Gemini/antigravity OpenAPI 3.0 schema subset, so
leaving it in function_declaration parameters triggered a hard upstream 400
("Unknown name multipleOf"). Add it to GEMINI_UNSUPPORTED_SCHEMA_KEYS so it is
stripped at every schema level; minimum/maximum stay (Gemini accepts them).

Reported-by: abil0321 (https://github.com/decolua/9router/issues/2309)

* fix(kimi-web, qwen-web): align model catalog with live /models + map scenario per model (#5915)

* fix(kimi-web): align catalog with live models

Update the kimi-web catalog and request scenario selection to match
www.kimi.com's live GetAvailableModels response.

* fix(qwen-web): stop aliasing qwen3-coder-plus

Keep qwen3-coder-plus as its own model because it is present in the
live Qwen web models catalog.

* feat(minimax): extract M3 <think> to reasoning_content on OpenAI-format tiers (#6050)

MiniMax M3 is registered with format:"openai" on 8 provider tiers (trae,
huggingchat, bazaarlink, ollama-cloud, opencode, cline, opencode-zen,
codebuddy-cn), where its raw <think>...</think> tags leaked directly into
`content` instead of surfacing as a separate `reasoning_content` field.

OmniRoute already has the extraction primitive
(extractThinkingFromContent in responseSanitizer/reasoning.ts); it was just
gated to deepseek-r1/r1-distill/qwq. Extend the allowlist
(isTextualReasoningTagNativeRoute) with a minimax-m3-only pattern, excluding
the two direct minimax/minimax-cn tiers, which stay on Anthropic's Messages
format (targetFormat: "claude") and already surface reasoning natively.


Inspired-by: https://github.com/decolua/9router/pull/2231

Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
Co-authored-by: zmf963 <19422469+zmf963@users.noreply.github.com>

* fix: unwrap Cline response envelope (#6046)

Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>

* refactor(sse): extract applyStrategyOrdering leaf from handleComboChat (Block J Task 3) (#6063)

* refactor(sse): extract applyStrategyOrdering leaf from handleComboChat

Block J Task 3: the ~177-line else-if chain covering every non-auto combo
strategy (lkgp / strict-random / random / fill-first / p2c / least-used /
cost-optimized / reset-aware / reset-window / context-optimized / headroom /
quota-share) is extracted into
open-sse/services/combo/applyStrategyOrdering.ts::applyStrategyOrdering.

Each branch only reorders orderedTargets (no early returns, no other mutable
state), so the extraction is a clean verbatim move returning the reordered list;
the host replaces the chain with `else { orderedTargets = await
applyStrategyOrdering(strategy, orderedTargets, deps); }`. Semantic diff vs the
original chain = only the leading `if` (was `} else if`), the trailing return
and the deeper getLKGP import path — no logic line changed. None of the 13 strategy
helpers live in combo.ts, so no DI/cycle (unlike the auto branch).

combo.ts 3065->2883 LOC (3309->2883 across Task 2+3). typecheck:core + check:cycles
clean; 9 dead host imports removed (targetSorters block emptied). 47/47 consumer
tests (router-strategies / combo-strategy-fallbacks / rr-session-stickiness /
tag-routing) cover the DB-backed branches end-to-end; new
tests/unit/combo-apply-strategy-ordering-split.test.ts pins random / fill-first /
unknown exits.

* test(sse): point #2359 modelStr-guard scans at applyStrategyOrdering leaf

The LKGP fallback + non-auto strategy ordering (the two target.modelStr string-
method call sites) were extracted verbatim from combo.ts into the
applyStrategyOrdering leaf (Block J Task 3). The #2359 source scans now read the
leaf that owns those usages; the guard and the no-unguarded-usage assertions are
unchanged in intent.

* chore(ci): scan combo strategy leaves in check:known-symbols

Block J decomposed the combo dispatch: the `strategy === "..."` branches for
the 12 non-auto strategies moved to combo/applyStrategyOrdering.ts and the auto
branch to combo/resolveAutoStrategy.ts. The known-symbols gate previously scanned
only combo.ts, so it would report those strategies as canonicalNotHandled. Scan
all three dispatch files. Verified: 18/18 canonical strategies via dispatch.

* fix(combo): fallback to sibling model on 500 for per-model-quota providers (#5976)

* fix(combo): fallback to sibling model on 500 for per-model-quota providers

Two issues prevented combo fallback when gemini/gemma-4-31b-it returned 500:

1. targetExhaustion: connection-level exhaustion marked the shared gemini
   connection as exhausted, skipping the sibling model (gemma-4-26b-a4b-it).
   Skip markConnectionLevelExhaustion for per-model-quota providers (gemini,
   github, passthrough, compatible) since a model-level 500 does not mean
   the connection is bad.

2. combo retry loop: the auth layer records a model lockout on 500, but the
   retry loop did not check isModelLocked before retrying — it retried the
   same locked model instead of falling back. Add isModelLocked guard before
   the transient-retry decision.

* fix tests timeout

* fix: clear quota fallback CI gates

* quality-gate: extract test SSE stream helpers

* drop scope creep

* fix(combo): retry sibling models only on 500 errors

* fix(combo): reconcile onto release/v3.8.44 — keep targetExhaustion 500 fix, drop slow integration test

Reconciled by maintainer onto the current release tip:
- kept the core fix (targetExhaustion.ts model-500 guard for per-model-quota
  providers + the isModelLocked retry early-return in combo.ts) and its unit test
- dropped tests/integration/combo-concurrent-failure-recovery.test.ts +
  _sseTestHelpers.ts: they use Math.random()-based delays and 30s timeouts, run
  >3min and are flake-prone in the test:integration CI job; the unit test
  (tests/unit/combo/combo-target-exhaustion.test.ts, 21 cases) fully covers the fix
- CHANGELOG entry added

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

---------

Co-authored-by: Koosha Pari <kooshapari@gmail.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@outlook.com>
Co-authored-by: hartmark <hartmark@users.noreply.github.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(xai): surface Grok usage on quota dashboard via local usageHistory aggregation (#5806)

xAI has no public per-account quota API (the billing console requires a
session cookie, not an API key). Add getXaiUsage(connectionId), mirroring
the existing Xiaomi MiMo self-track pattern: sum tokens routed to the
connection from usage_history via getMonthlyProviderTokensForConnection
and surface them as a cumulative, uncapped quota (unlimited: true,
remaining: 100 — xAI has no fixed monthly cap). Register 'xai' in
USAGE_FETCHER_PROVIDERS and wire a switch case in getUsageForProvider.


Inspired-by: https://github.com/decolua/9router/pull/2150

Co-authored-by: ron <devestacion@gmail.com>

* feat(services): add Mux managed embedded service (#6034)

Adds Mux (coder/mux — local agent-orchestration daemon) as a fourth-tier
embedded service built on the existing ServiceSupervisor framework, the
same shape as 9Router and CLIProxyAPI:

- Installer (src/lib/services/installers/mux.ts): npm install/update via
  runNpm (array args + env-based prefix, no shell interpolation), modeled
  on ninerouter.ts. Mux ships an npm package (`mux`) with a documented
  headless `mux server --host <host> --port <port>` mode, so no
  git-clone+build path was needed.
- Registered in bootstrap.ts (SERVICES[] + buildSpawnArgsFactory).
- DB seed migration 113 (version_manager row, not_installed/auto_start=0).
- 7 API endpoints under /api/services/mux/ (install/start/stop/restart/
  update/status/auto-start) plus the shared [name]/logs SSE endpoint,
  mirroring the cliproxy route shape and delegating errors through
  createErrorResponse().
- Dashboard tab (MuxServiceTab) reusing ServiceStatusCard,
  ServiceLifecycleButtons, AutoStartToggle, ServiceLogsPanel.
- Docs: EMBEDDED-SERVICES.md (service table, architecture diagram, API
  reference, key-injection section), openapi.yaml, ENVIRONMENT.md,
  .env.example.

Security:
- Every /api/services/mux/* route is covered by the existing
  LOCAL_ONLY_API_PREFIXES "/api/services/" prefix (Hard Rule #17);
  added an explicit isLocalOnlyPath regression test for all 8 routes.
- Mux binds to 127.0.0.1 explicitly (never 0.0.0.0) as defense-in-depth,
  since it orchestrates AI agents that can execute host commands.
- The bearer token is generated the same way as 9Router's key
  (getOrCreateApiKey) and injected via MUX_SERVER_AUTH_TOKEN (mux's
  documented env form) rather than a CLI flag, so it never appears in
  `ps`/process listings.
- No shell interpolation anywhere in the installer (Hard Rule #13): all
  npm/spawn args are static arrays; the install prefix and auth token
  travel via the env option.


Inspired-by: https://github.com/decolua/9router/pull/1802

Co-authored-by: Ansh7473 <Ansh7473@users.noreply.github.com>

* feat(services): promote Bifrost to embedded/supervised service (#5670) (#5817)

Promotes Bifrost (@maximhq/bifrost — Go AI-gateway) from an env-only relay
sidecar to a first-class embedded/supervised service, matching the existing
cliproxy/9router model. Implements item #2 of #5670; the broader RouterBackend
contract (items #1, #3-#5) stays out of scope.

- Installer (npm-style, ninerouter model): install/update/getInstalledVersion/
  getLatestVersion (1h cache)/resolveSpawnArgs (Go single-dash flags, pinned
  BIFROST_TRANSPORT_VERSION), needsApiKey=false
- Bootstrap SERVICES entry (healthPath /v1/models) + spawn-args factory branch
- Migration 113 seeds the version_manager row (not_installed, port 8080,
  auto_update=1, provider_expose=1)
- 7 lifecycle API routes under /api/services/bifrost/ (verbatim from cliproxy,
  errors sanitized) — loopback-only via existing LOCAL_ONLY_API_PREFIXES
- Shared [name]/logs branch for bifrost
- Dashboard tab + registration in the services page shell
- Relay auto-wiring: getBifrostRoutingConfig defaults BIFROST_BASE_URL to the
  supervised port when the instance is running; explicit env still wins; the
  env-only relay path (/v1/relay/.../bifrost) stays unchanged (compat layer)
- Docs (EMBEDDED-SERVICES, openapi) + unit tests (installer/route-guard/routing,
  19 tests) + RUN_SERVICES_INT-gated integration lifecycle

Note: the actual Go-binary install/start/health path requires a documented VPS
live-test before merge (Hard Rule #18 / spec section 7); the gated integration
harness is the vehicle for that run.

* fix(ci): document BIFROST_PORT to clear env-doc-sync base-red

The Bifrost embedded-service merge referenced process.env.BIFROST_PORT
(src/lib/services/bootstrap.ts, default 8080) without adding it to
.env.example / ENVIRONMENT.md, so check:env-doc-sync failed on the release
tip and reddened Fast Quality Gates for every open PR->release. Docs-only.

* fix(providers): emulate OpenAI tool_calls in GitLab Duo executor (#6051) (#6111)

Co-authored-by: felssxs <felssxs@users.noreply.github.com>

* fix(providers): strip orphan tool_result on Antigravity MITM path (#6026) (#6115)

* fix(registry): update grok-cli model context lengths (#5913)

grok-build 128k→256k, grok-composer-2.5-fast 128k→200k to match actual Grok CLI /context capacities so context-aware routing stops filtering these models out. Registry-only.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(proxy): batch delete, auto-test, health scheduler + transitive alias fix (#5918)

Proxy-registry batch management (batch-delete, auto-test, background health scheduler) + fix resolveProviderAlias to follow the alias chain transitively (oc -> opencode -> opencode-zen). Probe target now operator-configurable via PROXY_HEALTH_TEST_URL. Scope-creep files from the original branch dropped.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* feat(minimax): extract M3 reasoning_content on OpenAI-format tiers (#6073)

MiniMax M3 leaks raw <think>...</think> into content on 8 OpenAI-format provider tiers; extract it into reasoning_content, leaving the direct minimax/minimax-cn (Claude-format) tiers untouched. Replacement for the stale #5804 branch.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(ci): harden provider translate-path golden across CI runners (#6076)

Normalize OS/arch-derived request headers (X-Stainless-Os/Arch, (OS;arch) UAs, and Antigravity's os.platform()-derived platform substring) in the golden so the test is runner-independent. Fixes the Mac-literal Antigravity UA that would have failed on Linux CI. Supersedes stale #6002.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* test(embeddings): pin seeded connection to direct egress in route-edge-coverage (#5975 collateral)

#5975 made the embeddings service honor the connection-level proxy. The pre-existing
route-edge-coverage embeddings edge-case tests seed an openai connection while the
settings-proxy suite has left a provider-level proxy (provider.local:8080) in the shared
DATA_DIR that resetStorage() does not clear — inert before #5975, but now the leaked
proxy fast-fails the embedding upstream with PROXY_UNREACHABLE.

These tests do not exercise proxying, so seedOpenAIConnection now pins the connection to
proxyEnabled:false, making resolveProxyForConnection return a direct egress regardless of
leaked global proxyConfig. No assertions weakened; 16/16 in the file pass. Regression
surfaced by the concurrency=1 full-suite run; passes on #5975's parent, red after it.

* fix(config): externalize ws for copilot-m365-web executor (#6130, closes #6062)

Re-lands the #6098 ws-externalization fix onto release/v3.8.44 (it had merged to main by mistake and was reverted). Externalize ws/bufferutil/utf-8-validate so the copilot-m365-web WebSocket masking path works at runtime.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(providers): update Perplexity Web models (#6106)

Refresh the Perplexity Web model catalog + mode/model_preference mappings to the current live set. Regression guard: perplexity-web.test.ts.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(providers): update Gemini Web cookies and models (#6095)

Refresh Gemini Web cookie handling + model catalog. Regression guard: gemini-web.test.ts.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(models): normalize GLM-5.2 provider context (#6091)

Hosted GLM-5.2 provider aliases now respect their declared context caps instead of inheriting the native 1M; native/bare + verified OpenCode/ZenMux routes stay at 1M. Regression guards added.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(combo): prefer known context capacity over unknown (#6088)

When a combo filters a target for exceeding a known context limit, prefer remaining known-compatible targets over unknown-metadata ones. Regression guard: combo-context-window-filter.test.ts.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix: keep Claude tool results adjacent (#6035)

Reattach OpenAI tool_result adjacent to tool_use before Claude send (#6026). Integrated into release/v3.8.44.

* fix(security): persist IP filter config + enforce it in the authz pipeline (#6131) (#6132)

Integrated into release/v3.8.44 — IP filter persistence + authz-pipeline enforcement (closes #6131). HARD-neutro: validate-release-green on the merge shows the same 3 pre-existing base-reds as the release baseline (test-masking cycle-wide, unit red-herring, integration batch-E2E env); #6131's own tests + ip-filter/pipeline suites all green.

* fix(codex): use access_token.exp instead of id_token.exp for import expiresAt (#6075) (#6084)

Prefer access_token.exp over id_token.exp for Codex auth import (#6075). Integrated into release/v3.8.44.

* fix(compression): send patch-only to PUT /api/settings/compression in CompressionHub (#6039) (#6077)

Send patch-only to PUT /api/settings/compression in CompressionHub (#6039). Integrated into release/v3.8.44.

* fix: reqId ReferenceError in safety-net redirect, dead code, filename typo (#6097)

Fix reqId ReferenceError in safety-net combo redirect + dead-code + DESING→DESIGN rename. Integrated into release/v3.8.44.

* fix(combo): expand fingerprint-based providers into per-fingerprint combo targets (#6082)

Expand fingerprint-based providers into per-fingerprint combo targets. Integrated into release/v3.8.44.

* fix(auth): persist quota preflight account lockouts (#6090)

Persist quota preflight account lockouts until reset window. Integrated into release/v3.8.44.

* fix(combos): expand OpenCode/MiMo fingerprint accounts in combo builder (#6087) (#6092)

Expand OpenCode/MiMo fingerprint accounts in combo builder (#6087). Integrated into release/v3.8.44.

* chore(quality): rebaseline v3.8.44 release-green drift (eslint/cognitive/cyclomatic/file-size)

Measured on release tip 32e4c906e during the #6131/#5975 release-green pass:
eslintWarnings 4256->4270 (+14), cognitiveComplexity 861->867 (+6), cyclomatic
count 2015->2026 (+11), and testFrozen caps for models-catalog-route (1507->1600),
perplexity-web (959->999), route-edge-coverage (1234->1241, my #5975 comment +7).
Inherited cycle drift (the Quality Ratchet does not run on PR->release fast-gates);
compression 'bun not found' is a local-env false and codeql is within baseline, so
neither is rebaselined. No production code touched.

* fix(accountFallback): persist per-account 429 cascade + classify 'Monthly usage limit. Resets in N days.' (#6061)

Persist per-account 429 cascade + classify 'Monthly usage limit. Resets in N days'. Integrated into release/v3.8.44.

* feat(build): backend-only fast build (skip the dashboard frontend) (#6119)

Backend-only fast build (skip dashboard frontend). Integrated into release/v3.8.44.

* fix(provider-limits): clear transient rate-limit state when quota recovers (#6128)

Clear transient rate-limit state when quota recovers. Integrated into release/v3.8.44.

* docs: Normalize mixed-language documentation content (#6105)

Normalize mixed-language documentation to English. Integrated into release/v3.8.44.

* chore docs

* i18n(zh-CN): translate CHANGELOG entries and section headings (#6043)

Adopt zh-CN as a translated locale: translate CHANGELOG + supporting docs. Integrated into release/v3.8.44.

* chore(quality): rebaseline residual eslint + file-size drift (v3.8.44)

Residual drift on release tip 716041223 (moving target): eslintWarnings 4270->4279
(+9 as the branch advanced past the prior rebaseline) and testFrozen/frozen file-size
caps for providerLimits.ts (955->982), accountFallback.ts (1790->1864) and
sse-auth.test.ts (1553->1600). All inherited from parallel-session merges (e.g. #6128);
the two production god-files ideally warrant decomposition rather than a bump (tracked
as debt). No production code touched.

* fix(repo): remove Windows case-conflicting DESIGN duplicate (#6140)

Remove stale root DESIGN.md (Windows case-conflict with design.md). Integrated into release/v3.8.44.

* fix(provider-limits): close TOCTOU race in quota recovery clear (I2) (#6139)

Close TOCTOU race in quota recovery clear via CAS primitive (I2 from #6128). Integrated into release/v3.8.44.

* fix(glm): suppress </think> close marker leak in GLM Anthropic transport (#6133)

Suppress </think> close-marker leak in GLM Anthropic transport. Integrated into release/v3.8.44.

* fix(cli): give setup-claude a fallback profile generator like setup-codex (#6138)

Give setup-claude a fallback profile generator like setup-codex. Integrated into release/v3.8.44.

* fix(onboarding): route provider-details link by node id, not provider slug (#6145) (#6145)

Route onboarding provider-details link by node id (#6145). Integrated into release/v3.8.44.

* fix(translator): strip Responses-only truncation field before Chat Completions forwarding (#6109)

Strip Responses-only truncation field before Chat Completions forwarding (#2311). Integrated into release/v3.8.44.

* fix(mitm): guard against concurrent MITM server starts (#6107)

Guard against concurrent MITM server starts (#2316). Integrated into release/v3.8.44.

* feat(models): add claude-sonnet-5 to Antigravity catalog (#6103)

Add claude-sonnet-5 to Antigravity catalog. Integrated into release/v3.8.44.

* fix(providers): strip thinking param for minimax-m2.7 on NVIDIA NIM (#6102)

Strip unsupported thinking param for minimax-m2.7 on NVIDIA NIM. Integrated into release/v3.8.44.

* feat(providers): add Kenari OpenAI-compatible gateway (#6104)

Add Kenari OpenAI-compatible gateway (BYOK). Integrated into release/v3.8.44.

* feat(sse): per-request Auto-Combo controls (X-OmniRoute-Mode / X-OmniRoute-Budget) — closes #6023 #6024 #6025 (#6057)

Per-request Auto-Combo controls (X-OmniRoute-Mode / X-OmniRoute-Budget). Integrated into release/v3.8.44.

* feat(resilience): throttle concurrent upstream quota fetches — closes #6009 (#6058)

Throttle concurrent upstream quota fetches (#6009). Integrated into release/v3.8.44.

* fix(oauth): graceful 400 for keychain-import-only providers (zed) (#6041) (#6054)

Graceful 400 for keychain-import-only providers on OAuth route (zed, #6041). Integrated into release/v3.8.44.

* fix(dashboard): resolve broken Card import breaking next build (base-red from #6061) (#6155)

* fix(dashboard): resolve broken Card import breaking next build (base-red from #6061)

CoolingConnectionsPanel imported `Card` from `@/components/ui/card`, a path
that does not exist in this repo (there is no shadcn-style `src/components/ui/`).
The PR->release fast-gates do not run `next build`, so the broken import slipped
in and `next build` failed with:

  Module not found: Can't resolve '@/components/ui/card'

Fix: the <Card> here was only a styled container, so replace it with a <div>
carrying the equivalent Tailwind classes (border/bg/padding + rounded-card
shadow-sm). Also normalize the file from CRLF to LF (it shipped with CRLF).

Adds a vitest/jsdom regression test (tests/unit/ui/CoolingConnectionsPanel.test.tsx)
that fails-without-fix (Vite: 'Failed to resolve import @/components/ui/card')
and passes with it, plus renders/empty-state coverage. Rule #18.

* fix(dashboard): stop client CoolingConnectionsPanel dragging server DB barrel into browser bundle

Second base-red from #6061, surfaced once the broken Card import was fixed:

  ./node_modules/ioredis/built/connectors/StandaloneConnector.js
  Module not found: Can't resolve 'net'
  Import trace: ioredis <- rateLimiter.ts <- apiKeys.ts <- @/lib/localDb
                <- CoolingConnectionsPanel.tsx (a "use client" component)

The client panel imported `formatResetCountdown` from `@/lib/localDb` — the
server-side DB re-export barrel — which transitively pulls better-sqlite3/ioredis
(node:net) into the browser bundle. That violates the CLAUDE.md rule 'never
barrel-import from localDb'.

`formatResetCountdown` is a pure date-formatting function, so move its
implementation to the client-safe `@/shared/utils/formatting` (alongside
formatTime/formatDuration) and re-export it from db/providers/rateLimit.ts for the
existing server callers + barrel. The panel now imports it directly from the
shared util — no server code in the client bundle.

Tests (Rule #18):
- tests/unit/format-reset-countdown.test.ts (node:test, blocking test:unit) —
  pure-function coverage: null/past/invalid, s, m+s, h+m, ISO string.
- tests/unit/ui/CoolingConnectionsPanel.test.tsx mock updated to the new module.

* fix(release): v3.8.44 Phase-0 pre-flight — base-red sweep + ratchet absorption

- fix(models): stop resolveProviderAlias at registered provider ids so oc/
  reaches the no-auth opencode provider again (#2901 contract, regressed by
  #5918's transitive chain; transitivity kept across alias-only hops)
- fix(auggie): handle async EPIPE 'error' events on child stdin so a
  fast-exiting CLI surfaces a sanitized error instead of crashing (both
  spawn sites); deflakes auggie-executor tests
- test: align provider family count 166->167 (Kenari #6104), regenerate
  translate-path golden on Linux (+kenari), opencode quota scope
  provider->connection (#6061)
- quality(test-masking): add _deletedWithReplacement allowlist support to
  check-test-masking.mjs (deletion exempt ONLY when the declared replacement
  test exists in HEAD; 5 new gate unit tests) + reduction allowlist entries
  for the verified #5958/#6088/#5816 migrations + targetExhaustion->
  combo-target-exhaustion replacement (#5976, 21 cases/52 asserts vs 13/37)
- quality(file-size): absorb v3.8.44 cycle drift (oauth route 960,
  providerLimits 998, chat 1662, auth 2426) with justification; #6158 will
  restore the oauth-route freeze
- changelog: bullets for the above + the #6155 cooling-panel build fix

* chore(release): v3.8.44 — 2026-07-04

Release reconciliation + close (generate-release Phases 0a/1):
- CHANGELOG [3.8.44]: 21 PR refs added to existing bullets, 62 new bullets
  (incl. restoration of ~10 bullets erased by the stale-branch merge in
  1f6ec5bc8), 3 Maintenance rollups, #6061/#6130 credit fixes, 🙌 Contributors
  table (35 external contributors) — coverage 144/153 cycle commits by #ref
- 42 docs/i18n CHANGELOG mirrors synced (EN content; i18n workflow translates)
- README: What's New refreshed for v3.8.44 highlights
- build scope: exclude electron/node_modules + electron/dist-electron + .build
  from tsconfig (local build-output leak poisoned next build with 8GB OOM —
  same class as the 2026-06-25 incident; scope 14765→5207, gate green)
- quality: cyclomatic baseline 2026→2028 (+2 inherited end-of-cycle drift;
  verified the release-captain code fixes add 0 new violations)

* fix(release): v3.8.44 one-pass release-PR CI sweep

- fix(dashboard): /dashboard/system/proxy 500'd on EVERY render — #5918 put
  useProxyBatchOperations(load) before the const load declaration (TDZ
  ReferenceError, digest 539380095). Hook block moved after load; SSR
  renderToString regression test added (the exact crash mode).
- fix(server): TRACE/TRACK/CONNECT crashed Next's middleware adapter
  (undici cannot represent them) into a raw 500 on every route — the raw
  HTTP method guard now answers 405 + Allow up-front (dast-smoke
  Schemathesis finding on /api/keys/{id}/devices); guard test added.
- fix(api): restore Zod validation on the provider-scoped chat route via a
  .passthrough() schema preserving #5907's relaxed semantics (t06 gate).
- docs(openapi): /api/keys/{id}/devices 401 now refs the management error
  envelope (Schemathesis schema-conformance).
- quality: rebaseline i18nUiCoverage 77.5->76.8 (+~1352 new en.json UI keys
  from the cycle await the async translation workflow; v3.8.39 precedent).
- CodeQL: dismissed 2 incomplete-url-substring FPs on unit-test asserts
  (v3.8.35 precedent) with Hard Rule #14 justifications.
- changelog: bullets for the above + 42 i18n mirrors re-synced

* fix(release): round-2 CI findings — LocaleAutoDetect refresh gating + ratchet tighten

- fix(i18n): LocaleAutoDetect (#5979) refreshed the router on EVERY cookie-less
  first visit, even when the detected locale matched the server-rendered
  <html lang> — re-navigating mid-interaction (flaky e2e 'execution context
  destroyed' + visible flash for new visitors). Refresh now only fires when
  the locale actually differs; regression test added.
- quality: tighten openapiCoverage.pct 36.9->39.3 (require-tighten gate on the
  release PR; value measured by the CI Quality Ratchet on 00c55afcb)
- quality(file-size): shrink the ProxyRegistryManager TDZ note to fit the
  1117-line freeze (prettier reflow added a line at commit time)
- changelog bullet + 42 i18n mirrors re-synced

* test(release): collect the #6082 fingerprint-expansion ghost test

check:test-discovery (Lint job, layered behind the round-1 t06 fix) flagged
tests/e2e/fingerprint-expansion.test.ts as a NEW orphan — it is a node:test
server-boot test that no runner collected, so it had never run. Moved to
tests/integration/ (the collector for this shape), fixed the helper import,
and verified it actually passes (3/3 on first-ever run). CHANGELOG ref updated.

---------

Co-authored-by: Chirag Singhal <76880977+chirag127@users.noreply.github.com>
Co-authored-by: Hamsa_M <116961508+hamsa0x7@users.noreply.github.com>
Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com>
Co-authored-by: Vittor Guilherme Borges de Oliveira <vittoroliveira.dev@gmail.com>
Co-authored-by: nickwizard <35692452+nickwizard@users.noreply.github.com>
Co-authored-by: Ankit <177378174+anki1kr@users.noreply.github.com>
Co-authored-by: Fadhil Yusuf <33994304+yusufrahadika@users.noreply.github.com>
Co-authored-by: Giorgos Giakoumettis <giorgos@yiakoumettis.gr>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>
Co-authored-by: AgentKiller45 <jamalzzj45@gmail.com>
Co-authored-by: Nikolay Alafuzov <alafuzov_nn@rusklimat.ru>
Co-authored-by: ricatix <d.enistraju155@gmail.com>
Co-authored-by: Muhammad Mugni Hadi <mugni@rukita.co>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: backryun <bakryun0718@proton.me>
Co-authored-by: Ngô Tấn Tài <tantai@newnol.io.vn>
Co-authored-by: dopaemon <polarisdp@gmail.com>
Co-authored-by: yicone <yicone@gmail.com>
Co-authored-by: CườngNH <j2.cuong@gmail.com>
Co-authored-by: DuyPrX <93126969+DuyPrX@users.noreply.github.com>
Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>
Co-authored-by: eng2007 <aleksey.semenov@gmail.com>
Co-authored-by: chamdanilukman <16629923+chamdanilukman@users.noreply.github.com>
Co-authored-by: Umar Javed <114807145+tn5052@users.noreply.github.com>
Co-authored-by: JiangZhuo <jiangzhuo@qiniu.com>
Co-authored-by: Delynn Assistant <zhen@dkzhen.org>
Co-authored-by: whale9820 <87256750+whale9820@users.noreply.github.com>
Co-authored-by: whale <admin@dyntech.cc>
Co-authored-by: Rigel Ramadhani Waloni <rigel8911@gmail.com>
Co-authored-by: zocomputer <help@zocomputer.com>
Co-authored-by: aristorinjuang <aristorinjuang@gmail.com>
Co-authored-by: anmingwei <anmingwei@dobest.com>
Co-authored-by: janeza2 <49841619+janeza2@users.noreply.github.com>
Co-authored-by: zmf963 <19422469+zmf963@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Koosha Pari <kooshapari@gmail.com>
Co-authored-by: hartmark <hartmark@users.noreply.github.com>
Co-authored-by: ron <devestacion@gmail.com>
Co-authored-by: Ansh7473 <Ansh7473@users.noreply.github.com>
Co-authored-by: felssxs <felssxs@users.noreply.github.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Arthur Bodera <abodera@gmail.com>
Co-authored-by: Semianchuk Vitalii <fix20152@gmail.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: NOXX - Commiter <artur1992123@mail.ru>
Co-authored-by: Milan Soni <123074437+Iammilansoni@users.noreply.github.com>
Co-authored-by: Devin <studyzy@gmail.com>
Co-authored-by: Raxxoor <manker_lol@hotmail.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
2026-07-04 13:00:30 -03:00

120 KiB
Raw Blame History

🚀 OmniRoute — Gateway AI Gratis (Bahasa Indonesia)

🌐 Languages: 🇺🇸 English · 🇸🇦 ar · 🇧🇬 bg · 🇧🇩 bn · 🇨🇿 cs · 🇩🇰 da · 🇩🇪 de · 🇪🇸 es · 🇮🇷 fa · 🇫🇮 fi · 🇫🇷 fr · 🇮🇳 gu · 🇮🇱 he · 🇮🇳 hi · 🇭🇺 hu · 🇮🇩 id · 🇮🇹 it · 🇯🇵 ja · 🇰🇷 ko · 🇮🇳 mr · 🇲🇾 ms · 🇳🇱 nl · 🇳🇴 no · 🇵🇭 phi · 🇵🇱 pl · 🇵🇹 pt · 🇧🇷 pt-BR · 🇷🇴 ro · 🇷🇺 ru · 🇸🇰 sk · 🇸🇪 sv · 🇰🇪 sw · 🇮🇳 ta · 🇮🇳 te · 🇹🇭 th · 🇹🇷 tr · 🇺🇦 uk-UA · 🇵🇰 ur · 🇻🇳 vi · 🇨🇳 zh-CN


Jangan pernah berhenti ngoding. Routing cerdas ke model AI GRATIS & berbiaya rendah dengan fallback otomatis.

Proxy API universal Anda — satu endpoint, 100+ penyedia, tanpa downtime. Kini dengan MCP Server (25 alat), Protokol A2A, Sistem Memori/Skill & Aplikasi Desktop Electron.

Chat Completions • Embeddings • Pembuatan Gambar • Video • Musik • Audio • Reranking • Pencarian Web • MCP Server • Protokol A2A • 100% TypeScript


🌐 Available in: 🇺🇸 English | 🇧🇷 Português (Brasil) | 🇪🇸 Español | 🇫🇷 Français | 🇮🇹 Italiano | 🇷🇺 Русский | 🇨🇳 中文 (简体) | 🇩🇪 Deutsch | 🇮🇳 हिन्दी | 🇹🇭 ไทย | 🇺🇦 Українська | 🇸🇦 العربية | 🇯🇵 日本語 | 🇻🇳 Tiếng Việt | 🇧🇬 Български | 🇩🇰 Dansk | 🇫🇮 Suomi | 🇮🇱 עברית | 🇭🇺 Magyar | 🇮🇩 Bahasa Indonesia | 🇰🇷 한국어 | 🇲🇾 Bahasa Melayu | 🇳🇱 Nederlands | 🇳🇴 Norsk | 🇵🇹 Português (Portugal) | 🇷🇴 Română | 🇵🇱 Polski | 🇸🇰 Slovenčina | 🇸🇪 Svenska | 🇵🇭 Filipino | 🇨🇿 Čeština


🖼️ Dashboard Utama

OmniRoute Dashboard

📸 Pratinjau Dashboard

Klik untuk melihat tangkapan layar dashboard
Halaman Tangkapan Layar
Providers Providers
Combos Combos
Analytics Analytics
Health Health
Translator Translator
Settings Settings
CLI Tools CLI Tools
Usage Logs Usage
Endpoints Endpoints

🤖 Penyedia AI Gratis untuk agen coding favorit Anda

Hubungkan IDE atau alat CLI berbasis AI apa pun melalui OmniRoute — gateway API gratis untuk coding tanpa batas.

OpenClaw
OpenClaw

205K
NanoBot
NanoBot

20.9K
PicoClaw
PicoClaw

14.6K
ZeroClaw
ZeroClaw

9.9K
IronClaw
IronClaw

2.1K
OpenCode
OpenCode

106K
Codex CLI
Codex CLI

60.8K
Claude Code
Claude Code

67.3K
Kilo Code
Kilo Code

15.5K

📡 Semua agen terhubung melalui http://localhost:20128/v1 atau http://cloud.omniroute.online/v1 — satu konfigurasi, model dan kuota tak terbatas


🤔 Mengapa OmniRoute?

Berhenti membuang uang dan terus mencapai batas:

  • Kuota langganan kedaluwarsa tanpa digunakan setiap bulan
  • Batas rate menghentikan Anda di tengah sesi coding
  • API mahal ($20-50/bulan per penyedia)
  • Perpindahan manual antar penyedia

OmniRoute mengatasi ini:

  • Maksimalkan langganan - Pantau kuota, gunakan setiap bit sebelum reset
  • Fallback otomatis - Langganan → Kunci API → Murah → Gratis, tanpa downtime
  • Multi-akun - Round-robin antar akun per penyedia

📧 Dukungan

💬 Bergabunglah dengan komunitas kami! Grup WhatsApp — Dapatkan bantuan, berbagi tips, dan tetap terupdate.

🐛 Melaporkan Bug?

Saat membuka issue, jalankan perintah system-info dan lampirkan file yang dihasilkan:

npm run system-info

Perintah ini menghasilkan system-info.txt berisi versi Node.js, versi OmniRoute, detail OS, alat CLI yang terpasang (qoder, gemini, claude, codex, antigravity, droid, dll.), status Docker/PM2, dan paket sistem — semua yang dibutuhkan untuk mereproduksi masalah Anda dengan cepat. Lampirkan file tersebut langsung ke GitHub issue Anda.


🔄 Cara Kerjanya

┌─────────────┐
│  Your CLI   │  (Claude Code, Codex, OpenClaw, Cursor, Cline...)
│   Tool      │
└──────┬──────┘
       │ http://localhost:20128/v1
       ↓
┌─────────────────────────────────────────┐
│           OmniRoute (Router Cerdas)       │
│  • Translasi format (OpenAI ↔ Claude)   │
│  • Pelacakan kuota + Embeddings + Gambar│
│  • Refresh token otomatis               │
└──────┬──────────────────────────────────┘
       │
       ├─→ [Tier 1: LANGGANAN] Claude Code, Codex
       │   ↓ kuota habis
       ├─→ [Tier 2: KUNCI API] DeepSeek, Groq, xAI, Mistral, NVIDIA NIM, dll.
       │   ↓ batas anggaran
       ├─→ [Tier 3: MURAH] GLM ($0.6/1M), MiniMax ($0.2/1M)
       │   ↓ batas anggaran
       └─→ [Tier 4: GRATIS] Qoder, Qwen, Kiro (tidak terbatas)

Hasil: Tidak pernah berhenti coding, biaya minimal

🎯 Apa yang Diselesaikan OmniRoute — 30 Masalah Nyata & Kasus Penggunaan

Setiap developer yang menggunakan alat AI menghadapi masalah ini setiap hari. OmniRoute dibangun untuk menyelesaikannya semua — dari pembengkakan biaya hingga pemblokiran regional, dari alur OAuth yang rusak hingga operasi protokol dan observabilitas enterprise.

💸 1. "Saya membayar langganan mahal tapi masih terganggu oleh batas"

Developer membayar $20200/bulan untuk Claude Pro, Codex Pro, atau GitHub Copilot. Meski sudah membayar, kuota memiliki batas — 5 jam penggunaan, batas mingguan, atau batas rate per menit. Di tengah sesi coding, penyedia berhenti merespons dan developer kehilangan fokus dan produktivitas.

Cara OmniRoute menyelesaikannya:

  • Fallback 4-Tier Cerdas — Jika kuota langganan habis, secara otomatis mengarahkan ke Kunci API → Murah → Gratis tanpa intervensi manual
  • Pelacakan Batas Penyedia — Snapshot kuota yang di-cache diperbarui sesuai jadwal sisi server (default PROVIDER_LIMITS_SYNC_INTERVAL_MINUTES=70) dengan pembaruan manual tersedia di UI
  • Dukungan Multi-Akun — Beberapa akun per penyedia dengan round-robin otomatis — saat satu habis, beralih ke berikutnya
  • Combo Kustom — Rantai fallback yang dapat dikustomisasi dengan 13 strategi penyeimbangan (priority, weighted, fill-first, round-robin, P2C, random, least-used, cost-optimized, strict-random, auto, lkgp, context-optimized, context-relay)
  • Pembangun Combo Terstruktur — Buat combo langkah demi langkah dengan pemilihan penyedia + model + akun yang eksplisit, termasuk penyedia berulang dan target akun tetap
  • P2C Berbasis Kuota — Pemilihan akun power-of-two kini mempertimbangkan kapasitas kuota, backoff, kesalahan terkini, dan penggunaan berturut-turut
  • Kuota Bisnis Codex — Pemantauan kuota workspace Business/Team langsung di dashboard
🔌 2. "Saya perlu menggunakan beberapa penyedia tetapi masing-masing memiliki API berbeda"

OpenAI menggunakan satu format, Claude (Anthropic) menggunakan format lain, Gemini pun berbeda lagi. Jika seorang developer ingin menguji model dari penyedia berbeda atau melakukan fallback di antara mereka, mereka perlu mengonfigurasi ulang SDK, mengganti endpoint, dan menangani format yang tidak kompatibel. Penyedia kustom (FriendLI, NIM) memiliki endpoint model non-standar.

Cara OmniRoute menyelesaikannya:

  • Endpoint Terpadu — Satu http://localhost:20128/v1 berfungsi sebagai proxy untuk semua 100+ penyedia
  • Translasi Format — Otomatis dan transparan: OpenAI ↔ Claude ↔ Gemini ↔ Responses API
  • Sanitasi Respons — Menghapus field non-standar (x_groq, usage_breakdown, service_tier) yang merusak OpenAI SDK v1.83+
  • Normalisasi Peran — Mengonversi developersystem untuk penyedia non-OpenAI; systemuser untuk GLM/ERNIE
  • Ekstraksi Tag Think — Mengekstrak blok <think> dari model seperti DeepSeek R1 ke reasoning_content yang terstandarisasi
  • Output Terstruktur untuk Gemini — Konversi otomatis json_schemaresponseMimeType/responseSchema
  • stream default ke false — Selaras dengan spesifikasi OpenAI, menghindari SSE tak terduga di SDK Python/Rust/Go
🌐 3. "Penyedia AI saya memblokir wilayah/negara saya"

Penyedia seperti OpenAI/Codex memblokir akses dari wilayah geografis tertentu. Pengguna mendapat kesalahan seperti unsupported_country_region_territory saat OAuth dan koneksi API. Ini sangat membuat frustasi para developer dari negara berkembang.

Cara OmniRoute menyelesaikannya:

  • Konfigurasi Proxy 3-Level — Proxy yang dapat dikonfigurasi di 3 level: global (semua lalu lintas), per-penyedia (hanya satu penyedia), dan per-koneksi/kunci
  • Lencana Proxy Berkode Warna — Indikator visual: 🟢 proxy global, 🟡 proxy penyedia, 🔵 proxy koneksi, selalu menampilkan IP
  • Pertukaran Token OAuth Melalui Proxy — Alur OAuth juga melewati proxy, menyelesaikan masalah unsupported_country_region_territory
  • Uji Koneksi via Proxy — Uji koneksi menggunakan proxy yang dikonfigurasi (tidak ada lagi bypass langsung)
  • Dukungan SOCKS5 — Dukungan proxy SOCKS5 penuh untuk routing keluar
  • Spoofing Sidik Jari TLS — Sidik jari TLS seperti browser melalui wreq-js untuk melewati deteksi bot
  • 🔏 Pencocokan Sidik Jari CLI — Menyusun ulang header dan field body agar sesuai dengan tanda tangan biner CLI asli, sangat mengurangi risiko pemanduan akun. IP proxy tetap dipertahankan — Anda mendapatkan kesiluman dan penyamaran IP secara bersamaan
🆓 4. "Saya ingin menggunakan AI untuk coding tapi tidak punya uang"

Tidak semua orang bisa membayar $20200/bulan untuk langganan AI. Pelajar, developer dari negara berkembang, penghobi, dan freelancer membutuhkan akses ke model berkualitas tanpa biaya sama sekali.

Cara OmniRoute menyelesaikannya:

  • Ollama Cloud — Model Ollama yang di-host di cloud pada api.ollama.com dengan tier "Light usage" gratis; gunakan prefix ollamacloud/<model>
  • Combo Hanya Gratis — Rantai if/kimi-k2-thinking → qw/qwen3-coder-plus = $0/bulan tanpa downtime
  • Akses Gratis NVIDIA NIM — ~40 RPM akses gratis selamanya untuk 70+ model di build.nvidia.com (beralih dari kredit ke batas rate murni)
  • Strategi Optimasi Biaya — Strategi routing yang secara otomatis memilih penyedia termurah yang tersedia
🔒 5. "Saya perlu melindungi gateway AI saya dari akses tidak sah"

Saat mengekspos gateway AI ke jaringan (LAN, VPS, Docker), siapa pun yang memiliki alamat tersebut dapat mengonsumsi token/kuota developer. Tanpa perlindungan, API rentan terhadap penyalahgunaan, injeksi prompt, dan eksploitasi.

Cara OmniRoute menyelesaikannya:

  • Manajemen Kunci API — Pembuatan, rotasi, dan pembatasan lingkup per penyedia dengan halaman /dashboard/api-manager yang didedikasikan
  • Izin Tingkat Model — Batasi kunci API ke model tertentu (openai/*, pola wildcard), dengan toggle Izinkan Semua/Batasi
  • Perlindungan Endpoint API — Wajibkan kunci untuk /v1/models dan blokir penyedia tertentu dari daftar
  • Auth Guard + Perlindungan CSRF — Semua rute dashboard dilindungi dengan middleware withAuth + token CSRF
  • Pembatas Rate — Pembatasan rate per-IP dengan jendela yang dapat dikonfigurasi
  • Penyaringan IP — Allowlist/blocklist untuk kontrol akses
  • Penjaga Injeksi Prompt — Sanitasi terhadap pola prompt berbahaya
  • Enkripsi AES-256-GCM — Kredensial dienkripsi saat disimpan
🛑 6. "Penyedia saya mati dan saya kehilangan alur coding"

Penyedia AI bisa menjadi tidak stabil, mengembalikan kesalahan 5xx, atau mencapai batas rate sementara. Jika developer bergantung pada satu penyedia, mereka akan terganggu. Tanpa circuit breaker, percobaan ulang berulang dapat menyebabkan aplikasi crash.

Cara OmniRoute menyelesaikannya:

  • Antrian & Pacing Permintaan — Bucket permintaan per-koneksi memperhalus lonjakan sebelum mencapai batas rate upstream
  • Pendinginan Koneksi — Satu koneksi mendingin setelah kegagalan yang dapat dicoba ulang dengan petunjuk Retry-After upstream opsional dan backoff eksponensial
  • Circuit Breaker Penyedia — Penyedia hanya trip setelah fallback habis dan permintaan penyedia masih gagal dengan kesalahan transien seluruh penyedia; batas rate 429 yang terikat koneksi tetap di Pendinginan Koneksi
  • Tunggu Pendinginan — Server dapat menunggu pendinginan koneksi paling awal berakhir dan mencoba ulang permintaan klien yang sama secara otomatis
  • Anti-Thundering Herd — Perlindungan mutex + semaphore terhadap badai percobaan ulang bersamaan
  • Rantai Fallback Combo — Jika penyedia utama gagal, secara otomatis jatuh ke rantai berikutnya tanpa intervensi
  • Dashboard Kesehatan — Pemantauan uptime, status circuit breaker penyedia, pendinginan, statistik cache, latensi p50/p95/p99
🔧 7. "Mengonfigurasi setiap alat AI membosankan dan berulang"

Cara OmniRoute menyelesaikannya:

  • Dashboard Alat CLI — Halaman khusus dengan pengaturan satu klik untuk Claude Code, Codex CLI, OpenClaw, Kilo Code, Antigravity, Cline
  • Generator Konfigurasi GitHub Copilot — Menghasilkan chatLanguageModels.json untuk VS Code dengan pemilihan model massal
  • Wizard Orientasi — Pengaturan terpandu 4 langkah untuk pengguna pertama kali
  • Satu endpoint, semua model — Konfigurasi http://localhost:20128/v1 sekali, akses 100+ penyedia
🔑 8. "Mengelola token OAuth dari beberapa penyedia adalah mimpi buruk"

Cara OmniRoute menyelesaikannya:

  • Refresh Token Otomatis — Token OAuth diperbarui di latar belakang sebelum kedaluwarsa
  • OAuth Multi-Akun — Beberapa akun per penyedia melalui ekstraksi token JWT/ID
  • Perbaikan OAuth LAN/Jarak Jauh — Deteksi IP privat untuk redirect_uri + mode URL manual untuk server jarak jauh
  • OAuth di Balik Nginx — Menggunakan window.location.origin untuk kompatibilitas reverse proxy
  • Panduan OAuth Jarak Jauh — Panduan langkah demi langkah untuk kredensial Google Cloud di VPS/Docker
📊 9. "Saya tidak tahu berapa banyak yang saya belanjakan atau di mana"

Developer menggunakan beberapa penyedia berbayar tetapi tidak memiliki tampilan pengeluaran yang terpadu. Setiap penyedia memiliki dashboard penagihan sendiri, tetapi tidak ada tampilan konsolidasi. Biaya tak terduga bisa menumpuk.

Cara OmniRoute menyelesaikannya:

  • Dashboard Analitik Biaya — Pelacakan biaya per-token dan manajemen anggaran per penyedia
  • Batas Anggaran per Tier — Batas pengeluaran per tier yang memicu fallback otomatis
  • Konfigurasi Harga Per-Model — Harga yang dapat dikonfigurasi per model
  • Statistik Penggunaan Per Kunci API — Jumlah permintaan dan cap waktu terakhir digunakan per kunci
  • Dashboard Analitik — Kartu statistik, grafik penggunaan model, tabel penyedia dengan tingkat keberhasilan dan latensi
🐛 10. "Saya tidak dapat mendiagnosis kesalahan dan masalah dalam panggilan AI"

Saat panggilan gagal, pengembang tidak mengetahui apakah itu batas kecepatan, token kedaluwarsa, format salah, atau kesalahan penyedia. Log terfragmentasi di terminal yang berbeda. Tanpa observabilitas, debugging adalah trial-and-error.

Bagaimana OmniRoute menyelesaikannya:

  • Dasbor Log Terpadu — 4 tab: Log Permintaan, Log Proksi, Log Audit, Konsol
  • Penampil Log Konsol — Penampil gaya terminal real-time dengan level kode warna, gulir otomatis, pencarian, filter
  • Log Ringkasan SQLite — Indeks log permintaan dan proksi tetap dapat dikueri saat restart tanpa memuat blob payload besar ke SQLite
  • Translator Playground — 4 mode debugging: Playground (terjemahan format), Chat Tester (pulang pergi), Test Bench (batch), Live Monitor (real-time)
  • Telemetri Permintaan — latensi p50/p95/p99 + penelusuran X-Request-Id
  • Artefak Detail Berbasis File — Log aplikasi dirotasi berdasarkan ukuran, hari penyimpanan, dan jumlah arsip; payload permintaan/respons terperinci ada di DATA_DIR/call_logs/ dan diputar secara independen dari ringkasan SQLite
  • Laporan Info Sistemnpm run system-info menghasilkan system-info.txt dengan lingkungan lengkap Anda (versi Node, versi OmniRoute, OS, alat CLI, status Docker/PM2). Lampirkan saat melaporkan masalah untuk triase instan.
🏗️ 11. "Menyebarkan dan memelihara gateway itu rumit"

Menginstal, mengonfigurasi, dan memelihara proksi AI di berbagai lingkungan (lokal, VPS, Docker, cloud) membutuhkan banyak tenaga. Masalah seperti jalur hardcode, EACCES pada direktori, konflik port, dan pembangunan lintas platform menambah gesekan.

Bagaimana OmniRoute menyelesaikannya:

  • instal global npmnpm install -g omniroute && omniroute — selesai
  • Docker Multi-Platform — asli AMD64 + ARM64 (Apple Silicon, AWS Graviton, Raspberry Pi)
  • Docker Compose Profilesbase (tanpa alat CLI) dan cli (dengan Claude Code, Codex, OpenClaw)
  • Aplikasi Desktop Electron — Aplikasi asli untuk Windows/macOS/Linux dengan baki sistem, mulai otomatis, mode offline
  • Mode Port Terpisah — API dan Dasbor pada port terpisah untuk skenario tingkat lanjut (proksi terbalik, jaringan kontainer)
  • Cloud Sync — Konfigurasi sinkronisasi antar perangkat melalui Cloudflare Workers
  • DB Backups — Pencadangan otomatis, pemulihan, ekspor dan impor semua pengaturan, dengan DISABLE_SQLITE_AUTO_BACKUP untuk pencadangan yang dikelola secara eksternal
🌍 12. "Antarmuka hanya berbahasa Inggris dan tim saya tidak bisa berbahasa Inggris"

Tim di negara-negara yang tidak berbahasa Inggris, khususnya di Amerika Latin, Asia, dan Eropa, kesulitan dengan antarmuka yang hanya berbahasa Inggris. Hambatan bahasa mengurangi adopsi dan meningkatkan kesalahan konfigurasi.

Bagaimana OmniRoute menyelesaikannya:

  • Dasbor i18n — 30 Bahasa — 500+ tombol diterjemahkan termasuk Arab, Bulgaria, Denmark, Jerman, Spanyol, Finlandia, Prancis, Ibrani, Hindi, Hungaria, Indonesia, Italia, Jepang, Korea, Melayu, Belanda, Norwegia, Polandia, Portugis (PT/BR), Rumania, Rusia, Slovakia, Swedia, Thailand, Ukraina, Vietnam, China, Filipina, Inggris
  • Dukungan RTL — Dukungan kanan ke kiri untuk bahasa Arab dan Ibrani
  • README Multi-Bahasa — 30 terjemahan dokumentasi lengkap
  • Pemilih Bahasa — Ikon bola dunia di header untuk peralihan waktu nyata
🔄 13. "Saya memerlukan lebih dari sekedar chat — saya memerlukan embeddings, gambar, audio"

AI bukan hanya penyelesaian obrolan. Pengembang perlu membuat gambar, mentranskripsikan audio, membuat penyematan untuk RAG, mengubah peringkat dokumen, dan memoderasi konten. Setiap API memiliki titik akhir dan format yang berbeda.

Bagaimana OmniRoute menyelesaikannya:

  • Sematan/v1/embeddings dengan 6 penyedia dan 9+ model
  • Pembuatan Gambar/v1/images/generations dengan 10 penyedia dan 20+ model (OpenAI, xAI, Together, Fireworks, Nebius, Hyperbolic, NanoBanana, Antigravity, SD WebUI, ComfyUI)
  • Teks-ke-Video/v1/videos/generations — ComfyUI (AnimateDiff, SVD) dan SD WebUI
  • Teks-ke-Musik/v1/music/generations — ComfyUI (Audio Terbuka Stabil, MusicGen)
  • Transkripsi Audio/v1/audio/transcriptions — Whisper + Nvidia NIM, HuggingFace, Qwen3
  • Text-to-Speech/v1/audio/speech — ElevenLabs, Nvidia NIM, HuggingFace, Coqui, Tortoise, Qwen3, Inworld, Cartesia, PlayHT, + penyedia yang ada
  • Moderasi/v1/moderations — Pemeriksaan keamanan konten
  • Pemeringkatan ulang/v1/rerank — Pemeringkatan ulang relevansi dokumen
  • Respon API — Dukungan penuh /v1/responses untuk Codex
🧪 14. "Saya tidak punya cara untuk menguji dan membandingkan kualitas antar model"

Pengembang ingin mengetahui model mana yang terbaik untuk kasus penggunaan mereka — kode, terjemahan, penalaran — tetapi membandingkan secara manual itu lambat. Tidak ada alat evaluasi terintegrasi.

Bagaimana OmniRoute menyelesaikannya:

  • Evaluasi LLM — Pengujian set emas dengan 10 kasus yang dimuat sebelumnya yang mencakup salam, matematika, geografi, pembuatan kode, kepatuhan JSON, terjemahan, penurunan harga, penolakan keamanan
  • 4 Strategi Pertandinganexact, contains, regex, custom (fungsi JS)
  • Bangku Tes Taman Bermain Penerjemah — Pengujian batch dengan banyak masukan dan keluaran yang diharapkan, perbandingan lintas penyedia
  • Penguji Obrolan — Perjalanan bolak-balik penuh dengan rendering respons visual
  • Monitor Langsung — Aliran real-time dari semua permintaan yang mengalir melalui proxy
📈 15. "Saya perlu meningkatkan skala tanpa kehilangan performa"

Seiring bertambahnya volume permintaan, tanpa menyimpan pertanyaan yang sama akan menghasilkan biaya duplikat. Tanpa idempotensi, permintaan duplikat akan membuang-buang pemrosesan. Batasan tarif per penyedia harus dipatuhi.

Bagaimana OmniRoute menyelesaikannya:

  • Cache Semantik — Cache dua tingkat (tanda tangan + semantik) mengurangi biaya dan latensi
  • Request Idempoency — Jendela deduplikasi 5 detik untuk permintaan yang identik
  • Deteksi Batas Tarif — RPM per penyedia, selisih minimum, dan pelacakan serentak maks
  • Antrian & Kecepatan Permintaan — Antrean, kecepatan, dan kecepatan konkurensi yang dapat dikonfigurasi secara default di Pengaturan → Ketahanan
  • Cache Validasi Kunci API — cache 3 tingkat untuk kinerja produksi
  • Dasbor Kesehatan dengan Telemetri — latensi p50/p95/p99, statistik cache, waktu aktif
🤖 16. "Saya ingin mengontrol perilaku model secara global"

Pengembang yang menginginkan semua respons dalam bahasa tertentu, dengan nada tertentu, atau ingin membatasi token penalaran. Mengonfigurasi ini di setiap alat/permintaan tidak praktis.

Bagaimana OmniRoute menyelesaikannya:

  • Injeksi Perintah Sistem — Perintah global diterapkan ke semua permintaan
  • Validasi Anggaran Berpikir — Kontrol alokasi token penalaran per permintaan (passthrough, otomatis, kustom, adaptif)
  • 9 Strategi Perutean — Strategi global yang menentukan cara permintaan didistribusikan
  • Wildcard Router — pola provider/* merutekan secara dinamis ke penyedia mana pun
  • Combo Aktifkan/Nonaktifkan Toggle — Beralih kombo langsung dari dasbor
  • Pengurutan Kombo Manual — Seret kartu kombo berdasarkan pegangan dan pertahankan pesanan di SQLite
  • Toggle Penyedia — Mengaktifkan/menonaktifkan semua koneksi untuk penyedia dengan satu klik
  • Penyedia yang Diblokir — Kecualikan penyedia tertentu dari daftar /v1/models
🧰 17. "Saya membutuhkan alat MCP sebagai kemampuan produk kelas satu"

Many AI gateways expose MCP only as a hidden implementation detail. Teams need a visible, manageable operation layer.

Bagaimana OmniRoute menyelesaikannya:

  • MCP muncul di navigasi dasbor dan tab protokol titik akhir
  • Halaman manajemen MCP khusus dengan proses, alat, cakupan, dan audit
  • Mulai cepat bawaan untuk omniroute --mcp dan orientasi klien
🧠 18. "Saya memerlukan orkestrasi A2A dengan jalur tugas sinkronisasi + streaming"

Alur kerja agen memerlukan balasan langsung dan eksekusi streaming jangka panjang dengan kontrol siklus hidup.

Bagaimana OmniRoute menyelesaikannya:

  • Titik akhir A2A JSON-RPC (POST /a2a) dengan message/send dan message/stream
  • Streaming SSE dengan propagasi status terminal
  • API siklus hidup tugas untuk tasks/get dan tasks/cancel
🛰️ 19. "Saya membutuhkan kesehatan proses MCP yang nyata, bukan status yang dapat ditebak"

Tim operasional perlu mengetahui apakah MCP benar-benar aktif, bukan hanya apakah API dapat dijangkau.

Bagaimana OmniRoute menyelesaikannya:

  • File detak jantung runtime dengan PID, stempel waktu, transportasi, jumlah alat, dan mode cakupan
  • API status MCP menggabungkan detak jantung + aktivitas terkini
  • Kartu status UI untuk kesegaran proses/waktu aktif/detak jantung
📋 20. "Saya memerlukan eksekusi alat MCP yang dapat diaudit"

Saat alat mengubah konfigurasi atau memicu tindakan operasi, tim memerlukan kemampuan penelusuran forensik.

Bagaimana OmniRoute menyelesaikannya:

  • Pencatatan audit yang didukung SQLite untuk panggilan alat MCP
  • Filter berdasarkan alat, keberhasilan/kegagalan, kunci API, dan penomoran halaman
  • Tabel audit dasbor + titik akhir statistik untuk otomatisasi
🔐 21. "Saya memerlukan izin MCP terbatas per integrasi"

Different clients should have least-privilege access to tool categories.

Bagaimana OmniRoute menyelesaikannya:

  • 10 cakupan MCP granular untuk akses alat terkontrol
  • Penegakan cakupan dan visibilitas di UI manajemen MCP
  • Postur default yang aman untuk perkakas operasional
⚙️ 22. "Saya memerlukan kontrol operasional tanpa memindahkan"

Tim memerlukan perubahan runtime yang cepat selama insiden atau peristiwa biaya.

Bagaimana OmniRoute menyelesaikannya:

  • Beralih aktivasi kombo langsung dari dasbor MCP
  • Sesuaikan pengaturan antrean, cooldown, pemutus, dan tunggu dari halaman Ketahanan khusus
  • Tinjau status pemutus penyedia langsung dari dasbor Kesehatan
🔄 23. "Saya memerlukan visibilitas dan pembatalan siklus hidup tugas A2A langsung"

Without lifecycle visibility, task incidents become hard to triage.

Bagaimana OmniRoute menyelesaikannya:

  • Daftar tugas/pemfilteran berdasarkan status/keterampilan dengan penomoran halaman
  • Telusuri metadata tugas, peristiwa, dan artefak
  • Titik akhir pembatalan tugas dan tindakan UI dengan konfirmasi
🌊 24. "Saya memerlukan metrik aliran aktif untuk memuat A2A"

Alur kerja streaming memerlukan wawasan operasional tentang konkurensi dan koneksi langsung.

Bagaimana OmniRoute menyelesaikannya:

  • Penghitung aliran aktif terintegrasi ke dalam status A2A
  • Stempel waktu tugas terakhir dan jumlah per negara bagian
  • Kartu dasbor A2A untuk pemantauan operasi waktu nyata
🪪 25. "Saya memerlukan penemuan agen standar untuk klien"

Klien dan orkestra eksternal memerlukan metadata yang dapat dibaca mesin untuk orientasi.

Bagaimana OmniRoute menyelesaikannya:

  • Kartu Agen terungkap di /.well-known/agent.json
  • Kemampuan dan keterampilan yang ditunjukkan dalam manajemen UI
  • API status A2A mencakup metadata penemuan untuk otomatisasi
🧭 26. "Saya memerlukan kemampuan protokol untuk ditemukan di UX produk"

If users cannot discover protocol surfaces, adoption and support quality drop.

Bagaimana OmniRoute menyelesaikannya:

  • Halaman Endpoint terkonsolidasi dengan tab untuk Proxy, MCP, A2A, dan API Endpoints
  • Pengalih status layanan inline (Online/Offline) untuk MCP dan A2A
  • Tautan dari ikhtisar ke tab manajemen khusus
🧪 27. "Saya memerlukan validasi protokol end-to-end dengan klien nyata"

Mock tests are not enough to validate protocol compatibility before release.

Bagaimana OmniRoute menyelesaikannya:

  • Suite E2E yang mem-boot aplikasi dan menggunakan transportasi klien MCP SDK yang sebenarnya
  • Klien A2A menguji penemuan, pengiriman, streaming, dapatkan, dan pembatalan aliran
  • Periksa silang pernyataan terhadap audit MCP dan API tugas A2A
📡 28. "Saya memerlukan observabilitas terpadu di semua antarmuka"

Splitting observability by protocol creates blind spots and longer MTTR.

Bagaimana OmniRoute menyelesaikannya:

  • Dasbor/log/analitik terpadu dalam satu produk
  • Kesehatan + audit + permintaan telemetri di seluruh lapisan OpenAI, MCP, dan A2A
  • API Operasional untuk status dan otomatisasi
💼 29. "Saya memerlukan satu runtime untuk proxy + alat + orkestrasi agen"

Menjalankan banyak layanan terpisah akan meningkatkan biaya operasional dan mode kegagalan.

Bagaimana OmniRoute menyelesaikannya:

  • Proksi yang kompatibel dengan OpenAI, server MCP, dan server A2A dalam satu tumpukan
  • Otentikasi bersama, ketahanan, penyimpanan data, dan kemampuan observasi
  • Model kebijakan yang konsisten di seluruh platform interaksi
🚀 30. "Saya perlu mengirim workflow agentic tanpa tumpukan glue code"

Tim kehilangan kecepatan saat menggabungkan beberapa layanan dan skrip ad-hoc.

Bagaimana OmniRoute menyelesaikannya:

  • Strategi titik akhir terpadu untuk klien dan agen
  • UI manajemen protokol bawaan dan jalur validasi asap
  • Fondasi siap produksi (keamanan, logging, ketahanan, cadangan)
📚 31. "Sesi panjang saya crash karena batas 'context_length_exceeded'"

Selama proses debug mendalam, riwayat panjang dengan hasil alat dengan cepat melampaui jendela token penyedia, menyebabkan permintaan gagal dan konteks tidak ada lagi.

Bagaimana OmniRoute menyelesaikannya:

  • Kompresi Konteks Proaktif — Mengevaluasi anggaran token sebelum permintaan mencapai hulu dan secara proaktif memangkas riwayat percakapan lama dengan mekanisme pencarian biner yang cerdas.
  • Pengaman Integritas Struktural — Secara otomatis melacak definisi tool_use yang eksplisit dan memastikan bahwa jika masukan alat terpotong, tool_result yang terkait juga dihapus dengan aman, sehingga mencegah kesalahan validasi API.
  • Penghapusan Multi-Lapisan — Secara progresif menghapus pesan sistem, pesan biasa, dan akhirnya menerapkan batas panjang yang ketat tanpa merusak logika percakapan.

Contoh Playbook (Kasus Penggunaan Terintegrasi)

Playbook A: Maximize paid subscription + cheap backup

Combo: "maximize-claude"
  1. cc/claude-opus-4-7
  2. glm/glm-4.7
  3. if/kimi-k2-thinking

Monthly cost: $20 + small backup spend
Outcome: higher quality, near-zero interruption

Playbook B: Tumpukan coding tanpa biaya

Combo: "free-forever"
  1. if/kimi-k2-thinking       (unlimited free)
  2. qw/qwen3-coder-plus       (unlimited free)

Monthly cost: $0
Outcome: stable free coding workflow

Playbook C: Rantai fallback yang selalu aktif 24/7

Combo: "always-on"
  1. cc/claude-opus-4-7
  2. cx/gpt-5.2-codex
  3. glm/glm-4.7
  4. minimax/MiniMax-M2.1
  5. if/kimi-k2-thinking

Outcome: deep fallback depth for deadline-critical workloads

Playbook D: Operasi agen dengan MCP + A2A

1) Start MCP transport (`omniroute --mcp`) for tool-driven operations
2) Run A2A tasks via `message/send` and `message/stream`
3) Observe via /dashboard/endpoint (MCP and A2A tabs)
4) Toggle services via inline status controls

🆓 Mulai Gratis — Tanpa Biaya Konfigurasi

Siapkan pengkodean AI dalam hitungan menit di $0/bulan. Hubungkan akun gratis ini dan gunakan kombo Free Stack bawaan.

Step Action Providers Unlocked
1 Connect Kiro (AWS Builder ID OAuth) Claude Sonnet 4.5, Haiku 4.5 — unlimited
2 Connect Qoder (Google OAuth) kimi-k2-thinking, qwen3-coder-plus, deepseek-r1... — unlimited
3 Connect Qwen (Device Code) qwen3-coder-plus, qwen3-coder-flash... — unlimited
4 /dashboard/combosTemplat Tumpukan Gratis ($0) Round-robin semua penyedia gratis secara otomatis

Arahkan IDE/CLI apa pun ke: http://localhost:20128/v1 · Kunci API: any-string · Selesai.

Cakupan ekstra opsional (juga gratis): Kunci API Groq (gratis 30 RPM), NVIDIA NIM (gratis 40 RPM, 70+ model), Cerebras (1 juta tok/hari), kunci API LongCat (50 juta token/hari!), Cloudflare Workers AI (10 ribu neuron/hari, 50+ model).

Mulai Cepat

1) Instal dan jalankan

npm install -g omniroute
omniroute

pengguna pnpm: Pass --allow-build at install time to enable native build scripts required by better-sqlite3 and @swc/core (the approve-builds -g command is not supported for global installs on pnpm v11):

pnpm add -g omniroute@latest --allow-build=better-sqlite3 --allow-build=@swc/core
omniroute

Dasbor terbuka di http://localhost:20128 dan URL dasar API adalah http://localhost:20128/v1.

Arch Linux (AUR)

Pengguna Arch Linux dapat menginstal AUR package, yang menginstal OmniRoute dan menyediakan layanan pengguna systemd:

yay -S omniroute-bin
systemctl --user enable --now omniroute.service
Command Description
omniroute Mulai server (PORT=20128, API dan dasbor pada port yang sama)
omniroute --port 3000 Set canonical/API port to 3000
omniroute --mcp Mulai server MCP (stdio transport)
omniroute --no-open Don't auto-open browser
omniroute --help Show help

Optional split-port mode:

PORT=20128 DASHBOARD_PORT=20129 omniroute
# API:       http://localhost:20128/v1
# Dashboard: http://localhost:20129

2) Menghapus Instalasi

Saat Anda tidak lagi memerlukan OmniRoute, kami menyediakan dua skrip cepat untuk penghapusan bersih:

Command Action
npm run uninstall Menghapus aplikasi sistem tetapi menyimpan DB dan konfigurasi Anda di ~/.omniroute.
npm run uninstall:full Menghapus aplikasi DAN secara permanen menghapus semua konfigurasi, kunci, dan database.

Catatan: Untuk menjalankan perintah ini, navigasikan ke folder proyek OmniRoute (jika Anda mengkloningnya) dan jalankan. Alternatifnya, jika diinstal secara global, Anda cukup menjalankan npm uninstall -g omniroute.

Batas Waktu Streaming yang Berlangsung Lama

Untuk sebagian besar penerapan, Anda hanya memerlukan:

Variable Default Purpose
REQUEST_TIMEOUT_MS 600000 Garis dasar bersama untuk batas waktu mulai respons upstream, batas waktu Undici yang tersembunyi, permintaan sidik jari TLS, dan batas waktu permintaan/proksi jembatan API
STREAM_IDLE_TIMEOUT_MS inherits REQUEST_TIMEOUT_MS Kesenjangan maksimum antara potongan streaming sebelum OmniRoute membatalkan aliran SSE

Kompatibilitas mundur dipertahankan: FETCH_TIMEOUT_MS, API_BRIDGE_PROXY_TIMEOUT_MS, dan var batas waktu per lapisan lainnya yang ada masih berfungsi dan menggantikan garis dasar bersama.

Untuk upstream yang kompatibel dengan Kode Claude (anthropic-compatible-cc-*), OmniRoute juga memperoleh header X-Stainless-Timeout keluar dari batas waktu pengambilan yang diselesaikan sehingga batas waktu baca sisi penyedia tetap selaras dengan konfigurasi env Anda.

Untuk reverse proxy pihak ketiga yang kompatibel dengan Claude Code, OmniRoute tetap menggunakan default anthropic-beta disetel konservatif dan, ketika Client Cache Control tersisa di Auto, hanya meneruskan penanda cache_control yang disediakan klien. Jika permintaan tidak menyertakan cache_control, OmniRoute tidak memasukkan penanda milik jembatan.

Penggantian tingkat lanjut tersedia jika Anda memerlukan kontrol yang lebih baik:

Variable Default Purpose
FETCH_TIMEOUT_MS inherits REQUEST_TIMEOUT_MS Batas waktu mulai respons hulu digunakan hingga header respons tiba
FETCH_HEADERS_TIMEOUT_MS inherits FETCH_TIMEOUT_MS Batas waktu Undici untuk menerima header respons upstream
FETCH_BODY_TIMEOUT_MS inherits FETCH_TIMEOUT_MS Undici time limit between upstream body chunks (0 disables it)
FETCH_CONNECT_TIMEOUT_MS 30000 Undici TCP connect timeout
FETCH_KEEPALIVE_TIMEOUT_MS 4000 Undici idle keep-alive socket timeout
TLS_CLIENT_TIMEOUT_MS inherits FETCH_TIMEOUT_MS Batas waktu untuk permintaan sidik jari TLS yang dilakukan melalui wreq-js
API_BRIDGE_PROXY_TIMEOUT_MS inherits REQUEST_TIMEOUT_MS or 600000 Batas waktu untuk penerusan proxy /v1 dari port API ke port dasbor
API_BRIDGE_SERVER_REQUEST_TIMEOUT_MS max(API_BRIDGE_PROXY_TIMEOUT_MS, 300000) Batas waktu permintaan masuk di server jembatan API
API_BRIDGE_SERVER_HEADERS_TIMEOUT_MS 60000 Batas waktu header masuk di server jembatan API
API_BRIDGE_SERVER_KEEPALIVE_TIMEOUT_MS 5000 Batas waktu tetap hidup di server jembatan API
API_BRIDGE_SERVER_SOCKET_TIMEOUT_MS 0 Batas waktu ketidakaktifan soket di server jembatan API (0 menonaktifkannya)

Untuk permintaan streaming, FETCH_TIMEOUT_MS hanya mencakup pengaturan koneksi/menunggu respons upstream pertama. Setelah aliran aktif, OmniRoute hanya akan dibatalkan pada keadaan terhenti sebenarnya (STREAM_IDLE_TIMEOUT_MS) atau tubuh Undici tidak aktif (FETCH_BODY_TIMEOUT_MS).

Jika Anda menjalankan OmniRoute di belakang Nginx, Caddy, Cloudflare, atau proksi terbalik lainnya, pastikan proksi tersebut waktu tunggu juga lebih tinggi daripada waktu tunggu aliran/pengambilan OmniRoute Anda.

2) Hubungkan penyedia dan buat kunci API Anda

  1. Buka Dasbor → Providers dan sambungkan setidaknya satu penyedia (OAuth atau kunci API).
  2. Buka Dasbor → Endpoints dan buat kunci API.
  3. (Opsional) Buka Dasbor → Combos dan atur rantai cadangan Anda.

3) Arahkan alat pengkodean Anda ke OmniRoute

Base URL: http://localhost:20128/v1
API Key:  [copy from Endpoint page]
Model:    if/kimi-k2-thinking (or any provider/model prefix)

4) Mengaktifkan dan memvalidasi protokol (v2.0)

MCP (untuk operasi yang digerakkan oleh alat):

omniroute --mcp

Kemudian sambungkan klien MCP Anda melalui stdio dan uji alat seperti:

  • omniroute_get_health
  • omniroute_list_combos

A2A (untuk alur kerja agen-ke-agen):

curl http://localhost:20128/.well-known/agent.json
curl -X POST http://localhost:20128/a2a \
  -H 'content-type: application/json' \
  -d '{"jsonrpc":"2.0","id":"quickstart","method":"message/send","params":{"skill":"quota-management","messages":[{"role":"user","content":"Give me a short quota summary."}]}}'

5) Validasi semuanya end-to-end (direkomendasikan)

npm run test:protocols:e2e

Suite ini memvalidasi alur klien MCP dan A2A yang sebenarnya terhadap aplikasi yang sedang berjalan.

Alternatif: dijalankan dari sumber

cp .env.example .env
npm install
PORT=20128 DASHBOARD_PORT=20129 NEXT_PUBLIC_BASE_URL=http://localhost:20129 npm run dev
Void Linux (`xbps-src` template)

Untuk pengguna Void Linux, Anda dapat membuat paket asli menggunakan xbps-src. Simpan blok ini sebagai srcpkgs/omniroute/template:

# Template file for 'omniroute'
pkgname=omniroute
version=3.4.1
revision=1
hostmakedepends="nodejs python3 make"
depends="openssl"
short_desc="Universal AI gateway with smart routing for multiple LLM providers"
maintainer="zenobit <zenobit@disroot.org>"
license="MIT"
homepage="https://github.com/diegosouzapw/OmniRoute"
distfiles="https://github.com/diegosouzapw/OmniRoute/archive/refs/tags/v${version}.tar.gz"
checksum=009400afee90a9f32599d8fe734145cfd84098140b7287990183dde45ae2245b
system_accounts="_omniroute"
omniroute_homedir="/var/lib/omniroute"
export NODE_ENV=production
export npm_config_engine_strict=false
export npm_config_loglevel=error
export npm_config_fund=false
export npm_config_audit=false

do_build() {
	# Determine target CPU arch for node-gyp
	local _gyp_arch
	case "$XBPS_TARGET_MACHINE" in
		aarch64*) _gyp_arch=arm64 ;;
		armv7*|armv6*) _gyp_arch=arm ;;
		i686*) _gyp_arch=ia32 ;;
		*) _gyp_arch=x64 ;;
	esac

	# 1) Install all deps  skip scripts (no network in do_build, native modules
	#    compiled separately below; better-sqlite3 is serverExternalPackage so
	#    Next.js does not execute it during next build)
	NODE_ENV=development npm ci --ignore-scripts

	# 2) Build the Next.js standalone bundle
	npm run build

	# 3) Copy static assets into standalone
	cp -r .next/static .next/standalone/.next/static
	[ -d public ] && cp -r public .next/standalone/public || true

	# 4) Compile better-sqlite3 native binding for the target architecture.
	#    Use node-gyp directly so CC/CXX from xbps-src cross-toolchain are used
	#    without npm altering them.
	local _node_gyp=/usr/lib/node_modules/npm/node_modules/node-gyp/bin/node-gyp.js
	(cd node_modules/better-sqlite3 && node "$_node_gyp" rebuild --arch="$_gyp_arch")

	# 5) Place the compiled binding into the standalone bundle
	local _bs3_release=.next/standalone/node_modules/better-sqlite3/build/Release
	mkdir -p "$_bs3_release"
	cp node_modules/better-sqlite3/build/Release/better_sqlite3.node "$_bs3_release/"

	# 6) Remove arch-specific sharp bundles  upstream sets images.unoptimized=true
	#    so sharp is not used at runtime; x64 .so files would break aarch64 strip
	rm -rf .next/standalone/node_modules/@img

	# 7) Copy pino runtime deps omitted by Next.js static analysis:
	#    pino-abstract-transport  required by pino's worker thread
	#    split2  dep of pino-abstract-transport
	#    process-warning  dep of pino itself
	for _mod in pino-abstract-transport split2 process-warning; do
		cp -r "node_modules/$_mod" .next/standalone/node_modules/
	done
}

do_check() {
	npm run test:unit
}

do_install() {
	vmkdir usr/lib/omniroute/.next

	vcopy .next/standalone/. usr/lib/omniroute/.next/standalone

	# Prevent removal of empty Next.js app router dirs by the post-install hook
	for _d in \
		.next/standalone/.next/server/app/dashboard \
		.next/standalone/.next/server/app/dashboard/settings \
		.next/standalone/.next/server/app/dashboard/providers; do
		touch "${DESTDIR}/usr/lib/omniroute/${_d}/.keep"
	done

	cat > "${WRKDIR}/omniroute" <<'EOF'
#!/bin/sh
export PORT="${PORT:-20128}"
export DATA_DIR="${DATA_DIR:-${XDG_DATA_HOME:-${HOME}/.local/share}/omniroute}"
export APP_LOG_TO_FILE="${APP_LOG_TO_FILE:-false}"
mkdir -p "${DATA_DIR}"
exec node /usr/lib/omniroute/.next/standalone/server.js "$@"
EOF
	vbin "${WRKDIR}/omniroute"
}

post_install() {
	vlicense LICENSE
}

🐳 Docker

OmniRoute tersedia sebagai image Docker publik di Docker Hub.

Quick run:

docker run -d \
  --name omniroute \
  --restart unless-stopped \
  --stop-timeout 40 \
  -p 20128:20128 \
  -v omniroute-data:/app/data \
  diegosouzapw/omniroute:latest

Dengan file lingkungan:

# Copy and edit .env first
cp .env.example .env

docker run -d \
  --name omniroute \
  --restart unless-stopped \
  --stop-timeout 40 \
  --env-file .env \
  -p 20128:20128 \
  -v omniroute-data:/app/data \
  diegosouzapw/omniroute:latest

Menggunakan Docker Tulis:

# Base profile (no CLI tools)
docker compose --profile base up -d

# CLI profile (Claude Code, Codex, OpenClaw built-in)
docker compose --profile cli up -d

Dukungan dasbor untuk penerapan Docker kini mencakup Cloudflare Quick Tunnel sekali klik di Dashboard → Endpoints. Yang pertama mengaktifkan pengunduhan cloudflared hanya bila diperlukan, memulai terowongan sementara ke titik akhir /v1 Anda saat ini, dan menampilkan URL https://*.trycloudflare.com/v1 yang dihasilkan langsung di bawah URL publik normal Anda.

Notes:

  • URL Terowongan Cepat bersifat sementara dan berubah setelah setiap restart.
  • Terowongan Cepat tidak dipulihkan secara otomatis setelah OmniRoute atau kontainer dimulai ulang. Aktifkan kembali dari dasbor bila diperlukan.
  • Penginstalan terkelola saat ini mendukung Linux, macOS, dan Windows di x64 / arm64.
  • Terkelola Quick Tunnels default ke transportasi HTTP/2 untuk menghindari peringatan buffer UDP QUIC yang berisik di lingkungan kontainer yang terbatas. Setel CLOUDFLARED_PROTOCOL=quic atau auto jika Anda menginginkan transportasi lain.
  • Gambar Docker menggabungkan akar CA sistem dan meneruskannya ke cloudflared yang dikelola, yang menghindari kegagalan kepercayaan TLS ketika terowongan melakukan bootstrap di dalam wadah.
  • SQLite berjalan dalam mode WAL. docker stop harus dibiarkan selesai sehingga OmniRoute dapat memeriksa kembali perubahan terbaru ke storage.sqlite.
  • File Compose yang dibundel sudah menetapkan masa tenggang penghentian 40 detik. Jika Anda menjalankan image secara langsung, pertahankan --stop-timeout 40 (atau serupa) sehingga penghentian manual tidak menghentikan pembersihan pematian.
  • Setel CLOUDFLARED_BIN=/absolute/path/to/cloudflared jika Anda ingin OmniRoute menggunakan biner yang sudah ada alih-alih mengunduhnya.

Menggunakan Docker Compose dengan Caddy (HTTPS Auto-TLS):

OmniRoute dapat diekspos dengan aman menggunakan penyediaan SSL otomatis Caddy. Pastikan data DNS A domain Anda mengarah ke IP server Anda.

services:
  omniroute:
    image: diegosouzapw/omniroute:latest
    container_name: omniroute
    restart: unless-stopped
    volumes:
      - omniroute-data:/app/data
    environment:
      - PORT=20128
      - NEXT_PUBLIC_BASE_URL=https://your-domain.com

  caddy:
    image: caddy:latest
    container_name: caddy
    restart: unless-stopped
    ports:
      - "80:80"
      - "443:443"
    command: caddy reverse-proxy --from https://your-domain.com --to http://omniroute:20128

volumes:
  omniroute-data:
Image Tag Size Description
diegosouzapw/omniroute latest ~250MB Latest stable release
diegosouzapw/omniroute 3.6.2 ~250MB Current version

🖥️ Aplikasi Desktop — Offline & Selalu Aktif

🆕 BARU! OmniRoute kini tersedia sebagai aplikasi desktop asli untuk Windows, macOS, dan Linux.

Jalankan OmniRoute sebagai aplikasi desktop mandiri — tanpa terminal, tanpa browser, tanpa internet untuk model lokal. Aplikasi berbasis Electron meliputi:

  • 🖥️ Jendela Asli — Jendela aplikasi khusus dengan integrasi baki sistem
  • 🔄 Mulai Otomatis — Luncurkan OmniRoute saat login sistem
  • 🔔 Pemberitahuan Asli — Dapatkan peringatan jika kuota habis atau masalah penyedia
  • Instal Sekali Klik — NSIS (Windows), DMG (macOS), AppImage (Linux)
  • 🌐 Mode Offline — Bekerja sepenuhnya offline dengan server yang dibundel

Mulai Cepat

# Development mode
npm run electron:dev

# Build for your platform
npm run electron:build         # Current platform
npm run electron:build:win     # Windows (.exe)
npm run electron:build:mac     # macOS (.dmg) — x64 & arm64
npm run electron:build:linux   # Linux (.AppImage)

System Tray

Saat diminimalkan, OmniRoute ada di baki sistem Anda dengan tindakan cepat:

  • Buka dasbor
  • Ubah port server
  • Keluar dari aplikasi

📖 Full documentation: electron/README.md


💰 Harga Sekilas

Tier Provider Cost Quota Reset Best For
💳 SUBSCRIPTION Claude Code (Pro) $20/mo 5h + weekly Already subscribed
Codex (Plus/Pro) $20-200/mo 5h + weekly OpenAI users
GitHub Copilot $10-19/mo Monthly GitHub users
🔑 API KEY NVIDIA NIM GRATIS (pengembangan selamanya) ~40 RPM 70+ open models
Cerebras FREE (1M tok/day) 60K TPM / 30 RPM World's fastest
Groq FREE (30 RPM) 14.4K RPD Ultra-fast Llama/Gemma
DeepSeek V3.2 $0.27/$1.10 per 1M None Best price/quality reasoning
xAI Grok-4 Fast $0.20/$0.50 per 1M 🆕 None Fastest + tool calling, ultralow
xAI Grok-4 (standard) $0.20/$1.50 per 1M 🆕 None Penalaran andalan dari xAI
Mistral Uji coba gratis + berbayar Rate limited European AI
OpenRouter Bayar per penggunaan None 100+ models aggr.
💰 CHEAP GLM-5 (via Z.AI) 🆕 $0.5/1M Daily 10AM 128K output, newest flagship
GLM-4.7 $0.6/1M Daily 10AM Budget backup
MiniMax M2.5 🆕 $0.3/1M input 5-hour rolling Reasoning + agentic tasks
MiniMax M2.1 $0.2/1M 5-hour rolling Cheapest option
Kimi K2.5 (Moonshot API) 🆕 Bayar per penggunaan None Direct Moonshot API access
Kimi K2 $9/mo flat 10M tokens/mo Predictable cost
🆓 FREE Qoder $0 Unlimited 5 models unlimited
Qwen $0 Unlimited 4 models unlimited
Kiro $0 Unlimited Claude Sonnet/Haiku (AWS Builder)
LongCat Flash-Lite 🆕 $0 (50M tok/day 🔥) 1 RPS Kuota gratis terbesar di dunia
Pollinations AI 🆕 $0 (tidak perlu kunci) 1 req/15s GPT-5, Claude, DeepSeek, Llama 4
Cloudflare Workers AI 🆕 $0 (10K Neurons/day) ~150 resp/day 50+ model, keunggulan global
Scaleway AI 🆕 $0 (1M tokens total) Rate limited EU/GDPR, Qwen3 235B, Llama 70B

🆕 Model baru ditambahkan (Mar 2026): Keluarga Grok-4 Fast seharga $0,20/$0,50/M (dibandingkan pada 1143ms — 30% lebih cepat dibandingkan Gemini 2.5 Flash), GLM-5 melalui Z.AI dengan output 128K, penalaran MiniMax M2.5, harga DeepSeek V3.2 yang diperbarui, Kimi K2.5 melalui API langsung Moonshot.

💡 Tumpukan Kombo $0 — Penyiapan Gratis Lengkap:

# 🆓 Ultimate Free Stack 2026 — 11 Providers, $0 Forever
Kiro (kr/)             → Claude Sonnet/Haiku UNLIMITED
Qoder (if/)            → kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 UNLIMITED
LongCat Lite (lc/)     → LongCat-Flash-Lite — 50M tokens/day 🔥
Pollinations (pol/)    → GPT-5, Claude, DeepSeek, Llama 4 — no key needed
Qwen (qw/)             → qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next UNLIMITED
Gemini (gemini/)       → Gemini 2.5 Flash — 1,500 req/day free API key
Cloudflare AI (cf/)    → Llama 70B, Gemma 3, Mistral — 10K Neurons/day
Scaleway (scw/)        → Qwen3 235B, Llama 70B — 1M free tokens (EU)
Groq (groq/)           → Llama/Gemma ultra-fast — 14.4K req/day
NVIDIA NIM (nvidia/)   → 70+ open models — 40 RPM forever
Cerebras (cerebras/)   → Llama/Qwen world-fastest — 1M tok/day

Tanpa biaya. Jangan pernah berhenti melakukan pengkodean. Konfigurasikan ini sebagai satu kombo OmniRoute dan semua fallback terjadi secara otomatis — tidak pernah ada peralihan manual.



🆓 Model Gratis — Apa yang Sebenarnya Anda Dapatkan

Semua model di bawah 100% gratis tanpa memerlukan kartu kredit. OmniRoute melakukan rute otomatis di antara keduanya ketika satu kuota habis — gabungkan semuanya untuk kombo $0 yang tidak dapat dipecahkan.

🔵 MODEL CLAUDE (melalui Kiro — ID AWS Builder)

Model Prefix Limit Rate Limit
claude-sonnet-4.5 kr/ Unlimited No reported daily cap
claude-haiku-4.5 kr/ Unlimited No reported daily cap
claude-opus-4.6 kr/ Unlimited Latest Opus via Kiro

🟢 MODEL QODER (PAT Gratis melalui qodercli)

Model Prefix Limit Rate Limit
kimi-k2-thinking if/ Unlimited No reported cap
qwen3-coder-plus if/ Unlimited No reported cap
deepseek-r1 if/ Unlimited No reported cap
minimax-m2.1 if/ Unlimited No reported cap
kimi-k2 if/ Unlimited No reported cap

Metode koneksi yang disarankan: Token Akses Pribadi + qodercli. Peramban OAuth adalah eksperimental dan dinonaktifkan secara default kecuali variabel lingkungan QODER_OAUTH_* dikonfigurasi.

🟡 MODEL QWEN (Otentikasi Kode Perangkat)

Model Prefix Limit Rate Limit
qwen3-coder-plus qw/ Unlimited No reported cap
qwen3-coder-flash qw/ Unlimited No reported cap
qwen3-coder-next qw/ Unlimited No reported cap
vision-model qw/ Unlimited Multimodal (images)

NVIDIA NIM (Kunci API Gratis — build.nvidia.com)

Tier Daily Limit Rate Limit Notes
Free (Dev) No token cap ~40 RPM 70+ model; transisi ke batas tarif murni pada pertengahan tahun 2025

Model gratis populer: moonshotai/kimi-k2.5 (Kimi K2.5), z-ai/glm4.7 (GLM 4.7), deepseek-ai/deepseek-v3.2 (DeepSeek V3.2), nvidia/llama-3.3-70b-instruct, deepseek/deepseek-r1

CEREBRAS (Kunci API Gratis — inference.cerebras.ai)

Tier Daily Limit Rate Limit Notes
Free 1M tokens/day 60K TPM / 30 RPM World's fastest LLM inference; resets daily

Available free: llama-3.3-70b, llama-3.1-8b, deepseek-r1-distill-llama-70b

🔴 GROQ (Kunci API Gratis — console.groq.com)

Tier Daily Limit Rate Limit Notes
Free 14.4K RPD 30 RPM per model No credit card; 429 on limit, not charged

Available free: llama-3.3-70b-versatile, gemma2-9b-it, mixtral-8x7b, whisper-large-v3

🔴 LONGCAT AI (Kunci API Gratis — longcat.chat) 🆕

Model Prefix Kuota Gratis Harian Notes
LongCat-Flash-Lite lc/ 50M tokens 💥 Kuota gratis terbesar yang pernah ada
LongCat-Flash-Chat lc/ 500K tokens Multi-turn chat
LongCat-Flash-Thinking lc/ 500K tokens Reasoning / CoT
LongCat-Flash-Thinking-2601 lc/ 500K tokens Jan 2026 version
LongCat-Flash-Omni-2603 lc/ 500K tokens Multimodal

100% gratis saat dalam versi beta publik. Daftar di longcat.chat dengan email atau telepon. Reset setiap hari pukul 00:00 UTC.

🟢 POLLINASI AI (Tidak Perlu Kunci API) 🆕

Model Prefix Rate Limit Provider Behind
openai pol/ 1 req/15s GPT-5
claude pol/ 1 req/15s Anthropic Claude
gemini pol/ 1 req/15s Google Gemini
deepseek pol/ 1 req/15s DeepSeek V3
llama pol/ 1 req/15s Meta Llama 4 Scout
mistral pol/ 1 req/15s Mistral AI

Tanpa gesekan: Tanpa pendaftaran, tanpa kunci API. Tambahkan penyedia Penyerbukan dengan bidang kunci kosong dan itu langsung berfungsi.

🟠 AI CLOUDFLARE WORKERS (Kunci API Gratis — cloudflare.com) 🆕

Tier Daily Neurons Equivalent Usage Notes
Free 10,000 ~150 LLM resp / 500s audio / 15K embeds Keunggulan global, 50+ model

Model gratis populer: @cf/meta/llama-3.3-70b-instruct, @cf/google/gemma-3-12b-it, @cf/openai/whisper-large-v3-turbo (audio gratis!), @cf/qwen/qwen2.5-coder-15b-instruct

Membutuhkan Token API + ID Akun dari dash.cloudflare.com. Simpan ID Akun di pengaturan penyedia.

🟣 SCALEWAY AI (1 Juta Token Gratis — scaleway.com) 🆕

Tier Free Quota Location Notes
Free 1M tokens 🇫🇷 Paris, EU No credit card needed within limits

Tersedia gratis: qwen3-235b-a22b-instruct-2507 (Qwen3 235B!), llama-3.1-70b-instruct, mistral-small-3.2-24b-instruct-2506, deepseek-v3-0324

Sesuai dengan UE/GDPR. Dapatkan kunci API di console.scaleway.com.

💡 Tumpukan Gratis Terbaik (11 Penyedia, $0 Selamanya):

Kiro (kr/)             → Claude Sonnet/Haiku TANPA BATAS
Qoder (if/)            → kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 TANPA BATAS
LongCat Lite (lc/)     → LongCat-Flash-Lite — 50 juta token/hari 🔥
Pollinations (pol/)    → GPT-5, Claude, DeepSeek, Llama 4 — tidak perlu kunci
Qwen (qw/)             → model qwen3-coder TANPA BATAS
Gemini (gemini/)       → Gemini 2.5 Flash — 1.500 req/hari gratis
Cloudflare AI (cf/)    → 50+ model — 10 ribu Neurons/hari
Scaleway (scw/)        → Qwen3 235B, Llama 70B — 1 juta token gratis (EU)
Groq (groq/)           → Llama/Gemma — 14,4 ribu req/hari, sangat cepat
NVIDIA NIM (nvidia/)   → 70+ model terbuka — 40 RPM selamanya
Cerebras (cerebras/)   → Llama/Qwen tercepat di dunia — 1 juta tok/hari

🎙️ Kombo Transkripsi Gratis

Transkripsikan audio/video apa pun seharga $0 — Deepgram memimpin dengan $200 gratis, penggantian AssemblyAI $50, Groq Whisper sebagai cadangan darurat tanpa batas.

Provider Free Credits Best Model Rate Limit
🟢 Deepgram $200 free (signup) nova-3 — best accuracy, 30+ languages Tidak ada batasan RPM pada kredit gratis
🔵 AssemblyAI $50 free (signup) universal-3-pro — chapters, sentiment, PII Tidak ada batasan RPM pada kredit gratis
🔴 Groq Free forever whisper-large-v3 — OpenAI Whisper 30 RPM (rate limited)

Suggested combo in /dashboard/combos:

Name: free-transcription
Strategy: Priority
Nodes:
  [1] deepgram/nova-3          → uses $200 free first
  [2] assemblyai/universal-3-pro → fallback when Deepgram credits run out
  [3] groq/whisper-large-v3    → free forever, emergency fallback

Kemudian di tab /dashboard/mediaTranskripsi: unggah file audio atau video apa pun → pilih titik akhir kombo Anda → dapatkan transkripsi dalam format yang didukung.

💡 Fitur Utama

OmniRoute v3.6 dibangun sebagai platform operasional, bukan hanya proxy relai.

🆕 Baru — Sorotan v3.6.x (Apr 2026)

Feature Apa Fungsinya
🌐 V1 WebSocket Bridge Lalu lintas WebSocket yang kompatibel dengan OpenAI ditingkatkan dan diproksi melalui /v1/ws — streaming penuh melalui WS dengan autentikasi sesi (kunci API atau cookie sesi)
🔑 Sync Tokens & Config Bundle Menerbitkan/mencabut token sinkronisasi untuk titik akhir sinkronisasi konfigurasi. Paket konfigurasi diversi dengan ETag untuk polling hemat bandwidth
🧠 GLM Thinking (glmt) Preset GLM Thinking registered first-class: 65 536 max tokens, 24 576 thinking budget, 900s timeout, usage sync & pricing — Claude-compatible API
🔢 Hybrid Token Counting Menggunakan /messages/count_tokens sisi penyedia jika tersedia; kembali ke perkiraan — pelacakan penggunaan yang akurat tanpa menebak-nebak
🌱 Model Alias Benih Otomatis 30+ cross-proxy dialect aliases normalised at startup — no more routing mismatches
🛡️ Safe Outbound Fetch Semua validasi penyedia dan penemuan model melalui lapisan pengambilan yang dilindungi yang memblokir URL pribadi/lokal dengan percobaan ulang, batas waktu, dan perlindungan SSRF
Tunggu Masa Tenang Percobaan ulang obrolan sisi server ketika setiap koneksi kandidat sedang dingin; dapat dikonfigurasi enabled, maxRetries, dan maxRetryWaitSec
🔍 Validasi Env Runtime Startup memvalidasi semua env vars dengan skema Zod - menghapus kesalahan untuk rahasia yang hilang, URL yang tidak valid, atau tipe yang salah
📋 Compliance Audit Expansion Log audit terstruktur dengan penomoran halaman, konteks permintaan, peristiwa autentikasi, peristiwa CRUD penyedia, dan pencatatan validasi yang diblokir SSRF
🔐 TPS Log Metric Modal detail log menunjukkan Token Per Second (TPS) — sekilas kinerja cepat untuk setiap permintaan
🗑️ Uninstall / Full Uninstall npm run uninstall menyimpan data, npm run uninstall:full menghapus semuanya — penghapusan bersih untuk semua metode instalasi
🔧 OAuth Env Repair Tindakan "Perbaiki env" sekali klik untuk penyedia OAuth memulihkan vars env yang hilang dan memperbaiki status autentikasi yang rusak
🔒 Pematian Elektron yang Anggun Electron before-quit dimatikan Next.js dengan baik, mencegah penguncian database SQLite WAL pada penutupan desktop
👁️ Pengalih Visibilitas Model Pengalih visibilitas per model (ikon 👁) dengan filter pencarian dan lencana jumlah aktif (N/M active) di halaman penyedia
📧 Email Privacy Masking OAuth account emails masked (di*****@g****.com), full address visible on hover
🔗 Context Relay Strategy Strategi kombo menjaga kesinambungan sesi melalui ringkasan penyerahan terstruktur saat akun dirotasi di tengah percakapan
🛡️ Proxy Hardening Pemeriksaan kesehatan token, validasi kunci API, dan operator undici semuanya menghormati konfigurasi proxy
⚠️ Node.js 24 Login Warning Login page proactively detects incompatible Node.js versions and shows a clear warning banner
📎 Gemini PDF Attachments PDF attachments correctly routed to Gemini via inline_data and generic base64 detection
🔒 Pengerasan Keamanan CodeQL Resolved SSRF, insecure randomness, polynomial ReDoS, and incomplete URL sanitization alerts

🆕 Baru — Peningkatan Terinspirasi ClawRouter (Mar 2026)

Feature Apa Fungsinya
Grok-4 Fast Family xAI models at $0.20/$0.50/M — benchmarked 1143ms (30% faster than Gemini 2.5 Flash)
🧠 GLM-5 via Z.AI Konteks keluaran 128 ribu, $0,5/1 juta — andalan terbaru dari keluarga GLM
🔮 MiniMax M2.5 Penalaran + tugas agen seharga $0,30/1 juta — peningkatan signifikan dari M2.1
🎯 alat Memanggil Bendera per Model Per model toolCalling: true/false di registri — AutoCombo melewatkan model yang tidak mendukung alat
🌍 Multilingual Intent Detection Kata kunci PT/ZH/ES/AR dalam penilaian AutoCombo — pemilihan model yang lebih baik untuk konten non-Inggris
📊 Benchmark-Driven Fallbacks Latensi p95 nyata dari penilaian kombo umpan permintaan langsung — AutoCombo belajar dari data aktual
🔁 Request Deduplication Content-hash based dedup window — multi-agent safe, prevents duplicate charges
🔌 Pluggable RouterStrategy Antarmuka RouterStrategy yang dapat diperluas — tambahkan logika perutean khusus sebagai plugin

🚀 Sebelumnya v2.0.9+ — Playground, Fingerprint CLI & ACP

Feature Apa Fungsinya
🎮 Model Playground Halaman dasbor untuk menguji model apa pun secara langsung — pemilih penyedia/model/titik akhir, Editor Monaco, streaming, batalkan, pengaturan waktu
🔏 CLI Fingerprint Matching Pengurutan header/isi per penyedia agar sesuai dengan tanda tangan CLI asli — alihkan per penyedia di Pengaturan > Keamanan. IP proxy Anda dipertahankan
🤖 Dasbor Agen ACP Debug Halaman agen — kisi 14 agen dengan status pemasangan, versi, formulir agen khusus untuk alat CLI apa pun. Pengguna OpenCode mendapatkan tombol "Unduh opencode.json" yang secara otomatis menghasilkan konfigurasi siap pakai dengan semua model yang tersedia.
🔧 Custom Model apiFormat Routing Model khusus dengan apiFormat: "responses" sekarang dirutekan dengan benar ke penerjemah Responses API
🏢 Codex Workspace Isolation Multiple Codex workspaces per email — OAuth correctly separates connections by workspace ID
🔄 Pembaruan Otomatis Elektron Aplikasi desktop memeriksa pembaruan + instal otomatis saat restart

🤖 Operasi Agen & Protokol (v2.0)

Feature Apa Fungsinya
🔧 Server MCP (25 alat) IDE/agent tools via 3 transports: stdio, SSE (/api/mcp/sse), Streamable HTTP (/api/mcp/stream). 18 core + 3 memory + 4 skill tools
🤝 Server A2A (JSON-RPC + SSE) Eksekusi tugas agen-ke-agen dengan alur sinkronisasi dan streaming
🧭 Halaman Titik Akhir Konsolidasi Halaman manajemen bertab dengan tab Proksi Titik Akhir, MCP, A2A, dan Titik Akhir API
🎚️ Service Enable/Disable Toggles Sakelar ON/OFF untuk MCP dan A2A dengan pengaturan persistensi (default: OFF)
🛰️ Detak Jantung Waktu Proses MCP Real process status (pid, uptime, heartbeat age, transport, scope mode)
📋 MCP Audit Trail Log audit yang dapat difilter dengan keberhasilan/kegagalan dan atribusi kunci
🔐 MCP Scope Enforcement 10 izin cakupan terperinci untuk akses alat terkontrol
📡 Manajemen Siklus Hidup Tugas A2A List/filter tasks, inspect events/artifacts, cancel running tasks
📋 Agent Card Discovery /.well-known/agent.json untuk penemuan otomatis klien
🧪 Protocol E2E Test Harness Klien MCP SDK + A2A asli mengalir di test:protocols:e2e
⚙️ Operational Controls Ganti kombo, sesuaikan pengaturan ketahanan, dan tinjau status pemutus dari permukaan Kesehatan dan Pengaturan khusus

🧠 Routing & Kecerdasan

Feature Apa Fungsinya
🎯 Pengembalian 4 Tingkat Cerdas Rute otomatis: Berlangganan → Kunci API → Murah → Gratis
📊 Pelacakan Kuota Waktu Nyata Jumlah token langsung + setel ulang hitungan mundur per penyedia
🔄 Format Translation OpenAI ↔ Claude ↔ Gemini ↔ Respons dengan konversi skema-aman
👥 Dukungan Multi-Akun Banyak akun per penyedia dengan pilihan cerdas
🔄 Auto Token Refresh Token OAuth disegarkan secara otomatis dengan percobaan ulang
🎨 Custom Combos 13 strategi penyeimbangan + kontrol rantai mundur
🔗 Context Relay Penyerahan kesinambungan sesi ketika rotasi akun terjadi di tengah sesi
🌐 Wildcard Router provider/* dynamic routing
🧠 Thinking Budget Controls Batas penalaran passthrough, otomatis, kustom, dan adaptif
🔀 Model Aliases Alias model khusus + bawaan dan keamanan migrasi
Background Degradation Arahkan tugas latar belakang berprioritas rendah ke model yang lebih murah
🧪 Perutean Cerdas Sadar Tugas Pilih model secara otomatis berdasarkan jenis konten (pengkodean/visi/analisis/ringkasan)
🔄 A2A Agent Workflows Deterministic FSM orchestrator for stateful multi-step agent executions
🔀 Adaptive Routing Dynamic strategy override based on token volume and prompt complexity
🎲 Provider Diversity Shannon entropy scoring balancing auto-combo traffic distribution
💬 System Prompt Injection Global behavior controls applied consistently
📄 Kompatibilitas API Respons Dukungan penuh /v1/responses untuk Codex dan alur kerja agen tingkat lanjut

🎵 API Multi-Modal

Feature Apa Fungsinya
🖼️ Image Generation /v1/images/generations dengan cloud dan backend lokal
📐 Embeddings /v1/embeddings untuk saluran pencarian dan RAG
🎤 Audio Transcription /v1/audio/transcriptions — 7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support
🔊 Text-to-Speech /v1/audio/speech — 10 penyedia (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) dengan pesan kesalahan yang benar
🎬 Video Generation /v1/videos/generations (ComfyUI + SD WebUI workflows)
🎵 Music Generation /v1/music/generations (ComfyUI workflows)
🛡️ Moderations /v1/moderations safety checks
🔀 Reranking /v1/rerank untuk penilaian relevansi
🔍 Web Search 🆕 /v1/search — 5 providers (Serper, Brave, Perplexity, Exa, Tavily), 6,500+ free/month, auto-failover, cache

🛡️ Ketahanan, Keamanan & Tata Kelola

Feature Apa Fungsinya
🔌 Penyedia Pemutus Arus Perjalanan/pemulihan di seluruh penyedia setelah kelelahan fallback dengan ambang batas yang dapat dikonfigurasi
🔒 Kunci Kuota Harian 🆕 Mendeteksi sinyal kelelahan dan mengunci perutean untuk model tertentu hingga tengah malam
🎯 Model Sadar Titik Akhir Model khusus mendeklarasikan titik akhir + format API yang didukung
🛡️ Anti-Thundering Herd Mutex + semaphore protections on retry/rate events
🧠 Semantic + Signature Cache Pengurangan biaya/latensi dengan dua lapisan cache
Request Idempotency Duplicate protection window
🔒 TLS Fingerprint Spoofing Sidik jari TLS seperti browser — mengurangi deteksi bot dan penandaan akun
🔏 CLI Fingerprint Matching Matches native CLI request signatures — reduces ban risk while preserving proxy IP
🌐 IP Filtering Kontrol daftar yang diizinkan/daftar blokir untuk penerapan yang terbuka
🚦 Minta Antrian & Kecepatan Bucket permintaan per koneksi yang dapat dikonfigurasi untuk RPM, spasi, konkurensi, dan waktu tunggu maksimal
📉 Graceful Degradation Multi-layer capability fallbacks protecting core gateway operations
📜 Config Audit Trail Pelacakan perubahan berbasis diff mencegah penyimpangan operasional dengan rollback sederhana
Sinkronisasi Kesehatan Penyedia Proactive token expiration monitoring triggering alerts before authorization failures
❄️ Connection Cooldown Retryable 408/429/5xx failures cool down a single connection with optional upstream hints
🚪 Nonaktifkan Otomatis Akun yang Diblokir Akun token yang diblokir secara permanen dapat dinonaktifkan secara otomatis
🔑 Manajemen Kunci API + Pelingkupan Mengamankan penerbitan/rotasi kunci dan kontrol model/penyedia
👁️ Pengungkapan Kunci API Cakupan 🆕 Ikut serta dalam pemulihan kunci API melalui ALLOW_API_KEY_REVEAL
🛡️ Protected /models Gerbang autentikasi opsional dan penyembunyian penyedia untuk katalog model
🛡️ Safe Outbound Fetch 🆕 Pengambilan yang dijaga untuk panggilan penyedia — memblokir URL pribadi/lokal, percobaan ulang, perlindungan SSRF
Tunggu Cooldown 🆕 Coba ulang obrolan secara otomatis setelah cooldown koneksi; dapat dikonfigurasi enabled, maxRetries, dan maxRetryWaitSec
🔍 Validasi Env Runtime 🆕 Zod-based env schema validation at startup with actionable error messages
📋 Compliance Audit v2 🆕 Penomoran halaman, konteks permintaan, peristiwa autentikasi, CRUD penyedia, dan logging yang diblokir SSRF

📊 Observabilitas & Analitik

Feature Apa Fungsinya
📝 Permintaan + Pencatatan Proksi Permintaan/respons penuh dan pencatatan proksi
📉 Streamed Detailed Logs Merekonstruksi aliran muatan SSE dengan rapi ke dalam UI
🏷️ Lencana Model Real-Time 🆕 Status model langsung dan penghitung waktu mundur kuota harian
📋 Dasbor Log Terpadu Tampilan permintaan, proksi, audit, dan konsol dalam satu halaman
🔍 Request Telemetry latensi p50/p95/p99 dan penelusuran permintaan
🏥 Health Dashboard Uptime, breaker states, lockouts, cache stats
💰 Cost Tracking Kontrol anggaran dan visibilitas harga per model
📈 Analytics Visualizations Wawasan penggunaan model/penyedia dan tampilan tren
🧪 Evaluation Framework Pengujian set emas dengan strategi pencocokan yang dapat dikonfigurasi
📡 Live Diagnostics 🆕 Bypass cache semantik untuk pengujian langsung kombo yang akurat
🔐 TPS Log Metric 🆕 Tokens Per Second badge in log details modal

☁️ Deployment & Platform

Feature Apa Fungsinya
🌐 Deploy Anywhere Localhost, VPS, Docker, Cloud environments
🚇 Cloudflare Tunnel 🆕 Integrasi Quick Tunnel sekali klik dari dasbor
🔑 Pemfilteran Model Kunci API Respons asli /v1/models difilter melalui peran konteks Pembawa yang ditetapkan
Smart Cache Bypass Heuristik TTL yang dapat dikonfigurasi dan kontrol pengambilan ulang paksa
🔄 Backup/Restore Arus ekspor/impor dan pemulihan bencana
🧙 Onboarding Wizard Penyiapan terpandu yang dijalankan pertama kali
🔧 Dasbor Alat CLI Pengaturan sekali klik untuk alat pengkodean populer
🎮 Model Playground Uji penyedia/model/titik akhir apa pun dari dasbor
🔏 CLI Fingerprint Toggle Pencocokan sidik jari per penyedia di Pengaturan > Keamanan
🌐 i18n (30 languages) Dasbor lengkap + dukungan bahasa dokumen dengan cakupan RTL
🧹 Hapus Semua Model Pembersihan daftar model sekali klik di detail penyedia
👁️ Sidebar Controls 🆕 Sembunyikan komponen dan integrasi dari Pengaturan Penampilan
📋 Issue Templates Templat GitHub standar untuk bug dan fitur
📂 Custom Data Directory DATA_DIR penggantian untuk lokasi penyimpanan
🌐 V1 WebSocket Bridge 🆕 OpenAI-compatible WebSocket traffic proxied via /v1/ws
🔑 Sync Tokens & Bundle 🆕 Konfigurasikan token sinkronisasi + titik akhir bundel berversi dengan dukungan ETag

Fitur Penyelaman Mendalam

Penggantian cerdas dengan pengendalian biaya praktis

Combo: "my-coding-stack"
  1. cc/claude-opus-4-7
  2. nvidia/llama-3.3-70b
  3. glm/glm-4.7
  4. if/kimi-k2-thinking

Ketika kuota, tarif, atau kesehatan gagal, OmniRoute secara otomatis berpindah ke kandidat berikutnya tanpa peralihan manual.

Manajemen protokol yang terlihat dan dapat dioperasikan

  • MCP + A2A dapat ditemukan di UI dan dokumen (tidak disembunyikan)
  • API status protokol memaparkan data operasional langsung (/api/mcp/*, /api/a2a/*)
  • Dasbor mencakup tindakan untuk operasi hari ke-2 (pengalihan kombo, pengaturan ulang pemutus, pembatalan tugas)

Workflow translator + validasi

Area Penerjemah meliputi:

  • Playground: request transformation checks
  • Chat Tester: full request/response round-trip
  • Test Bench: multiple cases in one run
  • Live Monitor: real-time traffic view

Ditambah validasi protokol dengan klien nyata melalui npm run test:protocols:e2e.

📖 MCP Server README — Referensi alat, konfigurasi IDE, dan contoh klien

📖 A2A Server README — Keterampilan, metode JSON-RPC, streaming, dan siklus hidup tugas

🧪 Evaluasi (Evals)

OmniRoute menyertakan kerangka evaluasi bawaan untuk menguji kualitas respons LLM terhadap rangkaian emas. Akses melalui Analytics → Evals di dasbor.

Golden Set Bawaan

"OmniRoute Golden Set" yang dimuat sebelumnya berisi kasus uji untuk:

  • Salam, matematika, geografi, pembuatan kode
  • Kepatuhan format JSON, terjemahan, pembuatan penurunan harga
  • Penolakan keamanan (konten berbahaya), penghitungan, logika boolean

Strategi Evaluasi

Strategy Description Example
exact Output must match exactly "4"
contains Output must contain substring (case-insensitive) "Paris"
regex Output must match regex pattern "1.*2.*3"
custom Custom JS function returns true/false (output) => output.length > 10

📖 Panduan Setup

Pengaturan Protokol (MCP + A2A)

🧩 Penyiapan MCP (Protokol Konteks Model)

Start MCP transport in stdio mode:

omniroute --mcp

Recommended validation flow:

  1. Hubungkan klien MCP Anda melalui stdio.
  2. Jalankan omniroute_get_health.
  3. Jalankan omniroute_list_combos.
  4. Buka /dashboard/mcp untuk mengonfirmasi detak jantung, aktivitas, dan audit.

API yang berguna untuk otomatisasi:

  • GET /api/mcp/status
  • GET /api/mcp/tools
  • GET /api/mcp/audit
  • GET /api/mcp/audit/stats
🤝 Pengaturan A2A (Agen2Agen)

Temukan agennya:

curl http://localhost:20128/.well-known/agent.json

Send a task:

curl -X POST http://localhost:20128/a2a \
  -H 'content-type: application/json' \
  -d '{"jsonrpc":"2.0","id":"setup-a2a","method":"message/send","params":{"skill":"quota-management","messages":[{"role":"user","content":"Summarize quota status."}]}}'

Manage lifecycle:

  • GET /api/a2a/status
  • GET /api/a2a/tasks
  • GET /api/a2a/tasks/:id
  • POST /api/a2a/tasks/:id/cancel

Operational UI:

  • /dashboard/a2a untuk observasi tugas/status/aliran dan tindakan asap
🧪 Validasi protokol end-to-end

Validasi kedua protokol dengan klien nyata:

npm run test:protocols:e2e

This verifies:

  • Koneksi/daftar/panggilan klien MCP SDK
  • Penemuan A2A/kirim/aliran/dapatkan/batalkan
  • Periksa silang data dalam audit MCP dan API manajemen tugas A2A
💳 Penyedia Berlangganan

Claude Code (Pro/Max)

Dashboard → Providers → Connect Claude Code
→ OAuth login → Auto token refresh
→ 5-hour + weekly quota tracking

Models:
  cc/claude-opus-4-7
  cc/claude-sonnet-4-5-20250929
  cc/claude-haiku-4-5-20251001

Kiat Pro: Gunakan Opus untuk tugas kompleks, Soneta untuk kecepatan. OmniRoute melacak kuota per model!

OpenAI Codex (Plus/Pro)

Dashboard → Providers → Connect Codex
→ OAuth login (port 1455)
→ 5-hour + weekly reset

Models:
  cx/gpt-5.2-codex
  cx/gpt-5.1-codex-max

Manajemen Batas Akun Codex (5 jam + Mingguan)

Setiap akun Codex kini memiliki kebijakan yang dapat diubah di Dashboard -> Providers:

  • 5h (ON/OFF): menerapkan kebijakan ambang jendela 5 jam.
  • Weekly (ON/OFF): menerapkan kebijakan ambang jendela mingguan.
  • Perilaku ambang batas: ketika jendela yang diaktifkan mencapai >=90% penggunaan, akun tersebut dilewati.
  • Perilaku rotasi: OmniRoute merutekan ke akun Codex berikutnya yang memenuhi syarat secara otomatis.
  • Perilaku reset: ketika waktu resetAt penyedia telah berlalu, akun akan memenuhi syarat lagi secara otomatis.

Scenarios:

  • 5h ON + Weekly ON: akun dilewati ketika salah satu jendela mencapai ambang batas.
  • 5h OFF + Weekly ON: hanya penggunaan mingguan yang dapat memblokir akun.
  • 5h ON + Weekly OFF : hanya penggunaan 5 jam yang dapat memblokir akun.
  • resetAt lolos: akun masuk kembali ke rotasi secara otomatis (tidak ada pengaktifan ulang secara manual).

GitHub Copilot

Dashboard → Providers → Connect GitHub
→ OAuth via GitHub
→ Monthly reset (1st of month)

Models:
  gh/gpt-5
  gh/claude-4.5-sonnet
  gh/gemini-3.1-pro-preview
🔑 Penyedia Kunci API

NVIDIA NIM (akses pengembang GRATIS — 70+ model)

  1. Daftar: build.nvidia.com
  2. Dapatkan kunci API gratis (termasuk 1000 kredit inferensi)
  3. Dasbor → Tambah Penyedia → NVIDIA NIM:
    • Kunci API: nvapi-your-key

Model: nvidia/llama-3.3-70b-instruct, nvidia/mistral-7b-instruct, dan 50+ lainnya

Kiat Pro: API yang kompatibel dengan OpenAI — bekerja secara lancar dengan terjemahan format OmniRoute!

DeepSeek

  1. Daftar: platform.deepseek.com
  2. Dapatkan kunci API
  3. Dasbor → Tambah Penyedia → DeepSeek

Models: deepseek/deepseek-chat, deepseek/deepseek-coder

Groq (Tersedia Tingkat Gratis!)

  1. Daftar: console.groq.com
  2. Dapatkan kunci API (termasuk tingkat gratis)
  3. Dasbor → Tambah Penyedia → Groq

Models: groq/llama-3.3-70b, groq/mixtral-8x7b

Pro Tip: Ultra-fast inference — best for real-time coding!

OpenRouter (100+ Model)

  1. Daftar: openrouter.ai
  2. Dapatkan kunci API
  3. Dasbor → Tambah Penyedia → OpenRouter

Model: Akses 100+ model dari semua penyedia utama melalui satu kunci API.

Perilaku dasbor: Model OpenRouter dikelola dari Model yang Tersedia. Penambahan manual, impor, dan sinkronisasi otomatis semuanya memperbarui daftar yang sama.

💰 Penyedia Murah (Cadangan)

GLM-4.7 (Reset harian, $0.6/1M)

  1. Daftar: Zhipu AI
  2. Dapatkan kunci API dari Coding Plan
  3. Dasbor → Tambahkan Kunci API:
    • Penyedia: glm
    • Kunci API: your-key

Use: glm/glm-4.7

Tips Pro: Paket Coding menawarkan 3× kuota dengan biaya 1/7! Reset setiap hari pukul 10.00.

MiniMax M2.1 (Reset 5 jam, $0.20/1M)

  1. Daftar: MiniMax
  2. Dapatkan kunci API
  3. Dasbor → Tambahkan Kunci API

Use: minimax/MiniMax-M2.1

Kiat Pro: Opsi termurah untuk konteks panjang (1 juta token)!

Kimi K2 ($9/bulan flat)

  1. Berlangganan: Moonshot AI
  2. Dapatkan kunci API
  3. Dasbor → Tambahkan Kunci API

Use: kimi/kimi-latest

Kiat Pro: Memperbaiki $9/bulan untuk 10 juta token = biaya efektif $0,90/1 juta!

🆓 Penyedia GRATIS (Cadangan Darurat)

Qoder (5 model GRATIS melalui OAuth)

Dashboard → Connect Qoder
→ Qoder OAuth login
→ Unlimited usage

Models:
  if/kimi-k2-thinking
  if/qwen3-coder-plus
  if/glm-4.7
  if/minimax-m2
  if/deepseek-r1

Qwen (4 model GRATIS melalui Kode Perangkat)

Dashboard → Connect Qwen
→ Device code authorization
→ Unlimited usage

Models:
  qw/qwen3-coder-plus
  qw/qwen3-coder-flash

Kiro (Claude GRATIS)

Dashboard → Connect Kiro
→ AWS Builder ID or Google/GitHub
→ Unlimited usage

Models:
  kr/claude-sonnet-4.5
  kr/claude-haiku-4.5
🎨 Membuat Combo

Contoh 1: Maksimalkan Langganan → Cadangan Murah

Dashboard → Combos → Create New

Name: premium-coding
Models:
  1. cc/claude-opus-4-7 (Subscription primary)
  2. glm/glm-4.7 (Cheap backup, $0.6/1M)
  3. minimax/MiniMax-M2.1 (Cheapest fallback, $0.20/1M)

Use in CLI: premium-coding

Contoh 2: Gratis Saja (Tanpa Biaya)

Name: free-combo
Models:
  1. if/kimi-k2-thinking (unlimited)
  2. qw/qwen3-coder-plus (unlimited)

Cost: $0 forever!
🔧 Integrasi CLI

Cursor IDE

Settings → Models → Advanced:
  OpenAI API Base URL: http://localhost:20128/v1
  OpenAI API Key: [from OmniRoute dashboard]
  Model: cc/claude-opus-4-7

Claude Code

Gunakan halaman Alat CLI di dasbor untuk konfigurasi sekali klik, atau edit ~/.claude/settings.json secara manual.

Codex CLI

export OPENAI_BASE_URL="http://localhost:20128"
export OPENAI_API_KEY="your-omniroute-api-key"

codex "your prompt"

OpenClaw

Opsi 1 — Dasbor (disarankan):

Dashboard → CLI Tools → OpenClaw → Select Model → Apply

Option 2 — Manual: Edit ~/.openclaw/openclaw.json:

{
  "models": {
    "providers": {
      "omniroute": {
        "baseUrl": "http://127.0.0.1:20128/v1",
        "apiKey": "sk_omniroute",
        "api": "openai-completions"
      }
    }
  }
}

Catatan: OpenClaw hanya berfungsi dengan OmniRoute lokal. Gunakan 127.0.0.1 alih-alih localhost untuk menghindari masalah resolusi IPv6.

Cline / Continue / RooCode

Settings → API Configuration:
  Provider: OpenAI Compatible
  Base URL: http://localhost:20128/v1
  API Key: [from OmniRoute dashboard]
  Model: if/kimi-k2-thinking

OpenCode

Langkah 1: Tambahkan OmniRoute sebagai penyedia khusus:

opencode
/connect
# Select "Other" → Enter ID: "omniroute" → Enter your OmniRoute API key

Langkah 2: Buat/edit opencode.json di root proyek Anda:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "omniroute": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "OmniRoute",
      "options": {
        "baseURL": "http://localhost:20128/v1"
      },
      "models": {
        "cc/claude-sonnet-4-20250514": { "name": "Claude Sonnet 4" },
        "gg/gemini-2.5-pro": { "name": "Gemini 2.5 Pro" },
        "if/kimi-k2-thinking": { "name": "Kimi K2 (Free)" }
      }
    }
  }
}

Langkah 3: Pilih model di OpenCode:

/models
# Select any OmniRoute model from the list

Tips: Tambahkan model apa pun yang tersedia di titik akhir OmniRoute /v1/models Anda ke bagian models. Gunakan format provider/model-id dari dasbor OmniRoute Anda.


Pemecahan Masalah

Klik untuk memperluas panduan pemecahan masalah

"Model bahasa tidak memberikan pesan"

  • Kuota penyedia habis → Periksa dashboard pelacak kuota
  • Solusi: Gunakan combo fallback atau beralih ke tier yang lebih murah

Rate limiting

  • Kuota berlangganan habis → Penggantian ke GLM/MiniMax
  • Tambahkan kombo: cc/claude-opus-4-7 → glm/glm-4.7 → if/kimi-k2-thinking

OAuth token expired

  • Disegarkan secara otomatis oleh OmniRoute
  • Jika masalah terus berlanjut: Dasbor → Penyedia → Sambungkan kembali

High costs

  • Periksa statistik penggunaan di Dashboard → Biaya
  • Ganti model utama ke GLM/MiniMax

Port dasbor/API salah

  • PORT adalah port dasar kanonik (dan port API secara default)
  • API_PORT hanya menimpa pendengar API yang kompatibel dengan OpenAI
  • DASHBOARD_PORT hanya menimpa dashboard/pendengar Next.js
  • Setel NEXT_PUBLIC_BASE_URL ke dasbor/URL publik Anda (untuk panggilan balik OAuth)

Cloud sync errors

  • Verifikasi BASE_URL poin ke instance Anda yang sedang berjalan
  • Verifikasi CLOUD_URL poin ke titik akhir cloud yang Anda harapkan
  • Jaga agar nilai NEXT_PUBLIC_* selaras dengan nilai sisi server

First login not working

  • Periksa INITIAL_PASSWORD di .env
  • Jika tidak disetel, kata sandi cadangan adalah 123456

Tidak ada log permintaan

  • call_logs di SQLite menyimpan metadata ringkasan untuk tabel Log Permintaan dan tampilan analitik
  • Muatan permintaan/respons terperinci ditulis ke DATA_DIR/call_logs/ sebagai satu artefak JSON per permintaan
  • Aktifkan pengambilan saluran pipa dari Dasbor → Log → Log Permintaan jika Anda memerlukan muatan per tahap yang terperinci
  • Export Logs membaca file artefak sesuai permintaan, sementara Export All menyertakan direktori call_logs/ bersama storage.sqlite
  • Setel APP_LOG_TO_FILE=true jika Anda juga ingin log konsol aplikasi di logs/application/app.log
  • Sesuaikan APP_LOG_MAX_FILE_SIZE, APP_LOG_RETENTION_DAYS, APP_LOG_MAX_FILES, dan CALL_LOG_MAX_ENTRIES sesuai kebutuhan

Tes koneksi menunjukkan "Tidak Valid" untuk penyedia yang kompatibel dengan OpenAI

  • Banyak penyedia tidak mengekspos titik akhir /models
  • OmniRoute v1.0.6+ menyertakan validasi fallback melalui penyelesaian obrolan
  • Pastikan URL dasar menyertakan akhiran /v1

🔐 OAuth di Server Jarak Jauh

⚠️ Penting bagi pengguna yang menjalankan OmniRoute di VPS, Docker, atau server jarak jauh mana pun

Kredensial OAuth yang disertakan dalam OmniRoute didaftarkan hanya untuk localhost. Saat Anda mengakses OmniRoute di server jarak jauh (misalnya https://omniroute.myserver.com), Google menolak autentikasi dengan:

Error 400: redirect_uri_mismatch

Solusi: Konfigurasikan kredensial OAuth Anda sendiri

Anda perlu membuat ID Klien OAuth 2.0 di Google Cloud Console dengan URI server Anda.

Langkah demi langkah

1. Open Google Cloud Console

Go to: https://console.cloud.google.com/apis/credentials

2. Buat ID Klien OAuth 2.0 baru

  • Klik "+ Buat Kredensial""ID klien OAuth"
  • Jenis aplikasi: "Aplikasi web"
  • Nama: apa pun yang Anda suka (mis. OmniRoute Remote)

3. Add Authorized Redirect URIs

Di kolom "URI pengalihan resmi", tambahkan:

https://your-server.com/callback

Ganti your-server.com dengan domain atau IP server Anda (sertakan port jika diperlukan, misalnya http://45.33.32.156:20128/callback).

4. Simpan dan salin kredensial

Setelah pembuatan, Google akan menampilkan ID Klien dan Rahasia Klien.

5. Tetapkan variabel lingkungan

Di .env Anda (atau variabel lingkungan Docker):

# For Antigravity:
ANTIGRAVITY_OAUTH_CLIENT_ID=your-client-id.apps.googleusercontent.com
ANTIGRAVITY_OAUTH_CLIENT_SECRET=GOCSPX-your-secret

GEMINI_OAUTH_CLIENT_ID=your-client-id.apps.googleusercontent.com
GEMINI_OAUTH_CLIENT_SECRET=GOCSPX-your-secret

6. Restart OmniRoute

# npm:
npm run dev

# Docker:
docker restart omniroute

7. Try connecting again

Google will now redirect correctly to https://your-server.com/callback.


Solusi sementara (tanpa kredensial kustom)

Jika Anda tidak ingin menyiapkan kredensial Anda sendiri saat ini, Anda masih dapat menggunakan alur URL manual:

  1. OmniRoute membuka URL otorisasi Google
  2. Setelah otorisasi, Google mencoba mengalihkan ke localhost (yang gagal di server jauh)
  3. Salin URL lengkap dari bilah alamat browser Anda (meskipun halaman tidak dimuat)
  4. Tempelkan URL tersebut ke bidang yang ditampilkan di modal koneksi OmniRoute
  5. Klik "Hubungkan"

Ini berfungsi karena kode otorisasi di URL valid terlepas dari apakah halaman pengalihan dimuat.


🛠️ Stack Teknologi

Klik untuk membuka detail stack teknologi
  • Runtime: Node.js 1822 LTS (⚠️ Node.js 24+ tidak didukungbetter-sqlite3 biner asli tidak kompatibel)
  • Bahasa: TypeScript 5.9 — 100% TypeScript di src/ dan open-sse/ (nol any dalam modul inti sejak v2.0)
  • Kerangka Kerja: Next.js 16 + React 19 + Tailwind CSS 4
  • Database: lebih baik-sqlite3 (SQLite) + LowDB (JSON legacy) — status domain, log proksi, audit MCP, keputusan perutean, memori, keterampilan
  • Skema: Zod (validasi I/O alat MCP, kontrak API)
  • Protokol: MCP (stdio/HTTP) + A2A v0.3 (JSON-RPC 2.0 + SSE)
  • Streaming: Peristiwa Terkirim Server (SSE)
  • Auth: OAuth 2.0 (PKCE) + JWT + Kunci API + Otorisasi Cakupan MCP
  • Pengujian: Pelari pengujian Node.js + Vitest (900+ pengujian termasuk unit, integrasi, E2E)
  • CI/CD: Tindakan GitHub (publikasi npm otomatis + Docker Hub saat dirilis)
  • Situs Web: omniroute.online
  • Paket: npmjs.com/package/omniroute
  • Pekerja Pelabuhan: hub.docker.com/r/diegosouzapw/omniroute
  • Ketahanan: Pemutus arus, backoff eksponensial, kawanan anti-thundering, spoofing TLS, penyembuhan diri kombo otomatis

Dokumentasi

Document Description
User Guide Penyedia, kombo, integrasi CLI, penerapan
API Reference Semua titik akhir dengan contoh
MCP Server 25 alat MCP, konfigurasi IDE, klien Python/TS/Go
A2A Server Protokol JSON-RPC 2.0, keterampilan, streaming, manajemen tugas
Auto-Combo Engine 6-factor scoring, mode packs, self-healing
Context Relay Strategi penyerahan sesi untuk rotasi akun
Troubleshooting Masalah umum dan solusinya
Architecture Arsitektur sistem dan internal
Codebase Documentation Beginner-friendly codebase walkthrough
Uninstall Guide Penghapusan bersih untuk semua metode instalasi
Environment Config Lengkapi .env variabel dan referensi
Contributing Pengaturan dan pedoman pengembangan
OpenAPI Spec OpenAPI 3.0 specification
Security Policy Pelaporan kerentanan dan praktik keamanan
VM Deployment Panduan lengkap: pengaturan VM + nginx + Cloudflare
Features Gallery Tur dasbor visual dengan tangkapan layar
Release Checklist Pre-release validation steps

🗺️ Roadmap

OmniRoute memiliki 218+ fitur yang direncanakan di berbagai fase pengembangan. Berikut adalah bidang-bidang utamanya:

Category Planned Features Highlights
🧠 Routing & Intelligence 25+ Perutean latensi terendah, perutean berbasis tag, preflight kuota, P2C sadar kuota, perutean kombo berbasis langkah
🔒 Security & Compliance 20+ Pengerasan SSRF, penyelubungan kredensial, batas tarif per titik akhir, pelingkupan kunci manajemen
📊 Observability 15+ Integrasi OpenTelemetry, pemantauan kuota waktu nyata, kesehatan target kombo, pelacakan biaya per model
🔄 Provider Integrations 20+ Registri model dinamis, cooldown koneksi, Codex multi-akun, penguraian kuota Salinan
Performance 15+ Lapisan cache ganda, cache cepat, cache respons, streaming keepalive, API batch
🌐 Ecosystem 10+ WebSocket API, config hot-reload, distributed config store, commercial mode

🔜 Segera Hadir

  • 🔗 Integrasi OpenCode — Dukungan penyedia asli untuk IDE pengkodean AI OpenCode
  • 🔗 Integrasi TRAE — Dukungan penuh untuk kerangka pengembangan AI TRAE
  • 📦 Batch API — Pemrosesan batch asinkron untuk permintaan massal
  • 🎯 Perutean Berbasis Tag — Merutekan permintaan berdasarkan tag dan metadata khusus
  • 💰 Strategi Biaya Terendah — Secara otomatis memilih penyedia termurah yang tersedia

📝 Spesifikasi fitur lengkap tersedia di docs/new-features/ (217 spesifikasi detail)


👥 Kontributor

Contributors

Cara Berkontribusi

  1. Cabangkan repositori
  2. Buat cabang fitur Anda (git checkout -b feature/amazing-feature)
  3. Komit perubahan Anda (git commit -m 'Add amazing feature')
  4. Dorong ke cabang (git push origin feature/amazing-feature)
  5. Buka Permintaan Tarik

Lihat CONTRIBUTING.md untuk panduan detailnya.

Merilis Versi Baru

# Create a release — npm publish happens automatically
gh release create v2.0.0 --title "v2.0.0" --generate-notes

📊 Riwayat Star

Star History Chart

🌍 StarMapper

StarMapper

🙏 Ucapan Terima Kasih

Terima kasih khusus kepada CLIProxyAPI — implementasi Go asli yang menginspirasi port JavaScript ini.


Lisensi

Lisensi MIT - lihat LICENSE untuk detailnya.


Built with ❤️ for developers who code 24/7
omniroute.online