diff --git a/AGENTS.md b/AGENTS.md index bc9e7c2c3c..faaa956fdb 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -46,7 +46,7 @@ Repository map and Reference Documentation sections below. ## Project at a Glance -**OmniRoute** — unified AI proxy/router. One endpoint, 290 LLM providers, auto-fallback. +**OmniRoute** — unified AI proxy/router. One endpoint, 291 LLM providers, auto-fallback. | Layer | Location | Purpose | | ------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | diff --git a/README.md b/README.md index 888bd362a0..e0fba74af8 100644 --- a/README.md +++ b/README.md @@ -7,7 +7,7 @@ # 🚀 OmniRoute — The Free AI Gateway -OmniRoute — Never stop coding. Every AI tool → 290 providers — 90+ free — through one endpoint. Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini with auto-fallback. RTK + Caveman stacked compression saves 15–95% tokens (~89% avg) — never hit limits. 290 AI providers · 90+ free tiers · ~1.53B free tokens/mo · 19 routing strategies · $0 to start. +OmniRoute — Never stop coding. Every AI tool → 291 providers — 90+ free — through one endpoint. Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini with auto-fallback. RTK + Caveman stacked compression saves 15–95% tokens (~89% avg) — never hit limits. 291 AI providers · 90+ free tiers · ~1.53B free tokens/mo · 19 routing strategies · $0 to start. @@ -81,7 +81,7 @@ ⚙️ Features 🎯 Combos - 🌐 Providers + 🌐 Providers 🔌 CLI & MCP @@ -188,7 +188,7 @@ curl http://localhost:20128/v1/chat/completions \ -The Promise — One endpoint. 290 providers. Never stop building — OmniRoute picks the cheapest one that works. Six pillars: Never hit limits (auto-fallback across 290 providers in milliseconds, zero downtime) · Save up to 95% tokens (RTK + Caveman stacked compression cuts 15–95%, ~89% avg on tool-heavy sessions) · $0 to start (90+ free tiers, 40+ free forever — no card needed) · Every tool works (33 coding agents through one config) · One endpoint (OpenAI ↔ Claude ↔ Gemini ↔ Responses API at /v1) · Production-grade (circuit breakers, TLS stealth, MCP 104 tools, A2A, memory, guardrails, evals — 25,000+ tests). +The Promise — One endpoint. 291 providers. Never stop building — OmniRoute picks the cheapest one that works. Six pillars: Never hit limits (auto-fallback across 291 providers in milliseconds, zero downtime) · Save up to 95% tokens (RTK + Caveman stacked compression cuts 15–95%, ~89% avg on tool-heavy sessions) · $0 to start (90+ free tiers, 40+ free forever — no card needed) · Every tool works (33 coding agents through one config) · One endpoint (OpenAI ↔ Claude ↔ Gemini ↔ Responses API at /v1) · Production-grade (circuit breakers, TLS stealth, MCP 104 tools, A2A, memory, guardrails, evals — 25,000+ tests).

@@ -439,7 +439,7 @@ All **19** strategies — mix & match per combo step: -What sets OmniRoute apart — comparison table vs 9router, OpenRouter, CLIProxyAPI and LiteLLM across 13 capabilities. OmniRoute: 290 providers, 90+ free providers built-in, 19 routing strategies, 12-engine token compression, built-in MCP server with 104 tools, A2A agent protocol, persistent memory, guardrails, cloud agents, TLS fingerprint stealth, Desktop/Termux/PWA, 43 i18n UI locales, 100% MIT self-hosted. OmniRoute is the only one with the full set; competitors show a mix of checks, partials and crosses. Verified from each project's docs. +What sets OmniRoute apart — comparison table vs 9router, OpenRouter, CLIProxyAPI and LiteLLM across 13 capabilities. OmniRoute: 291 providers, 90+ free providers built-in, 19 routing strategies, 12-engine token compression, built-in MCP server with 104 tools, A2A agent protocol, persistent memory, guardrails, cloud agents, TLS fingerprint stealth, Desktop/Termux/PWA, 43 i18n UI locales, 100% MIT self-hosted. OmniRoute is the only one with the full set; competitors show a mix of checks, partials and crosses. Verified from each project's docs. 📊 Full methodology & per-feature detail vs 9router, OpenRouter, CLIProxyAPI & LiteLLM → [`docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md`](docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md) @@ -513,7 +513,7 @@ Pix copia-e-cola: - **🖼️ New endpoints** — `/v1/ocr` (Mistral OCR) and `/v1/audio/translations` (Whisper-style) round out the media surface. → [API Reference](docs/reference/API_REFERENCE.md) - **🎨 Image / video / audio generation** — one API for media: xAI Grok Imagine & Novita AI video, ComfyUI, Freepik, Adobe Firefly, Microsoft Designer, Google Imagen, Segmind, EdgeTTS. → [API Reference](docs/reference/API_REFERENCE.md) - **🌍 Deployment & ops** — reverse-proxy `basePath`, browser-language auto-detect, per-key device tracking, root-less MITM trust, zh-TW localization. → [Environment](docs/reference/ENVIRONMENT.md) -- **🤝 More providers & agents** — Cursor Cloud Agent, Grok Build (xAI) with browser + OAuth login, Ollama first-class card, Claude Opus 5 & Sonnet 5, Kimi official partnership (Code/Web/Moonshot), Zed, Requesty, SenseNova, Yuanbao, Agnes AI… and a refreshed **290-provider catalog**. → [Providers](docs/reference/PROVIDER_REFERENCE.md) +- **🤝 More providers & agents** — Cursor Cloud Agent, Grok Build (xAI) with browser + OAuth login, Ollama first-class card, Claude Opus 5 & Sonnet 5, Kimi official partnership (Code/Web/Moonshot), Zed, Requesty, SenseNova, Yuanbao, Agnes AI… and a refreshed **291-provider catalog**. → [Providers](docs/reference/PROVIDER_REFERENCE.md) - **📡 Routing transparency** — every response carries an `X-OmniRoute-Decision` header naming the strategy/provider/latency that served it, a new `cache-optimized` combo strategy + Auto-Combo `cacheAffinity` factor route repeat requests back to the connection holding the cached prefix, and a read-only `/v1/auto-combo/{channel}/candidates` endpoint exposes an `auto/*` channel's live candidate pool. → [Auto-Combo](docs/routing/AUTO-COMBO.md) - **⚡ Local performance & infra** — one-click local Redis, Cloudflare Workers / Deno Deploy relay deployers, Bifrost & Mux as supervised embedded services. → [Embedded Services](docs/frameworks/EMBEDDED-SERVICES.md) @@ -574,11 +574,11 @@ Pix copia-e-cola:
-## 🌐 290 AI Providers — 90+ Free +## 🌐 291 AI Providers — 90+ Free
-> The most complete catalog of any open-source router: **290 providers**, **90+ with a free tier**, **40+ free forever**. +> The most complete catalog of any open-source router: **291 providers**, **90+ with a free tier**, **40+ free forever**.
diff --git a/config/quality/file-size-baseline.json b/config/quality/file-size-baseline.json index 4034c147c2..9b8280ec8a 100644 --- a/config/quality/file-size-baseline.json +++ b/config/quality/file-size-baseline.json @@ -172,7 +172,7 @@ "_rebaseline_2026_07_25_8510_adobe_firefly_reference_images_tests": "#8510 (artickc, feat/adobe-firefly-reference-images) own test growth: tests/unit/adobe-firefly.test.ts 711->871 (+159, entirely this PR's diff — new referenceBlobs upload/dispatch coverage for handleAdobeFireflyImageGeneration, resolveAdobeSourceImageIds, and the storage-upload wire contract). Route-level /v1/images/edits coverage (credentials/rate-limit/4-ref-cap branches added to route.ts) lives in the new tests/unit/8510-adobe-firefly-edits-route.test.ts instead of growing this file further.", "_rebaseline_basered_codebuddy_cn": "Base-red fix (#4664 CodeBuddy CN): oauth-providers-config.test.ts 867->870 (+3) to align the EXPECTED provider list/config with the codebuddy-cn provider that #4664 added to the registry without updating this test (it asserts 'exactly once').", "_rebaseline_pr4613_compatible_provider_groups": "Reconcile #4613 already-merged growth: providers-page-utils.test.ts 1004->1052 (+48, buildCompatibleProviderGroups partition unit test). Fast-gate PR->release does not run check:file-size, so this surfaced post-merge.", - "tests/integration/chat-pipeline.test.ts": 1592, + "tests/integration/chat-pipeline.test.ts": 1598, "tests/integration/chatcore-compression-integration.test.ts": 1114, "tests/unit/account-fallback-service.test.ts": 1563, "tests/unit/batch_api.test.ts": 1324, @@ -190,7 +190,7 @@ "tests/unit/models-catalog-route.test.ts": 1636, "tests/unit/perplexity-web.test.ts": 1355, "tests/unit/provider-models-route.test.ts": 1784, - "tests/unit/provider-validation-specialty.test.ts": 2980, + "tests/unit/provider-validation-specialty.test.ts": 2985, "tests/unit/providers-page-utils.test.ts": 1106, "tests/unit/response-sanitizer.test.ts": 1063, "tests/unit/route-edge-coverage.test.ts": 1241, @@ -343,7 +343,7 @@ "_rebaseline_pr1043_minimax_tts": "Upstream port decolua/9router#1043 (toanalien) own growth: audioSpeech.ts 965->1061 (+96). Adds MiniMax T2A v2 TTS dispatch (handleMinimaxSpeech + hexToBytes helper) — provider entry was already in audioRegistry (format: minimax-tts) but no handler existed, falling through to the OpenAI-compatible default that fails (T2A has custom shape + hex-encoded audio + base_resp envelope). New branch sits next to the other inline provider branches (xiaomi-mimo, coqui, tortoise, aws-polly) — extracting would just create indirection. Covered by tests/unit/minimax-tts-1043.test.ts (3 tests, GREEN: success, base_resp error, invalid-hex).", "_rebaseline_pr4592_exclude_exhausted_auto": "Reconcile #4592 already-merged growth: combo.ts 2991->3036 (+45, terminal-status quota-cutoff exclusion in buildAutoCandidates + opt-in gate). Fast-gate PR->release does not run check:file-size.", "open-sse/executors/antigravity.ts": 1528, - "open-sse/executors/base.ts": 1578, + "open-sse/executors/base.ts": 1623, "open-sse/executors/chatgpt-web.ts": 3241, "open-sse/executors/codex.ts": 1534, "open-sse/executors/cursor.ts": 1560, @@ -363,8 +363,8 @@ "open-sse/services/claudeCodeCompatible.ts": 1202, "open-sse/services/combo.ts": 3648, "open-sse/services/compression/strategySelector.ts": 1060, - "open-sse/services/rateLimitManager.ts": 1060, - "open-sse/translator/response/openai-responses.ts": 1174, + "open-sse/services/rateLimitManager.ts": 1105, + "open-sse/translator/response/openai-responses.ts": 1204, "open-sse/utils/cursorAgentProtobuf.ts": 1505, "open-sse/utils/stream.ts": 2889, "src/app/(dashboard)/dashboard/HomePageClient.tsx": 1381, @@ -405,7 +405,7 @@ "src/sse/handlers/chat.ts": 1845, "src/sse/services/auth.ts": 2508, "tests/unit/account-fallback-service.test.ts": 1572, - "tests/unit/provider-validation-specialty.test.ts": 2980, + "tests/unit/provider-validation-specialty.test.ts": 2985, "open-sse/executors/hyperagent.ts": 1026 }, "_rebaseline_2026_07_27_v3849_train2": "Merge-train 2 (7 PRs) — owner-approved 2026-07-27. Single entry: chatCore.ts 4955->5006 (#8595, Responses multi-turn image compaction before the context hard-reject). Genuine irreducible growth at the existing compaction chokepoint in handleChatCore — the PR adds a last-resort retry against the concrete budget plus the estimateFinalInputTokens helper, both wired at the pre-existing call site rather than a new branch. Covered by tests/unit/8560-responses-image-compaction.test.ts (4 tests).", @@ -418,5 +418,7 @@ "_rebaseline_2026_07_29_8281_home_quickstart_prefetch": "Release v3.8.49 base-red fix (no PR — captain sweep): src/app/(dashboard)/dashboard/HomePageClient.tsx 1377->1381 (+4). #8292 added prefetch={false} to the sidebar but left /home's five quick-start Links prefetching, so first paint still fired 12 speculative RSC requests — caught by navigation.spec.ts only after the e2e helper bug (APP_ROUTE_PATTERN missing /home) was repaired in the same cycle. Growth is the five prefetch attributes; it was offset first by extracting the repeated className literals (INLINE_LINK x4, DOCS_LINK x1), which collapsed five wrapped blocks back to one line each — a naive fix measured 1391. Guard: tests/unit/sidebar-prefetch-policy-8281.test.ts.", "_rebaseline_2026_08_01_8964_xai_agent_tools": "PR #8964 own growth: chatCore.ts 5020->5034 at the existing native-passthrough chokepoint. Adds xAI Agent Tools passthrough for /v1/responses (xai/xai-oauth/xao): resolve nativeXaiResponsesPassthrough, force openai-responses targetFormat, stamp body marker, and OR into the existing nativeCodexPassthrough sites (web-search bypass + requestEndpointPath). Leaf logic in passthroughHelpers, responsesEndpoint, targetFormat, xai executor, responseSanitizer, usageTracking. Cohesive wiring at the Codex passthrough boundary.", "_rebaseline_2026_08_01_8964_response_sanitizer": "PR #8964 own growth: responseSanitizer.ts 1115->1128. Keep cost_in_usd_ticks / server_side_tool_usage(_details) through sanitizeResponsesApiResponse allowlists so native xAI tool responses retain usage.", - "_rebaseline_2026_08_02_v3850_agentrouter_responses": "Release v3.8.50 AgentRouter/Codex compatibility reconciliation. open-sse/executors/base.ts 1562->1578: #9190 wires AgentRouter's selected Claude/OpenAI/Responses protocol through the existing executor URL, auth, identity-header and fingerprint chokepoints; the reusable alternate resolver remains outside base.ts. open-sse/utils/stream.ts 2887->2889: #9213 evaluates Responses ID and usage normalization independently so response.completed always receives finite usage.total_tokens instead of short-circuiting after an ID rewrite. tests/unit/chatcore-translation-paths.test.ts 2769->2776: #9191 updates the existing Claude-Code bridge assertions for the dynamic AgentRouter wire image. PR #9224 offsets its own chatCore growth by extracting the AgentRouter protocol decisions into chatCore/agentRouterProtocol.ts, leaving chatCore below its frozen ceiling. Covered by agentrouter executor/chatCore protocol tests, chatcore translation-path tests, and responses-commentary-passthrough tests." + "_rebaseline_2026_08_02_v3850_agentrouter_responses": "Release v3.8.50 AgentRouter/Codex compatibility reconciliation. open-sse/executors/base.ts 1562->1578: #9190 wires AgentRouter's selected Claude/OpenAI/Responses protocol through the existing executor URL, auth, identity-header and fingerprint chokepoints; the reusable alternate resolver remains outside base.ts. open-sse/utils/stream.ts 2887->2889: #9213 evaluates Responses ID and usage normalization independently so response.completed always receives finite usage.total_tokens instead of short-circuiting after an ID rewrite. tests/unit/chatcore-translation-paths.test.ts 2769->2776: #9191 updates the existing Claude-Code bridge assertions for the dynamic AgentRouter wire image. PR #9224 offsets its own chatCore growth by extracting the AgentRouter protocol decisions into chatCore/agentRouterProtocol.ts, leaving chatCore below its frozen ceiling. Covered by agentrouter executor/chatCore protocol tests, chatcore translation-path tests, and responses-commentary-passthrough tests.", + "_rebaseline_2026_08_05_9323_agentrouter_waf_retry": "PR #9323 (fix(agentrouter): retry on 400 content-blocked + burst guard) own growth: open-sse/executors/base.ts 1578->1623 (check-file-size.mjs conta via split(\"\\n\").length; wc -l ve 1622). As +45 linhas sao o WAF_RETRY_CONFIG + o burst guard via gateOutboundRequest() para o WAF do agentrouter.org, com comentarios explicando o porque de cada mitigacao e cobertos por tests/unit/base-executor-waf-retry.test.ts e tests/unit/wafRateLimit.test.ts. Crescimento funcional legitimo, nao inchaco.", + "_rebaseline_2026_08_05_9529_own_growth": "PR #9529 own growth (base release/v3.8.50 medida EXATAMENTE nos frozen antigos, entao o modo base-relative #8522 nao cobre): open-sse/services/rateLimitManager.ts 1060->1105 (+45: helper applyLimiterSettings() que re-arma o heartbeat do reservoir apos updateSettings — fix do bug Bottleneck 2.19.5 que congelava a fila weighted; TDD em tests/unit/ratelimit-reservoir-refresh.test.ts); tests/integration/chat-pipeline.test.ts 1592->1598 (+6: User-Agent do codex derivado de getCodexClientVersion() em vez de literal pinado — teste-irmao alinhado ao contrato); tests/unit/provider-validation-specialty.test.ts 2980->2985 (+5: cobertura NOVA claude-web 429 -> valid:false, alinhamento #9406); open-sse/translator/response/openai-responses.ts 1174->1204 (+30: buildResponsesReasoningSummaryDelta MOVIDA do leaf pureHelpers.ts para o host — a funcao do #9500 muta stream state e violava o contrato do leaf puro; o LOC total do par host+leaf nao cresceu, o pureHelpers encolheu o mesmo tanto). Crescimento por fix de producao + cobertura adicional + realocacao arquitetural, nao inchaco." } diff --git a/config/quality/test-masking-allowlist.json b/config/quality/test-masking-allowlist.json index e01abf42cc..e15040cb47 100644 --- a/config/quality/test-masking-allowlist.json +++ b/config/quality/test-masking-allowlist.json @@ -109,5 +109,7 @@ "tests/unit/usage-service-hardening.test.ts": "v3.8.49 #7866/#8565/#8013: qwen removido (−3 asserts); o Kimi/Kiro builder-id (uso profileless) passou a ter SUCESSO real em vez de erro de ARN — supportsProfilelessKiroUsage(\"builder-id\") retorna true —, trocando 1 assert de regex de erro por 3 asserts de valor; e os ids de bucket de quota do Antigravity foram atualizados para o catálogo atual. Rodado no HEAD: 23/23 passam. Net 210→209. Verificado legítimo. Prune após v3.8.49 mergear para main.", "tests/unit/virtual-auto-combo.test.ts": "v3.8.49 #7928/#8183: o pooling de contas passou a agrupar conexões web-session do mesmo provider numa entrada lógica com allowedConnectionIds (campo confirmado em open-sse/services/autoCombo/virtualFactory.ts), e o pool no-auth virou uma allowlist fixa (AUTO_COMBO_NOAUTH_ALLOWLIST = opencode, felo-web) — os testes antigos esperavam duplicatas e a inclusão de duckduckgo-web/theoldllm/chipotle, que hoje são corretamente excluídos. Guard dedicado em noauth-autocombo-allowlist.test.ts. Rodado no HEAD: 10/10 passam. Net 39→31. Verificado legítimo. Prune após v3.8.49 mergear para main.", "open-sse/services/__tests__/tierResolver.test.ts": "v3.8.49 #7866: refactor(qwen) remove o provider OAuth legado — o teste \"classifies Qwen as free\" e a entrada de qwen na lista do batch saíram junto com o provider, e os índices do batch desceram de 10 para 9 elementos (net 61→59). Superfície extinta, não enfraquecimento. Verificado legítimo. Prune após v3.8.49 mergear para main.", - "tests/unit/plugins-welcome-banner-e2e.test.ts": "v3.8.50 #9126 (commit 8fac6bcd48): o teste único 'BUILTIN_EVENTS has all 14 events' (13 asserts .ok/.equal) foi reestruturado em 3 testes mais específicos — 'contains only emitted/public events' (assert.deepEqual da lista completa), 'does not advertise dead events' (7 asserts .equal(false) para eventos sem emissor real: onModelSelect/onComboResolve/onRateLimit/onQuotaExhaust/onProviderError/onStreamStart/onStreamEnd) e 'lifecycle events remain represented' (4 asserts .ok). Contrato mais forte (agora também nega presença dos eventos mortos), não mais fraco — a contagem líquida cai (73→61) porque o assert.deepEqual único substitui múltiplos assert.ok redundantes com a mesma cobertura. Asserts restruturados, não removidos sem substituição. Verificado legítimo." + "tests/unit/plugins-welcome-banner-e2e.test.ts": "v3.8.50 #9126 (commit 8fac6bcd48): o teste único 'BUILTIN_EVENTS has all 14 events' (13 asserts .ok/.equal) foi reestruturado em 3 testes mais específicos — 'contains only emitted/public events' (assert.deepEqual da lista completa), 'does not advertise dead events' (7 asserts .equal(false) para eventos sem emissor real: onModelSelect/onComboResolve/onRateLimit/onQuotaExhaust/onProviderError/onStreamStart/onStreamEnd) e 'lifecycle events remain represented' (4 asserts .ok). Contrato mais forte (agora também nega presença dos eventos mortos), não mais fraco — a contagem líquida cai (73→61) porque o assert.deepEqual único substitui múltiplos assert.ok redundantes com a mesma cobertura. Asserts restruturados, não removidos sem substituição. Verificado legítimo.", + "tests/unit/web-tools-translation-2820.test.ts": "v3.8.50 #9343 (commit d969555417): fix(security) exige envelope explicito — JSON puro NAO deve mais ser promovido a tool_calls. Os 5 testes foram REESCRITOS para o contrato oposto (antes: 'promove e valida name/arguments'; agora: 'toolCalls === null e content preservado'), o que naturalmente usa menos asserts: verificar a NAO-promocao custa 2 asserts, verificar o objeto promovido custava 4. Contrato mais restritivo, nao mais fraco (39->35). Verificado legitimo — a inversao esta explicita nos proprios nomes dos testes ('does NOT promote ... (#9343)').", + "tests/unit/deepseek-web-tools-execute-2820.test.ts": "v3.8.50 ed661f2126 (alinhamento ao #9343): o teste 'parses bare JSON reply into OpenAI tool_calls' foi reescrito para o contrato INVERTIDO do fix de seguranca #9343 — JSON puro sem envelope NAO deve mais ser promovido. Verificar a nao-promocao custa 3 asserts (finish_reason stop, sem tool_calls, content preservado verbatim) onde validar o objeto promovido custava 5 (23->21). Mesma classe da entrada web-tools-translation-2820 acima. Contrato mais restritivo, nao mais fraco." } diff --git a/open-sse/services/rateLimitManager.ts b/open-sse/services/rateLimitManager.ts index c39b39c18e..87f48a46fa 100644 --- a/open-sse/services/rateLimitManager.ts +++ b/open-sse/services/rateLimitManager.ts @@ -164,13 +164,58 @@ function buildLimiterDefaults() { }; } -function updateAllLimiterSettings() { - const defaults = buildLimiterDefaults(); - for (const limiter of limiters.values()) { - limiter.updateSettings(defaults); +/** + * Apply new settings to a Bottleneck limiter and re-arm its reservoir-refresh + * heartbeat. + * + * Bottleneck 2.19.5 (frozen upstream dependency, no release since 2019) has a + * bug in `LocalDatastore#_startHeartbeat()` + * (node_modules/bottleneck/lib/LocalDatastore.js:29,56): the guard + * `if (this.heartbeat == null && ...)` only (re)creates the periodic + * reservoir-refresh interval the FIRST time it runs. Every later call — + * including the one `updateSettings()` itself triggers internally — falls + * into the `else` branch and does `clearInterval(this.heartbeat)` WITHOUT + * resetting `this.heartbeat` back to `null`. Because the stale reference is + * left in place, every future `_startHeartbeat()` call keeps taking the same + * dead `else` branch: the periodic reservoir refresh is gone forever after + * the FIRST manual `updateSettings()` call on a limiter — every limiter here + * starts with a live heartbeat (buildLimiterDefaults() always sets + * reservoirRefreshInterval/reservoirRefreshAmount), so that "first call" is + * whichever of the 5 updateSettings() call sites in this file runs first. + * + * Work around it here instead of patching node_modules: null out the stale + * reference ourselves and re-invoke `_startHeartbeat()` so it takes the + * "start a fresh interval" branch again. Every `limiter.updateSettings(...)` + * call in this file MUST go through this helper, never Bottleneck's method + * directly. + */ +async function applyLimiterSettings( + limiter: Bottleneck, + updates: Bottleneck.ConstructorOptions +): Promise { + await limiter.updateSettings(updates); + const store = ( + limiter as unknown as { + _store?: { + heartbeat?: ReturnType | null; + _startHeartbeat?: () => void; + }; + } + )._store; + if (store && typeof store._startHeartbeat === "function") { + if (store.heartbeat != null) clearInterval(store.heartbeat); + store.heartbeat = null; + store._startHeartbeat(); } } +async function updateAllLimiterSettings() { + const defaults = buildLimiterDefaults(); + await Promise.all( + Array.from(limiters.values(), (limiter) => applyLimiterSettings(limiter, defaults)) + ); +} + function reconcileEnabledConnections( connectionsRaw: unknown[], requestQueueSettings: RequestQueueSettings @@ -381,7 +426,7 @@ export async function initializeRateLimits() { connections as unknown[], currentRequestQueueSettings ); - updateAllLimiterSettings(); + await updateAllLimiterSettings(); // Load per-connection rate limit overrides connectionRateLimitOverrides.clear(); @@ -414,7 +459,7 @@ export async function applyRequestQueueSettings(nextSettings: RequestQueueSettin const { getCachedProviderConnections } = await import("@/lib/localDb"); const connections = await getCachedProviderConnections(); reconcileEnabledConnections(connections as unknown[], currentRequestQueueSettings); - updateAllLimiterSettings(); + await updateAllLimiterSettings(); } /** @@ -779,9 +824,7 @@ export function updateFromHeaders(provider, connectionId, headers, status, model logRateLimit( `⚠️ [RATE-LIMIT] ${provider}:${connectionId.slice(0, 8)} — near capacity, slowing down` ); - limiter.updateSettings({ - minTime: 200, // Add 200ms between requests - }); + trackAsyncOperation(applyLimiterSettings(limiter, { minTime: 200 })); return; } @@ -812,7 +855,7 @@ export function updateFromHeaders(provider, connectionId, headers, status, model } } - limiter.updateSettings(updates); + trackAsyncOperation(applyLimiterSettings(limiter, updates)); // Persist learned limits (debounced) recordLearnedLimit( @@ -1014,7 +1057,7 @@ async function loadPersistedLimits() { const limiter = limiters.get(key); if (limiter && limit > 0) { const inferredMinTime = minTime || Math.max(0, Math.floor(60000 / limit) - 10); - limiter.updateSettings({ minTime: inferredMinTime }); + await applyLimiterSettings(limiter, { minTime: inferredMinTime }); count++; } } @@ -1050,10 +1093,12 @@ export function updateFromResponseBody(provider, connectionId, responseBody, sta `🚫 [RATE-LIMIT] ${provider}:${connectionId.slice(0, 8)} — body-parsed retry: ${Math.ceil(retryAfterMs / 1000)}s (${reason})` ); - limiter.updateSettings({ - reservoir: 0, - reservoirRefreshAmount: 60, - reservoirRefreshInterval: retryAfterMs, - }); + trackAsyncOperation( + applyLimiterSettings(limiter, { + reservoir: 0, + reservoirRefreshAmount: 60, + reservoirRefreshInterval: retryAfterMs, + }) + ); } } diff --git a/open-sse/translator/response/openai-responses.ts b/open-sse/translator/response/openai-responses.ts index 1b22d5ce3a..346c23ffa9 100644 --- a/open-sse/translator/response/openai-responses.ts +++ b/open-sse/translator/response/openai-responses.ts @@ -18,7 +18,6 @@ import { normalizeOutputIndex, normalizeUpstreamFailure, getVisibleResponsesReasoningSummaryText, - buildResponsesReasoningSummaryDelta, } from "./openai-responses/pureHelpers.ts"; import { createEventEmitter } from "./openai-responses/eventEmitter.ts"; import { buildResponsesToolCallItem } from "./responsesToolItem.ts"; @@ -727,6 +726,37 @@ function markResponsesReasoningDeltaEmitted(state, itemId) { state.reasoningItemsWithDelta.add(id); } +// #9500 — streaming separator helper. When summary_index increments mid-stream +// for a given item_id, a new reasoning segment begins; prefix "\n\n" so segments +// don't arrive back-to-back. Only prefixes when a delta was already emitted for +// the item AND the index advanced — never on the first segment. Lives here (not +// in pureHelpers.ts) because it reads and mutates stream state, which the pure +// leaf must not hold. +function buildResponsesReasoningSummaryDelta(state, data, reasoningDelta) { + const itemId = data.item_id != null ? String(data.item_id) : ""; + const summaryIndex = typeof data.summary_index === "number" ? data.summary_index : null; + if (!(state.reasoningSummaryIndex instanceof Map)) { + state.reasoningSummaryIndex = new Map(); + } + const lastIndex = itemId ? state.reasoningSummaryIndex.get(itemId) : undefined; + const alreadyEmittedForItem = itemId + ? state.reasoningItemsWithDelta instanceof Set && state.reasoningItemsWithDelta.has(itemId) + : Boolean(state.reasoningDeltaEmitted); + let deltaText = reasoningDelta; + if ( + summaryIndex !== null && + lastIndex !== undefined && + summaryIndex > lastIndex && + alreadyEmittedForItem + ) { + deltaText = `\n\n${reasoningDelta}`; + } + if (itemId && (lastIndex === undefined || summaryIndex > lastIndex)) { + state.reasoningSummaryIndex.set(itemId, summaryIndex); + } + return deltaText; +} + // #5786 — build a Chat-format reasoning delta chunk in the shape the client renders in // its thinking panel (`reasoning_content`, or `reasoning_text` for Copilot-compatible // clients). Mirrors the `response.reasoning_summary_text.delta` branch. diff --git a/open-sse/translator/response/openai-responses/pureHelpers.ts b/open-sse/translator/response/openai-responses/pureHelpers.ts index e010b2c6f4..e35ad24854 100644 --- a/open-sse/translator/response/openai-responses/pureHelpers.ts +++ b/open-sse/translator/response/openai-responses/pureHelpers.ts @@ -177,37 +177,6 @@ export function extractResponsesReasoningSummaryText(item) { .join("\n\n"); } -// #9500 — streaming separator helper. When summary_index increments mid-stream -// for a given item_id, a new reasoning segment begins; prefix "\n\n" so segments -// don't arrive back-to-back. Only prefixes when a delta was already emitted for -// the item AND the index advanced — never on the first segment. -export function buildResponsesReasoningSummaryDelta(state, data, reasoningDelta) { - const itemId = data.item_id != null ? String(data.item_id) : ""; - const summaryIndex = - typeof data.summary_index === "number" ? data.summary_index : null; - if (!(state.reasoningSummaryIndex instanceof Map)) { - state.reasoningSummaryIndex = new Map(); - } - const lastIndex = itemId ? state.reasoningSummaryIndex.get(itemId) : undefined; - const alreadyEmittedForItem = itemId - ? state.reasoningItemsWithDelta instanceof Set && - state.reasoningItemsWithDelta.has(itemId) - : Boolean(state.reasoningDeltaEmitted); - let deltaText = reasoningDelta; - if ( - summaryIndex !== null && - lastIndex !== undefined && - summaryIndex > lastIndex && - alreadyEmittedForItem - ) { - deltaText = `\n\n${reasoningDelta}`; - } - if (itemId && (lastIndex === undefined || summaryIndex > lastIndex)) { - state.reasoningSummaryIndex.set(itemId, summaryIndex); - } - return deltaText; -} - // #7095/#7176 — when Codex exposes a reasoning item only as encrypted private // reasoning (no plaintext summary), chat clients would otherwise see nothing in // their thinking panel. Reconciles two goals that used to be in tension: diff --git a/scripts/check/check-test-masking.mjs b/scripts/check/check-test-masking.mjs index 0f94b55d6a..9bb8832586 100644 --- a/scripts/check/check-test-masking.mjs +++ b/scripts/check/check-test-masking.mjs @@ -251,7 +251,7 @@ export function findReimplementedConditions(prodSources, testSource, testImports * (filtro D do git diff --diff-filter=MDR). * * `deletionAllowlist` (`_deletedWithReplacement` no test-masking-allowlist.json) - * isenta uma deleção de duas formas, cada uma com sua própria verificação: + * isenta uma deleção de três formas, cada uma com sua própria verificação: * 1. `replacement` (path string) — o substituto declarado existe no HEAD e é * ele próprio um arquivo de teste — o caso "reescrito em outro path sem * rename detectável" (conteúdo novo demais para o -M do git). @@ -259,12 +259,19 @@ export function findReimplementedConditions(prodSources, testSource, testImports * os arquivos de produção listados precisam estar ausentes no HEAD (sem * substituto porque não há mais código a testar). Usar apenas quando a * remoção do código-fonte está confirmada na mesma commit/PR. + * 3. `strayFromCommit` (hash) + `reason` (não-vazio) — o arquivo entrou no + * repositório POR ACIDENTE no commit declarado (ex.: um commit de docs + * que varreu artefatos de worktree de outra sessão, caso f4e93f339d) e a + * deleção devolve o arquivo ao seu fluxo dono (um PR/issue aberto). O + * gate verifica via git que o commit declarado é exatamente o que ADICIONOU + * o arquivo; o `reason` deve nomear o PR/issue dono para a revisão humana. * Qualquer entrada cuja condição declarada não se verifique continua flagada. */ export function evaluateDeletedFiles( deletedPaths, deletionAllowlist = {}, - fileExists = fs.existsSync + fileExists = fs.existsSync, + addedByCommit = lookupAddedByCommit ) { const flags = []; for (const f of deletedPaths) { @@ -285,6 +292,21 @@ export function evaluateDeletedFiles( ); continue; } + if (entry && typeof entry.strayFromCommit === "string" && entry.strayFromCommit.trim()) { + if (typeof entry.reason !== "string" || !entry.reason.trim()) { + flags.push( + `${f}: deleção allowlistada como stray mas sem \`reason\` — nomeie o PR/issue dono do arquivo` + ); + continue; + } + const actual = addedByCommit(f); + const declared = entry.strayFromCommit.trim(); + if (actual && (actual === declared || actual.startsWith(declared))) continue; + flags.push( + `${f}: deleção allowlistada como stray de ${declared} mas o commit que adicionou o arquivo é ${actual ?? "desconhecido"}` + ); + continue; + } flags.push( `${f}: arquivo de teste deletado — revisão humana obrigatória (mascaramento alto-sinal)` ); @@ -292,6 +314,26 @@ export function evaluateDeletedFiles( return flags; } +/** + * (subcheck 1, forma 3) Hash COMPLETO do commit que adicionou `path` (o add + * mais recente — cobre o caso deletado-e-readicionado). `null` quando o git + * não conhece o path. + */ +function lookupAddedByCommit(path) { + try { + const out = execFileSync("git", ["log", "--diff-filter=A", "--format=%H", "--", path], { + encoding: "utf8", + }); + const hashes = out + .split("\n") + .map((s) => s.trim()) + .filter(Boolean); + return hashes.length ? hashes[0] : null; + } catch { + return null; + } +} + /** * Parse `git diff --name-status -M --diff-filter=DR` output, separating TRUE * test-file deletions ("D\tpath") from RENAMES ("R\told\tnew"). diff --git a/scripts/quality/validate-release-green.mjs b/scripts/quality/validate-release-green.mjs index e1adca9927..f9693aa92e 100644 --- a/scripts/quality/validate-release-green.mjs +++ b/scripts/quality/validate-release-green.mjs @@ -96,9 +96,7 @@ export function firstFailureLine(out) { .split("\n") .map((l) => l.trim()) .filter(Boolean); - const hit = lines.find((l) => - /✖|✗|not ok|AssertionError|error TS|FAIL|Error:|REGRESS/i.test(l) - ); + const hit = lines.find((l) => /✖|✗|not ok|AssertionError|error TS|FAIL|Error:|REGRESS/i.test(l)); return (hit || lines[lines.length - 1] || "failed").slice(0, 200); } @@ -570,10 +568,20 @@ async function main() { // release — that is why it is a HARD pre-flight gate. const slow = [ { + // Raised 45→100min 2026-08-05: a hermetic-env run on the loaded devbox + // (load 7-26) was still inside invocation 1 of 3 at 76min when killed; + // contention factor 2-3× was measured against idle windows, and no idle + // measurement exists yet. The pre-flight's REAL condition is exactly + // this contended one (unit runs in Promise.all with integration+vitest + // plus whatever else the devbox carries), and there 45min provably + // killed a healthy suite and fabricated a false base-red. The ceiling's + // purpose — turning a genuine hang (stuck SQLite handle = zero progress + // forever) into a visible failure — survives at 100min. + // TODO: measure on the idle .113 box and re-tighten to ~1.8× measured. id: "unit", - label: "Unit tests (full suite, CI concurrency — runs ~20-35min silently)", + label: "Unit tests (full suite, CI concurrency — ~30-50min idle, up to ~100min under load)", args: ["run", "test:unit:ci"], - timeout: 45 * 60 * 1000, + timeout: 100 * 60 * 1000, }, { id: "vitest", @@ -582,10 +590,16 @@ async function main() { timeout: 15 * 60 * 1000, }, { + // Measured 2026-08-05 on an idle 16-core box: 22m08s hermetic (935 tests, + // 112 files at --test-concurrency=1, i.e. strictly serial because ~16 of + // them bind a port or share a DB). The old "~3-10min" estimate was stale by + // ~3x and the 20min ceiling killed a healthy run. 40min keeps the ceiling's + // real purpose — turning a genuine hang (unreleased DB handle) into a + // visible failure — without punishing a long-but-healthy suite. id: "integration", - label: "Integration tests (~3-10min)", + label: "Integration tests (~20-25min)", args: ["run", "test:integration"], - timeout: 20 * 60 * 1000, + timeout: 40 * 60 * 1000, }, ]; if (WITH_BUILD) { diff --git a/src/lib/guardrails/visionBridge.ts b/src/lib/guardrails/visionBridge.ts index 11f7555197..34fa2f10bd 100644 --- a/src/lib/guardrails/visionBridge.ts +++ b/src/lib/guardrails/visionBridge.ts @@ -241,7 +241,14 @@ export class VisionBridgeGuardrail extends BaseGuardrail { typeof settings.visionBridgeModel === "string" && settings.visionBridgeModel.trim() ? settings.visionBridgeModel.trim() : undefined; - const bestModel = await getBestVisionModel({ fixedModel: configuredModel }); + // Propagate the same resolved credential check used by the adjacent + // checkCreds() calls above/below (#8430) — without this, the router + // falls back to the real DB-backed hasUsableCredentialsForModel and + // ignores an injected `deps.hasUsableCredentials` test/DI override. + const bestModel = await getBestVisionModel( + { fixedModel: configuredModel }, + { hasUsableCredentials: checkCreds } + ); if (bestModel && bestModel !== model) { const bestUsable = await checkCreds(bestModel); // Only block the reroute when we KNOW the target is unusable (false). diff --git a/tests/integration/chat-pipeline.test.ts b/tests/integration/chat-pipeline.test.ts index a9568a0271..8144c972c0 100644 --- a/tests/integration/chat-pipeline.test.ts +++ b/tests/integration/chat-pipeline.test.ts @@ -689,7 +689,13 @@ test("chat pipeline applies Codex CLI fingerprint to OAuth responses requests", assert.equal(call.headers.Version, getCodexClientVersion()); assert.equal(call.headers["Openai-Beta"], "responses=experimental"); assert.equal(call.headers["X-Codex-Beta-Features"], "responses_websockets"); - assert.equal(call.headers["User-Agent"], "codex-cli/0.144.1 (Windows 10.0.26200; x64)"); + // Derive from the same source the code reads (see getCodexClientVersion() two + // lines above) instead of pinning the literal — #9323's version bump to 0.146.0 + // broke this assertion while the rest of the test kept passing. + assert.equal( + call.headers["User-Agent"], + `codex-cli/${getCodexClientVersion()} (Windows 10.0.26200; x64)` + ); assert.equal(call.headers["x-codex-window-id"], "conv_codex_fingerprint:0"); assert.ok(call.headers["x-client-request-id"], "expected Codex request id header"); assert.ok(call.headers["x-codex-turn-metadata"], "expected Codex turn metadata header"); diff --git a/tests/unit/8189-classifier-compat-auto-narrow.test.ts b/tests/unit/8189-classifier-compat-auto-narrow.test.ts index 4f6df06a6c..9070bf5c0b 100644 --- a/tests/unit/8189-classifier-compat-auto-narrow.test.ts +++ b/tests/unit/8189-classifier-compat-auto-narrow.test.ts @@ -11,6 +11,12 @@ * * Fix: in "auto" mode, the SECURITY_MONITOR_MARKER system-prompt text is now a * necessary condition. `stop_sequences` alone is no longer sufficient. + * + * Follow-up (#9276): "always" mode previously short-circuited EVERY Claude-format + * request unconditionally, regardless of signal shape. That let a normal chat + * request through /v1/messages be silently swallowed by an operator's "always" + * opt-in. The unconditional `if (mode === "always") return true` branch was + * removed — "always" now requires the same SECURITY_MONITOR_MARKER as "auto". */ import test from "node:test"; @@ -52,24 +58,37 @@ test("issue #8189: 'auto' mode still short-circuits when the security-monitor ma ); }); -test("issue #8189: 'always' mode is unaffected — every Claude-format request still short-circuits (operator opted in)", () => { +test("issue #9276: 'always' mode does NOT short-circuit without the security-monitor marker (narrowed to match 'auto')", () => { const body = { system: "You are a helpful assistant that writes CMS page templates.", stop_sequences: [""], messages: [{ role: "user", content: "hello" }], }; + assert.equal( + shouldDefaultAllowClassifier(FORMATS.CLAUDE, body, "always"), + false, + "mode='always' must NOT short-circuit a request with no security-monitor marker, even " + + "with stop_sequences=[''] — #9276 removed the unconditional always-mode return" + ); +}); + +test("issue #9276: 'always' mode still short-circuits when the security-monitor marker is present (operator opted in)", () => { + const body = { + system: + "You are a security monitor for autonomous AI coding agents. Evaluate the following action.", + stop_sequences: [], + messages: [{ role: "user", content: "Bash rm -rf /" }], + }; assert.equal( shouldDefaultAllowClassifier(FORMATS.CLAUDE, body, "always"), true, - "mode='always' must short-circuit every Claude-format request regardless of signal shape" + "mode='always' must still short-circuit when the security-monitor marker is present" ); }); test("issue #8189: 'off' mode (shipped default) never short-circuits", () => { const body = { - system: [ - { type: "text", text: "You are a security monitor for autonomous AI coding agents." }, - ], + system: [{ type: "text", text: "You are a security monitor for autonomous AI coding agents." }], stop_sequences: [""], }; assert.equal(shouldDefaultAllowClassifier(FORMATS.CLAUDE, body, "off"), false); diff --git a/tests/unit/deepseek-web-tools-execute-2820.test.ts b/tests/unit/deepseek-web-tools-execute-2820.test.ts index 49c01497d1..77854cebd6 100644 --- a/tests/unit/deepseek-web-tools-execute-2820.test.ts +++ b/tests/unit/deepseek-web-tools-execute-2820.test.ts @@ -132,7 +132,9 @@ test("execute (non-stream) parses reply into OpenAI tool_calls", async () assert.equal(choice.finish_reason, "tool_calls"); assert.equal(choice.message.tool_calls.length, 1); assert.equal(choice.message.tool_calls[0].function.name, "get_weather"); - assert.deepEqual(JSON.parse(choice.message.tool_calls[0].function.arguments), { city: "Paris" }); + assert.deepEqual(JSON.parse(choice.message.tool_calls[0].function.arguments), { + city: "Paris", + }); assert.ok( !String(choice.message.content || "").includes(""), "raw tool block stripped from content" @@ -142,8 +144,9 @@ test("execute (non-stream) parses reply into OpenAI tool_calls", async () } }); -test("execute (non-stream) parses bare JSON reply into OpenAI tool_calls", async () => { - const mock = installMock('{"name":"getWeather","arguments":{"city":"Paris"}}'); +test("execute (non-stream) does NOT promote bare JSON reply to tool_calls (#9343)", async () => { + const bareJson = '{"name":"getWeather","arguments":{"city":"Paris"}}'; + const mock = installMock(bareJson); try { const executor = new DeepSeekWebExecutor(); const result = await executor.execute({ @@ -156,11 +159,17 @@ test("execute (non-stream) parses bare JSON reply into OpenAI tool_calls", async assert.ok(result.response.ok); const json = JSON.parse(await result.response.text()); const choice = json.choices[0]; - assert.equal(choice.finish_reason, "tool_calls"); - assert.equal(choice.message.tool_calls.length, 1); - assert.equal(choice.message.tool_calls[0].function.name, "get_weather"); - assert.deepEqual(JSON.parse(choice.message.tool_calls[0].function.arguments), { city: "Paris" }); - assert.equal(choice.message.content, null, "bare JSON tool call is stripped from content"); + assert.equal( + choice.finish_reason, + "stop", + "bare JSON with no envelope must not be promoted to a tool call" + ); + assert.ok(!choice.message.tool_calls, "no tool_calls on a bare JSON reply (#9343)"); + assert.equal( + choice.message.content, + bareJson, + "bare JSON must be preserved verbatim as content, not stripped or promoted" + ); } finally { mock.restore(); } diff --git a/tests/unit/guardrails/visionBridge.test.ts b/tests/unit/guardrails/visionBridge.test.ts index b1b0d9f7ec..9dc0db18c0 100644 --- a/tests/unit/guardrails/visionBridge.test.ts +++ b/tests/unit/guardrails/visionBridge.test.ts @@ -428,7 +428,7 @@ test("VB-S07: reroutes base64 image to vision model", async () => { // ── VB-S03: Fail-open on vision error (via combo mapping path) ──────────── -test("VB-S03: preserves the original image when the vision API fails (#4012)", async () => { +test("VB-S03/#8430: combo-mapping describe failure replaces the image with an error stub (not preserved)", async () => { shouldVisionFail = true; const guardrail = createGuardrail({ deps: { @@ -464,12 +464,22 @@ test("VB-S03: preserves the original image when the vision API fails (#4012)", a text?: string; }>; - // #4012: a failed describe must NOT replace the image with an "(unavailable)" - // stub — the original image is preserved so a vision-capable upstream can see it. + // SEMANTIC CHANGE (#8430): in the combo describe path (forced here via + // checkModelHasComboMapping), when EVERY describe call fails, the upstream is + // a confirmed non-vision model that cannot handle raw images — the raw + // image_url part is now replaced with an "(unavailable)" error stub instead + // of being preserved. The original #4012 preserve-raw behavior still applies + // to the reroute path, where the upstream model might still be vision-capable + // (see tests/unit/vision-bridge-preserve-on-failure-4012.test.ts, updated by + // the same #8430 commit). const imagePart = content.find((p) => p.type === "image_url"); - assert.ok(imagePart, "original image_url part must be preserved on describe failure"); + assert.strictEqual( + imagePart, + undefined, + "raw image_url must be replaced when every describe call fails in the combo path" + ); const unavailPart = content.find((p) => p.type === "text" && p.text?.includes("unavailable")); - assert.strictEqual(unavailPart, undefined); + assert.ok(unavailPart, "an 'unavailable' error stub should be present when describe fails"); }); test("VB-S03: logs warning when vision API fails (via combo mapping)", async () => { @@ -771,21 +781,11 @@ test("VB-CRED-02: does NOT reroute to a vision model known to lack credentials", }); test("isProviderConnectionUsable rejects noauth without api key", async () => { - const { isProviderConnectionUsable } = await import( - "../../../src/lib/guardrails/visionBridge.ts" - ); - assert.strictEqual( - isProviderConnectionUsable({ authType: "noauth", apiKey: null }), - false - ); - assert.strictEqual( - isProviderConnectionUsable({ authType: "apikey", apiKey: "sk-real" }), - true - ); - assert.strictEqual( - isProviderConnectionUsable({ authType: "oauth", refreshToken: "rt" }), - true - ); + const { isProviderConnectionUsable } = + await import("../../../src/lib/guardrails/visionBridge.ts"); + assert.strictEqual(isProviderConnectionUsable({ authType: "noauth", apiKey: null }), false); + assert.strictEqual(isProviderConnectionUsable({ authType: "apikey", apiKey: "sk-real" }), true); + assert.strictEqual(isProviderConnectionUsable({ authType: "oauth", refreshToken: "rt" }), true); assert.strictEqual( isProviderConnectionUsable({ authType: "apikey", apiKey: "x", testStatus: "banned" }), false diff --git a/tests/unit/guardrails/visionBridgeHelpers.callVisionModel.test.ts b/tests/unit/guardrails/visionBridgeHelpers.callVisionModel.test.ts index 1465e927cb..8f812f5173 100644 --- a/tests/unit/guardrails/visionBridgeHelpers.callVisionModel.test.ts +++ b/tests/unit/guardrails/visionBridgeHelpers.callVisionModel.test.ts @@ -6,6 +6,8 @@ import test from "node:test"; import assert from "node:assert/strict"; import dns from "node:dns"; import { callVisionModel, type VisionModelConfig } from "@/lib/guardrails/visionBridgeHelpers"; +import { createProviderConnection } from "@/lib/db/providers"; +import { resetDbInstance } from "@/lib/db/core"; // Store original fetch const originalFetch = globalThis.fetch; @@ -30,6 +32,37 @@ process.on("exit", () => { (dns.promises as { lookup: unknown }).lookup = originalDnsLookup; }); +// (#8430) getBestVisionModel now validates that a `fixedModel` has a usable +// connection (via hasUsableCredentialsForModel, which queries the real DB) +// before returning it — an unreachable fixedModel falls through to +// auto-selection and, with nothing else configured either, resolves to `null`, +// which callVisionModel turns into a hard "No vision-capable provider +// connected" error before it ever reaches the HTTP call these tests mock. +// The router/credential-selection logic itself is already covered by +// visionBridgeRouter.test.ts and repro-8430.test.ts; these tests exercise +// callVisionModel's own request/response handling, so they just need one +// usable connection seeded per provider they use ("openai/gpt-4o-mini", +// "anthropic/claude-3-haiku") so getBestVisionModel resolves the requested +// fixedModel unchanged instead of null. +test.before(async () => { + await createProviderConnection({ + provider: "openai", + authType: "apikey", + name: "vision-bridge-test-openai", + apiKey: "sk-test-openai", + }); + await createProviderConnection({ + provider: "anthropic", + authType: "apikey", + name: "vision-bridge-test-anthropic", + apiKey: "sk-test-anthropic", + }); +}); + +test.after(() => { + resetDbInstance(); +}); + test("callVisionModel returns description on success", async () => { // Mock global fetch const mockResponse = { diff --git a/tests/unit/issue-7859-gemini-web-redirect-valid.test.ts b/tests/unit/issue-7859-gemini-web-redirect-valid.test.ts index 6944e288c7..bae36cd222 100644 --- a/tests/unit/issue-7859-gemini-web-redirect-valid.test.ts +++ b/tests/unit/issue-7859-gemini-web-redirect-valid.test.ts @@ -15,9 +15,8 @@ import test from "node:test"; import assert from "node:assert/strict"; -const { validateGeminiWebProvider } = await import( - "../../src/lib/providers/validation/webProvidersB.ts" -); +const { validateGeminiWebProvider } = + await import("../../src/lib/providers/validation/webProvidersB.ts"); const originalFetch = globalThis.fetch; @@ -25,13 +24,38 @@ test.afterEach(() => { globalThis.fetch = originalFetch; }); +// #9407 refined this contract: a redirect to accounts.google.com/ServiceLogin is +// specifically an EXPIRED session (valid:false with re-paste guidance), while other +// public accounts.google.com paths remain valid-with-warning. The original #7859 +// regression (public redirect must not fall through to the generic catch → invalid) +// is still covered — by the non-ServiceLogin variant below. +test("gemini-web validator: 302 redirect to ServiceLogin → expired session (#9407)", async () => { + globalThis.fetch = async (url) => { + const target = String(url); + if (target.includes("gemini.google.com/app")) { + return new Response(null, { + status: 302, + headers: { location: "https://accounts.google.com/ServiceLogin" }, + }); + } + throw new Error(`unexpected fetch: ${target}`); + }; + + const result = await validateGeminiWebProvider({ + apiKey: "__Secure-1PSID=eyJvalidsession", + }); + + assert.equal(result.valid, false); + assert.match(result.error || "", /Session expired/i); +}); + test("gemini-web validator: 302 redirect to a PUBLIC host → valid (regression #7859)", async () => { globalThis.fetch = async (url) => { const target = String(url); if (target.includes("gemini.google.com/app")) { return new Response(null, { status: 302, - headers: { location: "https://accounts.google.com/ServiceLogin" }, + headers: { location: "https://accounts.google.com/signin/continue" }, }); } throw new Error(`unexpected fetch: ${target}`); diff --git a/tests/unit/launch-codex-windows-spawn-6312.test.ts b/tests/unit/launch-codex-windows-spawn-6312.test.ts index a08019e40f..c463f1ef77 100644 --- a/tests/unit/launch-codex-windows-spawn-6312.test.ts +++ b/tests/unit/launch-codex-windows-spawn-6312.test.ts @@ -5,16 +5,26 @@ import { resolveCodexSpawn } from "../../bin/cli/commands/launch-codex.mjs"; // Regression guard for #6312: on Windows the `codex` binary is an npm `.cmd` // shim that `spawn` cannot resolve without a shell (bare "codex" → ENOENT). -test("resolveCodexSpawn: win32 spawns codex.cmd through a shell", () => { - const { command, shell } = resolveCodexSpawn("win32"); +// Since #9454 resolveCodexSpawn is async and probes PATH for a native .exe +// first; the .cmd+shell fallback below is the original #6312 contract (the +// .exe-preferred path is covered in cli/launch-claude-exe-windows-9454.test.ts). +test("resolveCodexSpawn: win32 spawns codex.cmd through a shell when no .exe is found", async () => { + const { command, shell } = await resolveCodexSpawn("win32", { probe: async () => null }); assert.equal(command, "codex.cmd"); assert.equal(shell, true); }); -test("resolveCodexSpawn: non-Windows platforms spawn the bare binary without a shell", () => { +test("resolveCodexSpawn: non-Windows platforms spawn the bare binary without a shell", async () => { for (const platform of ["linux", "darwin", "freebsd"]) { - const { command, shell } = resolveCodexSpawn(platform); + let probeCalls = 0; + const { command, shell } = await resolveCodexSpawn(platform, { + probe: async () => { + probeCalls++; + return null; + }, + }); assert.equal(command, "codex", `${platform} command`); assert.equal(shell, undefined, `${platform} shell`); + assert.equal(probeCalls, 0, `${platform} must not probe PATH off Windows`); } }); diff --git a/tests/unit/mitm-cert-install-mode-9442.test.ts b/tests/unit/mitm-cert-install-mode-9442.test.ts index 9c1fbb8fcd..979dee7d4c 100644 --- a/tests/unit/mitm-cert-install-mode-9442.test.ts +++ b/tests/unit/mitm-cert-install-mode-9442.test.ts @@ -31,6 +31,16 @@ import { execFileSync } from "node:child_process"; const originalPlatformDescriptor = Object.getOwnPropertyDescriptor(process, "platform")!; const originalPath = process.env.PATH; const originalNoSudo = process.env.OMNIROUTE_NO_SUDO; +// The global test harness (tests/_setup/isolateDataDir.ts) sets +// OMNIROUTE_SKIP_SYSTEM_TRUST=1 so no test mutates the host trust store — +// which makes installCert() return before issuing any command, so this file +// captures nothing and its install-gap assert can never pass under `npm run +// test:unit` (it only passed when invoked directly, without the harness). +// Clearing it here is safe: every spawned command (cp/mkdir/chmod/update-ca-*) +// is a logging stub on PATH and OMNIROUTE_NO_SUDO=1 strips sudo, so nothing +// touches the real system. Restored in test.after below. +const originalSkipSystemTrust = process.env.OMNIROUTE_SKIP_SYSTEM_TRUST; +delete process.env.OMNIROUTE_SKIP_SYSTEM_TRUST; const tmpRoot = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-9442-")); const binDir = path.join(tmpRoot, "bin"); @@ -68,6 +78,8 @@ test.after(() => { process.env.PATH = originalPath; if (originalNoSudo === undefined) delete process.env.OMNIROUTE_NO_SUDO; else process.env.OMNIROUTE_NO_SUDO = originalNoSudo; + if (originalSkipSystemTrust === undefined) delete process.env.OMNIROUTE_SKIP_SYSTEM_TRUST; + else process.env.OMNIROUTE_SKIP_SYSTEM_TRUST = originalSkipSystemTrust; fs.rmSync(tmpRoot, { recursive: true, force: true }); }); @@ -87,7 +99,10 @@ function fakeCertFile(seed: string): string { const der = crypto.createHash("sha256").update(seed).digest(); const pem = "-----BEGIN CERTIFICATE-----\n" + - der.toString("base64").match(/.{1,64}/g)!.join("\n") + + der + .toString("base64") + .match(/.{1,64}/g)! + .join("\n") + "\n-----END CERTIFICATE-----\n"; const certPath = path.join(tmpRoot, `${seed}.crt`); fs.writeFileSync(certPath, pem); @@ -156,11 +171,7 @@ test("filesystem proof: cp under umask 0077 creates mode 0600 (why the fix is ne // the bare `cp` on PATH below is a logging stub from the install tests. execFileSync("/usr/bin/cp", [src, dst]); const mode = fs.statSync(dst).mode & 0o777; - assert.equal( - mode, - 0o600, - "cp under umask 0077 must produce 0600 — the bug this fix repairs" - ); + assert.equal(mode, 0o600, "cp under umask 0077 must produce 0600 — the bug this fix repairs"); } finally { process.umask(oldUmask); } diff --git a/tests/unit/provider-validation-specialty.test.ts b/tests/unit/provider-validation-specialty.test.ts index a9317bcb38..4cdd871eef 100644 --- a/tests/unit/provider-validation-specialty.test.ts +++ b/tests/unit/provider-validation-specialty.test.ts @@ -2415,7 +2415,12 @@ test("claude-web validator: 401 → invalid session cookie", async () => { __setClaudeTlsFetchOverride(null); }); -test("claude-web validator: 429 → valid (rate limited means auth passed)", async () => { +// #9406 inverted this contract: a 429 session shows as UNHEALTHY (valid:false) +// so the dashboard stops painting rate-limited sessions green. The dedicated +// repro (tests/unit/repro-9406-claude-web-429-valid.test.ts) owns the full +// contract incl. Retry-After forwarding; this sibling keeps the validator-level +// assertion aligned with it. +test("claude-web validator: 429 → invalid (rate limited session is not healthy, #9406)", async () => { __setClaudeTlsFetchOverride(async () => makeClaudeTlsResponse(429, JSON.stringify({ error: "rate limited" })) ); @@ -2425,7 +2430,7 @@ test("claude-web validator: 429 → valid (rate limited means auth passed)", asy apiKey: "sessionKey=sk-ant-sid02-good-key", }); - assert.equal(result.valid, true); + assert.equal(result.valid, false); __setClaudeTlsFetchOverride(null); }); diff --git a/tests/unit/ratelimit-reservoir-refresh.test.ts b/tests/unit/ratelimit-reservoir-refresh.test.ts new file mode 100644 index 0000000000..7d218d31a6 --- /dev/null +++ b/tests/unit/ratelimit-reservoir-refresh.test.ts @@ -0,0 +1,132 @@ +/** + * TDD regression test — Bottleneck reservoir heartbeat death after updateSettings(). + * + * Bug: Bottleneck 2.19.5 (frozen upstream dependency, no release since 2019) has a + * defect in `LocalDatastore#_startHeartbeat()` + * (node_modules/bottleneck/lib/LocalDatastore.js:26-58). The guard + * `if (this.heartbeat == null && ...)` only (re)creates the periodic + * reservoir-refresh `setInterval` the FIRST time it runs. Every later call — + * including the one `updateSettings()` itself triggers internally via + * `__updateSettings__` — falls into the `else` branch and does + * `clearInterval(this.heartbeat)` WITHOUT resetting `this.heartbeat` back to + * `null`. Because the stale (now-invalid) reference is left in place, every + * future `_startHeartbeat()` call keeps taking the same dead `else` branch: + * the periodic reservoir refresh is gone forever after the FIRST manual + * `limiter.updateSettings()` call. + * + * Every limiter created by rateLimitManager.ts starts with a live heartbeat + * (the constructor call inside `getLimiter()` always sets + * reservoirRefreshInterval/reservoirRefreshAmount — see buildLimiterDefaults()), + * so the very first `updateFromHeaders()`/`updateFromResponseBody()`/ + * `applyRequestQueueSettings()` call against that limiter permanently kills its + * refresh. Once the reservoir then hits 0, it never refills again. + * + * Production symptom: an auto-enrolled apikey connection accumulates its + * default 60 requests, the reservoir zeroes, the request queue freezes for + * ~120s, the watchdog fires a synthetic 502 (RATE_LIMIT_QUEUE_WEDGED), the + * connection cools down and is excluded from weighted combo pools — turning a + * configured 70/30 split into ~50/50 (see + * tests/integration/combo-matrix/weighted.test.ts, the E2E proof for this + * same bug). + * + * This test drives the exact same sequence directly against + * open-sse/services/rateLimitManager.ts's public surface, without any DB or + * HTTP layer, to isolate the Bottleneck heartbeat defect on its own. + */ +import test from "node:test"; +import assert from "node:assert/strict"; + +const rateLimitManager = await import("../../open-sse/services/rateLimitManager.ts"); + +function wait(ms: number): Promise { + return new Promise((resolve) => setTimeout(resolve, ms)); +} + +const PROVIDER = "reservoir-refresh-test-provider"; +const CONNECTION_ID = "reservoir-refresh-test-conn"; + +test.after(async () => { + await rateLimitManager.__resetRateLimitManagerForTests(); +}); + +test("reservoir keeps refreshing after updateSettings() touches an already-heartbeating limiter", async () => { + rateLimitManager.enableRateLimitProtection(CONNECTION_ID); + + // 1. First call creates the limiter. Bottleneck's LocalDatastore constructor + // starts heartbeat #1 (alive) because the default reservoirRefreshInterval/ + // reservoirRefreshAmount are always set (buildLimiterDefaults()). + const warmup = await rateLimitManager.withRateLimit( + PROVIDER, + CONNECTION_ID, + null, + async () => "warmup" + ); + assert.equal(warmup, "warmup"); + + // 2. Header-learned update — the first *manual* updateSettings() call on this + // limiter. remaining(2) < limit(6000)*0.1 takes updateFromHeaders' "throttle" + // branch, which sets a real reservoir=2 with a 1s refresh window (limit=6000 + // keeps minTime at 0 so it doesn't pace the slot consumption below). This is + // exactly the call that kills the heartbeat under the unfixed Bottleneck bug. + rateLimitManager.updateFromHeaders( + PROVIDER, + CONNECTION_ID, + { + "x-ratelimit-limit-requests": "6000", + "x-ratelimit-remaining-requests": "2", + "x-ratelimit-reset-requests": "1s", + }, + 200 + ); + + // updateFromHeaders applies the limiter update asynchronously (fire-and-forget + // — see trackAsyncOperation in rateLimitManager.ts). Poll the test-only state + // hook until the reservoir actually lands at 2 instead of assuming a fixed + // number of event-loop ticks: Bottleneck's own updateSettings() goes through + // at least one real setTimeout(0) (yieldLoop) before storeOptions reflects the + // new value. + const pollDeadline = Date.now() + 2000; + let state = await rateLimitManager.__getLimiterStateForTests(PROVIDER, CONNECTION_ID, null); + while (state?.reservoir !== 2 && Date.now() < pollDeadline) { + await wait(10); + state = await rateLimitManager.__getLimiterStateForTests(PROVIDER, CONNECTION_ID, null); + } + assert.equal(state?.reservoir, 2, "reservoir must land at 2 before the slots below are consumed"); + + // 3. Consume both reservoir slots. + assert.equal( + await rateLimitManager.withRateLimit(PROVIDER, CONNECTION_ID, null, async () => "slot-1"), + "slot-1" + ); + assert.equal( + await rateLimitManager.withRateLimit(PROVIDER, CONNECTION_ID, null, async () => "slot-2"), + "slot-2" + ); + + // 4. Reservoir is now 0. A healthy Bottleneck heartbeat refills it ~1s later + // from the reservoirRefreshInterval/reservoirRefreshAmount configured above. + // Race a 3rd request against a 5s timer: if the heartbeat died (unfixed bug), + // the request stays QUEUED forever and the timer wins instead. + const RACE_TIMEOUT_MS = 5000; + let timeoutHandle: ReturnType | undefined; + const timeout = new Promise<"timed-out">((resolve) => { + timeoutHandle = setTimeout(() => resolve("timed-out"), RACE_TIMEOUT_MS); + }); + const request = rateLimitManager.withRateLimit( + PROVIDER, + CONNECTION_ID, + null, + async () => "slot-3" as const + ); + + const result = await Promise.race([request, timeout]); + if (timeoutHandle) clearTimeout(timeoutHandle); + + assert.equal( + result, + "slot-3", + 'reservoir must refresh ~1s after being exhausted; "timed-out" means the Bottleneck ' + + "heartbeat died after updateSettings() and the reservoir never refilled " + + "(node_modules/bottleneck/lib/LocalDatastore.js _startHeartbeat clearInterval-without-null bug)" + ); +});