mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-07-26 09:52:11 +03:00
* fix(sse): Gemini TPM classification, combo-cooldown-wait for auto/quota-share, and target-timeout floor
Gemini TPM/RPM 429s were misclassified as QUOTA_EXHAUSTED because
sanitizeErrorMessage() truncates to the first line, hiding Google's
metric name and retry hint on lines 2-3. Added a rawMessage field
(internal-only, never reaches the client) and classifyGeminiQuotaMetricFromText()
to classify from the untruncated text, reordered ahead of the generic
credits/daily-quota checks.
Widened comboCooldownWaitEnabled (wait out a short transient cooldown
instead of crystallizing a 429/503) from quota-share-only to also cover
auto-strategy combos, and raised the wait ceiling to 65s/130s-budget/90s-cap
to match Gemini's ~60s TPM/RPM windows.
The per-target timeout (DEFAULT_COMBO_TARGET_TIMEOUT_MS, 120s) was shorter
than the new 130s cooldown-wait budget, so a target could get cut off
mid-wait with a synthetic 524 instead of completing the retry. Added
resolveComboTargetTimeoutMsForCombo()/isComboCooldownWaitEligible() in
comboConfig.ts to raise the per-target floor to budgetMs+buffer only for
wait-eligible strategies (auto/quota-share), verified live: a 12-request
concurrent burst against a TPM-exhausted combo went from 2/12 succeeding
(10 x 524) to 12/12 succeeding with zero 503/524.
Also: liveGeminiShared.ts's sendAndValidate now fails fast on a 503
instead of retrying past it, and the health dashboard + request logger
surface TPM stats alongside RPM/RPD.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): combo-exhausted rejection logs now capture request body + attempted models
recordRejectedRequestUsage() (the fast path for combo requests that never
reach handleChatCore, e.g. all targets locked by resilience cooldown)
hardcoded provider: "-" and never passed a request body to saveCallLog(),
so /dashboard/logs entries for these failures were nearly useless for
debugging: no way to see the client's request or which models were tried.
- recordRejectedRequestUsage() now accepts requestBody and persists it
through the existing saveCallLog() artifact mechanism (same path
handleChatCore's own logging uses).
- Added summarizeComboAttemptedModels(), which reads the combo's own model
list (always available, unlike the response's combo-diagnostics headers —
a model-level resilience-lockout skip never touches the
exhaustedProviders/exhaustedConnections sets those headers are built
from) to populate a real "provider" value instead of "-".
- Wired both into the call site in src/sse/handlers/chat.ts.
NOTE: unrelated to the Gemini TPM/combo-cooldown-wait fix on this branch —
landed here per operator request, to be split into its own branch/PR.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* feat(sse): synthetic streaming keep-alive event + 5-minute Gemini cooldown-wait ceiling
Many clients enforce a first-SSE-byte timeout, which made it unsafe to wait
out a longer upstream rate-limit cooldown on a streaming request — the
client would abandon the connection before any bytes arrived. This landed
in two parts:
1. Synthetic startup "thinking" event (OpenAI chat/completions format):
the already-existing withEarlyStreamKeepalive wrapper (open-sse/utils/
earlyStreamKeepalive.ts, wired into /v1/chat/completions, /v1/messages,
/v1/responses since #2544) opens the SSE stream immediately once a
request runs past its threshold, but only ever sent empty/no-op
keepalive frames. Added a `startupFrame` option (defaults to
`keepaliveFrame` — zero behavior change unless a route opts in) so the
very first frame can carry real content instead. Wired
OPENAI_STARTUP_THINKING_FRAME (a reasoning_content delta: "OmniRoute:
got request, sending to provider") into /v1/chat/completions only —
Claude Messages and Responses API formats both require a preceding
envelope event (message_start / response.created) that a synthetic
pre-dispatch frame can't safely fabricate without risking a duplicate
envelope once the real stream arrives, so those two routes keep their
existing (safe, proven) keepalive frames unchanged.
2. Raised the "wait out a known cooldown, then retry" ceiling to 5 minutes
for both retry mechanisms, now that a client-side first-byte timeout is
no longer a risk on the (opted-in) route:
- comboCooldownWait (auto/quota-share combos, open-sse/services/combo.ts):
maxWaitMs hard clamp raised 90s -> 300s (src/lib/resilience/settings/
normalize.ts); defaults raised to maxWaitMs:90s/maxAttempts:5/
budgetMs:300s. comboConfig.ts's resolveComboTargetTimeoutMsForCombo
already derives the per-target timeout floor from budgetMs, so it
tracks the new ceiling with no further changes.
- waitForCooldown (direct, non-combo model requests, src/sse/handlers/
chat.ts): this mechanism had NO cumulative cap before — only a
per-wait cap (maxRetryWaitMs) and a retry count (maxRetries), so
maxRetries x maxRetryWaitMs could exceed 5 minutes with no ceiling.
Added a budgetMs field (mirrors comboCooldownWait) to
WaitForCooldownSettings/CooldownAwareRetrySettings, threaded a
requestRetryBudgetLeftMs tracker through chat.ts's requestAttemptLoop
(mirrors combo.ts's comboCooldownBudgetLeftMs), and made
getCooldownAwareRetryDecision refuse to wait once the cumulative
budget is exhausted even if the single wait is under maxRetryWaitMs.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): extend the synthetic keep-alive thinking event to /v1/responses
Live incident (OpenClaw, log id 1784407081908-cbc24f): a /v1/responses
request to gemini/gemma-4-31b-it took 56s to produce a first byte and the
client disconnected (499 request_signal_aborted) — the same client-first-byte-
timeout problem the previous commit fixed for /v1/chat/completions, but
/v1/responses only had the generic bare-comment keepalive (no content), so it
wasn't covered.
Added RESPONSES_STARTUP_THINKING_FRAME: a self-contained synthetic reasoning
item (response.output_item.added -> reasoning_summary_part.added ->
reasoning_summary_text.delta -> reasoning_summary_part.done), opened AND
closed within this one frame rather than left dangling — it never carries a
response_id, so it can't collide with the real upstream response's own
independent response.created lifecycle that follows. Mirrors the abbreviated
delta+part.done close pattern open-sse/utils/stream.ts's own
emitSyntheticResponsesReasoningSummary already uses for real mid-stream
reasoning content.
Wired into src/app/api/v1/responses/route.ts via the startupFrame option
added in the previous commit.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): combo cooldown-wait vars reset every setTry, crystallizing a bogus 503 instead of waiting
Live incident (log id 1784416706646-51): a request to the "default" combo
(strategy=auto, maxSetRetries=3) hit a real Gemini TPM 429 on both gemma-4
targets, correctly classified as a short 40s rate_limit lockout — then
crystallized a 503 "all upstream accounts are inactive" in 6.9s instead of
ever reaching the cooldown-aware wait.
Root cause: `lastError`/`earliestRetryAfter`/`lastStatus` were declared with
`let` INSIDE the `for (setTry...)` loop body, so they reset to null at the
start of every set-try. When both targets lock out on setTry 0, every
subsequent setTry (1..maxSetRetries) pre-skips both targets via the
isModelLocked check with no real dispatch — so on the FINAL setTry (the only
one whose values the post-loop decision reads, since it's gated behind
`if (setTry < maxSetRetries) continue`), lastStatus was null, hitting the
"!lastStatus" branch (ALL_ACCOUNTS_INACTIVE 503) and completely bypassing the
comboCooldownWaitEnabled / earliestRetryAfter wait logic — even though a
real 429 with a known ~40s retry-after WAS observed on setTry 0.
This bug predates today's Gemini TPM work (any combo with maxSetRetries > 0
whose targets all lock out on the first pass was affected) but was masked in
existing tests: the "auto strategy (2 models...)" regression test uses
maxSetRetries: 0, so it only ever runs ONE setTry iteration and never
exercises the reset-on-retry path. It also explains why the dedicated
12-concurrent-request burst test passed cleanly — with concurrent requests,
timing variance meant some request's FINAL setTry iteration still had a live
target to dispatch to, giving lastStatus/earliestRetryAfter fresh data. A
single isolated request has no such luck.
Fix: hoist lastError/earliestRetryAfter/lastStatus to just inside
dispatchWithCooldownRetry, before the setTry loop, so they persist across
set-tries (still reset fresh on each recursive dispatchWithCooldownRetry()
call after a wait, which is correct). recordedAttempts/fallbackCount/
exhaustedProviders etc. are intentionally left per-iteration (unrelated to
this bug).
New regression test in tests/unit/combo-quota-share-cooldown-wait.test.ts
reproduces the exact live scenario (2 targets, both lock out on setTry 0,
maxSetRetries: 3) — confirmed red (503) against the pre-fix code, green
(200, waits and retries) against the fix.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* test(sse): extend live Gemini workload to Responses API + add large-context TPM test
Two additions to the live Gemini test suite, both live-verified against the
dev instance:
1. sendAndValidate() (tests/integration/liveGeminiShared.ts) now accepts an
apiFormat: "chat" | "responses" parameter, building the Responses-API
request shape (input array, max_output_tokens) and parsing its SSE events
(response.output_text.delta / response.reasoning_summary_text.delta /
response.completed) via the new readResponsesSSEStream(). Wired into two
new tests in live-gemini-workload.test.ts ([30]/[31]), mirroring the
existing Chat Completions streaming coverage. Verified live: 24/25 + 5/5
payloads succeeded end-to-end through the new code path (the one failure
was a ~300s test-client fetch timeout unrelated to the Responses API code
itself — a separate, not-yet-addressed test-harness limitation).
2. genHugeContextMessage() builds a single message large enough (~4
chars/token estimate) to approach or exceed Gemini's free-tier TPM ceiling
(16000 input tokens/min for gemma-4) by itself. Every other prompt
generator in this file tops out around 1-2k tokens — nowhere near that
ceiling — so none of the existing workload tests ever exercised a REAL TPM
429, only RPM-style rate limiting. tests/integration/gemini-large-context-tpm.test.ts
sends two ~12-13k-token requests back-to-back (comfortably exceeding
16000/min together) to exercise the full path against production Gemini:
TPM classification, the comboCooldownWait retry, and the synthetic
keep-alive frame on a genuinely slow request.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): abandoned combo target dispatch now observes its own per-target timeout, fixing a permanent "pending" dashboard leak
Live incident (dashboard log id 1784418258231-14961a, reported as "an ongoing
request even though there's already a 200"): a combo target dispatch
abandoned by comboTargetTimeoutMs (open-sse/services/combo/targetTimeoutRunner.ts)
left a permanent phantom "pending" entry in the dashboard, even after the
overall combo request had already succeeded via a different retry.
Root cause: chatCore.ts's createStreamController — and everything downstream
that depends on it (withRateLimit's Promise.race against Bottleneck,
acquireAccountSemaphore) — only ever watches clientRawRequest.signal, which
is the ORIGINAL client's request signal (set once via buildClientRawRequest
and reused unchanged across every target dispatch in a combo). It has no
connection to targetTimeoutRunner.ts's OWN AbortController
(target.modelAbortSignal), which is what actually fires when
comboTargetTimeoutMs (300s) elapses. src/sse/handlers/chat.ts's
handleSingleModel bridge between combo.ts and handleSingleModelChat received
`target.modelAbortSignal` but silently dropped it — never forwarded it
anywhere. So when a target got abandoned (e.g. stuck inside a wedged
Bottleneck rate-limiter queue, see the WEDGED force-reset log line from the
same incident), its per-target timeout fired and let the COMBO move on and
retry successfully elsewhere — but the abandoned dispatch's own promise
chain never learned it had been superseded, so it hung forever waiting on a
signal that was never going to fire, and trackPendingRequest(false) (the
finalize call) never ran.
Fix: thread target.modelAbortSignal through as a new modelAbortSignal
runtimeOption, and merge it into clientRawRequest.signal (via the existing
mergeAbortSignals helper from open-sse/executors/base.ts) right before
dispatch, so an abandoned target's own promise chain now observes its abort
and can reach its cleanup path — new resolveDispatchClientRawRequest() makes
this mechanically testable in isolation.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): combo cooldown-wait state recording, rate-limit wedge recovery, OpenAI-format SSE error frames
Five related fixes surfaced by live incidents (dashboard log ids 1784457764961-73,
1784465227489-a2cbc0, 1784504040241-6f8b9a) while validating the Gemini TPM/cooldown-wait
work on this branch against real OpenClaw traffic:
- combo.ts: the model-lockout bail-out branches in dispatchWithCooldownRetry never
recorded lastStatus, so once every target in a set hit an existing lockout the final
check crystallized a bogus ALL_ACCOUNTS_INACTIVE 503 instead of reaching the
cooldown-wait decision, even with a real 429 + short retry-after observed.
- combo.ts/combo/types.ts: the "all credentials cooling down" pre-dispatch rejection
(buildModelCooldownBody) nests its retry hint as error.retry_after/reset_seconds, not
the top-level retryAfter every other 429 shape uses — combo's extraction only read the
latter, so earliestRetryAfter stayed null for this shape even after lastStatus was fixed.
- rateLimitManager.ts: the wedge-recovery watchdog used disconnect(), which releases the
heartbeat timer but never rejects jobs already QUEUED on that instance — orphaned
dispatches hung until the outer ~300s per-target timeout, well past real clients'
patience. Switched to stop({ dropWaitingJobs: true }), safe because the wedge condition
already requires RUNNING===0 && EXECUTING===0.
- earlyStreamKeepalive.ts: the in-band error frame emitted after committing to a 200 SSE
stream was hardcoded to Anthropic's `event: error` convention for every route, including
the OpenAI-format ones (/v1/chat/completions, /v1/responses) where that framing is
either invisible or malformed to a plain data-line parser. Added per-route
OPENAI_CHAT_ERROR_FRAME / OPENAI_RESPONSES_ERROR_FRAME and wired them in.
- chatCore.ts: persisted a synthetic clientResponse error body even when the client had
already disconnected (AbortError) before that body was ever computed — misleading the
dashboard into showing "what the client received" for a response that was never sent.
Also: RequestLoggerDetail.tsx — Provider/Client Event Stream panes lost their collapse
toggle when StreamSection replaced the collapsible PayloadSection (692d6be80, unifying
active/finished request views) without carrying the toggle over.
Each fix has a TDD regression test with a confirmed red-before-green cycle.
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
* test(sse): free-tier model + gemma-4 TPM-ceiling benchmark harness
Adds a live benchmark comparing free models OmniRoute exposes across
configured providers plus previously-unexercised no-auth providers
(felo-web, aihorde, opencode, duckduckgo-web — none need a connection
row, they were just never tried). Reuses liveGeminiShared.ts's SSE
parsers and CASE_BUILDERS instead of duplicating them.
Also adds a targeted TPM-stress test firing back-to-back large-context
prompts at the gemma-4-31b model across its 3 free hosts (Gemini,
NVIDIA, AI Horde) to isolate whether the documented 16k-tokens/minute
free-tier ceiling is Gemini-specific enforcement or an inherent
model property.
FORCE_TOOL_CHOICE_REQUIRED is a test-only, default-off env flag added
to liveGeminiShared.ts and live-gemini-agentic-loop.test.ts for an
earlier live A/B comparison of tool_choice: required vs unset — kept
as a reusable knob for future runs.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* test(sse): benchmark for the 2026-07-22 newly-enabled provider batch
Adds NEWLY_ENABLED_MODELS to freeModelBenchmarkShared.ts (Mistral
Leanstral, OpenRouter's live "free"-tagged roster, OpenCode Zen's
current free models — refetched live from
https://opencode.ai/zen/v1/models since the static catalog had
drifted) and a dedicated workload benchmark test for them.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* test(sse): sync geminiRateLimitTracker tests with e74a1722b's corrected Gemma 4 limits
e74a1722b updated geminiRateLimits.json's gemma-4-* entries from the stale
15/1500/-1 (rpm/rpd/tpm) to the real published free-tier values
16000/14400/16000, but never updated the tests asserting the old numbers.
Surfaced by running the full test:unit suite as a post-rebase sanity check.
Co-Authored-By: Markus Hartung <markus.hartream@gmail.com>
* chore(quality): file-size baseline for own-growth (#8213)
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
---------
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
438 lines
19 KiB
TypeScript
438 lines
19 KiB
TypeScript
/**
|
|
* Verifies that API routes sanitize error messages (CodeQL js/stack-trace-exposure)
|
|
* and that security-critical helpers behave correctly.
|
|
*/
|
|
import test from "node:test";
|
|
import assert from "node:assert/strict";
|
|
import fs from "node:fs";
|
|
import os from "node:os";
|
|
import path from "node:path";
|
|
|
|
const TEST_DATA_DIR = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-err-sanitize-"));
|
|
process.env.DATA_DIR = TEST_DATA_DIR;
|
|
process.env.API_KEY_SECRET = "test-api-key-secret-32chars-long!!";
|
|
|
|
const core = await import("../../src/lib/db/core.ts");
|
|
const combosDb = await import("../../src/lib/db/combos.ts");
|
|
const mappingsRoute = await import("../../src/app/api/model-combo-mappings/route.ts");
|
|
const mappingsIdRoute = await import("../../src/app/api/model-combo-mappings/[id]/route.ts");
|
|
const syncTokens = await import("../../src/lib/sync/tokens.ts");
|
|
|
|
const repoRoot = path.resolve(import.meta.dirname, "../..");
|
|
const read = (relPath: string) => fs.readFileSync(path.join(repoRoot, relPath), "utf8");
|
|
|
|
function makeRequest(url: string, options: { method?: string; body?: unknown } = {}) {
|
|
const { method = "GET", body } = options;
|
|
return new Request(url, {
|
|
method,
|
|
headers: body !== undefined ? { "content-type": "application/json" } : undefined,
|
|
body: body !== undefined ? JSON.stringify(body) : undefined,
|
|
});
|
|
}
|
|
|
|
async function resetStorage() {
|
|
core.resetDbInstance();
|
|
fs.rmSync(TEST_DATA_DIR, { recursive: true, force: true });
|
|
fs.mkdirSync(TEST_DATA_DIR, { recursive: true });
|
|
}
|
|
|
|
test.beforeEach(async () => {
|
|
await resetStorage();
|
|
});
|
|
|
|
test.after(() => {
|
|
core.resetDbInstance();
|
|
fs.rmSync(TEST_DATA_DIR, { recursive: true, force: true });
|
|
});
|
|
|
|
async function createCombo(name: string, model: string) {
|
|
return combosDb.createCombo({
|
|
name,
|
|
models: [{ provider: "openai", model }],
|
|
strategy: "priority",
|
|
config: {},
|
|
});
|
|
}
|
|
|
|
// ── model-combo-mappings routes ──────────────────────────────────────────────
|
|
|
|
test("GET /model-combo-mappings returns empty list on fresh DB", async () => {
|
|
const res = await mappingsRoute.GET(makeRequest("http://localhost/api/model-combo-mappings"));
|
|
assert.equal(res.status, 200);
|
|
const body = (await res.json()) as any;
|
|
assert.ok(Array.isArray(body.mappings), "body.mappings must be an array");
|
|
assert.equal(body.mappings.length, 0);
|
|
assert.ok(!("error" in body), "success response must not contain error field");
|
|
});
|
|
|
|
test("GET /model-combo-mappings error response never leaks raw error.message", async () => {
|
|
const res = await mappingsRoute.GET(makeRequest("http://localhost/api/model-combo-mappings"));
|
|
// In the success case, there is no error field at all
|
|
const body = (await res.json()) as any;
|
|
if (res.status >= 500) {
|
|
assert.equal(body.error, "Failed to list model-combo mappings");
|
|
assert.ok(!("stack" in body), "stack trace must not be present in response");
|
|
}
|
|
});
|
|
|
|
test("POST /model-combo-mappings returns 400 for empty pattern", async () => {
|
|
const res = await mappingsRoute.POST(
|
|
makeRequest("http://localhost/api/model-combo-mappings", {
|
|
method: "POST",
|
|
body: { pattern: "", comboId: "combo-1" },
|
|
})
|
|
);
|
|
assert.equal(res.status, 400);
|
|
const body = (await res.json()) as any;
|
|
assert.ok("error" in body);
|
|
assert.ok(!("stack" in body), "400 response must not contain stack trace");
|
|
});
|
|
|
|
test("POST /model-combo-mappings returns 400 for missing comboId", async () => {
|
|
const res = await mappingsRoute.POST(
|
|
makeRequest("http://localhost/api/model-combo-mappings", {
|
|
method: "POST",
|
|
body: { pattern: "gpt-*" },
|
|
})
|
|
);
|
|
assert.equal(res.status, 400);
|
|
});
|
|
|
|
test("POST /model-combo-mappings creates a mapping and response has no error field", async () => {
|
|
const combo = await createCombo("test-combo", "gpt-4o");
|
|
const res = await mappingsRoute.POST(
|
|
makeRequest("http://localhost/api/model-combo-mappings", {
|
|
method: "POST",
|
|
body: { pattern: "gpt-*", comboId: combo.id },
|
|
})
|
|
);
|
|
assert.equal(res.status, 201);
|
|
const body = (await res.json()) as any;
|
|
assert.ok("mapping" in body, "response must have mapping field");
|
|
assert.ok(!("error" in body), "success response must not contain error field");
|
|
assert.ok(!("stack" in body));
|
|
assert.equal(body.mapping.pattern, "gpt-*");
|
|
});
|
|
|
|
test("GET /model-combo-mappings/[id] returns 404 for non-existent id", async () => {
|
|
const res = await mappingsIdRoute.GET(
|
|
makeRequest("http://localhost/api/model-combo-mappings/nonexistent"),
|
|
{ params: Promise.resolve({ id: "nonexistent" }) }
|
|
);
|
|
assert.equal(res.status, 404);
|
|
const body = (await res.json()) as any;
|
|
assert.equal(body.error, "Mapping not found");
|
|
assert.ok(!("stack" in body), "404 response must not contain stack trace");
|
|
});
|
|
|
|
test("GET /model-combo-mappings/[id] error response never leaks internal details", async () => {
|
|
const res = await mappingsIdRoute.GET(
|
|
makeRequest("http://localhost/api/model-combo-mappings/some-id"),
|
|
{ params: Promise.resolve({ id: "some-id" }) }
|
|
);
|
|
const body = (await res.json()) as any;
|
|
if (res.status >= 500) {
|
|
assert.equal(body.error, "Failed to get mapping");
|
|
assert.ok(!body.error.includes("SQLITE"), "SQLite internals must not be exposed");
|
|
assert.ok(!("stack" in body));
|
|
}
|
|
});
|
|
|
|
test("DELETE /model-combo-mappings/[id] returns 404 for non-existent mapping", async () => {
|
|
const res = await mappingsIdRoute.DELETE(
|
|
makeRequest("http://localhost/api/model-combo-mappings/nonexistent", { method: "DELETE" }),
|
|
{ params: Promise.resolve({ id: "nonexistent" }) }
|
|
);
|
|
assert.equal(res.status, 404);
|
|
const body = (await res.json()) as any;
|
|
assert.equal(body.error, "Mapping not found");
|
|
assert.ok(!("stack" in body));
|
|
});
|
|
|
|
test("PUT /model-combo-mappings/[id] returns 404 for non-existent mapping", async () => {
|
|
const res = await mappingsIdRoute.PUT(
|
|
makeRequest("http://localhost/api/model-combo-mappings/nonexistent", {
|
|
method: "PUT",
|
|
body: { pattern: "new-*" },
|
|
}),
|
|
{ params: Promise.resolve({ id: "nonexistent" }) }
|
|
);
|
|
assert.equal(res.status, 404);
|
|
const body = (await res.json()) as any;
|
|
assert.equal(body.error, "Mapping not found");
|
|
assert.ok(!("stack" in body));
|
|
});
|
|
|
|
// ── sync token hashing (src/lib/sync/tokens.ts) ──────────────────────────────
|
|
|
|
test("hashSyncToken returns a 64-character hex string (SHA-256 output)", () => {
|
|
const token = syncTokens.generatePlaintextSyncToken();
|
|
const hash = syncTokens.hashSyncToken(token);
|
|
assert.match(hash, /^[0-9a-f]{64}$/, "hash must be 64 lowercase hex chars");
|
|
});
|
|
|
|
test("hashSyncToken is deterministic — same input always produces same output", () => {
|
|
const token = syncTokens.generatePlaintextSyncToken();
|
|
assert.equal(
|
|
syncTokens.hashSyncToken(token),
|
|
syncTokens.hashSyncToken(token),
|
|
"hashing the same token twice must yield the same result"
|
|
);
|
|
});
|
|
|
|
test("hashSyncToken produces different hashes for different tokens", () => {
|
|
const a = syncTokens.generatePlaintextSyncToken();
|
|
const b = syncTokens.generatePlaintextSyncToken();
|
|
assert.notEqual(
|
|
syncTokens.hashSyncToken(a),
|
|
syncTokens.hashSyncToken(b),
|
|
"different tokens must produce different hashes"
|
|
);
|
|
});
|
|
|
|
test("generatePlaintextSyncToken starts with osync_ prefix", () => {
|
|
const token = syncTokens.generatePlaintextSyncToken();
|
|
assert.ok(
|
|
token.startsWith("osync_"),
|
|
`token must start with 'osync_', got: ${token.slice(0, 10)}`
|
|
);
|
|
});
|
|
|
|
test("hashSyncToken output is never the plain token (not stored in clear text)", () => {
|
|
const token = syncTokens.generatePlaintextSyncToken();
|
|
const hash = syncTokens.hashSyncToken(token);
|
|
assert.notEqual(hash, token, "hash must differ from plaintext token");
|
|
assert.ok(!hash.startsWith("osync_"), "hash must not start with the token prefix");
|
|
});
|
|
|
|
test("sanitizeErrorMessage strips multi-line stack traces", async () => {
|
|
const { sanitizeErrorMessage } = await import("../../open-sse/utils/error.ts");
|
|
const input =
|
|
"Cannot read property 'foo' of undefined\n at handler (/srv/app/src/lib/x.ts:42:11)\n at next (internal)";
|
|
const out = sanitizeErrorMessage(input);
|
|
assert.equal(out, "Cannot read property 'foo' of undefined");
|
|
assert.ok(!out.includes("at handler"));
|
|
});
|
|
|
|
test("sanitizeErrorMessage replaces absolute paths with <path>", async () => {
|
|
const { sanitizeErrorMessage } = await import("../../open-sse/utils/error.ts");
|
|
const out1 = sanitizeErrorMessage("Failed to open /home/user/secret-project/src/config.ts:10");
|
|
assert.ok(!out1.includes("/home/user/secret-project"));
|
|
assert.ok(out1.includes("<path>"));
|
|
|
|
const out2 = sanitizeErrorMessage("Module not found: C:\\Users\\admin\\app\\index.js:1:1");
|
|
assert.ok(!out2.includes("C:\\Users\\admin"));
|
|
assert.ok(out2.includes("<path>"));
|
|
});
|
|
|
|
test("sanitizeErrorMessage handles non-string inputs safely", async () => {
|
|
const { sanitizeErrorMessage } = await import("../../open-sse/utils/error.ts");
|
|
assert.equal(sanitizeErrorMessage(undefined), "");
|
|
assert.equal(sanitizeErrorMessage(null), "");
|
|
assert.equal(sanitizeErrorMessage(42), "42");
|
|
assert.equal(sanitizeErrorMessage(new Error("boom")), "Error: boom");
|
|
});
|
|
|
|
test("buildErrorBody never exposes stack traces in its message", async () => {
|
|
const { buildErrorBody } = await import("../../open-sse/utils/error.ts");
|
|
const body = buildErrorBody(
|
|
500,
|
|
"Internal error\n at /opt/app/src/server.ts:99:7\n at next (internal)"
|
|
);
|
|
assert.equal(body.error.message, "Internal error");
|
|
assert.ok(!body.error.message.includes("at /opt"));
|
|
});
|
|
|
|
test("types barrel keeps the model cooldown payload export only", async () => {
|
|
const src = await read("src/types/index.ts");
|
|
assert.match(src, /ModelCooldownErrorPayload/);
|
|
assert.doesNotMatch(src, /ProviderConnection/);
|
|
assert.doesNotMatch(src, /ProviderNode/);
|
|
});
|
|
|
|
// ── sanitizeUpstreamDetails ──────────────────────────────────────────────────
|
|
|
|
test("sanitizeUpstreamDetails — basic pass-through for safe fields", async () => {
|
|
const { sanitizeUpstreamDetails } = await import("../../open-sse/utils/error.ts");
|
|
const input = { error: { message: "context_length_exceeded", type: "invalid_request_error" } };
|
|
const out = sanitizeUpstreamDetails(input) as any;
|
|
assert.equal(out.error.message, "context_length_exceeded");
|
|
assert.equal(out.error.type, "invalid_request_error");
|
|
});
|
|
|
|
test("sanitizeUpstreamDetails — sanitizes string values (absolute path)", async () => {
|
|
const { sanitizeUpstreamDetails } = await import("../../open-sse/utils/error.ts");
|
|
const input = { error: { message: "bad input at /srv/app/src/lib/db.ts:42" } };
|
|
const out = sanitizeUpstreamDetails(input) as any;
|
|
assert.ok(
|
|
!out.error.message.includes("/srv/app/src/lib/db.ts"),
|
|
"absolute path must be stripped"
|
|
);
|
|
assert.ok(out.error.message.includes("<path>"), "path placeholder must be present");
|
|
});
|
|
|
|
test("sanitizeUpstreamDetails — removes blocked keys (stack, apiKey)", async () => {
|
|
const { sanitizeUpstreamDetails } = await import("../../open-sse/utils/error.ts");
|
|
const input = {
|
|
error: { message: "oops" },
|
|
stack: "Error\n at foo.ts:1",
|
|
apiKey: "sk-secret",
|
|
};
|
|
const out = sanitizeUpstreamDetails(input) as any;
|
|
assert.ok(!("stack" in out), "stack key must be removed");
|
|
assert.ok(!("apiKey" in out), "apiKey key must be removed");
|
|
assert.equal(out.error.message, "oops");
|
|
});
|
|
|
|
test("sanitizeUpstreamDetails — depth cap replaces nested value at depth > 4", async () => {
|
|
const { sanitizeUpstreamDetails } = await import("../../open-sse/utils/error.ts");
|
|
// Build depth-6 nesting: a.b.c.d.e.f = "leaf"
|
|
const input = { a: { b: { c: { d: { e: { f: "leaf" } } } } } };
|
|
const out = sanitizeUpstreamDetails(input) as any;
|
|
// depth 0:a, 1:b, 2:c, 3:d, 4:e → e is at depth 4, f would be depth 5 → truncated
|
|
assert.equal(out.a.b.c.d.e, "[truncated]");
|
|
});
|
|
|
|
// ── buildErrorBody with upstreamDetails ──────────────────────────────────────
|
|
|
|
test("buildErrorBody — without upstream details omits upstream_details field", async () => {
|
|
const { buildErrorBody } = await import("../../open-sse/utils/error.ts");
|
|
const body = buildErrorBody(400, "bad request");
|
|
assert.ok(!("upstream_details" in body), "upstream_details must be absent when not provided");
|
|
});
|
|
|
|
test("buildErrorBody — with safe upstream details embeds upstream_details", async () => {
|
|
const { buildErrorBody } = await import("../../open-sse/utils/error.ts");
|
|
const body = buildErrorBody(400, "bad request", {
|
|
error: { message: "context_length_exceeded" },
|
|
});
|
|
assert.ok("upstream_details" in body, "upstream_details must be present");
|
|
assert.equal((body.upstream_details as any).error.message, "context_length_exceeded");
|
|
});
|
|
|
|
test("buildErrorBody — upstream details with stack key are stripped", async () => {
|
|
const { buildErrorBody } = await import("../../open-sse/utils/error.ts");
|
|
const body = buildErrorBody(500, "err", { stack: "Error\n at foo.ts:1", code: "internal" });
|
|
assert.ok("upstream_details" in body, "upstream_details must be present");
|
|
assert.ok(
|
|
!("stack" in (body.upstream_details as any)),
|
|
"stack must be stripped from upstream_details"
|
|
);
|
|
assert.equal((body.upstream_details as any).code, "internal");
|
|
});
|
|
|
|
// ── createErrorResult with upstreamDetails ───────────────────────────────────
|
|
|
|
test("createErrorResult — response body includes upstream_details when provided", async () => {
|
|
const { createErrorResult } = await import("../../open-sse/utils/error.ts");
|
|
const result = createErrorResult(
|
|
400,
|
|
"context too long",
|
|
null,
|
|
"context_length_exceeded",
|
|
"invalid_request_error",
|
|
{ error: { message: "context_length_exceeded" } }
|
|
);
|
|
const body = (await result.response.clone().json()) as any;
|
|
assert.ok("upstream_details" in body, "upstream_details must be in response body");
|
|
assert.equal(body.upstream_details.error.message, "context_length_exceeded");
|
|
});
|
|
|
|
test("createErrorResult — response body excludes upstream_details when not provided", async () => {
|
|
const { createErrorResult } = await import("../../open-sse/utils/error.ts");
|
|
const result = createErrorResult(400, "bad request", null, "bad_request");
|
|
const body = (await result.response.clone().json()) as any;
|
|
assert.ok(!("upstream_details" in body), "upstream_details must be absent when not provided");
|
|
});
|
|
|
|
test("createErrorResult — exposes error code/type on the result object", async () => {
|
|
const { createErrorResult } = await import("../../open-sse/utils/error.ts");
|
|
const result = createErrorResult(504, "upstream timeout", null, "UPSTREAM_TIMEOUT", "timeout");
|
|
assert.equal(result.errorCode, "UPSTREAM_TIMEOUT");
|
|
assert.equal(result.errorType, "timeout");
|
|
});
|
|
|
|
// ── createErrorResult.rawMessage (#7360) ──────────────────────────────────────
|
|
//
|
|
// `error` is sanitized to its first line (sanitizeErrorMessage) for the
|
|
// client-facing response body — correct per Hard Rule #12. But internal
|
|
// classification (checkFallbackError / Gemini TPM-vs-RPD metric detection)
|
|
// needs the FULL multi-line upstream text, since Google's metric name and
|
|
// retry hint live on lines 2-3. `rawMessage` carries the untruncated text on
|
|
// the returned object only — it must never leak into the HTTP response body.
|
|
|
|
test("createErrorResult — rawMessage preserves the full multi-line message untruncated", async () => {
|
|
const { createErrorResult } = await import("../../open-sse/utils/error.ts");
|
|
const fullMessage =
|
|
"You exceeded your current quota, please check your plan and billing details.\n" +
|
|
"* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_input_token_count, limit: 16000, model: gemma-4-31b\n" +
|
|
"Please retry in 8.093498133s.";
|
|
const result = createErrorResult(429, fullMessage);
|
|
|
|
assert.equal(result.rawMessage, fullMessage, "rawMessage must be the complete, untruncated text");
|
|
assert.ok(
|
|
result.error.length < fullMessage.length,
|
|
"error (client-facing) must still be truncated to the first line"
|
|
);
|
|
assert.ok(
|
|
!result.error.includes("generativelanguage.googleapis.com"),
|
|
"sanitized error must not include the metric name (line 2)"
|
|
);
|
|
});
|
|
|
|
test("createErrorResult — rawMessage never appears in the serialized response body", async () => {
|
|
const { createErrorResult } = await import("../../open-sse/utils/error.ts");
|
|
const fullMessage =
|
|
"You exceeded your current quota, please check your plan and billing details.\n" +
|
|
"* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_input_token_count, limit: 16000, model: gemma-4-31b";
|
|
const result = createErrorResult(429, fullMessage);
|
|
const bodyText = await result.response.clone().text();
|
|
|
|
assert.ok(
|
|
!bodyText.includes("generativelanguage.googleapis.com"),
|
|
"the raw multi-line metric text must never reach the HTTP response body"
|
|
);
|
|
});
|
|
|
|
test("buildModelCooldownBody returns the public cooldown error payload shape", async () => {
|
|
const { buildModelCooldownBody } = await import("../../open-sse/utils/error.ts");
|
|
const body = buildModelCooldownBody({ model: "gpt-4o", retryAfterSec: 1.2 });
|
|
|
|
assert.deepEqual(body, {
|
|
error: {
|
|
message: "All credentials for model gpt-4o are cooling down",
|
|
type: "rate_limit_error",
|
|
code: "model_cooldown",
|
|
model: "gpt-4o",
|
|
reset_seconds: 2,
|
|
},
|
|
});
|
|
});
|
|
|
|
test("regression: upstream_details never contains stack trace text", async () => {
|
|
const { createErrorResult } = await import("../../open-sse/utils/error.ts");
|
|
const upstream = { error: { message: "err" }, stack: "Error\n at /abs/path.ts:1:2" };
|
|
const result = createErrorResult(500, "upstream err", null, undefined, undefined, upstream);
|
|
const body = (await result.response.clone().json()) as any;
|
|
const serialized = JSON.stringify(body);
|
|
assert.ok(
|
|
!serialized.includes("at /abs/path.ts"),
|
|
"stack trace path must not appear in response body"
|
|
);
|
|
assert.ok(!("stack" in (body.upstream_details || {})), "stack key must not be present");
|
|
});
|
|
|
|
// ── existing tests continue ──────────────────────────────────────────────────
|
|
|
|
test("GET /token-health response never leaks stack frames or absolute paths", async () => {
|
|
const tokenHealthRoute = await import("../../src/app/api/token-health/route.ts");
|
|
const res = await tokenHealthRoute.GET();
|
|
const body = (await res.json()) as any;
|
|
assert.ok(!("stack" in body), "response must not contain stack trace");
|
|
if (typeof body.error === "string") {
|
|
assert.ok(!body.error.includes(" at "), "stack frame must not leak in error");
|
|
assert.ok(!/^\//.test(body.error), "absolute POSIX path must not leak");
|
|
assert.ok(!/^[A-Za-z]:[\\/]/.test(body.error), "absolute Windows path must not leak");
|
|
}
|
|
});
|