mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-02 21:32:10 +03:00
* fix(sse): collapse single-text-part Responses-API content to a plain string
Every /v1/responses request — even the simplest single-string input —
got 500'd by AI Horde's Aphrodite-backed facade. Root cause:
normalizeResponsesInputForChat() always wraps a plain string input as
`content: [{ type: "input_text", text: value }]` (a one-element array),
and openaiResponsesToOpenAIRequest() mapped that straight through to
`content: [{ type: "text", text: value }]` on the Chat Completions side
— an array. That's spec-valid (OpenAI's own API accepts both shapes),
but strict/naive OpenAI-compatible backends like AI Horde's only
implement the plain-string form and reject the array form outright.
A single-text-part array and a plain string are semantically
identical, so collapse is safe. Real multi-part messages (text+image,
text+file) are left untouched.
Regression test: tests/unit/openai-responses-single-text-content-string.test.ts
(RED before the fix — every collapsed-content assertion failed with an
object instead of a string; GREEN after).
Also adds a deeper AI Horde load-test suite (sequential/concurrent/
cross-model/sustained-throughput/new-capable-model-candidates) that
surfaced this bug via real live traffic after Behemoth-X-123B was
temporarily added to the "default" combo for evaluation.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): unsupportedParams provider-level fallback for aihorde's live-discovered models
Real OpenClaw traffic against the newly-added Behemoth-X-123B combo
target kept 500ing on every attempt even after the Responses-API
content-array fix landed. The pipeline artifact showed why: `tools`
was still present, unstripped, in the request actually sent to AI
Horde's Aphrodite backend.
Root cause: `unsupportedParams: ["tools", "tool_choice",
"parallel_tool_calls"]` was only declared on the 3 models statically
listed in the aihorde registry entry (Cydonia-24B, Skyfall-31B,
google/gemma-4-31b). AI Horde uses `passthroughModels: true` — its
live worker roster changes constantly — so Behemoth-X-123B, like every
other dynamically-discovered aihorde model, had no model-specific
unsupportedParams entry, and getUnsupportedParams() returned [] for
it. But "the workers run raw text-completion backends" (no tool
calling) is true of every model AI Horde serves, not just the 3
catalogued ones.
Adds a provider-level `unsupportedParams` fallback on RegistryEntry,
checked by getUnsupportedParams() after the per-model lookup misses.
Set on the aihorde entry so it covers its entire live-discovered
roster, present and future, without needing a static per-model catalog
entry for each one.
Regression test: tests/unit/aihorde-tools-unsupported-provider-fallback.test.ts
(RED before the fix — Behemoth-X and deepseek-v4-flash both returned
[] instead of the stripped param list; GREEN after, with a control
case confirming the fallback doesn't leak to unrelated providers).
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): flatten leftover tool-call history when stripping unsupported tools
Third bug in the same AI Horde/Behemoth-X saga: even after tools/
tool_choice were correctly stripped from the live request (previous
fix), real combo traffic still 500'd. The conversation history itself
carried a prior turn's role:"assistant" tool_calls and role:"tool"
result messages, left over from before the combo failed over from a
tool-capable model (Gemini) to a non-tool-capable one (AI Horde). Its
raw completion backend doesn't understand those message shapes at all,
independent of whether live `tools` is present — confirmed by
reproducing with a role:"tool" message and NO tools param at all.
flattenToolHistory() (open-sse/utils/flattenToolHistory.ts) already
existed for exactly this, fully unit-tested — it just had zero call
sites anywhere in the request pipeline. Extracts the unsupported-params
strip into a small testable module
(open-sse/handlers/chatCore/unsupportedParamsStrip.ts, following the
existing chatCore god-file decomposition pattern e.g.
executorClientHeaders.ts) that now also flattens tool-call history
whenever "tools" was among the stripped params.
Regression test: tests/unit/chatcore-unsupported-params-strip.test.ts
(RED before the fix — the flattening test failed with the raw
tool_calls array still present; GREEN after). All 434 existing
chatcore-*.test.ts tests still pass.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): gate tool-history flattening on unsupported, not on stripped-this-request
The previous commit's flattening only fired when "tools" was actually
present-and-stripped on THIS request. A second live reproduction
against AI Horde had no live `tools` param at all — only stale
tool_calls/tool-result messages inherited from before a combo
failover — and still 500'd, because that condition never triggered.
A model that can't do tool calling can't do it whether or not the
current request happens to carry a `tools` array. Gate on the
unsupported-params list itself (unsupported.includes("tools")) instead
of the subset that was actually present-and-deleted this time.
Regression test added to the same file (RED before — the no-live-tools
case left tool_calls/role:"tool" untouched; GREEN after). All 435
chatcore-*.test.ts still pass.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): skip tool-incapable combo targets, error clearly on direct requests
Two complementary fixes for a model that structurally can't do tool
calling at all (e.g. AI Horde's raw completion backends) rather than
silently degrading — following up on the earlier strip/flatten fix,
which stopped the crashes but let a tool-incapable target still get
selected and return a 200 that narrates a fake tool call in prose
instead of erroring or being skipped.
1. Root cause, combo routing: getResolvedModelCapabilities()'s
`supportsTools` resolution only checked per-model registry entries,
synced capabilities, and static specs — none of which exist for a
dynamically-discovered model (AI Horde's passthroughModels roster
changes as workers come and go). It fell through to
heuristicToolCalling(), which optimistically defaults to `true` for
any unrecognized model (TOOL_CALLING_UNSUPPORTED_PATTERNS is empty).
Added a provider-level fallback reusing the same unsupportedParams
signal the request-time strip already relies on. This makes the
EXISTING filterTargetsByRequestCompatibility (comboStructure.ts) —
which already correctly excludes non-tool-capable targets when a
request requires tools — actually work for these models; no combo.ts
changes were needed, it was only ever fed bad capability data.
2. Direct/pinned requests: filterTargetsByRequestCompatibility only
protects combo routing. A direct request naming an exact
tool-incapable model has no other target to fail over to — added
checkToolCallingRequiredButUnsupported (chatCore/toolCallingRequiredCheck.ts),
gated on isCombo: false, returning a clear 400 instead of a 200 that
silently can't do what was asked.
Regression tests (both RED before, GREEN after):
- tests/unit/model-capabilities-provider-unsupported-tools.test.ts
- tests/unit/chatcore-tool-calling-required-check.test.ts
All 463 chatcore-*/model-capabilities-*.test.ts and 31 combo
compatibility-filter tests still pass.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): correct handleChatCore return shape for the tool-calling-blocked error
handleChatCore's documented contract is `{ success, response, status,
error }`, not a raw Response — returning `new Response(...)` directly
(copied from a different early-return whose surrounding context turned
out not to share this function's top-level contract) produced "No
response is returned from route handler ... Expected a Response object
but received 'undefined'" and a bare 500 with an empty body, caught
immediately when verifying the previous commit live.
Uses createErrorResult() (already used by the adjacent
translation-failure branch a few lines up) instead of hand-building the
Response, matching the same pattern already established in this
function for early error returns.
Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>
* fix(sse): scope Responses single-text-content collapse to providers that need it
The single-text-part content array -> plain string collapse (added for AI
Horde's Aphrodite facade, which 500s on the array form) was applied
unconditionally to every provider, silently breaking the standard OpenAI
array-shaped content contract that other providers and existing tests
depend on. Added RegistryEntry.requiresPlainStringContent, gated the
collapse on it (true only for aihorde), and threaded modelInfo.provider
through responsesHandler -> responsesApiHelper -> the translator so the
real /v1/responses call site can identify the provider.
Co-Authored-By: Markus Hartung <markus.hartream@gmail.com>
---------
Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
289 lines
10 KiB
TypeScript
289 lines
10 KiB
TypeScript
/**
|
||
* tests/integration/aihorde-load-test.test.ts
|
||
*
|
||
* Deeper load test for AI Horde (aihorde.net) — a free, crowdsourced,
|
||
* kudos-priority-queued inference network — asking whether it's viable as
|
||
* an addition/alternative to the "default" combo (currently 2 Gemini
|
||
* gemma-4 targets, see DEFAULT_COMBO_CONFIG in liveGeminiShared.ts).
|
||
* Anonymous access (key "0000000000") gets the LOWEST queue priority and
|
||
* 0 starting kudos, so this specifically probes whether that translates to
|
||
* unacceptable latency/reliability for combo-style traffic, or whether it
|
||
* holds up.
|
||
*
|
||
* Three angles, all against the anonymous no-auth `aihorde` provider:
|
||
* 1. Sequential reliability — all 25 CASE_BUILDERS, one at a time, with
|
||
* full latency distribution (p50/p90/max) on gemma-4-31b (the same
|
||
* model family the real "default" combo runs, for a fair comparison).
|
||
* 2. Concurrent load — does firing several requests at once cause queue
|
||
* pile-up/timeouts, given combo routing doesn't pace requests apart?
|
||
* 3. Cross-model spot-check — same handful of cases against the other 2
|
||
* viable aihorde models, since a combo pool needs multiple healthy
|
||
* targets, not just one.
|
||
*
|
||
* Report, not a gate: kudos-priority queueing is a genuine, permanent
|
||
* property of anonymous AI Horde access, not a bug — a model or two being
|
||
* slow is data for the "is this viable" question, not a regression.
|
||
*
|
||
* Environment:
|
||
* OMNIROUTE_API_KEY — required (else test skips)
|
||
* OMNIROUTE_URL — defaults to http://localhost:3000
|
||
*/
|
||
import test from "node:test";
|
||
import assert from "node:assert/strict";
|
||
|
||
import { skip, CASE_BUILDERS, ensureTestEnvironment } from "./liveGeminiShared.ts";
|
||
import {
|
||
benchmarkRequest,
|
||
benchmarkTpmStress,
|
||
summarize,
|
||
formatBenchmarkTable,
|
||
BENCHMARK_CASES,
|
||
type FreeModelSpec,
|
||
type BenchmarkResult,
|
||
type ModelBenchmarkSummary,
|
||
} from "./freeModelBenchmarkShared.ts";
|
||
|
||
const GEMMA_4 = {
|
||
provider: "aihorde",
|
||
model: "aihorde/google/gemma-4-31b",
|
||
displayName: "Gemma 4 31B (AI Horde)",
|
||
};
|
||
const OTHER_MODELS: FreeModelSpec[] = [
|
||
{
|
||
provider: "aihorde",
|
||
model: "aihorde/aphrodite/TheDrummer/Cydonia-24B-v4.3",
|
||
displayName: "Cydonia 24B (AI Horde)",
|
||
},
|
||
{
|
||
provider: "aihorde",
|
||
model: "aihorde/aphrodite/TheDrummer/Skyfall-31B-v4.2",
|
||
displayName: "Skyfall 31B (AI Horde)",
|
||
},
|
||
];
|
||
|
||
// Bigger/more-capable candidates found via AI Horde's live status API
|
||
// (https://aihorde.net/api/v2/status/models?type=text — the 3 models above
|
||
// are just a static passthrough-discovery fallback, not the full live
|
||
// roster). Behemoth-X-123B has a healthy 8-worker pool; deepseek-v4-flash
|
||
// is the real DeepSeek V4 family but only had 1 worker online at discovery
|
||
// time, so expect it to be the fragile one of the two.
|
||
const NEW_CAPABLE_MODELS: FreeModelSpec[] = [
|
||
{
|
||
provider: "aihorde",
|
||
model: "aihorde/aphrodite/TheDrummer/Behemoth-X-123B-v2.1",
|
||
displayName: "Behemoth-X 123B (AI Horde)",
|
||
},
|
||
{
|
||
provider: "aihorde",
|
||
model: "aihorde/deepseek/deepseek-v4-flash",
|
||
displayName: "DeepSeek V4 Flash (AI Horde)",
|
||
},
|
||
];
|
||
|
||
const CONCURRENT_THREADS = 4;
|
||
const SPOT_CHECK_CASE_COUNT = 4;
|
||
|
||
test.before(async () => {
|
||
await ensureTestEnvironment();
|
||
});
|
||
|
||
function percentile(sortedMs: number[], p: number): number {
|
||
if (sortedMs.length === 0) return 0;
|
||
const idx = Math.min(sortedMs.length - 1, Math.ceil((p / 100) * sortedMs.length) - 1);
|
||
return sortedMs[Math.max(0, idx)];
|
||
}
|
||
|
||
test(
|
||
"AI Horde: sequential reliability across all 25 workload cases (gemma-4-31b)",
|
||
{ skip },
|
||
async () => {
|
||
console.log(
|
||
`\n AI Horde sequential: ${CASE_BUILDERS.length} cases, anonymous key, gemma-4-31b\n`
|
||
);
|
||
|
||
const results: BenchmarkResult[] = [];
|
||
for (const tc of CASE_BUILDERS) {
|
||
const r = await benchmarkRequest(GEMMA_4, tc.name, tc.build, 90_000);
|
||
results.push(r);
|
||
console.log(
|
||
` [${r.ok ? "OK " : "FAIL"}] ${tc.name.padEnd(40)} HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
|
||
(r.error ? ` | ${r.error}` : "")
|
||
);
|
||
}
|
||
|
||
const succeeded = results.filter((r) => r.ok);
|
||
const durations = results.map((r) => r.durationMs).sort((a, b) => a - b);
|
||
const successRate = Math.round((succeeded.length / results.length) * 100);
|
||
|
||
console.log(
|
||
`\n Sequential summary: ${succeeded.length}/${results.length} succeeded (${successRate}%) | ` +
|
||
`p50=${percentile(durations, 50)}ms p90=${percentile(durations, 90)}ms max=${durations[durations.length - 1]}ms\n`
|
||
);
|
||
|
||
assert.ok(
|
||
succeeded.length > 0,
|
||
"every sequential request failed — likely a harness/routing bug"
|
||
);
|
||
}
|
||
);
|
||
|
||
test(
|
||
"AI Horde: concurrent load — does anonymous queueing degrade under parallel requests?",
|
||
{ skip },
|
||
async () => {
|
||
console.log(
|
||
`\n AI Horde concurrent: ${CONCURRENT_THREADS} parallel requests, anonymous key, gemma-4-31b\n`
|
||
);
|
||
|
||
const start = performance.now();
|
||
const results = await Promise.all(
|
||
Array.from({ length: CONCURRENT_THREADS }, (_, i) =>
|
||
benchmarkRequest(
|
||
GEMMA_4,
|
||
`concurrent-${i + 1}`,
|
||
() => CASE_BUILDERS[i % CASE_BUILDERS.length].build(),
|
||
120_000
|
||
)
|
||
)
|
||
);
|
||
const wallClockMs = Math.round(performance.now() - start);
|
||
|
||
for (const r of results) {
|
||
console.log(
|
||
` [${r.ok ? "OK " : "FAIL"}] ${r.case.padEnd(20)} HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
|
||
(r.error ? ` | ${r.error}` : "")
|
||
);
|
||
}
|
||
|
||
const succeeded = results.filter((r) => r.ok);
|
||
console.log(
|
||
`\n Concurrent summary: ${succeeded.length}/${results.length} succeeded | ${wallClockMs}ms wall clock ` +
|
||
`(vs sum of individual durations: ${results.reduce((s, r) => s + r.durationMs, 0)}ms — a wall clock close to ` +
|
||
`the sum means requests queued rather than ran in parallel)\n`
|
||
);
|
||
|
||
assert.ok(
|
||
succeeded.length > 0,
|
||
"every concurrent request failed — likely a harness/routing bug"
|
||
);
|
||
}
|
||
);
|
||
|
||
test(
|
||
"AI Horde: cross-model spot-check — is more than one model viable for a combo pool?",
|
||
{ skip },
|
||
async () => {
|
||
console.log(
|
||
`\n AI Horde cross-model: ${OTHER_MODELS.length} models × ${SPOT_CHECK_CASE_COUNT} cases\n`
|
||
);
|
||
|
||
for (const spec of OTHER_MODELS) {
|
||
const results: BenchmarkResult[] = [];
|
||
for (let i = 0; i < SPOT_CHECK_CASE_COUNT; i++) {
|
||
const tc = CASE_BUILDERS[i];
|
||
const r = await benchmarkRequest(spec, tc.name, tc.build, 90_000);
|
||
results.push(r);
|
||
console.log(
|
||
` [${r.ok ? "OK " : "FAIL"}] ${spec.displayName.padEnd(28)} ${tc.name.padEnd(35)} ` +
|
||
`HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
|
||
(r.error ? ` | ${r.error}` : "")
|
||
);
|
||
}
|
||
const succeeded = results.filter((r) => r.ok).length;
|
||
console.log(` --> ${spec.displayName}: ${succeeded}/${results.length} succeeded\n`);
|
||
}
|
||
}
|
||
);
|
||
|
||
// #note: a live smoke test showed aihorde/google/gemma-4-31b narrates tool
|
||
// calls in plain text ("`write_file` with `path=...`") instead of emitting a
|
||
// real tool_calls array, finish_reason "stop" — no native/emulated
|
||
// tool-calling support for this model through OmniRoute today. A genuine
|
||
// tool-calling "agentic loop" (like live-gemini-agentic-loop.test.ts) isn't
|
||
// viable here, so this asks the underlying question directly instead: can
|
||
// AI Horde sustain ~60k tokens/minute of large-context throughput, the kind
|
||
// of volume a real agentic session's growing context would generate?
|
||
const SUSTAINED_TARGET_TOKENS_PER_PROMPT = 13_000;
|
||
const SUSTAINED_ROUNDS = 5; // 5 × 13k ≈ 65k tokens sent, back-to-back
|
||
|
||
test(
|
||
"AI Horde: sustained large-context throughput — can it hit ~60k tokens/minute?",
|
||
{ skip },
|
||
async () => {
|
||
console.log(
|
||
`\n AI Horde sustained throughput: ${SUSTAINED_ROUNDS} × ~${SUSTAINED_TARGET_TOKENS_PER_PROMPT} ` +
|
||
`estimated-token prompts, back-to-back, targeting ~60k tokens/minute\n`
|
||
);
|
||
|
||
const start = performance.now();
|
||
const results = await benchmarkTpmStress(
|
||
GEMMA_4,
|
||
SUSTAINED_TARGET_TOKENS_PER_PROMPT,
|
||
SUSTAINED_ROUNDS,
|
||
120_000
|
||
);
|
||
const wallClockMs = Math.round(performance.now() - start);
|
||
|
||
for (const r of results) {
|
||
console.log(
|
||
` [${r.ok ? "OK " : "FAIL"}] ${r.case.padEnd(45)} HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
|
||
(r.error ? ` | ${r.error}` : "")
|
||
);
|
||
}
|
||
|
||
const succeeded = results.filter((r) => r.ok);
|
||
// AI Horde doesn't report usage tokens (confirmed 0 across every prior test
|
||
// in this file), so use the same ~4 chars/token estimate genHugeContextMessage
|
||
// itself is built on, applied to every ATTEMPTED prompt (not just successful
|
||
// ones) — an attempted-but-failed send still occupied the queue's time.
|
||
const estimatedTokensSent = results.length * SUSTAINED_TARGET_TOKENS_PER_PROMPT;
|
||
const achievedTpm = Math.round((estimatedTokensSent / wallClockMs) * 60_000);
|
||
|
||
console.log(
|
||
`\n Sustained summary: ${succeeded.length}/${results.length} succeeded | ${wallClockMs}ms wall clock | ` +
|
||
`~${estimatedTokensSent} tokens sent | achieved ~${achievedTpm} tokens/minute ` +
|
||
`(target: 60000)\n`
|
||
);
|
||
|
||
assert.ok(
|
||
succeeded.length > 0,
|
||
"every sustained-throughput request failed — likely a harness/routing bug"
|
||
);
|
||
}
|
||
);
|
||
|
||
test(
|
||
"AI Horde: new capable model candidates — workload test (Behemoth-X 123B, DeepSeek V4 Flash)",
|
||
{ skip },
|
||
async () => {
|
||
console.log(
|
||
`\n AI Horde new candidates: ${NEW_CAPABLE_MODELS.length} models × ${BENCHMARK_CASES.length} workload case(s) ` +
|
||
`— the same 5-case slice every other benchmarked model was run through, for a direct comparison\n`
|
||
);
|
||
|
||
const summaries: ModelBenchmarkSummary[] = [];
|
||
|
||
for (const spec of NEW_CAPABLE_MODELS) {
|
||
const results: BenchmarkResult[] = [];
|
||
for (const tc of BENCHMARK_CASES) {
|
||
const r = await benchmarkRequest(spec, tc.name, tc.build, 120_000);
|
||
results.push(r);
|
||
console.log(
|
||
` [${r.ok ? "OK " : "FAIL"}] ${spec.displayName.padEnd(30)} ${tc.name.padEnd(35)} ` +
|
||
`HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
|
||
(r.error ? ` | ${r.error}` : "")
|
||
);
|
||
}
|
||
summaries.push(summarize(spec, results));
|
||
}
|
||
|
||
console.log(formatBenchmarkTable(summaries));
|
||
|
||
const totalSucceeded = summaries.reduce((s, m) => s + m.succeeded, 0);
|
||
assert.ok(
|
||
totalSucceeded > 0,
|
||
"every request to both new candidates failed — likely a harness/routing bug"
|
||
);
|
||
}
|
||
);
|