Files
OmniRoute/tests/integration/aihorde-load-test.test.ts
Markus Hartung 6e1e5c9a45 fix(sse): tool-incapable provider handling (AI Horde + Responses content-collapse scoping) (#8212)
* fix(sse): collapse single-text-part Responses-API content to a plain string

Every /v1/responses request — even the simplest single-string input —
got 500'd by AI Horde's Aphrodite-backed facade. Root cause:
normalizeResponsesInputForChat() always wraps a plain string input as
`content: [{ type: "input_text", text: value }]` (a one-element array),
and openaiResponsesToOpenAIRequest() mapped that straight through to
`content: [{ type: "text", text: value }]` on the Chat Completions side
— an array. That's spec-valid (OpenAI's own API accepts both shapes),
but strict/naive OpenAI-compatible backends like AI Horde's only
implement the plain-string form and reject the array form outright.

A single-text-part array and a plain string are semantically
identical, so collapse is safe. Real multi-part messages (text+image,
text+file) are left untouched.

Regression test: tests/unit/openai-responses-single-text-content-string.test.ts
(RED before the fix — every collapsed-content assertion failed with an
object instead of a string; GREEN after).

Also adds a deeper AI Horde load-test suite (sequential/concurrent/
cross-model/sustained-throughput/new-capable-model-candidates) that
surfaced this bug via real live traffic after Behemoth-X-123B was
temporarily added to the "default" combo for evaluation.

Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>

* fix(sse): unsupportedParams provider-level fallback for aihorde's live-discovered models

Real OpenClaw traffic against the newly-added Behemoth-X-123B combo
target kept 500ing on every attempt even after the Responses-API
content-array fix landed. The pipeline artifact showed why: `tools`
was still present, unstripped, in the request actually sent to AI
Horde's Aphrodite backend.

Root cause: `unsupportedParams: ["tools", "tool_choice",
"parallel_tool_calls"]` was only declared on the 3 models statically
listed in the aihorde registry entry (Cydonia-24B, Skyfall-31B,
google/gemma-4-31b). AI Horde uses `passthroughModels: true` — its
live worker roster changes constantly — so Behemoth-X-123B, like every
other dynamically-discovered aihorde model, had no model-specific
unsupportedParams entry, and getUnsupportedParams() returned [] for
it. But "the workers run raw text-completion backends" (no tool
calling) is true of every model AI Horde serves, not just the 3
catalogued ones.

Adds a provider-level `unsupportedParams` fallback on RegistryEntry,
checked by getUnsupportedParams() after the per-model lookup misses.
Set on the aihorde entry so it covers its entire live-discovered
roster, present and future, without needing a static per-model catalog
entry for each one.

Regression test: tests/unit/aihorde-tools-unsupported-provider-fallback.test.ts
(RED before the fix — Behemoth-X and deepseek-v4-flash both returned
[] instead of the stripped param list; GREEN after, with a control
case confirming the fallback doesn't leak to unrelated providers).

Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>

* fix(sse): flatten leftover tool-call history when stripping unsupported tools

Third bug in the same AI Horde/Behemoth-X saga: even after tools/
tool_choice were correctly stripped from the live request (previous
fix), real combo traffic still 500'd. The conversation history itself
carried a prior turn's role:"assistant" tool_calls and role:"tool"
result messages, left over from before the combo failed over from a
tool-capable model (Gemini) to a non-tool-capable one (AI Horde). Its
raw completion backend doesn't understand those message shapes at all,
independent of whether live `tools` is present — confirmed by
reproducing with a role:"tool" message and NO tools param at all.

flattenToolHistory() (open-sse/utils/flattenToolHistory.ts) already
existed for exactly this, fully unit-tested — it just had zero call
sites anywhere in the request pipeline. Extracts the unsupported-params
strip into a small testable module
(open-sse/handlers/chatCore/unsupportedParamsStrip.ts, following the
existing chatCore god-file decomposition pattern e.g.
executorClientHeaders.ts) that now also flattens tool-call history
whenever "tools" was among the stripped params.

Regression test: tests/unit/chatcore-unsupported-params-strip.test.ts
(RED before the fix — the flattening test failed with the raw
tool_calls array still present; GREEN after). All 434 existing
chatcore-*.test.ts tests still pass.

Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>

* fix(sse): gate tool-history flattening on unsupported, not on stripped-this-request

The previous commit's flattening only fired when "tools" was actually
present-and-stripped on THIS request. A second live reproduction
against AI Horde had no live `tools` param at all — only stale
tool_calls/tool-result messages inherited from before a combo
failover — and still 500'd, because that condition never triggered.

A model that can't do tool calling can't do it whether or not the
current request happens to carry a `tools` array. Gate on the
unsupported-params list itself (unsupported.includes("tools")) instead
of the subset that was actually present-and-deleted this time.

Regression test added to the same file (RED before — the no-live-tools
case left tool_calls/role:"tool" untouched; GREEN after). All 435
chatcore-*.test.ts still pass.

Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>

* fix(sse): skip tool-incapable combo targets, error clearly on direct requests

Two complementary fixes for a model that structurally can't do tool
calling at all (e.g. AI Horde's raw completion backends) rather than
silently degrading — following up on the earlier strip/flatten fix,
which stopped the crashes but let a tool-incapable target still get
selected and return a 200 that narrates a fake tool call in prose
instead of erroring or being skipped.

1. Root cause, combo routing: getResolvedModelCapabilities()'s
   `supportsTools` resolution only checked per-model registry entries,
   synced capabilities, and static specs — none of which exist for a
   dynamically-discovered model (AI Horde's passthroughModels roster
   changes as workers come and go). It fell through to
   heuristicToolCalling(), which optimistically defaults to `true` for
   any unrecognized model (TOOL_CALLING_UNSUPPORTED_PATTERNS is empty).
   Added a provider-level fallback reusing the same unsupportedParams
   signal the request-time strip already relies on. This makes the
   EXISTING filterTargetsByRequestCompatibility (comboStructure.ts) —
   which already correctly excludes non-tool-capable targets when a
   request requires tools — actually work for these models; no combo.ts
   changes were needed, it was only ever fed bad capability data.

2. Direct/pinned requests: filterTargetsByRequestCompatibility only
   protects combo routing. A direct request naming an exact
   tool-incapable model has no other target to fail over to — added
   checkToolCallingRequiredButUnsupported (chatCore/toolCallingRequiredCheck.ts),
   gated on isCombo: false, returning a clear 400 instead of a 200 that
   silently can't do what was asked.

Regression tests (both RED before, GREEN after):
- tests/unit/model-capabilities-provider-unsupported-tools.test.ts
- tests/unit/chatcore-tool-calling-required-check.test.ts
All 463 chatcore-*/model-capabilities-*.test.ts and 31 combo
compatibility-filter tests still pass.

Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>

* fix(sse): correct handleChatCore return shape for the tool-calling-blocked error

handleChatCore's documented contract is `{ success, response, status,
error }`, not a raw Response — returning `new Response(...)` directly
(copied from a different early-return whose surrounding context turned
out not to share this function's top-level contract) produced "No
response is returned from route handler ... Expected a Response object
but received 'undefined'" and a bare 500 with an empty body, caught
immediately when verifying the previous commit live.

Uses createErrorResult() (already used by the adjacent
translation-failure branch a few lines up) instead of hand-building the
Response, matching the same pattern already established in this
function for early error returns.

Co-Authored-By: Markus Hartung <markus.hartung@gmail.com>

* fix(sse): scope Responses single-text-content collapse to providers that need it

The single-text-part content array -> plain string collapse (added for AI
Horde's Aphrodite facade, which 500s on the array form) was applied
unconditionally to every provider, silently breaking the standard OpenAI
array-shaped content contract that other providers and existing tests
depend on. Added RegistryEntry.requiresPlainStringContent, gated the
collapse on it (true only for aihorde), and threaded modelInfo.provider
through responsesHandler -> responsesApiHelper -> the translator so the
real /v1/responses call site can identify the provider.

Co-Authored-By: Markus Hartung <markus.hartream@gmail.com>

---------

Co-authored-by: Markus Hartung <markus.hartung@gmail.com>
Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
2026-07-23 05:03:53 -03:00

289 lines
10 KiB
TypeScript
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/**
* tests/integration/aihorde-load-test.test.ts
*
* Deeper load test for AI Horde (aihorde.net) — a free, crowdsourced,
* kudos-priority-queued inference network — asking whether it's viable as
* an addition/alternative to the "default" combo (currently 2 Gemini
* gemma-4 targets, see DEFAULT_COMBO_CONFIG in liveGeminiShared.ts).
* Anonymous access (key "0000000000") gets the LOWEST queue priority and
* 0 starting kudos, so this specifically probes whether that translates to
* unacceptable latency/reliability for combo-style traffic, or whether it
* holds up.
*
* Three angles, all against the anonymous no-auth `aihorde` provider:
* 1. Sequential reliability — all 25 CASE_BUILDERS, one at a time, with
* full latency distribution (p50/p90/max) on gemma-4-31b (the same
* model family the real "default" combo runs, for a fair comparison).
* 2. Concurrent load — does firing several requests at once cause queue
* pile-up/timeouts, given combo routing doesn't pace requests apart?
* 3. Cross-model spot-check — same handful of cases against the other 2
* viable aihorde models, since a combo pool needs multiple healthy
* targets, not just one.
*
* Report, not a gate: kudos-priority queueing is a genuine, permanent
* property of anonymous AI Horde access, not a bug — a model or two being
* slow is data for the "is this viable" question, not a regression.
*
* Environment:
* OMNIROUTE_API_KEY — required (else test skips)
* OMNIROUTE_URL — defaults to http://localhost:3000
*/
import test from "node:test";
import assert from "node:assert/strict";
import { skip, CASE_BUILDERS, ensureTestEnvironment } from "./liveGeminiShared.ts";
import {
benchmarkRequest,
benchmarkTpmStress,
summarize,
formatBenchmarkTable,
BENCHMARK_CASES,
type FreeModelSpec,
type BenchmarkResult,
type ModelBenchmarkSummary,
} from "./freeModelBenchmarkShared.ts";
const GEMMA_4 = {
provider: "aihorde",
model: "aihorde/google/gemma-4-31b",
displayName: "Gemma 4 31B (AI Horde)",
};
const OTHER_MODELS: FreeModelSpec[] = [
{
provider: "aihorde",
model: "aihorde/aphrodite/TheDrummer/Cydonia-24B-v4.3",
displayName: "Cydonia 24B (AI Horde)",
},
{
provider: "aihorde",
model: "aihorde/aphrodite/TheDrummer/Skyfall-31B-v4.2",
displayName: "Skyfall 31B (AI Horde)",
},
];
// Bigger/more-capable candidates found via AI Horde's live status API
// (https://aihorde.net/api/v2/status/models?type=text — the 3 models above
// are just a static passthrough-discovery fallback, not the full live
// roster). Behemoth-X-123B has a healthy 8-worker pool; deepseek-v4-flash
// is the real DeepSeek V4 family but only had 1 worker online at discovery
// time, so expect it to be the fragile one of the two.
const NEW_CAPABLE_MODELS: FreeModelSpec[] = [
{
provider: "aihorde",
model: "aihorde/aphrodite/TheDrummer/Behemoth-X-123B-v2.1",
displayName: "Behemoth-X 123B (AI Horde)",
},
{
provider: "aihorde",
model: "aihorde/deepseek/deepseek-v4-flash",
displayName: "DeepSeek V4 Flash (AI Horde)",
},
];
const CONCURRENT_THREADS = 4;
const SPOT_CHECK_CASE_COUNT = 4;
test.before(async () => {
await ensureTestEnvironment();
});
function percentile(sortedMs: number[], p: number): number {
if (sortedMs.length === 0) return 0;
const idx = Math.min(sortedMs.length - 1, Math.ceil((p / 100) * sortedMs.length) - 1);
return sortedMs[Math.max(0, idx)];
}
test(
"AI Horde: sequential reliability across all 25 workload cases (gemma-4-31b)",
{ skip },
async () => {
console.log(
`\n AI Horde sequential: ${CASE_BUILDERS.length} cases, anonymous key, gemma-4-31b\n`
);
const results: BenchmarkResult[] = [];
for (const tc of CASE_BUILDERS) {
const r = await benchmarkRequest(GEMMA_4, tc.name, tc.build, 90_000);
results.push(r);
console.log(
` [${r.ok ? "OK " : "FAIL"}] ${tc.name.padEnd(40)} HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
(r.error ? ` | ${r.error}` : "")
);
}
const succeeded = results.filter((r) => r.ok);
const durations = results.map((r) => r.durationMs).sort((a, b) => a - b);
const successRate = Math.round((succeeded.length / results.length) * 100);
console.log(
`\n Sequential summary: ${succeeded.length}/${results.length} succeeded (${successRate}%) | ` +
`p50=${percentile(durations, 50)}ms p90=${percentile(durations, 90)}ms max=${durations[durations.length - 1]}ms\n`
);
assert.ok(
succeeded.length > 0,
"every sequential request failed — likely a harness/routing bug"
);
}
);
test(
"AI Horde: concurrent load — does anonymous queueing degrade under parallel requests?",
{ skip },
async () => {
console.log(
`\n AI Horde concurrent: ${CONCURRENT_THREADS} parallel requests, anonymous key, gemma-4-31b\n`
);
const start = performance.now();
const results = await Promise.all(
Array.from({ length: CONCURRENT_THREADS }, (_, i) =>
benchmarkRequest(
GEMMA_4,
`concurrent-${i + 1}`,
() => CASE_BUILDERS[i % CASE_BUILDERS.length].build(),
120_000
)
)
);
const wallClockMs = Math.round(performance.now() - start);
for (const r of results) {
console.log(
` [${r.ok ? "OK " : "FAIL"}] ${r.case.padEnd(20)} HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
(r.error ? ` | ${r.error}` : "")
);
}
const succeeded = results.filter((r) => r.ok);
console.log(
`\n Concurrent summary: ${succeeded.length}/${results.length} succeeded | ${wallClockMs}ms wall clock ` +
`(vs sum of individual durations: ${results.reduce((s, r) => s + r.durationMs, 0)}ms — a wall clock close to ` +
`the sum means requests queued rather than ran in parallel)\n`
);
assert.ok(
succeeded.length > 0,
"every concurrent request failed — likely a harness/routing bug"
);
}
);
test(
"AI Horde: cross-model spot-check — is more than one model viable for a combo pool?",
{ skip },
async () => {
console.log(
`\n AI Horde cross-model: ${OTHER_MODELS.length} models × ${SPOT_CHECK_CASE_COUNT} cases\n`
);
for (const spec of OTHER_MODELS) {
const results: BenchmarkResult[] = [];
for (let i = 0; i < SPOT_CHECK_CASE_COUNT; i++) {
const tc = CASE_BUILDERS[i];
const r = await benchmarkRequest(spec, tc.name, tc.build, 90_000);
results.push(r);
console.log(
` [${r.ok ? "OK " : "FAIL"}] ${spec.displayName.padEnd(28)} ${tc.name.padEnd(35)} ` +
`HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
(r.error ? ` | ${r.error}` : "")
);
}
const succeeded = results.filter((r) => r.ok).length;
console.log(` --> ${spec.displayName}: ${succeeded}/${results.length} succeeded\n`);
}
}
);
// #note: a live smoke test showed aihorde/google/gemma-4-31b narrates tool
// calls in plain text ("`write_file` with `path=...`") instead of emitting a
// real tool_calls array, finish_reason "stop" — no native/emulated
// tool-calling support for this model through OmniRoute today. A genuine
// tool-calling "agentic loop" (like live-gemini-agentic-loop.test.ts) isn't
// viable here, so this asks the underlying question directly instead: can
// AI Horde sustain ~60k tokens/minute of large-context throughput, the kind
// of volume a real agentic session's growing context would generate?
const SUSTAINED_TARGET_TOKENS_PER_PROMPT = 13_000;
const SUSTAINED_ROUNDS = 5; // 5 × 13k ≈ 65k tokens sent, back-to-back
test(
"AI Horde: sustained large-context throughput — can it hit ~60k tokens/minute?",
{ skip },
async () => {
console.log(
`\n AI Horde sustained throughput: ${SUSTAINED_ROUNDS} × ~${SUSTAINED_TARGET_TOKENS_PER_PROMPT} ` +
`estimated-token prompts, back-to-back, targeting ~60k tokens/minute\n`
);
const start = performance.now();
const results = await benchmarkTpmStress(
GEMMA_4,
SUSTAINED_TARGET_TOKENS_PER_PROMPT,
SUSTAINED_ROUNDS,
120_000
);
const wallClockMs = Math.round(performance.now() - start);
for (const r of results) {
console.log(
` [${r.ok ? "OK " : "FAIL"}] ${r.case.padEnd(45)} HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
(r.error ? ` | ${r.error}` : "")
);
}
const succeeded = results.filter((r) => r.ok);
// AI Horde doesn't report usage tokens (confirmed 0 across every prior test
// in this file), so use the same ~4 chars/token estimate genHugeContextMessage
// itself is built on, applied to every ATTEMPTED prompt (not just successful
// ones) — an attempted-but-failed send still occupied the queue's time.
const estimatedTokensSent = results.length * SUSTAINED_TARGET_TOKENS_PER_PROMPT;
const achievedTpm = Math.round((estimatedTokensSent / wallClockMs) * 60_000);
console.log(
`\n Sustained summary: ${succeeded.length}/${results.length} succeeded | ${wallClockMs}ms wall clock | ` +
`~${estimatedTokensSent} tokens sent | achieved ~${achievedTpm} tokens/minute ` +
`(target: 60000)\n`
);
assert.ok(
succeeded.length > 0,
"every sustained-throughput request failed — likely a harness/routing bug"
);
}
);
test(
"AI Horde: new capable model candidates — workload test (Behemoth-X 123B, DeepSeek V4 Flash)",
{ skip },
async () => {
console.log(
`\n AI Horde new candidates: ${NEW_CAPABLE_MODELS.length} models × ${BENCHMARK_CASES.length} workload case(s) ` +
`— the same 5-case slice every other benchmarked model was run through, for a direct comparison\n`
);
const summaries: ModelBenchmarkSummary[] = [];
for (const spec of NEW_CAPABLE_MODELS) {
const results: BenchmarkResult[] = [];
for (const tc of BENCHMARK_CASES) {
const r = await benchmarkRequest(spec, tc.name, tc.build, 120_000);
results.push(r);
console.log(
` [${r.ok ? "OK " : "FAIL"}] ${spec.displayName.padEnd(30)} ${tc.name.padEnd(35)} ` +
`HTTP ${r.status} | ${r.durationMs}ms | ${r.tokens} tok` +
(r.error ? ` | ${r.error}` : "")
);
}
summaries.push(summarize(spec, results));
}
console.log(formatBenchmarkTable(summaries));
const totalSucceeded = summaries.reduce((s, m) => s + m.succeeded, 0);
assert.ok(
totalSucceeded > 0,
"every request to both new candidates failed — likely a harness/routing bug"
);
}
);