Files
OmniRoute/tests/unit/resolve-model-alias-index-8697.test.ts
Dizzle e50f2329dc fix(lib): memoize catalog pricing/capability lookups to fix cold /v1/models freeze (#8697) (#8987)
Root cause: a cold GET /v1/models catalog rebuild froze the entire server 41-54s.
node --prof profiling found a systemic missing-memoization pattern — a per-model
function rescanning a static or synced data structure with Object.entries()/
Object.keys() (or hitting SQLite) on every call instead of once per rebuild. Fixed
6 instances of the same pattern, found by iteratively re-profiling the full catalog
sweep after each fix (plus a whitebox review pass) until no further hotspot of this
shape remained:

1. getModelsDevPricing() (modelsDevSync.ts) — re-ran a synchronous SQLite query and
   re-JSON.parse'd ~180 blobs on every call (up to ~6091x instead of once per
   request). Memoized via the existing modelCatalogCacheVersion invalidation signal
   (same pattern as getCachedRawProviderConnections/getCachedProviderNodes in
   db/readCache.ts). Dominant cost of the original 41-54s freeze.

2. findInsensitive() (modelMetadataRegistry.ts, resolveCatalogPricing) — rebuilt a
   full Object.entries() scan on every case-insensitive lookup miss, twice per
   model. Replaced with a lowercase-key index built once per distinct pricing
   object and cached by identity (WeakMap). Warns once at index-build time on a
   case-insensitive key collision instead of silently discarding the second value.

3. getSyncedCapability() (modelsDevSync.ts) — ran a per-model SQLite SELECT on cold
   cache instead of self-warming the whole-table cache; no caller in the
   /v1/models build path ever primed it, so a cold rebuild ran one SQLite
   round-trip per model per call site. Now self-warms via the existing bulk
   getSyncedCapabilities() on first miss. Measured as the dominant remaining cost
   after fixes 1-2 (~70% of a full catalog sweep).

4. getCanonicalModelSpecId() (shared/constants/modelSpecs.ts) — up to 3 separate
   linear scans over the static MODEL_SPECS table per call (exact ci, alias ci,
   prefix). Replaced with a lazy, lowercase-key index built once (MODEL_SPECS never
   changes at runtime); prefix-match iteration order preserved exactly so
   resolution outcomes are unchanged.

5. getStaticSpecCanonicalModelId() (modelCapabilities.ts) — duplicated the same
   exact+alias scan as (4) in a second, separate rescan. Now reuses the shared
   index via a new exported helper (findModelSpecIdByExactOrAlias) instead of
   maintaining a second cache over the same static table.
   reverseModelsDevProviders() (modelCapabilities.ts) — rescanned
   Object.entries(MODELS_DEV_PROVIDER_MAP) (also static) on every call; memoized
   by provider key. Result is frozen (readonly) since it is now shared across
   calls instead of freshly allocated each time.

6. resolveModelAlias() (shared/constants/modelSpecs.ts) — rescanned
   Object.entries(MODEL_SPECS) unconditionally once per model (verified 1:1 call
   ratio, no short-circuit). Case-sensitive exact match (Array.includes(), no
   .toLowerCase()) — uses a dedicated exact-match index, deliberately not the
   case-insensitive alias index from fix 4/5 (would silently broaden matches).

Measured on a 1940-pair real-catalog sample (static PROVIDER_MODELS registry):
cold sweep 828ms -> 356ms after fixes 3-5 on top of 1-2, extrapolating to roughly
1s on the real ~6091-model catalog, down from the original 41-54s freeze.

Complementary to the stale-serve fix in #8801 (upstream) — neither alone
eliminates the freeze.

Tests: call-count regression guards for every fix (DB prepare / Object.entries /
Object.keys call counts staying constant instead of scaling with iteration count),
plus correctness coverage for case-insensitive/case-sensitive resolution. All
pre-existing consumer suites re-verified passing (96 tests total across 19 files).

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
2026-08-04 14:34:10 -03:00

52 lines
1.8 KiB
TypeScript

import assert from "node:assert/strict";
import { describe, it } from "node:test";
import { resolveModelAlias } from "../../src/shared/constants/modelSpecs.ts";
describe("resolveModelAlias lookup index (#8697-adjacent)", () => {
it("still resolves a known exact alias", () => {
// Real MODEL_SPECS alias, case-sensitive exact match.
assert.equal(resolveModelAlias("openai/gpt-5.6"), "gpt-5.6");
});
it("does not match a case-varied alias (case-sensitive semantics preserved)", () => {
// resolveModelAlias uses Array.includes(), never .toLowerCase() — a case-varied
// input must NOT resolve, unlike the case-insensitive getCanonicalModelSpecId().
assert.equal(resolveModelAlias("OpenAI/GPT-5.6"), "OpenAI/GPT-5.6");
});
it("returns the input unchanged for an unknown alias", () => {
assert.equal(
resolveModelAlias("definitely-not-a-real-alias-xyz"),
"definitely-not-a-real-alias-xyz"
);
});
it("does not rescan MODEL_SPECS per call (regression guard for O(n) scans)", () => {
// Warm up outside the measured window.
resolveModelAlias("openai/gpt-5.6");
const originalEntries = Object.entries;
let calls = 0;
Object.entries = function patchedEntries(...args: Parameters<typeof Object.entries>) {
calls++;
return originalEntries.apply(this, args as never);
} as typeof Object.entries;
try {
for (let i = 0; i < 500; i++) {
resolveModelAlias("openai/gpt-5.6");
}
} finally {
Object.entries = originalEntries;
}
// Pre-fix: every call re-ran Object.entries(MODEL_SPECS). Indexed: the lazy
// index is built once and reused, so no further Object.entries calls happen.
assert.equal(
calls,
0,
`expected 0 Object.entries() calls across 500 repeated lookups, got ${calls}`
);
});
});