mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-14 11:12:17 +03:00
Root cause: a cold GET /v1/models catalog rebuild froze the entire server 41-54s. node --prof profiling found a systemic missing-memoization pattern — a per-model function rescanning a static or synced data structure with Object.entries()/ Object.keys() (or hitting SQLite) on every call instead of once per rebuild. Fixed 6 instances of the same pattern, found by iteratively re-profiling the full catalog sweep after each fix (plus a whitebox review pass) until no further hotspot of this shape remained: 1. getModelsDevPricing() (modelsDevSync.ts) — re-ran a synchronous SQLite query and re-JSON.parse'd ~180 blobs on every call (up to ~6091x instead of once per request). Memoized via the existing modelCatalogCacheVersion invalidation signal (same pattern as getCachedRawProviderConnections/getCachedProviderNodes in db/readCache.ts). Dominant cost of the original 41-54s freeze. 2. findInsensitive() (modelMetadataRegistry.ts, resolveCatalogPricing) — rebuilt a full Object.entries() scan on every case-insensitive lookup miss, twice per model. Replaced with a lowercase-key index built once per distinct pricing object and cached by identity (WeakMap). Warns once at index-build time on a case-insensitive key collision instead of silently discarding the second value. 3. getSyncedCapability() (modelsDevSync.ts) — ran a per-model SQLite SELECT on cold cache instead of self-warming the whole-table cache; no caller in the /v1/models build path ever primed it, so a cold rebuild ran one SQLite round-trip per model per call site. Now self-warms via the existing bulk getSyncedCapabilities() on first miss. Measured as the dominant remaining cost after fixes 1-2 (~70% of a full catalog sweep). 4. getCanonicalModelSpecId() (shared/constants/modelSpecs.ts) — up to 3 separate linear scans over the static MODEL_SPECS table per call (exact ci, alias ci, prefix). Replaced with a lazy, lowercase-key index built once (MODEL_SPECS never changes at runtime); prefix-match iteration order preserved exactly so resolution outcomes are unchanged. 5. getStaticSpecCanonicalModelId() (modelCapabilities.ts) — duplicated the same exact+alias scan as (4) in a second, separate rescan. Now reuses the shared index via a new exported helper (findModelSpecIdByExactOrAlias) instead of maintaining a second cache over the same static table. reverseModelsDevProviders() (modelCapabilities.ts) — rescanned Object.entries(MODELS_DEV_PROVIDER_MAP) (also static) on every call; memoized by provider key. Result is frozen (readonly) since it is now shared across calls instead of freshly allocated each time. 6. resolveModelAlias() (shared/constants/modelSpecs.ts) — rescanned Object.entries(MODEL_SPECS) unconditionally once per model (verified 1:1 call ratio, no short-circuit). Case-sensitive exact match (Array.includes(), no .toLowerCase()) — uses a dedicated exact-match index, deliberately not the case-insensitive alias index from fix 4/5 (would silently broaden matches). Measured on a 1940-pair real-catalog sample (static PROVIDER_MODELS registry): cold sweep 828ms -> 356ms after fixes 3-5 on top of 1-2, extrapolating to roughly 1s on the real ~6091-model catalog, down from the original 41-54s freeze. Complementary to the stale-serve fix in #8801 (upstream) — neither alone eliminates the freeze. Tests: call-count regression guards for every fix (DB prepare / Object.entries / Object.keys call counts staying constant instead of scaling with iteration count), plus correctness coverage for case-insensitive/case-sensitive resolution. All pre-existing consumer suites re-verified passing (96 tests total across 19 files). Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
33 lines
1.4 KiB
TypeScript
33 lines
1.4 KiB
TypeScript
import assert from "node:assert/strict";
|
|
import { describe, it, mock } from "node:test";
|
|
import { getDbInstance } from "../../src/lib/db/core.ts";
|
|
import { getSyncedCapability } from "../../src/lib/modelsDevSync.ts";
|
|
|
|
describe("getSyncedCapability warm-up (#8697-adjacent)", () => {
|
|
it("does not run a DB round-trip per distinct model lookup (regression guard for the missing bulk warm-up)", () => {
|
|
const db = getDbInstance();
|
|
const prepareSpy = mock.method(db, "prepare");
|
|
const callsBefore = prepareSpy.mock.calls.length;
|
|
|
|
// A catalog rebuild calls getSyncedCapability() once per distinct model — this
|
|
// used to run one SQLite SELECT per call on a cold cache (no warm-up caller sits
|
|
// in the /v1/models build path). Self-warmed, only the one-time bulk load (plus
|
|
// its CREATE TABLE IF NOT EXISTS guard) should touch the DB, regardless of how
|
|
// many distinct models are looked up afterward.
|
|
const N = 200;
|
|
for (let i = 0; i < N; i++) {
|
|
getSyncedCapability("openai", `synthetic-model-${i}`);
|
|
}
|
|
|
|
const callsAfter = prepareSpy.mock.calls.length;
|
|
prepareSpy.mock.restore();
|
|
|
|
assert.ok(
|
|
callsAfter - callsBefore <= 2,
|
|
`expected at most 2 db.prepare() calls (bulk load + table guard) across ${N} distinct ` +
|
|
`model lookups, got ${callsAfter - callsBefore} — getSyncedCapability() may have regressed ` +
|
|
`to a per-model SQLite round-trip`
|
|
);
|
|
});
|
|
});
|