mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-07 07:42:13 +03:00
Root cause: a cold GET /v1/models catalog rebuild froze the entire server 41-54s. node --prof profiling found a systemic missing-memoization pattern — a per-model function rescanning a static or synced data structure with Object.entries()/ Object.keys() (or hitting SQLite) on every call instead of once per rebuild. Fixed 6 instances of the same pattern, found by iteratively re-profiling the full catalog sweep after each fix (plus a whitebox review pass) until no further hotspot of this shape remained: 1. getModelsDevPricing() (modelsDevSync.ts) — re-ran a synchronous SQLite query and re-JSON.parse'd ~180 blobs on every call (up to ~6091x instead of once per request). Memoized via the existing modelCatalogCacheVersion invalidation signal (same pattern as getCachedRawProviderConnections/getCachedProviderNodes in db/readCache.ts). Dominant cost of the original 41-54s freeze. 2. findInsensitive() (modelMetadataRegistry.ts, resolveCatalogPricing) — rebuilt a full Object.entries() scan on every case-insensitive lookup miss, twice per model. Replaced with a lowercase-key index built once per distinct pricing object and cached by identity (WeakMap). Warns once at index-build time on a case-insensitive key collision instead of silently discarding the second value. 3. getSyncedCapability() (modelsDevSync.ts) — ran a per-model SQLite SELECT on cold cache instead of self-warming the whole-table cache; no caller in the /v1/models build path ever primed it, so a cold rebuild ran one SQLite round-trip per model per call site. Now self-warms via the existing bulk getSyncedCapabilities() on first miss. Measured as the dominant remaining cost after fixes 1-2 (~70% of a full catalog sweep). 4. getCanonicalModelSpecId() (shared/constants/modelSpecs.ts) — up to 3 separate linear scans over the static MODEL_SPECS table per call (exact ci, alias ci, prefix). Replaced with a lazy, lowercase-key index built once (MODEL_SPECS never changes at runtime); prefix-match iteration order preserved exactly so resolution outcomes are unchanged. 5. getStaticSpecCanonicalModelId() (modelCapabilities.ts) — duplicated the same exact+alias scan as (4) in a second, separate rescan. Now reuses the shared index via a new exported helper (findModelSpecIdByExactOrAlias) instead of maintaining a second cache over the same static table. reverseModelsDevProviders() (modelCapabilities.ts) — rescanned Object.entries(MODELS_DEV_PROVIDER_MAP) (also static) on every call; memoized by provider key. Result is frozen (readonly) since it is now shared across calls instead of freshly allocated each time. 6. resolveModelAlias() (shared/constants/modelSpecs.ts) — rescanned Object.entries(MODEL_SPECS) unconditionally once per model (verified 1:1 call ratio, no short-circuit). Case-sensitive exact match (Array.includes(), no .toLowerCase()) — uses a dedicated exact-match index, deliberately not the case-insensitive alias index from fix 4/5 (would silently broaden matches). Measured on a 1940-pair real-catalog sample (static PROVIDER_MODELS registry): cold sweep 828ms -> 356ms after fixes 3-5 on top of 1-2, extrapolating to roughly 1s on the real ~6091-model catalog, down from the original 41-54s freeze. Complementary to the stale-serve fix in #8801 (upstream) — neither alone eliminates the freeze. Tests: call-count regression guards for every fix (DB prepare / Object.entries / Object.keys call counts staying constant instead of scaling with iteration count), plus correctness coverage for case-insensitive/case-sensitive resolution. All pre-existing consumer suites re-verified passing (96 tests total across 19 files). Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
60 lines
1.8 KiB
TypeScript
60 lines
1.8 KiB
TypeScript
import assert from "node:assert/strict";
|
|
import { describe, it, before, after, mock } from "node:test";
|
|
import { getDbInstance } from "../../src/lib/db/core.ts";
|
|
import {
|
|
getModelsDevPricing,
|
|
saveModelsDevPricing,
|
|
clearModelsDevPricing,
|
|
type PricingByProvider,
|
|
} from "../../src/lib/modelsDevSync.ts";
|
|
|
|
describe("getModelsDevPricing memoization (#8697)", () => {
|
|
before(() => {
|
|
const pricing: PricingByProvider = {
|
|
openai: {
|
|
"gpt-4o": { input: 2.5, output: 10 },
|
|
},
|
|
};
|
|
saveModelsDevPricing(pricing);
|
|
});
|
|
|
|
after(() => {
|
|
try {
|
|
clearModelsDevPricing();
|
|
} catch {
|
|
// ignore
|
|
}
|
|
});
|
|
|
|
it("hits the DB once for repeated reads within the same cache version", () => {
|
|
const db = getDbInstance();
|
|
const prepareSpy = mock.method(db, "prepare");
|
|
const callsBefore = prepareSpy.mock.calls.length;
|
|
|
|
getModelsDevPricing();
|
|
getModelsDevPricing();
|
|
getModelsDevPricing();
|
|
|
|
const callsAfter = prepareSpy.mock.calls.length;
|
|
prepareSpy.mock.restore();
|
|
|
|
// The N+1 bug re-runs the SELECT + JSON.parse on every call — memoized,
|
|
// 3 calls should cost at most 1 real DB round-trip (0 if a prior test
|
|
// already warmed the cache at the same version).
|
|
assert.ok(
|
|
callsAfter - callsBefore <= 1,
|
|
`expected at most 1 db.prepare() call across 3 reads, got ${callsAfter - callsBefore}`
|
|
);
|
|
});
|
|
|
|
it("returns fresh data after a write invalidates the cache", () => {
|
|
getModelsDevPricing(); // warm the cache
|
|
saveModelsDevPricing({
|
|
anthropic: { "claude-x": { input: 1, output: 2 } },
|
|
});
|
|
const pricing = getModelsDevPricing();
|
|
assert.ok(pricing.anthropic, "cache should reflect the write, not a stale snapshot");
|
|
assert.equal(pricing.anthropic["claude-x"].input, 1);
|
|
});
|
|
});
|