Files
OmniRoute/open-sse/handlers/chatCore/semanticCacheStore.ts
Praveen K Palaniswamy 65e81158ab fix(ollama): route models by advertised capability (#11088)
Landed with the design call resolved per the owner's pick — **option 1**: the synced store is now endpoint-agnostic (persistDiscoveredModels and managedModelImport no longer drop non-chat models at write time), and chat selectability moved to read time (auto-pool expansion in autoStrategy applies filterChatSelectableModels; the models-route projection already had its chatOnly filter). Your discovery test now passes end-to-end (3/3): /api/show capabilities persist per connection and image/embedding requests route through the advertising host.

Reconciliation notes: conflicted areas merged onto the current tip (adobe discovery import, requestedModel preflight signature, resolvedProvider fast-path coexists with the synced-route override — explicit resolution wins); carried base-red drains (#10055 memoization, #11071 test variants) dropped as already-landed; the managed-model-import exclusion test was propagated to the new contract (image/video models persist; the read filter still hides them from chat pickers — pinned by a new assertion). Full battery: 205/206 focused (the one red is a confirmed periodic-timer timing flake on the loaded devbox — 20/20 isolated), autoCombo vitest 30/30, combo suites 46/46, gates + typecheck clean.

Thank you @yourspraveen — the capability probe + routing design was right; it just needed the store contract opened up. Fixes #11087.
2026-08-23 11:45:01 -03:00

74 lines
2.6 KiB
TypeScript

/**
* chatCore semantic-cache store (Quality Gate v2 / Fase 9 — chatCore god-file decomposition,
* #3501).
*
* Extracted from handleChatCore's non-streaming success path (Phase 9.1): when semantic caching is
* enabled and the request/response are cacheable, store the translated response under its signature
* so a later temp=0 request can be served from cache. Side-effect only (cache write + debug log);
* no early-return, no outer-variable reassignment. Behaviour is byte-identical to the previous
* inline block, including the `prompt + completion || 0` token-saved precedence.
*/
import {
generateSignature as defaultGenerateSignature,
setCachedResponse as defaultSetCachedResponse,
isCacheableForWrite as defaultIsCacheableForWrite,
} from "@/lib/semanticCache";
import { isSmallEnoughForSemanticCache as defaultIsSmallEnough } from "../../utils/estimateSize.ts";
type LoggerLike = { debug?: (...args: unknown[]) => void } | null | undefined;
type CacheBody = {
messages?: unknown;
input?: unknown;
temperature?: number;
top_p?: number;
};
type UsageLike = { prompt_tokens?: number; completion_tokens?: number } | null | undefined;
export interface SemanticCacheStoreDeps {
isCacheableForWrite: typeof defaultIsCacheableForWrite;
isSmallEnoughForSemanticCache: typeof defaultIsSmallEnough;
generateSignature: typeof defaultGenerateSignature;
setCachedResponse: typeof defaultSetCachedResponse;
}
const DEFAULT_DEPS: SemanticCacheStoreDeps = {
isCacheableForWrite: defaultIsCacheableForWrite,
isSmallEnoughForSemanticCache: defaultIsSmallEnough,
generateSignature: defaultGenerateSignature,
setCachedResponse: defaultSetCachedResponse,
};
export function storeSemanticCacheResponse(
args: {
enabled: boolean;
body: CacheBody;
headers: unknown;
translatedResponse: unknown;
model: string;
apiKeyId?: string;
usage?: UsageLike;
log?: LoggerLike;
},
deps: SemanticCacheStoreDeps = DEFAULT_DEPS
): void {
if (
!args.enabled ||
!deps.isCacheableForWrite(args.body, args.headers) ||
!deps.isSmallEnoughForSemanticCache(args.translatedResponse)
) {
return;
}
const signature = deps.generateSignature(
args.model,
args.body.messages ?? args.body.input,
args.body.temperature,
args.body.top_p,
args.apiKeyId ?? undefined
);
const tokensSaved = args.usage?.prompt_tokens + args.usage?.completion_tokens || 0;
deps.setCachedResponse(signature, args.model, args.translatedResponse, tokensSaved);
args.log?.debug?.("CACHE", `Stored response for ${args.model} (${tokensSaved} tokens)`);
}