mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-02 21:32:10 +03:00
* chore(release): open v3.8.35 development cycle
* fix db vacuum scheduler settings (#4726)
Scheduled VACUUM now follows Storage page settings (scheduledVacuum/vacuumHour) as single source of truth; env-flag control path removed. 11/11 vacuum-scheduler tests pass against release/v3.8.35 tip; no orphaned env refs. Integrated into release/v3.8.35.
* fix(tier): noAuth providers count as free; free filter returns empty … (#4753)
noAuth providers now classified free (union of legacy list + NOAUTH_PROVIDERS chat-tier derivation), -free arena_elo alias, and auto/<cat>:free returns an empty pool when no free candidate matches (opt-in legacy fallback via OMNIROUTE_AUTO_FREE_FALLBACK_TO_FULL_POOL). New env var documented in .env.example + ENVIRONMENT.md; CHANGELOG bullet added (maintainer co-author). 46/46 node + 56/56 vitest tests pass on release tip; env-doc-sync, docs-sync, typecheck:core, lint, file-size all green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai 11 helpers de nível superior para 6 leaves puros (#3501) (#4571)
chatCore god-file decomposition (#3501): extract 6 pure leaves (cacheUsageMeta, executorClientHeaders, nonStreamingResponseBody, skillsFormat, streamErrorResult, streamFinalize) from chatCore.ts. Rebased onto release/v3.8.35 tip (resolved single chatCore.ts conflict — removed now-extracted inline buildExecutorClientHeaders). 265/265 chatcore tests, 26/26 new leaf tests, typecheck:core, cycles, file-size all green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai resolveExecutorWithProxy + getExecutionCredentials para leaves (#3501) (#4646)
chatCore #3501: extract resolveExecutorWithProxy + getExecutionCredentials to leaves (executorProxy.ts, executionCredentials.ts). Clean cherry-pick onto release tip post-#4571. 12/12 new leaf tests, typecheck:core, cycles, file-size green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai transforms de mensagens Claude p/ leaf (#3501) (#4708)
chatCore #3501: extract Claude upstream-message transforms to leaf (claudeUpstreamMessages.ts + claudeMessageTypes.ts). Clean cherry-pick post-#4646. 8/8 new leaf tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai persistAttemptLogs para leaf (#3501) (#4717)
chatCore #3501: extract persistAttemptLogs to leaf (attemptLogging.ts). Rebased onto release tip post-#4708 (resolved imports conflict: kept tip's resolveCompressionHeader from compression Phase 3, dropped now-unused logTruncation import moved into the leaf). 288/288 chatcore tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai stageTrace + compressionUsageReceipt para leaves (#3501) (#4721)
chatCore #3501: extract stageTrace + compressionUsageReceipt to leaves. Clean cherry-pick post-#4717. 6/6 new leaf tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai prepareUpstreamBody (1ª sub-fatia do executeProviderRequest, #3501) (#4730)
chatCore #3501: extract prepareUpstreamBody (first sub-slice of executeProviderRequest) to leaf (upstreamBody.ts). Clean cherry-pick post-#4721. 7/7 new leaf tests, full 301/301 chatcore suite, typecheck/cycles/file-size green. Completes the 6-PR chatCore decomposition stack into release/v3.8.35.
* fix(db): make db-backup import size cap configurable (#4719) (#4757)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* chore(quality): expand check:release-green to the FULL release-PR gate set (#4758)
The release-green pre-flight (Solution C) previously covered only a subset of the
gates that run exclusively on the release PR (PR→main), so reds still accrued
silently on release/** and surfaced in ~40-min layers at release time (v3.8.34:
3 CI rounds — CodeQL sanitization, then the fail-fast Quality Ratchet revealing
openapi then cyclomatic-complexity one push at a time, plus zizmor/integration).
Now check:release-green reproduces the COMPLETE release-PR gate set and reports
EVERY red in one pass (collected, not fail-fast):
- New DRIFT ratchets (report-only, rebaselined at release, never block):
cyclomatic complexity, dead-code, type-coverage, compression-budget,
openapi-coverage, workflow-lint (zizmor), codeql-ratchet.
- New HARD gates (real defects): docs-all (fabricated-docs strict + i18n mirror
sync) and the integration test suite (gated behind !--quick).
The only release-PR gates it still cannot reproduce locally are GitHub-side CodeQL
semantic analysis and SonarQube/SonarCloud (external services).
The nightly-release-green workflow and /green-prs inherit the expanded coverage
automatically (they invoke this script), so cycle drift is now surfaced
continuously and the release PR is green on its first CI run.
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* fix(dashboard): add missing onboarding.tiers step title (#4698) (#4755)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* feat(compression): Output Styles registry + D0 telemetry (Phase 4A) (#4694)
Phase 4A: Output Styles registry + D0 telemetry. Integrated into release/v3.8.35.
* feat(compression): SLM tier for ultra (Phase 4B) [stacked on #4694] (#4707)
Phase 4B: SLM tier for ultra. Integrated into release/v3.8.35.
* feat(compression): context-budget adaptive compression (Phase 4C) [stacked on #4707] (#4716)
Phase 4C: adaptive context-budget compression. Integrated into release/v3.8.35.
* feat(compression): offline evaluation harness (Phase 4 D1) [stacked on #4716] (#4720)
Phase 4 D1: offline evaluation harness. Integrated into release/v3.8.35.
* fix(sse): deepseek-web folds role:tool results into prompt transcript (#4712) (#4756)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* fix(dashboard): remove dead unconditional useLiveRequests call in HomePageClient (#4759, #4745, #4596) (#4761)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* fix(dashboard): dedupe provider nodes by id on compatible-provider add (#4746) (#4768)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* chore(db): re-export compressionRunTelemetry from localDb to satisfy db-rules (#4775)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* docs(security): add canonical STRIDE-based threat model (#4783)
Canonical STRIDE threat model. Integrated into release/v3.8.35.
* test(dashboard): add smoke test for home client dashboard (#4793)
Smoke test guarding the dashboard home client render (regression #4745/#4759). Code fix already landed via #4761; this PR's jsdom smoke test is the net-new regression guard. Integrated into release/v3.8.35.
* fix(combos): auto-promote zeroLatencyOptimizationsEnabled so legacy configs (pre-3.8.33 fallbackCompressionMode="lite") round-trip on the first GUI edit (#4774)
Auto-promote zeroLatencyOptimizationsEnabled + strip v3.8.31-era removed keys so legacy combo configs round-trip through PUT /api/combos/{id} on first GUI edit (closes #4382 followup). Pre-merge: rewrote the now-stale reject test to assert auto-promotion + added passthrough/round-trip regression guards; reconciled combos/page.tsx file-size baseline. Integrated into release/v3.8.35.
* refactor(chatCore): extrai parse + usage-stats não-streaming do executeProviderRequest (#3501) (#4762)
chatCore #3501: extract parseNonStreamingResponseBody + recordNonStreamingUsageStats. Integrated into release/v3.8.35.
* refactor(chatCore): extrai recordContextEditingTelemetryHook (#3501) (#4779)
chatCore #3501: extract recordContextEditingTelemetryHook. Integrated into release/v3.8.35.
* refactor(chatCore): extrai recordCompressionCacheStats (#3501) (#4792)
chatCore #3501: extract recordCompressionCacheStats. Integrated into release/v3.8.35.
* refactor(chatCore): extrai writeCavemanOutputAnalytics (#3501) (#4794)
chatCore #3501: extract writeCavemanOutputAnalytics. Integrated into release/v3.8.35.
* refactor(chatCore): extrai scheduleQuotaShareConsumption (POST-hook não-streaming, #3501) (#4780)
chatCore #3501: extract scheduleQuotaShareConsumption (non-streaming POST-hook). Integrated into release/v3.8.35.
* refactor(chatCore): extrai emitRequestGamificationEvent (helper compartilhado DRY, #3501) (#4776)
chatCore #3501: extract emitRequestGamificationEvent (DRY streaming/non-streaming). Integrated into release/v3.8.35.
* refactor(chatCore): extrai runPluginOnResponseHook (#3501) (#4782)
chatCore #3501: extract runPluginOnResponseHook. Integrated into release/v3.8.35.
* refactor(chatCore): extrai scheduleStreamingQuotaShareConsumption (POST-hook streaming, #3501) (#4784)
chatCore #3501: extract scheduleStreamingQuotaShareConsumption (streaming POST-hook). Integrated into release/v3.8.35.
* refactor(chatCore): extrai recordStreamingUsageStats (analytics de usage streaming, #3501) (#4791)
chatCore #3501: extract recordStreamingUsageStats. Integrated into release/v3.8.35.
* refactor(chatCore): extrai recordStreamingCost (custo por-request streaming, #3501) (#4790)
chatCore #3501: extract recordStreamingCost (per-request streaming cost). Integrated into release/v3.8.35.
* docs(readme): credit ponytail + OmniCompress; restore env-doc-sync release-green (#4799)
README compression credits (ponytail/OmniCompress) + env-doc-sync ignore for eval-only OMNIROUTE_EVAL_CREDENTIALS (restores release-green after #4720). Integrated into release/v3.8.35.
* chore(quality): trim combo-config.test.ts comments under file-size cap (#4774 follow-up) (#4800)
Restore file-size release-green. Integrated into release/v3.8.35.
* feat(api-docs): Redoc-rendered /api/docs + consolidate OpenAPI spec to docs/openapi.yaml (#4781)
Redoc /api/docs + OpenAPI spec consolidated to docs/openapi.yaml (canonical 201-path complete spec; old path → legacy fallback). All refs/gates/tests/CI updated. Integrated into release/v3.8.35.
* docs(compression): declare Phase 4 layers — Output Styles, adaptive dial, per-request control (#4801)
The README compression section listed the 9 input engines but not the Phase 4
layers now in production:
- Output Styles (output-axis steering: terse-prose / less-code / terse-cjk, lite/full/ultra)
- adaptive context-budget dial (reserve-output|percentage|absolute · floor|replace-autotrigger|off)
- per-request x-omniroute-compression precedence + the offline eval harness
Also bumped the highlights range to v3.8.35, expanded the compression feature bullet,
and marked the GUIDE's Phase 4 row Shipped (was 'Planned' — it's merged on v3.8.35).
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(release): finalize v3.8.35 CHANGELOG + docs reconciliation
- CHANGELOG: complete 3.8.35 section (all 35 commits since v3.8.34,
contributor attribution: @rdself @megamen32 @KooshaPari @JxnLexn)
- docs(security): align THREAT_MODEL.md refs with real code
(routeGuard.ts, tokenLimits.ts, /api/monitoring/health) — fabricated-docs gate
- check:fabricated-docs: skip docs/superpowers/specs (dated research reports)
- i18n: sync 3.8.35 section into 41 CHANGELOG mirrors (docs-sync size gate)
- ratchet rebaseline: cyclomatic 1916->1920, eslintWarnings 3907->3912
(inherited cycle drift; release-finalize diff is docs-only)
* fix(release): resolve inherited base-reds surfaced by v3.8.35 release CI
Cycle base-reds that only run on PR→main (not the PR→release fast-path):
- test(autoCombo): suffixComposition-4517 used node:test in a vitest-only dir
(#4753) → vitest found no suite. Switch to the vitest API. (Vitest job)
- test(agentSkills): openapiParser fixture wrote docs/reference/openapi.yaml;
parser reads docs/openapi.yaml since #4781 → point fixture at the new path.
(Unit/Coverage/Node24/Node26 shard 4)
- test(integration): proxy-pipeline source-scan expected inline streaming-cost
code that #4790/#3501 extracted to the recordStreamingCost leaf → assert the
delegation instead. (Integration 1/2)
- fix(chatCore): derive the log trace id from crypto, not Math.random
(CodeQL js/insecure-randomness — log-correlation id, not a secret).
- test(resilience): circuit-breaker invalid-cooldown fallback asserted t>29000,
flaking on slow CI where ~1.6s elapsed gave t=28401 → tolerate wall-clock
drift (t>25000). (Unit 6/8)
* fix(usage): derive pending-request id from crypto, not Math.random
CodeQL js/insecure-randomness (#669): the pending-request id generated in
trackPendingRequest (usageHistory.ts) flows into attempt logging and was flagged
as insecure randomness in a security context. It's a log-correlation id, not a
secret — switch to crypto RNG to clear the alert. Pairs with the chatCore traceId
fix in 37c49781a (same sink).
---------
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: Demiurge The Single <megamen932@gmail.com>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: Jan Leon <Jan.gaschler@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
322 lines
12 KiB
TypeScript
322 lines
12 KiB
TypeScript
import { pruneByScore } from "./ultraHeuristic.ts";
|
|
import { extractPreservedBlocks } from "./preservation.ts";
|
|
import { DEFAULT_ULTRA_CONFIG } from "./types.ts";
|
|
import type { UltraConfig, CompressionStats, CompressionMode } from "./types.ts";
|
|
import { extractTextContent, mapTextContent, type ChatMessageLike } from "./messageContent.ts";
|
|
import {
|
|
slmAvailable,
|
|
runLlmlinguaUltra,
|
|
prewarmLlmlinguaUltra,
|
|
} from "./engines/llmlingua/ultraEntry.ts";
|
|
|
|
const COMPRESSED_PREFIX = "[COMPRESSED:";
|
|
|
|
/**
|
|
* Async sibling of `mapTextContent`: applies an async transform to each text part
|
|
* of a message's content (string content → single call; array content → each
|
|
* `{type:"text"}` part). Non-text parts and structure are preserved exactly.
|
|
*/
|
|
async function mapTextContentAsync(
|
|
msg: Message,
|
|
fn: (text: string) => Promise<string>
|
|
): Promise<Message> {
|
|
if (typeof msg.content === "string") {
|
|
return { ...msg, content: await fn(msg.content) };
|
|
}
|
|
if (Array.isArray(msg.content)) {
|
|
const next: unknown[] = [];
|
|
for (const part of msg.content) {
|
|
const p = part as Record<string, unknown>;
|
|
if (p && p["type"] === "text" && typeof p["text"] === "string") {
|
|
next.push({ ...p, text: await fn(p["text"] as string) });
|
|
} else {
|
|
next.push(part);
|
|
}
|
|
}
|
|
return { ...msg, content: next };
|
|
}
|
|
return msg;
|
|
}
|
|
|
|
/**
|
|
* Prune PROSE only. Fenced code, inline code, URLs, CONST_CASE, versions, etc. are
|
|
* tombstoned by `extractPreservedBlocks` and re-stitched verbatim, so the heuristic
|
|
* NEVER mangles structured content (mirrors caveman.ts / llmlingua/index.ts).
|
|
*
|
|
* Without this, `pruneByScore` tokenizes the whole text and drops low-score tokens
|
|
* (`b)`, `{`, `+`, …) inside code blocks, corrupting them while leaving the fence
|
|
* markers intact — output that looks like valid code but isn't (B-ULTRA-CODE).
|
|
*/
|
|
function pruneProseOnly(text: string, rate: number, minScore: number): string {
|
|
const { text: withPlaceholders, blocks } = extractPreservedBlocks(text);
|
|
if (blocks.length === 0) return pruneByScore(text, rate, minScore);
|
|
|
|
const placeholderToContent = new Map(blocks.map((b) => [b.placeholder, b.content]));
|
|
const escaped = blocks.map((b) => b.placeholder.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"));
|
|
const splitRe = new RegExp(`(${escaped.join("|")})`, "g");
|
|
|
|
return withPlaceholders
|
|
.split(splitRe)
|
|
.map((part) => {
|
|
if (!part) return "";
|
|
const preserved = placeholderToContent.get(part);
|
|
if (preserved !== undefined) return preserved; // verbatim — never pruned
|
|
return pruneByScore(part, rate, minScore); // prose only
|
|
})
|
|
.join("");
|
|
}
|
|
|
|
/**
|
|
* Compress one prose string with the SLM, preserving code/math/URLs verbatim.
|
|
* Reuses `extractPreservedBlocks` (same tombstoning as `pruneProseOnly`), sends
|
|
* ONLY prose to the worker backend, and re-stitches preserved blocks unchanged.
|
|
* Any backend failure (throw / no-op) falls back to the Tier-A heuristic for that
|
|
* segment, so the SLM NEVER touches structured content and NEVER fails the segment.
|
|
*/
|
|
/**
|
|
* One compressed prose segment plus whether the SLM (not the heuristic fallback)
|
|
* genuinely produced it. `usedSlm` is what the resolver records the tier from — it
|
|
* must reflect "Tier-B ran", NOT merely "the text changed" (a heuristic-fallback
|
|
* also shrinks the text, so deriving the tier from text inequality would mislabel
|
|
* a fallback as "slm").
|
|
*/
|
|
interface ProseSlmResult {
|
|
text: string;
|
|
usedSlm: boolean;
|
|
}
|
|
|
|
async function compressProseSlm(text: string, cfg: UltraConfig): Promise<ProseSlmResult> {
|
|
const { text: withPlaceholders, blocks } = extractPreservedBlocks(text);
|
|
if (blocks.length === 0) {
|
|
return slmOrHeuristic(text, cfg);
|
|
}
|
|
const placeholderToContent = new Map(blocks.map((b) => [b.placeholder, b.content]));
|
|
const escaped = blocks.map((b) => b.placeholder.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"));
|
|
const splitRe = new RegExp(`(${escaped.join("|")})`, "g");
|
|
const parts = withPlaceholders.split(splitRe);
|
|
const out: string[] = [];
|
|
let usedSlm = false;
|
|
for (const part of parts) {
|
|
if (!part) {
|
|
out.push("");
|
|
continue;
|
|
}
|
|
const preserved = placeholderToContent.get(part);
|
|
if (preserved !== undefined) {
|
|
out.push(preserved); // verbatim — never sent to the model
|
|
} else {
|
|
const seg = await slmOrHeuristic(part, cfg);
|
|
out.push(seg.text);
|
|
if (seg.usedSlm) usedSlm = true;
|
|
}
|
|
}
|
|
return { text: out.join(""), usedSlm };
|
|
}
|
|
|
|
/**
|
|
* Run the SLM on a prose segment; on throw/no-op, fall back to the Tier-A pruner for it.
|
|
* `usedSlm` is true ONLY when the SLM backend itself produced the output.
|
|
*/
|
|
async function slmOrHeuristic(prose: string, cfg: UltraConfig): Promise<ProseSlmResult> {
|
|
try {
|
|
const text = await runLlmlinguaUltra(prose, {
|
|
model: cfg.modelPath ? undefined : undefined,
|
|
compressionRate: cfg.compressionRate,
|
|
modelPath: cfg.modelPath,
|
|
});
|
|
return { text, usedSlm: true };
|
|
} catch {
|
|
return { text: pruneByScore(prose, cfg.compressionRate, cfg.minScoreThreshold), usedSlm: false };
|
|
}
|
|
}
|
|
|
|
export interface UltraCompressResult {
|
|
messages: Array<{ role: string; content?: string | unknown[]; [key: string]: unknown }>;
|
|
stats: CompressionStats;
|
|
}
|
|
|
|
type Message = ChatMessageLike;
|
|
|
|
/** Tier the ultra resolver records on the stats. */
|
|
export type UltraTier = "slm" | "heuristic-fallback" | "heuristic";
|
|
|
|
/**
|
|
* Tier-A heuristic ultra (PURE, SYNCHRONOUS). Identical to the pre-B behaviour.
|
|
* Used directly by the stacked sync engine (`cavemanAdapter`) and as the fallback
|
|
* tier inside `ultraCompress`. `tier` lets the async resolver tag the resolved
|
|
* tier as either "heuristic" (chosen directly) or "heuristic-fallback" (SLM failed).
|
|
*/
|
|
export function ultraCompressHeuristic(
|
|
messages: Message[],
|
|
config: Partial<UltraConfig> = {},
|
|
tier: UltraTier = "heuristic"
|
|
): UltraCompressResult {
|
|
const start = Date.now();
|
|
const effectiveConfig: UltraConfig = {
|
|
...DEFAULT_ULTRA_CONFIG,
|
|
...config,
|
|
};
|
|
const { compressionRate, minScoreThreshold, maxTokensPerMessage } = effectiveConfig;
|
|
|
|
let originalChars = 0;
|
|
let compressedChars = 0;
|
|
|
|
const compressed = messages.map((msg) => {
|
|
if (effectiveConfig.preserveSystemPrompt !== false && msg.role === "system") return msg;
|
|
const text = extractTextContent(msg.content);
|
|
if (!text) return msg;
|
|
if (text.startsWith(COMPRESSED_PREFIX)) return msg;
|
|
if (maxTokensPerMessage > 0 && Math.ceil(text.length / 4) <= maxTokensPerMessage) {
|
|
return msg;
|
|
}
|
|
|
|
let messageOriginalChars = 0;
|
|
let messageCompressedChars = 0;
|
|
const next = mapTextContent(msg, (textPart) => {
|
|
if (!textPart || textPart.startsWith(COMPRESSED_PREFIX)) return textPart;
|
|
messageOriginalChars += textPart.length;
|
|
const pruned = pruneProseOnly(textPart, compressionRate, minScoreThreshold);
|
|
messageCompressedChars += pruned.length;
|
|
return pruned;
|
|
}) as Message;
|
|
originalChars += messageOriginalChars;
|
|
compressedChars += messageCompressedChars;
|
|
return next;
|
|
});
|
|
|
|
const originalTokens = Math.ceil(originalChars / 4);
|
|
const compressedTokens = Math.ceil(compressedChars / 4);
|
|
const savingsPercent =
|
|
originalTokens > 0
|
|
? Math.round(((originalTokens - compressedTokens) / originalTokens) * 100 * 10) / 10
|
|
: 0;
|
|
|
|
const stats: CompressionStats = {
|
|
originalTokens,
|
|
compressedTokens,
|
|
savingsPercent,
|
|
techniquesUsed: ["ultra-heuristic-pruning"],
|
|
mode: "ultra" as CompressionMode,
|
|
timestamp: Date.now(),
|
|
durationMs: Date.now() - start,
|
|
ultraTier: tier,
|
|
};
|
|
|
|
return { messages: compressed, stats };
|
|
}
|
|
|
|
/**
|
|
* Ultra compression with the two-tier resolver (Phase 4, Sub-project B).
|
|
*
|
|
* - `ultraEngine: "slm"` AND `slmAvailable()` → route prose through the SLM worker
|
|
* backend (Tier-B). On timeout / worker error / load failure / no-op → fall back
|
|
* to the Tier-A heuristic for THIS request and record "heuristic-fallback".
|
|
* - otherwise → Tier-A heuristic ("heuristic").
|
|
*
|
|
* The structure-preservation wrapper (`extractPreservedBlocks` / re-stitch, inside
|
|
* `pruneProseOnly` for Tier-A and `splitProseAndPreserved` inside the worker engine
|
|
* path for Tier-B) wraps BOTH tiers, so code/math/URLs stay verbatim regardless.
|
|
* A request is NEVER failed or left uncompressed because the SLM was unavailable.
|
|
*/
|
|
export async function ultraCompress(
|
|
messages: Message[],
|
|
config: Partial<UltraConfig> & { ultraEngine?: "heuristic" | "slm" } = {}
|
|
): Promise<UltraCompressResult> {
|
|
if (config.ultraEngine !== "slm" || !slmAvailable()) {
|
|
return ultraCompressHeuristic(messages, config, "heuristic");
|
|
}
|
|
|
|
const start = Date.now();
|
|
const effectiveConfig: UltraConfig = { ...DEFAULT_ULTRA_CONFIG, ...config };
|
|
const { maxTokensPerMessage } = effectiveConfig;
|
|
|
|
let originalChars = 0;
|
|
let compressedChars = 0;
|
|
let anySlm = false;
|
|
|
|
try {
|
|
const compressed: Message[] = [];
|
|
for (const msg of messages) {
|
|
if (effectiveConfig.preserveSystemPrompt !== false && msg.role === "system") {
|
|
compressed.push(msg);
|
|
continue;
|
|
}
|
|
const text = extractTextContent(msg.content);
|
|
if (!text || text.startsWith(COMPRESSED_PREFIX)) {
|
|
compressed.push(msg);
|
|
continue;
|
|
}
|
|
if (maxTokensPerMessage > 0 && Math.ceil(text.length / 4) <= maxTokensPerMessage) {
|
|
compressed.push(msg);
|
|
continue;
|
|
}
|
|
|
|
let messageOriginalChars = 0;
|
|
let messageCompressedChars = 0;
|
|
const next = (await mapTextContentAsync(msg, async (textPart) => {
|
|
if (!textPart || textPart.startsWith(COMPRESSED_PREFIX)) return textPart;
|
|
messageOriginalChars += textPart.length;
|
|
const { text: out, usedSlm } = await compressProseSlm(textPart, effectiveConfig);
|
|
if (usedSlm) anySlm = true;
|
|
messageCompressedChars += out.length;
|
|
return out;
|
|
})) as Message;
|
|
originalChars += messageOriginalChars;
|
|
compressedChars += messageCompressedChars;
|
|
compressed.push(next);
|
|
}
|
|
|
|
const originalTokens = Math.ceil(originalChars / 4);
|
|
const compressedTokens = Math.ceil(compressedChars / 4);
|
|
const savingsPercent =
|
|
originalTokens > 0
|
|
? Math.round(((originalTokens - compressedTokens) / originalTokens) * 100 * 10) / 10
|
|
: 0;
|
|
|
|
const stats: CompressionStats = {
|
|
originalTokens,
|
|
compressedTokens,
|
|
savingsPercent,
|
|
techniquesUsed: anySlm ? ["ultra-slm"] : ["ultra-heuristic-pruning"],
|
|
mode: "ultra" as CompressionMode,
|
|
timestamp: Date.now(),
|
|
durationMs: Date.now() - start,
|
|
ultraTier: anySlm ? "slm" : "heuristic-fallback",
|
|
};
|
|
return { messages: compressed, stats };
|
|
} catch {
|
|
// Any unexpected error in the SLM path → whole-request fail-open to Tier-A.
|
|
return ultraCompressHeuristic(messages, config, "heuristic-fallback");
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Pure decision: should the ultra SLM model be pre-warmed for this config?
|
|
* True only when the SLM tier is selected AND pre-warm is enabled. The CALLER
|
|
* decides timing (enable-transition or cold-start) and fires `prewarmLlmlinguaUltra`
|
|
* best-effort; this helper stays clock-free / side-effect-free.
|
|
*/
|
|
export function shouldPrewarmUltraSlm(config: {
|
|
ultraEngine?: "heuristic" | "slm";
|
|
ultraSlmPrewarm?: boolean;
|
|
}): boolean {
|
|
return config.ultraEngine === "slm" && config.ultraSlmPrewarm === true;
|
|
}
|
|
|
|
/**
|
|
* Best-effort: when the resolved config selects the SLM tier WITH pre-warm,
|
|
* trigger a single warm call. Awaitable for tests; call sites fire-and-forget
|
|
* (`void maybePrewarmUltraSlmOnConfig(cfg)`). Never throws.
|
|
*/
|
|
export async function maybePrewarmUltraSlmOnConfig(config: {
|
|
ultraEngine?: "heuristic" | "slm";
|
|
ultraSlmPrewarm?: boolean;
|
|
}): Promise<void> {
|
|
if (!shouldPrewarmUltraSlm(config)) return;
|
|
try {
|
|
await prewarmLlmlinguaUltra();
|
|
} catch {
|
|
// best-effort
|
|
}
|
|
}
|