Files
OmniRoute/tests/unit/compression/ultra-slm-tier.test.ts
Diego Rodrigues de Sa e Souza cadc3f10b7 Release v3.8.35 (#4743)
* chore(release): open v3.8.35 development cycle

* fix db vacuum scheduler settings (#4726)

Scheduled VACUUM now follows Storage page settings (scheduledVacuum/vacuumHour) as single source of truth; env-flag control path removed. 11/11 vacuum-scheduler tests pass against release/v3.8.35 tip; no orphaned env refs. Integrated into release/v3.8.35.

* fix(tier): noAuth providers count as free; free filter returns empty … (#4753)

noAuth providers now classified free (union of legacy list + NOAUTH_PROVIDERS chat-tier derivation), -free arena_elo alias, and auto/<cat>:free returns an empty pool when no free candidate matches (opt-in legacy fallback via OMNIROUTE_AUTO_FREE_FALLBACK_TO_FULL_POOL). New env var documented in .env.example + ENVIRONMENT.md; CHANGELOG bullet added (maintainer co-author). 46/46 node + 56/56 vitest tests pass on release tip; env-doc-sync, docs-sync, typecheck:core, lint, file-size all green. Integrated into release/v3.8.35.

* refactor(chatCore): extrai 11 helpers de nível superior para 6 leaves puros (#3501) (#4571)

chatCore god-file decomposition (#3501): extract 6 pure leaves (cacheUsageMeta, executorClientHeaders, nonStreamingResponseBody, skillsFormat, streamErrorResult, streamFinalize) from chatCore.ts. Rebased onto release/v3.8.35 tip (resolved single chatCore.ts conflict — removed now-extracted inline buildExecutorClientHeaders). 265/265 chatcore tests, 26/26 new leaf tests, typecheck:core, cycles, file-size all green. Integrated into release/v3.8.35.

* refactor(chatCore): extrai resolveExecutorWithProxy + getExecutionCredentials para leaves (#3501) (#4646)

chatCore #3501: extract resolveExecutorWithProxy + getExecutionCredentials to leaves (executorProxy.ts, executionCredentials.ts). Clean cherry-pick onto release tip post-#4571. 12/12 new leaf tests, typecheck:core, cycles, file-size green. Integrated into release/v3.8.35.

* refactor(chatCore): extrai transforms de mensagens Claude p/ leaf (#3501) (#4708)

chatCore #3501: extract Claude upstream-message transforms to leaf (claudeUpstreamMessages.ts + claudeMessageTypes.ts). Clean cherry-pick post-#4646. 8/8 new leaf tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.

* refactor(chatCore): extrai persistAttemptLogs para leaf (#3501) (#4717)

chatCore #3501: extract persistAttemptLogs to leaf (attemptLogging.ts). Rebased onto release tip post-#4708 (resolved imports conflict: kept tip's resolveCompressionHeader from compression Phase 3, dropped now-unused logTruncation import moved into the leaf). 288/288 chatcore tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.

* refactor(chatCore): extrai stageTrace + compressionUsageReceipt para leaves (#3501) (#4721)

chatCore #3501: extract stageTrace + compressionUsageReceipt to leaves. Clean cherry-pick post-#4717. 6/6 new leaf tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.

* refactor(chatCore): extrai prepareUpstreamBody (1ª sub-fatia do executeProviderRequest, #3501) (#4730)

chatCore #3501: extract prepareUpstreamBody (first sub-slice of executeProviderRequest) to leaf (upstreamBody.ts). Clean cherry-pick post-#4721. 7/7 new leaf tests, full 301/301 chatcore suite, typecheck/cycles/file-size green. Completes the 6-PR chatCore decomposition stack into release/v3.8.35.

* fix(db): make db-backup import size cap configurable (#4719) (#4757)

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>

* chore(quality): expand check:release-green to the FULL release-PR gate set (#4758)

The release-green pre-flight (Solution C) previously covered only a subset of the
gates that run exclusively on the release PR (PR→main), so reds still accrued
silently on release/** and surfaced in ~40-min layers at release time (v3.8.34:
3 CI rounds — CodeQL sanitization, then the fail-fast Quality Ratchet revealing
openapi then cyclomatic-complexity one push at a time, plus zizmor/integration).

Now check:release-green reproduces the COMPLETE release-PR gate set and reports
EVERY red in one pass (collected, not fail-fast):

- New DRIFT ratchets (report-only, rebaselined at release, never block):
  cyclomatic complexity, dead-code, type-coverage, compression-budget,
  openapi-coverage, workflow-lint (zizmor), codeql-ratchet.
- New HARD gates (real defects): docs-all (fabricated-docs strict + i18n mirror
  sync) and the integration test suite (gated behind !--quick).

The only release-PR gates it still cannot reproduce locally are GitHub-side CodeQL
semantic analysis and SonarQube/SonarCloud (external services).

The nightly-release-green workflow and /green-prs inherit the expanded coverage
automatically (they invoke this script), so cycle drift is now surfaced
continuously and the release PR is green on its first CI run.

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>

* fix(dashboard): add missing onboarding.tiers step title (#4698) (#4755)

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>

* feat(compression): Output Styles registry + D0 telemetry (Phase 4A) (#4694)

Phase 4A: Output Styles registry + D0 telemetry. Integrated into release/v3.8.35.

* feat(compression): SLM tier for ultra (Phase 4B) [stacked on #4694] (#4707)

Phase 4B: SLM tier for ultra. Integrated into release/v3.8.35.

* feat(compression): context-budget adaptive compression (Phase 4C) [stacked on #4707] (#4716)

Phase 4C: adaptive context-budget compression. Integrated into release/v3.8.35.

* feat(compression): offline evaluation harness (Phase 4 D1) [stacked on #4716] (#4720)

Phase 4 D1: offline evaluation harness. Integrated into release/v3.8.35.

* fix(sse): deepseek-web folds role:tool results into prompt transcript (#4712) (#4756)

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>

* fix(dashboard): remove dead unconditional useLiveRequests call in HomePageClient (#4759, #4745, #4596) (#4761)

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>

* fix(dashboard): dedupe provider nodes by id on compatible-provider add (#4746) (#4768)

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>

* chore(db): re-export compressionRunTelemetry from localDb to satisfy db-rules (#4775)

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>

* docs(security): add canonical STRIDE-based threat model (#4783)

Canonical STRIDE threat model. Integrated into release/v3.8.35.

* test(dashboard): add smoke test for home client dashboard (#4793)

Smoke test guarding the dashboard home client render (regression #4745/#4759). Code fix already landed via #4761; this PR's jsdom smoke test is the net-new regression guard. Integrated into release/v3.8.35.

* fix(combos): auto-promote zeroLatencyOptimizationsEnabled so legacy configs (pre-3.8.33 fallbackCompressionMode="lite") round-trip on the first GUI edit (#4774)

Auto-promote zeroLatencyOptimizationsEnabled + strip v3.8.31-era removed keys so legacy combo configs round-trip through PUT /api/combos/{id} on first GUI edit (closes #4382 followup). Pre-merge: rewrote the now-stale reject test to assert auto-promotion + added passthrough/round-trip regression guards; reconciled combos/page.tsx file-size baseline. Integrated into release/v3.8.35.

* refactor(chatCore): extrai parse + usage-stats não-streaming do executeProviderRequest (#3501) (#4762)

chatCore #3501: extract parseNonStreamingResponseBody + recordNonStreamingUsageStats. Integrated into release/v3.8.35.

* refactor(chatCore): extrai recordContextEditingTelemetryHook (#3501) (#4779)

chatCore #3501: extract recordContextEditingTelemetryHook. Integrated into release/v3.8.35.

* refactor(chatCore): extrai recordCompressionCacheStats (#3501) (#4792)

chatCore #3501: extract recordCompressionCacheStats. Integrated into release/v3.8.35.

* refactor(chatCore): extrai writeCavemanOutputAnalytics (#3501) (#4794)

chatCore #3501: extract writeCavemanOutputAnalytics. Integrated into release/v3.8.35.

* refactor(chatCore): extrai scheduleQuotaShareConsumption (POST-hook não-streaming, #3501) (#4780)

chatCore #3501: extract scheduleQuotaShareConsumption (non-streaming POST-hook). Integrated into release/v3.8.35.

* refactor(chatCore): extrai emitRequestGamificationEvent (helper compartilhado DRY, #3501) (#4776)

chatCore #3501: extract emitRequestGamificationEvent (DRY streaming/non-streaming). Integrated into release/v3.8.35.

* refactor(chatCore): extrai runPluginOnResponseHook (#3501) (#4782)

chatCore #3501: extract runPluginOnResponseHook. Integrated into release/v3.8.35.

* refactor(chatCore): extrai scheduleStreamingQuotaShareConsumption (POST-hook streaming, #3501) (#4784)

chatCore #3501: extract scheduleStreamingQuotaShareConsumption (streaming POST-hook). Integrated into release/v3.8.35.

* refactor(chatCore): extrai recordStreamingUsageStats (analytics de usage streaming, #3501) (#4791)

chatCore #3501: extract recordStreamingUsageStats. Integrated into release/v3.8.35.

* refactor(chatCore): extrai recordStreamingCost (custo por-request streaming, #3501) (#4790)

chatCore #3501: extract recordStreamingCost (per-request streaming cost). Integrated into release/v3.8.35.

* docs(readme): credit ponytail + OmniCompress; restore env-doc-sync release-green (#4799)

README compression credits (ponytail/OmniCompress) + env-doc-sync ignore for eval-only OMNIROUTE_EVAL_CREDENTIALS (restores release-green after #4720). Integrated into release/v3.8.35.

* chore(quality): trim combo-config.test.ts comments under file-size cap (#4774 follow-up) (#4800)

Restore file-size release-green. Integrated into release/v3.8.35.

* feat(api-docs): Redoc-rendered /api/docs + consolidate OpenAPI spec to docs/openapi.yaml (#4781)

Redoc /api/docs + OpenAPI spec consolidated to docs/openapi.yaml (canonical 201-path complete spec; old path → legacy fallback). All refs/gates/tests/CI updated. Integrated into release/v3.8.35.

* docs(compression): declare Phase 4 layers — Output Styles, adaptive dial, per-request control (#4801)

The README compression section listed the 9 input engines but not the Phase 4
layers now in production:
- Output Styles (output-axis steering: terse-prose / less-code / terse-cjk, lite/full/ultra)
- adaptive context-budget dial (reserve-output|percentage|absolute · floor|replace-autotrigger|off)
- per-request x-omniroute-compression precedence + the offline eval harness
Also bumped the highlights range to v3.8.35, expanded the compression feature bullet,
and marked the GUIDE's Phase 4 row Shipped (was 'Planned' — it's merged on v3.8.35).

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(release): finalize v3.8.35 CHANGELOG + docs reconciliation

- CHANGELOG: complete 3.8.35 section (all 35 commits since v3.8.34,
  contributor attribution: @rdself @megamen32 @KooshaPari @JxnLexn)
- docs(security): align THREAT_MODEL.md refs with real code
  (routeGuard.ts, tokenLimits.ts, /api/monitoring/health) — fabricated-docs gate
- check:fabricated-docs: skip docs/superpowers/specs (dated research reports)
- i18n: sync 3.8.35 section into 41 CHANGELOG mirrors (docs-sync size gate)
- ratchet rebaseline: cyclomatic 1916->1920, eslintWarnings 3907->3912
  (inherited cycle drift; release-finalize diff is docs-only)

* fix(release): resolve inherited base-reds surfaced by v3.8.35 release CI

Cycle base-reds that only run on PR→main (not the PR→release fast-path):

- test(autoCombo): suffixComposition-4517 used node:test in a vitest-only dir
  (#4753) → vitest found no suite. Switch to the vitest API. (Vitest job)
- test(agentSkills): openapiParser fixture wrote docs/reference/openapi.yaml;
  parser reads docs/openapi.yaml since #4781 → point fixture at the new path.
  (Unit/Coverage/Node24/Node26 shard 4)
- test(integration): proxy-pipeline source-scan expected inline streaming-cost
  code that #4790/#3501 extracted to the recordStreamingCost leaf → assert the
  delegation instead. (Integration 1/2)
- fix(chatCore): derive the log trace id from crypto, not Math.random
  (CodeQL js/insecure-randomness — log-correlation id, not a secret).
- test(resilience): circuit-breaker invalid-cooldown fallback asserted t>29000,
  flaking on slow CI where ~1.6s elapsed gave t=28401 → tolerate wall-clock
  drift (t>25000). (Unit 6/8)

* fix(usage): derive pending-request id from crypto, not Math.random

CodeQL js/insecure-randomness (#669): the pending-request id generated in
trackPendingRequest (usageHistory.ts) flows into attempt logging and was flagged
as insecure randomness in a security context. It's a log-correlation id, not a
secret — switch to crypto RNG to clear the alert. Pairs with the chatCore traceId
fix in 37c49781a (same sink).

---------

Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: Demiurge The Single <megamen932@gmail.com>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: Jan Leon <Jan.gaschler@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 17:06:18 -03:00

353 lines
13 KiB
TypeScript

/**
* Ultra SLM tier — wiring the ultra mode's modelPath / slmFallbackToAggressive config
* to the real model path (the llmlingua engine).
*
* Until now ultra was a pure heuristic (pruneByScore) and modelPath /
* slmFallbackToAggressive were inert config. The async entry point now routes ultra
* through the llmlingua engine when modelPath is set, falling back per
* slmFallbackToAggressive when the model is unavailable / yields no gain.
*
* The llmlingua backend is injectable (setLlmlinguaBackend), so the tier is testable
* without the real ONNX model.
*/
import { describe, it, after, afterEach, test } from "node:test";
import assert from "node:assert/strict";
import { applyCompressionAsync } from "../../../open-sse/services/compression/index.ts";
import { setLlmlinguaBackend } from "../../../open-sse/services/compression/engines/llmlingua/index.ts";
import { DEFAULT_ULTRA_CONFIG } from "../../../open-sse/services/compression/types.ts";
// Comfortably above the llmlingua default 2000-token floor (estimate ≈ chars / 4).
const LARGE_PROSE = "The quick brown fox jumps over the lazy dog every morning. ".repeat(260);
function body() {
return { model: "gpt-4o", messages: [{ role: "user", content: LARGE_PROSE }] };
}
function ultraOpts(ultra: Record<string, unknown>) {
// Only config.ultra is read by the ultra SLM tier; the rest of CompressionConfig is unused here.
return { config: { ultra: { ...DEFAULT_ULTRA_CONFIG, ...ultra } } } as unknown as Parameters<
typeof applyCompressionAsync
>[2];
}
let backendCalls = 0;
function trackingCompressingBackend(text: string): Promise<string> {
backendCalls++;
return Promise.resolve(text.slice(0, Math.max(1, Math.floor(text.length / 3))));
}
function identityBackend(text: string): Promise<string> {
backendCalls++;
return Promise.resolve(text); // no gain → llmlingua reports compressed:false
}
function throwingBackend(_text: string): Promise<string> {
backendCalls++;
return Promise.reject(new Error("model unavailable"));
}
afterEach(() => {
backendCalls = 0;
});
after(() => setLlmlinguaBackend(null));
function techniques(stats: unknown): string[] {
return ((stats as { techniquesUsed?: string[] } | null)?.techniquesUsed ?? []) as string[];
}
describe("ultra SLM tier — modelPath routes through llmlingua", () => {
it("runs the SLM tier when modelPath is set and the model compresses", async () => {
setLlmlinguaBackend(trackingCompressingBackend);
const result = await applyCompressionAsync(
body(),
"ultra",
ultraOpts({ modelPath: "/models/fake.onnx", compressionRate: 0.5 })
);
assert.equal(backendCalls > 0, true, "backend was consulted");
assert.equal(result.compressed, true);
assert.equal((result.stats as { mode?: string } | null)?.mode, "ultra");
assert.ok(techniques(result.stats).includes("ultra-slm"), "tagged as the ultra SLM tier");
});
it("falls back to aggressive when the model yields no gain and slmFallbackToAggressive is on", async () => {
setLlmlinguaBackend(identityBackend);
const result = await applyCompressionAsync(
body(),
"ultra",
ultraOpts({ modelPath: "/models/fake.onnx", slmFallbackToAggressive: true })
);
assert.ok(techniques(result.stats).includes("aggressive"), "fell back to aggressive");
assert.ok(!techniques(result.stats).includes("ultra-slm"));
});
it("falls back to the heuristic when the model fails and slmFallbackToAggressive is off", async () => {
setLlmlinguaBackend(throwingBackend);
const result = await applyCompressionAsync(
body(),
"ultra",
ultraOpts({ modelPath: "/models/fake.onnx", slmFallbackToAggressive: false })
);
const techs = techniques(result.stats);
assert.equal((result.stats as { mode?: string } | null)?.mode, "ultra");
assert.ok(techs.includes("ultra"), "heuristic ultra ran");
assert.ok(!techs.includes("ultra-slm"), "not the SLM tier");
assert.ok(!techs.includes("aggressive"), "not the aggressive fallback");
});
it("uses the heuristic and never touches the model when modelPath is unset", async () => {
setLlmlinguaBackend(throwingBackend); // would blow up if (wrongly) consulted
const result = await applyCompressionAsync(body(), "ultra", ultraOpts({ modelPath: "" }));
assert.equal(backendCalls, 0, "model not consulted without modelPath");
assert.equal((result.stats as { mode?: string } | null)?.mode, "ultra");
assert.ok(!techniques(result.stats).includes("ultra-slm"));
});
});
// ─── Phase 4 (B): SLM-tier resolver, probe, telemetry, pre-warm ──────────────
// (Appended to the pre-existing legacy `modelPath` suite above, which must stay green.)
import { DEFAULT_COMPRESSION_CONFIG } from "../../../open-sse/services/compression/types.ts";
test("DEFAULT_COMPRESSION_CONFIG defaults ultraEngine to 'heuristic'", () => {
assert.equal(DEFAULT_COMPRESSION_CONFIG.ultraEngine, "heuristic");
});
test("DEFAULT_COMPRESSION_CONFIG defaults ultraSlmPrewarm to false", () => {
assert.equal(DEFAULT_COMPRESSION_CONFIG.ultraSlmPrewarm, false);
});
import type { CompressionStats } from "../../../open-sse/services/compression/types.ts";
test("CompressionStats accepts an optional ultraTier signal", () => {
const s = {
originalTokens: 10,
compressedTokens: 5,
savingsPercent: 50,
techniquesUsed: ["ultra-heuristic-pruning"],
mode: "ultra" as const,
timestamp: 1,
ultraTier: "heuristic" as const,
} satisfies CompressionStats;
assert.equal(s.ultraTier, "heuristic");
});
import {
ultraCompress,
ultraCompressHeuristic,
} from "../../../open-sse/services/compression/ultra.ts";
test("ultraCompressHeuristic is a synchronous pure heuristic (no SLM)", () => {
const cfg = {
enabled: true,
compressionRate: 0.5,
minScoreThreshold: 0.3,
slmFallbackToAggressive: false,
maxTokensPerMessage: 0,
};
const r = ultraCompressHeuristic(
[{ role: "user", content: "the quick brown fox jumps over the lazy dog" }],
cfg
);
assert.equal(r.stats.mode, "ultra");
assert.equal(r.stats.ultraTier, "heuristic");
assert.ok(r.stats.techniquesUsed.includes("ultra-heuristic-pruning"));
});
test("ultraCompress with default config (no ultraEngine) uses heuristic tier", async () => {
const r = await ultraCompress([{ role: "user", content: "the quick brown fox jumps" }], {
enabled: true,
compressionRate: 0.5,
minScoreThreshold: 0.3,
slmFallbackToAggressive: false,
maxTokensPerMessage: 0,
});
assert.equal(r.stats.ultraTier, "heuristic");
});
import {
__setUltraSlmTestHooks,
__resetUltraEntryForTests,
} from "../../../open-sse/services/compression/engines/llmlingua/ultraEntry.ts";
test("ultraEngine:'slm' with available stub backend records ultraTier:'slm'", async () => {
__setUltraSlmTestHooks({
available: true,
run: async (text) => text.slice(0, Math.ceil(text.length / 2)),
});
try {
const r = await ultraCompress(
[{ role: "user", content: "the quick brown fox jumps over the lazy dog repeatedly today" }],
{
enabled: true,
compressionRate: 0.5,
minScoreThreshold: 0.3,
slmFallbackToAggressive: false,
maxTokensPerMessage: 0,
ultraEngine: "slm",
}
);
assert.equal(r.stats.ultraTier, "slm");
assert.ok(r.stats.techniquesUsed.includes("ultra-slm"));
assert.ok(r.stats.compressedTokens <= r.stats.originalTokens);
} finally {
__resetUltraEntryForTests();
}
});
test("ultraEngine:'slm' but backend throws → ultraTier:'heuristic-fallback'", async () => {
__setUltraSlmTestHooks({
available: true,
run: async () => {
throw new Error("worker timeout");
},
});
try {
const r = await ultraCompress(
[{ role: "user", content: "the quick brown fox jumps over the lazy dog repeatedly today" }],
{
enabled: true,
compressionRate: 0.5,
minScoreThreshold: 0.3,
slmFallbackToAggressive: false,
maxTokensPerMessage: 0,
ultraEngine: "slm",
}
);
assert.equal(r.stats.ultraTier, "heuristic-fallback");
} finally {
__resetUltraEntryForTests();
}
});
test("ultraEngine:'slm' but slmAvailable() false → heuristic tier (no SLM attempt)", async () => {
__setUltraSlmTestHooks({
available: false,
run: async () => {
throw new Error("should not be called");
},
});
try {
const r = await ultraCompress([{ role: "user", content: "the quick brown fox jumps" }], {
enabled: true,
compressionRate: 0.5,
minScoreThreshold: 0.3,
slmFallbackToAggressive: false,
maxTokensPerMessage: 0,
ultraEngine: "slm",
});
assert.equal(r.stats.ultraTier, "heuristic");
} finally {
__resetUltraEntryForTests();
}
});
test("SLM tier preserves fenced code + URLs verbatim (structure wrapper)", async () => {
// Stub the SLM to lowercase prose — any leakage of code/URL into it would show.
__setUltraSlmTestHooks({
available: true,
run: async (text) => (text.trim() ? text.toLowerCase() + " x" : text),
});
try {
const code = "```js\nconst A = 1; // KEEP\n```";
const url = "https://Example.com/Path";
// The fenced block must open at line-start for `extractPreservedBlocks` to tombstone
// it (the same rule the heuristic Tier-A already relies on); both tiers share that
// wrapper, so this proves the SLM tier preserves structure identically.
const content = `Some PROSE here and ${url} trailing PROSE\n${code}\nmore PROSE after`;
const r = await ultraCompress([{ role: "user", content }], {
enabled: true,
compressionRate: 0.5,
minScoreThreshold: 0.3,
slmFallbackToAggressive: false,
maxTokensPerMessage: 0,
ultraEngine: "slm",
});
const out = r.messages[0].content as string;
assert.ok(out.includes(code), "fenced code block must survive verbatim");
assert.ok(out.includes(url), "URL must survive verbatim");
} finally {
__resetUltraEntryForTests();
}
});
import { ultraEngine } from "../../../open-sse/services/compression/engines/cavemanAdapter.ts";
test("stacked ultraEngine.apply stays synchronous and compresses via heuristic", () => {
const res = ultraEngine.apply(
{ messages: [{ role: "user", content: "the quick brown fox jumps over the lazy dog" }] },
{ config: { ultra: { compressionRate: 0.5 } } as never }
);
// Synchronous result object (not a Promise), with a real stats record.
assert.equal(typeof (res as { then?: unknown }).then, "undefined");
assert.ok(res.stats);
});
test("applyCompressionAsync ultra + ultraEngine:'slm' (stub) yields ultraTier in stats", async () => {
__setUltraSlmTestHooks({
available: true,
run: async (text) => text.slice(0, Math.ceil(text.length / 2)),
});
try {
const reqBody = {
messages: [
{ role: "user", content: "the quick brown fox jumps over the lazy dog more than once" },
],
};
const result = await applyCompressionAsync(reqBody, "ultra", {
config: {
enabled: true,
defaultMode: "ultra",
ultraEngine: "slm",
ultra: {
enabled: true,
compressionRate: 0.5,
minScoreThreshold: 0.3,
slmFallbackToAggressive: false,
maxTokensPerMessage: 0,
},
} as never,
});
assert.equal(result.stats?.ultraTier, "slm");
} finally {
__resetUltraEntryForTests();
}
});
import * as compression from "../../../open-sse/services/compression/index.ts";
test("compression index re-exports the ultra-SLM surface", () => {
assert.equal(typeof compression.ultraCompressHeuristic, "function");
assert.equal(typeof compression.slmAvailable, "function");
assert.equal(typeof compression.runLlmlinguaUltra, "function");
assert.equal(typeof compression.prewarmLlmlinguaUltra, "function");
});
import { shouldPrewarmUltraSlm } from "../../../open-sse/services/compression/ultra.ts";
test("shouldPrewarmUltraSlm: true only when slm + prewarm both on", () => {
assert.equal(shouldPrewarmUltraSlm({ ultraEngine: "slm", ultraSlmPrewarm: true }), true);
assert.equal(shouldPrewarmUltraSlm({ ultraEngine: "slm", ultraSlmPrewarm: false }), false);
assert.equal(shouldPrewarmUltraSlm({ ultraEngine: "heuristic", ultraSlmPrewarm: true }), false);
assert.equal(shouldPrewarmUltraSlm({}), false);
});
import { maybePrewarmUltraSlmOnConfig } from "../../../open-sse/services/compression/ultra.ts";
test("maybePrewarmUltraSlmOnConfig fires prewarm when slm+prewarm on (stub)", async () => {
let warmed = 0;
__setUltraSlmTestHooks({
available: true,
run: async (t) => {
warmed++;
return t.slice(0, 1);
},
});
try {
await maybePrewarmUltraSlmOnConfig({ ultraEngine: "slm", ultraSlmPrewarm: true });
assert.equal(warmed, 1);
await maybePrewarmUltraSlmOnConfig({ ultraEngine: "heuristic", ultraSlmPrewarm: true });
assert.equal(warmed, 1); // unchanged — heuristic does not prewarm
} finally {
__resetUltraEntryForTests();
}
});