* chore(release): open v3.8.34 development cycle * chore(quality): release-green pre-flight validator + nightly signal (C+D) (#4622) C — scripts/quality/validate-release-green.mjs (npm run check:release-green): reproduces the release-equivalent validation (typecheck, eslint, db-rules, public-creds, full unit, vitest, ratchets, optional --with-build package-artifact) against the current working tree and classifies each red as HARD (real defect, exit 1) vs DRIFT (ratchet — reported, never affects exit / never blocks). Pure helpers exported + orchestration behind a direct-run guard; unit-tested. D — .github/workflows/nightly-release-green.yml: runs C on the active release branch nightly (and on workflow_dispatch) and opens/updates a single tracking issue on HARD failures. Never a required check, never touches a contributor PR. Closes the gap where the full gate (ci.yml) only ran on the release PR, so reds accrued silently on release/** and surfaced in 40-min layers at release time. Non-blocking by construction; drift is the maintainer's to rebaseline at release. Co-authored-by: Diego Rodrigues de Sa e Souza <diego.souza@cdwasolutions.com.br> * fix(providers): show revealed connection API keys (#4583) Integrated into release/v3.8.34 * fix(resilience): respect upstream retry hint toggle (#4585) Integrated into release/v3.8.34 * feat(settings): expose stream recovery feature flags (#4586) Integrated into release/v3.8.34 * fix(logs): make active request stale sweep configurable (#4599) Integrated into release/v3.8.34 * fix(plugin): auto-prefix providerId with 'opencode-' for OC 1.17.8+ native gate (#4527) Integrated into release/v3.8.34 (supersedes #4445) * fix(models): treat unknown output caps as unset (#4584) Integrated into release/v3.8.34 * fix(executors): strip temperature for GitHub Copilot gpt-5.4 family (#4564) Integrated into release/v3.8.34 (rebuilt onto tip) * fix(oauth): update Qwen OAuth URLs from chat.qwen.ai to qwen.ai (#4561) Integrated into release/v3.8.34 (rebuilt onto tip) * fix(api/settings): prevent cached /api/settings responses (port from 9router#951) (#4566) Integrated into release/v3.8.34 (rebuilt onto tip) * feat(audio): MiniMax T2A v2 TTS dispatch in audioSpeech (port #1043) (#4553) Integrated into release/v3.8.34 (rebuilt onto tip) * fix(dashboard): surface manual config CTA when Open Claw CLI auto-detect fails (#4562) Integrated into release/v3.8.34 (rebuilt onto tip) * feat(providers): optional model ID for custom API-key validation (#4555) Integrated into release/v3.8.34 (rebuilt onto tip) * fix(cli): align data dir and env loading with runtime (#4607) Integrated into release/v3.8.34 (rebuilt onto tip) * fix(quota): expose Bailian quota windows (#4610) Integrated into release/v3.8.34 (rebuilt onto tip) * fix: retain provider cooldowns for configured max window (#4588) Integrated into release/v3.8.34 (rebuilt — bundled commits stripped) * fix: reject invalid provider cooldown bounds (#4589) Integrated into release/v3.8.34 (rebuilt — bundled commits stripped) * fix: preserve production combo metrics on shadow eviction (#4590) Integrated into release/v3.8.34 (rebuilt — bundled commits stripped) * fix(stream): estimate input tokens when upstream reports prompt_tokens=0 (#4615) Integrated into release/v3.8.34 (rebuilt onto tip) * fix(catalog): shorten no-thinking gateway prefix to no-think/ (#4525) Integrated into release/v3.8.34 (rebuilt — kept only the prefix rename, dropped stale-base reverts) * fix(relay): apply IP rate limit to bifrost sidecar (#4593) Integrated into release/v3.8.34 (rebuilt onto tip; merge before #4612) * fix(bifrost): finalize SSE relay usage after stream (#4612) Integrated into release/v3.8.34 (rebuilt + reconciled with #4593) * feat(compression): per-request `x-omniroute-compression` header (Phase 3) (#4645) * docs(compression): Phase 3 per-request header design spec Approved brainstorming output for the x-omniroute-compression header: header-first precedence, name-first combo matching (Decision A), explicit value bypasses auto-trigger (Decision B), DerivedPlan.source, and the X-OmniRoute-Compression response header. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(compression): Phase 3 per-request header implementation plan 4-task TDD plan (resolver header-first + source, parser, chatCore wiring + response header, docs/file-size) with full code and exact commands. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(compression): header-first resolver + plan source (Phase 3 core) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(compression): resolveCompressionHeader parser (Phase 3) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(compression): wire x-omniroute-compression header + response header (Phase 3) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(compression): extract plan-resolution leaf (planResolution.ts) under size cap (Phase 3) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(compression): document x-omniroute-compression header (Phase 3) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(compression): harden named-combo map + trim engine: header id (Phase 3 review) Addresses gemini-code-assist review on #4645: - Extract buildNamedComboLookup (pure) so a blank/whitespace/null combo name contributes only its id key (no '' key, no throw that disables all combos). - Trim the engine:<id> header value so 'engine: rtk' resolves. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Diego Rodrigues de Sa e Souza <diego.souza@cdwasolutions.com.br> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com> * fix: exclude exhausted connections from auto scoring (#4592) Integrated into release/v3.8.34 (rebuilt + opt-in gate fix) * fix(dashboard): memoize compatible provider groups (#4613) Integrated into release/v3.8.34 (rebuilt + test added) * fix(dashboard): isolate quota widget refresh clock (#4611) Integrated into release/v3.8.34 (rebuilt + jsdom test) * fix(dashboard): gate topology side effects behind widget visibility (#4606) Integrated into release/v3.8.34 (rebuilt + jsdom test) * fix(dashboard): keep play_arrow spinning on provider Test All buttons (#4563) Integrated into release/v3.8.34 (rebuilt onto tip; UI-cosmetic per owner) * fix(db): schedule retention cleanup + fix cleanup table/column names (extracted from #4428) (#4691) Integrated into release/v3.8.34 (cleanup core extracted from #4428, credit @oyi77) * fix(telemetry): back off live-WS event forwarding when the sidecar is unreachable (#4604) (#4687) Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com> * fix(api): serve GET /v1/models/{model} as JSON, not the HTML dashboard (#4674) (#4677) Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com> * feat(opencode): add go deepseek reasoning variants (#4647) Integrated into release/v3.8.34 * fix(executors): robust deepseek-web tool-call parsing and agentic context retention (#4644) Integrated into release/v3.8.34 * fix(cli): authenticate `omniroute logs` and honor active context (#4638) Integrated into release/v3.8.34 (authored by Rahul Sharma, AI co-author trailer stripped per project policy) * fix(proxy): apply pipelining:0 + connections cap to the direct dispatcher (#4580) (#4684) Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com> * fix(executors): Firecrawl web_fetch 500 with include_metadata=true (#4692) Integrated into release/v3.8.34 * fix(routing): include all noAuth models in auto-combos + add reka-flash + best-free template (#4621) Integrated into release/v3.8.34 (dead getFirstRegistryModelId dropped, rebuilt onto tip) * fix(dashboard): gate home topology live-WS networking (#4596) (#4618) Integrated into release/v3.8.34 (adapted onto #4606's extracted topology section: default-hidden flip + enabled gate on useLiveDashboard) * fix(cli): align `omniroute` env loading with the runtime data dir (#4597) (#4619) Integrated into release/v3.8.34 (data-dir.mjs refactor reconciled with #4607; loadEnvFile aligned to getDefaultDataDir) * chore(quality): reconcile file-size baseline for #4644 (deepseek-web.ts 1117->1125) (#4695) file-size reconcile for #4644 * Support quota scraping for OpenCode Go and Ollama Cloud (#4642) Integrated into release/v3.8.34 (Ollama Cloud + OpenCode Go dashboard quota scraping; rebuilt onto tip, gates green: typecheck/public-creds/file-size/lint/docs-sync + 31 tests) * feat(executors): land M365 Copilot pure framing + connection helpers (#4042) (#4696) Land M365 pure modules ahead of draft #4400 * deps: bump production + development groups; migrate js-yaml to v5 ESM (#4697) Incorporates Dependabot #4667 + #4668 + js-yaml v5 ESM migration into release/v3.8.34 * fix: noAuth provider validation + kimi executor routing (#4699) Integrated into release/v3.8.34 (noAuth in NOAUTH_PROVIDERS dynamic check + remove misrouted kimi web alias; 9 tests) * refactor(imageGeneration): extract 8 provider families to co-located files (#4609) Integrated into release/v3.8.34 (extraction completed: added missing imports/exports per module, main imports handlers locally; 145 image-gen tests pass, typecheck/cycles/file-size green) * chore(release): v3.8.34 — finalize changelog, rebaseline drift, fix release-green reds - Finalize CHANGELOG [3.8.34] (43 bullets, full contributor attribution) + seed i18n mirrors - Rebaseline inherited cycle drift surfaced by release-green pre-flight: eslint warnings 3900->3907, cognitive-complexity 797->801 (release-finalize touches no prod code; all drift is from this cycle's contributor merges) - fix(providers): keep reka-flash-3 as the Reka provider default. #4621 inserted reka-flash at the head of the model list, silently changing the default from reka-flash-3 (the free-tier model) to reka-flash; reorder so reka-flash-3 stays default, reka-flash retained. - test: align provider-models-config / provider-models-route / web-cookie-providers-new with #4621 (reka-flash now in the Reka catalog) and #4699 (the `kimi` API-key provider correctly falls through to DefaultExecutor instead of KimiWebExecutor) - chore(quality): allowlist the COMPRESSION_GUIDE doc name in check-fabricated-docs (false-positive env-var match; docs/compression/COMPRESSION_GUIDE.md exists) * fix(release-green): resolve release-PR full-CI reds for v3.8.34 Surfaced only on the release PR (these gates don't run on PR->release fast-gates): - fix(quota): complete HTML-comment sanitization in opencodeOllamaUsage SSR reset-time parsing — strip any <!--...--> generically instead of the two literal React hydration markers, so no partial "<!--" can survive (CodeQL js/incomplete-multi-character- sanitization, HIGH, introduced by #4642). Regression test added. - test(codex): correct the Codex-fingerprint body key order assertion to match the canonical bodyFieldOrder (prompt_cache_key precedes include); #4584 flipped the two and integration tests don't run on fast-gates so it never executed until the release PR. - chore(quality): rebaseline inherited cycle drift surfaced by full CI — zizmorFindings 152->155 (+3 unpinned-uses in nightly-release-green.yml from #4622, same @vN convention as ci.yml) and openapiCoverage.pct 38.4->37.8 (-0.6, contributor routes added faster than openapi docs). Release-finalize touches no prod routes. * fix(release-green): complete CodeQL sanitization + rebaseline complexity drift - fix(quota): handle unterminated HTML comments in opencodeOllamaUsage SSR reset-time parsing — the `(?:-->|$)` arm consumes a trailing "<!--" with no closing "-->", so no partial "<!--" can survive (CodeQL js/incomplete-multi-character-sanitization persisted with the plain <!--...--> form because an unclosed comment could still leave "<!--"). - chore(quality): rebaseline cyclomatic complexity 1915->1916 (+1) — inherited v3.8.34 cycle drift (contributor feature branches); check:complexity does not run on PR->release fast-gates so it surfaced only on the release PR. Release-finalize adds 0 complexity (measured 1916 with/without the regex tweak). dead-code/cognitive/type-coverage/ compression-budget/codeql ratchets all pass. --------- Co-authored-by: Diego Rodrigues de Sa e Souza <diego.souza@cdwasolutions.com.br> Co-authored-by: Randi <55005611+rdself@users.noreply.github.com> Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com> Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com> Co-authored-by: Abhishek Divekar <adivekar@utexas.edu> Co-authored-by: Rahul sharma <sharmaR0810@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com> Co-authored-by: Ronald Estacion <DevEstacion@users.noreply.github.com> Co-authored-by: Igor <60442260+BugsBag@users.noreply.github.com> Co-authored-by: Oonishi <275808243+ponkcore@users.noreply.github.com> Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com> Co-authored-by: Jan Leon <Jan.gaschler@gmail.com>
11 KiB
title, version, lastUpdated
| title | version | lastUpdated |
|---|---|---|
| Compression Config Panel — Phase 3: Per-Request Header | 3.8.33 | 2026-06-21 |
Compression Config Panel — Phase 3: Per-Request Header
Status: approved direction (2026-06-21), pending spec review
Base branch: release/v3.8.33 (Phase 1 #4432, Phase 2 #4521 merged)
Goal: Let a client override the resolved compression plan for a single request via the
x-omniroute-compression HTTP header, taking precedence over every operator-configured
layer (routing override, active profile, auto-trigger, Default). The compression resolver
is already header-aware in shape; Phase 3 adds the header parsing, threads the value to the
top of the resolver, and surfaces the resolved plan back to the client via a response header.
Phase 1 (#4432) built the engines map, deriveDefaultPlan, resolveCompressionPlan
(header/active-combo-aware in signature), and persisted activeComboId. Phase 2 (#4521)
wired the active-profile selector and lifted active-combo resolution into resolveBasePlan.
Phase 3 is the final piece of the original three-phase plan.
1. Background — current state
The resolver is partially header-aware but the header never reaches the top of the decision. From the code map:
open-sse/services/compression/resolveCompressionPlan.tsalready acceptsResolveCtx.headerand interpretsoff/default/engine:<id>/<combo>via a privateheaderToPlanhelper. But no caller ever passesheadertoday, so the branch is dormant.open-sse/services/compression/strategySelector.tsis the real entry point (selectCompressionStrategy,selectCompressionPlan,getEffectiveMode,applyCompression). ItsresolveBasePlanonly callsresolveCompressionPlanin two of five precedence paths (routing-combo override and theenginesExplicitderived-default). The active-profile and auto-trigger paths short-circuit andreturnbefore those calls. So merely threadingheaderinto the existingresolveCompressionPlancalls would not give the header top precedence — it would be silently ignored whenever an active profile is set (the common case after Phase 2). The header must be evaluated at the top ofresolveBasePlan.open-sse/handlers/chatCore/headers.tsalready has the parsing precedent:isNoMemoryRequested(x-omniroute-no-memory, PR #4290) + the case-insensitivegetHeaderValueCaseInsensitivereader.src/lib/db/compressionCombos.ts: a combo'sidis auuidv4()(except the seededdefault-caveman), and itsnameisTEXT NOT NULLwith no UNIQUE constraint (migration 042). So a header value that names a combo is far more usable as the name than the opaque UUIDid.open-sse/handlers/chatCore.tsbuilds the response headers (X-OmniRoute-Model,X-OmniRoute-Cache, …) viabuildStreamingResponseHeaders+attachOmniRouteMetaHeaderson the main response path. There is currently no response header reporting the applied compression plan.
2. The header contract
x-omniroute-compression: <value> — mirrors the x-omniroute-no-memory / no-cache
convention. Parsed alongside the other omniroute request headers. Keyword values and the
engine: prefix are case-insensitive. Values:
| Value | Meaning |
|---|---|
off |
No compression for this request. |
default |
The panel-derived Default plan, deterministically — ignores active profile, routing override, and auto-trigger. |
engine:<id> |
A single engine, when that engine is enabled in config (e.g. engine:rtk). |
<combo> |
A named combo. Matched by name (case-insensitive) first, then by exact id. |
Decision A — combo matched by name: because the stored id is a UUID, the ergonomic
header value is the combo's name (e.g. my-fast-combo). Names are not unique in the DB,
so the contract is documented as first-match-wins; clients wanting determinism can pass
the exact id.
Decision B — an explicit header value is authoritative: any valid value
(off / default / engine:<id> / <combo>) bypasses auto-trigger. For example,
default on a very large prompt keeps the panel Default rather than auto-escalating. The
mental model is "the header decides, full stop."
Invalid / unknown value → ignored. Resolution falls through to the normal operator
precedence; the request is never rejected. A debug log line under the COMPRESSION
channel records the unrecognized value for observability.
3. Architecture
3.1 Resolution model (per request, most-specific wins)
x-omniroute-compression header (per-request) <- NEW top of precedence
-> routing-combo override (comboOverrides[comboId]) (per-route)
-> active profile (activeComboId) (global, Phase 2)
-> auto-trigger (large prompt -> autoTriggerMode)
-> Default = derived from panel engines map
-> off (master disabled, or zero engines on)
The header is evaluated at the top of resolveBasePlan. A valid value returns its plan
immediately; an unknown value falls through to the existing precedence unchanged.
3.2 The resolver source
resolveBasePlan (and the public selectCompressionPlan) return the existing
DerivedPlan ({ mode, stackedPipeline }) extended with an optional source field —
which precedence layer decided the plan:
request-header | routing-override | active-profile | auto-trigger | default | off
mode answers what compression runs; source answers who decided. The field is
optional so Phase 1/2 callers and snapshots are unaffected. chatCore reads plan.source
(+ plan.mode) to build the response header.
3.3 Header parsing helper
A new pure function in open-sse/handlers/chatCore/headers.ts, mirroring
isNoMemoryRequested:
export function resolveCompressionHeader(
headers: Record<string, unknown> | Headers | null | undefined
): string | null {
const value = (getHeaderValueCaseInsensitive(headers, "x-omniroute-compression") || "").trim();
return value || null;
}
It returns the raw trimmed value (or null); the resolver owns interpretation and casing
rules (so the single source of truth for "what a value means" stays in the resolver, with
the parser only reading the wire).
3.4 Threading
chatCore reads the header from clientRawRequest?.headers, then passes it as a new
header?: string | null argument (default undefined) through
selectCompressionStrategy / selectCompressionPlan / getEffectiveMode ->
resolveBasePlan, exactly the pattern Phase 2 used for combos. Phase 1/2 call sites that
omit the argument are byte-for-byte unchanged.
For the <combo> form, chatCore builds the named-combo map keyed by both combo id and
lowercased name ({ [c.id]: c.pipeline, [c.name.toLowerCase()]: c.pipeline }). The
active-profile lookup keys on config.activeComboId (always a UUID/slug id), so the added
name keys are inert for it — one map serves both paths. The resolver matches <combo>
name-first (per Decision A): it looks up the value lowercased (hitting a name key, or an
already-lowercase id), then falls back to the value as-is (an exact id) —
combos[value.toLowerCase()] ?? combos[value]. All combo ids are lowercase
(uuidv4() hex or the default-caveman slug), so an exact id still resolves on the first
lookup.
3.5 Where the header logic lives
The existing headerToPlan interpretation (off / default / engine:<id> / <combo>)
is reused. resolveBasePlan evaluates the header before the routing-override branch.
default routes to deriveDefaultPlanFromConfig (which already handles both
enginesExplicit and legacy defaultMode), so default means "the Default profile" for
every install type. The resolver remains pure — no src/lib/db import (enforced by the
existing cycle/source guard).
4. Observability
A new compression key in OMNIROUTE_RESPONSE_HEADERS
(src/shared/constants/headers.ts) -> response header:
X-OmniRoute-Compression: <mode>; source=<source>
Examples: aggressive; source=request-header, off; source=request-header,
stacked; source=active-profile, lite; source=auto-trigger, off; source=off.
chatCore captures { mode, source } from the compression resolution (computed early,
~line 1530) into an outer-scope variable and injects the header when building
responseHeaders (~line 4697), so it appears on both streaming and non-streaming
responses. The header is informational only and never affects routing.
5. Error handling & safety
- Header absent / blank ->
null; behaviour is byte-identical to Phase 2. - Unknown value -> silent fall-through to normal resolution + a
COMPRESSIONdebug log; never a 4xx. The response header reflects the layer that actually won (e.g.source=auto-trigger), notrequest-header. engine:<id>naming a disabled / unknown engine -> fall-through (same rule as the currentheaderToPlan: returnsnull).- Gating: the header is honored unconditionally, like
x-omniroute-no-memory. Rationale: it only affects the compression of the client's own request. The worst case is a client opting itself out of compression (off), which increases only that client's own upstream token count — there is no cross-tenant, security, or cost-shifting concern.
6. Components (units & responsibilities)
| Unit | Responsibility | Depends on |
|---|---|---|
resolveCompressionHeader (chatCore/headers.ts) |
Read the raw header value off the wire. | getHeaderValueCaseInsensitive |
resolveBasePlan + headerToPlan (strategySelector.ts / resolveCompressionPlan.ts) |
Interpret the value, evaluate header-first precedence, return { mode, stackedPipeline, source }. |
config, combos map (pure) |
DerivedPlan.source (deriveDefaultPlan.ts type) |
Carry which layer decided. | — |
OMNIROUTE_RESPONSE_HEADERS.compression (shared/constants/headers.ts) |
Name the response header. | — |
chatCore wiring (chatCore.ts) |
Parse, thread header, build id+name combo map, capture {mode,source}, emit response header. |
all of the above |
7. Testing
- Parser unit (
headers.ts): each value form, absent, blank, mixed casing, value with surrounding whitespace -> correct raw/null. - Resolver unit (
strategySelector/resolveCompressionPlan): each form resolves the expected plan andsource; header beats an active profile and a routing override; unknown value falls through to normal resolution; a valid value bypasses auto-trigger on a large prompt;defaultreturns the derived Default for bothenginesExplicitand legacy installs;<combo>matches by name and by id. - Integration / fetch-capture: per-request precedence end-to-end (same config, header
present vs absent yields different applied plans) and the
X-OmniRoute-Compressionresponse header value. - Source guard: the resolver still has no
src/lib/dbimport.
8. Scope
In scope: header parsing, top-of-precedence wiring, source on DerivedPlan, the
response header, docs (API_REFERENCE / COMPRESSION_GUIDE), tests.
Out of scope (YAGNI): bare mode names in the header (lite/aggressive/…); panel UI for
the header; honoring the header on non-chat paths (combo.ts proactive-fallback, the
preview route). The header is a chat-request feature.