mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-06 15:22:12 +03:00
* fix(quality): resolve net-new lint errors and allowlist #9343 assert rewrite Two `no-explicit-any` errors landed with #9407 and #9320 after the suppressions inventory was generated. Project policy is to fix new violations rather than freeze them, so both are typed instead: - #9407: `executor as unknown as Record<string, unknown>` - #9320: `(k: { name?: string })` Also allowlists the net-assert reduction in web-tools-translation-2820 (39->35). #9343 inverted the contract — bare JSON must no longer be promoted to tool_calls without an explicit <tool> envelope — so the tests were rewritten to assert non-promotion, which costs fewer asserts than validating a promoted object. More restrictive, not weaker. * fix(quality): raise integration ceiling to 40min and unpin codex-cli version in test The integration gate's 20min ceiling killed a healthy run: measured 22m08s hermetic on an idle 16-core box (935 tests across 112 files, strictly serial at --test-concurrency=1 because ~16 of them bind a port or share a DB). The "~3-10min" estimate in the code was stale by ~3x. 40min keeps the ceiling's real purpose — turning a genuine hang into a visible failure — without failing a long-but-healthy suite. Also fixes a base-red in chat-pipeline:564c204efebumped DEFAULT_CODEX_CLIENT_VERSION to 0.146.0 but the User-Agent assertion still pinned 0.144.1. The line two above already read the constant via getCodexClientVersion(); this one duplicated the literal. Deriving it from the same source stops the next bump from breaking the test again. * fix(ratelimit): re-arm Bottleneck reservoir heartbeat after updateSettings Bottleneck 2.19.5 (frozen upstream dependency, no release since 2019) has a bug in LocalDatastore#_startHeartbeat() (node_modules/bottleneck/lib/ LocalDatastore.js:29,56): the guard `if (this.heartbeat == null && ...)` only (re)creates the periodic reservoir-refresh interval the first time it runs. Every later call -- including the one updateSettings() itself triggers internally -- falls into the else branch and does clearInterval(this.heartbeat) WITHOUT resetting the reference back to null. Because the stale reference sticks around, every future _startHeartbeat() call keeps taking the same dead else branch: the periodic reservoir refresh is gone forever after the first manual updateSettings() call on a limiter. Every limiter created by this file starts with a live heartbeat (buildLimiterDefaults() always sets reservoirRefreshInterval/ reservoirRefreshAmount), so the very first updateFromHeaders() / updateFromResponseBody() / applyRequestQueueSettings() call against a limiter permanently kills its refresh. In production this wedges the request queue once the reservoir hits 0: an auto-enrolled apikey connection accumulates its default 60 requests, the reservoir zeroes, the queue freezes for ~120s, the watchdog fires a synthetic 502 (RATE_LIMIT_QUEUE_WEDGED), the connection cools down and gets excluded from weighted combo pools -- turning a configured 70/30 split into ~50/50. Add applyLimiterSettings(), a module-local wrapper around limiter.updateSettings() that nulls the stale heartbeat reference and re-invokes _startHeartbeat() afterward so it takes the "start a fresh interval" branch again. Route all 5 updateSettings() call sites through it (updateAllLimiterSettings, both updateFromHeaders() branches, loadPersistedLimits(), and updateFromResponseBody()). updateAllLimiterSettings is now async and awaited by its two callers (initializeRateLimits, applyRequestQueueSettings); the sync call sites use the existing trackAsyncOperation() fire-and-forget tracking pattern. tests/integration/combo-matrix/weighted.test.ts is the E2E proof: the "weighted: 70/30" case now passes with zero WEDGED/RATE_LIMIT_QUEUE/502 log lines across 200 sequential requests (previously the wedge/recovery cycle inflated its runtime and skewed the distribution toward ~50/50). Refs #8213 * fix(tests): remove stray TDD probes committed by accident inf4e93f339dThree TDD repro/probe test files landed on the release tip viaf4e93f339d(docs: add management authentication terminology guide, files from a worktree. Each file is a pre-fix TDD probe that belongs to a *different*, still-in-flight fix branch/PR and duplicates a file path that PR already owns and will properly update on merge: - tests/unit/authz/probe-9033-repro.test.ts: probe for #9033 (IP blacklist direct-connection bypass). 3/4 asserts fail against this tree (D1, D2, Bonus — all assert the not-yet-implemented target behavior); D3 passes (pre-existing behavior). Owned by PR #9385 (open, unmerged), which modifies this exact path. - tests/unit/repro-8522.test.ts: probe for #8522 (absolute file-size baseline reds innocent PRs on inherited drift). First test fails against this tree's evaluateFileSizes (still absolute-only); second (sanity: real growth still flags) passes. #8522 is actually CLOSED upstream — PR #9355 merged the real fix into release/v3.8.50 today (2026-08-05T15:53Z) modifying this exact path — but this branch's merge-base with release/v3.8.50 (6b0e11e378) predates that merge, so the fix has not synced into this tree yet. - tests/unit/repro-8956.test.ts: probe for #8956 (resolveProjectRoot stops at synthetic Next.js standalone package.json). First test fails against this tree; second (sanity: named package.json still resolves) passes. Owned by PR #9354 (open, unmerged), which modifies this exact path. Each deleted file's real implementation + passing version already exists in its owning PR and will land normally through that PR's own merge — deleting the premature copy here does not lose any coverage. No config/quality/test-masking-allowlist.json entry was added: the _deletedWithReplacement schema only supports `replacement` (a test file that must already exist in this tree's HEAD — none does, the real versions live in the unmerged sibling PRs above) or `sourceRemoved` (production files that must be absent from HEAD — they are not, none of the three issues are implemented in this tree). Neither shape fits an "owned by an in-flight sibling PR" deletion, so the CI test-masking gate will flag these 3 deletions for mandatory human review on this branch's next PR diff against release/v3.8.50 — flagged for the owner rather than inventing a new allowlist shape. Refs #9033, #8522, #8956, #7786 * fix(tests): align 8189-classifier-compat with #9276 always-mode semantics tests/unit/8189-classifier-compat-auto-narrow.test.ts was a test-sibling forgotten when #9276 (commit6b531fbacd) removed the unconditional `if (mode === "always") return true` branch from shouldDefaultAllowClassifier(). tests/unit/claude-classifier-compat.test.ts was updated in that same commit; this file was not. Old contract: 'always' mode short-circuited every Claude-format request unconditionally (operator opt-in was treated as sufficient on its own). New contract: 'always' now requires the same SECURITY_MONITOR_MARKER system-prompt text as 'auto' — the marker-optional behavior let a normal chat request through /v1/messages be silently swallowed by an operator's 'always' opt-in. The single 'always' test (1 assert, no-marker body expecting true) is replaced by two tests mirroring the depth already used for 'auto' mode in the same file: no-marker/false and marker-present/true. Net effect is +1 assert, not a reduction — the new pair verifies both directions of the narrowed contract instead of only the now-incorrect unconditional case. Before: 3/4 pass (the 'always' test failed: expected true, got false). After: 5/5 pass. Refs #9276 * fix(tests): align deepseek-web-tools-execute with #9343 tool envelope contract tests/unit/deepseek-web-tools-execute-2820.test.ts (executor level) was a test-sibling forgotten when #9343 (commitd969555417) hardened tool-call parsing: bare JSON with no explicit <tool>/<tool_call> envelope is never promoted to tool_calls anymore (previously it was, whenever a tools[] set was requested — a security gap allowing prose/code-fenced JSON echoed back by the model, or a copy-attack, to trigger real tool execution). Three siblings were updated in the same commit: web-tools-translation.test.ts and web-tools-translation-2820.test.ts (parseToolCallsFromText, the shared translator), and deepseek-web-tools-variants.test.ts (parseDeepSeekToolCalls, deepseek-specific parser) — all inverted their bare-JSON assertions to `toolCalls === null` + `content === text` (preserved verbatim, not stripped). This file calls the executor's execute() (full HTTP round trip through buildToolAwareResult), so it was not touched by that diff and kept asserting the old contract (finish_reason: "tool_calls", content: null). Verified against source (open-sse/executors/deepseek-web.ts buildToolAwareResult): when parseDeepSeekToolCalls returns toolCalls=null, hasCalls is false, so finish_reason is "stop", message.tool_calls is never set, and message.content is the parser's returned content — which for text with no <tool>/<tool_call> tag at all is the original string, unchanged (parseToolCallsFromText's early-return branch). The test now asserts exactly that shape, at the same executor level as the rest of the file's tool_calls that make sense at that level as the rest of the file's tool_calls Refs #9343 * fix(tests): align visionBridge tests with #8430 contract (partial — see note) Two test-siblings were forgotten when #8430 (commit7e55abbc41) hardened Vision Bridge's vision-model selection: getBestVisionModel() now validates that a candidate has a usable active connection (hasUsableCredentialsForModel, DB-backed) before returning it, instead of unconditionally returning the fixedModel or a hardcoded "openai/gpt-4o-mini" default. Three siblings were updated in the same commit (visionBridgeRouter.test.ts, the new repro-8430.test.ts, vision-bridge-preserve-on-failure-4012.test.ts); these two were not. tests/unit/guardrails/visionBridgeHelpers.callVisionModel.test.ts (8 failures, all "No vision-capable provider connected"): callVisionModel()'s `routerConfig` param only merges into getBestVisionModel's CONFIG argument, never its `deps` argument, so there is no way to inject a credentials stub through this function's public signature (unlike the guardrail class and getBestVisionModel itself, which do accept an injectable `hasUsableCredentials`). These tests exercise callVisionModel's own request/response handling, not credential routing (already covered elsewhere), so the fix seeds one real usable `provider_connections` row per provider the file exercises (openai, anthropic) via createProviderConnection in a test.before() hook, with resetDbInstance() in test.after() per the DB-handle-cleanup convention. All 8 now pass. tests/unit/guardrails/visionBridge.test.ts (7 failures): 1 of the 7 (VB-S03) is a genuine forgotten-contract case, fixed here — same semantic flip already applied to vision-bridge-preserve-on-failure-4012.test.ts: in the combo describe path, when EVERY describe call fails, the raw image is now replaced with an "(unavailable)" stub instead of preserved, because that path is only reached for confirmed non-vision targets. Assertions inverted to match (imagePart undefined, unavailable-stub present), same assert count, no weakening. *** THE OTHER 6 (VB-S12, VB-S12b, VB-S01, VB-S13, VB-S07, VB-S10) ARE DELIBERATELY LEFT FAILING. *** These are NOT a #8430 contract change — root- caused to what looks like a separate, unintentional regression: the ONE call to getBestVisionModel() in visionBridge.ts's whole-request-reroute path (line 244, `getBestVisionModel({ fixedModel: configuredModel })`) does not pass a `deps` second argument, so it always uses the real DB-backed hasUsableCredentialsForModel instead of this.deps.hasUsableCredentials — even though the two adjacent checks in the very same function (`checkCreds(model)` at line 226, `checkCreds(bestModel)` at line 246) DO honor the injectable override. In this suite's empty-but-readable isolated test DB, that real check deterministically returns `false` (not the indeterminate `null` the file's own createGuardrail() comment says these tests rely on: "Fail-open (null) so classic VB-S01/S07/S10 reroute tests keep working without a live credential DB"), so getBestVisionModel silently returns null, the reroute branch's `if (bestModel && ...)` guard never fires, and every test that expects a reroute observes a silent no-op instead. Evidence this is a source gap, not a test that needs updating: - The file's own pre-existing comment names VB-S01/S07/S10 as tests the `null` fail-open default is SUPPOSED to keep green. - VB-CRED-01/02 (the file's only two tests that actually inject a non-default hasUsableCredentials mock) both pass today, but neither one's assertions distinguish "mock honored" from "mock ignored, real check also says no" — they don't prove the threading works, they just don't happen to notice it's missing. - visionBridgeRouter.test.ts, repro-8430.test.ts, and vision-bridge-preserve-on-failure-4012.test.ts (22 tests, all green) all either call getBestVisionModel directly with explicit deps, or mock callVisionModel wholesale (bypassing getBestVisionModel entirely) — none of them exercises this exact call site through the guardrail's own deps. Per instructions, this was intentionally NOT "fixed" by weakening these 6 tests' assertions (that would mask the gap) or by seeding fake DB credentials to route around it (that would hide a real production DI inconsistency behind a test-only workaround) or by touching src/lib/guardrails/visionBridge.ts (a production behavior change outside a test-alignment task's scope, and Hard Rule #18 requires its own TDD/validation cycle). Flagging for the owner: the likely one-line fix is threading `{ hasUsableCredentials: this.deps.hasUsableCredentials }` as getBestVisionModel's second argument at visionBridge.ts:244, mirroring the two adjacent call sites in the same function. Before: 15 failures (7 + 8). After: 9 pass added (1 + 8), 6 still fail (unchanged, by design). Refs #8430 * feat(quality): add strayFromCommit deletion allowlist form to test-masking gate The deletion allowlist supported two shapes: replacement (test rewritten elsewhere) and sourceRemoved (feature deleted). Neither fits a third legitimate case surfaced today: test files that entered the repo BY ACCIDENT — commitf4e93f339d(#7786 docs) swept another session's worktree artifacts into the release, including TDD probes owned by open fix PRs (probe-9033-repro -> PR #9385, repro-8956 -> PR #9354, repro-8522 -> PR #9355). Those probes fail by design until their owning PR merges, so every unit run on the release tip broke on them. The new strayFromCommit form is verified, not trusted: the gate asks git which commit actually ADDED the file (git log --diff-filter=A) and only exempts the deletion when it matches the declared hash; a non-empty reason naming the owning PR/issue is mandatory. Also allowlists the deepseek-web-tools-execute assert reduction (23->21) fromed661f2126— same #9343 contract-inversion class as the existing web-tools-translation entry. Gate unit tests: 55/55 pass. Full gate vs main: OK. * fix(guardrails): pass credential deps to getBestVisionModel at reroute call site The individual-model reroute path in VisionBridgeGuardrail.preCall() calls getBestVisionModel({ fixedModel: configuredModel }) without its second `deps` argument, so the router always falls back to the real DB-backed hasUsableCredentialsForModel instead of an injected `deps.hasUsableCredentials` override. The two adjacent credential checks in the same function (the original-model check and the best-model check, both via the local `checkCreds` binding) already thread deps correctly — only this middle call, added in #8430, was left out. Pass the same resolved `checkCreds` used by those two adjacent checks as `getBestVisionModel`'s deps argument so all three credential checks in this reroute path stay consistent. Fixes 6 tests in tests/unit/guardrails/visionBridge.test.ts that depended on the injected hasUsableCredentials mock being honored on this path: VB-S12, VB-S12b, VB-S01, VB-S13, VB-S07, VB-S10. Refs #8430 * fix(quality): raise unit ceiling to 100min and align 2 more forgotten sibling tests Unit ceiling 45->100min: a hermetic-env measurement on the loaded devbox (load 7-26) was still inside invocation 1 of 3 at 76min when killed; contention factor 2-3x measured, no idle measurement exists. The pre-flight's real condition is exactly that contended one (unit runs in Promise.all with integration+vitest), and there 45min provably killed a healthy suite and fabricated a false base-red. The 45min value came from v3.8.43 as an estimate never validated by measurement. TODO in-code: re-tighten after an idle run on the .113 box. Also aligns the 6th and 7th occurrences of the same systemic pattern (behavior change merged updating only part of the sibling tests): - issue-7859-gemini-web-redirect-valid: #9407 refined ServiceLogin redirects to mean expired session; the #7859 regression coverage is preserved via a non-ServiceLogin public redirect variant. - provider-validation-specialty claude-web 429: #9406 inverted the contract (rate-limited session is unhealthy); the dedicated repro file owns the full contract, this sibling now matches it. Also carries the file-size rebaseline for #9323's base.ts growth (1578->1623, WAF retry + burst guard) and the eslintWarnings baseline tightened 5000->0 (real measured value with the TS7 suppressions in place — 5000 left the ratchet inert). Refs #9407, #9406, #9323 * fix(tests): restore the 3 TDD probes now owned by merged fixes and drop their stray allowlist entries The base advanced while this PR was open: the real fixes for the three issues behind the stray probes all merged into release/v3.8.50 — #9385 (issue 9033), #9355 (issue 8522) and #9354 (issue 8956). - probe-9033-repro / repro-8522: the base rewrote both probes into the regression tests of their merged fixes, so the delete side of the rebase conflict was dropped and the base versions kept. - repro-8956: #9354 only realigned one fixture line in auto-update.test.ts (package.json marker now needs a name field) and added no test for the new skip-synthetic behavior — the probe is the ONLY regression coverage of that merged fix (2/2 green on the base), so deleting it would remove real coverage. Restored. With no test-file deletions left in the PR diff, the three strayFromCommit allowlist entries are stale and removed. The strayFromCommit form support in check-test-masking.mjs stays (covered by its own fixtures). * fix(quality): rebaseline file-size for PR #9529 own growth The base sits exactly at the old frozen values, so the base-relative mode (#8522) does not cover this growth — it is this PR's own: - open-sse/services/rateLimitManager.ts 1060->1105: the applyLimiterSettings() helper that re-arms the reservoir heartbeat after updateSettings (Bottleneck 2.19.5 fix, TDD in ratelimit-reservoir-refresh.test.ts). - tests/integration/chat-pipeline.test.ts 1592->1598: codex User-Agent derived from getCodexClientVersion() instead of a pinned literal. - tests/unit/provider-validation-specialty.test.ts 2980->2985: new claude-web 429 -> valid:false coverage (#9406). * fix(docs): sync provider count to 291 in README and CLAUDE The live catalog counts 291 providers but README.md/CLAUDE.md still said 290, so the STRICT 'Docs Gates (fast-path)' check reds EVERY open PR against release/v3.8.50 (verified on #9537/#9539 as well — inherited base-red, not introduced by this PR). Updated all provider-count mentions including the section anchor. * fix(tests): align launch-codex 6312 guard with the async #9454 spawn contract #9454 made resolveCodexSpawn async (PATH-probes a native codex.exe before the .cmd shim) and updated its own tests, but left this older sibling calling the function synchronously — destructuring the Promise yields undefined and reds Unit fast-path (1/4) for EVERY open PR against the release (verified on #9537/#9539; inherited base-red). Realigned to the async contract with an injected probe; keeps the original #6312 fallback guard plus the only non-Windows codex coverage (now also asserting the probe never runs off Windows). * fix(translator): move state-mutating reasoning summary helper out of the pure leaf #9500 added buildResponsesReasoningSummaryDelta(state, ...) to pureHelpers.ts, but the function reads AND mutates stream state (reasoningSummaryIndex map) — violating the leaf contract declared in the file header ('no host imports, no stream state') and guarded by response-openai-responses-purehelpers-split.test.ts, which reds Unit fast-path (4/4) for every open PR (inherited base-red, verified on #9537/#9539). Moved verbatim to the host next to the other stream-state helpers (markResponsesReasoningDeltaEmitted); the host was its only consumer. Behavior unchanged: repro-9500-reasoning-separator 3/3 green, leaf/host architecture tests green. * fix(quality): rebaseline openai-responses.ts for the leaf-state relocation The #9500 helper moved from pureHelpers.ts into the host (previous commit) grows the host file 1174->1204 while the leaf shrinks by the same amount — net-zero LOC across the pair, but the per-file frozen ratchet only sees the growing side. * fix(tests): let the 9442 cert-mode test see past the harness trust-store guard tests/_setup/isolateDataDir.ts sets OMNIROUTE_SKIP_SYSTEM_TRUST=1 globally, which makes installCert() return before issuing any command — so the #9442 install-gap test captured nothing and could NEVER pass under npm run test:unit (it only passed invoked directly, harness-less; inherited base-red on Unit fast-path 3/4, verified on #9537/#9539). Clear the flag for this file only (restored in test.after): safe because every spawned command is a logging stub on PATH and OMNIROUTE_NO_SUDO=1 strips sudo, so nothing touches the real trust store. 6/6 under the CI harness including system-trust-test-guard. --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
719 lines
30 KiB
JavaScript
719 lines
30 KiB
JavaScript
#!/usr/bin/env node
|
||
// scripts/quality/validate-release-green.mjs
|
||
//
|
||
// "Release-green" pre-flight validator (Solution C).
|
||
//
|
||
// WHY: the full gate (ci.yml — unit shards, vitest, ratchets, package-artifact)
|
||
// runs ONLY on the release PR (PR → main). PRs into release/** only get the
|
||
// fast-gates (quality.yml: TIA-impacted tests + typecheck + lint checks). So
|
||
// reds accumulate silently on the release branch and explode — in layers — at
|
||
// release time. This script reproduces the release-equivalent validation against
|
||
// the CURRENT working tree so the maintainer (or the nightly, Solution D) can see
|
||
// the real state of the release branch at any time.
|
||
//
|
||
// DESIGN — never blocking to contributors:
|
||
// • HARD checks (typecheck, lint errors, db-rules, public-creds, docs-all,
|
||
// unit, vitest, integration, optionally package-artifact) → a failure here is
|
||
// a real defect; exit 1.
|
||
// • DRIFT checks (eslint WARNINGS, cognitive-complexity, file-size, cyclomatic
|
||
// complexity, dead-code, type-coverage, compression-budget, openapi-coverage,
|
||
// workflow-lint/zizmor, codeql-ratchet) → ratchet drift accrued across the
|
||
// cycle is NOT a contributor's fault; it is reported and rebaselined by the
|
||
// maintainer at release. Drift NEVER changes the exit code, so wiring this as
|
||
// a check can never block anyone on drift.
|
||
//
|
||
// COMPLETENESS: this mirrors the FULL release-PR gate set (quality-gate +
|
||
// quality-extended + docs-sync-strict + integration), not a subset — and reports
|
||
// EVERY red in one pass (the report is collected, not fail-fast), so the release
|
||
// PR is green on its first CI run instead of revealing reds in ~40-min layers. The
|
||
// only release-PR gates it cannot reproduce locally are GitHub-side CodeQL semantic
|
||
// analysis and SonarQube/SonarCloud (external services).
|
||
//
|
||
// This script DIAGNOSES + REPORTS only (no auto-fix). The fix-to-green
|
||
// orchestration lives in the /green-prs + review-prs flows that call it.
|
||
//
|
||
// Usage:
|
||
// node scripts/quality/validate-release-green.mjs [--json] [--with-build] [--quick] [--full-ci] [--hermetic]
|
||
// --json emit machine-readable JSON to stdout (report goes to stderr)
|
||
// --with-build also run check:pack-artifact (needs a dist/ build — slow)
|
||
// --quick skip the slow unit + vitest + integration suites (drift + fast
|
||
// gates only)
|
||
// --full-ci ALSO run every static gate declared in ci.yml's gate jobs (lint,
|
||
// quality-gate, quality-extended, docs-sync-strict, pr-test-policy) —
|
||
// read straight from ci.yml so the set never drifts. Catches the whole
|
||
// "static base-red" category the curated list missed (v3.8.46: 11 of 16
|
||
// leaked reds). Pair with --quick for the fast "1 command, 0 CI layers" pass.
|
||
// --hermetic scrub OMNIROUTE_API_KEY/OMNIROUTE_URL from gate env so live
|
||
// tests self-skip exactly like CI (dev machines otherwise run
|
||
// them against localhost and produce false-positive reds)
|
||
//
|
||
// Per-gate output is saved to _artifacts/release-green/<gate>.log (gitignored) —
|
||
// diagnose a red from the file instead of re-running the gate.
|
||
|
||
import { execFile, execFileSync } from "node:child_process";
|
||
import { promisify } from "node:util";
|
||
import { mkdirSync, readFileSync, writeFileSync } from "node:fs";
|
||
import { dirname, join } from "node:path";
|
||
import { fileURLToPath } from "node:url";
|
||
import { parse as parseYaml } from "yaml";
|
||
|
||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||
const ROOT = join(__dirname, "..", "..");
|
||
const npmCmd = process.platform === "win32" ? "npm.cmd" : "npm";
|
||
|
||
// Per-gate captured output. execFileSync buffers everything and the report only
|
||
// shows a one-line summary, so without these files every red requires RE-RUNNING
|
||
// the gate just to see the detail (the dominant cost of the 2026-07-05 pre-flight).
|
||
const LOG_DIR = join(ROOT, "_artifacts", "release-green");
|
||
function saveGateLog(id, out) {
|
||
try {
|
||
mkdirSync(LOG_DIR, { recursive: true });
|
||
writeFileSync(join(LOG_DIR, `${id}.log`), String(out ?? ""));
|
||
} catch {
|
||
/* log persistence is best-effort — never fails a gate */
|
||
}
|
||
}
|
||
|
||
// ─── Pure helpers (exported for tests) ──────────────────────────────────────
|
||
|
||
/** Read the committed ratchet baseline value for a metric (null if unknown). */
|
||
export function baselineValue(metric, root = ROOT) {
|
||
try {
|
||
const raw = JSON.parse(
|
||
readFileSync(join(root, "config/quality/quality-baseline.json"), "utf8")
|
||
);
|
||
const metrics = raw.metrics || raw;
|
||
const v = metrics?.[metric]?.value;
|
||
return typeof v === "number" ? v : null;
|
||
} catch {
|
||
return null;
|
||
}
|
||
}
|
||
|
||
/** Best-effort "first meaningful failure line" from captured command output. */
|
||
export function firstFailureLine(out) {
|
||
const lines = String(out || "")
|
||
.split("\n")
|
||
.map((l) => l.trim())
|
||
.filter(Boolean);
|
||
const hit = lines.find((l) => /✖|✗|not ok|AssertionError|error TS|FAIL|Error:|REGRESS/i.test(l));
|
||
return (hit || lines[lines.length - 1] || "failed").slice(0, 200);
|
||
}
|
||
|
||
/** Sum {errorCount,warningCount} across an eslint --format json result array. */
|
||
export function eslintCounts(parsed) {
|
||
let errors = 0;
|
||
let warnings = 0;
|
||
for (const f of parsed || []) {
|
||
errors += f.errorCount || 0;
|
||
warnings += f.warningCount || 0;
|
||
}
|
||
return { errors, warnings };
|
||
}
|
||
|
||
/**
|
||
* Parse the eslint JSON array out of mixed stdout (tolerates a leading banner AND trailing
|
||
* non-JSON text, e.g. ESLint 9.x's `--suppressions-location` "unpruned suppressions" stderr
|
||
* sentence glued onto the report when stdout+stderr are concatenated — #7837).
|
||
*/
|
||
export function parseEslintJson(out) {
|
||
const str = String(out || "");
|
||
const start = str.indexOf("[");
|
||
if (start < 0) return null;
|
||
// Fast path: the whole remainder is valid JSON (no trailing text).
|
||
try {
|
||
return JSON.parse(str.slice(start));
|
||
} catch {
|
||
// fall through to bracket-depth scan below
|
||
}
|
||
// Slow path: find the matching closing "]" for the array that starts at `start`, tolerating
|
||
// any non-JSON text appended after it. Depth-tracks brackets while skipping over string
|
||
// literals (so a "]" or "[" inside a message string doesn't miscount).
|
||
let depth = 0;
|
||
let inString = false;
|
||
let escaped = false;
|
||
for (let i = start; i < str.length; i++) {
|
||
const ch = str[i];
|
||
if (inString) {
|
||
if (escaped) {
|
||
escaped = false;
|
||
} else if (ch === "\\") {
|
||
escaped = true;
|
||
} else if (ch === '"') {
|
||
inString = false;
|
||
}
|
||
continue;
|
||
}
|
||
if (ch === '"') {
|
||
inString = true;
|
||
} else if (ch === "[") {
|
||
depth++;
|
||
} else if (ch === "]") {
|
||
depth--;
|
||
if (depth === 0) {
|
||
try {
|
||
return JSON.parse(str.slice(start, i + 1));
|
||
} catch {
|
||
return null;
|
||
}
|
||
}
|
||
}
|
||
}
|
||
return null;
|
||
}
|
||
|
||
/** Pull the cognitive-complexity violation count from the gate's output. */
|
||
export function parseCognitiveCount(out) {
|
||
const s = String(out || "");
|
||
// `check:complexity-ratchets` runs ONE shared ESLint walk and prints BOTH ratchets, with the
|
||
// cyclomatic "N violações" summary emitted FIRST — so a bare `\d+ violações` regex would grab
|
||
// the cyclomatic count. Prefer the unambiguous machine-readable `cognitiveComplexity=N` line
|
||
// (mirrors the cyclomatic `complexity=N` parse used for cycCurrent below).
|
||
const machine = s.match(/(?:^|\n)cognitiveComplexity=(\d+)/);
|
||
if (machine) return Number(machine[1]);
|
||
const m = s.match(/(\d+)\s+(?:function\(s\) exceed|violações|violations)/i);
|
||
return m ? Number(m[1]) : null;
|
||
}
|
||
|
||
/**
|
||
* Drift verdict for a ratchet: a metric that grew past its committed baseline is
|
||
* "drift" (reported, never blocking). `direction:"down"` metrics (warnings,
|
||
* complexity, file-size counts) regress when current > baseline.
|
||
*/
|
||
export function isDrift(current, baseline) {
|
||
if (typeof current !== "number" || typeof baseline !== "number") return false;
|
||
return current > baseline;
|
||
}
|
||
|
||
/** releaseGreen iff there are zero failing HARD checks (drift never blocks). */
|
||
export function computeVerdict(results) {
|
||
const hardFailures = results.filter((r) => r.kind === "hard" && !r.ok);
|
||
const drift = results.filter((r) => r.kind === "drift" && !r.ok);
|
||
return { releaseGreen: hardFailures.length === 0, hardFailures, drift };
|
||
}
|
||
|
||
// ─── --full-ci: reproduce the EXACT ci.yml gate set (P0, v3.8.46 post-mortem) ──
|
||
//
|
||
// WHY: the curated HARD/DRIFT lists above are a hand-maintained SUBSET. The v3.8.46
|
||
// release leaked 11 static/gate base-reds (route-validation:t06, docs-counts --strict,
|
||
// docs-symbols, bundle-size --ratchet, test-masking, …) that the pre-flight never ran
|
||
// because they live only in the ci.yml gate JOBS, not in this script. --full-ci reads
|
||
// ci.yml itself and runs every `npm run check:*` / `npm run lint` from those jobs, so the
|
||
// set stays current as gates are added (no drift between this script and CI). One command
|
||
// → zero CI layers for the whole static category.
|
||
|
||
/** ci.yml jobs whose npm-run gate steps --full-ci reproduces locally. */
|
||
export const FULL_CI_GATE_JOBS = [
|
||
"lint",
|
||
"quality-gate",
|
||
"quality-extended",
|
||
"docs-sync-strict",
|
||
"pr-test-policy",
|
||
];
|
||
|
||
// Gates that cannot run meaningfully in a local working-tree pre-flight:
|
||
// • check:pr-evidence — inspects the open PR body (no PR locally)
|
||
// • check:codeql-ratchet — queries GitHub's code-scanning alerts for the REMOTE main
|
||
// branch (CodeQL Default Setup only analyzes main/PRs→main; a local run reflects
|
||
// post-merge server state the pre-flight can't change). Checked on the release PR.
|
||
export const FULL_CI_SKIP = new Set(["check:pr-evidence", "check:codeql-ratchet"]);
|
||
|
||
// Gates that need a specific env to behave like CI (else they compare against the wrong base).
|
||
export const FULL_CI_ENV = { "check:test-masking": { GITHUB_BASE_REF: "main" } };
|
||
|
||
/**
|
||
* Parse a ci.yml text and return the ordered, de-duplicated list of gate commands to run.
|
||
* Each entry: { id, job, args:["run", <script>, ...("--" + args)], env }.
|
||
* Only `npm run lint` and `npm run check:*` steps are taken (build/install/test-run npm
|
||
* scripts are ignored); a `run: |` block is scanned line-by-line so multi-command steps work.
|
||
* Exported pure (no side effects) so the extraction has a fixture-driven unit test.
|
||
*/
|
||
export function extractCiGates(
|
||
yamlText,
|
||
{ jobs = FULL_CI_GATE_JOBS, skip = FULL_CI_SKIP, envMap = FULL_CI_ENV } = {}
|
||
) {
|
||
const doc = parseYaml(yamlText) || {};
|
||
const gates = [];
|
||
const seen = new Set();
|
||
for (const job of jobs) {
|
||
const steps = doc?.jobs?.[job]?.steps;
|
||
if (!Array.isArray(steps)) continue;
|
||
for (const step of steps) {
|
||
const runStr = typeof step?.run === "string" ? step.run : "";
|
||
if (!runStr) continue;
|
||
for (const rawLine of runStr.split("\n")) {
|
||
const m = rawLine.trim().match(/^npm run (\S+)(?:\s+--\s+(.+?))?\s*$/);
|
||
if (!m) continue;
|
||
const script = m[1];
|
||
if (script !== "lint" && !script.startsWith("check:")) continue; // gates only
|
||
if (skip.has(script) || seen.has(script)) continue; // dedup + skip non-local
|
||
seen.add(script);
|
||
const extra = m[2] ? m[2].split(/\s+/).filter(Boolean) : [];
|
||
gates.push({
|
||
id: script,
|
||
job,
|
||
// preserve the `--` so args reach the script (npm run x -- --ratchet)
|
||
args: extra.length ? ["run", script, "--", ...extra] : ["run", script],
|
||
env: envMap[script],
|
||
});
|
||
}
|
||
}
|
||
}
|
||
return gates;
|
||
}
|
||
|
||
// ─── Orchestration (only when run directly) ─────────────────────────────────
|
||
|
||
/**
|
||
* Map a thrown `execFileSync` error to a {code, out} gate result. Exported as a pure helper
|
||
* so the timeout/hang path has a regression test: a gate that exceeds its ceiling (e.g. the unit
|
||
* suite wedged on an unreleased SQLite handle — see CLAUDE.md "Database Handles in Tests") is
|
||
* killed by `execFileSync` (`err.killed === true`) and MUST surface as a visible non-zero gate,
|
||
* never an infinite block that the release captain mistakes for a hang and kills the pre-flight.
|
||
*/
|
||
export function classifyRunError(err, timeoutMs) {
|
||
if (err && err.killed && timeoutMs) {
|
||
return {
|
||
code: 124,
|
||
out: `gate exceeded its ${Math.round(timeoutMs / 1000)}s ceiling and was killed — treat as a hung/failed gate (e.g. an unreleased DB handle in the unit suite); does NOT pass`,
|
||
};
|
||
}
|
||
return {
|
||
code: typeof err?.status === "number" ? err.status : 1,
|
||
out: `${err?.stdout || ""}${err?.stderr || ""}`,
|
||
};
|
||
}
|
||
|
||
// --hermetic: scrub the live-test trigger vars so the pre-flight behaves like CI
|
||
// (a dev machine with OMNIROUTE_API_KEY set runs 17+ live tests that CI skips —
|
||
// every one a false-positive red against the release branch).
|
||
const HERMETIC_SCRUB = ["OMNIROUTE_API_KEY", "OMNIROUTE_URL"];
|
||
let hermetic = false;
|
||
function buildGateEnv(extra) {
|
||
const env = { ...process.env, FORCE_COLOR: "0", ...(extra || {}) };
|
||
if (hermetic) for (const k of HERMETIC_SCRUB) delete env[k];
|
||
return env;
|
||
}
|
||
|
||
function run(cmd, cmdArgs, opts = {}) {
|
||
try {
|
||
const out = execFileSync(cmd, cmdArgs, {
|
||
cwd: ROOT,
|
||
encoding: "utf8",
|
||
stdio: ["ignore", "pipe", "pipe"],
|
||
maxBuffer: 256 * 1024 * 1024,
|
||
env: buildGateEnv(opts.env),
|
||
// A hard ceiling for the long, silent test suites (execFileSync buffers all output until
|
||
// exit, so they show no progress while running). undefined = no timeout for fast gates.
|
||
...(opts.timeout ? { timeout: opts.timeout } : {}),
|
||
});
|
||
return { code: 0, out };
|
||
} catch (err) {
|
||
return classifyRunError(err, opts.timeout);
|
||
}
|
||
}
|
||
|
||
const execFileAsync = promisify(execFile);
|
||
|
||
// Async twin of run() — same {code, out} contract, so the slow suites (unit /
|
||
// vitest / integration / pack-artifact) can run CONCURRENTLY instead of in
|
||
// series. Sequentially they dominate the pre-flight wall time (~2h in the
|
||
// v3.8.45 run); they are independent processes with per-process DATA_DIR
|
||
// isolation, so overlapping them cuts the pre-flight to ~the slowest single one.
|
||
async function runAsync(cmd, cmdArgs, opts = {}) {
|
||
try {
|
||
const { stdout, stderr } = await execFileAsync(cmd, cmdArgs, {
|
||
cwd: ROOT,
|
||
encoding: "utf8",
|
||
maxBuffer: 256 * 1024 * 1024,
|
||
env: buildGateEnv(opts.env),
|
||
...(opts.timeout ? { timeout: opts.timeout } : {}),
|
||
});
|
||
return { code: 0, out: `${stdout || ""}${stderr || ""}` };
|
||
} catch (err) {
|
||
return classifyRunError(err, opts.timeout);
|
||
}
|
||
}
|
||
|
||
async function main() {
|
||
const args = new Set(process.argv.slice(2));
|
||
const JSON_OUT = args.has("--json");
|
||
const WITH_BUILD = args.has("--with-build");
|
||
const QUICK = args.has("--quick");
|
||
const FULL_CI = args.has("--full-ci");
|
||
hermetic = args.has("--hermetic");
|
||
|
||
const results = [];
|
||
const record = (r) => {
|
||
results.push(r);
|
||
const icon = r.ok ? "✅" : r.kind === "drift" ? "🟡" : "❌";
|
||
process.stderr.write(`${icon} [${r.kind}] ${r.label}${r.detail ? ` — ${r.detail}` : ""}\n`);
|
||
};
|
||
|
||
// Announce a gate BEFORE running it. The long suites (unit/vitest/integration) run silently
|
||
// for many minutes (execFileSync buffers their output until exit), which previously looked like
|
||
// a hang and got the pre-flight killed before it surfaced the unit reds (the v3.8.42 miss).
|
||
const announce = (label) => process.stderr.write(`▶ ${label}…\n`);
|
||
|
||
const hardCmd = (id, label, cmd, cmdArgs, opts) => {
|
||
announce(label);
|
||
const { code, out } = run(cmd, cmdArgs, opts);
|
||
saveGateLog(id, out);
|
||
record({
|
||
id,
|
||
label,
|
||
kind: "hard",
|
||
ok: code === 0,
|
||
detail: code === 0 ? "pass" : firstFailureLine(out),
|
||
});
|
||
};
|
||
|
||
// A ratchet command (check:complexity, check:dead-code, …) exits 1 ONLY on a
|
||
// measured regression and self-skips (exit 0) when its tooling is absent — so a
|
||
// non-zero exit here is drift to rebaseline at release, never a contributor block.
|
||
// ALL checks run regardless of earlier failures (the report is collected, not
|
||
// fail-fast) so one pass surfaces every red instead of revealing them in layers.
|
||
const driftCmd = (id, label, cmd, cmdArgs, okDetail = "within baseline", opts) => {
|
||
announce(label);
|
||
const { code, out } = run(cmd, cmdArgs, opts);
|
||
saveGateLog(id, out);
|
||
record({
|
||
id,
|
||
label,
|
||
kind: "drift",
|
||
ok: code === 0,
|
||
detail: code === 0 ? okDetail : firstFailureLine(out),
|
||
});
|
||
};
|
||
|
||
process.stderr.write("🔎 Release-green validation (current working tree)\n\n");
|
||
|
||
hardCmd("typecheck", "Typecheck (core)", npmCmd, ["run", "typecheck:core"]);
|
||
|
||
// ESLint: ONE pass → errors (hard) + warnings (drift)
|
||
{
|
||
announce("ESLint (errors + warnings — ~5-15min)");
|
||
// Suppressions-aware, matching `npm run lint` (Pacote 4 no-new-warnings): the frozen
|
||
// pre-existing debt in config/quality/eslint-suppressions.json must not count as
|
||
// errors here — only NET-NEW violations are release reds. Timeout raised: a full
|
||
// repo pass takes ~14min alone and this pre-flight often runs alongside test suites.
|
||
const { out } = run(
|
||
"npx",
|
||
[
|
||
"eslint",
|
||
".",
|
||
"--format",
|
||
"json",
|
||
"--suppressions-location",
|
||
"config/quality/eslint-suppressions.json",
|
||
// An "unpruned" suppression means a previously-frozen violation was legitimately
|
||
// fixed — release-time housekeeping (same bucket as ratchet drift), never a
|
||
// contributor-blocking defect. Without this flag ESLint 9.x exits 2 for that
|
||
// reason alone, which used to mask the real `--format json` report (#7837).
|
||
"--pass-on-unpruned-suppressions",
|
||
],
|
||
{ timeout: 30 * 60 * 1000 }
|
||
);
|
||
saveGateLog("lint", out);
|
||
const parsed = parseEslintJson(out);
|
||
if (!parsed) {
|
||
record({
|
||
id: "lint",
|
||
label: "ESLint",
|
||
kind: "hard",
|
||
ok: false,
|
||
detail: "could not parse eslint json",
|
||
});
|
||
} else {
|
||
const { errors, warnings } = eslintCounts(parsed);
|
||
record({
|
||
id: "lint-errors",
|
||
label: "ESLint errors",
|
||
kind: "hard",
|
||
ok: errors === 0,
|
||
detail: `${errors} error(s)`,
|
||
});
|
||
const base = baselineValue("eslintWarnings");
|
||
const over = isDrift(warnings, base);
|
||
record({
|
||
id: "eslint-warnings",
|
||
label: "ESLint warnings (ratchet)",
|
||
kind: "drift",
|
||
ok: !over,
|
||
detail:
|
||
base == null
|
||
? `${warnings} (no baseline)`
|
||
: `${warnings} vs baseline ${base}${over ? ` (+${warnings - base} drift → rebaseline at release)` : ""}`,
|
||
});
|
||
}
|
||
}
|
||
|
||
hardCmd("db-rules", "DB rules", npmCmd, ["run", "check:db-rules"]);
|
||
hardCmd("public-creds", "Public creds", npmCmd, ["run", "check:public-creds"]);
|
||
|
||
// Complexity + cognitive (one ESLint walk; both still recorded as drift)
|
||
{
|
||
announce("Complexity + cognitive ratchets (shared ESLint walk)");
|
||
const { out } = run(npmCmd, ["run", "check:complexity-ratchets"]);
|
||
saveGateLog("complexity-ratchets", out);
|
||
const cogCurrent = parseCognitiveCount(out);
|
||
const cogBase = baselineValue("cognitiveComplexity");
|
||
const cogOver = isDrift(cogCurrent, cogBase);
|
||
const cycMatch = /(?:^|\n)complexity=(\d+)/.exec(out);
|
||
const cycOkMatch = /\[complexity\] OK — (\d+)/.exec(out);
|
||
const cycRegMatch = /\[complexity\] REGRESSÃO — (\d+)/.exec(out);
|
||
const cycCurrent = cycMatch
|
||
? Number(cycMatch[1])
|
||
: cycOkMatch
|
||
? Number(cycOkMatch[1])
|
||
: cycRegMatch
|
||
? Number(cycRegMatch[1])
|
||
: null;
|
||
const cycRegressed = /\[complexity\] REGRESSÃO/.test(out);
|
||
record({
|
||
id: "cognitive-complexity",
|
||
label: "Cognitive complexity (ratchet)",
|
||
kind: "drift",
|
||
ok: !cogOver,
|
||
detail:
|
||
cogCurrent == null
|
||
? "could not parse count"
|
||
: `${cogCurrent} vs baseline ${cogBase}${cogOver ? ` (+${cogCurrent - cogBase} drift → rebaseline at release)` : ""}`,
|
||
});
|
||
record({
|
||
id: "complexity",
|
||
label: "Cyclomatic complexity (ratchet)",
|
||
kind: "drift",
|
||
ok: !cycRegressed,
|
||
detail:
|
||
cycCurrent == null
|
||
? firstFailureLine(out) || "measured via check:complexity-ratchets"
|
||
: `complexity=${cycCurrent} (shared walk with cognitive)${cycRegressed ? " REGRESSED" : ""}`,
|
||
});
|
||
}
|
||
|
||
// file-size (drift)
|
||
{
|
||
const { code, out } = run(npmCmd, ["run", "check:file-size"]);
|
||
record({
|
||
id: "file-size",
|
||
label: "File-size ratchet",
|
||
kind: "drift",
|
||
ok: code === 0,
|
||
detail: code === 0 ? "within frozen caps" : firstFailureLine(out),
|
||
});
|
||
}
|
||
|
||
// test-masking (hard) — a PR-context gate: it only runs on the release PR (PR→main) in CI, so
|
||
// net-assert reductions accrue unseen on release/** and explode on the release PR. Reproduce it
|
||
// here against origin/main so a non-allowlisted reduction surfaces in the pre-flight, not in a
|
||
// ~40-min CI layer (v3.8.43 cost 3 such round-trips). Legitimate reductions get allowlisted in
|
||
// config/quality/test-masking-allowlist.json; tautology/skip/deletion signals are never allowlistable.
|
||
if (!QUICK) {
|
||
announce("Test-masking (weakened-assert guard vs main)");
|
||
// best-effort fetch so the merge-base diff is accurate; ignore fetch failure (offline pre-flight)
|
||
run("git", ["fetch", "--no-tags", "origin", "main", "--depth=200"], { timeout: 60 * 1000 });
|
||
const { code, out } = run(npmCmd, ["run", "check:test-masking"], {
|
||
env: { GITHUB_BASE_REF: "main" },
|
||
});
|
||
saveGateLog("test-masking", out);
|
||
record({
|
||
id: "test-masking",
|
||
label: "Test-masking (weakened-assert guard)",
|
||
kind: "hard",
|
||
ok: code === 0,
|
||
detail: code === 0 ? "no weakening" : firstFailureLine(out),
|
||
});
|
||
}
|
||
|
||
// Remaining quality-gate / quality-extended ratchets that the PR→release
|
||
// fast-gates skip and that historically surfaced — one at a time, because the
|
||
// CI Quality Ratchet job is fail-fast — only on the release PR. Running them all
|
||
// here (drift, never blocking) means a single rebaseline pass at release.
|
||
// complexity recorded above with cognitive (check:complexity-ratchets)
|
||
driftCmd("dead-code", "Dead-code (ratchet)", npmCmd, ["run", "check:dead-code"]);
|
||
driftCmd("type-coverage", "Type coverage (ratchet)", npmCmd, ["run", "check:type-coverage"]);
|
||
driftCmd("compression-budget", "Compression budget (ratchet)", npmCmd, [
|
||
"run",
|
||
"check:compression-budget",
|
||
]);
|
||
driftCmd("openapi-coverage", "OpenAPI route coverage (ratchet)", npmCmd, [
|
||
"run",
|
||
"check:openapi-coverage",
|
||
]);
|
||
driftCmd("workflow-lint", "Workflow lint (zizmor ratchet)", npmCmd, [
|
||
"run",
|
||
"check:workflows",
|
||
"--",
|
||
"--ratchet",
|
||
]);
|
||
driftCmd("codeql-ratchet", "CodeQL alerts (ratchet)", npmCmd, ["run", "check:codeql-ratchet"]);
|
||
|
||
// Docs sync + fabricated-docs (strict) is a real-defect gate (invented env vars /
|
||
// routes, i18n mirror drift) — HARD.
|
||
hardCmd("docs-all", "Docs sync + fabricated-docs (strict)", npmCmd, ["run", "check:docs-all"]);
|
||
|
||
if (!QUICK) {
|
||
// These are the gates that catch inherited base-red tests from cycle PRs (the fast-path
|
||
// PR→release does NOT run unit/vitest/integration per-PR — the v3.8.42 release PR exploded
|
||
// with 15 such reds). They run SILENTLY for many minutes; the announce line above + these
|
||
// hard ceilings keep a long-but-healthy run from being mistaken for a hang (the ceiling also
|
||
// converts a genuine DB-handle hang into a visible failure instead of an infinite block).
|
||
// The slow suites are INDEPENDENT processes (each self-isolates DATA_DIR) with
|
||
// no shared state, so they run CONCURRENTLY — the pre-flight wall time becomes
|
||
// ~the slowest single suite instead of their sum (unit ~25-35min + vitest
|
||
// ~3-8min + integration ~3-10min + pack-artifact ~15min was ~1h serial in the
|
||
// v3.8.45 run). pack-artifact (--with-build) joins the same wave. Integration
|
||
// runs ONLY on the release PR full CI, so a regression here is invisible until
|
||
// release — that is why it is a HARD pre-flight gate.
|
||
const slow = [
|
||
{
|
||
// Raised 45→100min 2026-08-05: a hermetic-env run on the loaded devbox
|
||
// (load 7-26) was still inside invocation 1 of 3 at 76min when killed;
|
||
// contention factor 2-3× was measured against idle windows, and no idle
|
||
// measurement exists yet. The pre-flight's REAL condition is exactly
|
||
// this contended one (unit runs in Promise.all with integration+vitest
|
||
// plus whatever else the devbox carries), and there 45min provably
|
||
// killed a healthy suite and fabricated a false base-red. The ceiling's
|
||
// purpose — turning a genuine hang (stuck SQLite handle = zero progress
|
||
// forever) into a visible failure — survives at 100min.
|
||
// TODO: measure on the idle .113 box and re-tighten to ~1.8× measured.
|
||
id: "unit",
|
||
label: "Unit tests (full suite, CI concurrency — ~30-50min idle, up to ~100min under load)",
|
||
args: ["run", "test:unit:ci"],
|
||
timeout: 100 * 60 * 1000,
|
||
},
|
||
{
|
||
id: "vitest",
|
||
label: "Vitest (MCP / autoCombo / cache — ~3-8min)",
|
||
args: ["run", "test:vitest"],
|
||
timeout: 15 * 60 * 1000,
|
||
},
|
||
{
|
||
// Measured 2026-08-05 on an idle 16-core box: 22m08s hermetic (935 tests,
|
||
// 112 files at --test-concurrency=1, i.e. strictly serial because ~16 of
|
||
// them bind a port or share a DB). The old "~3-10min" estimate was stale by
|
||
// ~3x and the 20min ceiling killed a healthy run. 40min keeps the ceiling's
|
||
// real purpose — turning a genuine hang (unreleased DB handle) into a
|
||
// visible failure — without punishing a long-but-healthy suite.
|
||
id: "integration",
|
||
label: "Integration tests (~20-25min)",
|
||
args: ["run", "test:integration"],
|
||
timeout: 40 * 60 * 1000,
|
||
},
|
||
];
|
||
if (WITH_BUILD) {
|
||
slow.push({
|
||
id: "pack-artifact",
|
||
label: "Package artifact (npm pack policy)",
|
||
args: ["run", "check:pack-artifact"],
|
||
timeout: 20 * 60 * 1000,
|
||
});
|
||
// WS1.2 (#7065 class): boot the REAL packed tarball from a clean install —
|
||
// the runtime gate structure checks cannot provide. Reuses the same dist/ build.
|
||
slow.push({
|
||
id: "pack-boot",
|
||
label: "Tarball boot-smoke (installed CLI serves /health)",
|
||
args: ["run", "check:pack-boot"],
|
||
timeout: 15 * 60 * 1000,
|
||
});
|
||
}
|
||
slow.forEach((g) => announce(`${g.label} [parallel]`));
|
||
const slowResults = await Promise.all(
|
||
slow.map((g) => runAsync(npmCmd, g.args, { timeout: g.timeout }))
|
||
);
|
||
slow.forEach((g, i) => {
|
||
const { code, out } = slowResults[i];
|
||
saveGateLog(g.id, out);
|
||
record({
|
||
id: g.id,
|
||
label: g.label,
|
||
kind: "hard",
|
||
ok: code === 0,
|
||
detail: code === 0 ? "pass" : firstFailureLine(out),
|
||
});
|
||
});
|
||
} else if (WITH_BUILD) {
|
||
// --with-build without the suites (--quick): still verify the package artifact.
|
||
const { code, out } = await runAsync(npmCmd, ["run", "check:pack-artifact"], {
|
||
timeout: 20 * 60 * 1000,
|
||
});
|
||
saveGateLog("pack-artifact", out);
|
||
record({
|
||
id: "pack-artifact",
|
||
label: "Package artifact (npm pack policy)",
|
||
kind: "hard",
|
||
ok: code === 0,
|
||
detail: code === 0 ? "pass" : firstFailureLine(out),
|
||
});
|
||
}
|
||
|
||
// --full-ci: run every static gate declared in ci.yml's gate jobs (superset of the
|
||
// curated HARD list above). Combine with --quick to run ONLY these + drift ratchets
|
||
// (skip the slow suites) — the "1 command, 0 CI layers" pre-flight for the static category.
|
||
if (FULL_CI) {
|
||
let gates = [];
|
||
try {
|
||
gates = extractCiGates(readFileSync(join(ROOT, ".github/workflows/ci.yml"), "utf8"));
|
||
} catch (err) {
|
||
record({
|
||
id: "full-ci-extract",
|
||
label: "full-ci: parse .github/workflows/ci.yml",
|
||
kind: "hard",
|
||
ok: false,
|
||
detail: `could not read/parse ci.yml: ${err?.message || err}`,
|
||
});
|
||
}
|
||
const already = new Set(results.map((r) => r.id));
|
||
process.stderr.write(`\n──── full-ci gates from ci.yml (${gates.length}) ────\n`);
|
||
for (const g of gates) {
|
||
// Skip a gate the curated pass already ran with the same id (avoid double-running lint).
|
||
if (already.has(g.id)) continue;
|
||
const { code, out } = run(npmCmd, g.args, { env: g.env, timeout: 10 * 60 * 1000 });
|
||
saveGateLog(`fullci-${g.id.replace(/[^a-z0-9]+/gi, "-")}`, out);
|
||
record({
|
||
id: g.id,
|
||
label: `ci.yml:${g.job} → npm ${g.args.join(" ")}`,
|
||
kind: "hard",
|
||
ok: code === 0,
|
||
detail: code === 0 ? "pass" : firstFailureLine(out),
|
||
});
|
||
}
|
||
}
|
||
|
||
const { releaseGreen, hardFailures, drift } = computeVerdict(results);
|
||
|
||
process.stderr.write("\n──────── verdict ────────\n");
|
||
process.stderr.write(`HARD failures (block — real defects): ${hardFailures.length}\n`);
|
||
hardFailures.forEach((r) => process.stderr.write(` ❌ ${r.label}: ${r.detail}\n`));
|
||
process.stderr.write(`Ratchet drift (non-blocking — rebaseline at release): ${drift.length}\n`);
|
||
drift.forEach((r) => process.stderr.write(` 🟡 ${r.label}: ${r.detail}\n`));
|
||
process.stderr.write(
|
||
releaseGreen
|
||
? "\n✅ RELEASE-GREEN (no hard failures). Any drift above is rebaselined at release, not a contributor concern.\n"
|
||
: "\n❌ NOT release-green — hard failures must be fixed (in the originating PR branch, via co-authorship).\n"
|
||
);
|
||
|
||
if (JSON_OUT) {
|
||
process.stdout.write(
|
||
JSON.stringify(
|
||
{
|
||
releaseGreen,
|
||
hardFailures: hardFailures.map((r) => ({ id: r.id, label: r.label, detail: r.detail })),
|
||
drift: drift.map((r) => ({ id: r.id, label: r.label, detail: r.detail })),
|
||
checks: results.map((r) => ({ id: r.id, kind: r.kind, ok: r.ok, detail: r.detail })),
|
||
},
|
||
null,
|
||
2
|
||
) + "\n"
|
||
);
|
||
}
|
||
|
||
process.exit(releaseGreen ? 0 : 1);
|
||
}
|
||
|
||
// Run only when invoked directly (so tests can import the pure helpers).
|
||
if (process.argv[1] && fileURLToPath(import.meta.url) === process.argv[1]) {
|
||
main();
|
||
}
|