* test(infra): retry recursive temp-dir removal instead of failing a shard on ENOTEMPTY (#11966)
Two shards on release/v3.8.51 went red in one day with the same signature —
"ENOTEMPTY, Directory not empty: /tmp/omniroute-<test>-XXXXXX" — from
combo-same-provider-cascade (Unit Tests fast-path 4/4, on a PR that touches only
.github/) and auth-policy-embeddings-webfetch-7785 (the 20k-test TIA step). Both pass
alone and on re-run: the cleanup races something still writing into the directory
(SQLite WAL/-shm checkpoint, a worker, the backup) and under a loaded hosted runner
the window opens. 1154 test files do their own cleanup with
fs.rmSync(dir, { recursive: true, force: true }); 57 already asked for retries.
One-shot codemod (scripts/ad-hoc/codemod-rm-maxretries.mjs, kept for the record):
every rm / rmSync / rmdirSync option object with `recursive: true` and no
`maxRetries` gains `maxRetries: 5, retryDelay: 100` — Node itself then retries
ENOTEMPTY/EBUSY/EPERM for up to ~0.5 s before giving up. 2243 call sites in 1292
files under tests/, the shared tests/_setup/isolateDataDir.ts exit hook included.
Only the option object changes: no call site, assertion or import is touched.
Validation: prettier and ESLint (with the frozen suppressions) clean on all 1292
files; a random 20-file sample runs green (quota-redis-store hangs identically on
the untouched tree — it needs a Redis on localhost, an environment matter). The
four unit shards on this PR are the full run.
* fix(quality): let check-forgotten-sibling-tests read a 1,000-file diff
The gate shells out to `git diff` through execFileSync with Node's default 1 MB
maxBuffer; the 1,292-file codemod in this PR is the first diff large enough to
overflow it, and the gate died with `spawnSync git ENOBUFS` before comparing
anything. 64 MB is far above any real PR and costs nothing when unused.
Two base-reds on the v3.8.50 tip, found by the release pre-flight.
1. #11355 regressed #10534. It replaced the per-window recovery check with an
unconditional `hasActiveCooldown()` stop, which is right for an
upstream-derived cooldown but also blocks the case #10534 exists for: a
Claude-subscription 429 persists a SYNTHETIC 1h rateLimitedUntil because the
upstream sends no parseable reset. When the later poll shows every governing
window has really reset with quota left, holding that synthetic cooldown just
deadlocks the connection for an hour.
The orphaned `windowStillExhaustedAfterRealReset()` helper and the three
unused claudeExtraUsage imports that ESLint flagged were the fingerprint of
this regression, not dead code: they are the two halves of the original gate.
Re-wired as `isQuotaExhaustedCooldownReleasable()`, deliberately narrow —
only lastErrorType "quota_exhausted" is eligible, one still-exhausted or
unknown-reset window keeps the lock, and an extra-usage POLICY block stays
locked even though its quota windows do look recovered in the same fetch.
#11277/#11355 semantics are untouched (both guards still pass).
Regression guard: tests/unit/provider-limits-recovery.test.ts already pinned
this contract and was red on the tip. 15/15 now.
2. The three volcengine-plan connect routes read `request.json()` and handed the
raw fields to a headless-browser login service after ad-hoc typeof checks
(`check:route-validation:t06`, Hard Rule #7). `String(body.code ?? "")` turned
123 into "123" and an absent code into "", both reaching the service as a
plausible SMS code. Now parsed with Zod schemas, before the session lookup, so
a malformed body answers 400 instead of a misleading 404.
New: tests/unit/volcengine-plan-connect-validation.test.ts (8 cases, red
before the fix). Gate: 687 route files scanned, PASS.
Also drops a genuinely dead import (formatVideoTimestamp in videoBridge.ts —
only used inside the helpers module that defines it).