chore(ops): make the runner janitor act on what it can prove, not advise (#11893)

* chore(ops): make the runner janitor act on what it can prove, not advise

The .113 janitor already knew the rules and had been shouting them into a log
nobody reads: on 2026-08-28 12:00Z it reported "10 listeners > ceiling 8" and
"disk 85%" — for hours — while 6.7 GB of dead-run leftovers sat on the 12 GB
tmpfs (RAM) because its patterns matched neither e2e-build.tar.gz nor
next-build/, its 24 h fuse is a day too long for memory, and its _work/_temp
base (/home/*/actions-runner*) does not exist on this box (runners live under
/opt). Measured while draining the v3.8.50 npm publish (postmortem, Parte III).

What changes:

- idle is PROVEN before removal, with ONE lsof snapshot filtered to the swept
  bases (lsof +D per path walked whole trees and took minutes; 460 candidates
  grepping a re-printed 83k-line string was the other half). 20 s on the box.
  Without lsof the janitor removes nothing and says why (exit 1).
- tmpfs leftovers go after 3 h, disk _work/_temp after 24 h; both overridable.
  Patterns gain next-build* and e2e-build.tar.gz; /opt/actions-runner* is swept.
- zombie builds: a next-build older than 75 min has no job (a real Build step is
  ~26 min). On 2026-08-27 one ran 70 min after GitHub had declared its job lost,
  holding 3.6 GB. KillMode=mixed on the units covers systemctl stop/restart;
  this covers the lost-connection path.
- prunes 48 h-old checkouts under _work of runners whose unit is STOPPED — an
  active runner is never touched.
- alerts on memory PSI (full/avg60) and reports the listener ceiling with an
  omniroute/other breakdown (the box also hosts OmniHeuris and OmniMind).
  Enforcing the ceiling stays an operator decision (label split), not cron's.
- --dry-run prints exactly what it would do and touches nothing; unknown
  arguments are rejected.

Dry-run on the real box: 460 stale omniroute-* test fixtures (930 MB of RAM) it
would reclaim, 0 busy, 0 false "removed" lines, 20 s. The unit suite drives the
script against a fixture tree with every base redirected; the sweep branch runs
where lsof exists (hosted CI images) and the without-lsof contract everywhere.

docs/ops/RUNNER_BOX.md reconciled to the measured box: 31 GB (it said 16), ten
listeners, the 14 GB next-build ceiling, the KillMode drop-in, and the rule that
nothing is cleaned by hand while a runner is busy.

* docs(ops): restore the frontmatter fumadocs requires on RUNNER_BOX.md

Rewriting the page whole dropped its `title:` frontmatter, and docs/ is
compiled into the Next build by fumadocs-mdx — so Build, Fast Production Build
and dast-smoke all died with "[MDX] invalid frontmatter in
docs/ops/RUNNER_BOX.md". Same block as before, verbatim.
This commit is contained in:
Diego Rodrigues de Sa e Souza
2026-08-28 15:44:29 -03:00
committed by GitHub
parent 09de69edc7
commit e71be03398
4 changed files with 403 additions and 53 deletions

View File

@@ -0,0 +1,5 @@
- `scripts/ops/runner-janitor.sh` now proves a path is idle with one `lsof`
snapshot and removes stale leftovers itself (tmpfs after 3 h — it is RAM — disk
after 24 h), kills orphan `next-build` processes, prunes checkouts of stopped
runners, and alerts on memory pressure; `--dry-run` shows exactly what it would
do. `docs/ops/RUNNER_BOX.md` reconciled to the measured box (31 GB, 10 listeners).

View File

@@ -4,32 +4,62 @@ title: Self-Hosted Runner Box Operations
# Self-Hosted Runner Box Operations (.113 pool) # Self-Hosted Runner Box Operations (.113 pool)
The self-hosted pool (`self-hosted, omni-release` labels) runs on the 16 GB box at The self-hosted pool (`self-hosted, omni-release` labels) runs on the **.113** box.
`192.168.0.113`. Two failure modes recurred on release days and were, until v3.8.49, Measured 2026-08-28 (v3.8.50 postmortem, Parte III):
manual discipline; the **janitor script codifies them** (WS3.3 of the quality plan):
1. **Orphaned temp/work dirs** filling the disk → disk-full SQLite errors mid-job. | resource | value | what it means for scheduling |
2. **>4 concurrent runners** → OOM-killed jobs (8-wide killed jobs twice on the | --------- | ---------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
v3.8.47 release day; 4-wide is the proven ceiling). | RAM / CPU | **31 GB / 32 cores** (was 16 GB when this doc was first written) | one `next-build` peaks at **~14 GB** → 2 concurrent heavy builds saturate the box, 3 take it down (2026-08-28 06:42Z: load 56, two jobs lost) |
| swap | 15 GB | it swapped its way through the v3.8.50 publish; pressure shows in `/proc/pressure/memory` |
| `/tmp` | **12 GB tmpfs = RAM** | anything parked there is memory; leftovers are swept after 3 h |
| disk | 188 GB | `_work` checkouts of 8 runners reach ~70 GB with no cap |
| runners | **10 listeners**: 8 OmniRoute + OmniHeuris + OmniMind | all share the memory above |
## Install the janitor (one-time, on the box) ## Install the janitor (one-time, on the box)
```bash ```bash
sudo mkdir -p /opt/omniroute-ops scp scripts/ops/runner-janitor.sh root@192.168.0.113:/opt/omniroute-ops/runner-janitor.sh
sudo cp scripts/ops/runner-janitor.sh /opt/omniroute-ops/ ssh root@192.168.0.113 'chmod +x /opt/omniroute-ops/runner-janitor.sh; apt-get install -y lsof'
sudo chmod +x /opt/omniroute-ops/runner-janitor.sh # cron (root): every 30 min, log to /var/log/runner-janitor.log
( sudo crontab -l 2>/dev/null; echo '*/30 * * * * /opt/omniroute-ops/runner-janitor.sh >> /var/log/runner-janitor.log 2>&1' ) | sudo crontab - */30 * * * * MAX_ACTIVE_RUNNERS=8 /opt/omniroute-ops/runner-janitor.sh >> /var/log/runner-janitor.log 2>&1
``` ```
What it does every 30min: sweeps runner temp leftovers older than 24h, alerts at `lsof` is required: the janitor proves a path is idle with one snapshot of open
≥85% root-disk usage, and alerts when more than the runner ceiling (default 4, tunable files before removing it, and without the tool it removes nothing and says so
via the script's own environment) of `Runner.Listener` processes are up. Alerts land in `/var/log/runner-janitor.log` (exit 1). Try any change with `--dry-run` first — it prints exactly what it would
with a non-zero exit (grep for `⚠`). do and touches nothing.
What it does every run: sweeps our own leftovers (`runner-*`, `omniroute-*`,
`next-build*`, `e2e-build.tar.gz`) after **3 h on tmpfs** and 24 h on disk
`_work/_temp`; kills a `next-build` older than 75 min (no job runs that long — on
2026-08-27 one ran 70 min after GitHub had declared its job lost); prunes 48 h-old
checkouts of runners whose unit is **stopped**; alerts on disk ≥ 85 %, memory PSI
`full/avg60` ≥ 10 %, and more listeners than `MAX_ACTIVE_RUNNERS` (with an
omniroute/other breakdown). Exit 1 = attention needed; read the log.
## Runner units: KillMode
The runner's default `KillMode=process` leaves `Runner.Worker → npm → next-build`
alive when a unit is stopped or restarted — an orphan build keeps eating RAM and
CPU with no job attached. Every OmniRoute unit carries a drop-in
(`/etc/systemd/system/actions.runner.diegosouzapw-OmniRoute.<name>.service.d/10-killmode.conf`)
with `KillMode=mixed`: SIGTERM to the listener first, SIGKILL to the whole cgroup at
`TimeoutStop`. It takes effect on the unit's next restart — restart **one runner at
a time, only when idle**, with the idle check and the restart in the same command.
## Operating rules ## Operating rules
- **Ceiling: 4 runners** on the 16 GB box. Runners 58 stay STOPPED except for - **Heavy-build ceiling: 2 at a time.** The listener ceiling (`MAX_ACTIVE_RUNNERS=8`
explicit off-peak experiments — never during a release window. in cron) is a proxy until jobs are split by label — `omni-build` on 2 runners for
- Stopping a runner mid-job cancels the job (observed live): `systemctl stop` Build/publish/heavy shards, `omni-light` on the rest — which is an operator
only when its runner is idle (`Runner.Listener` without a `Runner.Worker` child). decision, not something cron should enforce by killing listeners.
- **Never clean `/tmp` or `_work` by hand while any runner is busy.** A
check-then-delete with a gap between the two is how a live Build job lost its
`_work` on 2026-08-27. The janitor does the check and the removal in one step;
let it.
- Stopping a runner mid-job cancels the job (observed live): `systemctl stop` only
when its listener has no `Runner.Worker` child — and do it in one command.
- Workflows must not park artefacts in `/tmp` (it is RAM). Download to
`$RUNNER_TEMP` (on disk, per runner) — the 1.3 GB `next-build` artefact took 2732
minutes to land on the tmpfs and 2 minutes to upload from disk.
- The `.15` VPS is homologation-only — never runs CI runners. - The `.15` VPS is homologation-only — never runs CI runners.

View File

@@ -1,53 +1,172 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# runner-janitor — self-hosted runner box hygiene (WS3.3, v3.8.49 quality plan). # runner-janitor — self-hosted runner box hygiene for the .113 pool.
# #
# The .113 runner box has recurring failure modes that until now were manual # Runs from cron every 30 min (see docs/ops/RUNNER_BOX.md). It ACTS on what it
# discipline: orphaned tmpfs/work dirs filling the disk, and >4 concurrent # can prove is safe and ALERTS on what needs an operator decision. Reads of
# runners OOM-killing jobs (16 GB box; incidents on the v3.8.47 release day). # "is this in use?" and the removal happen in the same command, never in two
# Install via cron on the box (see docs/ops/RUNNER_BOX.md): # passes: a check-then-delete with a gap is how a live Build job lost its _work
# */30 * * * * /opt/omniroute-ops/runner-janitor.sh >> /var/log/runner-janitor.log 2>&1 # on 2026-08-27.
# #
# Measured box (2026-08-28): 31 GB RAM, 32 cores, 15 GB swap, /tmp = 12 GB
# tmpfs (RAM!), 188 GB disk. A single `next-build` peaks at ~14 GB, so two
# concurrent heavy builds saturate the box and three take it down (06:42Z that
# day: load 56, two jobs lost). The v3.8.50 postmortem (Parte III) has the numbers.
#
# What it does, in order:
# 1) sweep stale artefacts our tooling leaves behind — tmpfs bases after 3 h
# (they hold RAM), disk _work/_temp bases after 24 h; only names we create,
# only when no process has them open
# 2) kill zombie builds: a `next-build` older than ZOMBIE_BUILD_MAX_MIN has no
# job attached (a real Build step measures ~26 min). On 2026-08-27 one ran
# 70 minutes after GitHub had already declared its job lost, eating 3.6 GB
# and a full core set. KillMode=mixed on the units covers systemctl
# stop/restart; this covers the lost-connection path.
# 3) prune 48 h-old checkouts under _work of runners whose unit is INACTIVE
# (stopped runners cannot be mid-job; active ones are never touched)
# 4) alert: root disk >= DISK_ALERT_PCT, memory PSI full/avg60 >= threshold,
# Runner.Listener count above the ceiling (with a per-project breakdown —
# the box also hosts OmniHeuris and OmniMind runners)
#
# Usage: runner-janitor.sh [--dry-run] [--help]
# Exit codes: 0 healthy · 1 attention needed (printed to stdout for the log). # Exit codes: 0 healthy · 1 attention needed (printed to stdout for the log).
set -euo pipefail set -euo pipefail
MAX_ACTIVE_RUNNERS="${MAX_ACTIVE_RUNNERS:-4}" DRY_RUN=0
DISK_ALERT_PCT="${DISK_ALERT_PCT:-85}" for arg in "$@"; do
WORK_DIR_MAX_AGE_HOURS="${WORK_DIR_MAX_AGE_HOURS:-24}" case "$arg" in
STATUS=0 --dry-run) DRY_RUN=1 ;;
-h|--help)
echo "[janitor] $(date -u +%FT%TZ) start" sed -n '2,32p' "$0" | sed 's/^# \{0,1\}//'
exit 0 ;;
# 1) Sweep stale runner temp/work leftovers (>24h — no legitimate job runs that long). *) echo "unknown argument: $arg" >&2; exit 2 ;;
# Hardened for a root cron on world-writable paths: never follow a symlinked base esac
# (a compromised runner could plant one), -P + -xdev so the sweep cannot traverse
# out of the filesystem, and patterns narrowed to names OUR tooling creates
# (no generic tmp* — unrelated system temp files are out of scope).
for base in /tmp /home/*/actions-runner*/_work/_temp; do
[ -d "$base" ] || continue
[ -L "$base" ] && { echo "[janitor] skip symlinked base: $base"; continue; }
find -P "$base" -xdev -maxdepth 1 \( -name 'runner-*' -o -name 'omniroute-*' \) \
! -type l -mmin +$((WORK_DIR_MAX_AGE_HOURS * 60)) -exec rm -rf {} + 2>/dev/null || true
done done
echo "[janitor] stale temp sweep done"
# 2) Disk pressure — alert loudly before SQLITE_FULL kills jobs mid-run. MAX_ACTIVE_RUNNERS="${MAX_ACTIVE_RUNNERS:-8}"
USAGE=$(df --output=pcent / | tail -1 | tr -dc '0-9') DISK_ALERT_PCT="${DISK_ALERT_PCT:-85}"
if [ "$USAGE" -ge "$DISK_ALERT_PCT" ]; then TMPFS_MAX_AGE_HOURS="${TMPFS_MAX_AGE_HOURS:-3}"
echo "[janitor] ⚠ ROOT DISK ${USAGE}% >= ${DISK_ALERT_PCT}% — clean before the next heavy run" WORK_TEMP_MAX_AGE_HOURS="${WORK_TEMP_MAX_AGE_HOURS:-24}"
STATUS=1 WORK_CHECKOUT_MAX_AGE_HOURS="${WORK_CHECKOUT_MAX_AGE_HOURS:-48}"
ZOMBIE_BUILD_MAX_MIN="${ZOMBIE_BUILD_MAX_MIN:-75}"
ZOMBIE_BUILD_COMM="${ZOMBIE_BUILD_COMM:-next-build}"
PSI_FULL_AVG60_ALERT="${PSI_FULL_AVG60_ALERT:-10}"
# Overridable so the unit test can point everything at a fixture tree.
JANITOR_TMP_BASES="${JANITOR_TMP_BASES-/tmp}"
JANITOR_WORK_TEMP_BASES="${JANITOR_WORK_TEMP_BASES-/opt/actions-runner*/_work/_temp /home/*/actions-runner*/_work/_temp}"
JANITOR_RUNNER_DIRS="${JANITOR_RUNNER_DIRS-/opt/actions-runner*}"
JANITOR_PSI_FILE="${JANITOR_PSI_FILE:-/proc/pressure/memory}"
JANITOR_DF_PATH="${JANITOR_DF_PATH:-/}"
STATUS=0
say() { echo "[janitor] $*"; }
# "Is anything using this?" — ONE snapshot of every open path on the box
# (lsof -Fn), then a prefix match per candidate. `lsof +D <dir>` walks the whole
# tree instead and took minutes on a 5 GB leftover — unusable from cron. An
# absent lsof means "cannot prove idle": the sweep keeps the path and says so.
LSOF_BIN="${JANITOR_LSOF:-lsof}"
have_busy_tools() { command -v "$LSOF_BIN" >/dev/null 2>&1; }
SNAP=""
cleanup() { [ -n "$SNAP" ] && rm -f -- "$SNAP"; }
trap cleanup EXIT
# One lsof for the whole run (~13 s / 83k lines on the box), kept ONLY for the
# bases we sweep — 460 candidates grepping a re-printed 83k-line string was the
# slow part, not lsof itself.
snapshot_open_paths() {
have_busy_tools || return 0
SNAP=$(mktemp) || return 0
local prefixes="" b
for b in $JANITOR_TMP_BASES $JANITOR_WORK_TEMP_BASES; do [ -d "$b" ] && prefixes="$prefixes"$'\n'"$b/"; done
# -F n: one "n<path>" line per open file; -w: no warnings
"$LSOF_BIN" -w -Fn 2>/dev/null | sed -n 's/^n//p' | grep -F -f <(printf '%s' "$prefixes" | sed '/^$/d') > "$SNAP" 2>/dev/null || true
}
is_busy() {
local p="$1"
[ -n "$SNAP" ] && [ -s "$SNAP" ] || return 1
# exact path, or anything beneath it when it is a directory
grep -qxF -- "$p" "$SNAP" && return 0
[ -d "$p" ] && grep -qF -- "$p/" "$SNAP"
}
# sweep <base> <max-age-minutes>: only names our tooling creates, never through
# a symlinked base, never across a filesystem, and remove+check in one step.
sweep() {
local base="$1" max_min="$2" p
[ -d "$base" ] || return 0
[ -L "$base" ] && { say "skip symlinked base: $base"; return 0; }
while IFS= read -r -d '' p; do
if ! have_busy_tools; then say "cannot prove idle (lsof missing — apt install lsof), kept: $p"; STATUS=1; continue; fi
if is_busy "$p"; then say "busy, kept: $p"; continue; fi
if [ "$DRY_RUN" -eq 1 ]; then say "would remove ($(( max_min / 60 ))h+): $p"; else rm -rf -- "$p" && say "removed ($(( max_min / 60 ))h+): $p"; fi
done < <(find -P "$base" -xdev -mindepth 1 -maxdepth 1 \
\( -name 'runner-*' -o -name 'omniroute-*' -o -name 'next-build*' -o -name 'e2e-build.tar.gz' \) \
! -type l -mmin "+$max_min" -print0 2>/dev/null || true)
}
say "$(date -u +%FT%TZ) start${DRY_RUN:+ (dry-run=$DRY_RUN)} busy-tools=$(have_busy_tools && echo ok || echo MISSING)"
# 1) stale artefacts — tmpfs is RAM, so it gets the short fuse
snapshot_open_paths
for base in $JANITOR_TMP_BASES; do sweep "$base" $(( TMPFS_MAX_AGE_HOURS * 60 )); done
for base in $JANITOR_WORK_TEMP_BASES; do sweep "$base" $(( WORK_TEMP_MAX_AGE_HOURS * 60 )); done
say "stale temp sweep done"
# 2) zombie builds
ZOMBIES=0
while read -r pid etimes comm; do
[ -n "${pid:-}" ] || continue
if [ "$etimes" -gt $(( ZOMBIE_BUILD_MAX_MIN * 60 )) ]; then
say "⚠ zombie build pid=$pid comm=$comm age=$(( etimes / 60 ))min > ${ZOMBIE_BUILD_MAX_MIN}min — no job runs this long"
if [ "$DRY_RUN" -eq 1 ]; then say "[dry-run] would: kill -TERM $pid (then -KILL)"; else
kill -TERM "$pid" 2>/dev/null || true; sleep 10
kill -0 "$pid" 2>/dev/null && { kill -KILL "$pid" 2>/dev/null || true; say " needed SIGKILL"; }
fi
ZOMBIES=$(( ZOMBIES + 1 )); STATUS=1
fi
done < <(ps -eo pid=,etimes=,comm= 2>/dev/null | awk -v c="$ZOMBIE_BUILD_COMM" '$3 ~ ("^" c) {print $1, $2, $3}' || true)
say "zombie builds: $ZOMBIES"
# 3) old checkouts of STOPPED runners
for d in $JANITOR_RUNNER_DIRS; do
[ -d "$d" ] && [ -f "$d/.runner" ] || continue
agent=$(grep -o '"agentName": *"[^"]*"' "$d/.runner" 2>/dev/null | sed 's/.*"\([^"]*\)"$/\1/')
[ -n "$agent" ] || continue
unit=$(systemctl list-units --plain --no-legend "actions.runner.*.${agent}.service" 2>/dev/null | awk 'NR==1{print $1}')
[ -n "$unit" ] || continue
if systemctl is-active --quiet "$unit"; then continue; fi
while IFS= read -r -d '' co; do
if [ "$DRY_RUN" -eq 1 ]; then say "would prune checkout of stopped runner $agent: $co"; else rm -rf -- "$co" && say "pruned checkout of stopped runner $agent: $co"; fi
done < <(find -P "$d/_work" -xdev -mindepth 2 -maxdepth 2 -type d -mmin "+$(( WORK_CHECKOUT_MAX_AGE_HOURS * 60 ))" -print0 2>/dev/null || true)
done
# 4a) disk
USAGE=$(df --output=pcent "$JANITOR_DF_PATH" 2>/dev/null | tail -1 | tr -dc '0-9')
if [ "${USAGE:-0}" -ge "$DISK_ALERT_PCT" ]; then
say "⚠ ROOT DISK ${USAGE}% >= ${DISK_ALERT_PCT}% — clean before the next heavy run"; STATUS=1
else else
echo "[janitor] disk ${USAGE}% OK" say "disk ${USAGE:-?}% OK"
fi fi
# 3) Concurrency ceiling — 8-wide OOMed the 16 GB box twice on release day; # 4b) memory pressure (PSI) — the box swapped its way through the v3.8.50 publish
# 4 is the proven ceiling. This CODIFIES the rule that was manual discipline. if [ -r "$JANITOR_PSI_FILE" ]; then
FULL60=$(awk '/^full/ {for(i=1;i<=NF;i++) if ($i ~ /^avg60=/) {sub("avg60=","",$i); print $i}}' "$JANITOR_PSI_FILE" 2>/dev/null || echo "")
if [ -n "$FULL60" ] && awk -v v="$FULL60" -v t="$PSI_FULL_AVG60_ALERT" 'BEGIN{exit !(v+0 >= t+0)}'; then
say "⚠ MEMORY PRESSURE psi full/avg60=${FULL60}% >= ${PSI_FULL_AVG60_ALERT}% — too many heavy jobs at once"; STATUS=1
else
say "memory psi full/avg60=${FULL60:-n/a}% OK"
fi
fi
# 4c) concurrency ceiling — alert with a breakdown; the fix is fewer/labelled
# runners (an operator decision), not killing listeners from cron.
ACTIVE=$(pgrep -fc "Runner.Listener" || true) ACTIVE=$(pgrep -fc "Runner.Listener" || true)
OMNI=$(pgrep -fc "actions-runner-omniroute[^ ]*/bin[^ ]*/Runner.Listener" || true)
if [ "${ACTIVE:-0}" -gt "$MAX_ACTIVE_RUNNERS" ]; then if [ "${ACTIVE:-0}" -gt "$MAX_ACTIVE_RUNNERS" ]; then
echo "[janitor] ${ACTIVE} Runner.Listener processes > ceiling ${MAX_ACTIVE_RUNNERS} — stop the extra runners (systemctl stop actions.runner.<name>)" say "${ACTIVE} Runner.Listener processes (omniroute=${OMNI:-0}, other=$(( ${ACTIVE:-0} - ${OMNI:-0} ))) > ceiling ${MAX_ACTIVE_RUNNERS} — stop idle extras: systemctl stop <unit> only when it has no Runner.Worker child"
STATUS=1 STATUS=1
else else
echo "[janitor] runners active: ${ACTIVE:-0}/${MAX_ACTIVE_RUNNERS} OK" say "runners active: ${ACTIVE:-0}/${MAX_ACTIVE_RUNNERS} (omniroute=${OMNI:-0}) OK"
fi fi
echo "[janitor] done status=$STATUS" say "done status=$STATUS"
exit "$STATUS" exit "$STATUS"

View File

@@ -0,0 +1,196 @@
import { describe, it } from "node:test";
import assert from "node:assert/strict";
import { spawnSync } from "node:child_process";
import {
existsSync,
mkdirSync,
mkdtempSync,
readFileSync,
rmSync,
statSync,
utimesSync,
writeFileSync,
} from "node:fs";
import os from "node:os";
import path from "node:path";
/**
* scripts/ops/runner-janitor.sh runs from cron on the .113 runner box. This
* suite pins its safety contract against a fixture tree — never the real /tmp:
* every base, the runner dirs, the PSI file and the df path are redirected, the
* zombie pattern is set to a name no process has, and the ceilings are lifted
* so the outcome does not depend on the box the test happens to run on.
*/
const ROOT = path.resolve(import.meta.dirname, "..", "..");
const SCRIPT = path.join(ROOT, "scripts", "ops", "runner-janitor.sh");
const HOUR = 3_600_000;
// The sweep needs lsof to PROVE a path is idle (one snapshot of open paths). Hosted CI
// images ship both; a bare devbox may not. Each branch below asserts what must
// hold in that environment — without the tools the contract is "delete nothing,
// say why", which is exactly the behaviour worth pinning.
const HAVE_BUSY_TOOLS =
spawnSync("bash", ["-c", "command -v lsof"], { stdio: "ignore" }).status === 0;
function fixture() {
const base = mkdtempSync(path.join(os.tmpdir(), "janitor-fixture-"));
const old = new Date(Date.now() - 5 * HOUR);
const mk = (name: string, dir: boolean, when: Date | null) => {
const p = path.join(base, name);
if (dir) {
mkdirSync(p);
writeFileSync(path.join(p, "x"), "x");
} else writeFileSync(p, "x");
if (when) utimesSync(p, when, when);
return p;
};
return {
base,
staleTar: mk("e2e-build.tar.gz", false, old), // fixed-name artefact ci.yml/npm-publish leave behind
staleBuild: mk("next-build-abc", true, old),
staleUpgrade: mk("omniroute-install-upgrade-xyz", true, old),
fresh: mk("omniroute-batch-api-fresh", true, null), // in use right now
unrelated: mk("somebody-elses.log", false, old), // not ours — never touched
};
}
function run(args: string[], base: string, extraEnv: Record<string, string> = {}) {
return spawnSync("bash", [SCRIPT, ...args], {
encoding: "utf8",
stdio: ["ignore", "pipe", "pipe"],
env: {
...process.env,
JANITOR_TMP_BASES: base,
JANITOR_WORK_TEMP_BASES: "",
JANITOR_RUNNER_DIRS: path.join(base, "no-runners-here-*"),
JANITOR_PSI_FILE: path.join(base, "no-psi"),
JANITOR_DF_PATH: base,
ZOMBIE_BUILD_COMM: "janitor-test-no-such-process",
MAX_ACTIVE_RUNNERS: "9999",
DISK_ALERT_PCT: "101",
...extraEnv,
},
});
}
describe("runner-janitor.sh", () => {
it("is executable bash with strict mode and prints usage on --help", () => {
assert.ok(existsSync(SCRIPT));
assert.ok(statSync(SCRIPT).mode & 0o111, "must be chmod +x (cron runs it directly)");
const body = readFileSync(SCRIPT, "utf8");
assert.ok(body.startsWith("#!/usr/bin/env bash"));
assert.ok(body.includes("set -euo pipefail"));
const help = run(["--help"], os.tmpdir());
assert.equal(help.status, 0, help.stderr);
assert.match(help.stdout, /--dry-run/);
});
it("without lsof it cannot prove idle, so it deletes nothing and says why (exit 1)", () => {
const f = fixture();
try {
const r = run([], f.base, { JANITOR_LSOF: "/nonexistent/lsof" });
assert.equal(r.status, 1, "a janitor that cannot do its job must show up in the cron log");
assert.match(r.stdout, /busy-tools=MISSING/);
assert.match(
r.stdout,
/cannot prove idle \(lsof missing — apt install lsof\), kept: .*e2e-build\.tar\.gz/
);
for (const p of [f.staleTar, f.staleBuild, f.staleUpgrade, f.fresh, f.unrelated]) {
assert.ok(existsSync(p), `must not delete ${p} when idleness cannot be proven`);
}
} finally {
rmSync(f.base, { recursive: true, force: true });
}
});
it("--dry-run names what it WOULD remove and removes nothing", (t) => {
if (!HAVE_BUSY_TOOLS) return t.skip("lsof absent on this box — sweep branch covered in CI");
const f = fixture();
try {
const r = run(["--dry-run"], f.base);
assert.equal(r.status, 0, r.stderr + r.stdout);
assert.match(r.stdout, /busy-tools=ok/);
for (const p of [f.staleTar, f.staleBuild, f.staleUpgrade]) {
assert.match(
r.stdout,
new RegExp(`would remove \\(3h\\+\\): ${p.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")}`)
);
assert.doesNotMatch(
r.stdout,
new RegExp(`removed \\(3h\\+\\): ${p.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")}`),
"dry-run must never claim it removed something"
);
assert.ok(existsSync(p), `dry-run must not delete ${p}`);
}
assert.doesNotMatch(
r.stdout,
/omniroute-batch-api-fresh/,
"a fresh dir is never a candidate"
);
assert.doesNotMatch(r.stdout, /somebody-elses\.log/, "only names our tooling creates");
assert.match(r.stdout, /zombie builds: 0/);
assert.match(r.stdout, /done status=0/);
} finally {
rmSync(f.base, { recursive: true, force: true });
}
});
it("for real: sweeps the three stale artefacts, keeps the fresh one and the stranger", (t) => {
if (!HAVE_BUSY_TOOLS) return t.skip("lsof absent on this box — sweep branch covered in CI");
const f = fixture();
try {
const r = run([], f.base);
assert.equal(r.status, 0, r.stderr + r.stdout);
assert.ok(!existsSync(f.staleTar), "stale e2e-build.tar.gz must go (it is RAM on tmpfs)");
assert.ok(!existsSync(f.staleBuild), "stale next-build dir must go");
assert.ok(!existsSync(f.staleUpgrade), "stale install-upgrade dir must go");
assert.ok(existsSync(f.fresh), "a fresh dir must survive");
assert.ok(existsSync(f.unrelated), "files we did not create must survive even when old");
} finally {
rmSync(f.base, { recursive: true, force: true });
}
});
it("tmpfs fuse is shorter than the disk fuse (RAM vs disk), both overridable", () => {
const f = fixture();
try {
// With a 6h tmpfs fuse the 5h-old artefacts are NOT stale yet.
const r = run(["--dry-run"], f.base, { TMPFS_MAX_AGE_HOURS: "6" });
assert.doesNotMatch(
r.stdout,
/would remove|removed \(|cannot prove idle/,
"nothing is stale under a 6h fuse, so no candidate is even examined"
);
const body = readFileSync(SCRIPT, "utf8");
assert.match(body, /TMPFS_MAX_AGE_HOURS:-3\}/, "tmpfs default must stay short — it is RAM");
assert.match(body, /WORK_TEMP_MAX_AGE_HOURS:-24\}/);
} finally {
rmSync(f.base, { recursive: true, force: true });
}
});
it("alerts (exit 1) on disk and memory pressure thresholds without touching files", () => {
const f = fixture();
try {
writeFileSync(
path.join(f.base, "psi"),
"some avg10=0.00 avg60=0.00 avg300=0.00 total=1\nfull avg10=0.00 avg60=23.50 avg300=9.00 total=1\n"
);
const r = run(["--dry-run"], f.base, {
JANITOR_PSI_FILE: path.join(f.base, "psi"),
DISK_ALERT_PCT: "0",
});
assert.equal(r.status, 1, "attention needed must be exit 1 for the cron log");
assert.match(r.stdout, /MEMORY PRESSURE psi full\/avg60=23\.50%/);
assert.ok(existsSync(f.fresh) && existsSync(f.unrelated));
assert.match(r.stdout, /ROOT DISK \d+% >= 0%/);
assert.ok(existsSync(f.staleTar), "alerting never deletes");
} finally {
rmSync(f.base, { recursive: true, force: true });
}
});
it("rejects unknown arguments instead of silently running", () => {
const r = run(["--yolo"], os.tmpdir());
assert.equal(r.status, 2);
});
});