--- title: Self-Hosted Runner Box Operations --- # Self-Hosted Runner Box Operations (.113 pool) The self-hosted pool (`self-hosted, omni-release` on all eight runners; `omni-build` on two) runs on the **.113** box. Measured 2026-08-28 (v3.8.50 postmortem, Parte III): | resource | value | what it means for scheduling | | --------- | ---------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | RAM / CPU | **31 GB / 32 cores** (was 16 GB when this doc was first written) | one `next-build` peaks at **~14 GB** → 2 concurrent heavy builds saturate the box, 3 take it down (2026-08-28 06:42Z: load 56, two jobs lost) | | swap | 15 GB | it swapped its way through the v3.8.50 publish; pressure shows in `/proc/pressure/memory` | | `/tmp` | **12 GB tmpfs = RAM** | anything parked there is memory; leftovers are swept after 3 h | | disk | 188 GB | `_work` checkouts of 8 runners reach ~70 GB with no cap | | runners | **10 listeners**: 8 OmniRoute + OmniHeuris + OmniMind | all share the memory above | ## Install the janitor (one-time, on the box) ```bash scp scripts/ops/runner-janitor.sh root@192.168.0.113:/opt/omniroute-ops/runner-janitor.sh ssh root@192.168.0.113 'chmod +x /opt/omniroute-ops/runner-janitor.sh; apt-get install -y lsof' # cron (root): every 30 min, log to /var/log/runner-janitor.log */30 * * * * MAX_ACTIVE_RUNNERS=8 /opt/omniroute-ops/runner-janitor.sh >> /var/log/runner-janitor.log 2>&1 ``` `lsof` is required: the janitor proves a path is idle with one snapshot of open files before removing it, and without the tool it removes nothing and says so (exit 1). Try any change with `--dry-run` first — it prints exactly what it would do and touches nothing. What it does every run: sweeps our own leftovers (`runner-*`, `omniroute-*`, `next-build*`, `e2e-build.tar.gz`) after **3 h on tmpfs** and 24 h on disk `_work/_temp`; kills a `next-build` older than 75 min (no job runs that long — on 2026-08-27 one ran 70 min after GitHub had declared its job lost); prunes 48 h-old checkouts of runners whose unit is **stopped**; alerts on disk ≥ 85 %, memory PSI `full/avg60` ≥ 10 %, and more listeners than `MAX_ACTIVE_RUNNERS` (with an omniroute/other breakdown). Exit 1 = attention needed; read the log. ## Runner units: KillMode The runner's default `KillMode=process` leaves `Runner.Worker → npm → next-build` alive when a unit is stopped or restarted — an orphan build keeps eating RAM and CPU with no job attached. Every OmniRoute unit carries a drop-in (`/etc/systemd/system/actions.runner.diegosouzapw-OmniRoute..service.d/10-killmode.conf`) with `KillMode=mixed`: SIGTERM to the listener first, SIGKILL to the whole cgroup at `TimeoutStop`. It takes effect on the unit's next restart — restart **one runner at a time, only when idle**, with the idle check and the restart in the same command. ## Operating rules - **Heavy-build ceiling: 2 at a time — enforced by label.** Every job that runs a `next build` (`ci.yml` `build`, `npm-publish.yml` `publish`, both `nightly-release-green` validations) targets `[self-hosted, omni-build]`, and only **two** runners carry that label (`omniroute-113-5`, `omniroute-113-6`, added through the runners API — no re-registration). The other six keep `omni-release` and take nothing heavy; GitHub queues a third build instead of the kernel killing one. Pair with the `heavy-build-*` concurrency lanes in `ci.yml`. To add capacity, label another runner — never raise the count past what 31 GB holds (one next-build ≈ 14–16 GB). - **Never clean `/tmp` or `_work` by hand while any runner is busy.** A check-then-delete with a gap between the two is how a live Build job lost its `_work` on 2026-08-27. The janitor does the check and the removal in one step; let it. - Stopping a runner mid-job cancels the job (observed live): `systemctl stop` only when its listener has no `Runner.Worker` child — and do it in one command. - Workflows must not park artefacts in `/tmp` (it is RAM). Download to `$RUNNER_TEMP` (on disk, per runner) — the 1.3 GB `next-build` artefact took 27–32 minutes to land on the tmpfs and 2 minutes to upload from disk. - The `.15` VPS is homologation-only — never runs CI runners.