Four nightly jobs run a backend-only `next build` on ubuntu-latest (7 GB): Schemathesis, promptfoo injection guard, garak probes and the axe a11y suite (self-building webServer). On release/v3.8.51 three of them died with the hosted VM shutdown signature and nobody saw it — nightlies have no audience — and the fourth passes by a margin of minutes. They now target [self-hosted, omni-light] (hosted fallback when USE_VPS_RUNNER is off), a new two-listener label on the .113 box for jobs that need ~6 GB, not the 14-16 GB of a full build; they run once a day in the 04:00-06:00 UTC window, when the box is idle. Fleet reshaped the same day and documented in docs/ops/RUNNER_BOX.md: 4 active OmniRoute listeners (omniroute-113-5/-6 omni-build, omniroute-113/-2 omni-light), omniroute-113-3/-4/-7/-8 disabled (systemctl enable --now brings one back), janitor ceiling MAX_ACTIVE_RUNNERS=4. The remaining headroom limit is the VM's 31 GB of RAM (2 heavy + 2 light ≈ 42 GB peak, inside the 16 GB swap); more RAM on the Proxmox VM is the lever that turns the label ceilings into 3 heavy + 2 light. check:workflows --ratchet unchanged (194/194); check-workflows and backend-only-smoke-workflows suites pass; docs-sync PASS.
6.0 KiB
title
| title |
|---|
| Self-Hosted Runner Box Operations |
Self-Hosted Runner Box Operations (.113 pool)
The self-hosted pool (self-hosted, omni-release on all eight runners; omni-build on two) runs on the .113 box.
Measured 2026-08-28 (v3.8.50 postmortem, Parte III):
| resource | value | what it means for scheduling |
|---|---|---|
| RAM / CPU | 31 GB / 32 cores (was 16 GB when this doc was first written) | one next-build peaks at ~14 GB → 2 concurrent heavy builds saturate the box, 3 take it down (2026-08-28 06:42Z: load 56, two jobs lost) |
| swap | 15 GB | it swapped its way through the v3.8.50 publish; pressure shows in /proc/pressure/memory |
/tmp |
12 GB tmpfs = RAM | anything parked there is memory; leftovers are swept after 3 h |
| disk | 188 GB | _work checkouts of 8 runners reach ~70 GB with no cap |
| runners | 6 listeners: 4 OmniRoute (2 omni-build + 2 omni-light) + OmniHeuris + OmniMind |
all share the memory above; omniroute-113-3/-4/-7/-8 are disabled (systemctl enable --now brings one back) |
Install the janitor (one-time, on the box)
scp scripts/ops/runner-janitor.sh root@192.168.0.113:/opt/omniroute-ops/runner-janitor.sh
ssh root@192.168.0.113 'chmod +x /opt/omniroute-ops/runner-janitor.sh; apt-get install -y lsof'
# cron (root): every 30 min, log to /var/log/runner-janitor.log
*/30 * * * * MAX_ACTIVE_RUNNERS=4 /opt/omniroute-ops/runner-janitor.sh >> /var/log/runner-janitor.log 2>&1
lsof is required: the janitor proves a path is idle with one snapshot of open
files before removing it, and without the tool it removes nothing and says so
(exit 1). Try any change with --dry-run first — it prints exactly what it would
do and touches nothing.
What it does every run: sweeps our own leftovers (runner-*, omniroute-*,
next-build*, e2e-build.tar.gz) after 3 h on tmpfs and 24 h on disk
_work/_temp; kills a next-build older than 75 min (no job runs that long — on
2026-08-27 one ran 70 min after GitHub had declared its job lost); prunes 48 h-old
checkouts of runners whose unit is stopped; alerts on disk ≥ 85 %, memory PSI
full/avg60 ≥ 10 %, and more listeners than MAX_ACTIVE_RUNNERS (with an
omniroute/other breakdown). Exit 1 = attention needed; read the log.
Runner units: KillMode
The runner's default KillMode=process leaves Runner.Worker → npm → next-build
alive when a unit is stopped or restarted — an orphan build keeps eating RAM and
CPU with no job attached. Every OmniRoute unit carries a drop-in
(/etc/systemd/system/actions.runner.diegosouzapw-OmniRoute.<name>.service.d/10-killmode.conf)
with KillMode=mixed: SIGTERM to the listener first, SIGKILL to the whole cgroup at
TimeoutStop. It takes effect on the unit's next restart — restart one runner at
a time, only when idle, with the idle check and the restart in the same command.
Operating rules
- Heavy-build ceiling: 2 at a time — enforced by label. Every job that runs a
next build(ci.ymlbuild,npm-publish.ymlpublish, bothnightly-release-greenvalidations) targets[self-hosted, omni-build], and only two runners carry that label (omniroute-113-5,omniroute-113-6, added through the runners API — no re-registration). The other six keepomni-releaseand take nothing heavy; GitHub queues a third build instead of the kernel killing one. Pair with theheavy-build-*concurrency lanes inci.yml. To add capacity, label another runner — never raise the count past what 31 GB holds (one next-build ≈ 14–16 GB). - Light pool:
omni-light(2026-08-29, #11965).omniroute-113andomniroute-113-2carryomni-lightfor jobs that need a backend-onlynext build(~5–6 GB) but not a full one: the nightly Schemathesis, promptfoo, garak and axe-a11y jobs. They ran on the hosted 7 GB runner and died onrelease/v3.8.51with nobody watching. Worst case on the box is 2 heavy + 2 light ≈ 30 + 12 GB — over 31 GB of RAM, inside the 16 GB of swap; the real fix for headroom is more RAM on the Proxmox VM (tomni-proxmox-113), which turns the label ceilings into 3 heavy + 2 light. - Fewer listeners on purpose. Four OmniRoute units were disabled on 2026-08-29 — with only
ci.ymlBuildand the nightlies using the box, 8 listeners were idle and each extra one is a potential 14 GB tenant. The janitor ceiling is 4 (MAX_ACTIVE_RUNNERS=4in cron). - Never clean
/tmpor_workby hand while any runner is busy. A check-then-delete with a gap between the two is how a live Build job lost its_workon 2026-08-27. The janitor does the check and the removal in one step; let it. - Stopping a runner mid-job cancels the job (observed live):
systemctl stoponly when its listener has noRunner.Workerchild — and do it in one command. - Workflows must not park artefacts in
/tmp(it is RAM). Download to$RUNNER_TEMP(on disk, per runner) — the 1.3 GBnext-buildartefact took 27–32 minutes to land on the tmpfs and 2 minutes to upload from disk. - The
.15VPS is homologation-only — never runs CI runners.