* fix(ci): stop hosted docker-publish OOM and unpaint Build (advisory) docker-publish was firing 8 concurrent hosted builds on every merge storm; each died ResourceExhausted in npm run build (#11976). One publish per ref, webpack instead of Turbopack so native RSS stays inside the V8 heap we can cap. Build (advisory) is skipped: continue-on-error still reports FAILURE and was painting every fork PR red. Closes #11976 * fix(ci): run docker-publish amd64 on omni-build and share the heavy lane The .113 box is 31 GB / 32 cores — enough for one next-build. Hosted ubuntu-24.04 is ~7 GB and ResourceExhausted every publish (#11976). amd64 now targets [self-hosted, omni-build] (Turbopack) when USE_VPS_RUNNER is on, joins the existing heavy-build-main group so it queues beside ci.yml Build instead of becoming a third heavy, and falls back to hosted + webpack if the VPS is off. arm64 stays on ubuntu-24.04-arm with webpack (no ARM box). * test(ci): align the advisory-build contract with the hosted-OOM skip if: ${{ false }} tripped zizmor obfuscation (194→195). Bare if: false skips the job without a new finding. The #7307 test now pins the skip and keeps the job body as the restore recipe.
6.6 KiB
title
| title |
|---|
| Self-Hosted Runner Box Operations |
Self-Hosted Runner Box Operations (.113 pool)
The self-hosted pool (self-hosted, omni-release on all eight runners; omni-build on two) runs on the .113 box.
Measured 2026-08-28 (v3.8.50 postmortem, Parte III):
| resource | value | what it means for scheduling |
|---|---|---|
| RAM / CPU | 31 GB / 32 cores (was 16 GB when this doc was first written) | one next-build peaks at ~14 GB → 2 concurrent heavy builds saturate the box, 3 take it down (2026-08-28 06:42Z: load 56, two jobs lost) |
| swap | 15 GB | it swapped its way through the v3.8.50 publish; pressure shows in /proc/pressure/memory |
/tmp |
12 GB tmpfs = RAM | anything parked there is memory; leftovers are swept after 3 h |
| disk | 188 GB | _work checkouts of 8 runners reach ~70 GB with no cap |
| runners | 6 listeners: 4 OmniRoute (2 omni-build + 2 omni-light) + OmniHeuris + OmniMind |
all share the memory above; omniroute-113-3/-4/-7/-8 are disabled (systemctl enable --now brings one back) |
Install the janitor (one-time, on the box)
scp scripts/ops/runner-janitor.sh root@192.168.0.113:/opt/omniroute-ops/runner-janitor.sh
ssh root@192.168.0.113 'chmod +x /opt/omniroute-ops/runner-janitor.sh; apt-get install -y lsof'
# cron (root): every 30 min, log to /var/log/runner-janitor.log
*/30 * * * * MAX_ACTIVE_RUNNERS=4 /opt/omniroute-ops/runner-janitor.sh >> /var/log/runner-janitor.log 2>&1
lsof is required: the janitor proves a path is idle with one snapshot of open
files before removing it, and without the tool it removes nothing and says so
(exit 1). Try any change with --dry-run first — it prints exactly what it would
do and touches nothing.
What it does every run: sweeps our own leftovers (runner-*, omniroute-*,
next-build*, e2e-build.tar.gz) after 3 h on tmpfs and 24 h on disk
_work/_temp; kills a next-build older than 75 min (no job runs that long — on
2026-08-27 one ran 70 min after GitHub had declared its job lost); prunes 48 h-old
checkouts of runners whose unit is stopped; alerts on disk ≥ 85 %, memory PSI
full/avg60 ≥ 10 %, and more listeners than MAX_ACTIVE_RUNNERS (with an
omniroute/other breakdown). Exit 1 = attention needed; read the log.
Runner units: KillMode
The runner's default KillMode=process leaves Runner.Worker → npm → next-build
alive when a unit is stopped or restarted — an orphan build keeps eating RAM and
CPU with no job attached. Every OmniRoute unit carries a drop-in
(/etc/systemd/system/actions.runner.diegosouzapw-OmniRoute.<name>.service.d/10-killmode.conf)
with KillMode=mixed: SIGTERM to the listener first, SIGKILL to the whole cgroup at
TimeoutStop. It takes effect on the unit's next restart — restart one runner at
a time, only when idle, with the idle check and the restart in the same command.
Operating rules
- Heavy-build ceiling: 2 at a time — enforced by label. Every job that runs a
full
next buildtargets[self-hosted, omni-build], and only two runners carry that label (omniroute-113-5,omniroute-113-6, added through the runners API — no re-registration):ci.ymlBuildandnpm-publish.ymlpublish- both
nightly-release-greenvalidations docker-publish.ymlamd64 (the hosted 7 GB runner ResourceExhausted this tree — #11976). The arm64 leg stays onubuntu-24.04-arm(no ARM box) with webpack. The other listeners keepomni-release/omni-lightand take nothing heavy; GitHub queues a third build instead of the kernel killing one. Pair with theheavy-build-*concurrency lanes:ci.ymlBuildonmainand docker-publish amd64 shareheavy-build-main(cancel-in-progress: false) so a:nextpublish waits beside the artefact instead of sitting next to it. PR builds useheavy-build-pr. To add capacity, label another runner — never raise the count past what 31 GB holds (one next-build ≈ 14–16 GB; a Docker amd64 build is the same class plus the daemon — still one slot). Docker Engine must be on those two units (docker infois the first step of the publish job). If it is missing, the job fails closed instead of hanging onsetup-buildx.
- Light pool:
omni-light(2026-08-29, #11965).omniroute-113andomniroute-113-2carryomni-lightfor jobs that need a backend-onlynext build(~5–6 GB) but not a full one: the nightly Schemathesis, promptfoo, garak and axe-a11y jobs. They ran on the hosted 7 GB runner and died onrelease/v3.8.51with nobody watching. Worst case on the box is 2 heavy + 2 light ≈ 30 + 12 GB — over 31 GB of RAM, inside the 16 GB of swap; the real fix for headroom is more RAM on the Proxmox VM (tomni-proxmox-113), which turns the label ceilings into 3 heavy + 2 light. - Fewer listeners on purpose. Four OmniRoute units were disabled on 2026-08-29 — with only
ci.ymlBuildand the nightlies using the box, 8 listeners were idle and each extra one is a potential 14 GB tenant. The janitor ceiling is 4 (MAX_ACTIVE_RUNNERS=4in cron). - Never clean
/tmpor_workby hand while any runner is busy. A check-then-delete with a gap between the two is how a live Build job lost its_workon 2026-08-27. The janitor does the check and the removal in one step; let it. - Stopping a runner mid-job cancels the job (observed live):
systemctl stoponly when its listener has noRunner.Workerchild — and do it in one command. - Workflows must not park artefacts in
/tmp(it is RAM). Download to$RUNNER_TEMP(on disk, per runner) — the 1.3 GBnext-buildartefact took 27–32 minutes to land on the tmpfs and 2 minutes to upload from disk. - The
.15VPS is homologation-only — never runs CI runners.