mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-07-26 09:52:11 +03:00
* chore(ops): runner-box janitor script + operations runbook (WS3.3) Codifies what was manual discipline on the .113 self-hosted pool (two live incidents on the v3.8.47 release day): 30min cron sweeping stale runner temp/work dirs (>24h), disk-pressure alert at >=85% (SQLITE_FULL killed shards mid-run), and the proven 4-runner ceiling on the 16 GB box (8-wide OOM'd jobs; stopping a busy runner cancels its job — documented). Script smoke-tested live (disk 82%, 1 active runner, exit 0); bash -n clean. * docs(ops): reword error-code/bash-env mentions the fabricated-docs env detector misreads * fix(ops): harden janitor sweep — no symlink follow, -xdev, narrowed patterns (root-cron on world-writable /tmp)
1.6 KiB
1.6 KiB
title
| title |
|---|
| Self-Hosted Runner Box Operations |
Self-Hosted Runner Box Operations (.113 pool)
The self-hosted pool (self-hosted, omni-release labels) runs on the 16 GB box at
192.168.0.113. Two failure modes recurred on release days and were, until v3.8.49,
manual discipline; the janitor script codifies them (WS3.3 of the quality plan):
- Orphaned temp/work dirs filling the disk → disk-full SQLite errors mid-job.
- >4 concurrent runners → OOM-killed jobs (8-wide killed jobs twice on the v3.8.47 release day; 4-wide is the proven ceiling).
Install the janitor (one-time, on the box)
sudo mkdir -p /opt/omniroute-ops
sudo cp scripts/ops/runner-janitor.sh /opt/omniroute-ops/
sudo chmod +x /opt/omniroute-ops/runner-janitor.sh
( sudo crontab -l 2>/dev/null; echo '*/30 * * * * /opt/omniroute-ops/runner-janitor.sh >> /var/log/runner-janitor.log 2>&1' ) | sudo crontab -
What it does every 30min: sweeps runner temp leftovers older than 24h, alerts at
≥85% root-disk usage, and alerts when more than the runner ceiling (default 4, tunable
via the script's own environment) of Runner.Listener processes are up. Alerts land in /var/log/runner-janitor.log
with a non-zero exit (grep for ⚠).
Operating rules
- Ceiling: 4 runners on the 16 GB box. Runners 5–8 stay STOPPED except for explicit off-peak experiments — never during a release window.
- Stopping a runner mid-job cancels the job (observed live):
systemctl stoponly when its runner is idle (Runner.Listenerwithout aRunner.Workerchild). - The
.15VPS is homologation-only — never runs CI runners.