Files
OmniRoute/docs/ops/RUNNER_BOX.md
Diego Rodrigues de Sa e Souza 5b8d63c094 chore(ops): runner-box janitor + operations runbook (WS3.3) (#7115)
* chore(ops): runner-box janitor script + operations runbook (WS3.3)

Codifies what was manual discipline on the .113 self-hosted pool (two live
incidents on the v3.8.47 release day): 30min cron sweeping stale runner
temp/work dirs (>24h), disk-pressure alert at >=85% (SQLITE_FULL killed shards
mid-run), and the proven 4-runner ceiling on the 16 GB box (8-wide OOM'd jobs;
stopping a busy runner cancels its job — documented). Script smoke-tested live
(disk 82%, 1 active runner, exit 0); bash -n clean.

* docs(ops): reword error-code/bash-env mentions the fabricated-docs env detector misreads

* fix(ops): harden janitor sweep — no symlink follow, -xdev, narrowed patterns (root-cron on world-writable /tmp)
2026-07-14 13:51:06 -03:00

1.6 KiB
Raw Blame History

title
title
Self-Hosted Runner Box Operations

Self-Hosted Runner Box Operations (.113 pool)

The self-hosted pool (self-hosted, omni-release labels) runs on the 16 GB box at 192.168.0.113. Two failure modes recurred on release days and were, until v3.8.49, manual discipline; the janitor script codifies them (WS3.3 of the quality plan):

  1. Orphaned temp/work dirs filling the disk → disk-full SQLite errors mid-job.
  2. >4 concurrent runners → OOM-killed jobs (8-wide killed jobs twice on the v3.8.47 release day; 4-wide is the proven ceiling).

Install the janitor (one-time, on the box)

sudo mkdir -p /opt/omniroute-ops
sudo cp scripts/ops/runner-janitor.sh /opt/omniroute-ops/
sudo chmod +x /opt/omniroute-ops/runner-janitor.sh
( sudo crontab -l 2>/dev/null; echo '*/30 * * * * /opt/omniroute-ops/runner-janitor.sh >> /var/log/runner-janitor.log 2>&1' ) | sudo crontab -

What it does every 30min: sweeps runner temp leftovers older than 24h, alerts at ≥85% root-disk usage, and alerts when more than the runner ceiling (default 4, tunable via the script's own environment) of Runner.Listener processes are up. Alerts land in /var/log/runner-janitor.log with a non-zero exit (grep for ).

Operating rules

  • Ceiling: 4 runners on the 16 GB box. Runners 58 stay STOPPED except for explicit off-peak experiments — never during a release window.
  • Stopping a runner mid-job cancels the job (observed live): systemctl stop only when its runner is idle (Runner.Listener without a Runner.Worker child).
  • The .15 VPS is homologation-only — never runs CI runners.