1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
|
# sbo-testbuild image chain, root's crontab on the docker host.
#
# Copyright (C) 2026 Danilo M. <danix@danix.xyz>
# GPLv2 only; see LICENSE.
#
# Reference copy of what the VM actually runs. Nothing reads this file: install
# it with `crontab -` (or paste into `crontab -e`) and keep the two in step.
# It is recorded here because the schedule is as much a part of the build
# system as the scripts, and it used to live only on the VM, where a rebuilt
# host would have lost it.
#
# Timings are not arbitrary. The chain is gated stage to stage, so each stage
# must finish before the next starts, and every reclaim must land outside a
# build window or it strips the working set mid-export.
#
# 02:50 reclaim dangling images, then build cache to a 5G floor
# 03:00 -current bootstrap -> full -> testbuild
# 05:00 15.0 bootstrap -> full -> testbuild
# 07:00 reclaim dangling images (catches both variants)
# 07:30 registry blob GC, daily
# 09:00 staleness alert if any tag stopped moving
#
# Repos sync at 01:00/02:00, so the chain starts after that and the images are
# ready by 09:00.
#
# Every build runs under `notify.sh run`, which posts to Gotify on a non-zero
# exit and passes the status through. That is necessary but not sufficient:
# see the staleness check at the bottom for why.
# ---------------------------------------------------------------------------
# Pre-build reclaim
# ---------------------------------------------------------------------------
# Exporting the -current full image needs roughly its own size (~33G) in
# transient space on top of what is already resident. An afternoon cache prune
# with a 20G floor left the volume short by 03:20, and build-full-image.sh
# failed ten nights running (2026-09-13 to 09-22) with "no space left on
# device", always in the same export phase. build-sbo-testbuild.sh then saw an
# unchanged parent and skipped silently, so the -current tags sat six days
# stale while 15.0 rebuilt fine.
#
# Reclaiming just before the build, not hours after it, is what makes the
# headroom exist when it is needed. Images first (debris from a previous
# failure), then the cache down to a 5G floor.
50 2 * * * docker image prune -f >> /var/log/sbo-testbuild.log 2>&1
55 2 * * * docker builder prune -f --reserved-space 5g >> /var/log/sbo-testbuild.log 2>&1
# ---------------------------------------------------------------------------
# Build chain
# ---------------------------------------------------------------------------
# No --force: each script self-gates (bootstrap=ChangeLog hash, full=base
# digest, testbuild=full digest + .txz hash), so an unchanged night is a cheap
# no-op that exits in seconds.
#
# -current (moves daily):
0 3 * * * /opt/sbo-testbuild/image-builder/notify.sh run "bootstrap current" /opt/sbo-testbuild/image-builder/bootstrap.sh --version current >> /var/log/sbo-testbuild.log 2>&1
20 3 * * * /opt/sbo-testbuild/image-builder/notify.sh run "full current" /opt/sbo-testbuild/image-builder/build-full-image.sh --version current >> /var/log/sbo-testbuild.log 2>&1
30 4 * * * /opt/sbo-testbuild/image-builder/notify.sh run "testbuild current" /opt/sbo-testbuild/image-builder/build-sbo-testbuild.sh --version current >> /var/log/sbo-testbuild.log 2>&1
# 15.0 (frozen stable; rebuilds only on a real repo update):
0 5 * * * /opt/sbo-testbuild/image-builder/notify.sh run "bootstrap 15.0" /opt/sbo-testbuild/image-builder/bootstrap.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1
20 5 * * * /opt/sbo-testbuild/image-builder/notify.sh run "full 15.0" /opt/sbo-testbuild/image-builder/build-full-image.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1
30 6 * * * /opt/sbo-testbuild/image-builder/notify.sh run "testbuild 15.0" /opt/sbo-testbuild/image-builder/build-sbo-testbuild.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1
# ---------------------------------------------------------------------------
# Post-build cleanup
# ---------------------------------------------------------------------------
# Since 1.1.2 each build prunes its own superseded image right after pushing,
# so this is a backstop for anything those missed (a failed run, a manual
# build). Cheap when there is nothing to do.
0 7 * * * docker image prune -f >> /var/log/sbo-testbuild.log 2>&1
# The registry never reclaims on its own: every push adds blobs and nothing
# removes them, so its store grows until the disk fills. Weekly was enough
# while the store lived on sdb1; on sda2 (since 2026-09-22) it has ~55G and
# grew 13G -> 53G in one week, filling / and failing every build. Daily.
# registry-gc.sh has its own safety gates; see the script.
30 7 * * * /opt/sbo-testbuild/image-builder/notify.sh run "registry GC" /opt/sbo-testbuild/image-builder/registry-gc.sh >> /var/log/sbo-testbuild.log 2>&1
# ---------------------------------------------------------------------------
# Staleness check
# ---------------------------------------------------------------------------
# The exit-status alerts above would not have caught the September 2026
# outage on their own. build-full-image.sh failed for ten nights and did
# report non-zero, but build-sbo-testbuild.sh exited 0 every single night:
# it saw an unchanged parent digest and skipped, which is correct. So after
# the first alert the chain went quiet while its tags aged six days.
#
# This asks the registry a different question: not "did anything error" but
# "is anything still current". It catches a skipped stage, a stopped cron and
# a wedged mirror alike. Runs after the chain has had its chance.
0 9 * * * /opt/sbo-testbuild/image-builder/notify.sh stale >> /var/log/sbo-testbuild.log 2>&1
|