diff options
| author | Danilo M. <danix@danix.xyz> | 2026-09-22 11:10:35 +0200 |
|---|---|---|
| committer | Danilo M. <danix@danix.xyz> | 2026-09-22 11:10:35 +0200 |
| commit | 48882edb25df1dabb6197edda8ffb092c32731fd (patch) | |
| tree | dde7f70fe28595030eaaa187bb155f03c7c02108 | |
| parent | 88b55ee4d2920f1bfbf9fe75bfe3ceec6cb9ba9b (diff) | |
| download | sbo-dockerbuild-48882edb25df1dabb6197edda8ffb092c32731fd.tar.gz sbo-dockerbuild-48882edb25df1dabb6197edda8ffb092c32731fd.zip | |
release 1.1.2v1.1.2
Also records the host config the chain depends on. The schedule and the
storage layout existed only on the VM, so a rebuilt host would have lost
both, and the README's inline copy of the schedule had already drifted
from what actually runs. crontab.example is byte-identical to the
deployed crontab; fstab.example carries the two-disk layout and the
dockerd mount-namespace trap that makes moving the registry store
non-obvious.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| -rw-r--r-- | ChangeLog.md | 34 | ||||
| -rw-r--r-- | image-builder/README | 52 | ||||
| -rwxr-xr-x | image-builder/bootstrap.sh | 2 | ||||
| -rwxr-xr-x | image-builder/build-full-image.sh | 2 | ||||
| -rwxr-xr-x | image-builder/build-sbo-testbuild.sh | 2 | ||||
| -rw-r--r-- | image-builder/crontab.example | 68 | ||||
| -rw-r--r-- | image-builder/fstab.example | 45 | ||||
| -rw-r--r-- | image-builder/registry-gc.sh | 2 | ||||
| -rwxr-xr-x | install.sh | 2 | ||||
| -rwxr-xr-x | test-build | 2 |
10 files changed, 182 insertions, 29 deletions
diff --git a/ChangeLog.md b/ChangeLog.md index cb817ed..1cdbe3d 100644 --- a/ChangeLog.md +++ b/ChangeLog.md @@ -4,6 +4,40 @@ All notable changes to this project are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/), versioning is [SemVer](https://semver.org/). +## [1.1.2] - 2026-09-22 + +Disk exhaustion on the build host, from three directions. The nightly +`-current` chain had been failing for ten nights with "no space left on +device", and because each stage is gated on the previous one's digest, the +later stages skipped silently: the `-current` tags sat six days stale while +`15.0` rebuilt fine and nothing reported an error. + +### Added +- `image-builder/registry-gc.sh`: reclaim unreferenced blobs from the + registry store, which never shrinks on its own. Stops the registry for a + stable blob graph, refuses to run against OCI-index manifests (distribution + 2.8.x deletes their children), and verifies a tag still resolves afterwards. +- `image-builder/crontab.example` and `image-builder/fstab.example`: reference + copies of the host schedule and storage layout, recorded so a rebuilt VM is + reproducible. They previously existed only on the VM. + +### Changed +- Each build now prunes its own superseded image right after pushing, instead + of leaving it resident until a daily prune. A replaced `sbo-full:current` is + ~33G of dead weight that the next variant had to build around. +- The registry store belongs on a different disk from the docker volume. It + grows with every push while the build needs a large transient peak at a + fixed hour; sharing one volume pits a slow leak against a hard failure. + `fstab.example` documents the layout, including the dockerd mount-namespace + trap when moving an existing store. +- The README points at `crontab.example` rather than repeating the schedule, + which is how the inline copy had gone stale. + +### Fixed +- `bootstrap.sh`, `build-full-image.sh` and `build-sbo-testbuild.sh` abort + early when the docker daemon is unreachable, instead of failing deeper in + with a less obvious error. + ## [1.1.1] - 2026-07-16 ### Fixed diff --git a/image-builder/README b/image-builder/README index 86c463f..5910497 100644 --- a/image-builder/README +++ b/image-builder/README @@ -13,10 +13,19 @@ Three scripts, chained (see docs/specs/2026-07-13-image-builder-design.md): Plus one maintenance script (not part of the chain): registry-gc.sh reclaim unreferenced blobs from the registry store +And two reference copies of the host config the chain depends on. Nothing +reads them; they are here so a rebuilt VM is reproducible: + crontab.example the nightly schedule, as deployed + fstab.example the two-disk storage layout, as deployed + All settings live in ./config. -VM setup (docker.noland.dnx, Slackware x86_64, 4 vCPU / 4 GB / 80 GB) --------------------------------------------------------------------- +VM setup (docker.noland.dnx, Slackware x86_64, 4 vCPU / 4 GB) +------------------------------------------------------------- +Two disks, and which is which matters: a system disk (80 GB, also holding the +registry store) and a separate docker volume (160 GB) for images and build +scratch. See fstab.example. + 1. Install docker; enable the daemon. 2. NFS-mount the two NAS trees read-only, named to match: @@ -25,10 +34,17 @@ VM setup (docker.noland.dnx, Slackware x86_64, 4 vCPU / 4 GB / 80 GB) Each is a full mirror (PACKAGES.TXT, ChangeLog.txt, slackware64/, patches/, extra/). Root must be able to read them (bootstrap runs installpkg as root). -3. Run a LAN registry (storage on the same disk as docker, bind-mounted): +3. Run a LAN registry. Put its storage on a DIFFERENT disk from docker's: docker run -d --restart=always -p 5000:5000 \ -v /opt/sbo-testbuild/registry:/var/lib/registry --name registry registry:2 + The bind target belongs on the system disk, not the docker volume. The + registry grows with every push and never shrinks on its own, while the + build needs a large transient peak at a fixed hour; sharing one volume + pits a slow leak against a hard failure, and the build loses. See + fstab.example for the layout and the dockerd-namespace trap if you move + an existing store. + 4. Mark the registry insecure (plain HTTP) on the VM AND every pulling client (this dev box, the buildsystem VM). In /etc/docker/daemon.json: { "insecure-registries": ["docker.noland.dnx:5000"] } @@ -44,26 +60,16 @@ VM setup (docker.noland.dnx, Slackware x86_64, 4 vCPU / 4 GB / 80 GB) full-image on the base-image digest, build-sbo-testbuild on the full-image digest + tools .txz hash), so an unchanged night is a cheap no-op. -current moves daily and rebuilds most nights; 15.0 is frozen stable and rebuilds only - on a real repo update. Deployed schedule on docker.noland.dnx: - # -current (ready ~04:35) - 0 3 * * * /path/to/sbo-dockerbuild/image-builder/bootstrap.sh --version current >> /var/log/sbo-testbuild.log 2>&1 - 20 3 * * * /path/to/sbo-dockerbuild/image-builder/build-full-image.sh --version current >> /var/log/sbo-testbuild.log 2>&1 - 30 4 * * * /path/to/sbo-dockerbuild/image-builder/build-sbo-testbuild.sh --version current >> /var/log/sbo-testbuild.log 2>&1 - # 15.0 (ready ~06:35) - 0 5 * * * /path/to/sbo-dockerbuild/image-builder/bootstrap.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1 - 20 5 * * * /path/to/sbo-dockerbuild/image-builder/build-full-image.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1 - 30 6 * * * /path/to/sbo-dockerbuild/image-builder/build-sbo-testbuild.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1 - - Post-build cleanup, after the chain (which ends ~06:30) and before the 15:00 - cache prune: - # daily: drop dangling images left behind when a tag moves to a new build - 0 7 * * * docker image prune -f >> /var/log/sbo-testbuild.log 2>&1 - # weekly (Sunday): reclaim unreferenced blobs from the registry store - 0 8 * * 0 /path/to/sbo-dockerbuild/image-builder/registry-gc.sh >> /var/log/sbo-testbuild.log 2>&1 - - The registry never reclaims blobs on its own, so without the weekly GC its - storage grows until the disk fills and the nightly builds fail with - "no space left on device" (see the section below). + on a real repo update. + + The schedule lives in crontab.example, which is a copy of what the VM runs: + crontab crontab.example # or paste it into `crontab -e` + + Install it rather than retyping it. The timings are load-bearing, not + cosmetic: reclaim runs at 02:50, immediately before the 03:00 chain, so the + headroom exists when the build needs it. An earlier schedule pruned in the + afternoon instead and the -current build failed ten nights running with + "no space left on device". The file explains each window. 7. Ensure docker.noland.dnx resolves on the LAN (static IP or DNS). diff --git a/image-builder/bootstrap.sh b/image-builder/bootstrap.sh index e632e7a..7bf87ac 100755 --- a/image-builder/bootstrap.sh +++ b/image-builder/bootstrap.sh @@ -14,7 +14,7 @@ # bootstrap.sh — build the sbo-base:{ver} image FROM scratch from NAS trees. # Run as root (installpkg). Adapted from forge slackware/docker-images. set -euo pipefail -PROJECT_VERSION="1.1.1" # bump via sed across all scripts; see CLAUDE.md Releases +PROJECT_VERSION="1.1.2" # bump via sed across all scripts; see CLAUDE.md Releases HERE="$(cd "$(dirname "$0")" && pwd)" source "${HERE}/config" LOG_TAG=bootstrap diff --git a/image-builder/build-full-image.sh b/image-builder/build-full-image.sh index 6ca0288..47e034b 100755 --- a/image-builder/build-full-image.sh +++ b/image-builder/build-full-image.sh @@ -14,7 +14,7 @@ # build-full-image.sh — sbo-full:{ver} FROM sbo-base:{ver}, all series. # Adapted from forge slackware/docker-images. No root needed. set -euo pipefail -PROJECT_VERSION="1.1.1" # bump via sed across all scripts; see CLAUDE.md Releases +PROJECT_VERSION="1.1.2" # bump via sed across all scripts; see CLAUDE.md Releases HERE="$(cd "$(dirname "$0")" && pwd)" source "${HERE}/config" LOG_TAG=full diff --git a/image-builder/build-sbo-testbuild.sh b/image-builder/build-sbo-testbuild.sh index d40ccb1..2bb8487 100755 --- a/image-builder/build-sbo-testbuild.sh +++ b/image-builder/build-sbo-testbuild.sh @@ -15,7 +15,7 @@ # Adds sbopkg + sbo-maintainer-tools from prebuilt .txz in PKGDIR. # Rebuilds when the full image OR the .txz set changes. set -euo pipefail -PROJECT_VERSION="1.1.1" # bump via sed across all scripts; see CLAUDE.md Releases +PROJECT_VERSION="1.1.2" # bump via sed across all scripts; see CLAUDE.md Releases HERE="$(cd "$(dirname "$0")" && pwd)" source "${HERE}/config" LOG_TAG=testbuild diff --git a/image-builder/crontab.example b/image-builder/crontab.example new file mode 100644 index 0000000..174ace0 --- /dev/null +++ b/image-builder/crontab.example @@ -0,0 +1,68 @@ +# sbo-testbuild image chain, root's crontab on the docker host. +# +# Copyright (C) 2026 Danilo M. <danix@danix.xyz> +# GPLv2 only; see LICENSE. +# +# Reference copy of what the VM actually runs. Nothing reads this file: install +# it with `crontab -` (or paste into `crontab -e`) and keep the two in step. +# It is recorded here because the schedule is as much a part of the build +# system as the scripts, and it used to live only on the VM, where a rebuilt +# host would have lost it. +# +# Timings are not arbitrary. The chain is gated stage to stage, so each stage +# must finish before the next starts, and every reclaim must land outside a +# build window or it strips the working set mid-export. +# +# 02:50 reclaim dangling images, then build cache to a 5G floor +# 03:00 -current bootstrap -> full -> testbuild +# 05:00 15.0 bootstrap -> full -> testbuild +# 07:00 reclaim dangling images (catches both variants) +# 08:00 registry blob GC, Sundays only +# +# Repos sync at 01:00/02:00, so the chain starts after that and the images are +# ready by 09:00. + +# --------------------------------------------------------------------------- +# Pre-build reclaim +# --------------------------------------------------------------------------- +# Exporting the -current full image needs roughly its own size (~33G) in +# transient space on top of what is already resident. An afternoon cache prune +# with a 20G floor left the volume short by 03:20, and build-full-image.sh +# failed ten nights running (2026-09-13 to 09-22) with "no space left on +# device", always in the same export phase. build-sbo-testbuild.sh then saw an +# unchanged parent and skipped silently, so the -current tags sat six days +# stale while 15.0 rebuilt fine. +# +# Reclaiming just before the build, not hours after it, is what makes the +# headroom exist when it is needed. Images first (debris from a previous +# failure), then the cache down to a 5G floor. +50 2 * * * docker image prune -f >> /var/log/sbo-testbuild.log 2>&1 +55 2 * * * docker builder prune -f --reserved-space 5g >> /var/log/sbo-testbuild.log 2>&1 + +# --------------------------------------------------------------------------- +# Build chain +# --------------------------------------------------------------------------- +# No --force: each script self-gates (bootstrap=ChangeLog hash, full=base +# digest, testbuild=full digest + .txz hash), so an unchanged night is a cheap +# no-op that exits in seconds. +# +# -current (moves daily): +0 3 * * * /opt/sbo-testbuild/image-builder/bootstrap.sh --version current >> /var/log/sbo-testbuild.log 2>&1 +20 3 * * * /opt/sbo-testbuild/image-builder/build-full-image.sh --version current >> /var/log/sbo-testbuild.log 2>&1 +30 4 * * * /opt/sbo-testbuild/image-builder/build-sbo-testbuild.sh --version current >> /var/log/sbo-testbuild.log 2>&1 +# 15.0 (frozen stable; rebuilds only on a real repo update): +0 5 * * * /opt/sbo-testbuild/image-builder/bootstrap.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1 +20 5 * * * /opt/sbo-testbuild/image-builder/build-full-image.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1 +30 6 * * * /opt/sbo-testbuild/image-builder/build-sbo-testbuild.sh --version 15.0 >> /var/log/sbo-testbuild.log 2>&1 + +# --------------------------------------------------------------------------- +# Post-build cleanup +# --------------------------------------------------------------------------- +# Since 1.1.2 each build prunes its own superseded image right after pushing, +# so this is a backstop for anything those missed (a failed run, a manual +# build). Cheap when there is nothing to do. +0 7 * * * docker image prune -f >> /var/log/sbo-testbuild.log 2>&1 +# The registry never reclaims on its own: every push adds blobs and nothing +# removes them, so its store grows until the disk fills. Weekly is enough. +# registry-gc.sh has its own safety gates; see the script. +0 8 * * 0 /opt/sbo-testbuild/image-builder/registry-gc.sh >> /var/log/sbo-testbuild.log 2>&1 diff --git a/image-builder/fstab.example b/image-builder/fstab.example new file mode 100644 index 0000000..79c9473 --- /dev/null +++ b/image-builder/fstab.example @@ -0,0 +1,45 @@ +# sbo-testbuild storage layout, /etc/fstab on the docker host. +# +# Copyright (C) 2026 Danilo M. <danix@danix.xyz> +# GPLv2 only; see LICENSE. +# +# Reference copy of the two entries this build system depends on. Nothing +# reads this file; it is recorded so the layout survives a lost VM, because +# where these live is not a detail. Getting it wrong is what broke the +# nightly builds for ten days. +# +# The UUID below is this host's; use your own (`blkid /dev/sdb1`). + +# --------------------------------------------------------------------------- +# Docker data: its own disk +# --------------------------------------------------------------------------- +# Build scratch space wants room to breathe. Exporting the -current full image +# writes roughly its own size (~33G) transiently, on top of the images already +# resident, so this volume is sized for the peak and not the steady state. +# Everything on it is reproducible from the mirror, so it is excluded from +# vzdump. +UUID=0e5d008d-0ee0-41bc-9222-aedba1e088d0 /var/lib/docker ext4 defaults 0 2 + +# --------------------------------------------------------------------------- +# Registry store: NOT on the docker disk +# --------------------------------------------------------------------------- +# The registry's blob store used to be a bind mount from the docker disk +# (/var/lib/docker/registry-data). That put ~25G of permanently-resident data +# on the same volume as the transient build peak, and the two competed: the +# registry grows with every push, the build needs headroom at 03:20, and the +# build lost. Moved to the system disk on 2026-09-22, which had 49G idle. +# +# Keep them separate. The registry is small, static and I/O-light; the build +# disk is large, churning and latency-insensitive. Sharing one volume couples +# a slow leak to a hard failure. +# +# If this host's backups cover the system disk, exclude /opt/registry-data: +# the contents are reproducible by re-pushing the images. +# +# Moving it is not just an fstab edit. dockerd caches the mount in its own +# namespace, so after remounting you must restart the daemon and recreate the +# registry container, or pushes keep silently landing on the old disk. Verify +# with: grep ' /var/lib/registry ' /proc/$(docker inspect registry \ +# --format '{{.State.Pid}}')/mountinfo +# and check the device is the system disk, not the docker one. +/opt/registry-data /opt/sbo-testbuild/registry none bind 0 0 diff --git a/image-builder/registry-gc.sh b/image-builder/registry-gc.sh index fab37bb..ef7a6ab 100644 --- a/image-builder/registry-gc.sh +++ b/image-builder/registry-gc.sh @@ -26,7 +26,7 @@ # Intended as a weekly cron job, after the build chain and before the daily # cache prune. See README. set -euo pipefail -PROJECT_VERSION="1.1.1" # bump via sed across all scripts; see CLAUDE.md Releases +PROJECT_VERSION="1.1.2" # bump via sed across all scripts; see CLAUDE.md Releases HERE="$(cd "$(dirname "$0")" && pwd)" source "${HERE}/config" LOG_TAG=registry-gc @@ -22,7 +22,7 @@ # ./install.sh --uninstall remove it set -euo pipefail -PROJECT_VERSION="1.1.1" # bump via sed across all scripts; see CLAUDE.md Releases +PROJECT_VERSION="1.1.2" # bump via sed across all scripts; see CLAUDE.md Releases HERE="$(cd "$(dirname "$0")" && pwd)" BINDIR="${BINDIR:-$HOME/bin}" PROG=test-build @@ -23,7 +23,7 @@ # # No em dashes in prose by author convention. -PROJECT_VERSION="1.1.1" # bump via sed across all scripts; see CLAUDE.md Releases +PROJECT_VERSION="1.1.2" # bump via sed across all scripts; see CLAUDE.md Releases # ============================================================================= # CONFIG (do not edit here; real values live in the external config file) |
