| Age | Commit message (Collapse) | Author | Files | Lines |
|
On sda2 the registry grew 13G -> 53G between weekly GCs and filled /.
That broke /tmp, the build log, buildx state and the GC itself, so every
build failed from 2026-10-04. Run registry-gc.sh daily at 07:30 instead
of Sundays only.
The store has since moved to a dedicated disk with backup=0, out of
vzdump. Record that in fstab.example along with two traps hit during
the fix: `crontab -` on a full disk silently writes a 0-byte crontab,
and a plain docker stop/start leaves the registry serving the old,
deleted directory.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
|
A dockerd restarted from a root login shell inherits
XDG_RUNTIME_DIR=/run/user/0. BuildKit's overlay differ creates its temp
dir there, and once elogind removes it at logout every image export
fails with a misleading "ref ... locked" error. This stalled the
-current full image for three nights. The fix on the VM is
`unset XDG_RUNTIME_DIR` in /etc/default/docker, which rc.docker sources.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
|
The -current full build failed on 2026-09-23 after all ten steps had
passed, at layer export, losing a race on the containerd snapshotter's
content lock ("failed to open writer: ref ... locked for 74ms ...
unavailable"). Disk was not the issue this time (73G free). The build
cache pre-prune that works around this only runs under --force, so the
normal gated nightly build had no guard.
Add docker_build() to lib.sh: docker build, retried once. By the time
the export fails every step is cached, so the retry is little more than
a re-export. Used by build-full-image.sh and build-sbo-testbuild.sh,
the two scripts that export ~20-33G images; bootstrap.sh's image is
small and has never hit this.
Self-check covers success, one failure, and two failures with a stub.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
|
The docker host boots BIOS/legacy with LILO in the MBR; GRUB is installed
but only as EFI binaries the firmware never loads, so it is inert rather
than competing. Worth writing down, because "two bootloaders installed"
reads as ambiguous until you check which one the firmware actually uses.
The part that bites is that LILO maps the kernel by physical block, so
`lilo -v` has to be re-run by hand after a kernel or initrd change. Skip
it and the machine looks fine until it reboots, which on a headless VM
means recovering from the console.
Noted next to the mount entries, and noted that editing those entries
specifically does not need it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
The chain had no failure notification at all, which is why the September
outage ran ten days: every failure was in the log from the first night
and nobody read the log.
notify.sh has two modes because the chain fails in two ways and only one
of them has a non-zero exit status:
run <label> <cmd...> runs the command, posts on non-zero, and passes
the status through so cron still sees the truth.
stale posts if any tag is older than its budget.
The second exists because exit status alone would not have caught what
happened. build-full-image.sh did report non-zero for ten nights, but
build-sbo-testbuild.sh exited 0 every one of them: it saw an unchanged
parent digest and skipped, which is correct behaviour. After the first
alert the chain would have gone quiet while its tags aged six days. The
staleness check asks the registry a different question, "is anything
still current", and catches a skipped stage, a stopped cron or a wedged
mirror alike.
Budgets are split: -current rebuilds nightly (2 days), 15.0 is frozen
and legitimately sits still for weeks (30). A single budget would either
cry wolf on the stable tags or go blind on the rolling ones.
The token is read from /root/.gotify-token (mode 600, not in the repo).
Posting is best-effort throughout: a notifier that fails a build because
the notifier is down would be worse than no notifier.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Also records the host config the chain depends on. The schedule and the
storage layout existed only on the VM, so a rebuilt host would have lost
both, and the README's inline copy of the schedule had already drifted
from what actually runs. crontab.example is byte-identical to the
deployed crontab; fstab.example carries the two-disk layout and the
dockerd mount-namespace trap that makes moving the registry store
non-obvious.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
A rebuild leaves the image it replaced untagged but still resident, which
is ~33G for sbo-full:current. Nothing reclaimed it until the daily 07:00
prune, so the next variant built on top of that dead weight.
That was enough to break the chain. Exporting a 33G image needs roughly
its own size again in transient space, and with the old cache floor the
volume was short at 03:20: build-full-image.sh failed ten nights running
(2026-09-13 through 09-22) with "no space left on device", always in the
same export/unpack phase. build-sbo-testbuild.sh then saw an unchanged
parent and skipped, correctly and silently, so the -current tags sat six
days stale while 15.0 kept rebuilding fine.
Prune at the point where the superseded image becomes dangling: after
the push, while the tag just pushed is the only thing referencing its
layers. `docker image prune -f` only removes untagged images, so the
chain's own tags are never at risk. Best-effort, since a failed prune
costs space, not correctness.
bootstrap.sh is left alone: its images are ~150MB, so a superseded one
is noise next to the transient peak this addresses.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
The registry never reclaims blobs, so its store grows until the disk fills and the nightly builds fail with "no space left on device". Add registry-gc.sh, run weekly (Sunday 08:00), plus a daily dangling-image prune.
registry-gc.sh refuses to run while a build is active, stops the registry for a stable blob graph, deletes only untagged manifests (-m) and their blobs, restarts via an EXIT trap, and verifies a tag still pulls.
distribution 2.8.x GC does not follow OCI image indexes, so -m deletes their child manifests (distribution#3178). Default BuildKit provenance made every pushed tag an OCI index, which made -m destructive. Build scripts now pass --provenance=false (plain schema2), and registry-gc.sh refuses to run if any tag is still an index.
|
|
A dead dockerd is a global failure, not a per-variant one, but every stage
treated it as the latter: the digest probes read an unreachable daemon as
"no digest, rebuilding to be safe", the build then failed, and the variant
loop logged "WARNING: variant X failed; continuing" and moved on. Cron kept
exiting nonzero into a log nobody read, so a daemon that died in August went
unnoticed for 25 days while no image was ever rebuilt.
Add require_docker() to lib.sh, following the existing require_mount()
precedent, and call it before the variant loop in all three stages. One loud
error, exit 1, instead of a nightly pile of warnings.
In build-full-image.sh the check goes before the --force cache prune, since
that prune also talks to the daemon.
Self-check stubs `docker` so it never contacts a real daemon.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AD4jP4wFMpBgEvd9ZP9Xhu
|
|
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
|
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
|
Bake a PROJECT_VERSION="1.0.0" const into every script (test-build,
install.sh, and the three image-builder/*.sh). test-build and install.sh
expose -V/--version; the image-builder scripts already use --version for
the Slackware target, so their project-version flag is -V only. bootstrap.sh
handles -V before its root check so it is queryable as any user.
No VERSION file: a release bump is one sed over the consts, anchored on
^PROJECT_VERSION= so it never hits the SBo package VERSION/OPT_VERSION vars.
Release procedure documented in CLAUDE.md; user-facing note in README.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
|
Extracted from the .extras/ dir of the sbo-slackbuilds package repo, where it
grew from a throwaway helper into a standalone tool. Git history starts fresh;
the original 41-commit development log is preserved in HISTORY.md.
Contents: the test-build CLI (resolves a local SlackBuild's deps from a
configured SBo tree and builds it in a throwaway container, lints, caches deps),
the image-builder chain that produces the images it consumes, their pure-logic
self-checks, design specs/plans, an install.sh for ~/bin, and docs.
Licensed GPLv2-only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|