aboutsummaryrefslogtreecommitdiffstats
path: root/image-builder/README
blob: a1ea0c485bfe33eb145a05423dfc9b493165adce (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
sbo-testbuild image builder
===========================

Builds the docker images that test-build consumes:
  docker.noland.dnx:5000/sbo-testbuild:current
  docker.noland.dnx:5000/sbo-testbuild:15.0

Three scripts, chained (see docs/specs/2026-07-13-image-builder-design.md):
  bootstrap.sh           sbo-base:{ver}       FROM scratch, base pkgs from NAS
  build-full-image.sh    sbo-full:{ver}       FROM base, all series
  build-sbo-testbuild.sh sbo-testbuild:{ver}  FROM full, + sbopkg + tools

Plus one maintenance script (not part of the chain):
  registry-gc.sh         reclaim unreferenced blobs from the registry store

And two reference copies of the host config the chain depends on. Nothing
reads them; they are here so a rebuilt VM is reproducible:
  crontab.example        the nightly schedule, as deployed
  fstab.example          the two-disk storage layout, as deployed

All settings live in ./config.

VM setup (docker.noland.dnx, Slackware x86_64, 4 vCPU / 4 GB)
-------------------------------------------------------------
Two disks, and which is which matters: a system disk (80 GB, also holding the
registry store) and a separate docker volume (160 GB) for images and build
scratch. See fstab.example.

1. Install docker; enable the daemon. Then add to /etc/default/docker
   (rc.docker sources it before starting dockerd):
     unset XDG_RUNTIME_DIR

   A dockerd restarted from a root login shell inherits
   XDG_RUNTIME_DIR=/run/user/0, and BuildKit's overlay differ puts its temp
   dir there. elogind removes that dir at logout, so every later image
   export fails. dockerd logs "failed to create temp dir: stat /run/user/0:
   no such file or directory", and the build prints a misleading
   "failed to open writer: ref moby/1/... locked ... unavailable". Retrying
   does not help; only a restart from a clean env does. This stalled the
   -current full image from 2026-09-23 to 09-25. Check a running daemon with:
     tr '\0' '\n' < /proc/$(pidof dockerd)/environ | grep XDG_RUNTIME_DIR

2. NFS-mount the two NAS trees read-only, named to match:
     /mnt/nas/slackware64-current   -> -current mirror tree
     /mnt/nas/slackware64-15.0      -> 15.0 mirror tree
   Each is a full mirror (PACKAGES.TXT, ChangeLog.txt, slackware64/, patches/,
   extra/). Root must be able to read them (bootstrap runs installpkg as root).

3. Run a LAN registry. Put its storage on a DIFFERENT disk from docker's:
     docker run -d --restart=always -p 5000:5000 \
       -v /opt/sbo-testbuild/registry:/var/lib/registry --name registry registry:2

   The bind target belongs on the system disk, not the docker volume. The
   registry grows with every push and never shrinks on its own, while the
   build needs a large transient peak at a fixed hour; sharing one volume
   pits a slow leak against a hard failure, and the build loses. See
   fstab.example for the layout and the dockerd-namespace trap if you move
   an existing store.

4. Mark the registry insecure (plain HTTP) on the VM AND every pulling client
   (this dev box, the buildsystem VM). In /etc/docker/daemon.json:
     { "insecure-registries": ["docker.noland.dnx:5000"] }
   then restart docker.

5. Drop the two prebuilt packages (built once, re-drop on upstream bumps):
     /opt/sbo-testbuild/pkgs/sbopkg-*.txz
     /opt/sbo-testbuild/pkgs/sbo-maintainer-tools-*.txz

6. Install the nightly cron (root). The NAS repos sync at 01:00 and 02:00, so
   the chain runs after and both variants are ready well before the ~09:00 work
   start. No --force: each script self-gates (bootstrap on the ChangeLog hash,
   full-image on the base-image digest, build-sbo-testbuild on the full-image
   digest + tools .txz hash), so an unchanged night is a cheap no-op. -current
   moves daily and rebuilds most nights; 15.0 is frozen stable and rebuilds only
   on a real repo update.

   The schedule lives in crontab.example, which is a copy of what the VM runs:
     crontab crontab.example        # or paste it into `crontab -e`

   Install it rather than retyping it. The timings are load-bearing, not
   cosmetic: reclaim runs at 02:50, immediately before the 03:00 chain, so the
   headroom exists when the build needs it. An earlier schedule pruned in the
   afternoon instead and the -current build failed ten nights running with
   "no space left on device". The file explains each window.

7. Ensure docker.noland.dnx resolves on the LAN (static IP or DNS).

Manual first run
----------------
   ./bootstrap.sh --version current --force
   ./build-full-image.sh --version current --force
   ./build-sbo-testbuild.sh --version current --force
Then confirm:
   docker pull docker.noland.dnx:5000/sbo-testbuild:current
   docker run --rm docker.noland.dnx:5000/sbo-testbuild:current sbopkg -V

Flags: --force (rebuild unconditionally), --version <current|15.0> (one variant).

Registry garbage collection (registry-gc.sh)
--------------------------------------------
The registry keeps every blob ever pushed; it never reclaims on its own. Left
alone, the store grows until the disk fills and the nightly builds fail. Two
cleanups keep it bounded:

  docker image prune -f   (daily) removes dangling images left in the docker
                          store when a tag moves to a freshly built image.
  registry-gc.sh          (weekly) reclaims unreferenced blobs from the
                          registry's own store.

registry-gc.sh is deliberately conservative:
  * it refuses to run while any build script is active, so it can never race a
    push (cron runs it at 08:00 Sunday, well after the ~06:30 chain);
  * it stops the registry so the manifest/blob graph is stable, and restarts it
    via an EXIT trap even if collection fails part-way;
  * it deletes only untagged manifests (-m) and the blobs they alone
    reference, so every tag keeps resolving;
  * it verifies afterwards that a tag still pulls.

Why the build scripts pass --provenance=false: with default BuildKit
provenance, `docker push` stores an OCI image index (the image plus an
attestation manifest). Distribution 2.8.x garbage collection does not follow
OCI indexes, so `-m` would delete their child manifests and orphan the layer
blobs (distribution issue #3178). Disabling provenance keeps each tag a plain
Docker schema2 manifest, which the collector handles correctly. registry-gc.sh
refuses to run if it finds any tag that is still an index, so this cannot
regress silently.

Preview without touching anything (registry stays up, nothing is deleted):
  ./registry-gc.sh --dry-run

Storage path is resolved from the running container's /var/lib/registry mount,
so the script follows the registry wherever it is mounted.

Tests
-----
   bash test-image-builder.sh     # pure-logic self-check, no docker