Skip to content

verify-release-binaries

Proves that every release binary published to the estate's object store is still there, and still the bytes it claimed to be.

include:
  - component: gitlab.com/phpboyscout/cicd/[email protected]

Runs only on a pipeline schedule carrying VERIFY_TASK=binaries, in this repository. It is an estate-wide chore, so it wants one schedule in one place rather than a copy per project.

Why it exists

Release binaries live on pkg.phpboyscout.uk and carry a public guarantee: a published file is never changed and never removed. Nothing enforces that.

The object store's bucket lock, which did enforce it absolutely, was dropped deliberately — under it an accidentally published secret would have been unremovable, which is a worse position than the tampering it defended against. The credential that writes releases holds DELETE, and the store has no object versioning. So enforcement was replaced with detection, and this is the detection. See colophon spec 0025 D13 and infra spec 0019 D4.

It depends on nothing

The engine imports only Python's standard library. No pip, no install step, nothing fetched at run time. This job is the only thing standing behind the release store's guarantee and it runs unattended, so it must not be able to fail because a package was missing from an image — which it did, once, on ci-base, which carries python3 and no pip.

It holds no bucket credential

It reads the public URL like any other consumer. That is deliberate twice over:

  • it verifies what a user actually receives, rather than what the bucket believes it holds;
  • a verifier that cannot write can never be the cause of the damage it reports.

The only credential is a GitLab token with api scope, to enumerate releases across the group and to raise an issue when something is wrong. CI_JOB_TOKEN will not do — it cannot read other projects' releases.

Three tiers

The estate does not publish uniformly, so a verifier written to one shape would fail on a third of it.

tier what it proves where it runs
reachability every published link resolves everywhere, every run
integrity the bytes match the published digest where checksums.txt exists
authenticity the digest list is signed where a signature exists
project publishes
colophon, go-tool-base, ffmpeg-wasi checksums.txt + a PGP .sig
keryx, skillup checksums.txt, no signature
rust-tool-base per-file .minisig, no checksums.txt

A tier with nothing to run against is counted, not failed. Reporting rust-tool-base as broken every night because it signs differently would train everyone to ignore this job, which is the only failure mode that really matters for a scheduled check.

Integrity rotates; reachability does not

Verifying every digest means downloading everything the estate has ever released — about 23 GiB, growing roughly 1 GiB a month. Nightly, that is hours of runner1 spent re-reading bytes that were correct yesterday.

So each run digest-verifies one share, and integrity_parts runs cover everything. Shares are sliced on a hash of the release's identity rather than its position, so adding a release does not reshuffle every other release into a different share.

Reachability still covers everything, every run, because it is a HEAD. That is the tier that catches deletion, which is the failure this job exists for.

It is allowed to fail

allow_failure defaults to false, unlike most scheduled chores. A missing or altered release binary is a broken public guarantee and possibly a compromised credential. A check that cannot fail is not a check.

Issues are filed on the affected project and are idempotent: a fault that persists does not file a new issue every night.

Inputs

input default purpose
stage test stage to run in
image ci-base:v0.1.3 any image with python3; the engine is stdlib-only
if schedule + $VERIFY_TASK == "binaries" the gate
group phpboyscout group to enumerate
projects (empty) space-separated paths to check instead of the group
base_url https://pkg.phpboyscout.uk only links under this are checked
integrity_parts 7 cover all digests over this many runs
skip_integrity false reachability only, the cheap pass
open_issues true file issues on findings
token_variable GITLAB_TOKEN CI/CD variable holding the token
allow_failure false whether a finding fails the pipeline

Scheduling it

Off-peak. runner1 is concurrent = 2 and is the whole Linux estate while the AWS fleet is off; the nightly Renovate scan holds 00:00 UTC.

The schedule must set VERIFY_TASK=binaries. That variable is load-bearing: schedules all run against the default branch, so without it every schedule on this repository would fire every scheduled job.