Decision log¶
Every design decision behind these components is recorded as a dated spec
under docs/development/specs/, per spec-first development.
This page is the exhaustive index — every spec, in date order. The major
design decisions are also distilled into readable
Explanation pages; this table is the complete,
unfiltered record those pages draw from.
| Date | Spec | Status | Decides | Component(s) |
|---|---|---|---|---|
| 2026-05-15 | cicd v0.1 | approved | Four initial components shipped as v0.1.0: single-file form, the infra-tools image, the core input rules. |
all |
| 2026-05-16 | tofu-plan-apply v0.2 | approved | OIDC-authenticated tofu plan/apply, the GitLab HTTP state backend, manual-gated apply. |
tofu-plan, tofu-apply |
| 2026-05-16 | tofu-apply-plan-sources v0.3 | approved | plan_source (job/ref); the hidden-job extends: pattern; mode-coupled rules:; tag_pattern. |
tofu-plan, tofu-apply |
| 2026-05-18 | gate-component-rules v0.4 | approved | Gate jobs carry an explicit rules: so they run in merge-request pipelines. |
tofu-lint, tofu-security, tofu-validate, zensical-pages |
| 2026-05-19 | token-inputs v0.5 | approved | Token-requiring inputs default to $CI_JOB_TOKEN; the consumer overrides. |
tofu-plan, tofu-apply |
| 2026-05-24 | provider-cache v0.6 | approved | OpenTofu provider plugin cache via TF_PLUGIN_CACHE_DIR, reduces re-download flakes. |
tofu-plan, tofu-apply, tofu-validate |
| 2026-05-27 | module-publish v0.7 | approved | tofu-module-publish publishes the tagged tree to the GitLab Terraform Module Registry. |
tofu-module-publish |
| 2026-08-01 | 0061-releaser-pleaser-ff-tag-commit | approved | On a project combining merge_method: ff with GitLab 19.2's automatic rebase before merge, releaser-pleaser cuts the tag at the Release MR's recorded head. Automatic rebase is what lets a Release MR that is behind its target be merged at all — GitLab rebases it during the merge and never writes that back to the MR; without the setting GitLab blocks the merge until an explicit rebase, which does update the head, so such a project is not exposed. The fleet keeps it on deliberately, for Renovate volume. A fast-forward merge produces neither a merge commit nor a squash commit, so upstream (v0.9.0, latest) falls through to pr.SHA — and GitLab never writes the rebase it performs DURING the merge back to the MR, so that head is permanently pre-rebase. The tag then names a commit that is not on the default branch and the release silently omits whatever landed while the MR was open. Measured: 9 of 231 recent tags across phpboyscout/go, 8 repos. The MR only goes stale when rp declines to force-push, which it does whenever the landed commits do not change the changelog — so the loss set is always the trailing run of docs/chore/ci/style/test commits and released code cannot be dropped this way. Upstream added that fallback deliberately in PR #210 ("support Fast-forward merge", v0.6.1), replacing a loud "pull request is missing the merge commit" failure with a silent wrong release; it merged unreviewed with 0% patch coverage, and no upstream issue tracks it. Adds releaser-pleaser:verify: sets squash: true on the open Release MR so the merge records a squash_commit_sha (a commit GitLab creates on top of the current target head, hence always on the branch — no project-wide squash setting, so the never-squash house rule is untouched) and asserts the tag is on the release branch after each release, exit 2. A separate job because the upstream image has no HTTP client at all. Rebasing the stale MR, and setting squash_option=require fleet-wide, both considered and rejected; forking rp deferred. |
releaser-pleaser |
| 2026-05-30 | module-cache v0.7.2 | approved | .terraform/modules/ cache reduces flakes from transient upstream-registry outages. |
tofu-validate |
| 2026-06-02 | renovate-self v0.8 | approved | renovate-self component + the shared Renovate preset with a cicd-component-pin custom manager. |
renovate-self |
| 2026-06-03 | go-track v0.9 | approved | Four Go components extracted from go-tool-base. | go-lint, go-test, go-security, goreleaser |
| 2026-06-03 | rust-track v0.10 | approved | Five Rust components; disk-pressure tuning; cargo-binstall bootstrap. | rust-lint, rust-test, rust-security, rust-docs, release-plz |
| 2026-06-08 | gitleaks-scan-scoping v0.10.3 | approved | Gitleaks scans scoped to the MR's commit range, avoiding cross-branch false positives on shared runners. | go-security, rust-security, tofu-security |
| 2026-06-12 | release-plz-split-jobs v0.10.3 | approved | release-plz split into separate pr and release jobs to avoid working-tree mutation blocking publish. |
release-plz |
| 2026-06-16 | goreleaser-retry v0.10.5 | approved | goreleaser gains retry_max (default 2) auto-retry on transient failures. |
goreleaser |
| 2026-06-16 | release-plz-order-pr-after-release v0.10.4 | approved | release-plz:pr ordered after release-plz:release to avoid a spurious same-version MR. |
release-plz |
| 2026-06-17 | release-plz-checkout-pinned-sha v0.10.6 | approved | Checkout forced to the pipeline's commit SHA, avoiding a stale branch ref on reused runners. | release-plz |
| 2026-06-19 | renovate-self-token-self-ref v0.10.7 | approved | Token aliased to a non-colliding runtime variable name, avoiding a self-referencing job/group-variable collision. | renovate-self |
| 2026-06-21 | schedule-pipeline-scoping v0.10.8 | approved | Every component except renovate-self gains a leading schedule → never guard. |
all |
| 2026-06-21 | gate-components-tag-scoping v0.11.1 | approved | Gate jobs also skip tag pipelines ($CI_COMMIT_TAG → never) — tags are for publish jobs, not gates. |
tofu-lint, tofu-security, tofu-validate, zensical-pages |
| 2026-06-21 | go-image-default-1.26.4 v0.11.2 | approved | Image defaults bumped to match go.mod's toolchain requirement. |
go-test, go-security |
| 2026-06-21 | goreleaser-gotoolchain-auto v0.11.3 | approved | GOTOOLCHAIN default changed from local to auto, so goreleaser resolves go.mod's toolchain directive. |
goreleaser |
| 2026-06-21 | releaser-pleaser-component v0.11 | approved | releaser-pleaser wrapper component: schedule-never guard + token convention over apricote's image. |
releaser-pleaser |
| 2026-06-21 | self-test-churn-scoping | approved | Self-test triggers scoped to changed components, not run unconditionally. | all |
| 2026-06-22 | components-use-dev-tools v0.14 | approved | go-*/rust-* components default to the consolidated dev-tools image; runtime tool installs deleted. |
go-lint, go-test, go-security, goreleaser, rust-lint, rust-test, rust-security, rust-docs |
| 2026-06-22 | hugo-pages-component v0.12 | approved | hugo-pages component: build+deploy split, an MR build gate hand-rolled Hugo sites lacked. |
hugo-pages |
| 2026-06-22 | hugo-pages-mr-gate v0.13 | approved | mr_gate boolean input — a loose, direct-to-main mode for content sites. |
hugo-pages |
| 2026-06-22 | pipeline-churn-interruptible-cache v0.11.5 | approved | interruptible: true on gate jobs; terminal cache readers go pull-only; stable per-project cache keys. |
all |
| 2026-06-22 | renovate-image-default-43 v0.11.4 | approved | Image default bumped to match the preset's managerFilePatterns usage (≥39). |
renovate-self |
| 2026-06-23 | change-detection | implemented | changes array input on gate components; security category exempt (always-on). |
go-*, rust-*, tofu-*, zensical-pages, hugo-pages, svelte-* |
| 2026-06-23 | skill-security-component v0.19 | approved | skill-security: hidden-char, injection-heuristic, plugin-schema, and gitleaks scans for AI instruction-file repos. |
skill-security |
| 2026-06-23 | svelte-frontend-track | implemented | svelte-build/svelte-lint/svelte-test/svelte-security; the Go↔Svelte embed coupling. |
svelte-build, svelte-lint, svelte-test, svelte-security |
| 2026-06-30 | docs-diataxis-restructure | approved | This site's restructure into Diátaxis — the spec behind the page you're reading. | docs (no component) |
| 2026-07-06 | goreleaser-pro-toggle-v0.20.0 | draft | A pro: true input runs goreleaser-pro (bundled in dev-tools) for Pro-only config — app_bundles/dmg/native-notarize. Driven by krites' signed macOS .dmg. |
goreleaser |
| 2026-07-12 | image-pins-track-latest v0.20.2 | approved | Renovate custom managers (gitlab-tags) auto-track the dev-tools / infra-tools image tags pinned in component defaults + reference docs; stale pins bumped to latest in one pass. |
go-*, rust-*, svelte-*, tofu-*, zensical-pages |
| 2026-07-12 | track-scanner-image-pins v0.21.1 | approved | Renovate custom manager (docker) auto-tracks the scanner tool images (trivy, gitleaks, osv-scanner, semgrep) in the *-security components + reference docs; skewed/stale pins reconciled to latest. |
go-security, rust-security, svelte-security, skill-security |
| 2026-07-12 | preset-default-json v0.21.2 | approved | Renovate preset renamed default.json5 → default.json: Renovate resolves the unnamed gitlab>phpboyscout/cicd preset from default.json only, so consumers were silently missing the component-pin manager. |
renovate-self (preset) |
| 2026-07-12 | osv-scanner-toolchain-and-waiver | approved | osv-scanner: disable call-analysis by default (osv_scanner_call_analysis) to end the GOTOOLCHAIN=local exit-127; always apply the repo .osv-scanner.toml via --config; ship a component-wide GO-2026-5932 (x/crypto) waiver; bump image to v2.4.0. |
go-security |
| 2026-07-13 | bake-zensical-into-image | approved | Bake the Zensical docs toolchain into the infra-tools image (pipx, like checkov); zensical-pages runs the baked CLI instead of pip install --require-hashes -r requirements-lock.txt. Removes the un-rehashable hand-rolled lockfile (and 18 fleet-wide stuck Renovate MRs), drops the python_lock input, pins the toolchain in one place. |
zensical-pages |
| 2026-07-13 | tag-pipeline-workflow-guard | approved | The fleet-wide workflow: dedup rule keyed on $CI_OPEN_MERGE_REQUESTS && push also matched tag pushes, intermittently suppressing the whole tag pipeline so release publish jobs never ran (issue #2). Swap to the tag-safe $CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS guard; fix the root pipeline + document it; remediate all 16 consumer repos. No release. |
root pipeline, docs |
| 2026-07-20 | releaser-pleaser-project-scoping | approved | releaser-pleaser v0.8.0's PendingReleases called the instance-wide GET /merge_requests, so one merged-but-untagged Release MR jammed the release job in every project the group token could see (issue #4). Bump the pinned image to v0.9.0, where it is project-scoped; the --owner/--repo fix proposed in the issue is inert under GitLab CI. Add the missing Renovate manager for the apricote + release-plz image pins that let this bump go unnoticed. |
releaser-pleaser, release-plz (pin), renovate.json |
| 2026-07-20 | tflint-authenticated-ruleset-lookup | approved | tflint --init resolves a non-baked ruleset version from api.github.com unauthenticated — rate-limited per source IP, so it 403s on shared runners. Adds a github_token input (default $GITHUB_COM_TOKEN) exported as GITHUB_TOKEN on the tflint job only. The consumer-side half of infra-tools' plugin-dir fix: that one authenticates the bake, this the fallback. |
tofu-lint |
| 2026-07-20 | renovate-terraform-docs-postupgrade | approved | Renovate bumps version = in .tf but never regenerates the terraform-docs README tables, so every Terraform module bump failed terraform-docs-drift by construction. Add a postUpgradeTasks hook via an opt-in named preset, plus terraform_docs_version (the binary is absent from the Renovate image and not Containerbase-installable) and allowed_commands (global-only, so it cannot ship in the preset) inputs on renovate-self. Allowlist anchored to the exact terraform-docs shape — no shell. |
renovate-self, preset, tofu-lint (backstop) |
| 2026-07-20 | renovate-cadence-soak-not-window | approved | Seven repos gated Renovate behind schedule: ["before 6am on monday"] while running it daily, so six of seven runs were no-ops and pin lag was 0–7 days — why the infra-tools v0.4.1 hotfix could not reach consumers. Drop the window for minimumReleaseAge: "3 days" (the Go track's value), and exempt first-party phpboyscout/** releases from the soak via the shared preset. | preset (default.json), renovate.json ×7 |
| 2026-07-21 | rust-resource-group-input | approved | The compile-heavy Rust jobs (clippy, test-linux, test-integration, coverage, cargo-doc) ran concurrently and filled every slot of the shared self-hosted runner with parallel full-workspace compiles, starving other projects. Add a resource_group input to rust-lint/rust-test/rust-docs, defaulting to $CI_PROJECT_PATH_SLUG-rust-compile so a consumer's compile jobs serialise (project-scoped — never affects other projects). Fleet-wide form of rust-tool-base's P0. | rust-lint, rust-test, rust-docs |
| 2026-07-22 | go-test-dind-testcontainers | approved | New config-* projects draft testcontainers-go integration tests that need a Docker daemon. Add an opt-in go-test-integration job (docker:dind service + DOCKER_HOST, Ryuk disabled) mirroring rust-test's test-integration. No runner change (runner1 already privileged) and no dev-tools change (testcontainers-go uses the Docker API). Self-tested by cicd's first Go module fixture — stdlib-only, drives the dind Engine API to run a container, so it validates the privileged runner with zero dependency surface. | go-test, tests/go-test/fixture |
| 2026-07-23 | centralized-renovate-presets | approved | An audit found 51 byte-identical renovate.json files and 69/70 missing osvVulnerabilityAlerts (the gap that hid a HIGH grpc CVE). Move policy into COMPOSABLE presets: a base every repo extends (now carrying OSV alerts + patch+minor automerge + majors-drafted + house posture — live fleet-wide on merge, no per-repo change), plus orthogonal leaves — ecosystem (:go :rust :javascript :tofu :docker :python), static-site supersets (:hugo :zensical), role (:library :application). A project extends its bag of leaves; matchManagers-scoping makes them compose. Adds a renovate-config-validator self-test over every preset. Migration to thin extends is a canary-then-rest follow-on. Pairs with the group-wide ff+pipeline-must-pass merge settings (D9). | default.json (base) + leaf presets, root pipeline |
| 2026-07-22 | renovate-group-run-headroom | approved | The first live renovate-group run processed all 73 phpboyscout/** repos but took 53.7 of the project's 60-min job timeout. The scan is sequential and Renovate's repo order is deterministic, so a run that ever exceeds the timeout drops the SAME tail repos every night, silently. Add a timeout input (default "2 hours") lifting the job ceiling above the 60-min project default — viable because runner1 sets no maximum_timeout. Timeout is the safety fix; persisting Renovate's repositoryCache (to keep the run fast, not just un-capped) is a deferred follow-on. | renovate-group |
| 2026-07-22 | automerge-first-party-cicd-pins | approved | With renovate-group opening phpboyscout/cicd pin bumps across the whole group at once, hand-merging our own CI-verified component pins is the bottleneck. Add one narrow packageRule to the shared preset (default.json) automerging phpboyscout/cicd pins via platformAutomerge (GitLab merge-when-pipeline-succeeds) — the consumer's own green pipeline is the gate, a breaking change leaves the MR open+red. Scoped to phpboyscout/cicd exactly (not image pins or third-party); consumer stricter rules still override. Pairs with the existing first-party soak-skip carve-out. | preset (default.json) |
| 2026-07-22 | renovate-group-autodiscover-component | approved | renovate-self is single-repo and needs a per-repo schedule nobody added — 37/53 go repos silently drifted behind cicd v0.26.0 with the wiring in place but dormant. Add a new renovate-group component running Renovate in autodiscover mode over a required autodiscover_filter (e.g. phpboyscout/**), config-gated (onboarding=false + requireConfig=required) so unconfigured repos are skipped silently. One schedule (dogfooded in cicd, whose filter includes cicd itself) keeps the whole group current and picks up new repos automatically. Reuses renovate-self's token-alias, image-≥39 floor, and terraform-docs postUpgrade mechanics verbatim. Follow-up: dogfood switch after the tag, then a fleet sweep decommissioning the dormant renovate-self usages. | renovate-group (new), renovate-self (sibling), root pipeline |
| 2026-07-23 | go-test-exclude-generated-coverage | implemented | go-test drives its coverage badge off the full cover.out, so mockery/protoc/stringer generated packages (~0% coverage) dilute it well below real hand-written coverage — measured 98.2% → 73.0% from one generated package in go/observability, forcing fragile per-repo paths scoping. Add an exclude_generated input (default true) filtering files that carry Go's // Code generated … DO NOT EDIT. marker out of the profile before the badge is computed, in shell (no image change). paths stays ./...; the raw artifact is untouched; false reproduces today's figure. Issue #5 Part A. | go-test, tests/go-test/fixture |
| 2026-07-23 | releaser-pleaser-stage-ordering | implemented | releaser-pleaser emits needs: [] (from a needs input defaulting to the empty array), which in GitLab means "start immediately, ignore all stages" — so the release job runs concurrently with the lint/test/security gate stages instead of after them, in cicd and fleet-wide (proven: releaser-pleaser started 21s before go-test finished on a main pipeline). Remove the needs: line and the input so the job falls back to stage ordering, matching the already-correct release-plz:release. Consumers wanting DAG order can still override the job. fix, folded into v0.29.0. tofu-apply's own needs: [] (ref-mode tag apply) flagged out of scope. | releaser-pleaser |
| 2026-07-23 | go-changelog-changes-release-ci | implemented | REVERSED 2026-09-03 by 0079 D3: CHANGELOG.md left both default changes lists once colophon's fast-forward release MR made re-testing it redundant (2,030 min / 21 days). Original: A bootstrap v0.1.0 Release MR changes only CHANGELOG.md, which matches nothing in go-test/go-lint's default changes array, so the first release of a new module is tagged never having passed test or lint and ships an unknown coverage badge (security is exempt — it has no changes filter). Add CHANGELOG.md + **/CHANGELOG.md to both defaults so every Release MR (not just v0.1.0) runs a real test+lint pass before the tag is cut; ordinary code MRs already match **/*.go and are unaffected. Defaults-only patch. Issue #5 Part B. | go-test, go-lint |
| 2026-07-24 | trivy-db-shared-runner-cache | implemented | The v0.29.0 release burst (125 MRs) showed the security jobs are the worst network offenders: every trivy fs re-downloads the 1.2 GB Aqua DB, and every osv-scanner calls api.osv.dev live. Add a shared runner cache (/opt/ci-cache) refreshed once nightly (22:40 UTC, 20 min before the midnight Renovate run): trivy skips its DB update when a DB is present, and osv-scanner runs OFFLINE against a cached ~11 MB Go / ~200 MB npm DB — making the scan hermetic under burst. Guarded on DB presence (osv offline-without-DB hard-errors 127, never a false pass), so with no shared cache both fall back to today's online behaviour. New scanner-db-refresh jobs + scanner-cache-refresh schedule; runner1 volume applied out-of-band. Verified all download/offline commands locally at CI tool versions. | go-security, rust-security, svelte-security, root pipeline |
| 2026-07-25 | tofu-deploy-generate-component | approved | Extract the release-tag promotion generator proven e2e in phpboyscout/infra (selective-env spec D8/D15/D16, prove-then-extract) into a reusable component. A consumer-owned JSON catalog (schema v1 — the versioned public interface) drives a fail-closed stdlib engine emitting one deploy child pipeline per target; children plan from the tag checkout and apply same-pipeline artifacts, killing the stale-plan/rebase race class. Engine bundled in this repo, fetched at exactly the component version (template+engine = one immutable bundle), generator_path overrides with a consumer-local variant. Tier guards: production/other targets forced manual+confirmation; the strict-semver ref gate is hardcoded, never an input. Two modes (generate on tags / verify-fixtures golden-diff on MRs); trigger jobs stay consumer-side (templates cannot loop). | tofu-deploy-generate |
| 2026-07-25 | renovate-gomod-import-paths | approved | Every Go major bump the fleet's Renovate bot opened landed as a broken draft: it wrote the new major into go.mod alongside the old one but never rewrote the /vN import paths, so the code kept compiling against the OLD major and CI went green on a no-op (proven across go-github v88→v89 ×2, azappconfig v1→v2, consul/api v1→v2 — the last caught only by its own TestDependencyFootprint). Cause: gomodUpdateImportPaths was off, so Renovate did the manifest half of a major and left the source half undone. Add postUpdateOptions: ["gomodTidy", "gomodUpdateImportPaths"] to the :go leaf (go.json) so Renovate rewrites import paths itself and lands a go mod tidy-clean manifest — a real, verifiable diff instead of a false-green draft. Go-specific → leaf not base; majors stay drafted (not auto-merged); the full renovate/renovate:43 image already supplies the Go toolchain + mod helper, so no image change. | preset (go.json) |
| 2026-07-26 | renovate-merge-sweep | approved | Renovate cannot complete an automerge on any GitLab project using the fast-forward (ff) merge method: GitLab's merge API requires a sha for an FF merge and Renovate never sends one, so every automerge returns 400 "SHA must be provided when merging" and the MR stays open (confirmed 2026-07-26 by a DEBUG renovate-group run; upstream renovate #26972). The fleet keeps ff (only it gives merge-commit-free linear history), so the automation is fixed around it. Add renovate-merge, a scheduled sibling of renovate-group: one job enumerates the group's open, non-draft, mergeable renovate/* MRs and PUT …/merges each with its head sha (the call that works). Safety-first selection (source-branch prefix guard never touches human MRs; majors excluded as drafts); never rebases; per-MR failures skip, only auth/enumeration fails the job. stdlib-python on infra-tools, own RENOVATE_TASK=merge schedule, dry_run + exclude inputs. Dogfood include + hourly schedule are a post-tag follow-up (as renovate-group did). | renovate-merge (new) |
| 2026-07-25 | stoppable-environments | approved | Non-production environments cost money overnight for nothing. Stop is an overlay apply (scale-to-zero tfvars), never a destroy, so state and data survive; catalog schema v2 adds stoppable + stop_tfvars, tier-guarded so production can never be stopped; tofu-deploy-generate emits a stop lane into each deploy child pipeline; tofu-stop ships as the trunk-pipeline stop lane; tofu-apply gains optional environment recording. | tofu-stop (new), tofu-apply, tofu-deploy-generate |
| 2026-07-27 | renovate-gomod-tool-directives | approved | gomodUpdateImportPaths closed the source-import half of a Go major bump, but Go 1.24 tool directives in go.mod name an import path too and nothing rewrites it to /vN. The old path's dependency tree no longer resolves, so go mod tidy fails and the draft is unbuildable — presenting as a broken upstream dependency rather than an incomplete rewrite (keryx !164, golangci-lint v1→v2). Renovate cannot know which modules a repo lists as tools, so add a prBodyNotes warning to every gomod major on the :go leaf rather than a hand-maintained list. | renovate presets (go.json) |
| 2026-08-01 | 0060-module-publish-stage-default | approved | tofu-module-publish defaulted to stage: .post. Every tofu gate carries if: $CI_COMMIT_TAG → when: never, so on a release tag the publish job was the pipeline's only job — and GitLab does not create a pipeline whose jobs are all in .pre/.post. No tag pipeline was created at all, so the module silently never published and nothing went red (green Release MR, tag present, Release page present, registry empty); three of four iac/* module repos lost their latest tag for two months (issue #8). Default becomes release — what every consumer already declares, and what the one working repo already passed. Guarded twice: a repo-level lint forbidding any component defaulting stage to a built-in stage, and the self-test now takes the default instead of overriding it. Post-publish registry verification considered and deferred. | tofu-module-publish |
| 2026-08-01 | 0062-renovate-enable-pre-commit-manager | approved | Renovate's pre-commit manager is opt-in — documented as beta and off unless a config enables it — so a repo carrying a .pre-commit-config.yaml gets no hook updates and no signal it is missing them: the failure mode is an ABSENCE of MRs, indistinguishable from "nothing to update". Measured 2026-08-01: 13 of 98 active projects carry one, 38 remote hook repos in total, and only 4 enable the manager locally (sigillum, scoutdm, krites, go-tool-base) — the other 9, including all four iac/terraform-aws-* module repos and infra, have been frozen since creation. None restricts enabledManagers, so nothing local blocks a preset-level enable. Enabled once in the BASE preset rather than an ecosystem leaf, because .pre-commit-config.yaml is language-agnostic and the affected repos span Go, Tofu and mixed tracks; it is a no-op where no such file exists. The four local enables are left in place (identical value, no conflict). Expect a one-off burst that the existing automerge policy + hourly merge sweep drain. | renovate presets (default.json) |
| 2026-08-02 | 0063-goreleaser-retry-reasons | approved | goreleaser is the only component setting retry, and it listed stuck_or_timeout_failure — deprecated in GitLab 19.1, so every consumer including the component got a CI lint warning on every validation. The migration is NOT a value swap: 19.1 also moved some failures OUT of runner_system_failure, a value being kept, into runner_external_dependency_failure (registry unreachable — a Docker Hub rate limit lands here) and runner_interrupted (restart, shutdown, spot reclamation). Replacing only the deprecated value would have silently narrowed coverage on runners already running 19.1.0. The deprecated reason expands to its four documented successors, and the two that split away are added alongside the one they came from. Deliberately excluded: runner_configuration_error (an invalid image is deterministic — retrying reaches the same failure slower) and the job_execution_timeout successors (a genuine timeout is not transient and was never covered). The valid value set was established by submitting each candidate to the CI lint API rather than inferred from prose, which also revealed job_execution_timeout is deprecated too. | goreleaser |
| 2026-08-09 | 0064-discord-release-announcements | superseded | The Discord announcements channel had nothing publishing to it, and GitLab's own Discord integration cannot fix that without making it worse: its only filtering parameter is branches_to_be_notified (all/default/protected/default_and_protected), a BRANCH filter, so a tag_push_events subscription fires on every tag with no semver awareness — every patch bump from every project, in a group where Renovate churn produces a steady drip of them. It also posts GitLab's generic tag-push message rather than the release notes releaser-pleaser/release-plz already generate. New discord-release component instead: parses the semver tag and announces only when the patch component is 0 (which maps onto Conventional Commits as used here — feat: is a minor, fix: a patch, chore(deps): nothing — so it approximates "something was added"), with announce_patches for headline projects and prereleases quiet by default. Nothing may break a release: allow_failure defaults true and every recoverable condition exits 0, including a MISSING webhook — which is what lets the component roll out across the estate before the Discord webhook exists and start working the day it does. Release notes are best-effort against a real race (the tag pipeline can start before the Release object lands). Deliberate deviation from authoring rule 6: webhook_url names $DISCORD_RELEASE_WEBHOOK as its default, traded for a two-line consumer include across ~20 repos. Filter verified by running the extracted script body locally over an 11-case matrix before it reached CI, which caught two set -e aborts on the degraded no-notes path. SUPERSEDED 2026-09-20 by 0100: announcements moved to colophon's own verb and the component was removed in cicd !383 after 111 consumers moved. | discord-release |
| 2026-08-14 | 0065-renovate-untracked-pin-families | implemented | Two families of version pin were annotated as if Renovate tracked them and were tracked by NOTHING — the failure mode being an ABSENCE of MRs, indistinguishable from "already current". (1) Renovate's native dockerfile manager reads FROM lines, NOT # renovate: comments on ARG/ENV lines; that convention needs a customManager and none existed. The 2026-08-13 group scan extracted "dockerfile": {"fileCount": 1, "depCount": 2} from a dev-tools Dockerfile carrying TEN annotations — the two seen were the FROM and # syntax= lines. Six of the eight missed had rotted, GOVULNCHECK_VERSION by six minors (v1.1.4 vs v1.7.0) — the gate every Go repo's security job runs. infra-tools had the identical hole, all three of its ARG pins stale. Both repos' own comments assert the mechanism exists; it was never written. (2) An image passed as a component inputs: image: is not a job-level image:, so gitlabci never extracts it — go-tool-base sat on dev-tools v0.2.0 for five releases while the SAME image was bumped normally wherever it appears job-level. Dockerfile manager goes in the :docker leaf (all three Dockerfile-carrying repos already extend it; base would misfile container policy across 100+ repos with no Dockerfile); component-input manager goes in the BASE preset, scoped to our own registry so any overlap with gitlabci is a duplicate rather than a conflict. Both match the ANNOTATION/shape rather than a list of names, so a new pin is tracked the day it is written. D4 records that customManagers is REPLACED not merged by an extending repo, so go-tool-base — the repo D3 was written for — must mirror it. D6 records that turning a dormant manager on is not a no-op: TFLINT_AWS_VERSION is interpolated bare into tflint HCL while upstream tags are v-prefixed, so it needs extractVersion or the first bump breaks the image build. | renovate presets (docker.json, default.json) |
| 2026-08-14 | 0066-renovate-group-log-cap | implemented | The nightly group scan failed 3 of the last 10 nights, ALWAYS after logging Repository finished for all 110 repos — so not the mid-run cut-off 0049 addressed — and the reason is unreadable: the trace hit GitLab's 4 MB collection ceiling (4,194,426 bytes, over by 122) and stopped collecting BEFORE the failure. Budget goes on 32 getChangeLogJSON error records, each serialising a full HTTP response plus a 20-frame stack, all GitHub 502 (11) / 504 (21) fetching release notes for aws/aws-sdk-go-v2. NOT auth or rate limiting: GITHUB_COM_TOKEN exists as a group var, no x-ratelimit header appears, and 502/504 are gateway errors not the 403 a limit gives. Fix is to stop losing the tail, not to log less: LOG_FILE/LOG_FILE_LEVEL/LOG_FILE_FORMAT (UNDOCUMENTED on the self-hosted-configuration page — confirmed by reading the shipped code in renovate/renovate:43, whose file stream writes SYNCHRONOUSLY specifically to survive process.exit()) plus an artifact with when: always, since without that clause it would upload only on the runs nobody needs. File level defaults to info, not Renovate's debug, because it uploads every run. D4 explicitly REJECTS suppressing the changelog errors — release notes are the most useful part of a dep MR and the errors are honest. The underlying non-zero exit remains unexplained by design; this makes the next occurrence readable rather than guessing. | renovate-group |
| 2026-08-15 | 0067-release-train-orchestrator | approved | The Go estate is a 76-module DAG EIGHT TIERS deep; a change to go/errors reaches 68 of 76 repos. Renovate is per-repository and stateless ACROSS repositories, so it cannot sequence a cascade — it opens hop N+1 only after a human has released hop N, and it emits incompatible intermediate sets because direct vs indirect is the boundary it uses while core vs adapter is the one that matters (keryx!269, phpbotscout!44 both broke exactly this way). Measured: upstream-first costs 69 releases in 7 ordered rounds, unordered up to 178. Collapsing adapter families into monorepos was measured and REJECTED — it removes width (76 nodes -> 36) but not depth (8 tiers -> 7), so sequencing is still required, and the separation is intentional anyway. New release-train component derives the DAG from go.mod EVERY run (a checked-in tier list would sequence confidently and wrongly), longest-path tiers, cycles are a hard error. plan is side-effect free; run walks the tiers driving a targeted renovate-group pass per tier, and REFUSES to execute on a scheduled pipeline because cutting releases is a per-train decision, never standing. First whole-estate rehearsal found go/artifacts — created the same day as a hand-enumerated sweep and therefore invisible to it, shipping on a Go with two known advisories. AMENDED 2026-08-25 with D9 after the FIRST REAL TRAIN, which was driven by hand from plan mode's ordering rather than by run and so skipped steps 1, 2 and 4: config v0.17.3 was tagged at 12:28, the first config-* adapter released at 13:02, and 23 OF 23 ADAPTERS TAGGED PINNING v0.17.2 — a version number carrying an upstream already six hours stale. The overnight scan then delivered the real pin and opened a second Release MR for every one; next morning 46 of 52 open Release MRs were pure dependency churn and 39 were a repo's SECOND release inside 24 hours. So the ordering was never the weak part: releasing an upstream does not put it into the downstream, and the only thing doing that was a nightly scan, which caps the cascade at ONE TIER PER NIGHT however fast a human merges. The walk had the same hole — step 2 landed a tier's bumps and step 3 released it immediately, but a merged bump moves the target branch and leaves the Release MR behind until releaser-pleaser regenerates it, so the release could omit the very bump the tier just landed. Now containment is MEASURED, never asked for: a hold after any bump lands (release_timeout, default 1800s — a queue wait, not a job runtime) plus a gate on every Release MR whether or not the hold ran. detailed_merge_status cannot answer this (it reported mergeable for 59 of 87 genuinely-behind MRs, cicd#35), so the test is compare(from=branch, to=target, straight) and an UNANSWERABLE payload RAISES rather than reading as current — failing open would rebuild the bug the check closes. The rehearsal could never have caught it: with execute=False step 2 merges nothing, so step 3 never met a behind Release MR; merge_one now reports would-merge so the hold is exercised, and on the first rehearsal after the change the new gate found afmpeg!186 and krites!176 behind main by 1 and 3 commits, both reported WOULD MERGE (green, mergeable) half an hour earlier. Same class as cicd#20 — a check reasoning from cached or derived state reports the cache, not the world. AMENDED AGAIN the same day with D10, after the FIRST EXECUTING run died in tier 0: step 1 TRIGGERS Renovate, Renovate opened ffmpeg-wasi!116 at 06:00:10, and the walk halted on it seconds later because its pipeline was running — halting on work the walk had just caused ITSELF. Not a corner case: any tier where Renovate opens anything halts by construction, so the engine had shipped in August with an executing path that had never once got past its own first side effect. Running is not a verdict but the absence of one, so the walk now waits (pipeline_timeout, default 2700s for go-tool-base's 25-minute e2e), re-reads mergeability after, and blocks ONLY on running out of patience — which it must, because a train that waits forever is no safer than one that stops and says why. The cross-cutting lesson, now its own section: A CLEAN REHEARSAL IS EVIDENCE THE TRAIN WILL START, NOT THAT IT WILL FINISH. A dry run changes nothing, so it reaches no state that exists only BECAUSE something changed — with execute=False step 2 merges no bumps (so D9's behind-Release-MR never occurs) and the walk starts no pipelines (so D10's own in-flight work never occurs). Worse, the two modes deliberately DISAGREED at exactly that point: a running pipeline is a rehearsal non-event and an executing fatality, and only the path the dry run could not take exposed it. Prefer read-only checks that run identically in both modes — which is why D9's gate, being read-only, found afmpeg!186 and krites!176 behind main on the very next rehearsal after reporting both green half an hour earlier. | release-train |
| 2026-08-17 | 0068-docs-verify-component | draft | Every quality component is change-detected against code paths, correctly, which leaves DOCUMENTATION with no gate at all: on krites!162 (the MR fixing three stale pages) the only jobs that ran were the security set and zensical-build. Gating the check on docs paths is the obvious design and is WRONG — krites' documented commands went stale because the CODE gained subcommands, so a docs-path filter runs the check exactly when the docs are already right and skips the MRs that break them. So docs-verify is always-on and offers NO changes input, with the self-test asserting its absence structurally. Specced and built CONCURRENTLY by two sessions (this and 0070) which reached the no-changes-input decision independently; the superset implementation shipped as cicd!217. | docs-verify |
| 2026-08-17 | 0069-npm-install-script-execution | implemented | Six of eight npm installs across the svelte-* components ran lifecycle scripts, so a hostile postinstall in any transitive dependency executed on the shared runner during ordinary build/test/lint jobs — and the unprotected installs run earliest and most often. Protection was per-invocation, which is why it drifted. Moved to the npm_config_ignore_scripts JOB VARIABLE so a ninth install site inherits it. Blanket application is NOT sufficient: measured on node 24, the variable SILENTLY SKIPS prebuild/postbuild while npm run build still succeeds, so a consumer generating types in prebuild would lose it with no error. Boundary drawn as 'dependency code must not execute at install', not 'no scripts' — the install is protected job-wide and first-party build commands re-enable explicitly. Both the component header and reference page had claimed the vector was already neutralised; true of two jobs, false of the pipeline. | svelte-build, svelte-test, svelte-lint, svelte-security |
| 2026-08-17 | 0071-retire-db-authenticated-fetch | implemented | retire fetches jsrepository-v5.json from raw.githubusercontent.com ANONYMOUSLY every run and fails the job on HTTP 429; three retries over ~40 minutes all 429'd, and it blocked an unrelated MR plus three in keryx. Applying the 0054 shared-cache pattern does NOT fix it alone — the nightly refresh still has to download, and anonymously hits the same limit. Measured: anonymous 429, authenticated GitHub contents API 200 (571,672 bytes), so the AUTHENTICATED FETCH is load-bearing and the cache is an optimisation. Three tiers, each logged, degrading rather than failing. --jsrepo not --cachedir (cachedir relocates retire's own cache but leaves it deciding whether to re-fetch). The fetcher is DETECTED not assumed — a first version hardcoded wget on the strength of checking node:24-alpine, but npm_image defaults to dev-tools which has curl and NO wget; it shipped red. Verified the gate still DETECTS with a local DB (exit 13 against a jquery 1.4.2 fixture) rather than passing quietly. | svelte-security |
| 2026-08-17 | 0072-opentofu-required-version | implemented | Renovate's terraform manager HARDCODES hashicorp/terraform as the depName for every required_version — not a default, unconfigurable in the extractor — so an OpenTofu estate has its core constraint bumped against a different product on a different version line (Terraform 1.15.8 vs OpenTofu 1.12.5). The resulting ~> 1.15.0 satisfies no OpenTofu release and broke tofu init/validate/plan on every stack in phpboyscout/infra. The ticket proposed DISABLING the bump; re-pointed instead via overrideDepName/overridePackageName, since disabling leaves a hand-maintained pin nobody tracks — the failure 0065 exists to prevent. Proven with a controlled --platform=local pair: without the rule ~> 1.10.0 resolves to ~> 1.15.0, with it to ~> 1.12.0. Also tracks .opentofu-version, which no built-in manager reads. Renumbered from 0068 after a concurrent session claimed that number. Corrected a FALSE belief held in 0065 and go-tool-base's config: customManagers/packageRules are mergeable:true and CONCATENATE across presets; labels/extends do not. | renovate presets (tofu.json) |
| 2026-08-17 | 0073-release-stamp | implemented | org 0001 D3 moved issue closure to MERGE — the right boundary, since a merged MR means no engineering action remains, but it drops which release actually SHIPPED the work. On cicd that gap is 0–2 days; on keryx it has run to a fortnight, during which a ticket reads done and the thing is not yet usable from a tag. NEITHER release tool can close it: releaser-pleaser's Forge interface has no comment method at all and release-plz's forge client has no /comments endpoint — feature requests upstream, not configuration. The chain IS derivable from Free-tier endpoints (compare -> commits/:sha/merge_requests -> merge_requests/:iid/closes_issues), verified live against v0.36.0..v0.37.0: 18 commits, 13 MRs, exactly the 4 issues that release closed. A COMMENT not a label (org 0001 D10 — a label needs a value per version, needs creating per release, and goes stale). IDEMPOTENT on a hidden marker, because a retried tag pipeline would otherwise re-stamp every issue. allow_failure defaults TRUE like discord-release: the tag and artefacts already exist by then, so a red release pipeline over a missing annotation is worse signal than a yellow one. Refuses rather than guessing when the tag is absent from the version-ordered tag list, and filters cross-project issues out by numeric project_id. Two real defects the live dry run caught that unit tests had not: closes_issues returns NO references field (the partition keyed on an invented shape and classified every real issue foreign), and an iid-only dedup key collides across projects. | release-stamp |
| 2026-08-17 | 0074-hardened-base-images | implemented | Measuring before deciding INVERTED the ticket's premise. dev-tools carries 2582 CVEs (2472 os-pkgs) and infra-tools 665 (399 os-pkgs), but the images' own gate is --severity HIGH,CRITICAL --ignore-unfixed and EVERY gate-visible finding in both — 66 and 158 — is lang-pkgs: Go stdlib and golang.org/x/* compiled into third-party release binaries (goreleaser-pro, golangci-lint, syft; terraform-docs, gitleaks, tflint-ruleset-aws, trivy, tofu). Not one OS package reaches the gate, because Debian rates its base CVEs below HIGH or ships them unfixed. So the base swap removes NONE of the churn #9 was raised to fix — that needs its own remedy (D8). It is still worth doing: a controlled pair built and scanned with the same package set took the OS layer from 2468 findings (20 CRITICAL, 235 HIGH) to ZERO. But it does NOT shrink the image — measured 257 MB on Wolfi against 176 MB on Debian, so a size claim is false. Wolfi over Docker Hardened Images: DHI went free (Apache 2.0, Dec 2025) so cost is not the discriminator; DHI's value is concentrated in DISTROLESS runtime images and a CI image cannot be distroless (it needs a shell, a package manager and compilers by definition), and DHI's glibc line is Debian, which keeps the archive being escaped. glibc is the hard constraint that rules out every musl option — cargo-binstall runs --disable-strategies compile (prebuilt-only, no fallback), and aws-cli v2 publishes no musl build. Change the base but NOT the acquisition mechanism: Renovate has no apk manager, so moving ~12 pins to apk add pkg=1.2.3 would blind them all — the exact failure 0065 exists to prevent. wolfi-base has ONE real tag (latest; the rest are cosign artefacts), so it must be digest-pinned with pinDigests or it regresses the just-landed floating-image fix. AMENDED 2026-08-18 to fold in DECOMPOSITION, whose case turns out to be stronger than the rebase's: attributing every megabyte to the components that use it shows all 66 of dev-tools' gate-visible findings sit in Go-track and Node tooling (goreleaser-pro 19, golangci-lint 17, syft 14, goreleaser 9, Node 7) and NOT ONE in the Rust toolchain — which is the largest thing in the image at 920 MB. So a Rust repo pulls 2.7 GB and inherits 66 HIGH/CRITICAL findings from tools it never invokes. Split by track into go-tools / rust-tools / node-tools / tofu-tools / docs-tools / ci-base: a Go repo drops 60%, Svelte 83%, Tofu 87%, an automation-only job 94% (discord-release pulls 2.36 GB TO RUN CURL AND JQ; release-stamp/release-train/renovate-merge pull 2.72 GB for stdlib-only python3). ONE repo, one release stream, several build targets — splitting the repository would mean six release streams and is the wrong half to split. Spec 0024's consolidation win SURVIVES because it was about BAKING (deleting runtime cargo/go installs), not about one-image-ness, and 0024 D3 already establishes the pattern by running each security scanner from its own upstream image. Three deletions needing no split: 141 MB of Go build cache is published inside dev-tools, aws-cli is 257 MB and NO component invokes aws (OIDC goes through the Tofu provider), and tofu-security is the last security component still scanning from the monolith (438 MB). D16 (2026-08-18): keryx and krites embed a Svelte SPA via //go:embed built by goreleaser's go generate before-hook, so their goreleaser job needs Go AND Node in one image. This does NOT block go-tools, because of HOW it fails: build-web.sh exits 0 when npm is missing and leaves the committed placeholder.html, and spaFromFS serves that placeholder — so a Node-less image would build, SIGN, NOTARIZE and PUBLISH a release with a placeholder UI, entirely green. That is a latent hazard TODAY; it only works because dev-tools happens to carry Node. Fix: make the generator CI-aware — outside CI unchanged (build if npm present, placeholder if not, the local-dev affordance), inside CI never build and FAIL LOUDLY if the prebuilt bundle is absent. No copy step needed: svelte-build's output_dir already defaults to embed, its artifact is **/embed, and GitLab restores to the same relative path, so the bundle lands in place — the CI branch is a verification. Verified: svelte-build's if default already includes $CI_COMMIT_TAG so it already runs on tag pipelines; pkg/studio/web/embed/* is gitignored except placeholder.html so the restored artifact is IGNORED not untracked and goreleaser's dirty-tree check still passes (krites lost v0.5.0 to exactly that check); and krites is unaffected by the split anyway since its goreleaser runs tags:[macos] on a shell runner. Side benefit: a broken SPA build is currently found only at tag time as a failed release — svelte-build on MRs moves it pre-merge. IMPLEMENTED 2026-08-19 as cicd v0.38.0: six images published, all twenty components moved, verified against the tag. Measured — ci-base 48 MB/0 findings, node-tools 112 MB (-89%), docs-tools 67 MB, tofu-tools 131 MB (-79%), go-tools 423 MB (-57%, 66->46), rust-tools 589 MB with ZERO findings of any severity, which is the empirical form of D9's argument that none of dev-tools' 66 came from Rust. Four asserted claims became measurements: the Wolfi base is clean; glibc is load-bearing (cargo-binstall prebuilt-only ran on it); the Rust toolchain contributes none of the findings; and RENOVATE OPENED THE ci-base BASE BUMP UNPROMPTED OVERNIGHT — the load-bearing claim under one-repo-per-image. Five defects the spec could not have predicted: Wolfi creates neither /usr/local/bin, /usr/local/sbin nor /usr/local/include though the first two are on PATH; SHELL must follow the apk add that installs bash; Go's .sha256 sidecar returns HTTP 200 and an HTML page; mktemp -t name-XXXXXX.py fails on busybox; and CHANGING THE USER A JOB RUNS AS INVALIDATES ANY CACHE THE PREVIOUS USER WROTE (npm's root-owned cache broke non-root node-tools with EACCES — fixed by changing the cache key, not by reaching for root). The near-miss worth keeping: release-stamp and release-train self-tests tolerate exit 1 to assert a documented refusal, and the mktemp breakage also exits 1, so both reported SUCCESS while the engine had never run — a failure-path test must assert the failure it EXPECTS, not merely that a failure occurred. | dev-tools, infra-tools → go-tools, rust-tools, node-tools, tofu-tools, docs-tools, ci-base; component image defaults; keryx + krites build-web.sh |
| 2026-08-20 | 0075-go-release-orchestrator | draft | Replace releaser-pleaser with a Go release orchestrator we own, rather than forking it. Three measured findings drive it. UPSTREAM IS NOT MAINTAINED: 2 human commits in the last 60, the most recent 2026-07-05 and it was the v0.9.0 release itself; everything since is renovate[bot]; our issue 463 and its accompanying PR 462 (+398/-1, mergeable) have sat unacknowledged for three weeks. THE BUG IT WILL NOT MERGE IS A CORRECTNESS BUG: GitLab 19.2 rebases before a fast-forward merge and does not write the result back to the MR record, so the recorded head stays at the pre-rebase commit; upstream picks MergeCommitSHA, else SquashCommitSHA, else that head, so FF-without-squash tags a commit not on the target branch — measured 9 of 231 tags across 62 projects, omitting docs/chore/ci/style/test commits. WE ARE NOT EXPOSED ONLY BECAUSE spec 0061 forces squash:true, which makes the squash commit on-branch by construction — verified 5/5 open release MRs carry it and 48/48 recent tags are reachable from main, so ensure_release_mr_squash: false re-opens the bug and the input's description should say so. AND WE ALREADY OWN THE PARTS: go/forge is 8,308 LOC + 7,499 test with four adapters, auth and release reads — roughly 4x the whole tool it would replace — plus go-tool-base's cmd/changelog and gtb sign. D1 build not fork (a fork's only gain is dropping the squash mitigation, which currently holds, and it buys nothing toward the goreleaser seam). D2 BEHAVIOURAL PARITY IS NOT A REQUIREMENT — we own all 96 consumers and define the interface; what replaces parity as the constraint is migration sequencing, since 96 projects release automatically and a new tool's bugs land as permanent tags, so rollout is cicd first, then the forge family, then fan out, never a flag day. D3 GitLab and Go ONLY, with Rust/release-plz (12 projects, and a materially healthier upstream at 1,453 stars) and the other three forges as NON-GOALS not deferrals — the unification argument is acknowledged and may become compelling, but scoping for it now is what would prevent it later. D4 the goreleaser seam is IN SCOPE: a tag cut by one tool and binaries attached by another leaves a window where a release object exists with no assets (a live broken krites update path, seen 2026-08-18), and owning both ends removes the seam rather than managing it forever — the strongest argument for greenfield over fork. D5 never trust a forge's MR record for what landed; resolve from the target branch. D6 merging a release MR stays a human action. BLOCKED ON go/forge#12 (merge-request lifecycle, GitLab first) — the only capability the orchestrator needs that the forge modules lack, and the core of what it does. Open: where it lives (gtb subcommand vs new module), how MR state is carried, the changelog contract, migration mechanics (a dry-run that agrees with the old tool before cutover), and releaser-pleaser.yml's deprecation. | releaser-pleaser, goreleaser; future new component |
| 2026-08-21 | 0076-playwright-tools-image | approved | svelte-test's playwright job has not run since the 0074 image split, and allow_failure reported the 12-second failure as a GREEN pipeline for three runs (krites: green through 08-18 17:56, failed 08-20 x3, two of them on main) — both e2e consumers, krites and keryx, shipped without e2e coverage and without a signal. THE CAUSE IS NOT THE ONE IT LOOKS LIKE: the visible error is su: must be suid, which reads as a non-root problem, but running as ROOT only moves the failure to sh: apt-get: not found — --with-deps CAN NEVER WORK ON WOLFI, because Playwright has no apk path, detects an unsupported OS, falls back to ubuntu24.04 and shells to apt-get. So 'run the job as root' is the obvious-looking fix that would have traded away a deliberate 0074 decision to buy nothing. MEASURED (du -sm /usr on node-tools:v0.1.0): baseline 321 MB; +26 apk packages for Chromium's runtime = 711 MB (+390); playwright chromium-headless-shell+ffmpeg 267 MB; full chromium+shell+ffmpeg 655 MB; Wolfi's own chromium package +1,080 MB — the last because it is a DESKTOP browser: only ~380 MB is chrome, the rest is Qt6/GTK4/systemd/x265/librsvg/two Pythons plus 192 MB of libLLVM so Mesa's llvmpipe can JIT shaders. D1 separate playwright-tools image, NOT a fatter node-tools: baking would take 117 MB -> ~780 MB, a 6.5x inflation of the image serving svelte-build/svelte-lint/vitest estate-wide to serve one opt-in job two projects enable, spending the split's headline -89% win; 0024 D3 already establishes per-tool images. D2 provide the libs via apk as root then drop to ci, and DROP --with-deps; the working 26-package set was found by walking the missing-library chain to a successful LAUNCH, not copied from docs — verified by chrome-headless-shell --dump-dom reaching Chromium startup (the dbus socket errors are benign in a container). Wolfi naming traps that each look like 'no such package': nss/nspr are libnss/libnspr (a plain nss exists and is NOT the one with libnspr4.so), atk is libatk-1.0 + libatk-bridge-2.0 (versioned in the package name), font-dejavu is ttf-dejavu. D3 KEEP FULL MESA: libgbm.so.1 is genuinely required, and mesa-gbm is a false economy — it still pulls libLLVM via mesa-libgallium, saves only 40 MB of 390, and adds a libudev.so.1 dependency the full package satisfies. D4 RESOLVED — bake chromium-headless-shell (267 MB) not full chromium (655 MB): pure CI, never headed. Needed checking rather than asserting because 'headless CI' != 'the headless-shell binary' — Playwright's chromium project historically resolved to the FULL browser and both consumers use the default channel with no channel: set (projects: [{name:'chromium', use:{...devices['Desktop Chrome']}}]). Tested directly with ONLY the shell installed against that exact config: 1 passed. Modern Playwright auto-selects the shell when headless on the default channel, so NEITHER CONSUMER NEEDS A CONFIG CHANGE. This holds because the versions match exactly, not approximately: krites and keryx both pin @playwright/test 1.62.1 and the test ran 1.62.1 — a version-dependent behaviour, which is what makes D5's pin load-bearing rather than tidy. D5 PLAYWRIGHT_BROWSERS_PATH + PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD so a consumer's npm ci cannot silently fetch a second copy; the image's Playwright version is the pin. D6 RESOLVED — SPLIT THE JOB IN TWO rather than dropping or parameterising allow_failure: playwright:preflight is hard-gated and launches Chromium through the CONSUMER'S OWN @playwright/test (the pairing that must hold, since Playwright refuses to drive a browser build it does not recognise), and playwright keeps allow_failure:true with needs:[preflight] so a broken browser means the suite does not run at all instead of running, failing and being forgiven. Verified negatively against an empty browser path: exits 1 in seconds naming the executable it wanted. Dropping allow_failure outright was rejected because it lands two neglected suites' accumulated drift on consumers' merges on day one; an input was rejected because it defers the decision, and the current default is what hid three failures. D7 (NEW, found while writing the self-test) — THE E2E GATE BECOMES AN INPUT BECAUSE IT WAS UNTESTABLE: svelte-test's self-test only ever covered the when: never disabled path, and could not do otherwise — the e2e rule was HARDCODED to $RUN_E2E on merge-request/default-branch pipelines, while a component self-test runs in a CHILD pipeline whose CI_PIPELINE_SOURCE is parent_pipeline. A hardcoded gating expression made the component's most fragile path permanently unreachable by its own tests — not why the job broke, but why nothing caught it. e2e_if now defaults to exactly the old expression (no consumer behaviour change) and the new tests/svelte-test/e2e.gitlab-ci.yml sets it true against a real browser. IMPLEMENTATION MEASURED on the built image: /usr 711 MB, browsers 267 MB, and the 27 Chromium runtime packages introduce ZERO new OS scan findings over the node-tools baseline. SEQUENCING: playwright-tools cannot publish its first tag until ci-base -> node-tools clears two INHERITED busybox HIGHs (CVE-2026-38753/4) that node-tools has been failing scheduled rebuilds on since 08-20; current wolfi-base already carries the fixed busybox 1.38.0-r1. Rollout note: krites/keryx e2e suites have not run since 08-18 and should be EXPECTED to fail on accumulated drift before they pass. | svelte-test, node-tools → new playwright-tools; krites, keryx |
| 2026-08-30 | 0077-a-colophon-component-and-the-swap-away-from-releaser-pleaser | implemented | Add a colophon component and begin swapping off releaser-pleaser. Supersedes the delivery half of 0075, whose decision moved to phpboyscout/colophon (all three verbs built, proven end to end against real GitLab and GitHub, v0.1.0 released 2026-08-30). D1 TWO JOBS AND publish MUST RUN BEFORE propose — a correctness constraint, measured not reasoned: after a Release MR merges the release commit is on the target branch but the tag does not exist yet, and IN THAT WINDOW COLOPHON COMPUTES THE SAME RELEASE AGAIN (verified on colophon at a817445: plan says 'No release due' with the tag present and '0.0.0 -> 0.1.0, 30 commits' with it removed), so a pipeline that proposed first would re-open a Release MR for the version about to be tagged and leave it open forever once publish tagged it. AMENDED 2026-08-31 — the spec's own D1 text said the ordering is 'expressed as two stages rather than needs:', which is a shape that was NEVER SHIPPED; the concern behind it is real (needs: overrides stage ordering and a release job must fall in behind the consumer's gates, per 0053) and is solved by putting needs: on the PROPOSE job only. AND THE SPLIT CANNOT BE COLLAPSED INTO ONE JOB, for a reason no spec stated at source and which is stronger than D1's original design argument: both jobs check out the SAME commit, and A COMMIT DOES NOT CARRY ITS OWN TAGS — publish tags via its own in-memory clone and pushes to the REMOTE, never writing into the job's checkout, while propose reads the last release from the tags in that checkout, which are whatever the runner fetched at job start. So propose sees the new tag ONLY because it is a separate job whose fetch happens after publish pushed: a TIMING property, not a cloning one, and one job means one fetch taken before the tag exists. Measured on sigillum v0.4.1 (publish finished 20:25:34, propose started 20:25:36; the release carried four fix(deps) commits, so 'Nothing to release' was only possible if propose saw the tag). Two changes this decision predates: the job is now colophon-propose-next (cicd v0.42.0), and colophon v0.1.3 carries ErrAlreadyReleased because NOTHING TESTED THE ORDERING — tests/colophon/ runs with a dummy token against a branch that cannot exist. Originally expressed as needs: on the PROPOSE job and never on publish — publish must stay free of needs: so it falls in behind the consumer's gates by stage ordering (0053), and propose inherits that transitively; a single job running both verbs was rejected because colophon separates them precisely so each blast radius can be stated in a sentence. D2 NO INPUT FOR ANYTHING COLOPHON READS FROM THE REPOSITORY (Release-As: trailer, Release-Note: trailer/block, .colophon.yaml placement/heading/hold) — a second place to say it is a second thing to disagree with; inputs cover only what CI owns. D3 $CI_JOB_TOKEN refused as for releaser-pleaser (a CI_JOB_TOKEN-pushed tag does not fire downstream tag pipelines). AMENDED 2026-08-31 — THE RENAME HAPPENED: the spec text said the default stays $RELEASER_PLEASER_TOKEN with renaming deferred, but the component's token input now defaults to $COLOPHON_TOKEN, that group variable exists (masked + protected) and neither adopter overrides it. The original reasoning was right and is why it was done as its own change ahead of any fleet move rather than alongside one — it cost nothing per repository because the token is a component DEFAULT plus a group variable: one change in cicd, zero consumer edits. RELEASER_PLEASER_TOKEN remains while releaser-pleaser still uses it. D4 PROVE ON COLOPHON, THEN ONE MORE, THEN STOP: 96 projects releasing through a tool that has cut two releases is not a rollout, it is a wager. D5 THE DIVERGENCE IS STATED IN THE COMPONENT, NOT DISCOVERED — colophon releases on perf and refactor and releaser-pleaser on neither; measured head-to-head over colophon's own history BOTH COMPUTE v0.1.0 and the changelogs differ by EXACTLY ONE ENTRY (the single refactor: commit), with nothing releaser-pleaser lists missing from colophon's. WHAT GETS SIMPLER: releaser-pleaser:verify, the forced squash: true and the second container image all exist to work around upstream picking the merge request's RECORDED HEAD as the release commit, which GitLab 19.2's automatic rebase never updates (9 of 231 tags off-branch, 8 repos, measured 2026-08-01); colophon resolves via ResolveMergedCommit against the target branch and forge.PullRequest carries no head-SHA field at all, so that class is unrepresentable and the component needs none of the three. IMAGE: a new phpboyscout/images/release-tools deriving from ci-base — NOT ci-base itself (the binary is 75 MB and ci-base is the base every toolchain image derives from, so it would land in node/rust/tofu/docs-tools which will never run it — the exact failure 0074 D10 undid) and NOT go-tools (colophon is forge-neutral and will release Rust and Node projects, which should not pull a Go toolchain to cut a tag). | releaser-pleaser -> new colophon; new release-tools image; colophon |
| 2026-08-31 | 0078-fleet-move-to-colophon | implemented | Move the fleet from releaser-pleaser to colophon, which 0077 deliberately left out of scope. SCOPE MEASURED THREE TIMES AND WRONG TWICE: 96 (predated several repos, never counted), then 103 (searched only for OUR component), finally 104 — scoutdm includes the UPSTREAM apricote/releaser-pleaser/[email protected] directly, so a grep for cicd/releaser-pleaser@ never saw it. A count is only as good as the pattern it greps for. scoutdm therefore never got spec 0061's forced squash: true, the only thing protecting this estate from upstream's off-branch-tag bug (0075: 9 of 231 tags off-branch), has NO TAGS at all, and still carries a renovate-self@ include the fleet otherwise stripped — it needs a look independently of this spec. Of the 104: 6 build binaries with goreleaser and are HELD (D2); 98 move under this spec (go/** 80, images/** 10, iac/** 5, top-level 3 = afmpeg, cicd, infra). Plus 12 on release-plz, 2 already on colophon, 6 with no CI file. D1 RUST IS A BOUNDARY, NOT A PHASE: the 12 release-plz projects are excluded permanently, because release-plz publishes to crates.io and colophon computes versions and writes changelogs — it does not publish packages, so moving Rust means building a publish path that does not exist to replace one that works. D2 THE SIX GORELEASER REPOSITORIES ARE HELD UNTIL COLOPHON SPEC 0004 LANDS (go-tool-base, keryx, krites, phpbotscout, scoutdm, skillup): 0004 is about to change how a release and its assets are produced for exactly that shape — colophon tags on the default branch, goreleaser builds and uploads to the package registry without creating a release, and colophon creates the release linking what was uploaded. Moving them now would move them onto a shape about to change and then move them again: two swaps, two windows for something to go wrong, on the six repositories where a release is most complicated. Holding costs nothing — they stay on releaser-pleaser, which works. THE 98 THAT REMAIN HAVE NO BINARIES AND NO ASSETS: their release is a tag, a changelog and a release object with nothing attached, which is the shape colophon has already done for two projects since 2026-08-30, so nothing in 0004 changes anything for them and nothing is gained by making them wait. colophon and sigillum, already moved, are BOTH goreleaser repos — they are 0004's pilot, not an exception. D3 THE PRE-FLIGHT COMPARISON RUNS ON ALL 98 BEFORE ANYTHING MOVES, and it is nearly free: 54 of the 104 currently have an open releaser-pleaser Release MR whose title names the version releaser-pleaser decided, so colophon plan (local git, ~100ms, no forge writes, no pipelines, no runner slots) can be diffed against it for over half the fleet in one sitting. Agreement is the baseline; a perf/refactor difference is the documented 0077 D5 divergence; ANY OTHER DIFFERENCE STOPS THAT REPO, and stops the move if the cause is general. The 49 without an open MR get a weaker check, AMENDED 2026-08-31 because the first draft got it backwards: requiring 'nothing due' would false-alarm on exactly the divergence the pre-flight exists to tolerate, since a repo whose only unreleased commits are perf/refactor has NO releaser-pleaser Release MR but colophon WILL propose one. Corrected to: nothing due, OR a proposal whose deciding commits are all perf/refactor; anything driven by a feat/fix that rp did not act on is a genuine discrepancy. Piloted against go/config, whose open Release MR says v0.17.4 and where colophon plan printed '0.17.3 -> 0.17.4 (patch), decided by 1 commit: fix(deps) ... yamldoc v0.2.2' — agreement in seconds, with the deciding commit and its bump named. D4 THE BOOTSTRAP CIRCULARITY IS REAL BUT THE ESCAPE HATCH IS A TAG, NOT A COMPONENT. colophon ships in images/release-tools, which is released by releaser-pleaser, and cicd defines the component and is also on releaser-pleaser, so the worry was that the tool needed to ship a colophon fix is the broken tool. AMENDED 2026-08-31 — THAT WORRY DOES NOT SURVIVE CHECKING: every artefact-producing job in the estate is gated on THE TAG EXISTING, not on which tool cut it (goreleaser if: $CI_COMMIT_TAG =~ /<tag_pattern>/; release-tools' image push if: $CI_COMMIT_TAG =~ /^v\d+\.\d+\.\d+$/). A release tool only decides a version and cuts a tag; binaries, image push and the latest retag all key off the tag alone. So recovery from a wholly broken colophon is one command any maintainer can run (v* is protected at Maintainers, and a human qualifies): git tag -a vX.Y.Z && git push. Carrying a whole component — image, CVE surface, docs, self-test, Renovate pins — as insurance against a one-command failure is a bad trade, so release-tools MOVES WITH EVERYTHING ELSE and the hand-tag recovery goes in the runbook. One thing to prove once on a throwaway tag before batch 1: colophon's publish creates the release object and goreleaser attaches assets, so whether goreleaser CREATES that object on a hand-tagged release is untested here (it normally does). D5 BATCHES ARE RISK-ORDERED AND THE RUNNER FLEET SETS THEIR SIZE — seven concurrent slots estate-wide and a pipeline per branch push, so a 103-repo sweep is a daytime denial of service. AMENDED 2026-08-31, REVERSING THE FIRST DRAFT: cicd moves FIRST as the canary, not last. Measured release cadence over 90 days — cicd 65, go/config 28, go/forge 26, go/output 6 — so cicd releases ~10x a typical library and would surface a colophon defect within a day, where go/output might take weeks and 75 repos would already have swapped; a repo that has swapped but not released has proved nothing (D8's own logic applied to sequencing). The downside that justified 'last' does not survive inspection: consumers pin the component by tag so a bad cicd tag affects nobody automatically, and if cicd cannot cut a release the recovery is hand-pushing one tag. Order: cicd, then 5 go/config-* adapters with open Release MRs, then the rest of go/, then iac/ + top-level, then images/** except release-tools. BATCH SIZE IS MEASURED, NOT GUESSED: a swap edits .gitlab-ci.yml, which is IN the components' rules:changes path lists by design, so it fires the FULL suite — measured on the v0.42.0 pin bumps at 4 jobs/0.5 min (colophon) and 7 jobs/22.7 min (sigillum, where go-cross-arch, go-mutation and go-fuzz all ran). At ~23 min per Go repo, 103 repos is ~39 hours of runner time, ~5.6 h wall clock across seven slots; the escape-hatch fast-merge does NOT rescue it, because it can skip the MR pipeline but not the post-merge pipeline where colophon-propose-next runs, which is the signal being bought. So the go/** batch splits into groups of 15-20, off-hours. D6 ROLLBACK IS CHEAP BEFORE THE FIRST PUBLISH AND EXPENSIVE AFTER — reverting an include is one line, but v* tags are protected, may already be consumed by a downstream pin, and a release object may carry assets; so the strategy is prevention, in three existing layers (D3's pre-flight, colophon-propose-next proposing rather than publishing so a wrong version is an MR a human reads, and colophon v0.1.3's ErrAlreadyReleased). D7 THE TOKEN RENAME IS ALREADY DONE — AMENDED 2026-08-31, having originally argued it should be done first. Verified: the component's token input defaults to $COLOPHON_TOKEN, the group variable exists (masked + protected), no adopter overrides it, and 0077 D3 has been corrected to record it. The reasoning stands and is why it was cheap: a component DEFAULT plus a group variable is one change in cicd and ZERO consumer edits, where doing it after the move would have meant touching 103 repositories twice and doing it during would have given every failed swap a second candidate cause. So the 103 swaps land on a correctly-named token with nothing further to do. D8 releaser-pleaser IS KEPT AS A LEGACY COMPONENT — AMENDED 2026-09-02, reversing the original decision that it is removed once unused. Matt's call, in his words: 'it's a good tool, just not our tool', so deletion is low priority and it stays. THE REASONING THAT CHANGED: 'unused' is a state this estate reaches, but reaching it is not an argument for deletion. releaser-pleaser works; the reasons for moving off it were specific to how THIS ESTATE MERGES — fast-forward without squash, where upstream can pick a merge request's RECORDED HEAD and tag a commit that is not on the target branch (9 of 231 tags measured off-branch). That makes it wrong for us, not wrong. Components are published by URL, so one nobody here includes costs nothing to leave in place and may serve somebody who is not us. It is not maintained ahead of need either, but spec 0061's forced squash: true stays, because that is the mitigation making it safe under this estate's merge method. The release-cycle criterion is dropped: nothing now waits on every migrated repo having released. D9 0077 D1 MUST BE AMENDED BEFORE IT LEAVES DRAFT — it states the ordering is 'expressed as two stages rather than needs:', and the shipped component does the opposite (needs: on propose, none on publish, per 0053); the implementation is correct and the spec text is stale. Also records two changes 0077 predates: the job is now colophon-propose-next (cicd v0.42.0), and colophon v0.1.3 carries ErrAlreadyReleased because NOTHING TESTED THE ORDERING — tests/colophon/ runs with a dummy token against a branch that cannot exist, so it cannot observe it. OQ1-OQ3 all RESOLVED 2026-08-31 by measurement (see D3 and D5). OQ4 RESOLVED 2026-08-31 BY REMOVING ITS PREMISE (see D4): the escape hatch need not be a retained component, because every artefact job keys off the tag rather than the tool that cut it, so recovery is a hand-pushed tag; no repo is permanently retained and D8's 'removed when unused' is reachable. Leaves one small thing to verify once: that goreleaser creates the release object on a hand-tagged release. IMPLEMENTED 2026-09-03: measured 106 projects on the colophon component (the 98 movable, the 6 held, the 2 already there), 1 deferred (go/objectstore, no Go files), 12 on release-plz by D1. The D3 pre-flight ran across all 98 before batch 1 (cicd !270: 50 agree, 48 nothing due, 0 diverge). D2's six moved once colophon 0004 landed; keryx and skillup have since released with assets. D4'S PROOF WAS NEVER RUN AND IS SUPERSEDED: colophon's publish starts from the last merged release request, so a hand-pushed tag yields a tag and NO release object — colophon spec 0005 (release from a tag) delivers the verb and the runbook, tracked in cicd#41. scoutdm's leftovers are scoutdm#6. | releaser-pleaser, colophon; 103 consumers |
| 2026-09-03 | 0079-pipeline-churn-on-three-runner-slots | implemented | The AWS runner fleet is off after a $270 bill; runner1's slots serve the estate. FAST-FORWARD STAYS. The founding hypothesis was wrong and the measurement is the point of this spec (reports/2026-09-03-mr-pipeline-minutes, 3,373 MR pipelines / 31,247 jobs / 21 days): the heavy specialist jobs are 5.7% across three repos, while Renovate branches are 49.6% — cicd component-pin bumps alone 23.4%, release MRs a further 20.4%. D1 prConcurrentLimit: 2 (its original quadratic rationale later corrected — sibling rebases measured at 1% of minutes). D2 skip the Catalog and pin the fleet to @main (OQ1: the Catalog needs a release: job, and colophon creates releases via the API). D3 five items, all closed: CHANGELOG.md out of go-test/go-lint changes (reverses 0051); heavy jobs INVERTED — they stay on MRs and the default-branch re-run went instead, with go-fuzz to nightly schedules; ordinary Go gate left alone; ffmpeg-wasi's matrix gated by profile. D4 explicit interruptible. D5 runner1 8 vCPU, GOMAXPROCS=4, concurrent=2. D6 one-job release-MR pipeline. D7 shared Go cache on the runner (go-test 60s→18s, golangci-lint 60s→17s). D8 coverage read from the test run, never a second one. D9 records two rejections (skipping security jobs on bot branches, and spreading the nightly Renovate burst) so the same numbers do not produce the same proposals twice. OQ3: Premium not bought — push rules were the strongest case and were rejected because a push rule binds every push while the estate's AI-attribution rule binds only Matt and those under his direct supervision. Outcome: ~247-430 MR min/day vs 888 baseline; queue p90 509s→116s. Watch the tail, not the peak. | renovate presets; every component; runner1 |
| 2026-09-04 | 0080-go-singleuse | implemented | go/controls spec 0005 D9 decides a cicd component runs the singleuse analyzer in consumers, advisory, so a finding is visible and fails nothing. Records the component's shape only: a version input for the nested lint module, a dir input for a consumer that is itself a nested module, a tolerated exit code rather than a blanket allow_failure, and a success-path self-test. GOTCHA: go run exits 1 for any non-zero child, so allow_failure: exit_codes: [3] never fires — go install into GOBIN and run the binary. | go-singleuse |
| 2026-09-05 | 0081-a-tool-bump-that-changes-an-image-must-release-it | implemented | Renovate types a baked-tool bump in an images repository as chore(deps), and colophon does not release a chore — so the tool lands on main, no tag is cut, no image is published, and every consumer keeps pulling the old one while main looks current. Measured on release-tools, which shipped colophon 0.2.1 for a day after 0.3.0 was merged. A dep read from a Dockerfile under images/** now commits as fix in the :docker preset; scoped to images/** so scoutdm (never released) is excluded. GOTCHA: a local Renovate dry run CANNOT test this — it reports the repository as local, so matchRepositories never matches; exercise Renovate's own dist/util/package-rules/index.js matcher instead. | renovate presets (:docker) |
| 2026-09-05 | 0082-pins-nothing-tracks-and-retired-fixture-images | implemented | Issues #24/#27/#28/#30 claimed Renovate tracked no per-track image pin; the dashboard's Detected Dependencies says it tracks all eight across 30 template files, so !282 had fixed them and nobody closed them (now closed with that evidence). What is genuinely untracked is different: D1 go-singleuse's version pins the nested go/controls/lint module — pin v0.7.0 and track phpboyscout/go/controls with the lint/ prefix stripped, noting that the prefix exists only because colophon's tag_also writes it, so the manager is coupled to colophon's tag naming the same way the 0074 rename orphaned every image pin. D2 retire and cyclonedx are npx --yes fetched per run, so their pins are live and unmanaged. D3 cargo_deny_version stays UNTRACKED deliberately — its own description says it is ignored on the default rust-tools image, so bumping it would prove nothing. D4 seven self-test fixtures still pin dev-tools/infra-tools, retired by 0074 and untagged since August, because tests/ is in no manager's file patterns — repoint each to its component's own default image, THEN add the paths (adding them first would raise seven MRs bumping dead images). D5 each repointed fixture is verified by running it. IMPLEMENTED 2026-09-05 (!303, !305) with three corrections worth carrying: D1 used the go datasource rather than gitlab-tags+extractVersion (the Go proxy already maps lint/v0.7.0 for a nested module); D2 was short by one — there are THREE npx tools, lockfile_lint_version having been eaten by a greedy block regex, and all three were stale; and D4's "use the component's default image" rule was wrong, because these are fixture ASSERTION jobs rather than the component's job, so the rule is the lightest image carrying what the job needs (only go-test needs a Go toolchain; the rest took ci-base). D5 justified itself by finding that check:engine-tests in tests/tofu-deploy-generate had NEVER run — needs: with no rules:, silently dropped from every child pipeline, 56 unit tests dormant. AMENDED 2026-09-11 and FIXED 2026-09-14 (cicd#45): D4's tracking half never matched anything for the nine days it was in the file. config:recommended supplies an ignorePaths containing **/tests/**, and ignorePaths is mergeable: false, so a repository inherits the whole list or replaces it - adding tests/ to a manager's file patterns while that is inherited is configuration that reports success and covers nothing, the same shape as 0074's orphaned pins and as the ratchet 0085 describes. The fix sets ignorePaths explicitly to the standard list minus **/tests/** and disables npm, gomod, terraform and gitlabci under tests/**, because a fixture is tracked for its IMAGE PIN and nothing else. The amendment's own measurement was wrong on one point and it needed a SECOND rule: gitlabci is not what produces the duplicate branch. It is still worth disabling, for a different reason proved by probe rather than argued - it reads ANY image: in a fixture's .gitlab-ci.yml, not only ours, and a temporary image: node:20.0.0 in a fixture raises two branches without that entry and none with it, an effect that had to be provoked because it changes nothing about the repository as it stands. That comes from default.json, this repo's own shared preset, whose fourth custom manager matches image: in ANY .gitlab-ci.yml and names the dep by full registry path; it exists for consumers and cannot be narrowed, so it is disabled here BY DEP NAME rather than by manager, since it and manager 0 both report as custom.regex. Measured before/after on Renovate 43 with the baseline re-run rather than quoted: 25 to 29 files with updates, 0 to 4 of them under tests/, 0 to 43 deps disabled, 8 to 10 branches. The two new branches are a real go-tools fixture pin and an OSV alert on vitest in a fixture, which is DELIBERATE and cannot be rule-suppressed because vulnerabilityAlerts.enabled takes precedence over enabled: false. Proof it was not theoretical: !316 was already red, the tofu-tools pin having moved in twelve files and not in the two generator goldens under tests/, and the fixed config puts those goldens in the same branch. Two measurement traps recorded: a disabled dependency is still logged with a branch name, so is disabled does not prove a branch went away; and under --platform=local the repo config is read from the checkout while RENOVATE_CONFIG_FILE is the GLOBAL config, so pointing it at an old copy gives a false baseline that matches the fixed run. A FOURTH untracked copy surfaced the same afternoon, found by the bump it broke: tofu-deploy-generate keeps its engine in TWO places, scripts/tofu-deploy-generate/generate.py and a copy embedded in its template, both carrying a DEFAULT_IMAGE literal, and scripts/ was in no manager's file patterns. Renovate moved the template's copy and the goldens and left the script's, so the standalone engine emitted v0.1.0 against goldens carrying v0.1.2 and check:engine-tests went red on a pin bump that touched neither. The two-engine design is what CAUGHT it: the goldens are the shared contract, tofu-deploy-generate-verify checks the template's engine against them and check:engine-tests checks the script's, so they cannot drift silently. Only the tracking was missing, so /^scripts/.*\.py$/ joins manager 0; measured, the tofu-tools branch goes 14 files to 15, gaining that one file and nothing else. An audit of every image literal in the repo finds pins in four directories only - templates/ 18, tests/ 12, docs/ 2, scripts/ 1 - and that was the sole file outside a tracked pattern. THREE untracked-pin defects in ten days, each found by a different accident (0074's rename by a manual audit, tests/ by a spec review, scripts/ by a red pipeline). The lesson is not any one pattern: ask what in the repository carries a version and which of those a manager SEES, which is one grep and would have found all three. | go-singleuse, svelte-security, rust-security, renovate presets, all self-test fixtures |
| 2026-09-07 | 0083-a-merge-time-apply-job | implemented | Implements cicd#44 (a colophon-session handover) and through it colophon spec 0012 D12. colophon 0.6.0 ships apply, a verb that runs when a merge lands rather than when a release is cut: it reads the commits a push added, takes their Updates: trailers and carries each to the issue, merge request or wiki page it names (shipped / implemented, a closed vocabulary; on a wiki page it sets one frontmatter field and nothing else). It is what colophon#14 asked for and the durable fix for the spec-status drift 0082 documented, and the estate cannot reach it because templates/colophon.yml has no job for it. D1 a SEPARATE colophon-apply job rather than a flag on publish/propose, whose span, exit-code contract and ordering it shares none of — and FORCED rather than preferred, because colophon-publish is interruptible: true while D3 requires interruptible: false, and one job cannot be both. D2 the apply input defaults to TRUE, fleet-wide — every project in the estate uses specs and tickets, and being selective about which of them can say so is harder to manage long-term (Matt, 2026-09-07); the estate already ran the opposite experiment, where per-repo renovate-self currency depended on a schedule 51 of 53 repos never added. Cost measured first: 1,199 default-branch merges in seven days (~171/day across 128 repos, 44% of them bot branches that can never carry a trailer), ~85 min/day, which is ~2% of the three-slot capacity believed available at the time and ~3% of the two slots actually running. An earlier draft reported 20-35% by comparing against MR-pipeline minutes rather than slot capacity — the wrong denominator, and it nearly bought a per-repo opt-in nobody wanted. D3 interruptible: false, unlike every other job here, because cancelling LOSES DATA: the span is bounded by the push that started it, so a superseding push's run never covers the cancelled one's commits and their assertions are dropped rather than deferred. D4 its own resource_group, not colophon's — it writes to forge objects, not the repository, and the upstream idempotency key is explicitly not atomic. D5 GIT_DEPTH: "0", because history.Between resolves both span endpoints in the local checkout and GitLab's default depth of 20 makes that a latent failure on a busy branch. D6 the default image moves to release-tools v0.1.8 in the SAME release, since apply needs colophon >= 0.6.0 and shipping the job against v0.1.7 fails on an unknown verb. D7 reuse the existing token input, naming what a project-scoped token cannot reach (cross-project targets fail and the run still exits zero). D8 omit needs: rather than needs: [], which would ignore stage order and run before the consumer's gates. D9 allow_failure: true via an input, matching release-stamp — asked for by cicd#44 and missed by the first draft: apply exits non-zero only when it cannot START, and by then the merge has already happened, so there is nothing in the pipeline to go back and fix and a red default branch nobody can act on is how a red default branch stops meaning anything. OQ1 asks whether the 44% of runs on bot merges can be skipped, and records the two obvious gates as REJECTED rather than open: $CI_COMMIT_MESSAGE =~ /Updates:/ and a $CI_COMMIT_AUTHOR test both read the TIP commit only, and under fast-forward a push carries every commit on the branch, so a trailer on an earlier one would be dropped silently — D3's failure reached from the other direction. OQ2 colophon's pipeline-requirements.md was not updated for the new verb; OQ3 whether the component wants a dry_run input, deferred because apply has no --dry-run flag to drive one yet. IMPLEMENTED 2026-09-08 (!313) with one decision changed and one confirmed. D6's floor is release-tools v0.1.9, NOT v0.1.8: colophon 0.6.0 shipped the apply verb but could not START it in a repository with no colophon config file — which is most of them — because publish/propose were exempt from the missing-config gate and apply was not (colophon#15). The honest floor is the release that fixed it, and upstream returned better than the ask: the exempt list is now DERIVED from the registered verbs, so the next verb cannot repeat it. D9's self-test override is what caught it — the job ships allow_failure: true, so under the shipped default this would have been a yellow job in a green pipeline, merged, then yellow on every push to main in 106 repositories; the fixture sets apply_allow_failure: false on the argument, made before there was anything to catch, that a check which cannot fail is not a check. Confirmed in the trace: colophon-apply succeeded with allow_failure FALSE, the span resolved under GIT_DEPTH 0, and the all-zero CI_COMMIT_BEFORE_SHA was read as the tip alone with no special-casing. Also turned up, and not this spec's to fix: Renovate ratchets DOCUMENTARY version floors (it rewrote the assets row's historical minimum from v0.1.7 to v0.1.9, because the manager cannot tell a current pin from a recorded minimum), and the estate has TWO concurrent job slots, not the three AGENTS.md claimed until 2026-09-14. Related and NOT implemented here: cicd#43 (widen tag_pattern for component tags, teach the release check to read a component changelog) touches the same file, and states in its own text that it SUPERSEDES cicd#42, which is still open and should be closed rather than implemented. | colophon |
| 2026-09-07 | 0084-component-tags-and-the-release-check | implemented | Implements cicd#43 and through it colophon spec 0011 D5. colophon 0011 gave a declared component its own tag and its own release cut from its own changelog; two assumptions in templates/colophon.yml predate that. D1 tag_pattern widens from strict root semver to one RE2 admitting the root plus the three shapes colophon's tag templates produce (<dir>/vX.Y.Z, <name>-vX.Y.Z, <name>@X.Y.Z) — which REVERSES 0083's description of the same input, corrected in the same MR rather than left to contradict the template. D2 the component passes $CI_COMMIT_TAG through and takes NO view on which tags carry a release object: that view lives in colophon, where 0011 D6 makes release --tag report and exit 0 for a tag carrying none, and duplicating it here would be a second place to disagree with the manifest — which is also why cicd#42 (skip every tag containing /) was wrong and is closed as superseded. D3 colophon-release-check reads the tag list from the request title (colophon 0010 D5) instead of grepping one fixed file. D4 a tag maps to its changelog BY HEADING rather than path arithmetic — a heading carries the whole tag (0010 D6), so searching the already-computed changed-file list handles all three shapes with no knowledge of any of them, where deriving the path would need the component to read the manifest. D5 the check stays INFORMATIONAL: the version line echoes and does not fail, and making a missing heading fatal would fail the legitimate cases 0011 D6 names, on the job that gates every release in the estate. D6 goreleaser's tag_pattern is untouched — colophon 0008 D6 defers per-component assets, so widening it would start a build job with nothing to build. Sequencing VERIFIED rather than assumed: the ticket says do not land before the image ships 0011, and 0011's write-side commits are ancestors of colophon v0.5.0, carried by v0.6.0, baked into release-tools v0.1.8. OQ1 whether a titled tag with no changelog heading should fail; OQ2 whether an edited MR title desynchronises the tag list. REVIEW (first pass, NOT READY) caught a BLOCKING defect carried straight from cicd#43's suggested pattern: the RE2 contains a literal /, the value lands inside rules:if: ... =~ /.../ where / delimits the regex literal, and GitLab validates EVERY rules:if at parse time, so the unescaped default would stop the pipeline starting for every consumer rather than merely misbehaving. Confirmed against the real POST /ci/lint: unescaped INVALID (rule if invalid expression syntax), escaped \\/ VALID, current default VALID. Review also found the sequencing section misattributed the v0.1.8->v0.1.9 floor correction to 0083 (0083 shipped at v0.1.9 because 0.6.0 could not start apply), that TWO places carry the stale "nested tags are not releases" claim (the reference page AND the input's own inline description, which additionally says to keep the default aligned with goreleaser's, made false by D6), and that OQ2 was answerable from source: propose re-renders the title on every push to the target branch and compares byte-for-byte, so a hand-edited title does not survive, making the question moot. IMPLEMENTED 2026-09-14 (!325), every decision as written and D1 on the escaped form it gained in review. The new check's shell carried TWO defects on its first run, both found by exercising it against fabricated request titles before committing rather than by CI: the tag-list extraction was over-escaped and matched nothing, so every title reported no tags, and a title that is not a release title parsed as a tag named title, because the extraction was not anchored to the whole prefix. Fixed with grep -F for the heading search and s/^chore(.*): release //p for the title parse, then re-exercised across five titles: root only, root plus a component, component only, a shape carrying @, and a non-release title. A check whose output is INFORMATIONAL by D5 is exactly the kind that can be wrong for a long time without anyone noticing, which is the argument for running it by hand against made-up input before it ever sees a real one. Two stale claims were corrected rather than one: the reference page and the input's own INLINE description both carried the "nested tags are not releases" text D1 reverses, and the inline one additionally told the reader to keep the default aligned with goreleaser's, made false by D6. A claim repeated in the template outlives one repeated only in the docs, because the next person to edit the input reads the description and not the reference page. NOT yet exercised against a real component tag: no repository in the estate both declares components and ships assets, which is why nothing was broken before this and also why D1's shapes are verified against POST /ci/lint and against the tag list rather than against a pipeline that cut one. The first repository to declare a component is the real test. | colophon |
| 2026-09-12 | 0085-a-documented-floor-is-not-a-pin | implemented | Spec 0082 put docs/reference/**/*.md into the image managers' file patterns so a reference page's inputs table stays in step with its template, which works: 36 of the 38 matches Renovate actually sees are current pins. Two are not. docs/reference/automation/colophon.md records a MINIMUM beside assets and apply ("requires colophon >= 0.2.0, image at least release-tools v0.1.7"), and manager 0's docs matchString was unanchored, so it rewrote those floors on every image release. The assets floor has been ratcheted TWICE (aecab53 on 2026-09-07, restored by hand during !313, bf75ae2 again on 2026-09-08) with nothing reporting either, which is the argument for fixing the manager rather than the prose. It degrades one way only: a ratcheted floor always overstates, so the failure is a consumer concluding they cannot use a feature they could have used all along. D1 anchors the docs matchString to the inputs-table row shape (name cell, then string, then the backticked image) the way managers 5-8 already are, using [a-z0-9_]+ because this repo names inputs with digits (e2e_command, svelte-test.md:83); measured 38 matches before and 36 after, losing exactly the two floors, gaining none, with no file+line matched twice. D2 writes a floor so it cannot look like a pin (release-tools v0.1.7, no backticks, no colon) and restores the assets floor to its true value. BOTH, not either: D1 alone leaves a prose form one regex change from being captured again, D2 alone leaves the manager free to grab the next documentary tag anyone writes. D3 states the rule generally: a version in documentation is a claim about the past unless it sits in the inputs table, which Renovate owns. Managers 1 and 3 are looser than 5-8 but far tighter than 0 (1 wants three co-occurring tokens, 3 wants two and carries no image token, its depNameTemplate being fixed). NOT here, split out after review: the discovery that config:recommended ignores **/tests/** so spec 0082 D4's tracking half has never matched anything - that is an amendment to 0082's Outcome plus OQ3, tracked as cicd#45. IMPLEMENTED 2026-09-12 (!323); verified after the edit at 36 matches with neither floor line touched. Three review passes, NOT READY each time, and every finding real: a measurement over a scope Renovate ignores, an anchor excluding digits against a naming convention already in the repo, two wrong claims about which managers are anchored (the second written while fixing the first), a retracted figure left standing in an earlier section, and a quote of a file that had changed underneath the spec. THE assets FLOOR WAS RATCHETED A SECOND TIME WHILE THIS SPEC WAS OPEN ABOUT THE FIRST (bf75ae2, 2026-09-08), which is the evidence that a hand-correction is not a fix. Generalised rule recorded in the Outcome: three of the five findings here and in 0082's amendment are the same shape, configuration or a measurement that reports success while covering nothing, so count what a pattern matches and name the scope the count is over. | renovate presets, colophon docs |
| 2026-09-15 | 0086-a-project-that-names-its-own-assets | implemented | Implements cicd#49, raised by a session working colophon and ffmpeg-wasi. colophon-release always passes --assets "$[[ inputs.dist ]]" and dist defaults to dist, so a project that names its release files in .colophon.yaml (colophon >= 0.10.0) inherits a goreleaser-shaped default it cannot satisfy and is refused at the end of its first release pipeline with an error naming a flag it never passed. THE FIX IS SMALLER THAN THE TICKET ASSUMES, and the finding is what shrank it: --assets "" is IDENTICAL to omitting the flag, because colophon declares it flags.StringVar(&opts.Dist, "assets", "", ...) with an empty zero value (read at v0.10.1, pkg/cmd/publish/cmd.go:93 and pkg/cmd/release/cmd.go:80). So the ticket's suggested direction, have the component skip --assets when dist is empty, is a NO-OP - the component would emit a different command line and colophon would behave identically. The ticket's "workaround" is therefore the correct configuration, and the defect is that nothing says so. D1 assets STAYS A BOOLEAN and the tri-state true/false/named is REJECTED on a count rather than a principle: eight repositories pass assets: true as a YAML boolean (colophon, ffmpeg-wasi, go-tool-base, keryx, krites, phpbotscout, scoutdm, skillup, swept across all 138 non-archived projects), retyping to string for an options: list asks all eight to quote a value they write unquoted, and it buys no behaviour dist: "" does not already express. D2 fix it WHERE A CONSUMER LOOKS - the dist description, which is the input they must change; the assets description, which is what they turn on first; and the reference page, which gains both plus a worked example beside the goreleaser one - stated as a supported shape rather than as a caveat. D3 the component adds NO VALIDATION of its own, on 0084 D2's argument reused unchanged: it would have to read .colophon.yaml and its reading could drift, making a second place to disagree with the manifest. The poor error message belongs to colophon, where there is one copy of the check. Nothing is sacrificed by that because colophon-release runs with needs: goreleaser, so any check it performed would already be after the build. Verified at the consumer: ffmpeg-wasi runs exactly the configuration D2 documents, and needed a nine-line comment in its own CI file to explain what the input description now carries. The self-test cannot prove this one - tests/colophon/ is a goreleaser-less library that does not set assets: true, and D1 and D3 change no behaviour, so a green pipeline here would say nothing. IMPLEMENTED 2026-09-15 (!338), all three decisions as written. OQ1 is closed by colophon#27, which asks for a message naming both sources it found and where each came from rather than naming a flag, and records the asymmetry that makes the current wording hard to act on: the manifest list is deliberate and the dist arrived from a default the consumer never wrote. THE STATUS FLIP USED THE TOOLING AND THE IMPLEMENTING COMMIT DID NOT. Spec 0083 shipped colophon-apply on 2026-09-08 so that a merge moves a spec's status, and 03c3fd1 carries no Updates: trailer so it moved nothing; the correct form, read from colophon's own tests at v0.10.1, is Updates: cicd/wikis/specs/<slug> implemented, and it has to be on the IMPLEMENTING commit before it merges because the commit is immutable afterwards - the same timing trap that governs Release-Note:. This flip carries the trailer as a demonstration: apply set the frontmatter field, while specs/home and the Outcome were still written by hand, because it sets one field and nothing else. Worth recording because 0082 documented that 59 of 82 specs sat APPROVED with every one shipped: the rule alone did not fix it, the tooling was built, and the first spec written after the tooling existed still forgot to use it. | colophon |
| 2026-09-15 | 0087-the-unit-test-jobs-timeout | implemented | Implements cicd#34, raised from afmpeg whose -race suite sits at 70-110% of the ceiling. go-test's UNIT job ran go test -race with no -timeout, so Go's 10m default applied and could not be changed, while both sibling jobs in the same file expose one (e2e_timeout 5m, integration_timeout 10m). The failure is loud in a MISLEADING way: a suite that runs out of clock prints a goroutine dump, mostly inside wazero's register allocator, that reads exactly like a deadlock and is not one. D1 a timeout input, named BARE not unit_timeout, because the file's convention is that the unit job's inputs are unprefixed (paths, coverage_regex) and only the other two jobs prefix theirs - a prefix here would imply a fourth job. D2 the default is 20m, NOT Go's 10m. The ticket framed keeping 10m as "nothing changes for anyone", and that does not survive the evidence it did not have: all three of its passing runs were on the AWS SPOT FLEET, switched off since 2026-09-03 after a $270 month (0079), so every one ran on a path that no longer exists. runner1 is now the whole Linux estate and pins GOMAXPROCS=4 on an 8-core 7GB box with concurrent=2 - not an accident, it is what fixed four OOM kills when the VM went 4 to 8 cores (0079 D5). So -race, the most parallel and memory-hungry thing Go runs, is permanently capped at half the cores on the slower machine, sharing it. The workload did not get slower; the hardware did, by choice. D3 A HANG COSTS THE INPUT'S VALUE, NOT THE JOB TIMEOUT, recorded as a decision because the ticket's own counter-argument was WRONG BY A FACTOR OF FIVE and nearly bought the wrong default: it said raising the default means a hung test burns the 1h GitLab job timeout, but go test -timeout 20m still panics at 20m with its dump, and the 3600s project timeout (checked on cicd, afmpeg and go/config) is only reached if -timeout is removed entirely. The real cost of 10m to 20m on a hang is TEN EXTRA MINUTES. With the cost stated correctly the trade is one-sided: too low fails misleadingly and fast, sending the reader to debug something that is not happening; too high fails honestly and slow, and for a gate a wrong diagnosis costs more than a slow one. Testing strategy is explicit that the self-test CANNOT prove the timeout - a fixture that sat for twenty minutes would be the most expensive job in the repo and would prove only that Go's flag works - so the fixture overrides it to 2m to exercise the INTERPOLATION, which a broken $[[ inputs.timeout ]] fails immediately on an unparseable duration. OQ1 whether e2e_timeout and integration_timeout want the same re-examination, left open because nobody has hit either and raising a limit nobody has hit is how a timeout stops meaning anything; OQ2 whether the estate wants a shared statement of what a CI timeout is for. IMPLEMENTED 2026-09-15 (!341), all three decisions as written. THIS SPEC'S STATUS FIELD WAS SET BY colophon-apply, the first time in the estate that 0083's job has done what it was built for: the implementing commit carried Updates: /wikis/specs/<slug> implemented in the BARE form, which defaults to the commit's own repository. 0086's flip hours earlier used cicd/wikis/... and 404'd, because the project before /wikis/ is taken verbatim and the wrong form came from colophon's own unit tests, where no forge is involved. AND THE JOB REPORTED THE SUCCESS AS A FAILURE - "gitlab's wiki API takes no commit message: the operation succeeded, but some of what was asked for could not be applied", headed failed and counted toward needed attention. Added to colophon#28, which already argues a FAILED instruction is invisible because the job exits zero; the report is now wrong in both directions, so the only reliable check is to open the page, which is what was done before claiming it worked. specs/home, the mirror and this log were still moved by hand: apply sets one frontmatter field and nothing else, so the three-index rule is unchanged. | go-test |
| 2026-09-15 | 0088-a-tool-a-component-runs-is-part-of-what-it-ships | approved | Implements cicd#48. Four component inputs name a version the component EXECUTES rather than one describing itself: retire_version, lockfile_lint_version and cyclonedx_version on svelte-security (all npx --yes <tool>@…) and version on go-singleuse (go install …/controls/lint@…). Each default is baked into a released component, so it is part of what that release DOES, and Renovate typed all four chore(deps): which colophon weighs as no release. It fails two ways from one cause: the bump NEVER SHIPS until something unrelated cuts a release, and when one does the bump SHIPS INVISIBLY with no changelog entry. Both happened in one week - @cyclonedx/cyclonedx-npm went 5.0.0 to 6.0.1 in cdedc31 and shipped in v0.48.0, a MAJOR carrying a reworked npm-detection path and advisory GHSA-q69g-4hcv-6jg4, and v0.48.0's notes do not mention it; retire went to 5.7.0 in e6a6d29 and shipped in v0.47.0 unmentioned, and only shipped because two fix(images): commits cut that release. D1 TWO rules scoped by COMPONENT, not one. Adding the four to the existing fix(images): rule was the obvious move and is rejected - retire is an npm package and fix(images): Update dependency retire would be false in the one artefact this spec exists to make truthful; a generic fix(tools): was also rejected because a Conventional Commit's scope is what tells a reader WHICH PART changed and tools asks them to already know what retire is. So fix(svelte-security): for the three npx tools and fix(go-singleuse): for the lint module. D2 fix not feat: a tool moving changes what a component does without changing what it offers, and patch is the right weight. D3 EXACT dep names, not a glob - …/go/controls/** reads naturally and is wrong, because Renovate also extracts gitlab.com/phpboyscout/go/controls here from the go-singleuse fixture's go.mod, confirmed present in a real run, and a glob would sweep in a different dependency; the four names were taken from a run rather than read off the templates. TESTING: the resolved prefix CANNOT be simulated, written down because it costs a run to discover - under --platform=local Renovate computes no commit message in either dry-run mode, --dry-run=lookup stopping before branching and --dry-run=full behaving identically because the local platform creates no branches, so commitMessagePrefix appears only in the echoed config. The evidence is production history instead and is stronger for it: cdedc31, e6a6d29 and a803c58 prove the current typing, every fix(images): Update dependency phpboyscout/images/… commit proves the mechanism, and what neither proves is the SPELLING, which is what D3 addresses. OQ1 ghcr.io/apricote/releaser-pleaser is still in the fix(images): rule with nothing in the estate running it; OQ2 phpboyscout/cicd's self-include correctly stays chore(deps): because bumping the version cicd pins of itself changes nothing a consumer receives. | renovate presets, svelte-security, go-singleuse |
| 2026-09-16 | 0089-every-docs-site-carries-a-map-for-machine-readers | implemented | The blog's discoverability analysis of 2026-09-15 found the 66 zensical microsites are 71% of the estate's search impressions, that verified AI crawlers (GPTBot, ClaudeBot, OAI-SearchBot) and assistant fetches (ChatGPT-User, Claude-User) already read them, and that NONE serves /llms.txt while the blog does. An assistant fetches one page at a time and has no way to learn what the other sixty are. D1 the map is written by zensical-pages after the build, for every consumer on the next docs rebuild, rather than per repo: 66 MRs to add a file that every site derives identically is the wrong shape, and a per-repo file drifts from the nav the moment a page moves. D2 derived from zensical.toml and the docs tree, never authored - sections follow the site's own nav when it has one and the Diátaxis directory order otherwise, lines carry front-matter title and description, _-prefixed paths are skipped so _template/ scaffolding never appears - so the map cannot disagree with the site. D3 embedded in the template as a heredoc with the canonical copy under scripts/zensical-llms-txt/, exactly as release-stamp ships its engine (0073), because a component cannot fetch files from its own repo at run time and a tool baked into docs-tools would need an image release for a 260-line script; the self-test byte-diffs the two. D4 site_url is required and the build fails without it, since relative links in a file whose whole purpose is to be fetched out of context are worse than no file. D5 the Optional block links the repository's releases, the estate changelog filtered to the project and the estate map at phpboyscout.uk/llms.txt, which are the three facts an assistant is most often asked for and the join that makes 66 sites read as one author. IMPLEMENTED 2026-09-16 (!345), all five decisions as written; colophon-apply flipped the wiki status from the commit's Updates: trailer and reported it as failed (colophon#28 again), and the approved/implemented dates, this row and the mirror were moved by hand. | zensical-pages |
| 2026-09-16 | 0090-exclude-docs-is-honoured-by-the-component | implemented | The discoverability analysis found /_template/ indexed by Google on the aws-bootstrap and aws-security-baseline sites. All three terraform modules carrying docs/_template.md, and go-tool-base with docs/reference/migration/_template.md, declare exclude_docs = "_template.md" and the file's own header says that excludes it. IT DOES NOT: Zensical accepts the MkDocs key and builds the file regardless, reproduced with 0.0.57 locally and 0.0.62 (latest, 2026-09-16) in a clean venv, and neither wheel's Python mentions the string. D1 the component honours the key, removing matches from the docs tree BEFORE the build so sitemap and search index stay consistent; fixing four repos was rejected because they already declare the right thing the documented way and the fix should hold for every site using the key. D2 the useful subset of gitignore semantics: bare name at any depth, slash anchors to the docs root, trailing slash for a directory, no negation; a pattern matching nothing is logged and ignored because a typo in an exclusion must not stop a docs site publishing. D3 fifteen inline lines, not a third embedded engine, since there is nothing to unit-test. D4 reported upstream via cicd#52, ready-for-human, per the estate rule on third-party trackers. APPROVED 2026-09-16; IMPLEMENTED in the same merge (!347), the in-repo indexes carried in the implementing MR this time rather than a trailing docs MR, and the wiki flipped by colophon-apply from the commit's Updates: trailer. | zensical-pages |
| 2026-09-16 | 0091-every-docs-page-names-its-site-software-and-author | implemented | Recommendation 6 of the blog's discoverability analysis: the microsites emit no structured data, so an answer engine deciding who wrote them and whether 66 hosts are one estate has only the text. Zensical's main.html is a bare extends, base.html exposes an empty extrahead block, and there is no partial hook in the head, so a main.html is the only way in and only a site's custom_dir supplies one (two sites have one, 64 do not). D1 the component writes the override into the site's custom_dir in the job workspace: the declared one when the site has it (its own main.html, if any, wins), otherwise .zensical-overrides with the key inserted into zensical.toml by a one-line textual edit that is re-parsed with tomllib before it is trusted. REPLACING THE PACKAGE'S OWN main.html WAS THE FIRST DESIGN AND FAILED ON THE FIRST PIPELINE: /opt/zensical is root-owned and the job runs as ci, Permission denied; recorded so nobody tries it again. Baking into docs-tools was rejected because every change to the graph would be an image release (0089 D3). D2 the graph: WebSite, SoftwareSourceCode and one Person on every page, TechArticle on every page but the home, every field derived from zensical.toml or the page's front matter, the home detected by canonical URL because page.is_homepage was tested and is false there. D3 programmingLanguage from [project.extra] language or a .go./.rust./.iac. host. D4 the author is the first authors: entry's name before <, else site_author. D5 sameAs is the site's social links minus the site, the repository (replaced by its GitLab namespace) and chat invites, which lands on the same seven profiles the blog's Person carries: one entity across 67 hosts. D6 the Person is emitted once per page and referenced by @id. OQ1 the bare-host flagships should declare a language; OQ2 the blog About page's BlogPosting should be a ProfilePage; OQ3 Zensical's default edit_uri points every edit link at edit/master/ on an estate whose branch is main. APPROVED 2026-09-16; IMPLEMENTED in the same merge, the in-repo rows carried in the implementing MR. | zensical-pages |
| 2026-09-16 | 0092-the-component-fills-in-edit-uri-and-language | implemented | 0091's OQ1 and OQ3. Thirteen sites on bare hosts emit no programmingLanguage unless each declares [project.extra] language, a fact go.mod or Cargo.toml already states; and Zensical's default edit_uri is edit/master/<docs>/ on an estate whose branch is main, so every edit link on 65 sites 301s to a branch that does not exist (go/repo alone sets it by hand). D1 edit_uri = -/edit/$CI_DEFAULT_BRANCH/<docs path from repo root>/ when unset and repo_url is on gitlab.com, the path computed from the working directory so monorepo sites link correctly; GitHub repos keep Zensical's default because that IS GitHub's form. D2 language from the manifest, working directory then repo root: go.mod Go, Cargo.toml Rust, .tf HCL, pyproject.toml Python; no manifest, no key, and 0091's host rule still applies. D3 each value a one-line insert under its header or a new table appended after its own sub-tables (TOML permits it), re-parsed with tomllib before replacing the file, skipped with a log line if it would not parse, because a docs build must not fail over a nicety. Same shape as 0090's exclude_docs edit. APPROVED 2026-09-16; IMPLEMENTED in the same merge, rows carried in the implementing MR. | zensical-pages |
| 2026-09-16 | 0093-tofu-components-run-a-stack-with-no-aws-identity | draft | role_arn on tofu-plan, tofu-apply and tofu-stop becomes optional (default empty), and a deploy-generate catalog target may carry role_arn: null. When no role is named the job mints no OIDC token and exports no AWS_* variable; the stack's providers authenticate from their own CI variables. Chosen over a second cloud-agnostic component (two ways to configure one thing) and over a generic credentials block (designing for clouds nobody has). Raised by infra's Cloudflare stack (infra spec 0019), answers cicd issue 26. | tofu-plan, tofu-apply, tofu-stop, tofu-deploy-generate |
| 2026-09-17 | 0094-spike-retag-instead-of-rebuild | implemented | Whether a release should RETAG the sha- image the main pipeline already built and scanned, rather than rebuilding from the same commit. The correctness half of the original case did not survive measurement and is withdrawn: with the Buildah switch and per-image scrubs of spec 0098, a rebuild of the same commit demonstrably produces the SAME BYTES, proven on every push by each repo's reproducible job. What stands is time - six to eleven minutes per release, and more now that Buildah is 14% to 67% slower than kaniko on every image, against a fleet with two concurrent slots. Reproducibility also makes a retag auditable rather than trusted: rebuild the tagged commit and compare the digest. APPROVED 2026-09-17 once the open questions were answered: trivy scans a REGISTRY REFERENCE fine (verified against ci-base v0.2.1), so the scan jobs stop needing image.tar and the dependency the spike called the real design question dissolves. D1 the two paths are two CHILD PIPELINES named retag and rebuild, not a branch inside one job - the path taken becomes a pipeline object visible in the graph and the API rather than a log line nobody reads, and the fallback has never fired (0 of 22 semver tags lacked a sha- tag) which is exactly why it would go unnoticed when it does. D2 the release GATE is colophon's release MR, not the tag pipeline: it builds, scans and runs reproducible against the EXACT commit that gets tagged, because under fast-forward the release MR head becomes main's head (ci-base !34 at 0d6df19d, v0.2.1 tagged at 0d6df19d). A hand-cut tag has no release MR and so no gate - that is the emergency path, and it is precisely what the rebuild child covers. Measured saving: 43.8 min of redundant work per release round across the eight repos, down to ~40s of crane tag, and about 22 min of wall-clock contention on a two-slot fleet. IMPLEMENTED on release-tools 2026-09-17 - the first repo in the estate where a release is a registry operation, verified in the live registry: v0.1.13, v0.1, latest and sha-d41a5a81 all resolve to sha256:db1d8ee1, and the release path is ~15s against ~126s of build, scan and push. D1 IS SUPERSEDED: two static children selected by rules: CANNOT work, because rules are evaluated at pipeline creation and a dotenv value written by a job in that same pipeline is empty - both triggers were dropped and release-tools v0.1.13 was tagged in git with NO IMAGE behind it. The child config is now GENERATED and included as an artifact, which is the dynamic form originally proposed. FOUR DEFECTS SHIPPED AND EVERY ONE FAILED GREEN, which is a property of the design rather than bad luck: a release path that publishes nothing looks exactly like one that had nothing to do. They were the omitted --format docker (buildah discarded SHELL, so piped RUNs stopped failing), the stage-order race (colophon tags from release, repos published sha- from publish, so the tag always won and the probe would have rebuilt every time), wait_seconds typed integer which GitLab does not accept (pipeline creation failed with zero jobs and no yaml_errors, breaking a consumer's default branch), and the rules/dotenv defect above. The third is why the authoring convention exists: the self-tests never INCLUDED the component, so nothing parsed its spec header - they do now, from $CI_COMMIT_SHA. What no test can reach: release:resolve, the trigger and the child exist ONLY on a tag pipeline, so the component is only ever exercised by cutting a release; three of the four defects were found that way. hadolint also dropped from tag pipelines (19 runs in two days, two of them on tags telling nobody anything). Rollout is ONE repo, not eight - the other seven wait until release-tools completes a release unattended, and ci-base goes LAST because every other image derives from it. | image repos (phpboyscout/images/*) |
| 2026-09-16 | 0095-verify-release-binaries | in progress | A scheduled, credential-free verifier for release binaries on pkg.phpboyscout.uk. Three tiers — reachability everywhere, integrity where a checksums.txt exists, authenticity where a signature does — each counted rather than failed where a project's publishing shape gives it nothing to run against. Reads the public URL and holds no bucket credential, so it verifies what users receive and can never cause the damage it reports. Integrity rotates on sha256(project@tag) so a 23 GiB estate is covered weekly without a nightly full download. allow_failure defaults false, against the house default for scheduled chores: a missing binary is a broken public guarantee. One engine copy in scripts/, no template embed, so there is nothing to drift. Replaces the R2 bucket lock dropped in colophon spec 0025 D13. | verify-release-binaries |
| 2026-09-16 | 0096-goreleaser-refuses-to-republish | in progress | The goreleaser job refuses to run for a tag already published to the release binary store. Needed because the store's bucket lock was dropped (colophon 0025 D13) AND goreleaser's blobs pipe cannot write conditionally — its eighteen config keys include no If-None-Match, skip-existing or overwrite refusal — so a re-run of an old tag would silently replace published bytes under a recorded checksum. Guards at tag granularity, sentinel checksums.txt because goreleaser writes it last, so a completed publish is protected while a half-uploaded tag can still be finished. Unauthenticated (reads the public URL), and an unreachable store warns rather than blocking, so a CDN blip cannot stop every release in the estate. | goreleaser |
| 2026-09-16 | 0097-one-signing-binary-in-the-base-image | implemented | Two projects compiled the WHOLE go-tool-base framework from source in CI (go build ... cmd/gtb, on every release) to reach one subcommand, gtb sign; sigillum is that subcommand as a standalone binary. D1 the swap cannot change the command surface and that is VERIFIABLE: gtb sign and sigillum sign both take their command definitions from go/signing-cli and at the versions in play it is the same version of the same module (v0.6.1), so cmd: gtb -> cmd: sigillum is the whole change to the invocation. D2 ci-base carries it, by the maintainer's call; the first draft's reach argument (that ci-base would reach the non-Go signers) was measured and found WRONG, since rust-tool-base signs in dev-tools which derives from nothing of ours, and a full census found TEN projects sign, not four, with the estate already half on sigillum - so this finishes a migration already half done. AS BUILT 2026-09-17, each step verified in the published artefact: sigillum v0.4.2 (first release able to sign at all), ci-base v0.2.1 (sigillum installed, plus the Buildah --format docker fix so SHELL and pipefail are honoured), go-tools v0.4.0 rebuilt on it, cicd !361 moved ONLY the goreleaser component to that image, colophon !122 deleted the build step. go-lint/go-test/go-security/go-singleuse STAY on go-tools v0.2.0 because the same jump carries govulncheck v1.7.0 -> v1.8.0 and a newer analyser finds more - moving the gates could turn MRs red across the estate for findings that were always there; that pilot is separate work. The OIDC token-file write stays (measured: the AWS SDK wants a file, GitLab hands a variable). STILL UNPROVEN: the signature itself - KMS admits only tag pipelines on the owning project, so the test that matters is colophon's next release producing a checksums.txt.sig that verifies; implemented and unverified is not the same as working. | goreleaser; ci-base, go-tools images; colophon |
| 2026-09-17 | 0098-spike-replace-kaniko | implemented | kaniko was archived 2025-06-03 and all eight image repos pinned a 2024-07-10 build, so nothing would ever fix a CVE in it. Buildah replaces it because it keeps the property kaniko was chosen for - daemonless, no docker:dind, no mounted socket - and reads the existing Dockerfiles unchanged. BuildKit's ordinary buildx form needs a daemon; apko cannot express RUN and so cannot build the tool images; ko is for Go applications, not toolchains. FOUR FINDINGS THE SPIKE GOT WRONG, each of which cost a rebuild cycle. (1) Buildah is SLOWER than kaniko on all eight images, +14% to +67% measured on runner1, where the spike's ~100s figure was a workstation number; the switch is justified on maintenance, not speed, and reproducible makes an MR cost about three builds where it cost one, which is why it is excluded from the nightly schedule. (2) --format docker is MANDATORY and omitting it is silent: the OCI image format has no SHELL instruction, so buildah DISCARDS SHELL [... -o pipefail ...] with only a warning and a piped RUN stops failing - a piped RUN whose FIRST command fails exits 0 under oci and 1 under docker. It shipped to ci-base and go-tools before being caught, making ci-base v0.2.0 the one release built without it. (3) SEVEN causes of non-determinism, not the three the spike named from go-tools alone, and every one is a tool caching its own state BESIDE the content rather than the content itself - ldconfig's aux-cache, Go telemetry and the sumdb tree head, pip's mtime-keyed .pyc headers, Node's V8 compile cache, tflint and sigstore metadata, rustup's parallel component ordering, a stray mktemp directory. release-tools needed NOTHING, which is the counter-example that makes copying one image's scrub to another a mistake. (4) Two failure modes a digest comparison cannot see: rm -rf /tmp/* does not match dotfiles (Chromium's /tmp dotfile survived it), and a random FILENAME changes tar entry order - rust-tools had a 634 MB layer differ with zero differing file contents across 9,023 entries. Each repo now carries a reproducible job that builds twice and compares, proven able to fail. It cannot catch date-keyed state (both builds run minutes apart) and reproducibility is run-to-run, not across time, since apk resolves against a rolling Wolfi repository. | image repos (phpboyscout/images/*) |
| 2026-09-17 | 0099-the-nightly-scan-tells-someone-when-its-answer-changes | implemented | Six of the eight image repos rebuild and rescan nightly against a fresh vulnerability database and NOTHING reports the result: go-tools' scan:tools has failed every night for weeks on an upstream-baked grpc CVE and nobody was told. D1 NOT a chat channel and NOT the image repo - a vulnerability finding is not an announcement, and the eight image repos are PUBLIC, so filing there would publish an unpatched advisory against what the estate ships. The destination is the PRIVATE phpboyscout/org. A Discord webhook was proposed first and is recorded as rejected so it is not proposed again. D2 one long-lived issue per image, edited in place like Renovate's dependency dashboard: a per-run notification is noise, a per-finding issue is an ungroomed backlog, and both get muted. D3 the issue body IS the store, so there is nothing else to persist or expire. D4 allow_failure, and no scan threshold or exit code changes - it makes an existing answer visible, it does not gate. D5 playwright-tools and release-tools get the nightly schedule they lack. D6 this is SURVEILLANCE, not a release gate - the release MR already gates properly (0094 D2); what this covers is the weeks-wide gap BETWEEN releases, measured at 17 to 24 days for five of the eight repos, during which a newly-disclosed CVE against an already-shipped image is invisible. AMENDED 2026-09-18 - THE SCAN ITSELF BECOMES A COMPONENT: surveying the eight repos found seven carrying the SAME hand-written .scan-base -> scan:os (hard) / scan:tools (advisory) pair, so the copy goes and image-scan replaces it with the reporting as an INPUT (D7); two deviations are decisions, not drift - docs-tools keeps scan:tools as a hard gate via tools_allow_failure: false, and ci-base takes the same three jobs rather than a single-job variant. Job names are preserved so the consumers' publish -> needs: [scan:os] gate is untouched. D8 scan:report exists only when report_project is set (surfaced via a job variable so rules: can see it) and runs only on DEFAULT-BRANCH pipelines: an MR pipeline scans a branch build and must not write a branch's findings into the image's dashboard. D9 target phpboyscout/org, title Image scan: <image>, GROUP label security-advisory; finding identity is (advisory, package, installed version) per section; the body is ALWAYS rewritten with a last scanned stamp - on a CLOSED issue too, because a clean image and a dead token otherwise look identical - a comment carries new/cleared only on change, and the issue CLOSES when both sections are empty and REOPENS on a finding (reversing the first draft's lean). D10 the self-test proves scanning (crane-pulled public image) and the report in dry-run; the LIVE WRITE is not self-testable (protected GITLAB_TOKEN invisible to MR pipelines, job token cannot create issues) and is verified at the consumer, go-tools, whose first default-branch pipeline after adoption must produce the issue with the known finding set. OQ1 RESOLVED: the group GITLAB_TOKEN - GitLab Free has no project or group access tokens, and a job token cannot write another project's issues, so the token input defaults to $GITLAB_TOKEN, a stated departure from authoring rule 6 with renovate-merge and colophon as precedent. OQ2 yes (diff comment on change). OQ3 two sections, one issue. OQ4 closes when clean. IMPLEMENTED 2026-09-18: cicd !370 + !371 (v0.50.0); adopted the same day by seven of eight image repos (docs-tools !30 blocked by its own new reproducible gate, docs-tools#2). OUTCOME: the pilot found the case D9 did not name - a SKIPPED scan (build lost a GitHub 504) read as clean and created org#13 CLOSED, because when: always was written for a FAILED scan; !371 makes a missing scan file exit 2 and touch nothing. A pipeline retry does NOT re-run a job that succeeded wrongly; scan:report had to be retried alone. Every hand-written scan:os re-exported image.tar and publish took it from there, so each swap also pointed publish at build. The dashboard after adoption: go-tools 14, node-tools 4, playwright-tools 4, tofu-tools 81 tool-binary findings (open); rust-tools, ci-base, release-tools clean (closed); OS packages zero everywhere. | new image-scan; image repos (phpboyscout/images/*) |
| 2026-09-19 | colophon 0025-release-binaries-in-object-storage-on-a-vanity-URL | implemented (D12, D15, D16 here) | A colophon spec, implemented on the cicd side by the goreleaser component: the release binary store's keys reach goreleaser's blobs pipe as AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY through two inputs whose defaults name the R2_RELEASE_* group variables, so a consumer carries no credential lines. The credential seam was MEASURED, not argued: static env keys outrank the OIDC web-identity token file in the AWS SDK's chain, a named AWS_PROFILE does NOT shield them (web identity still wins over a profile's static keys), and blanking the two keys restores web identity — so each signing project blanks them on its sign command with signs.env, which goreleaser overlays onto a copy of the env. Six projects proved the shape on real tags the same day (colophon v0.12.0, sigillum v0.5.0, go-tool-base v0.44.0, keryx v0.20.0, skillup v0.3.0, phpbotscout v0.5.0) carrying the two lines per project before they moved here. The two hand-rolled uploaders (rust-tool-base, ffmpeg-wasi) take the curl route (0025 D19): curl --aws-sigv4 with If-None-Match: *, which gets the per-object 412 goreleaser's blobs cannot express. | goreleaser |
| 2026-09-19 | 0100-release-announcements-move-to-colophon-announce | implemented | colophon v0.11.0 ships announce (discord + file adapters, configured in .colophon.yaml, reads the release object, exits 0 whatever happens) and nothing in cicd runs it; discord-release (0064) re-implements the Discord half in shell on six consumers and cannot feed the blog's seeded release feed. D1 announcements are colophon's; discord-release goes and 0064 becomes SUPERSEDED. D2 announce runs WHERE THE RELEASE OBJECT IS CREATED, which is two places: after colophon-release on the tag pipeline for assets: true projects, and after step 5 of colophon-publish on the DEFAULT BRANCH for everyone else - the brief's single tag-pipeline job with needs: would race the object on ~100 projects because publish pushes the tag (starting that pipeline) BEFORE it creates the release (publish.go steps 4-5); no separate announce job, because the verb exits 0 so the graph would be green on a failed announcement anyway. D3 adapter config is the manifest's (0077 D2), Discord per headline project. D4 the blog feed is ESTATE POLICY identical across ~106 consumers, so 106 manifest blocks is the wrong shape - colophon to accept it once (OQ2), sweep as fallback riding the R2 rollout. D5 BLOG_TOKEN cannot be a project access token (GitLab Free has none) - a PAT-backed group variable, masked+hidden+protected. D6 the decommission list. D7 0064's notes-retry is removed by placement, never re-added. Written against colophon spec 0026 (DRAFT, same day), which answers OQ1 (publish/release announce at the end, --no-announce), OQ2 (the feed block is OPERATOR config the component ships, the manifest merges over it by adapter name) and OQ5 (a local-filesystem file target for tests). OQ3 DECIDED: TEN headline projects (afmpeg, colophon, ffmpeg-wasi, go-tool-base, keryx, krites, phpbotscout, rust-tool-base, sigillum, skillup) to Discord AND the feed, when: any; every other PUBLIC project to the feed only, gated on $CI_PROJECT_VISIBILITY so a private release never reaches the public blog. OQ4 DECIDED: the feed writes with the group GITLAB_TOKEN, no BLOG_TOKEN minted, until project access tokens exist. AMENDED on approval: announcing is OPT-IN per project, always (Matt on colophon 0026 OQ2) - the component ships the feed SETTINGS, a project enables the feed with announce: { file: {} }, so the consumer pass is every public consumer (~100 one-line MRs, batched off-hours) plus the discord block on the ten. Its own pass, not the R2 rollout. APPROVED 2026-09-19. IMPLEMENTED 2026-09-20: component in cicd !378 (colophon 0.13.0 / release-tools v0.1.14), canary colophon v0.14.0, then the estate pass - 111 MRs, 108 merged in 25 MINUTES via the escape-hatch fast path (a one-line manifest change gains nothing from five scanners; 103 pipelines would have been 12 hours on two slots), 3 on pipelines overnight, 0 flagged - and the decommission in !383. OUTCOME: every COLOPHON_* env var is configuration (an empty COLOPHON_ANNOUNCE_FLAGS became announce.flags; colophon#32, job variables are CICD_*); the feed token is aliased $CICD_FEED_TOKEN because the job exports GITLAB_TOKEN as the colophon token; the config file needs a .yaml extension; opt-in made the pass 111 MRs not ten; krites kept the include (parked, krites#12); a gate-restoring sweep MUST be scoped to touched repos. | colophon; discord-release (removed); six consumers + the blog feed |
| 2026-09-22 | colophon 0005-release-from-a-tag | implemented (D7 here) | A colophon spec, implemented on the cicd side by switching colophon-release from publish --assets to release --tag $CI_COMMIT_TAG. publish starts from the last merged release MR, so on a tag a maintainer pushed by hand it found the previous release and created nothing; release starts from the tag, so the hand-cut recovery tag gets its release object, assets and announcement from the tag pipeline like any other. No new image floor: the job already needed colophon 0.13.0 for the announce flags, and release has existed since 0.3.0. Closes cicd#41, whose runbook is colophon's *Recover with a hand-cut tag, now linked from the publish-a-release how-to. | colophon |
| 2026-09-23 | 0101-go-test-says-how-many-tests-it-skipped | implemented | go test without -v prints the same ok line whether a package ran its tests or skipped them, so a gated suite that runs nothing is green. The test jobs run through gotestsum: the unit job keeps its one-line-per-package log and closes with the skipped tests and their reasons, the failures, and a count; all three jobs publish a JUnit report GitLab renders as a Tests tab and an MR widget; a stock Go image without gotestsum falls back to plain go test. Nothing new gates. gotestsum is built from phpboyscout's fork in go-tools, because upstream's release carries 49 advisories by trivy and the fork scans clean. The other half of cicd#21, a helper that fails when its gate is on and its dependency is missing, is guidance (Explanation: a skip is not a pass). | go-test |
| 2026-10-04 | 0102-family-core-currency-guard | implemented | A family member can cut a release still pinning the previous release of its core: 0067 D9 measured 23 of 23 config adapters tagging against a stale config. A standalone Go-track component (amended the day it was approved, out of colophon, because the rule is about Go dependencies and the default would have imposed this estate's families on every colophon user) runs on the release MR, reads the family cores from a feed given as a required feed input (a URL or a repository path, in a published schema-v1 format; the estate uses the blog's projects feed, blog spec 0003), and compares every core the project's go.mod files require against the Go proxy's @latest. mode: warn (default) reports a stale core as an allowed failure, enforce fails it; an unreadable source is always a warning. Membership comes from go.mod, never from a list. | go-core-currency |
| 2026-10-05 | 0103-go-core-currency-estate-modules | approved | 0102 checked only family cores and deferred the rest (its D7). A first_party prefix input adds an estate tier: any direct requirement under it that is not a core is checked too, and stale, deprecated or moved is a warning in every mode; only a stale core can block, under mode: enforce. A "latest" is trusted only if its own go.mod declares the module, because a renamed repository let the proxy answer go/chat-platform's @latest with a go/comms tag. Deprecation is read from that go.mod, as the go command does, not from pkg.go.dev's flag, which missed it. | go-core-currency |
| 2026-10-05 | 0105-security-components-stop-colliding-on-shared-job-names | implemented | go-security and svelte-security both defined osv-scanner and gitleaks; GitLab merges same-named jobs across includes and the later wins, so keryx and krites never osv-scanned go.mod (cicd#57). svelte-security's two jobs become svelte-osv-scanner and svelte-gitleaks, go-security is untouched (97 consumers, and go-tool-base overrides its osv-scanner), and a gitleaks boolean input (default on) lets a project that already runs gitleaks switch svelte's off. A self-test:job-names lint fails cicd when two templates define the same top-level key, hidden keys included, with dormant collisions allowlisted by date and reason; a combined self-test needs: both osv jobs so a missing one stops the pipeline being created. | svelte-security, go-security |
How to read this¶
- Status mirrors each spec's frontmatter:
draft(under discussion, not yet implementable) →approved(safe to implement) →rejected(considered and declined — kept, not deleted) →implemented(shipped). Most specs here go straight fromapprovedto shipped without an explicitimplementedflip; treatapprovedentries whose described behaviour matches the currenttemplates/as shipped. - This table is generated by hand and kept current as part of the contributing workflow — every new spec gets a row here when it lands.