Two sessions resolved the same all-checks-passed conflict differently. The
remote's version is the one kept: the selector with '|| ubuntu-latest'
covers a standard-hosted outage too, which a bare ubuntu-latest pin does
not, and its runbook wording is internally consistent at four failover
jobs. Reverted this side's three-job doc downgrade.
One semantic conflict on the `all-checks-passed` verdict job. Master moved
it to `ubuntu-latest`; this branch had routed it through the failover
selector so a hosted-pool outage could not leave the required verdict
queued. Master's resolution satisfies that requirement more directly:
standard hosted capacity is independent of both custom pools, so the
verdict is reachable whichever pool is degraded, and the selector is no
longer needed. Kept `ubuntu-latest` and folded the failover reasoning
into its comment.
The failover runbook accordingly documents three failover jobs (the three
required Linux workers), not four, and states why the verdict stays on
standard hosted capacity in both states.
config.sh only registers; the runner stays offline until svc.sh
install/start. Both language sides updated so emergency capacity
actually comes online.
- serial-linux-selfhosted checks out fetch-depth 0: depth 2 misses
github.event.before on multi-commit or force pushes, failing the
archive verifier on a valid tree. Full fetch is cheap against the
VM's local mirror.
- Runbook (both languages): every remaining admin phrasing (problem
statement, switch heading, alternatives, consequences) now says
writer; and the 'composes with this mechanism' claim about a
master-ref-pinned runner group is replaced with the truth observed
live on 2026-07-27 — master-ref pinning blocks PR failover, and the
shipped posture is repository-scoped all-workflow group access.
Static gate green locally: 32 passed, 0 failed.
The standby lane itself is push-only, but under failover pull_request
jobs do reach these runners with the PR merge ref's workflow. The
workflow comment and the larger-runner note (both languages) now state
that plainly and name the actual boundary — repository membership
(private, forking disabled, Dependabot excluded) — matching the
runbook. Static gate green locally: 32 passed, 0 failed.
- Sweep every remaining 'admin-only' claim (workflow comments, runbook
lines 13/40, topology note, all zh pairs): the variable is
writer-manageable, and the boundary against untrusted code is
repository membership (private, forking disabled, Dependabot
excluded) — stated identically at every site instead of only in the
'who can flip' paragraph.
- Serial cross-platform reference note (both languages): master now
runs four references — the three hosted OS legs plus the self-hosted
standby drill, linked to the failover runbook.
Static gate green locally: 32 passed, 0 failed.
- serial-linux-selfhosted now fetches depth 2 and passes
DSH_ARCHIVE_BASE_REF=github.event.before, running the same
frozen-archive comparison as serial-linux instead of diffing the
new manifest against itself.
- Runbook (both languages): documents the deliberate dependabot
exception (queued-on-hosted during failover is expected, not a
failed switch); corrects the emergency-capacity bootstrap to
exclude .runner/.credentials when cloning a runner directory; and
replaces the 'admin-only' variable claim with the accurate
trust-model statement — repository variables are writer-manageable,
which in this private fork-disabled repo with an all-workflows
runner group is routing among members, not an escalation.
Static gate green locally: 32 passed, 0 failed.
- All four failover selectors (three workers + the verdict job) and the
paired env/cache expressions now exclude dependabot[bot]: under
failover, dependency-supplied code keeps queueing for the hosted pool
instead of executing on the persistent VM. A delayed Dependabot PR
during an outage is an acceptable cost; dependency code on the
privileged host is not.
- Runbook (both languages): records the shipped failover bounds
(coverage 8, snapshots 12, sized for six instances) and documents
that the verdict job follows the selector too — operators previously
had no explanation for a verdict queued after all workers passed.
- Local static gate green: 32 passed, 0 failed (translation pairing
519 pairs consistent).
The pairing gate requires link target #9 to be byte-identical between
the language sides; my earlier 'fix' pointed the zh side at the zh
runbook and broke the contract. Reverted to the shared target and
re-recorded the pairing hash.
- all-checks-passed now resolves its pool through the same
DSH_CI_FAILOVER expression as the worker jobs it aggregates.
Pinned to the hosted pool it would leave the branch-protection
verdict queued on the failed pool after every failover job passed —
observed live during the 2026-07-27 outage as a required check
looping against dead capacity.
- Coverage worker bound under failover drops 12 → 8 and snapshot
concurrency 16 → 12: the pool now runs six always-on instances (the
spare tier was retired), so worst case is 6 × 8 = 48 coverage
workers on the shared 64-core VM.