Platform operations · DC Bar
Six documents covering how this platform was built, how the two sites relate, and what to do when one of them is gone. Written while doing the work, so the traps are recorded where they were found.
What to do when HQ is gone and AWS has to take over — and which parts of that are proven rather than assumed.
Read →How the HQ and AWS sites fit together: what runs where, what replicates, and which single points of failure remain.
Read →How the HQ cluster was built, from bare Ubuntu to a working three-master control plane with Longhorn and MetalLB.
Read →The full path from a repository to a working production URL — ten stages, and which check actually proves each one.
Read →Making Argo CD respond to pushes in seconds instead of waiting out its three-minute poll.
Read →What went wrong bringing four applications into production, and why every one of them reported success somewhere.
Read →Each document distinguishes what has been verified from what is merely intended, because that distinction is the whole value of a runbook. Where something has been tested, it says so and gives the command that tested it. Where it has not, it says that too — a step marked as unproven is a step that will surprise you at three in the morning.
Every one of them also carries a “traps” section. Those entries are not general advice; each one cost real time to find in this specific environment, and none of them produced an obvious error message. They are the parts most worth reading before you need them.
These pages are generated from source in this repository and served from the cluster, so a correction is a commit rather than an edit somewhere nobody else can see. If you fix something in the platform, fix the document in the same change — a runbook that has drifted is worse than no runbook, because it is trusted.
Anything with a date attached is a claim about a moment. Revised in a document’s masthead is the date someone last checked it against reality, not the date it was last touched.