DC Bar · 10.20.100.0/24 · Ubuntu 26.04 LTS

prod-k8s Build Runbook

One script, four roles, six machines. This is the address plan and run order for the HA cluster — three control-plane nodes behind a floating API VIP, replicated storage on the workers themselves, and Argo CD, Portainer, and Traefik each on their own LoadBalancer address. No external storage host.

6 cluster nodes Kubernetes 1.36 Cilium 1.20.1 Longhorn 1.12.1 Argo CD 3.5.1 HA Traefik v3 etcd quorum 2 of 3 volume replicas 3

Before you start

Five things will stop the build cold if they are wrong. Two of them — the VRRP password and the replica disks — fail quietly rather than loudly, so check them first.

  • Two new VMs at .228 and .229 Ubuntu 26.04 LTS, 2 vCPU / 4 GB minimum for a control-plane node. Hostnames must be unique across the cluster — Longhorn keys NFS lock recovery on them.
  • .227 is genuinely unused Nothing may answer ARP for it. Check with arping -c3 10.20.100.227 from another host on the segment.
  • The same VRRP_PASS on all three control-plane nodes Mismatched passwords produce two nodes that both believe they hold the VIP. Also confirm VRRP router ID 51 is unused on this L2 segment.
  • A spare disk on each of .231, .232, .233 One unpartitioned device per worker, same size on all three, sized for 3× your total volume needs. Confirm the device name with lsblk — the script formats what you name and refuses if it already holds a filesystem.
  • Node-to-node traffic is open TCP 6443 and 8443 (API + haproxy), 2379–2380 (etcd), 10250/10257/10259 (kubelet, controller, scheduler), 9500–9504 and 8500–8501 (Longhorn), 30000–32767 (NodePorts), VRRP protocol 112, and ARP for the MetalLB range.

Run order

The same file on every node, with a different ROLE. Phase 1 prints the two join commands; you paste the right one into phases 2 and 3. Each phase refuses to run on the wrong host, and rejects a join command meant for the other role.

on 10.20.100.230 · master-01

Initialise the control plane

sudo VRRP_PASS='your-shared-secret' \
  ROLE=master-init ./k8s-ha-stack.sh

Prepares the node, installs keepalived and haproxy, brings up the VIP, then runs kubeadm init against 10.20.100.227:8443.

thenTwo join commands print, and are saved to /root/join-*.txt. The cluster has one node, NotReady — expected until Cilium lands in phase 4.

on .228, then .229 · master-02, master-03

Join the other control-plane nodes

sudo VRRP_PASS='your-shared-secret' \
  JOIN_CMD='kubeadm join … --control-plane --certificate-key …' \
  ROLE=master-join ./k8s-ha-stack.sh

One at a time, not in parallel — each new etcd member must be admitted and synchronised before the next one asks to join.

thenThree etcd members, quorum of 2. The cluster now survives losing any single control-plane node.

on .231, .232, .233 · workers

Join the workers and claim their disks

sudo LONGHORN_DISK=/dev/sdb \
  JOIN_CMD='kubeadm join …' \
  ROLE=worker ./k8s-ha-stack.sh

These can run in parallel. No VRRP password — workers hold no VIP. The named disk is formatted ext4 and mounted at /var/lib/longhorn before the join.

thenkubectl get nodes lists all six. All still NotReady without a CNI.

back on 10.20.100.230

Install the platform

sudo ROLE=platform ./k8s-ha-stack.sh

Cilium, then MetalLB and the .235–.239 pool, then Longhorn, Traefik, Argo CD in HA mode, and Portainer. Roughly 15–25 minutes, mostly image pulls.

thenAll six nodes Ready. The three service IPs answer. The Argo CD password and the generated Longhorn UI password both print at the end — save the Longhorn one, it is stored nowhere else.

Verify

Do not treat the build as finished because the script exited zero. The one test that matters most is the failover — an HA cluster that has never failed over is only theoretically HA.

# every node Ready, three of them control-plane
kubectl get nodes -o wide

# all three etcd members healthy and one leader
kubectl -n kube-system exec -it etcd-prod-k8s-master-01 -- etcdctl \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key endpoint status --cluster -w table

# the service IPs were actually assigned from the pool
kubectl get svc -A --field-selector spec.type=LoadBalancer

# three storage nodes, each Ready and schedulable
kubectl -n longhorn-system get nodes.longhorn.io

# shared-volume round-trip: Bound within seconds, with a share-manager pod serving NFSv4
kubectl create -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: rwx-smoke }
spec:
  accessModes: [ReadWriteMany]
  storageClassName: longhorn-rwx
  resources: { requests: { storage: 1Gi } }
EOF
kubectl get pvc rwx-smoke
kubectl -n longhorn-system get pods -l longhorn.io/component=share-manager
kubectl delete pvc rwx-smoke

Failover test

Stop haproxy on whichever node holds the VIP and confirm it moves. Run the probe loop in the background so the result does not depend on your typing speed — expect roughly 5–9 seconds of failed probes, which is the interval 2 × fall 3 detection window plus one VRRP advertisement.

If the VIP does not move, check the keepalived weight. The health check needs weight -50: a positive weight only adds a bonus while healthy and cannot drop a node below its peers, so the node keeps the VIP with a dead haproxy behind it. The log line to look for is Changing effective priority from 150 to 100, followed by Entering BACKUP STATE.

# on .230 — find the holder, then drop it
ip -4 addr show | grep 10.20.100.227
sudo systemctl stop haproxy

# from your workstation — should keep answering throughout
while true; do kubectl get --raw='/readyz'; sleep 1; done

# put it back
sudo systemctl start haproxy

Access

Every UI terminates at Traefik on 10.20.100.237 using the *.dcbar.org wildcard, loaded once as Traefik's default certificate. All four hostnames are A records pointing at that one address.

ServiceAddressCredentials
Kubernetes API https://10.20.100.227:8443 ~/.kube/config on any master
Argo CD https://prod-argocd.dcbar.org admin · rotate after install
Portainer https://prod-portainer.dcbar.org set on first visit
Longhorn UI https://prod-longhorn.dcbar.org admin · basic auth, generated
Traefik dashboard https://prod-traefik.dcbar.org admin · basic auth, read-only

Certificate renewal

One secret, one command, no restart — Traefik reloads it live. Build the chain leaf first; the issuer's own file has only the leaf, and the intermediate comes from the CA Issuers URI inside the certificate.

# fetch the intermediate the cert itself points to
openssl x509 -in NEW.crt -noout -text | grep -A2 'Authority Information Access'
curl -sSLo inter.der <that URI> && openssl x509 -inform DER -in inter.der -out inter.pem
cat NEW.crt inter.pem > fullchain-complete.crt
openssl verify -untrusted inter.pem NEW.crt          # expect: OK

kubectl -n traefik create secret tls wildcard-tls \
  --cert=fullchain-complete.crt --key=NEW.key \
  --dry-run=client -o yaml | kubectl apply -f -

Known trade-offs

Four things in this build are deliberate compromises rather than oversights. Each is fine for an internal-tools cluster and each has a clear upgrade path.

not yet configured

No backup target

Three replicas survive a lost node. They do not survive a deleted volume, a bad upgrade, or ransomware — all three replicas take the same write. Point Longhorn at S3 or MinIO and schedule recurring snapshots before anything real lands here.

expires

Bootstrap credentials

The join token lasts 24 hours, the certificate key only 2. Delete /root/join-*.txt from the control-plane nodes once every node has joined — those two lines grant cluster membership.

substitution

Traefik, not ingress-nginx

ingress-nginx was archived in March 2026 with no further security patches, and its last release was never tested above Kubernetes 1.35. Traefik reads the same standard Ingress objects, so nothing downstream changes.

expires Nov 2026

The wildcard is a manual renewal

Nothing automates this — no cert-manager, no ACME. Put the expiry in a calendar and follow the renewal steps above. Note also that Portainer's HTTPS backend annotations live on its Service, which Helm owns: a helm upgrade portainer drops them and the site returns 500 until they are reapplied.

unencrypted

Plaintext replication

Longhorn replicates and serves RWX over the pod network without encryption. Treat this VLAN as trusted, or enable volume encryption for sensitive data.

install quirks

Two things that bite on first run

Argo CD's ApplicationSet CRD exceeds the 256KB annotation limit that client-side kubectl apply uses — it needs --server-side --force-conflicts. And Portainer now prints a one-time setup token to its pod log that the first-login form demands: kubectl -n portainer logs deploy/portainer | grep setup_token.

brief stall

RWX failover is not seamless

Each shared volume is served by one share-manager pod. Lose its node and the pod is rescheduled — clients stall for roughly a minute rather than failing, thanks to softerr. Design RWX consumers to tolerate a pause.