PostgreSQL
Replica replaying HQ's WAL from S3, ~5 min behind. Endpoint, certificate and hostname verification all tested from HQ.
Disaster recovery · HQ to AWS
What to do when the Washington DC datacentre is gone and the AWS site has to take over — including the parts that are not built yet.
Applications now run at AWS and connect back to the HQ database over the VPN, using the name prod-db.dcbar.org. That makes AWS a second front end for HQ today — useful, but it means a HQ failure takes both sites down until this procedure is run.
The AWS PostgreSQL replica is not in the serving path. It exists solely to be promoted, and the failover is: promote it, repoint prod-db.dcbar.org, restart the applications. No connection strings change.
The database half of this procedure has been exercised without a cutover. Both clusters now serve the *.dcbar.org certificate, and a client verifying with sslmode=verify-full sslrootcert=system connects to either site under the name prod-db.dcbar.org. Phase 3, step 6 is no longer a leap of faith.
PostgreSQL
Replica replaying HQ's WAL from S3, ~5 min behind. Endpoint, certificate and hostname verification all tested from HQ.
Ingress & tunnel
Traefik and a healthy dcbar-aws connector, with hello-world-tim and nom-portal routed to it.
Applications
nom-portal runs at both sites, deployed per-overlay. No ApplicationSet, so failover is a manual deploy rather than a replica count.
Argo CD
Runs only at HQ. A HQ outage takes your deployment tooling with it.
Everything here is reversible. Nothing after Phase 2 is.
Not slow. Not partitioned. Down, and not coming back within your tolerance for being offline.
If HQ is actually alive but unreachable from where you are standing, promoting AWS gives you two writable databases accepting different writes. Reconciling that is manual, slow, and lossy. A network partition looks exactly like a dead site from one side.
Check from a second vantage point — a different network, a phone hotspot — before deciding.
kubectl --context kubernetes-admin@kubernetes get nodes
curl -s -o /dev/null -w 'http=%{http_code} verify=%{ssl_verify_result}\n' https://prod-argocd.dcbar.org/
ping 10.20.100.227 # the API-server VIP
No -k. The certificate is publicly trusted, so verify=0 is part of the signal — and -k is how a real chain problem stays hidden for a week.
Scaling the HQ cluster to zero guarantees it cannot accept another write or archive another segment while you promote. Skip this only if HQ is genuinely unreachable.
kubectl --context kubernetes-admin@kubernetes -n database \
patch cluster app-db --type merge -p '{"spec":{"instances":0}}'
The point of no return. After this, HQ is behind and stays behind.
Whatever WAL already reached S3 is still recoverable. Watch the lag until it stops shrinking — that is everything HQ managed to archive before it died.
kubectl --context aws -n database exec app-db-1 -c postgres -- psql -X -c \
"SELECT pg_is_in_recovery() AS replica,
pg_last_xact_replay_timestamp() AS last_replayed,
now() - pg_last_xact_replay_timestamp() AS lag;"
Run it two or three times a minute apart. When last_replayed stops advancing, there is nothing left to replay.
One field. The cluster stops following HQ and becomes writable.
kubectl --context aws -n database \
patch cluster app-db --type merge -p '{"spec":{"replica":{"enabled":false}}}'
Then confirm it actually accepts a write, rather than assuming:
kubectl --context aws -n database exec app-db-1 -c postgres -- \
psql -d nomination-portal-prod -X -c \
"CREATE TABLE IF NOT EXISTS failover_marker(at timestamptz default now());
INSERT INTO failover_marker DEFAULT VALUES;
SELECT pg_is_in_recovery(), count(*) FROM failover_marker;"
pg_is_in_recovery returns f and the insert succeeds. AWS now begins archiving its own WAL into s3://dcbar-pg-backups/aws, deliberately separate from HQ's history so nothing overwrites it.
Written for the world where the ApplicationSet exists. Today this phase is done by hand.
Once the ApplicationSet is built: change the AWS replica count in Git and let Argo CD apply it — assuming a second Argo CD runs at AWS, since the HQ one is gone.
Today: deploy by hand with kubectl --context aws. Whatever you apply, write it down as you go; that record becomes the ApplicationSet.
kubectl --context aws get deploy -A
Applications at both sites connect to prod-db.dcbar.org, never to a cluster-local service. So the database failover is a single DNS change: point that record at the AWS load balancer.
prod-db.dcbar.org → a710caeabde924bd78dd0b5cf5767523-f05aa628024352eb.elb.us-east-1.amazonaws.com
(currently 10.40.100.162, but use the name — NLB addresses change)
Since 2 Sep 2026 both clusters present the same *.dcbar.org certificate, signed by Network Solutions and trusted by every client out of the box. The name prod-db.dcbar.org matches the wildcard at either site, so verification survives the switch with no CA bundle, no mounted certificate and no redeploy.
Two things this depends on. The record's TTL must be short (60s) or clients cache the old address for as long as the TTL says. And existing pooled connections survive a DNS change, so the application deployments must be restarted to pick up the new address.
kubectl --context aws -n <app-namespace> rollout restart deploy --all
hostaddr chooses where the connection goes; host is still the name used for certificate verification. So the entire cutover path — routing, allow list, TLS, hostname match, data — can be proven against AWS while HQ stays live:
psql "host=prod-db.dcbar.org hostaddr=10.40.100.162 \
sslmode=verify-full sslrootcert=system \
dbname=nomination-portal-prod user=dbadmin" \
-c 'SELECT pg_is_in_recovery(), count(*) FROM submissions'
t plus a row count is a pass. A write attempt failing read-only is correct before promotion, not a fault.
The cutover itself. Fast, and reversible in one click.
In Cloudflare, delete the public hostname from dcbar-hq and add the same name to dcbar-aws, pointing at http://traefik.traefik.svc.cluster.local:80.
Cloudflare creates one DNS record per hostname. That constraint is what makes the cutover a single deliberate act rather than a race between two sites — and it is why the two tunnels were built with separate tokens.
Names such as prod-db.dcbar.org and prod-portainer.dcbar.org resolve to 10.20.100.x addresses that are now gone. Anything on the office network still trying to reach the old site will hang rather than fail fast.
Address AWS directly with the Host header, over the VPN. This works at any time — including right now — and tells you whether a cutover would land somewhere healthy.
curl -sI -H 'Host: hello-world-tim.dcbar.org' http://10.40.100.222/ | head -3
A 200 means the whole chain is intact. A 404 from Traefik means the request arrived but no Ingress matched — the app is not deployed at AWS.
Test over HTTP, not HTTPS. Addressing Traefik by IP sends no SNI, so it answers with its own self-signed default certificate and every HTTPS check looks broken whether it is or not.
To ask Traefik itself what it will match, rather than inferring it from YAML:
kubectl --context aws -n traefik port-forward deploy/traefik 9000:8080
# then, in another shell
curl -s localhost:9000/api/http/routers | python3 -c "
import sys, json
for r in json.load(sys.stdin):
print('%-55s -> %s' % (r.get('rule',''), r.get('service','')))
"
Not from a machine on the office network — from a phone on cellular, or anywhere that has to traverse Cloudflare like a real user.
curl -sS -o /dev/null -w '%{http_code}\n' https://<your-public-hostname>/
The AWS watchdog is looking for replication lag on a cluster that is no longer a replica. Left alone it will email that AWS "is no longer a replica" every ten minutes — correct, but no longer news.
kubectl --context aws -n database set env cronjob/db-watchdog MODE=hq
kubectl --context aws -n database patch cronjob pg-heartbeat \
-p '{"spec":{"suspend":false}}'
The heartbeat matters again now: without writes, archive_timeout never fires and AWS stops archiving its own backups.
Plan for this before you need it, not during.
The moment AWS accepts one write, HQ is stale and cannot resume as primary — its timeline has diverged. Failing back means rebuilding HQ as a replica of AWS, waiting for it to catch up, and then performing this same procedure in reverse, deliberately, during a quiet window.
Do not attempt to restart the old HQ primary and let the two reconcile. They will not.
Each of these turns a manual step above into an automatic one.
kubectl and say so here.prod-db.dcbar.orgStill on Auto. The cutover is a DNS edit, so the TTL is the floor on how long clients keep talking to a dead site. Lower it before you need it, not during.kubectl patch are reverted on the next Argo CD sync, producing new ReplicaSets at unpredictable moments — each one re-rolling the dice on which node a pod lands on, so the same fault reappears looking like a different problem. Also still in Git: the leading space in destination.namespace.Each of these cost real time to find. None produced an obvious error message.
kubectl.kubernetes.io/restartedAt. An unrecognised one is ignored silently — the cluster never restarts, and every downstream check still passes.externalClusters and a restore looks in an empty catalog: no target backup found.service.spec.type, not service.type. The wrong key is accepted, stored in the release values where it looks correct, and read by nothing — leaving a LoadBalancer, which on EKS means an ELB nobody asked for.web entrypoint already binds it inside the pod. A second entrypoint on the same container port crash-loops.aws eks get-token and turn a working kubeconfig into Unauthorized.SELF_SIGNED_CERT_IN_CHAIN — a TLS error at the worst moment with nothing wrong with the data. Both clusters now serve the public *.dcbar.org certificate instead, so no client carries a CA at all. If a cluster is ever rebuilt, re-apply db-server-tls / db-server-ca and the certificates block, or it silently reverts to its own CA.serverAltDNSNames and serverTLSSecret are mutually exclusive. CloudNativePG only adds names to certificates it issues itself, so supplying your own is rejected until the alt-names list is removed in the same patch ("serverAltDNSNames":null).sslmode=verify-full alone looks for ~/.postgresql/root.crt and fails when it is absent, even for a publicly trusted certificate. Add sslrootcert=system (Postgres 16+).10.244.0.0/16. Omit it and a pod scheduled onto the node that owns the MetalLB VIP hairpins into kube-proxy's reject rule and gets EHOSTUNREACH, while pods on every other node work normally via the service short-circuit. It presents as one replica of three crash-looping, and it moves whenever the VIP moves.dcbar-hq sends the request to HQ no matter how correct the AWS routes look — verify with the DNS tab, or the tunnel's Live logs, which stay empty if the request never arrives.pg_stat_replication. It replays WAL from S3 rather than streaming, so the primary has no connection to report. An empty replication view is expected here, not evidence that replication is broken — check pg_last_xact_replay_timestamp() on the replica instead.10.40.100.0/24, so five of its six addresses were unreachable. Pin it with aws-load-balancer-subnets rather than widening the tunnel.aws-load-balancer-internal: "true". The newer aws-load-balancer-scheme annotation belongs to the AWS Load Balancer Controller, which eksctl does not install — so it is ignored, and the default is internet-facing. Always set loadBalancerSourceRanges too, and check the address you get back is a VPC one.weight only adds a bonus while healthy. To move a VIP on failure it must be negative.| What | Where | Note |
|---|---|---|
| HQ context | kubernetes-admin@kubernetes | 3 masters, 3 workers |
| AWS context | aws | EKS, 2 nodes, us-east-1 |
| API-server VIP | 10.20.100.227 | keepalived + haproxy |
| Traefik | 10.20.100.237 | Also Portainer's edge tunnel on 8000 |
| Database (LAN) | prod-db.dcbar.org → 10.20.100.238:5432 | MetalLB → PgBouncer. nomination-portal-prod / appuser; dbadmin for administration. Source ranges must include 10.244.0.0/16. |
| Database (AWS) | a710caeabde924bd78dd0b5cf5767523-f05aa628024352eb.elb.us-east-1.amazonaws.com:5432 | Internal NLB → PgBouncer, pinned to subnet-0bebfb4e758075240 (10.40.100.0/24), currently 10.40.100.162. Read-only until promoted. |
| Database certificate | db-server-tls / db-server-ca | The *.dcbar.org wildcard, in the database namespace of both clusters. Clients need only sslmode=verify-full sslrootcert=system. |
| Replication | barman → s3://dcbar-pg-backups/hq | Archive replay, not streaming. ~5 min behind under load; unbounded when idle, which is what the heartbeat exists to prevent. That lag is the RPO. |
| Traefik (AWS) | aa9888774376443ebabbc395bc4e868b-4e0bdcb481c32ca5.elb.us-east-1.amazonaws.com → 10.40.100.222 | Internal NLB, same subnet. Reach a site before cutover with curl -H 'Host: <name>' http://10.40.100.222/ — HTTP, since addressing by IP defeats SNI. |
| VPN reach | 10.0.0.0/16, 10.40.0.0/16 | Extended 2 Sep 2026. Previously carried only 10.40.100.0/24, so pods scheduled onto other nodes silently lost all access to HQ — verify after any VPC or node-group change. |
| Backups | s3://dcbar-pg-backups/hq | AWS archives to /aws after promotion |
| Tunnels | dcbar-hq / dcbar-aws | Separate tokens, 2 connectors each |
| Alerts | itdiagnostics@dcbar.org | Via SendGrid, tested from both sites |