Disaster recovery · HQ to AWS

Failing over to AWS

What to do when the Washington DC datacentre is gone and the AWS site has to take over — including the parts that are not built yet.

Written 1 Sep 2026 Revised 2 Sep 2026 HQ: 10.20.100.0/24 AWS: EKS, us-east-1 Database path rehearsed

Read this before anything else

Applications now run at AWS and connect back to the HQ database over the VPN, using the name prod-db.dcbar.org. That makes AWS a second front end for HQ today — useful, but it means a HQ failure takes both sites down until this procedure is run.

The AWS PostgreSQL replica is not in the serving path. It exists solely to be promoted, and the failover is: promote it, repoint prod-db.dcbar.org, restart the applications. No connection strings change.

Proven 2 Sep 2026

The database half of this procedure has been exercised without a cutover. Both clusters now serve the *.dcbar.org certificate, and a client verifying with sslmode=verify-full sslrootcert=system connects to either site under the name prod-db.dcbar.org. Phase 3, step 6 is no longer a leap of faith.

Ready

PostgreSQL

Replica replaying HQ's WAL from S3, ~5 min behind. Endpoint, certificate and hostname verification all tested from HQ.

Ready

Ingress & tunnel

Traefik and a healthy dcbar-aws connector, with hello-world-tim and nom-portal routed to it.

By hand

Applications

nom-portal runs at both sites, deployed per-overlay. No ApplicationSet, so failover is a manual deploy rather than a replica count.

Missing

Argo CD

Runs only at HQ. A HQ outage takes your deployment tooling with it.

Phase 1 — Decide, then stop the primary

Everything here is reversible. Nothing after Phase 2 is.

1

Confirm HQ is down and staying down

Not slow. Not partitioned. Down, and not coming back within your tolerance for being offline.

Why this gate exists

If HQ is actually alive but unreachable from where you are standing, promoting AWS gives you two writable databases accepting different writes. Reconciling that is manual, slow, and lossy. A network partition looks exactly like a dead site from one side.

Check from a second vantage point — a different network, a phone hotspot — before deciding.

kubectl --context kubernetes-admin@kubernetes get nodes
curl -s -o /dev/null -w 'http=%{http_code} verify=%{ssl_verify_result}\n' https://prod-argocd.dcbar.org/
ping 10.20.100.227   # the API-server VIP

No -k. The certificate is publicly trusted, so verify=0 is part of the signal — and -k is how a real chain problem stays hidden for a week.

2

If HQ is reachable at all, stop it writing

Scaling the HQ cluster to zero guarantees it cannot accept another write or archive another segment while you promote. Skip this only if HQ is genuinely unreachable.

kubectl --context kubernetes-admin@kubernetes -n database \
  patch cluster app-db --type merge -p '{"spec":{"instances":0}}'

Phase 2 — Promote the database

The point of no return. After this, HQ is behind and stays behind.

3

Let the replica finish replaying

Whatever WAL already reached S3 is still recoverable. Watch the lag until it stops shrinking — that is everything HQ managed to archive before it died.

kubectl --context aws -n database exec app-db-1 -c postgres -- psql -X -c \
  "SELECT pg_is_in_recovery() AS replica,
          pg_last_xact_replay_timestamp() AS last_replayed,
          now() - pg_last_xact_replay_timestamp() AS lag;"

Run it two or three times a minute apart. When last_replayed stops advancing, there is nothing left to replay.

4

Promote

One field. The cluster stops following HQ and becomes writable.

kubectl --context aws -n database \
  patch cluster app-db --type merge -p '{"spec":{"replica":{"enabled":false}}}'

Then confirm it actually accepts a write, rather than assuming:

kubectl --context aws -n database exec app-db-1 -c postgres -- \
  psql -d nomination-portal-prod -X -c \
  "CREATE TABLE IF NOT EXISTS failover_marker(at timestamptz default now());
   INSERT INTO failover_marker DEFAULT VALUES;
   SELECT pg_is_in_recovery(), count(*) FROM failover_marker;"
What good looks like

pg_is_in_recovery returns f and the insert succeeds. AWS now begins archiving its own WAL into s3://dcbar-pg-backups/aws, deliberately separate from HQ's history so nothing overwrites it.

Phase 3 — Bring applications up

Written for the world where the ApplicationSet exists. Today this phase is done by hand.

5

Scale the AWS applications from zero

Once the ApplicationSet is built: change the AWS replica count in Git and let Argo CD apply it — assuming a second Argo CD runs at AWS, since the HQ one is gone.

Today: deploy by hand with kubectl --context aws. Whatever you apply, write it down as you go; that record becomes the ApplicationSet.

kubectl --context aws get deploy -A
6

Repoint the database name — this is the cutover

Applications at both sites connect to prod-db.dcbar.org, never to a cluster-local service. So the database failover is a single DNS change: point that record at the AWS load balancer.

prod-db.dcbar.org  →  a710caeabde924bd78dd0b5cf5767523-f05aa628024352eb.elb.us-east-1.amazonaws.com
                       (currently 10.40.100.162, but use the name — NLB addresses change)
Why this works without touching any application

Since 2 Sep 2026 both clusters present the same *.dcbar.org certificate, signed by Network Solutions and trusted by every client out of the box. The name prod-db.dcbar.org matches the wildcard at either site, so verification survives the switch with no CA bundle, no mounted certificate and no redeploy.

Two things this depends on. The record's TTL must be short (60s) or clients cache the old address for as long as the TTL says. And existing pooled connections survive a DNS change, so the application deployments must be restarted to pick up the new address.

kubectl --context aws -n <app-namespace> rollout restart deploy --all
Rehearse this without touching DNS

hostaddr chooses where the connection goes; host is still the name used for certificate verification. So the entire cutover path — routing, allow list, TLS, hostname match, data — can be proven against AWS while HQ stays live:

psql "host=prod-db.dcbar.org hostaddr=10.40.100.162 \
      sslmode=verify-full sslrootcert=system \
      dbname=nomination-portal-prod user=dbadmin" \
  -c 'SELECT pg_is_in_recovery(), count(*) FROM submissions'

t plus a row count is a pass. A write attempt failing read-only is correct before promotion, not a fault.

Phase 4 — Move the traffic

The cutover itself. Fast, and reversible in one click.

7

Move the public hostname to the AWS tunnel

In Cloudflare, delete the public hostname from dcbar-hq and add the same name to dcbar-aws, pointing at http://traefik.traefik.svc.cluster.local:80.

Why one name can only live on one tunnel

Cloudflare creates one DNS record per hostname. That constraint is what makes the cutover a single deliberate act rather than a race between two sites — and it is why the two tunnels were built with separate tokens.

8

Fix internal DNS

Names such as prod-db.dcbar.org and prod-portainer.dcbar.org resolve to 10.20.100.x addresses that are now gone. Anything on the office network still trying to reach the old site will hang rather than fail fast.

Phase 5 — Verify, then watch

9

Prove the AWS site serves, before moving any traffic

Address AWS directly with the Host header, over the VPN. This works at any time — including right now — and tells you whether a cutover would land somewhere healthy.

curl -sI -H 'Host: hello-world-tim.dcbar.org' http://10.40.100.222/ | head -3

A 200 means the whole chain is intact. A 404 from Traefik means the request arrived but no Ingress matched — the app is not deployed at AWS.

Test over HTTP, not HTTPS. Addressing Traefik by IP sends no SNI, so it answers with its own self-signed default certificate and every HTTPS check looks broken whether it is or not.

To ask Traefik itself what it will match, rather than inferring it from YAML:

kubectl --context aws -n traefik port-forward deploy/traefik 9000:8080
# then, in another shell
curl -s localhost:9000/api/http/routers | python3 -c "
import sys, json
for r in json.load(sys.stdin):
    print('%-55s -> %s' % (r.get('rule',''), r.get('service','')))
"
10

Prove it from outside

Not from a machine on the office network — from a phone on cellular, or anywhere that has to traverse Cloudflare like a real user.

curl -sS -o /dev/null -w '%{http_code}\n' https://<your-public-hostname>/
11

Switch the watchdog to primary mode

The AWS watchdog is looking for replication lag on a cluster that is no longer a replica. Left alone it will email that AWS "is no longer a replica" every ten minutes — correct, but no longer news.

kubectl --context aws -n database set env cronjob/db-watchdog MODE=hq
kubectl --context aws -n database patch cronjob pg-heartbeat \
  -p '{"spec":{"suspend":false}}'

The heartbeat matters again now: without writes, archive_timeout never fires and AWS stops archiving its own backups.

Failing back costs a re-seed

Plan for this before you need it, not during.

The moment AWS accepts one write, HQ is stale and cannot resume as primary — its timeline has diverged. Failing back means rebuilding HQ as a replica of AWS, waiting for it to catch up, and then performing this same procedure in reverse, deliberately, during a quiet window.

Do not attempt to restart the old HQ primary and let the two reconcile. They will not.

What has to be built before this is a real plan

Each of these turns a manual step above into an automatic one.

Traps already discovered

Each of these cost real time to find. None produced an obvious error message.

Reference

WhatWhereNote
HQ contextkubernetes-admin@kubernetes3 masters, 3 workers
AWS contextawsEKS, 2 nodes, us-east-1
API-server VIP10.20.100.227keepalived + haproxy
Traefik10.20.100.237Also Portainer's edge tunnel on 8000
Database (LAN)prod-db.dcbar.org → 10.20.100.238:5432MetalLB → PgBouncer. nomination-portal-prod / appuser; dbadmin for administration. Source ranges must include 10.244.0.0/16.
Database (AWS)a710caeabde924bd78dd0b5cf5767523-f05aa628024352eb.elb.us-east-1.amazonaws.com:5432Internal NLB → PgBouncer, pinned to subnet-0bebfb4e758075240 (10.40.100.0/24), currently 10.40.100.162. Read-only until promoted.
Database certificatedb-server-tls / db-server-caThe *.dcbar.org wildcard, in the database namespace of both clusters. Clients need only sslmode=verify-full sslrootcert=system.
Replicationbarman → s3://dcbar-pg-backups/hqArchive replay, not streaming. ~5 min behind under load; unbounded when idle, which is what the heartbeat exists to prevent. That lag is the RPO.
Traefik (AWS)aa9888774376443ebabbc395bc4e868b-4e0bdcb481c32ca5.elb.us-east-1.amazonaws.com → 10.40.100.222Internal NLB, same subnet. Reach a site before cutover with curl -H 'Host: <name>' http://10.40.100.222/ — HTTP, since addressing by IP defeats SNI.
VPN reach10.0.0.0/16, 10.40.0.0/16Extended 2 Sep 2026. Previously carried only 10.40.100.0/24, so pods scheduled onto other nodes silently lost all access to HQ — verify after any VPC or node-group change.
Backupss3://dcbar-pg-backups/hqAWS archives to /aws after promotion
Tunnelsdcbar-hq / dcbar-awsSeparate tokens, 2 connectors each
Alertsitdiagnostics@dcbar.orgVia SendGrid, tested from both sites