DC Bar · target architecture

prod-k8s Two-Site Design

HQ is primary and owns the database. AWS is a warm standby that replays HQ's write-ahead log from S3, promotable by decision rather than by accident. Cloudflare is the only public door, and the VPN carries nothing a user is waiting on.

The shape

Internet users Cloudflare DNS · WAF · load balancing no inbound ports live traffic failover HQ on-prem · 10.20.100.0/24 PRIMARY AWS · 10.40.100.0/24 STANDBY cloudflared → Traefik application pods PgBouncer < 1 ms PostgreSQL · 3 instances 1 sync standby · local disk Longhorn serves internal apps, never the database Argo CD cloudflared application pods PgBouncer PostgreSQL replica read-only · promotable pods scaled to zero until failover kube-apiserver site-to-site VPN GitOps only S3 bucket WAL archive · base backups archives WAL continuously replays WAL Replication rides S3, not the VPN. Neither user traffic nor replication depends on the tunnel.
The database lives at HQ beside the application, so a request's twenty queries stay local. HQ archives its write-ahead log to S3; the AWS cluster replays it from the same bucket over its own network. Both sites reach Cloudflare by dialling out, so no inbound port is open at either end, and the VPN's only job is letting Argo CD deploy to AWS.

Every node

Nine machines in total: six at HQ that already exist, and two in AWS with the control plane handed to EKS. What each one runs, and which PostgreSQL instance sits where.

Cloudflare the only public entry live traffic failover HQ on-prem · 10.20.100.0/24 PRIMARY · 6 VMs API VIP 10.20.100.227 · keepalived + haproxy prod-k8s-master-01 10.20.100.230 etcd · apiserver prod-k8s-master-02 10.20.100.228 etcd · apiserver prod-k8s-master-03 10.20.100.229 etcd · apiserver tainted — quorum 2 of 3, no application workloads prod-k8s-worker-01 10.20.100.231 postgres-1 primary 8 vCPU · 32 GB prod-k8s-worker-02 10.20.100.232 postgres-2 synchronous standby 8 vCPU · 32 GB prod-k8s-worker-03 10.20.100.233 postgres-3 async standby 8 vCPU · 32 GB every worker also runs application pods, Traefik, cloudflared and a Longhorn replica PostgreSQL uses a dedicated local disk per worker — never Longhorn Argo CD · Portainer · Longhorn UI — internal tools, HQ only AWS · 10.40.100.0/24 STANDBY · 2 nodes EKS managed control plane · multi-AZ · no etcd to operate AWS runs the control plane — the part most likely to rot unattended prod-aws-k8s-node-01 10.40.100.67 postgres replica-1 read-only EBS gp3 prod-aws-k8s-node-02 10.40.100.93 postgres replica-2 read-only EBS gp3 application pods and cloudflared deployed, scaled to zero no Longhorn, no keepalived, no MetalLB — nothing stateful promoted by decision, never automatically S3 bucket WAL archive · base backups archives WAL continuously replays WAL Nine machines: six at HQ, two in AWS, and a control plane AWS operates.
The three HQ masters stay tainted, so every workload — including all three PostgreSQL instances — lands on the three workers, one instance each. A worker failure therefore costs exactly one instance. In AWS the control plane is EKS's problem rather than yours, which matters for a site you touch rarely. The GitOps path from Argo CD to the AWS API server runs over the site-to-site VPN and is shown in the topology figure above.

Why replication goes through S3

A streaming standby would give lower lag, and it is the obvious choice until you look at what the primary has to hold on the replica's behalf.

CHOSEN HQ primary archive S3 restore AWS replica Nothing points back at HQ. The primary archives and forgets; the replica's health is its own problem. AVOIDED HQ primary AWS replica streams through a replication slot over the VPN The slot makes HQ retain WAL for the replica. Replica disconnects, WAL accumulates, primary's disk fills.
The difference is the direction of dependency. In the chosen design the primary owes the replica nothing; in the avoided one, a problem at the DR site becomes an outage at the production site. The price is recovery-point objective — roughly one WAL segment behind instead of a few seconds.

Who owns what

ConcernHQ on-premAWS
Live user trafficprimaryfailover
DatabasePostgreSQL, 3 instances, 1 sync standbyRead-only replica, promotable
Database storageDedicated local disk per workerEBS gp3
Application podsAll workloadsDeployed, scaled to zero
LonghornInternal apps onlynone
GitOpsArgo CD lives hereDeployed to, not from
BackupsArchives WAL + nightly base to S3Reads the same bucket
Public entrycloudflared tunnelcloudflared tunnel

What breaks when

FailureEffectRecovery
One HQ worker dies no data loss PostgreSQL fails over to the synchronous standby Automatic, seconds. This is what the sync standby buys
VPN tunnel drops serving continues Replication unaffected — it goes via S3. Argo CD cannot deploy to AWS Nothing urgent
S3 unreachable degraded HQ keeps serving, but WAL accumulates locally and the replica falls behind Fix before the WAL volume fills. Alert on archiving failure
HQ site lost outage Until AWS is promoted. RPO ≈ one WAL segment Manual promotion plus a Cloudflare change, 5–10 min. Failback needs a re-seed
AWS site lost serving continues No DR target until it returns Rebuild the replica from S3
Cloudflare unreachable outage No public path to either site Outside your control — the trade for closing every inbound port

Build order

  1. A disk per HQ workerOne dedicated device on each of .231–.233 for database volumes. Not Longhorn — Postgres does its own replication, and stacking the two means nine copies of every write.
  2. PostgreSQL at HQCloudNativePG, three instances, one synchronous standby, PgBouncer, WAL archiving to S3. Then take a backup and restore it somewhere before trusting it.
  3. cloudflared at HQTwo replicas, hostname routed to Traefik. Public traffic starts flowing with no inbound firewall rule.
  4. The AWS clusterThree nodes, kubeadm, Cilium. No keepalived, no MetalLB, no Longhorn — nothing here is stateful and the endpoint is internal to the VPC.
  5. The replica clusterBootstrap from HQ's backup in S3, then replay continuously. Watch the lag for a few days before believing it.
  6. Register AWS in Argo CDOne ApplicationSet targeting both clusters, with AWS replicas at zero. This is where two clusters stop being twice the work.
  7. Rehearse the promotionPromote AWS, serve from it, fail back. Untested failover is not failover, and the cheapest time to find a missing IAM permission is on a quiet afternoon.