DC Bar · target architecture
HQ is primary and owns the database. AWS is a warm standby that replays HQ's write-ahead log from S3, promotable by decision rather than by accident. Cloudflare is the only public door, and the VPN carries nothing a user is waiting on.
Nine machines in total: six at HQ that already exist, and two in AWS with the control plane handed to EKS. What each one runs, and which PostgreSQL instance sits where.
A streaming standby would give lower lag, and it is the obvious choice until you look at what the primary has to hold on the replica's behalf.
| Concern | HQ on-prem | AWS |
|---|---|---|
| Live user traffic | primary | failover |
| Database | PostgreSQL, 3 instances, 1 sync standby | Read-only replica, promotable |
| Database storage | Dedicated local disk per worker | EBS gp3 |
| Application pods | All workloads | Deployed, scaled to zero |
| Longhorn | Internal apps only | none |
| GitOps | Argo CD lives here | Deployed to, not from |
| Backups | Archives WAL + nightly base to S3 | Reads the same bucket |
| Public entry | cloudflared tunnel | cloudflared tunnel |
| Failure | Effect | Recovery |
|---|---|---|
| One HQ worker dies | no data loss PostgreSQL fails over to the synchronous standby | Automatic, seconds. This is what the sync standby buys |
| VPN tunnel drops | serving continues Replication unaffected — it goes via S3. Argo CD cannot deploy to AWS | Nothing urgent |
| S3 unreachable | degraded HQ keeps serving, but WAL accumulates locally and the replica falls behind | Fix before the WAL volume fills. Alert on archiving failure |
| HQ site lost | outage Until AWS is promoted. RPO ≈ one WAL segment | Manual promotion plus a Cloudflare change, 5–10 min. Failback needs a re-seed |
| AWS site lost | serving continues No DR target until it returns | Rebuild the replica from S3 |
| Cloudflare unreachable | outage No public path to either site | Outside your control — the trade for closing every inbound port |