Procedure
The full path from a repository to a working production URL at HQ. Ten stages, in order. The generator does the first part; most of what follows is manual, and the steps nobody writes down are the ones that cost a day.
Everything the generator asks for comes from the application's own code. Collect it before opening the form, not while filling it in.
Read these out of the repository rather than assuming the house defaults. An application that listens on 8080 and is deployed as though it listens on 3000 produces a healthy pod behind a dead Service, which is a slow thing to diagnose.
| What | Where it comes from | Typical value |
|---|---|---|
| App name | Repository name, lowercase with hyphens | billing-api |
| Container port | EXPOSE in the Dockerfile, or the server's listen call | 3000 |
| Health check path | The route the app actually serves | /healthz |
| Namespace | Convention: <app>-production | billing-api-production |
| Ingress hostname | Agreed with the requester | billing.dcbar.org |
| Images | The GitHub Actions workflow's push targets | ghcr.io/dc-bar-web/billing-api |
| Secret names | Every environment variable the app reads | database-url, session-secret |
| Needs a database? | Presence of migrations or an ORM config | Usually yes |
An application is frequently three images — an API, a web front end and a worker. Each one needs its own entry in the Image Updater annotations later. Check the workflow file for how many it pushes, because a missing alias means that component silently never updates while the other two do.
Also note whether the application sends PostgreSQL timeout parameters on connect. Anything built on node-postgres or Drizzle usually does, and that matters at stage 5.
k8s-deployment-generator.dcbar.org — internal only, so you need
to be on the network or the VPN.
Fill in the APPLICATION block from the facts you collected, then the production overlay. Leave a field blank to accept the grey default, which is derived from the app name. Tick Include Azure Key Vault secrets — every production application here reads its configuration from a vault.
Tick Persistent storage only if the application writes files it needs to keep. A stateless API does not, and an unnecessary PVC is one more thing tying the pod to a node.
The Generate & email button zips the manifests and sends them via
SendGrid from k8s@dcbar.org to the addresses you list, with a
copy to itdiagnostics@dcbar.org. Nothing is written to
disk — if the mail does not arrive, the run is lost and you start again.
Download setup.ps1 gives you the setup commands as a script instead.
ImageUpdater. Those
are stage 6 and are written by hand.The CLUSTER SETUP COMMANDS pane updates live as you fill in the form, so what it shows is correct for the names you actually typed. The commands reproduced below are here to explain what each one is for — take the real text from the screen.
Four things the cluster needs before any manifest will work: an identity that can read the vault, a namespace, a way to pull the image, and a certificate.
The Secrets Store CSI driver reads Key Vault as an Azure service principal. The generator's first block creates one:
read -r SP_CLIENT_ID SP_CLIENT_SECRET <<< "$(az ad sp create-for-rbac \
--name app-kv-reader --years 2 --query '[appId,password]' -o tsv)"
It is labelled "one-time, run FIRST". Confirm with Infrastructure whether that means once ever for the estate, or once per application. If a single principal is shared and you create another, yours will have access to no vault at all and the failure appears later as a pod stuck on a CSI mount. The client secret is shown once and is valid two years — put it in the password manager the moment it appears, and record the expiry date somewhere that gets looked at.
kubectl create namespace billing-api-production
kubectl -n billing-api-production create secret docker-registry ghcr-pull-secret \
--docker-server=ghcr.io \
--docker-username=<gh-username> \
--docker-password=<PAT with read:packages>
The generator says <your-gh-username>. Do not take that
literally. A personal token means every image pull for this application
breaks when that person's token expires or they leave, and the estate
already carries two personal tokens covering around forty entries between
them. Ask Infrastructure for the shared credential.
The driver looks for a secret named secrets-store-creds in the
application's own namespace. It is not cluster-wide and it is not copied
automatically — every new namespace needs its own:
kubectl -n billing-api-production create secret generic secrets-store-creds \
--from-literal=clientid="$SP_CLIENT_ID" \
--from-literal=clientsecret="$SP_CLIENT_SECRET"
kubectl -n billing-api-production label secret secrets-store-creds \
secrets-store.csi.k8s.io/used=true
The label is required. Without it the driver ignores the secret and the pod
waits on a mount that never completes, reporting nothing more useful than
ContainerCreating.
All production hostnames are served from the *.dcbar.org
wildcard. Copy it from a namespace that already has it rather than
re-importing the PFX:
kubectl -n helpdesk-production get secret production-helpdesk -o yaml \
| sed -e 's/namespace: helpdesk-production/namespace: billing-api-production/' \
-e 's/name: production-helpdesk/name: production-billing-api/' \
-e '/resourceVersion\|uid\|creationTimestamp\|selfLink/d' \
| kubectl apply -f -
Then confirm what you actually copied:
kubectl -n billing-api-production get secret production-billing-api \
-o jsonpath='{.data.tls\.crt}' | base64 -d | openssl x509 -noout -subject -dates
Every production hostname depends on it. Renewal replaces the secret in every namespace that holds a copy, so the list of namespaces is itself something to keep current.
The generator creates the vault and the SecretProviderClass that reads from it. What goes inside is always typed by a person.
Each secret named in the SecretProviderClass must exist in the vault before the pod starts. A missing one fails the whole mount, not just that value, so the pod never starts and the message does not say which secret was absent.
az keyvault secret set --vault-name dcbar-billing-api-prod-kv \
--name database-url --value 'postgresql://...'
az keyvault secret list --vault-name dcbar-billing-api-prod-kv \
--query '[].name' -o tsv
The second command is the check — compare its output against the
objects list in the SecretProviderClass. They must match
exactly, including hyphens.
The CSI driver refreshes mounted files every two minutes, but a running
process never re-reads them. So changing a secret needs a
rollout restart — and the restart must come after the
vault is updated. Restart first and the new pods read the old value into
memory, then the file changes underneath them: the pod hash changes, the
behaviour does not, and it looks as though the new secret is wrong.
kubectl -n billing-api-production rollout restart deploy/billing-api
One CloudNativePG cluster serves every production application. Each application gets its own database and its own role.
prod-db.dcbar.org resolves to 10.20.100.238, which
is the PgBouncer pooler app-db-pooler-rw, in front of the
app-db cluster in namespace database. Create the
database and role on the primary — a replica returns
cannot execute CREATE DATABASE in a read-only transaction:
kubectl -n database get cluster app-db -o jsonpath='{.status.currentPrimary}{"\n"}'
kubectl -n database exec <that primary> -- psql -c \
"CREATE ROLE \"billing-api-production_user\" LOGIN PASSWORD '<generated>'"
kubectl -n database exec <that primary> -- psql -c \
"CREATE DATABASE \"billing-api-production\" OWNER \"billing-api-production_user\""
Follow the existing naming exactly: database <app>-production,
role <app>-production_user. Then put the connection string
into the vault as database-url, pointing at
prod-db.dcbar.org:5432 — the pooler, not a cluster member.
The pooler runs in transaction mode and refuses any startup parameter not
on its allow list, with
FATAL 08P01 unsupported startup parameter. Clients built on
node-postgres or Drizzle commonly send statement_timeout,
idle_in_transaction_session_timeout and
lock_timeout. The application cannot run a single query, and
presents as a generic error on sign-in.
The durable fix is to set them on the role instead, on the primary, so the client never sends them:
ALTER ROLE "billing-api-production_user" SET statement_timeout = '30s';
The repository also needs changing, or the next deployment reintroduces it.
Two objects, both written by hand, both required. Either one missing is a silent failure.
Model it on an existing production Application rather than writing from scratch:
kubectl -n argocd get app helpdesk-production -o yaml > model.yaml
The fields that change per application are repoURL,
path (normally k8s/overlays/production),
targetRevision (normally production),
destination.namespace and the name. Keep
syncPolicy.automated with prune and
selfHeal enabled, as the others have.
This cluster runs Image Updater v1, where an ImageUpdater
resource decides which Applications are processed. Without one, nothing is
processed and no error is logged anywhere:
cat <<'EOF' | kubectl apply -f -
apiVersion: argocd-image-updater.argoproj.io/v1alpha1
kind: ImageUpdater
metadata:
name: billing-api-production-updater
namespace: argocd
spec:
applicationRefs:
- namePattern: billing-api-production
useAnnotations: true
EOF
useAnnotations: true means the rules come from the
Application's annotations, which are in Git. Set to false, the
annotations become inert text and the resource's own
images block takes over — rules that live only in the cluster,
where nobody reviews them.
kubectl -n argocd annotate application billing-api-production --overwrite \
argocd-image-updater.argoproj.io/image-list='api=ghcr.io/dc-bar-web/billing-api' \
argocd-image-updater.argoproj.io/write-back-method='argocd' \
argocd-image-updater.argoproj.io/api.update-strategy='newest-build' \
argocd-image-updater.argoproj.io/api.allow-tags='regexp:^production-' \
argocd-image-updater.argoproj.io/api.pull-secret='pullsecret:argocd/ghcr-pull-secret'
The alias before the = in image-list is the prefix
every other annotation must use. Get that wrong and the annotations parse
cleanly and apply to nothing. One alias per image, so a three-image
application has three sets.
Without a tag filter, the updater takes whatever is newest in the
registry — which is how a staging build reached a development environment
on 13 September. regexp:^production- is the minimum;
regexp:^production-[0-9a-f]{7,40}$ is better where the
pipeline tags by commit, because it also excludes moving tags like
production-latest.
kubectl -n argocd get imageupdater -o json | jq -r \
'.items[] | . as $i | $i.spec.applicationRefs[] |
"\(.namePattern)\t\($i.metadata.name)\t\(.useAnnotations)"' | sort
Your application should appear exactly once, with
true. Appearing twice means two resources are competing and
the Git annotations may not be what is in force.
kubectl -n argocd logs deploy/argocd-image-updater-controller --tail=200 \
| grep -i billing-api
You want a line reading images_considered=1 ... errors=0.
Silence means the Application is not being processed at all.
A completed Job is not a migrated database.
Migration Jobs here have exited zero after failing to connect at all. The application then deploys, reports healthy, and serves against an empty schema — one did so for sixteen hours. Read the log before deleting the Job, because deleting it destroys the only evidence:
kubectl -n billing-api-production logs job/billing-api-migrate --tail=100
Then verify the database rather than the Job:
kubectl -n database exec <primary> -- psql -d billing-api-production -c \
"select count(*) as tables from pg_tables where schemaname='public'"
A non-zero count is the confirmation. To re-run after a fix, delete the Job and sync — Argo recreates it:
kubectl -n billing-api-production delete job billing-api-migrate
argocd app sync billing-api-production --core
--core talks to the Kubernetes API directly and needs no
login, but it reads argocd-cm from your context's current
namespace — so run kubectl config set-context --current --namespace=argocd
first or it fails with configmap "argocd-cm" not found.
Internal first. Public only if it has been asked for, because the proxy chain changes when it is.
Two annotations decide whether anything works:
traefik.ingress.kubernetes.io/router.entrypoints: web,websecure
traefik.ingress.kubernetes.io/router.middlewares: billing-api-production-redirect-https@kubernetescrd
websecure alone means plain HTTP returns 404 and any redirect
middleware can never fire — it is attached to a router that HTTP never
reaches. The middleware reference is
<namespace>-<name>@kubernetescrd, and the Middleware
must exist in that namespace:
kubectl -n billing-api-production get middleware
A reference to a Middleware that does not exist produces 404 on every entrypoint, with nothing logged against the Ingress. It looks like a routing problem and is a missing object.
Add an A record for the hostname pointing at 10.20.100.237, the
Traefik VIP. Five production hostnames already share it.
Add a route in the Cloudflare dashboard — the HQ tunnel's configuration is held there, not in a ConfigMap and not in Git. Service URL:
http://traefik.traefik.svc.cluster.local:80
Port 80, not 443. Traefik serves the *.dcbar.org certificate,
which cannot validate against a cluster.local name, and the
result is a 502 at the edge. There is no redirect loop on port 80: Traefik's
web entrypoint trusts X-Forwarded-Proto from
10.244.0.0/16, which contains the cloudflared pods.
A public hostname has four proxies ahead of the application from the internet and two from the LAN. An application that trusts a fixed hop count will mis-attribute one of those paths, which breaks sign-in rate limiting and writes wrong addresses into audit trails. Trusting by CIDR rather than by depth is the durable answer.
Each check exercises the dependency rather than the thing in front of it. That distinction is why these are the checks and not others.
| # | Check | Passes when |
|---|---|---|
| 1 | kubectl -n <ns> get pods | All Running, restart counts not climbing |
| 2 | kubectl -n <ns> exec deploy/<app> -- ls /mnt/secrets-store | Every secret named in the SecretProviderClass is present |
| 3 | Table count on the database | Non-zero |
| 4 | argocd app get <app>-production --core | Synced and Healthy |
| 5 | Image Updater log | images_considered equals your image count, errors=0 |
| 6 | curl -sS -o /dev/null -w '%{http_code}' https://<host>/ | 200 |
| 7 | Sign in as a real user | You reach an authenticated page |
Step 7 is not optional. Every failure in the September and October notes passed steps 1 to 6 and failed at 7 — healthy pods, successful syncs, a responding hostname, and an application nobody could actually log into.
If the hostname is public, test it from outside the network. Internal DNS
resolves it straight to the Traefik VIP, so a curl from a
cluster node returns 200 whatever state the tunnel is in:
curl -sS --resolve billing.dcbar.org:443:104.20.45.170 \
-o /dev/null -w "%{http_code} ip=%{remote_ip}\n" https://billing.dcbar.org/
ip= showing an internal address means the test was invalid.
Each of these cost at least half a day the first time. None of them produce an error that names the cause.
| Symptom | Actual cause |
|---|---|
Pod stuck ContainerCreating | secrets-store-creds missing from the namespace, or missing its used=true label |
| Pod stuck, CSI mount error naming no secret | A secret listed in the SecretProviderClass does not exist in the vault |
| Generic error on sign-in, no useful log | PgBouncer rejecting a startup parameter |
| Application healthy, every query fails | Migration Job exited zero against an empty database |
| 404 on both HTTP and HTTPS | Ingress references a Middleware that does not exist |
| HTTP 404, HTTPS fine | router.entrypoints is websecure only |
| New images never deploy | No ImageUpdater resource, or the alias prefix does not match image-list |
| Wrong environment's image deploys | No allow-tags filter |
| Rotated secret has no effect | Pods restarted before the vault was updated |
| 502 from the public hostname | Tunnel route points at port 443 instead of 80 |
| Rate limiting treats everyone as one visitor | trust proxy set for the wrong hop count |
The ImageUpdater resources, the CSI credentials, the TLS
secrets and the tunnel routes are not in Git and not in any repository.
A rebuilt cluster needs every one of them recreated by hand, so record
what you created in the application's infrastructure request.