Procedure

Deploying a new application to production

The full path from a repository to a working production URL at HQ. Ten stages, in order. The generator does the first part; most of what follows is manual, and the steps nobody writes down are the ones that cost a day.

Cluster HQ production, master 10.20.100.230 Takes about half a day, first time Prerequisites Azure CLI, kubectl, Argo CD, GHCR access

1Gather the facts from the repository

Everything the generator asks for comes from the application's own code. Collect it before opening the form, not while filling it in.

Read these out of the repository rather than assuming the house defaults. An application that listens on 8080 and is deployed as though it listens on 3000 produces a healthy pod behind a dead Service, which is a slow thing to diagnose.

WhatWhere it comes fromTypical value
App nameRepository name, lowercase with hyphensbilling-api
Container portEXPOSE in the Dockerfile, or the server's listen call3000
Health check pathThe route the app actually serves/healthz
NamespaceConvention: <app>-productionbilling-api-production
Ingress hostnameAgreed with the requesterbilling.dcbar.org
ImagesThe GitHub Actions workflow's push targetsghcr.io/dc-bar-web/billing-api
Secret namesEvery environment variable the app readsdatabase-url, session-secret
Needs a database?Presence of migrations or an ORM configUsually yes
Count the images, not the applications

An application is frequently three images — an API, a web front end and a worker. Each one needs its own entry in the Image Updater annotations later. Check the workflow file for how many it pushes, because a missing alias means that component silently never updates while the other two do.

Also note whether the application sends PostgreSQL timeout parameters on connect. Anything built on node-postgres or Drizzle usually does, and that matters at stage 5.


2Run the generator

k8s-deployment-generator.dcbar.org — internal only, so you need to be on the network or the VPN.

Fill in the APPLICATION block from the facts you collected, then the production overlay. Leave a field blank to accept the grey default, which is derived from the app name. Tick Include Azure Key Vault secrets — every production application here reads its configuration from a vault.

Tick Persistent storage only if the application writes files it needs to keep. A stateless API does not, and an unnecessary PVC is one more thing tying the pod to a node.

What it produces

The Generate & email button zips the manifests and sends them via SendGrid from k8s@dcbar.org to the addresses you list, with a copy to itdiagnostics@dcbar.org. Nothing is written to disk — if the mail does not arrive, the run is lost and you start again. Download setup.ps1 gives you the setup commands as a script instead.

What it does not produce

Copy the commands from the generator, not from this page

The CLUSTER SETUP COMMANDS pane updates live as you fill in the form, so what it shows is correct for the names you actually typed. The commands reproduced below are here to explain what each one is for — take the real text from the screen.


3Cluster setup commands

Four things the cluster needs before any manifest will work: an identity that can read the vault, a namespace, a way to pull the image, and a certificate.

The service principal

The Secrets Store CSI driver reads Key Vault as an Azure service principal. The generator's first block creates one:

read -r SP_CLIENT_ID SP_CLIENT_SECRET <<< "$(az ad sp create-for-rbac \
  --name app-kv-reader --years 2 --query '[appId,password]' -o tsv)"
Check before you run this

It is labelled "one-time, run FIRST". Confirm with Infrastructure whether that means once ever for the estate, or once per application. If a single principal is shared and you create another, yours will have access to no vault at all and the failure appears later as a pod stuck on a CSI mount. The client secret is shown once and is valid two years — put it in the password manager the moment it appears, and record the expiry date somewhere that gets looked at.

Namespace and image pull secret

kubectl create namespace billing-api-production

kubectl -n billing-api-production create secret docker-registry ghcr-pull-secret \
  --docker-server=ghcr.io \
  --docker-username=<gh-username> \
  --docker-password=<PAT with read:packages>
Use the organisation's machine account, not your own

The generator says <your-gh-username>. Do not take that literally. A personal token means every image pull for this application breaks when that person's token expires or they leave, and the estate already carries two personal tokens covering around forty entries between them. Ask Infrastructure for the shared credential.

CSI credentials

The driver looks for a secret named secrets-store-creds in the application's own namespace. It is not cluster-wide and it is not copied automatically — every new namespace needs its own:

kubectl -n billing-api-production create secret generic secrets-store-creds \
  --from-literal=clientid="$SP_CLIENT_ID" \
  --from-literal=clientsecret="$SP_CLIENT_SECRET"

kubectl -n billing-api-production label secret secrets-store-creds \
  secrets-store.csi.k8s.io/used=true

The label is required. Without it the driver ignores the secret and the pod waits on a mount that never completes, reporting nothing more useful than ContainerCreating.

The TLS certificate

All production hostnames are served from the *.dcbar.org wildcard. Copy it from a namespace that already has it rather than re-importing the PFX:

kubectl -n helpdesk-production get secret production-helpdesk -o yaml \
  | sed -e 's/namespace: helpdesk-production/namespace: billing-api-production/' \
        -e 's/name: production-helpdesk/name: production-billing-api/' \
        -e '/resourceVersion\|uid\|creationTimestamp\|selfLink/d' \
  | kubectl apply -f -

Then confirm what you actually copied:

kubectl -n billing-api-production get secret production-billing-api \
  -o jsonpath='{.data.tls\.crt}' | base64 -d | openssl x509 -noout -subject -dates
The wildcard expires 27 November 2026

Every production hostname depends on it. Renewal replaces the secret in every namespace that holds a copy, so the list of namespaces is itself something to keep current.


4Put the secrets in the Key Vault

The generator creates the vault and the SecretProviderClass that reads from it. What goes inside is always typed by a person.

Each secret named in the SecretProviderClass must exist in the vault before the pod starts. A missing one fails the whole mount, not just that value, so the pod never starts and the message does not say which secret was absent.

az keyvault secret set --vault-name dcbar-billing-api-prod-kv \
  --name database-url --value 'postgresql://...'

az keyvault secret list --vault-name dcbar-billing-api-prod-kv \
  --query '[].name' -o tsv

The second command is the check — compare its output against the objects list in the SecretProviderClass. They must match exactly, including hyphens.

Rotation order: vault first, restart second

The CSI driver refreshes mounted files every two minutes, but a running process never re-reads them. So changing a secret needs a rollout restart — and the restart must come after the vault is updated. Restart first and the new pods read the old value into memory, then the file changes underneath them: the pod hash changes, the behaviour does not, and it looks as though the new secret is wrong.

kubectl -n billing-api-production rollout restart deploy/billing-api

5The database

One CloudNativePG cluster serves every production application. Each application gets its own database and its own role.

prod-db.dcbar.org resolves to 10.20.100.238, which is the PgBouncer pooler app-db-pooler-rw, in front of the app-db cluster in namespace database. Create the database and role on the primary — a replica returns cannot execute CREATE DATABASE in a read-only transaction:

kubectl -n database get cluster app-db -o jsonpath='{.status.currentPrimary}{"\n"}'
kubectl -n database exec <that primary> -- psql -c \
  "CREATE ROLE \"billing-api-production_user\" LOGIN PASSWORD '<generated>'"

kubectl -n database exec <that primary> -- psql -c \
  "CREATE DATABASE \"billing-api-production\" OWNER \"billing-api-production_user\""

Follow the existing naming exactly: database <app>-production, role <app>-production_user. Then put the connection string into the vault as database-url, pointing at prod-db.dcbar.org:5432 — the pooler, not a cluster member.

PgBouncer rejects unknown startup parameters

The pooler runs in transaction mode and refuses any startup parameter not on its allow list, with FATAL 08P01 unsupported startup parameter. Clients built on node-postgres or Drizzle commonly send statement_timeout, idle_in_transaction_session_timeout and lock_timeout. The application cannot run a single query, and presents as a generic error on sign-in.

The durable fix is to set them on the role instead, on the primary, so the client never sends them:

ALTER ROLE "billing-api-production_user" SET statement_timeout = '30s';

The repository also needs changing, or the next deployment reintroduces it.


6Argo CD Application and ImageUpdater

Two objects, both written by hand, both required. Either one missing is a silent failure.

The Application

Model it on an existing production Application rather than writing from scratch:

kubectl -n argocd get app helpdesk-production -o yaml > model.yaml

The fields that change per application are repoURL, path (normally k8s/overlays/production), targetRevision (normally production), destination.namespace and the name. Keep syncPolicy.automated with prune and selfHeal enabled, as the others have.

The ImageUpdater

This cluster runs Image Updater v1, where an ImageUpdater resource decides which Applications are processed. Without one, nothing is processed and no error is logged anywhere:

cat <<'EOF' | kubectl apply -f -
apiVersion: argocd-image-updater.argoproj.io/v1alpha1
kind: ImageUpdater
metadata:
  name: billing-api-production-updater
  namespace: argocd
spec:
  applicationRefs:
  - namePattern: billing-api-production
    useAnnotations: true
EOF

useAnnotations: true means the rules come from the Application's annotations, which are in Git. Set to false, the annotations become inert text and the resource's own images block takes over — rules that live only in the cluster, where nobody reviews them.

The annotations

kubectl -n argocd annotate application billing-api-production --overwrite \
  argocd-image-updater.argoproj.io/image-list='api=ghcr.io/dc-bar-web/billing-api' \
  argocd-image-updater.argoproj.io/write-back-method='argocd' \
  argocd-image-updater.argoproj.io/api.update-strategy='newest-build' \
  argocd-image-updater.argoproj.io/api.allow-tags='regexp:^production-' \
  argocd-image-updater.argoproj.io/api.pull-secret='pullsecret:argocd/ghcr-pull-secret'

The alias before the = in image-list is the prefix every other annotation must use. Get that wrong and the annotations parse cleanly and apply to nothing. One alias per image, so a three-image application has three sets.

Always set allow-tags

Without a tag filter, the updater takes whatever is newest in the registry — which is how a staging build reached a development environment on 13 September. regexp:^production- is the minimum; regexp:^production-[0-9a-f]{7,40}$ is better where the pipeline tags by commit, because it also excludes moving tags like production-latest.

Confirm it is actually being processed

kubectl -n argocd get imageupdater -o json | jq -r \
  '.items[] | . as $i | $i.spec.applicationRefs[] |
   "\(.namePattern)\t\($i.metadata.name)\t\(.useAnnotations)"' | sort

Your application should appear exactly once, with true. Appearing twice means two resources are competing and the Git annotations may not be what is in force.

kubectl -n argocd logs deploy/argocd-image-updater-controller --tail=200 \
  | grep -i billing-api

You want a line reading images_considered=1 ... errors=0. Silence means the Application is not being processed at all.


7Migrations

A completed Job is not a migrated database.

Migration Jobs here have exited zero after failing to connect at all. The application then deploys, reports healthy, and serves against an empty schema — one did so for sixteen hours. Read the log before deleting the Job, because deleting it destroys the only evidence:

kubectl -n billing-api-production logs job/billing-api-migrate --tail=100

Then verify the database rather than the Job:

kubectl -n database exec <primary> -- psql -d billing-api-production -c \
  "select count(*) as tables from pg_tables where schemaname='public'"

A non-zero count is the confirmation. To re-run after a fix, delete the Job and sync — Argo recreates it:

kubectl -n billing-api-production delete job billing-api-migrate
argocd app sync billing-api-production --core

--core talks to the Kubernetes API directly and needs no login, but it reads argocd-cm from your context's current namespace — so run kubectl config set-context --current --namespace=argocd first or it fails with configmap "argocd-cm" not found.


8Making it reachable

Internal first. Public only if it has been asked for, because the proxy chain changes when it is.

The Ingress

Two annotations decide whether anything works:

traefik.ingress.kubernetes.io/router.entrypoints: web,websecure
traefik.ingress.kubernetes.io/router.middlewares: billing-api-production-redirect-https@kubernetescrd

websecure alone means plain HTTP returns 404 and any redirect middleware can never fire — it is attached to a router that HTTP never reaches. The middleware reference is <namespace>-<name>@kubernetescrd, and the Middleware must exist in that namespace:

kubectl -n billing-api-production get middleware

A reference to a Middleware that does not exist produces 404 on every entrypoint, with nothing logged against the Ingress. It looks like a routing problem and is a missing object.

Internal DNS

Add an A record for the hostname pointing at 10.20.100.237, the Traefik VIP. Five production hostnames already share it.

If it must be public

Add a route in the Cloudflare dashboard — the HQ tunnel's configuration is held there, not in a ConfigMap and not in Git. Service URL:

http://traefik.traefik.svc.cluster.local:80

Port 80, not 443. Traefik serves the *.dcbar.org certificate, which cannot validate against a cluster.local name, and the result is a 502 at the edge. There is no redirect loop on port 80: Traefik's web entrypoint trusts X-Forwarded-Proto from 10.244.0.0/16, which contains the cloudflared pods.

Tell the developers the hop count changed

A public hostname has four proxies ahead of the application from the internet and two from the LAN. An application that trusts a fixed hop count will mis-attribute one of those paths, which breaks sign-in rate limiting and writes wrong addresses into audit trails. Trusting by CIDR rather than by depth is the durable answer.


9Verification

Each check exercises the dependency rather than the thing in front of it. That distinction is why these are the checks and not others.

#CheckPasses when
1kubectl -n <ns> get podsAll Running, restart counts not climbing
2kubectl -n <ns> exec deploy/<app> -- ls /mnt/secrets-storeEvery secret named in the SecretProviderClass is present
3Table count on the databaseNon-zero
4argocd app get <app>-production --coreSynced and Healthy
5Image Updater logimages_considered equals your image count, errors=0
6curl -sS -o /dev/null -w '%{http_code}' https://<host>/200
7Sign in as a real userYou reach an authenticated page

Step 7 is not optional. Every failure in the September and October notes passed steps 1 to 6 and failed at 7 — healthy pods, successful syncs, a responding hostname, and an application nobody could actually log into.

If the hostname is public, test it from outside the network. Internal DNS resolves it straight to the Traefik VIP, so a curl from a cluster node returns 200 whatever state the tunnel is in:

curl -sS --resolve billing.dcbar.org:443:104.20.45.170 \
  -o /dev/null -w "%{http_code} ip=%{remote_ip}\n" https://billing.dcbar.org/

ip= showing an internal address means the test was invalid.


10Traps worth knowing before you start

Each of these cost at least half a day the first time. None of them produce an error that names the cause.

SymptomActual cause
Pod stuck ContainerCreatingsecrets-store-creds missing from the namespace, or missing its used=true label
Pod stuck, CSI mount error naming no secretA secret listed in the SecretProviderClass does not exist in the vault
Generic error on sign-in, no useful logPgBouncer rejecting a startup parameter
Application healthy, every query failsMigration Job exited zero against an empty database
404 on both HTTP and HTTPSIngress references a Middleware that does not exist
HTTP 404, HTTPS finerouter.entrypoints is websecure only
New images never deployNo ImageUpdater resource, or the alias prefix does not match image-list
Wrong environment's image deploysNo allow-tags filter
Rotated secret has no effectPods restarted before the vault was updated
502 from the public hostnameTunnel route points at port 443 instead of 80
Rate limiting treats everyone as one visitortrust proxy set for the wrong hop count
Things that exist only in the cluster

The ImageUpdater resources, the CSI credentials, the TLS secrets and the tunnel routes are not in Git and not in any repository. A rebuilt cluster needs every one of them recreated by hand, so record what you created in the application's infrastructure request.