Operations notes

What we learned bringing four applications into production

Nine failures from September and October 2026, each of which looked like something it was not. Written down because every one of them was silent: the deployment reported success, the pods reported healthy, and the fault surfaced only when a person tried to use the thing.

Period 23 September – 7 October 2026 Applications perform, contracts, crm, helpdesk, runbooks Status all fixes applied and verified

1PgBouncer rejects connection parameters

An application that cannot run a single database query, presenting as a generic error on sign-in.

prod-db.dcbar.org resolves to 10.20.100.238, which is the CloudNativePG pooler app-db-pooler-rw in the database namespace — PgBouncer, not Postgres. It runs in transaction pooling mode, which is the correct mode for these workloads and also the reason it is strict about what a client may send when it connects.

PgBouncer refuses any startup parameter that is not on its ignore_startup_parameters allowlist, and it refuses the whole connection rather than the parameter. The node-postgres client sends three of them when the pool is configured with timeouts:

error: unsupported startup parameter: statement_timeout
severity: FATAL   code: 08P01

The application caught that, discarded it, and returned {"error":{"code":"INTERNAL"}} to the browser. Nothing was logged above info until log level was raised.

Why it took nine rounds to find

Every pod reported 1/1 Running with zero restarts for sixteen hours, because the readiness probe hits /healthz, which does not touch the database. Argo CD reported Synced and Healthy throughout. A health check that does not exercise the application's dependencies will report health it has not verified.

The fix, in two parts

First, so that the parameters stop being fatal — note that all four values must be listed, because a merge patch replaces the string and extra_float_digits is CloudNativePG's own default:

kubectl -n database patch pooler app-db-pooler-rw --type merge -p \
  '{"spec":{"pgbouncer":{"parameters":{
     "default_pool_size":"25",
     "max_client_conn":"500",
     "ignore_startup_parameters":"extra_float_digits,statement_timeout,idle_in_transaction_session_timeout,lock_timeout"
   }}}}'

kubectl -n database rollout status deploy/app-db-pooler-rw

Be clear about what that does: PgBouncer now discards those three settings. The application believes it has timeouts and does not. That is an acceptable trade to restore service, and it is not a finished fix.

Second, so the timeouts actually exist — set them on the role, where pooling cannot interfere. Run this on the primary; a replica answers cannot execute ALTER ROLE in a read-only transaction:

# which instance is primary
kubectl -n database get cluster app-db -o jsonpath='{.status.currentPrimary}{"\n"}'

kubectl -n database exec app-db-2 -- psql -d perform-production \
  -c "ALTER ROLE \"perform-production_user\" SET statement_timeout = '120s'" \
  -c "ALTER ROLE \"perform-production_user\" SET idle_in_transaction_session_timeout = '60s'" \
  -c "ALTER ROLE \"perform-production_user\" SET lock_timeout = '10s'"

These apply to new sessions only, and PgBouncer holds server connections open, so kubectl -n database rollout restart deploy/app-db-pooler-rw if you want them live immediately. Verify with:

kubectl -n database exec app-db-2 -- psql -d perform-production \
  -c "select rolname, rolconfig from pg_roles where rolname = 'perform-production_user'"
Choose 120s deliberately

perform's AI extraction calls run to roughly ninety seconds. A 30-second statement timeout would trade this outage for a subtler one. The same number is why Traefik's responding timeout on websecure needs raising from its 60-second default.

Reproducing it without a browser

This runs inside a pod using its own mounted credentials and sets the parameters exactly as the application does, so it reproduces the failure or proves the fix:

kubectl -n perform-production exec deploy/perform-api -- node -e "
const {Client} = require('pg');
const fs = require('fs');
const c = new Client({
  connectionString: fs.readFileSync(process.env.DATABASE_URL_FILE,'utf8').trim(),
  statement_timeout: 30000,
  idle_in_transaction_session_timeout: 60000,
  lock_timeout: 10000
});
c.connect()
 .then(() => c.query('select current_user, current_database()'))
 .then(r => { console.log('OK', r.rows); return c.end(); })
 .catch(e => console.log('ERR', e.message));
"
For the application teams

An application behind PgBouncer in transaction mode must not set statement_timeout, idle_in_transaction_session_timeout or lock_timeout in its client configuration. Set them with ALTER ROLE instead. This affects every application the generator produces, not just the one where it was found.


2SecretProviderClass as a sync hook

A whole namespace unable to start any pod, with an error that names a missing object rather than the thing that removed it.

MountVolume.SetUp failed for volume "kv-secret":
  SecretProviderClass "dcbar-perform-prod-kv" not found

Pods sit in ContainerCreating indefinitely. Already-running pods are unaffected, because they mounted while the object still existed — so the application keeps serving and the fault is invisible until something triggers a restart. Then it cannot come back.

The mechanism

Every SecretProviderClass in this estate carries Argo CD hook annotations:

annotations:
  argocd.argoproj.io/hook: PreSync
  argocd.argoproj.io/sync-wave: "-2"
finalizers:
  - argocd.argoproj.io/hook-finalizer

Argo's default hook delete policy is BeforeHookCreation: on every sync it deletes the existing hook resource and then creates the one from Git. During that window no pod in the namespace can mount its secrets. If the sync fails partway, or if the resource is not actually in Git, it never comes back.

Hooks are invisible to the obvious check

Hook resources do not appear in an Application's status.resources. Listing an app's resources and seeing no SecretProviderClass does not mean it is absent from Git. That misreading cost time twice, and the second time led to a correct manifest being overwritten by a hand-written guess.

# this does NOT list hooks — do not conclude from it
kubectl -n argocd get app perform-production \
  -o jsonpath='{range .status.resources[*]}{.kind}/{.name}{"\n"}{end}'

Recovering a lost one

If it is in Git, a sync brings it back. Without the argocd CLI, write an operation onto the Application — which is what the CLI does:

kubectl -n argocd patch app <app> --type merge -p \
  '{"operation":{"initiatedBy":{"username":"itadmin"},
    "sync":{"revision":"HEAD","syncStrategy":{"hook":{}}}}}'

A sync with nothing else to do reports successfully synced (no more tasks) and may not run hooks at all. If the object does not reappear, recreate it by hand. Everything needed is recoverable: the tenant is constant, the vault follows a naming pattern, and a running pod's mount directory lists the object names.

# object names, from any pod that still has the mount
kubectl -n <ns> exec deploy/<app> -- ls -la /mnt/secrets

# the naming pattern across the estate
kubectl get secretproviderclass -A \
  -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,VAULT:.spec.parameters.keyvaultName'

Current inventory

NamespaceSecretProviderClassKey VaultHook annotated
perform-productiondcbar-perform-prod-kvdcbar-perform-prod-kvNo — recreated without them
contracts-productioncontracts-secretsdcbar-contracts-prod-kvYes
crm-productioncrm-secretsdcbar-crm-prod-kvYes
nom-portal-productionnom-portal-kvdcbar-nomportal-prod-kvYes
helpdesk-productionhelpdesk-kvdcbar-helpdesk-prod-kvYes

Tenant for all of them: 89df4ea7-2a11-4867-8a15-e2e6102deb04. Note nomportal has no hyphen while the namespace does — the convention is not perfectly mechanical, so confirm the vault name in Azure rather than deriving it.

Four namespaces are still one sync away from this

A SecretProviderClass should be an ordinary tracked resource in Git, not a PreSync hook. Until the four rows above are changed, syncing any of those applications briefly removes the object every pod in the namespace depends on. The secrets-store-creds secret is the other half of the same gap: it is hand-copied, so a rebuilt namespace never gets it back.


3Key Vault rotation does not reach a running process

Rotating a credential in Azure changes the file on disk within two minutes and changes nothing about the application using it.

The Secrets Store CSI driver is installed on all six nodes with rotation enabled:

kubectl -n kube-system get ds secrets-store-csi-driver \
  -o jsonpath='{range .spec.template.spec.containers[*]}{.name}{": "}{.args}{"\n"}{end}'

# secrets-store: … --enable-secret-rotation=true --rotation-poll-interval=2m …

So the mounted files under /mnt/secrets do refresh on their own. But a Node process that read its key at startup holds the old value for the life of the process, and anything consumed as an environment variable never changes in a running container at all.

The rule

Any credential change needs a rollout restart of the workloads that use it. That applies to database passwords and the GHCR token as much as to API keys. A fresh pod fetches from Key Vault at mount time, so there is no need to wait out the poll interval.

kubectl -n perform-production rollout restart deploy/perform-ai deploy/perform-api
kubectl -n contracts-production rollout restart deploy/contracts

Verifying without exposing the value

Hashes are safe to record and compare, and they answer the question a rotation actually raises — did this pod get the new value?

kubectl -n perform-production exec deploy/perform-ai -- sh -c \
  'for f in /mnt/secrets/*; do printf "%s  " "$(basename $f)"; md5sum < "$f"; done'

Take the baseline before touching the vault. Two caveats learned the hard way: run it only after rollout status reports the rollout finished, or you may be reading a surviving old pod; and an unchanged hash on a freshly created pod means the vault object you edited is not the one this application reads.

Two findings from the first rotation

perform has no OpenAI key — it uses Anthropic, via ANTHROPIC_API_KEY_FILE. Only contracts holds an openai-api-key. And in perform's vault, entra-client-secret and graph-client-secret have identical hashes: either one app registration deliberately serves both OIDC sign-in and Graph mail, or one was pasted over the other. Rotating one would then silently break the other. Worth settling.


4Client IP and the proxy chain

What an application must trust to record a visitor's address correctly — and why the answer differs between the two clusters.

Production

visitor → Cloudflare edge (L7, terminates TLS) → cloudflared tunnel (2 replicas, namespace "cloudflare") → Traefik (3 replicas, 10.20.100.237) → the app's own nginx pod → the API

There is no external load balancer. 10.20.100.237 is internal and unreachable from the internet; the only inbound path is the tunnel, which is an outbound connection initiated by cloudflared. Traefik trusts and extends the chain arriving from the tunnel:

--entryPoints.web.forwardedHeaders.trustedIPs=10.244.0.0/16

That is the pod network, which contains the cloudflared pods. It is set on web only — websecure trusts nothing, which is correct for traffic arriving directly on the LAN. The applications' own nginx then appends with proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for.

Dev and staging

visitor (LAN / VPN) → ingress-nginx (10.20.100.251) → the app's own nginx pod → the API

A different ingress controller, and no Cloudflare. The controller ConfigMap holds only large-client-header-buffers, so use-forwarded-headers and compute-full-forwarded-for are both unset and default to false — ingress-nginx discards a visitor-supplied header and replaces it with the real connecting address.

ProductionDev / staging
Proxies ahead of the API42
X-Forwarded-For at the API<visitor>, <cloudflared pod>, <traefik pod><visitor>, <ingress-nginx pod>
Express trust proxy32
CF-Connecting-IPPresentNot present
Ingress controllerTraefikingress-nginx
Too high fails silently, in the attacker's favour

If an application trusts more hops than actually exist, a visitor can forge their own apparent address by sending the header themselves. Where this feeds sign-in rate limiting and an audit log, that defeats the limiter and writes attacker-chosen addresses into the trail. Too low fails safe — visitors share a rate-limit bucket. Where uncertain, err low.

Prefer CF-Connecting-IP in production: a single value holding the real address, with nothing to count and nothing to break when a proxy is added. Since dev and staging have no such header, this belongs in configuration rather than in code — a header-name setting with a hop depth as fallback, so environments differ by a ConfigMap value.


5Migrations that report success

An application deployed and healthy for sixteen hours against a completely empty database.

perform-migrate showed Completed. The perform-production database existed, was correctly owned, and contained zero tables. The job had hit the PgBouncer rejection above and exited zero regardless. Nothing in the deploy pipeline noticed, because a completed Job is a successful Job.

Verify the database, not the job

kubectl -n database exec app-db-2 -- psql -d <database> \
  -c "select count(*) as tables from pg_tables where schemaname='public'"

A migration is confirmed by the schema existing, not by an exit code. This belongs in the deployment notes for every new application.

Re-running one

Jobs are immutable, so there is no restart. Read the log before deleting anything — a deleted Job takes the only evidence with it. Then pick a route:

Whether Git owns the Job differs per application

perform's migrate Job is a hook — invisible in status.resources, and it only runs when a sync has real work to do. helpdesk's is an ordinary tracked resource, so deleting it and syncing genuinely does recreate it. Check before assuming; the two behave differently.

A second failure mode, from helpdesk

TypeError: Invalid URL from pg-connection-string is not a connectivity problem — it means the connection string never parsed. In helpdesk's case the pods had no mounted secrets at all, because the SecretProviderClass was missing. The library masks the value as *****REDACTED*****, so the error deliberately tells you nothing about the content. Check the mount before suspecting the password.


6The dev and staging cluster

A second cluster, sharing nothing with production but the domain name. Surveyed for the first time on 29 September.

Dev and staging are namespaces on one cluster, reached at 10.20.100.241 (k8smaster) and 10.20.100.242 (k8sworker). It is not a smaller copy of production — the differences matter when it is used to validate a change before production sees it.

Production (HQ)Dev / staging
Kubernetesv1.36.4v1.30.14
Nodes3 control-plane, 3 workers1 control-plane, 1 worker
IngressTraefik, 3 replicasingress-nginx, 1 healthy pod
Ingress address10.20.100.23710.20.100.251
Cloudflare tunnelYesNone
TLS on app ingressesWildcard, per namespaceWildcard
The ingress controller is a single point of failure

Four ingress-nginx-controller pods exist and only one is Running; the rest are Error or Completed, one of them 119 days old. If that pod dies, every dev and staging hostname goes down together — around twenty applications.

kubectl -n ingress-nginx get pods
kubectl -n ingress-nginx delete pod --field-selector status.phase!=Running
Six minor versions behind

An environment whose purpose is to catch problems before production cannot reliably do so when it runs a different ingress controller and a Kubernetes release six minors older. Neither difference caused a failure this month, but both limit what a passing test in staging actually proves.

Note also that topologySpreadConstraints are effectively inert here. With one worker — and a control-plane node that is tainted unless someone removed it — there is nowhere to spread to. Confirm with:

kubectl get nodes -o custom-columns='NAME:.metadata.name,TAINTS:.spec.taints[*].key'

7One node announces every address

Five hostnames unreachable for far longer than the 55 seconds a worker node was off the network — and two separate reasons why.

On 6 October prod-k8s-worker-03 left the network for roughly 55 seconds. perform, contracts, crm, helpdesk and support all stopped answering and stayed down well past its return. docs and nominations were unaffected throughout.

That split is the first clue. The five that failed share one MetalLB address; the two that survived arrive over the Cloudflare tunnel and never touch it.

HostnameReached via6 October
perform, contracts, crm, helpdesk, supportMetalLB VIP 10.20.100.237Down
docs, nominationsCloudflare tunnelUnaffected

Fault 1 — a speaker bound to an address its host no longer owned

MetalLB in layer-2 mode has exactly one node announcing each address at any moment. Which node is decided by a memberlist gossip mesh on port 7946, and each speaker binds that port to the pod IP it was given at start, taken from status.podIP via METALLB_ML_BIND_ADDR.

In September the nodes were renumbered following a MAC address change on the hypervisor. prod-k8s-worker-01 moved from 10.20.100.225 to 10.20.100.231. Its speaker pod was not restarted, so it went on listening on an address the host no longer had:

ssh itadmin@10.20.100.231 'ip -4 addr show eth0 | grep inet; ss -lntup | grep 7946'

    inet 10.20.100.231/24 brd 10.20.100.255 scope global eth0
udp   UNCONN 0      0      10.20.100.225:7946       0.0.0.0:*
tcp   LISTEN 0      4096   10.20.100.225:7946       0.0.0.0:*

Every other speaker said the same thing from the other side:

memberlist: Conflicting address for prod-k8s-worker-01.
  Mine: 10.20.100.225:7946  Theirs: 10.20.100.231:7946
Failed to join 10.20.100.231:7946: connect: connection refused

The mesh had been running five members of six for eight days. Failure detection and re-election are both slower with a member permanently in the suspect state, which is why a brief node outage took minutes rather than seconds to recover.

Renumbering a node is not finished when the node comes back

Any process that bound to the old address at start keeps that binding until it is restarted, and nothing reports it as wrong — the pod is 1/1 Running throughout. After changing a node's IP, restart the MetalLB speaker on that node and check the socket actually matches the interface.

Fault 2 — ingress concentrated on the node that left

Traefik runs three replicas with podAntiAffinity expressed as preferredDuringSchedulingIgnoredDuringExecution. That is a scheduling hint, not a rule, and it is never re-evaluated after placement. On 6 October the distribution was:

prod-k8s-worker-01 0 replicas prod-k8s-worker-02 1 replica prod-k8s-worker-03 2 replicas ← the node that left

So the node holding the announced VIP was also holding two thirds of ingress capacity. Both failures landed together.

Remediation, 7 October

kubectl -n metallb-system delete pod metallb-speaker-rxq8s
kubectl -n traefik delete pod traefik-847dcd6f89-8xgft

The speaker returned on 10.20.100.231, the mesh converged within 30 seconds, and both VIPs re-announced. The Traefik replacement was scheduled onto the empty worker, giving one replica per worker node. No hostname was interrupted during either step.

Still true after the fix

10.20.100.237 (ingress) and 10.20.100.238 (the database pooler) are in the same address pool and are currently announced by the same node. Losing that node still takes ingress and the database path together. Splitting the pool so the two cannot share an announcer is the change that removes this, and it has not been made.


8Image Updater v1: the resource decides

Annotations that read like configuration and configure nothing — and the opposite failure, where deleting the wrong object stops deployments with no error anywhere.

Argo CD Image Updater v0.x was configured entirely by annotations on each Argo CD Application. v1.x, which is what both clusters run, is not. In v1 an ImageUpdater custom resource selects which Applications the controller processes, and its useAnnotations field decides whether the Application's annotations are read at all.

useAnnotationsWhat is in force
trueThe Application's annotations
falseThe images block inside the resource. The annotations are inert text.
No resource at allNothing. The Application is never considered, and no error is logged.

Two Applications had two updaters each

On the dev/staging cluster, contracts-dev, perform-dev and helpdesk-dev were each matched by two ImageUpdater resources — one with useAnnotations: true and one with its own image rules. Reading the Application's annotations told you nothing about what was actually running. This is one command:

kubectl -n argocd get imageupdater -o json | jq -r \
  '.items[] | . as $i | $i.spec.applicationRefs[] |
   "\(.namePattern)\t\($i.metadata.name)\t\(.useAnnotations)"' | sort

Every Application should appear exactly once, with useAnnotations set to true. A namePattern appearing twice is a conflict; anything set to false is configured somewhere other than where the repository says it is.

Deleting an ImageUpdater stops deployments silently

On 6 October four production updaters — contracts, helpdesk, nom-portal and perform — were deleted on the production cluster during a cleanup intended for the dev/staging cluster. Nothing broke, nothing alerted, and no running workload changed. Production simply stopped picking up new images. The only symptom available is a build that does not deploy, days later.

The platform standard still describes v0.x behaviour and instructs repository owners to ensure no ImageUpdater resource exists. Followed literally on these clusters, that stops image updates for the application concerned. The standard needs correcting, as does its claim that production is one cluster spanning two sites — production is two separate clusters.


9Publishing a hostname through the tunnel

Making support.dcbar.org reachable from the internet, and the two things that make it look broken when it is not.

The HQ tunnel's routes are held in the Cloudflare dashboard, not in a ConfigMap and not in Git. cloudflared picks up changes live and logs each new revision:

kubectl -n cloudflare logs deploy/cloudflared --tail=100 | grep -i 'Updated to new configuration'

Note the namespace is cloudflare, not cloudflared. The working routes all point at the same origin:

http://traefik.traefik.svc.cluster.local:80
Do not point a route at port 443

Traefik serves the *.dcbar.org certificate, which cannot validate against a cluster.local name. The result is a 502 at the edge and this in the cloudflared log: tls: failed to verify certificate: x509: certificate is valid for *.dcbar.org, dcbar.org, not traefik.traefik.svc.cluster.local. Port 80 avoids the handshake entirely.

Port 80 does not cause a redirect loop, despite the redirect-https middleware. Traefik's web entrypoint trusts X-Forwarded-Proto from 10.244.0.0/16, the cloudflared pods are inside it, and the edge sets that header to https. The middleware sees the request as already secure and passes it through.

Split-horizon DNS will tell you the tunnel works when it does not

Internal DNS resolves these hostnames to 10.20.100.237 directly. A curl run from inside the estate never reaches Cloudflare, so it returns 200 whatever state the tunnel is in. Pin the connection to an edge address to test the real path:

curl -sS --resolve support.dcbar.org:443:104.20.45.170 \
  -o /dev/null -w "%{http_code} redirects=%{num_redirects} ip=%{remote_ip}\n" \
  -L https://support.dcbar.org/

ip= showing an internal address means the test was invalid, not that the site is healthy.

One consequence worth carrying into the applications. A hostname published this way now has two different proxy chain lengths depending on where the visitor is: four hops from the internet, two from the LAN. An application that trusts a fixed hop count will mis-attribute one of them. Trusting by CIDR rather than by depth is the durable answer — see section 4.


10Outstanding

Work created or revealed by the above, not yet done. Written here so it is somewhere other than one person's memory.

In the cluster but not in Git

Each of these is one kubectl apply or one sync away from disappearing, and each disappearance is an outage:

Changes owed to the application repositories

Platform

Network and ingress resilience

All four follow from section 7 and none were in scope for the 7 October change:

Deployment pipeline

To tell the application teams

Credentials pending rotation

The GitHub personal access token used for the initial push, the Postgres superuser account (enableSuperuserAccess: false once nothing needs it), and the dbadmin role. All three were exposed in working transcripts during this period and should be treated as known.