Operations notes
Nine failures from September and October 2026, each of which looked like something it was not. Written down because every one of them was silent: the deployment reported success, the pods reported healthy, and the fault surfaced only when a person tried to use the thing.
An application that cannot run a single database query, presenting as a generic error on sign-in.
prod-db.dcbar.org resolves to 10.20.100.238, which
is the CloudNativePG pooler app-db-pooler-rw in the
database namespace — PgBouncer, not Postgres. It runs in
transaction pooling mode, which is the correct mode for these
workloads and also the reason it is strict about what a client may send when
it connects.
PgBouncer refuses any startup parameter that is not on its
ignore_startup_parameters allowlist, and it refuses the whole
connection rather than the parameter. The node-postgres client sends three
of them when the pool is configured with timeouts:
error: unsupported startup parameter: statement_timeout
severity: FATAL code: 08P01
The application caught that, discarded it, and returned
{"error":{"code":"INTERNAL"}} to the browser. Nothing was
logged above info until log level was raised.
Every pod reported 1/1 Running with zero restarts for
sixteen hours, because the readiness probe hits /healthz,
which does not touch the database. Argo CD reported Synced and
Healthy throughout. A health check that does not exercise the
application's dependencies will report health it has not verified.
First, so that the parameters stop being fatal — note that all four values
must be listed, because a merge patch replaces the string and
extra_float_digits is CloudNativePG's own default:
kubectl -n database patch pooler app-db-pooler-rw --type merge -p \
'{"spec":{"pgbouncer":{"parameters":{
"default_pool_size":"25",
"max_client_conn":"500",
"ignore_startup_parameters":"extra_float_digits,statement_timeout,idle_in_transaction_session_timeout,lock_timeout"
}}}}'
kubectl -n database rollout status deploy/app-db-pooler-rw
Be clear about what that does: PgBouncer now discards those three settings. The application believes it has timeouts and does not. That is an acceptable trade to restore service, and it is not a finished fix.
Second, so the timeouts actually exist — set them on the role, where pooling
cannot interfere. Run this on the primary; a replica answers
cannot execute ALTER ROLE in a read-only transaction:
# which instance is primary
kubectl -n database get cluster app-db -o jsonpath='{.status.currentPrimary}{"\n"}'
kubectl -n database exec app-db-2 -- psql -d perform-production \
-c "ALTER ROLE \"perform-production_user\" SET statement_timeout = '120s'" \
-c "ALTER ROLE \"perform-production_user\" SET idle_in_transaction_session_timeout = '60s'" \
-c "ALTER ROLE \"perform-production_user\" SET lock_timeout = '10s'"
These apply to new sessions only, and PgBouncer holds server connections
open, so kubectl -n database rollout restart deploy/app-db-pooler-rw
if you want them live immediately. Verify with:
kubectl -n database exec app-db-2 -- psql -d perform-production \
-c "select rolname, rolconfig from pg_roles where rolname = 'perform-production_user'"
perform's AI extraction calls run to roughly ninety seconds. A 30-second
statement timeout would trade this outage for a subtler one. The same
number is why Traefik's responding timeout on websecure needs
raising from its 60-second default.
This runs inside a pod using its own mounted credentials and sets the parameters exactly as the application does, so it reproduces the failure or proves the fix:
kubectl -n perform-production exec deploy/perform-api -- node -e "
const {Client} = require('pg');
const fs = require('fs');
const c = new Client({
connectionString: fs.readFileSync(process.env.DATABASE_URL_FILE,'utf8').trim(),
statement_timeout: 30000,
idle_in_transaction_session_timeout: 60000,
lock_timeout: 10000
});
c.connect()
.then(() => c.query('select current_user, current_database()'))
.then(r => { console.log('OK', r.rows); return c.end(); })
.catch(e => console.log('ERR', e.message));
"
An application behind PgBouncer in transaction mode must not set
statement_timeout, idle_in_transaction_session_timeout
or lock_timeout in its client configuration. Set them with
ALTER ROLE instead. This affects every application the
generator produces, not just the one where it was found.
A whole namespace unable to start any pod, with an error that names a missing object rather than the thing that removed it.
MountVolume.SetUp failed for volume "kv-secret":
SecretProviderClass "dcbar-perform-prod-kv" not found
Pods sit in ContainerCreating indefinitely. Already-running
pods are unaffected, because they mounted while the object still existed —
so the application keeps serving and the fault is invisible until something
triggers a restart. Then it cannot come back.
Every SecretProviderClass in this estate carries Argo CD hook annotations:
annotations:
argocd.argoproj.io/hook: PreSync
argocd.argoproj.io/sync-wave: "-2"
finalizers:
- argocd.argoproj.io/hook-finalizer
Argo's default hook delete policy is BeforeHookCreation: on
every sync it deletes the existing hook resource and then creates the
one from Git. During that window no pod in the namespace can mount its
secrets. If the sync fails partway, or if the resource is not actually in
Git, it never comes back.
Hook resources do not appear in an Application's
status.resources. Listing an app's resources and seeing no
SecretProviderClass does not mean it is absent from Git. That misreading
cost time twice, and the second time led to a correct manifest being
overwritten by a hand-written guess.
# this does NOT list hooks — do not conclude from it
kubectl -n argocd get app perform-production \
-o jsonpath='{range .status.resources[*]}{.kind}/{.name}{"\n"}{end}'
If it is in Git, a sync brings it back. Without the argocd CLI,
write an operation onto the Application — which is what the CLI does:
kubectl -n argocd patch app <app> --type merge -p \
'{"operation":{"initiatedBy":{"username":"itadmin"},
"sync":{"revision":"HEAD","syncStrategy":{"hook":{}}}}}'
A sync with nothing else to do reports successfully synced (no more
tasks) and may not run hooks at all. If the object does not reappear,
recreate it by hand. Everything needed is recoverable: the tenant is
constant, the vault follows a naming pattern, and a running pod's mount
directory lists the object names.
# object names, from any pod that still has the mount
kubectl -n <ns> exec deploy/<app> -- ls -la /mnt/secrets
# the naming pattern across the estate
kubectl get secretproviderclass -A \
-o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,VAULT:.spec.parameters.keyvaultName'
| Namespace | SecretProviderClass | Key Vault | Hook annotated |
|---|---|---|---|
| perform-production | dcbar-perform-prod-kv | dcbar-perform-prod-kv | No — recreated without them |
| contracts-production | contracts-secrets | dcbar-contracts-prod-kv | Yes |
| crm-production | crm-secrets | dcbar-crm-prod-kv | Yes |
| nom-portal-production | nom-portal-kv | dcbar-nomportal-prod-kv | Yes |
| helpdesk-production | helpdesk-kv | dcbar-helpdesk-prod-kv | Yes |
Tenant for all of them: 89df4ea7-2a11-4867-8a15-e2e6102deb04.
Note nomportal has no hyphen while the namespace does — the
convention is not perfectly mechanical, so confirm the vault name in Azure
rather than deriving it.
A SecretProviderClass should be an ordinary tracked resource in Git, not a
PreSync hook. Until the four rows above are changed, syncing any of those
applications briefly removes the object every pod in the namespace depends
on. The secrets-store-creds secret is the other half of the
same gap: it is hand-copied, so a rebuilt namespace never gets it back.
Rotating a credential in Azure changes the file on disk within two minutes and changes nothing about the application using it.
The Secrets Store CSI driver is installed on all six nodes with rotation enabled:
kubectl -n kube-system get ds secrets-store-csi-driver \
-o jsonpath='{range .spec.template.spec.containers[*]}{.name}{": "}{.args}{"\n"}{end}'
# secrets-store: … --enable-secret-rotation=true --rotation-poll-interval=2m …
So the mounted files under /mnt/secrets do refresh on their own.
But a Node process that read its key at startup holds the old value for the
life of the process, and anything consumed as an environment variable never
changes in a running container at all.
Any credential change needs a rollout restart of the
workloads that use it. That applies to database passwords and the GHCR
token as much as to API keys. A fresh pod fetches from Key Vault at mount
time, so there is no need to wait out the poll interval.
kubectl -n perform-production rollout restart deploy/perform-ai deploy/perform-api
kubectl -n contracts-production rollout restart deploy/contracts
Hashes are safe to record and compare, and they answer the question a rotation actually raises — did this pod get the new value?
kubectl -n perform-production exec deploy/perform-ai -- sh -c \
'for f in /mnt/secrets/*; do printf "%s " "$(basename $f)"; md5sum < "$f"; done'
Take the baseline before touching the vault. Two caveats learned the hard
way: run it only after rollout status reports the rollout
finished, or you may be reading a surviving old pod; and an unchanged hash
on a freshly created pod means the vault object you edited is not the one
this application reads.
perform has no OpenAI key — it uses Anthropic, via
ANTHROPIC_API_KEY_FILE. Only contracts holds an
openai-api-key. And in perform's vault,
entra-client-secret and graph-client-secret have
identical hashes: either one app registration deliberately serves
both OIDC sign-in and Graph mail, or one was pasted over the other.
Rotating one would then silently break the other. Worth settling.
What an application must trust to record a visitor's address correctly — and why the answer differs between the two clusters.
There is no external load balancer. 10.20.100.237 is internal
and unreachable from the internet; the only inbound path is the tunnel,
which is an outbound connection initiated by cloudflared.
Traefik trusts and extends the chain arriving from the tunnel:
--entryPoints.web.forwardedHeaders.trustedIPs=10.244.0.0/16
That is the pod network, which contains the cloudflared pods.
It is set on web only — websecure trusts nothing,
which is correct for traffic arriving directly on the LAN. The applications'
own nginx then appends with
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for.
A different ingress controller, and no Cloudflare. The controller ConfigMap
holds only large-client-header-buffers, so
use-forwarded-headers and compute-full-forwarded-for
are both unset and default to false — ingress-nginx discards a
visitor-supplied header and replaces it with the real connecting address.
| Production | Dev / staging | |
|---|---|---|
| Proxies ahead of the API | 4 | 2 |
X-Forwarded-For at the API | <visitor>, <cloudflared pod>, <traefik pod> | <visitor>, <ingress-nginx pod> |
Express trust proxy | 3 | 2 |
CF-Connecting-IP | Present | Not present |
| Ingress controller | Traefik | ingress-nginx |
If an application trusts more hops than actually exist, a visitor can forge their own apparent address by sending the header themselves. Where this feeds sign-in rate limiting and an audit log, that defeats the limiter and writes attacker-chosen addresses into the trail. Too low fails safe — visitors share a rate-limit bucket. Where uncertain, err low.
Prefer CF-Connecting-IP in production: a single value holding
the real address, with nothing to count and nothing to break when a proxy is
added. Since dev and staging have no such header, this belongs in
configuration rather than in code — a header-name setting with a hop depth
as fallback, so environments differ by a ConfigMap value.
An application deployed and healthy for sixteen hours against a completely empty database.
perform-migrate showed Completed. The
perform-production database existed, was correctly owned, and
contained zero tables. The job had hit the PgBouncer rejection above
and exited zero regardless. Nothing in the deploy pipeline noticed, because
a completed Job is a successful Job.
kubectl -n database exec app-db-2 -- psql -d <database> \
-c "select count(*) as tables from pg_tables where schemaname='public'"
A migration is confirmed by the schema existing, not by an exit code. This belongs in the deployment notes for every new application.
Jobs are immutable, so there is no restart. Read the log before deleting anything — a deleted Job takes the only evidence with it. Then pick a route:
node dist/db/migrate.js for perform,
node packages/api/dist/db/migrate.js for helpdesk. Output is
live and there is no object to clean up.
jq, stripping
selector, uid, resourceVersion,
ownerReferences and status. Avoids Argo
entirely.
perform's migrate Job is a hook — invisible in status.resources,
and it only runs when a sync has real work to do. helpdesk's is an ordinary
tracked resource, so deleting it and syncing genuinely does recreate it.
Check before assuming; the two behave differently.
TypeError: Invalid URL from pg-connection-string
is not a connectivity problem — it means the connection string never
parsed. In helpdesk's case the pods had no mounted secrets at all, because
the SecretProviderClass was missing. The library masks the value as
*****REDACTED*****, so the error deliberately tells you
nothing about the content. Check the mount before suspecting the password.
A second cluster, sharing nothing with production but the domain name. Surveyed for the first time on 29 September.
Dev and staging are namespaces on one cluster, reached at
10.20.100.241 (k8smaster) and
10.20.100.242 (k8sworker). It is not a smaller
copy of production — the differences matter when it is used to validate a
change before production sees it.
| Production (HQ) | Dev / staging | |
|---|---|---|
| Kubernetes | v1.36.4 | v1.30.14 |
| Nodes | 3 control-plane, 3 workers | 1 control-plane, 1 worker |
| Ingress | Traefik, 3 replicas | ingress-nginx, 1 healthy pod |
| Ingress address | 10.20.100.237 | 10.20.100.251 |
| Cloudflare tunnel | Yes | None |
| TLS on app ingresses | Wildcard, per namespace | Wildcard |
Four ingress-nginx-controller pods exist and only one is
Running; the rest are Error or Completed, one of
them 119 days old. If that pod dies, every dev and staging hostname goes
down together — around twenty applications.
kubectl -n ingress-nginx get pods
kubectl -n ingress-nginx delete pod --field-selector status.phase!=Running
An environment whose purpose is to catch problems before production cannot reliably do so when it runs a different ingress controller and a Kubernetes release six minors older. Neither difference caused a failure this month, but both limit what a passing test in staging actually proves.
Note also that topologySpreadConstraints are effectively inert
here. With one worker — and a control-plane node that is tainted unless
someone removed it — there is nowhere to spread to. Confirm with:
kubectl get nodes -o custom-columns='NAME:.metadata.name,TAINTS:.spec.taints[*].key'
Five hostnames unreachable for far longer than the 55 seconds a worker node was off the network — and two separate reasons why.
On 6 October prod-k8s-worker-03 left the network for roughly
55 seconds. perform, contracts, crm,
helpdesk and support all stopped answering and
stayed down well past its return. docs and
nominations were unaffected throughout.
That split is the first clue. The five that failed share one MetalLB address; the two that survived arrive over the Cloudflare tunnel and never touch it.
| Hostname | Reached via | 6 October |
|---|---|---|
perform, contracts, crm, helpdesk, support | MetalLB VIP 10.20.100.237 | Down |
docs, nominations | Cloudflare tunnel | Unaffected |
MetalLB in layer-2 mode has exactly one node announcing each address at any
moment. Which node is decided by a memberlist gossip mesh on port 7946, and
each speaker binds that port to the pod IP it was given at start, taken from
status.podIP via METALLB_ML_BIND_ADDR.
In September the nodes were renumbered following a MAC address change on the
hypervisor. prod-k8s-worker-01 moved from
10.20.100.225 to 10.20.100.231. Its speaker pod was
not restarted, so it went on listening on an address the host no longer had:
ssh itadmin@10.20.100.231 'ip -4 addr show eth0 | grep inet; ss -lntup | grep 7946'
inet 10.20.100.231/24 brd 10.20.100.255 scope global eth0
udp UNCONN 0 0 10.20.100.225:7946 0.0.0.0:*
tcp LISTEN 0 4096 10.20.100.225:7946 0.0.0.0:*
Every other speaker said the same thing from the other side:
memberlist: Conflicting address for prod-k8s-worker-01.
Mine: 10.20.100.225:7946 Theirs: 10.20.100.231:7946
Failed to join 10.20.100.231:7946: connect: connection refused
The mesh had been running five members of six for eight days. Failure detection and re-election are both slower with a member permanently in the suspect state, which is why a brief node outage took minutes rather than seconds to recover.
Any process that bound to the old address at start keeps that binding
until it is restarted, and nothing reports it as wrong — the pod is
1/1 Running throughout. After changing a node's IP, restart
the MetalLB speaker on that node and check the socket actually matches the
interface.
Traefik runs three replicas with
podAntiAffinity expressed as
preferredDuringSchedulingIgnoredDuringExecution. That is a
scheduling hint, not a rule, and it is never re-evaluated after placement.
On 6 October the distribution was:
So the node holding the announced VIP was also holding two thirds of ingress capacity. Both failures landed together.
kubectl -n metallb-system delete pod metallb-speaker-rxq8s
kubectl -n traefik delete pod traefik-847dcd6f89-8xgft
The speaker returned on 10.20.100.231, the mesh converged
within 30 seconds, and both VIPs re-announced. The Traefik replacement was
scheduled onto the empty worker, giving one replica per worker node. No
hostname was interrupted during either step.
10.20.100.237 (ingress) and 10.20.100.238
(the database pooler) are in the same address pool and are currently
announced by the same node. Losing that node still takes ingress and the
database path together. Splitting the pool so the two cannot share an
announcer is the change that removes this, and it has not been made.
Annotations that read like configuration and configure nothing — and the opposite failure, where deleting the wrong object stops deployments with no error anywhere.
Argo CD Image Updater v0.x was configured entirely by annotations on each
Argo CD Application. v1.x, which is what both clusters run, is not. In v1 an
ImageUpdater custom resource selects which Applications the
controller processes, and its useAnnotations field decides
whether the Application's annotations are read at all.
useAnnotations | What is in force |
|---|---|
true | The Application's annotations |
false | The images block inside the resource. The annotations are inert text. |
| No resource at all | Nothing. The Application is never considered, and no error is logged. |
On the dev/staging cluster, contracts-dev,
perform-dev and helpdesk-dev were each matched by
two ImageUpdater resources — one with
useAnnotations: true and one with its own image rules. Reading
the Application's annotations told you nothing about what was actually
running. This is one command:
kubectl -n argocd get imageupdater -o json | jq -r \
'.items[] | . as $i | $i.spec.applicationRefs[] |
"\(.namePattern)\t\($i.metadata.name)\t\(.useAnnotations)"' | sort
Every Application should appear exactly once, with useAnnotations
set to true. A namePattern appearing twice is a
conflict; anything set to false is configured somewhere other
than where the repository says it is.
On 6 October four production updaters — contracts,
helpdesk, nom-portal and perform —
were deleted on the production cluster during a cleanup intended for the
dev/staging cluster. Nothing broke, nothing alerted, and no running
workload changed. Production simply stopped picking up new images. The
only symptom available is a build that does not deploy, days later.
The platform standard still describes v0.x behaviour and instructs repository
owners to ensure no ImageUpdater resource exists. Followed
literally on these clusters, that stops image updates for the application
concerned. The standard needs correcting, as does its claim that production
is one cluster spanning two sites — production is two separate clusters.
Making support.dcbar.org reachable from the internet, and the
two things that make it look broken when it is not.
The HQ tunnel's routes are held in the Cloudflare dashboard, not in a
ConfigMap and not in Git. cloudflared picks up changes live and
logs each new revision:
kubectl -n cloudflare logs deploy/cloudflared --tail=100 | grep -i 'Updated to new configuration'
Note the namespace is cloudflare, not cloudflared.
The working routes all point at the same origin:
http://traefik.traefik.svc.cluster.local:80
Traefik serves the *.dcbar.org certificate, which cannot
validate against a cluster.local name. The result is a 502 at
the edge and this in the cloudflared log:
tls: failed to verify certificate: x509: certificate is valid for
*.dcbar.org, dcbar.org, not traefik.traefik.svc.cluster.local.
Port 80 avoids the handshake entirely.
Port 80 does not cause a redirect loop, despite the
redirect-https middleware. Traefik's web entrypoint
trusts X-Forwarded-Proto from 10.244.0.0/16, the
cloudflared pods are inside it, and the edge sets that header to
https. The middleware sees the request as already secure and
passes it through.
Internal DNS resolves these hostnames to 10.20.100.237
directly. A curl run from inside the estate never reaches
Cloudflare, so it returns 200 whatever state the tunnel is in.
Pin the connection to an edge address to test the real path:
curl -sS --resolve support.dcbar.org:443:104.20.45.170 \
-o /dev/null -w "%{http_code} redirects=%{num_redirects} ip=%{remote_ip}\n" \
-L https://support.dcbar.org/
ip= showing an internal address means the test was invalid,
not that the site is healthy.
One consequence worth carrying into the applications. A hostname published this way now has two different proxy chain lengths depending on where the visitor is: four hops from the internet, two from the LAN. An application that trusts a fixed hop count will mis-attribute one of them. Trusting by CIDR rather than by depth is the durable answer — see section 4.
Work created or revealed by the above, not yet done. Written here so it is somewhere other than one person's memory.
Each of these is one kubectl apply or one sync away from
disappearing, and each disappearance is an outage:
ignore_startup_parameters — applied live to an
object whose source manifest lives in a file somewhere.perform-production's redirect-https Middleware.dcbar-perform-prod-kv SecretProviderClass.router.entrypoints should be web,websecure, not
websecure alone — an HTTPS-only router means HTTP 404s and any
attached redirect middleware can never fire.tls: block must not reference a secret that nothing
creates.:production-newest-build are an Image
Updater strategy name, not a version. That is a mutable tag by
another name.websecure responding timeout to 120s.docs.dcbar.org.
A 200 from an unauthenticated context means this site — which
names every internal address in the estate — is public.All four follow from section 7 and none were in scope for the 7 October change:
10.20.100.237 and
10.20.100.238 cannot be announced by the same node. This is
the one that still represents a live single point of failure.--node-ip on all six nodes, so a node cannot
report a floating address as its own after a restart.podAntiAffinity preference with a
topology spread constraint, so even distribution is maintained rather
than restored by hand.metallb-frr-k8s
components are already deployed on all six nodes and unused. BGP removes
the single-announcer design entirely, and needs a peering session from the
network team.ImageUpdater resources
deleted on 6 October — contracts, helpdesk,
nom-portal and perform. Until then production
does not pick up new images, and says nothing about it.imageName values:
several point at git URLs or at a registry namespace with no image, and
one uses a capitalised path that no registry will resolve.support.dcbar.org is public as of 6 October. Its proxy
chain is now four hops from the internet and two from the LAN, which
affects sign-in rate limiting and anything that records a visitor
address.bull:email.outbound and
bull:reports.build from a site-local Redis. No repeatable job
schedulers are registered, so the only active/active duplication risk is
the queues themselves.
The GitHub personal access token used for the initial push, the Postgres
superuser account (enableSuperuserAccess: false once nothing
needs it), and the dbadmin role. All three were exposed in
working transcripts during this period and should be treated as known.