Advanced DevOps Interview Questions
194 questions
How would you design scheduled Terraform drift detection so it alerts the right team without becoming noise nobody reads?
AdvancedTerraform8 min
You need to make a breaking change to a shared infrastructure module used by 30 different projects. How do you version and roll this out without breaking everyone simultaneously?
AdvancedTerraform8 min
You've extracted common pipeline logic into a Jenkins shared library used by 40+ Jenkinsfiles. How do you version it so you can improve the library without breaking every pipeline that depends on it?
AdvancedJenkins7 min
How would Kaniko or rootless BuildKit avoid the Docker socket mounting problem entirely, and what do you give up by switching to them?
AdvancedJenkinsDocker7 min
A new mutating webhook accidentally intercepted kube-system pod creation and broke core cluster components — how would you design its scoping to prevent this?
AdvancedKubernetes7 min
An admission webhook's failurePolicy is set to Fail — what happens if the webhook itself becomes unavailable, and why might that be the wrong default?
AdvancedKubernetes7 min
A validating webhook is rejecting pod creations that look completely valid — how do you diagnose what it's actually objecting to?
AdvancedKubernetes7 min
kubectl apply --dry-run=server behaves unexpectedly for requests a webhook handles — what does a webhook's sideEffects field have to do with it?
AdvancedKubernetes6 min
How would you design backup/DR for etcd, and what's actually recoverable from a snapshot versus what isn't?
AdvancedKubernetesEtcd8 min
How would you design multi-cluster architecture — when does an org actually need multiple clusters instead of namespaces?
AdvancedKubernetes8 min
A cluster upgrade needs zero workload downtime — walk through sequencing control-plane and node upgrades safely.
AdvancedKubernetes8 min
An HPA scales up rapidly during a spike, then flaps up and down repeatedly for the next hour — what's causing it, and how do you fix it?
AdvancedKubernetes7 min
Users report increasing latency under load, but the HPA isn't scaling the Deployment at all — how do you figure out why?
AdvancedKubernetes8 min
How would you design autoscaling for a workload with a sharp, predictable daily spike versus one with genuinely unpredictable bursty traffic?
AdvancedKubernetes7 min
A security scan flags the API server's anonymous authentication as enabled — what does that actually expose, and how would you harden it safely?
AdvancedKubernetes7 min
How would you design a Kubernetes audit logging policy that's actually useful for a security investigation, without drowning in log volume?
AdvancedKubernetes7 min
How would you design network-level segmentation between the control plane and worker nodes, beyond what Kubernetes' own RBAC and NetworkPolicy provide?
AdvancedKubernetes7 min
You already enforce preventive admission policies (Kyverno/Gatekeeper) — why would you also need runtime security tooling like Falco?
AdvancedKubernetesFalco6 min
What does a seccomp profile actually add on top of SecurityContext's capability restrictions, and when do you need one?
AdvancedKubernetes6 min
How would you design a workflow so a ConfigMap change automatically triggers a rolling restart of the Deployments that depend on it?
AdvancedKubernetes7 min
How would you manage Secrets across dev/staging/prod without committing plaintext to Git, while staying GitOps-declarative?
AdvancedKubernetes8 min
A Secret manifest with real credentials was committed to a public repo — how does remediation differ from a generic leaked-secret response?
AdvancedKubernetes7 min
A custom resource is stuck in Terminating status indefinitely after being deleted — what's a finalizer, and how does it cause this?
AdvancedKubernetes7 min
Custom resources are being created and updated, but the operator managing them appears to have silently stopped reconciling — how do you diagnose it?
AdvancedKubernetes8 min