Troubleshooting Interview Questions — All Categories
141 questions across 32 categories
Real production incidents, spanning Kubernetes, AWS, Docker, Terraform, CI/CD, networking, and databases — diagnosing a specific symptom back to its actual cause, not textbook definitions.
A ServiceAccount that worked fine before a deployment suddenly gets Forbidden errors on the Kubernetes API — how do you diagnose it?
IntermediateKubernetes7 min
A Deployment runs fine in the staging namespace but fails with RBAC errors in production — what's actually different?
IntermediateKubernetes6 min
How would you distinguish a genuine memory leak from a legitimately growing in-memory cache, using only Kubernetes-level metrics?
AdvancedKubernetesPrometheus8 min
Your container runs fine locally but repeatedly gets OOMKilled after you deploy it to Kubernetes. How would you investigate it?
IntermediateKubernetesContainers10 min
Two pods that should never co-locate keep landing on the same node — what's wrong with the anti-affinity rule?
AdvancedKubernetes7 min
A critical pod gets preempted by a seemingly lower-priority pod during a resource crunch — how do you investigate and prevent it?
AdvancedKubernetes7 min
Why might a pod with a very specific nodeSelector never get scheduled, even though matching nodes exist with capacity?
IntermediateKubernetes6 min
A pod stays Pending with node(s) had untolerated taint — how do you diagnose it and decide toleration vs. removing the taint?
IntermediateKubernetes6 min
How would you troubleshoot a pod stuck Pending even though kubectl describe pod shows no scheduling errors at all?
ExpertKubernetes8 min
A CSI driver upgrade causes new attach operations to fail while already-mounted volumes keep working — how do you investigate, and how would you roll this out more safely next time?
ExpertKubernetes8 min
A pod is stuck Pending with an event about its PVC failing to bind — how do you diagnose why?
IntermediateKubernetes7 min
Why might mounting the same ReadWriteOnce PVC work for two pods on some clusters but fail on others?
AdvancedKubernetes6 min
A StatefulSet pod is rescheduled to a new node but its volume won't attach — what's happening, and how do you fix it?
AdvancedKubernetes8 min
A pod goes into CrashLoopBackOff immediately after you roll out a ConfigMap change, but only in one namespace. How do you investigate it?
IntermediateKubernetesContainers10 min
A CronJob has been silently creating thousands of failed Jobs over several days — how did this happen, and how would you prevent it?
AdvancedKubernetes7 min
A DaemonSet pod is missing from exactly one node while running fine everywhere else — how do you find out why?
IntermediateKubernetes6 min
A Job is supposed to run to completion exactly once, but it created multiple pods — why, and is that actually a bug?
IntermediateKubernetes6 min
A Deployment's new pods keep getting scheduled but immediately evicted, while old pods keep running fine — what changed?
AdvancedKubernetes7 min
A rolling update to a Deployment is stuck at 50% — how do you determine whether it's a bad readiness probe, insufficient capacity, or a PodDisruptionBudget blocking it?
AdvancedKubernetes8 min
A StatefulSet pod is deleted but isn't recreated with the same identity fast enough — what's actually blocking it?
AdvancedKubernetes7 min