>DevOps Interview KB

Troubleshooting Interview Questions — All Categories

141 questions across 32 categories

Real production incidents, spanning Kubernetes, AWS, Docker, Terraform, CI/CD, networking, and databases — diagnosing a specific symptom back to its actual cause, not textbook definitions.

Why might a pod with a very specific nodeSelector never get scheduled, even though matching nodes exist with capacity?

IntermediateKubernetes6 min

A pod stays Pending with node(s) had untolerated taint — how do you diagnose it and decide toleration vs. removing the taint?

IntermediateKubernetes6 min

How would you troubleshoot a pod stuck Pending even though kubectl describe pod shows no scheduling errors at all?

ExpertKubernetes8 min

A CSI driver upgrade causes new attach operations to fail while already-mounted volumes keep working — how do you investigate, and how would you roll this out more safely next time?

ExpertKubernetes8 min

A pod is stuck Pending with an event about its PVC failing to bind — how do you diagnose why?

IntermediateKubernetes7 min

Why might mounting the same ReadWriteOnce PVC work for two pods on some clusters but fail on others?

AdvancedKubernetes6 min

A StatefulSet pod is rescheduled to a new node but its volume won't attach — what's happening, and how do you fix it?

AdvancedKubernetes8 min

A pod goes into CrashLoopBackOff immediately after you roll out a ConfigMap change, but only in one namespace. How do you investigate it?

IntermediateKubernetesContainers10 min

A CronJob has been silently creating thousands of failed Jobs over several days — how did this happen, and how would you prevent it?

AdvancedKubernetes7 min

A DaemonSet pod is missing from exactly one node while running fine everywhere else — how do you find out why?

IntermediateKubernetes6 min

A Job is supposed to run to completion exactly once, but it created multiple pods — why, and is that actually a bug?

IntermediateKubernetes6 min

A Deployment's new pods keep getting scheduled but immediately evicted, while old pods keep running fine — what changed?

AdvancedKubernetes7 min

A rolling update to a Deployment is stuck at 50% — how do you determine whether it's a bad readiness probe, insufficient capacity, or a PodDisruptionBudget blocking it?

AdvancedKubernetes8 min

A StatefulSet pod is deleted but isn't recreated with the same identity fast enough — what's actually blocking it?

AdvancedKubernetes7 min

df says a server's disk is 100% full, but running du -sh on every top-level directory only adds up to a fraction of that. Where did the rest of the space go?

IntermediateLinux8 min

How would you find the same df/du phantom-disk-usage issue on a containerized workload, where the culprit process might be in a different mount namespace?

AdvancedLinuxDocker8 min

A containerized process shows only 40% average CPU usage, well under its configured limit, but application metrics show frequent latency spikes correlating with CPU throttling events. How is this possible?

AdvancedLinuxKubernetes8 min

A service takes 25 seconds to shut down gracefully, but your orchestrator sends SIGKILL after a 10-second grace period, causing corrupted in-progress writes on every deploy. How do you actually fix this?

IntermediateLinux7 min

A service running fine for months suddenly throws 'Too many open files' errors under increased load. How do you fix it correctly, rather than just raising the limit blindly?

IntermediateLinux7 min

A team added a 'user_id' label to a metric to debug a specific customer's issue, and now your Prometheus instance is running out of memory and queries are timing out. What happened?

AdvancedPrometheus8 min

How would you differentiate a client-side DNS resolution problem from the authoritative or upstream DNS server itself being unreliable?

IntermediateDNS7 min

An application intermittently fails to resolve a hostname — most requests succeed, but a small percentage fail with a DNS error. How would you track this down?

IntermediateDNSLinux8 min

During every deployment, a handful of in-flight requests get dropped with connection-reset errors right as old servers are terminated. How do you fix this?

IntermediateNetworking6 min

Your load balancer keeps marking healthy backend servers as unhealthy and removing them from rotation, causing capacity to drop and requests to concentrate on fewer servers. How do you diagnose and fix this?

AdvancedNetworking8 min