Senior DevOps Interview Questions
516 questions across 39 categories
Deeper trade-off and design questions — why a specific approach was chosen, what breaks at scale, and how to reason about a system you didn't build.
A critical pod gets preempted by a seemingly lower-priority pod during a resource crunch — how do you investigate and prevent it?
AdvancedKubernetes7 min
How would you dedicate a set of nodes exclusively to one team's workloads, while still letting that team's pods run elsewhere too?
AdvancedKubernetes7 min
How does the Kubernetes scheduler actually decide which node to place a pod on — walk through filtering and scoring?
IntermediateKubernetes6 min
Why might a pod with a very specific nodeSelector never get scheduled, even though matching nodes exist with capacity?
IntermediateKubernetes6 min
A pod stays Pending with node(s) had untolerated taint — how do you diagnose it and decide toleration vs. removing the taint?
IntermediateKubernetes6 min
What's the difference between requiredDuringSchedulingIgnoredDuringExecution and preferredDuringSchedulingIgnoredDuringExecution, and what incident does confusing them cause?
IntermediateKubernetes5 min
How would you design pod topology spread constraints to keep a Deployment's replicas evenly distributed across availability zones?
AdvancedKubernetes7 min
How would you back up and restore persistent volume data for a stateful app, given kubectl alone doesn't capture volume contents?
IntermediateKubernetes7 min
How would you safely expand a PVC for a running StatefulSet without downtime, and what does the StorageClass need to support?
AdvancedKubernetes7 min
A pod is stuck Pending with an event about its PVC failing to bind — how do you diagnose why?
IntermediateKubernetes7 min
Why might mounting the same ReadWriteOnce PVC work for two pods on some clusters but fail on others?
AdvancedKubernetes6 min
A team deletes a PVC expecting the data gone, but it's later recovered from the underlying disk — why, and how should reclaim policy be chosen deliberately?
IntermediateKubernetes6 min
A StatefulSet pod is rescheduled to a new node but its volume won't attach — what's happening, and how do you fix it?
AdvancedKubernetes8 min
How would you design a StorageClass for a high-IOPS database workload versus a cheap-capacity logging workload?
IntermediateKubernetes6 min
What does volumeBindingMode: WaitForFirstConsumer actually solve, and what breaks in a multi-zone cluster if you don't set it?
AdvancedKubernetes6 min
How would you set up alerting to catch a CrashLoopBackOff-class issue before it reaches production traffic, rather than discovering it via a user-facing outage?
IntermediateKubernetesPrometheus7 min
A pod goes into CrashLoopBackOff immediately after you roll out a ConfigMap change, but only in one namespace. How do you investigate it?
IntermediateKubernetesContainers10 min
How do liveness and readiness probes interact with a Pod that's already crash-looping on startup?
IntermediateKubernetes6 min
How would you decide between a Deployment, a StatefulSet, and a DaemonSet for three different real services (a stateless API, a database, a node agent)?
IntermediateKubernetes6 min
What's the difference between a CronJob's concurrencyPolicy: Forbid and Replace, and what production incident does picking the wrong one cause?
IntermediateKubernetes5 min
A CronJob has been silently creating thousands of failed Jobs over several days — how did this happen, and how would you prevent it?
AdvancedKubernetes7 min
A DaemonSet pod is missing from exactly one node while running fine everywhere else — how do you find out why?
IntermediateKubernetes6 min
A Job is supposed to run to completion exactly once, but it created multiple pods — why, and is that actually a bug?
IntermediateKubernetes6 min
A Deployment's new pods keep getting scheduled but immediately evicted, while old pods keep running fine — what changed?
AdvancedKubernetes7 min