>DevOps Interview KB

Senior DevOps Interview Questions

516 questions across 39 categories

Deeper trade-off and design questions — why a specific approach was chosen, what breaks at scale, and how to reason about a system you didn't build.

Kubernetes has built-in admission controllers compiled into the API server, separate from webhook-based ones — what's the actual difference and when does it matter?

IntermediateKubernetes5 min

How would you design and roll out a policy blocking :latest image tags cluster-wide, without breaking every existing deployment on day one?

IntermediateKubernetes7 min

How do OPA Gatekeeper and Kyverno actually differ as policy engines for Kubernetes admission control, and which would you choose?

IntermediateKubernetesOpaKyverno6 min

A new mutating webhook accidentally intercepted kube-system pod creation and broke core cluster components — how would you design its scoping to prevent this?

AdvancedKubernetes7 min

What's the difference between a validating and a mutating admission webhook, and in what order do they actually run?

IntermediateKubernetes6 min

An admission webhook's failurePolicy is set to Fail — what happens if the webhook itself becomes unavailable, and why might that be the wrong default?

AdvancedKubernetes7 min

A validating webhook is rejecting pod creations that look completely valid — how do you diagnose what it's actually objecting to?

AdvancedKubernetes7 min

kubectl apply --dry-run=server behaves unexpectedly for requests a webhook handles — what does a webhook's sideEffects field have to do with it?

AdvancedKubernetes6 min

For enforcing a new policy, when does it belong in a cluster admission webhook versus a CI-time check before deployment even happens?

IntermediateKubernetes6 min

What's the difference between kube-apiserver, kube-scheduler, and kube-controller-manager, and what breaks if each is unavailable?

IntermediateKubernetes6 min

How would you design a highly-available control plane, and what breaks with only one control-plane node?

IntermediateKubernetes7 min

How would you design backup/DR for etcd, and what's actually recoverable from a snapshot versus what isn't?

AdvancedKubernetesEtcd8 min

How does etcd's quorum requirement affect control-plane node count, and why is an even number a bad choice?

IntermediateKubernetesEtcd6 min

How would you safely drain and remove a node without disrupting running workloads?

IntermediateKubernetes6 min

What's the difference between a static pod and a normal pod, and why does the control plane often run as static pods?

IntermediateKubernetes5 min

How would you design multi-cluster architecture — when does an org actually need multiple clusters instead of namespaces?

AdvancedKubernetes8 min

A cluster upgrade needs zero workload downtime — walk through sequencing control-plane and node upgrades safely.

AdvancedKubernetes8 min

What's the difference between HPA scaling on CPU utilization versus a custom metric like queue depth, and when is CPU actually the wrong signal?

IntermediateKubernetes6 min

How does HPA's scaling decision actually get computed from raw metrics — walk through what happens between a CPU spike and a new replica appearing?

IntermediateKubernetes6 min

Why might an HPA be unable to scale a Deployment even with plenty of spare CPU capacity on existing nodes?

IntermediateKubernetes6 min

An HPA scales up rapidly during a spike, then flaps up and down repeatedly for the next hour — what's causing it, and how do you fix it?

AdvancedKubernetes7 min

Users report increasing latency under load, but the HPA isn't scaling the Deployment at all — how do you figure out why?

AdvancedKubernetes8 min

A workload is constantly OOMKilled despite having an HPA configured — why doesn't horizontal scaling fix this, and what should you actually do?

IntermediateKubernetes6 min

How would you design autoscaling for a workload with a sharp, predictable daily spike versus one with genuinely unpredictable bursty traffic?

AdvancedKubernetes7 min