>DevOps Interview KB

Staff / Principal Interview Questions

228 questions across 34 categories

Expert-level and system-design questions — the scope, trade-offs, and organizational reasoning expected at the most senior technical levels.

Two mutating webhooks both touch a pod's containers field, and the final result isn't what either webhook intended alone — how do you diagnose and resolve this?

ExpertKubernetes8 min

A new mutating webhook accidentally intercepted kube-system pod creation and broke core cluster components — how would you design its scoping to prevent this?

AdvancedKubernetes7 min

Every pod creation cluster-wide suddenly starts failing with an admission webhook TLS error — what happened, and how do you recover quickly?

ExpertKubernetes8 min

An admission webhook's failurePolicy is set to Fail — what happens if the webhook itself becomes unavailable, and why might that be the wrong default?

AdvancedKubernetes7 min

A validating webhook is rejecting pod creations that look completely valid — how do you diagnose what it's actually objecting to?

AdvancedKubernetes7 min

kubectl apply --dry-run=server behaves unexpectedly for requests a webhook handles — what does a webhook's sideEffects field have to do with it?

AdvancedKubernetes6 min

The API server responds slowly to all requests — how do you determine whether etcd, the API server, or something else is the bottleneck?

ExpertKubernetes9 min

How would you design backup/DR for etcd, and what's actually recoverable from a snapshot versus what isn't?

AdvancedKubernetesEtcd8 min

How would you design multi-cluster architecture — when does an org actually need multiple clusters instead of namespaces?

AdvancedKubernetes8 min

A cluster upgrade needs zero workload downtime — walk through sequencing control-plane and node upgrades safely.

AdvancedKubernetes8 min

A team wants to run HPA and VPA on the same Deployment for both CPU and memory — what breaks if you're not careful, and how do you combine them safely?

ExpertKubernetes8 min

An HPA scales up rapidly during a spike, then flaps up and down repeatedly for the next hour — what's causing it, and how do you fix it?

AdvancedKubernetes7 min

Users report increasing latency under load, but the HPA isn't scaling the Deployment at all — how do you figure out why?

AdvancedKubernetes8 min

A Deployment's pods get evicted during scale-up because new nodes take too long to become ready — how would you close that gap?

ExpertKubernetes8 min

How would you design autoscaling for a workload with a sharp, predictable daily spike versus one with genuinely unpredictable bursty traffic?

AdvancedKubernetes7 min

A security scan flags the API server's anonymous authentication as enabled — what does that actually expose, and how would you harden it safely?

AdvancedKubernetes7 min

Runtime security tooling alerts that a specific pod is exhibiting behavior consistent with compromise — walk through your immediate containment response.

ExpertKubernetes8 min

How would you design a Kubernetes audit logging policy that's actually useful for a security investigation, without drowning in log volume?

AdvancedKubernetes7 min

How would you design a policy requiring every image deployed to a cluster be cryptographically signed, and what does that actually protect against?

ExpertKubernetes8 min

A security scan found the kubelet's API port reachable without authentication on some nodes — what can an attacker actually do with that, and how do you fix it?

ExpertKubernetes8 min

How would you design network-level segmentation between the control plane and worker nodes, beyond what Kubernetes' own RBAC and NetworkPolicy provide?

AdvancedKubernetes7 min

You already enforce preventive admission policies (Kyverno/Gatekeeper) — why would you also need runtime security tooling like Falco?

AdvancedKubernetesFalco6 min

What does a seccomp profile actually add on top of SecurityContext's capability restrictions, and when do you need one?

AdvancedKubernetes6 min

How would you design a workflow so a ConfigMap change automatically triggers a rolling restart of the Deployments that depend on it?

AdvancedKubernetes7 min