>DevOps Interview KB

Scenario Interview Questions — All Categories

92 questions across 28 categories

Open-ended, realistic situations that combine multiple technologies and force genuine engineering judgment.

A new mutating webhook accidentally intercepted kube-system pod creation and broke core cluster components — how would you design its scoping to prevent this?

AdvancedKubernetes7 min

Every pod creation cluster-wide suddenly starts failing with an admission webhook TLS error — what happened, and how do you recover quickly?

ExpertKubernetes8 min

An admission webhook's failurePolicy is set to Fail — what happens if the webhook itself becomes unavailable, and why might that be the wrong default?

AdvancedKubernetes7 min

How would you design multi-cluster architecture — when does an org actually need multiple clusters instead of namespaces?

AdvancedKubernetes8 min

A cluster upgrade needs zero workload downtime — walk through sequencing control-plane and node upgrades safely.

AdvancedKubernetes8 min

A team wants to run HPA and VPA on the same Deployment for both CPU and memory — what breaks if you're not careful, and how do you combine them safely?

ExpertKubernetes8 min

A Deployment's pods get evicted during scale-up because new nodes take too long to become ready — how would you close that gap?

ExpertKubernetes8 min

How would you design autoscaling for a workload with a sharp, predictable daily spike versus one with genuinely unpredictable bursty traffic?

AdvancedKubernetes7 min

Runtime security tooling alerts that a specific pod is exhibiting behavior consistent with compromise — walk through your immediate containment response.

ExpertKubernetes8 min

A Secret manifest with real credentials was committed to a public repo — how does remediation differ from a generic leaked-secret response?

AdvancedKubernetes7 min

A CRD needs a breaking schema change, but existing custom resources and consumers depend on the old shape — how do you version a CRD safely?

ExpertKubernetes8 min

A team wants to automate a repetitive operational task with a custom Kubernetes operator — when is that actually the right tool versus overkill?

AdvancedKubernetes7 min

How would you migrate a cluster from one CNI plugin to another without a full cluster rebuild — what's actually risky about it?

ExpertKubernetes8 min

A security team rejects a pod spec requesting privileged: true — what SecurityContext alternatives would you propose to meet the actual requirement?

AdvancedKubernetes7 min

A pod's ServiceAccount token was found in a public repo — what's your incident response, and how do you reduce blast radius for next time?

ExpertKubernetes8 min

How would you design least-privilege RBAC for a CI/CD pipeline that deploys to multiple namespaces, without granting cluster-admin?

AdvancedKubernetes8 min

A Pod Security Standard (restricted) rejects a legacy workload that needs to run as root — how do you handle this without disabling the standard cluster-wide?

AdvancedKubernetes7 min

How would you dedicate a set of nodes exclusively to one team's workloads, while still letting that team's pods run elsewhere too?

AdvancedKubernetes7 min

A CSI driver upgrade causes new attach operations to fail while already-mounted volumes keep working — how do you investigate, and how would you roll this out more safely next time?

ExpertKubernetes8 min

A team deletes a PVC expecting the data gone, but it's later recovered from the underlying disk — why, and how should reclaim policy be chosen deliberately?

IntermediateKubernetes6 min

A production workload needs persistent storage, and its pods may be rescheduled to different nodes — how do you design storage so data survives that?

BeginnerKubernetes6 min

A CronJob has been silently creating thousands of failed Jobs over several days — how did this happen, and how would you prevent it?

AdvancedKubernetes7 min

How would you design a Job for a task that must never run twice, even if a pod fails partway through (e.g., a billing charge)?

ExpertKubernetes8 min

How would you safely roll out a breaking change to a DaemonSet running a critical node-level agent across a large production cluster?

AdvancedKubernetes7 min