>DevOps Interview KB

Staff / Principal Interview Questions

228 questions across 34 categories

Expert-level and system-design questions — the scope, trade-offs, and organizational reasoning expected at the most senior technical levels.

How would you troubleshoot connectivity that works within a node but fails between pods on different nodes?

AdvancedKubernetes8 min

A pod resolves a Service's DNS name intermittently but not consistently — how do you investigate CoreDNS itself?

AdvancedKubernetes8 min

How would you migrate a cluster from one CNI plugin to another without a full cluster rebuild — what's actually risky about it?

ExpertKubernetes8 min

What does a pod's ndots DNS setting default to, and how can it cause unexpectedly slow external DNS lookups?

AdvancedKubernetes7 min

A NetworkPolicy is applied but pods that should be blocked can still communicate — why might it not be enforced?

AdvancedKubernetes7 min

How would you debug a Service routing failure differently if you were using a service mesh like Istio instead of plain kube-proxy-based Services?

AdvancedKubernetesIstio7 min

A security team rejects a pod spec requesting privileged: true — what SecurityContext alternatives would you propose to meet the actual requirement?

AdvancedKubernetes7 min

How would you audit an entire cluster to find ServiceAccounts with effectively cluster-admin permissions before a security review?

AdvancedKubernetes7 min

A pod's ServiceAccount token was found in a public repo — what's your incident response, and how do you reduce blast radius for next time?

ExpertKubernetes8 min

How would you design least-privilege RBAC for a CI/CD pipeline that deploys to multiple namespaces, without granting cluster-admin?

AdvancedKubernetes8 min

A Pod Security Standard (restricted) rejects a legacy workload that needs to run as root — how do you handle this without disabling the standard cluster-wide?

AdvancedKubernetes7 min

How would you distinguish a genuine memory leak from a legitimately growing in-memory cache, using only Kubernetes-level metrics?

AdvancedKubernetesPrometheus8 min

Two pods that should never co-locate keep landing on the same node — what's wrong with the anti-affinity rule?

AdvancedKubernetes7 min

A critical pod gets preempted by a seemingly lower-priority pod during a resource crunch — how do you investigate and prevent it?

AdvancedKubernetes7 min

How would you dedicate a set of nodes exclusively to one team's workloads, while still letting that team's pods run elsewhere too?

AdvancedKubernetes7 min

How would you troubleshoot a pod stuck Pending even though kubectl describe pod shows no scheduling errors at all?

ExpertKubernetes8 min

How would you design pod topology spread constraints to keep a Deployment's replicas evenly distributed across availability zones?

AdvancedKubernetes7 min

A CSI driver upgrade causes new attach operations to fail while already-mounted volumes keep working — how do you investigate, and how would you roll this out more safely next time?

ExpertKubernetes8 min

How would you safely expand a PVC for a running StatefulSet without downtime, and what does the StorageClass need to support?

AdvancedKubernetes7 min

Why might mounting the same ReadWriteOnce PVC work for two pods on some clusters but fail on others?

AdvancedKubernetes6 min

A StatefulSet pod is rescheduled to a new node but its volume won't attach — what's happening, and how do you fix it?

AdvancedKubernetes8 min

What does volumeBindingMode: WaitForFirstConsumer actually solve, and what breaks in a multi-zone cluster if you don't set it?

AdvancedKubernetes6 min

A CronJob has been silently creating thousands of failed Jobs over several days — how did this happen, and how would you prevent it?

AdvancedKubernetes7 min

How would you design a Job for a task that must never run twice, even if a pod fails partway through (e.g., a billing charge)?

ExpertKubernetes8 min