Advanced DevOps Interview Questions
194 questions
For distributing a complex application, when would you package it as a Helm chart versus building a dedicated operator for it?
AdvancedKubernetesHelm7 min
How would you test a custom operator's reconcile logic without needing a full live cluster for every test run?
AdvancedKubernetes7 min
A team wants to automate a repetitive operational task with a custom Kubernetes operator — when is that actually the right tool versus overkill?
AdvancedKubernetes7 min
How would you troubleshoot connectivity that works within a node but fails between pods on different nodes?
AdvancedKubernetes8 min
A pod resolves a Service's DNS name intermittently but not consistently — how do you investigate CoreDNS itself?
AdvancedKubernetes8 min
What does a pod's ndots DNS setting default to, and how can it cause unexpectedly slow external DNS lookups?
AdvancedKubernetes7 min
A NetworkPolicy is applied but pods that should be blocked can still communicate — why might it not be enforced?
AdvancedKubernetes7 min
How would you debug a Service routing failure differently if you were using a service mesh like Istio instead of plain kube-proxy-based Services?
AdvancedKubernetesIstio7 min
A security team rejects a pod spec requesting privileged: true — what SecurityContext alternatives would you propose to meet the actual requirement?
AdvancedKubernetes7 min
How would you audit an entire cluster to find ServiceAccounts with effectively cluster-admin permissions before a security review?
AdvancedKubernetes7 min
How would you design least-privilege RBAC for a CI/CD pipeline that deploys to multiple namespaces, without granting cluster-admin?
AdvancedKubernetes8 min
A Pod Security Standard (restricted) rejects a legacy workload that needs to run as root — how do you handle this without disabling the standard cluster-wide?
AdvancedKubernetes7 min
How would you distinguish a genuine memory leak from a legitimately growing in-memory cache, using only Kubernetes-level metrics?
AdvancedKubernetesPrometheus8 min
Two pods that should never co-locate keep landing on the same node — what's wrong with the anti-affinity rule?
AdvancedKubernetes7 min
A critical pod gets preempted by a seemingly lower-priority pod during a resource crunch — how do you investigate and prevent it?
AdvancedKubernetes7 min
How would you dedicate a set of nodes exclusively to one team's workloads, while still letting that team's pods run elsewhere too?
AdvancedKubernetes7 min
How would you design pod topology spread constraints to keep a Deployment's replicas evenly distributed across availability zones?
AdvancedKubernetes7 min
How would you safely expand a PVC for a running StatefulSet without downtime, and what does the StorageClass need to support?
AdvancedKubernetes7 min
Why might mounting the same ReadWriteOnce PVC work for two pods on some clusters but fail on others?
AdvancedKubernetes6 min
A StatefulSet pod is rescheduled to a new node but its volume won't attach — what's happening, and how do you fix it?
AdvancedKubernetes8 min
What does volumeBindingMode: WaitForFirstConsumer actually solve, and what breaks in a multi-zone cluster if you don't set it?
AdvancedKubernetes6 min
A CronJob has been silently creating thousands of failed Jobs over several days — how did this happen, and how would you prevent it?
AdvancedKubernetes7 min
A Deployment's new pods keep getting scheduled but immediately evicted, while old pods keep running fine — what changed?
AdvancedKubernetes7 min
A rolling update to a Deployment is stuck at 50% — how do you determine whether it's a bad readiness probe, insufficient capacity, or a PodDisruptionBudget blocking it?
AdvancedKubernetes8 min