Kubernetes Interview Questions — All Categories
169 questions tagged with Kubernetes as a technology, across every category it appears in
Related categories:Argo CDAWSAzureCI/CDCloud ArchitectureContainersDockerGitOpsHelmJenkinsKubernetesLinuxNetworkingSystem DesignYAML
How would you design a highly-available control plane, and what breaks with only one control-plane node?
IntermediateKubernetes7 min
How would you design backup/DR for etcd, and what's actually recoverable from a snapshot versus what isn't?
AdvancedKubernetesEtcd8 min
How does etcd's quorum requirement affect control-plane node count, and why is an even number a bad choice?
IntermediateKubernetesEtcd6 min
What's etcd's actual role in a cluster, and what happens if it becomes unavailable?
BeginnerKubernetesEtcd5 min
How would you safely drain and remove a node without disrupting running workloads?
IntermediateKubernetes6 min
What's the difference between a static pod and a normal pod, and why does the control plane often run as static pods?
IntermediateKubernetes5 min
How would you design multi-cluster architecture — when does an org actually need multiple clusters instead of namespaces?
AdvancedKubernetes8 min
A cluster upgrade needs zero workload downtime — walk through sequencing control-plane and node upgrades safely.
AdvancedKubernetes8 min
A team wants to run HPA and VPA on the same Deployment for both CPU and memory — what breaks if you're not careful, and how do you combine them safely?
ExpertKubernetes8 min
What's the difference between HPA scaling on CPU utilization versus a custom metric like queue depth, and when is CPU actually the wrong signal?
IntermediateKubernetes6 min
How does HPA's scaling decision actually get computed from raw metrics — walk through what happens between a CPU spike and a new replica appearing?
IntermediateKubernetes6 min
Why might an HPA be unable to scale a Deployment even with plenty of spare CPU capacity on existing nodes?
IntermediateKubernetes6 min
An HPA scales up rapidly during a spike, then flaps up and down repeatedly for the next hour — what's causing it, and how do you fix it?
AdvancedKubernetes7 min
Users report increasing latency under load, but the HPA isn't scaling the Deployment at all — how do you figure out why?
AdvancedKubernetes8 min
Why does setting only limits (no requests) on a container break HPA's CPU-based scaling calculation?
BeginnerKubernetes5 min
A workload is constantly OOMKilled despite having an HPA configured — why doesn't horizontal scaling fix this, and what should you actually do?
IntermediateKubernetes6 min
A Deployment's pods get evicted during scale-up because new nodes take too long to become ready — how would you close that gap?
ExpertKubernetes8 min
How would you design autoscaling for a workload with a sharp, predictable daily spike versus one with genuinely unpredictable bursty traffic?
AdvancedKubernetes7 min
A security scan flags the API server's anonymous authentication as enabled — what does that actually expose, and how would you harden it safely?
AdvancedKubernetes7 min
What does running kube-bench against a cluster actually check, and how would you prioritize the findings rather than trying to fix everything at once?
IntermediateKubernetes6 min
Runtime security tooling alerts that a specific pod is exhibiting behavior consistent with compromise — walk through your immediate containment response.
ExpertKubernetes8 min
How would you design a Kubernetes audit logging policy that's actually useful for a security investigation, without drowning in log volume?
AdvancedKubernetes7 min
How would you design a policy requiring every image deployed to a cluster be cryptographically signed, and what does that actually protect against?
ExpertKubernetes8 min
A security scan found the kubelet's API port reachable without authentication on some nodes — what can an attacker actually do with that, and how do you fix it?
ExpertKubernetes8 min