Kubernetes Interview Guide
A structured path through Kubernetes interview prep — what to learn first, what interviewers probe for at each level, and which questions to practice.
Introduction
Kubernetes interviews rarely stay on the surface for long. A question that opens as "how do Services work" is usually a setup for "now here's a Service that isn't routing traffic — where do you look first." This guide is built around that pattern: it sequences the concepts in the order they actually build on each other, tells you what each interview level is really being screened for, and points you at the specific questions in this site's Kubernetes bank (128 of them, spanning every difficulty and topic area) that are worth your limited prep time.
This is not a copy of the question bank. Every section below explains why a topic matters and how to approach it — the actual answers, commands, and investigation steps live in the linked questions themselves, which this guide never duplicates.
Who This Guide Is For
This guide is written for engineers preparing for a DevOps, SRE, or Platform Engineering interview where Kubernetes is a significant part of the role — which today is most of them. It's useful whether you're:
- New to Kubernetes interviews and need a sequence to follow rather than randomly sampling questions.
- Experienced with Kubernetes operationally but have never had to articulate your reasoning out loud under interview pressure, where "I'd just check the logs" isn't a complete answer.
- Targeting a specific level — the further down this guide you go, the more it shifts from "does this candidate know the objects" toward "does this candidate reason correctly about failure modes and trade-offs."
It assumes you're preparing for a role that actually operates Kubernetes clusters day to day, not an academic quiz on Kubernetes internals for their own sake. If your interview is purely about container orchestration theory with no production angle, some of the later, judgment-heavy sections will be less relevant to you — the fundamentals sections still apply.
Prerequisites
Before working through this guide, you should already be comfortable with:
- Basic container concepts (what an image is, what a container actually isolates, the difference between an image and a running container) — if these are shaky, container fundamentals belong on your list before Kubernetes-specific prep.
- Reading and writing simple YAML.
- Basic Linux process and networking concepts (processes, ports, DNS resolution) — Kubernetes networking questions assume this baseline and won't re-explain it.
- Comfort with a command-line shell —
kubectlfluency is itself part of what's being evaluated, but you should already be generally comfortable typing commands and reading their output before adding Kubernetes-specific syntax on top.
You do not need prior production Kubernetes experience to use this guide productively — the Learning / Interview Path below is sequenced to build that understanding from the ground up. You do need the prerequisites above already in place, since this guide (and the question bank it links to) doesn't re-teach them.
Learning / Interview Path
Work through these in order. Each one builds on assumptions from the ones before it — jumping straight to Networking or Autoscaling without a solid grip on Scheduling will leave gaps that show up as vague answers later, even if you technically "know" the later topic in isolation.
- Fundamentals & Core Objects — Pods, Deployments, and the controller pattern that underlies almost everything else in Kubernetes. If you only remember one thing from this stage, make it this: Kubernetes is a collection of controllers continuously reconciling observed state toward desired state. Nearly every "why did X happen" question in later stages traces back to that one idea.
- Scheduling — how the scheduler actually places pods, and the specific mechanisms (taints/tolerations, affinity/anti-affinity, resource requests) that constrain or redirect that placement. This is a prerequisite for autoscaling, which is really "scheduling plus a feedback loop."
- Resource Management & Autoscaling — requests vs. limits, QoS classes, HPA/VPA, and the cluster-level capacity questions (cluster autoscaler, node provisioning latency) that only make sense once scheduling is solid.
- Networking — Services, Endpoints, DNS, and CNI-level cross-node connectivity. This is consistently one of the areas where candidates who've only used Kubernetes casually (rather than debugged it under pressure) fall short, because networking failures are frequently silent rather than erroring loudly.
- Storage — PersistentVolumes, PersistentVolumeClaims, StorageClasses, and the access-mode subtleties (
ReadWriteOnceis node-scoped, not pod-scoped) that trip up even experienced candidates. - Configuration — ConfigMaps, Secrets, and the propagation-timing gotchas (a Secret consumed via
envFromneeds a restart; a volume-mounted one doesn't, but lags) that show up constantly in real incidents. - RBAC & Cluster Security — authorization, ServiceAccounts, and cluster-hardening concerns like kubelet API exposure and read-only root filesystems. Treat this as one continuous topic, not two — access control and cluster hardening are the same discipline applied at different layers.
- Admission Control — mutating and validating webhooks, and the operational risks specific to them (certificate expiry taking down the whole cluster, multiple webhooks interacting unexpectedly). This is consistently the area where staff/principal-level questions concentrate, because webhook failures have outsized blast radius.
- CRDs & Operators — how custom resources, controllers, finalizers, and owner references extend Kubernetes' own reconciliation model to manage things Kubernetes doesn't natively understand (a database, a certificate, an Argo CD Application).
- Architecture & System Design — pulling every earlier topic together into cluster-level design decisions: control plane sizing and etcd performance, multi-tenancy models, and how you'd actually design a platform on top of everything above.
Notice troubleshooting isn't its own numbered stage. It's the single largest question_type in this bank (44 of the 128 Kubernetes questions), and it's deliberately spread across every topic above rather than collected in one place — because that's how it actually shows up in interviews and in production. Once you've worked through a topic's fundamentals, immediately practice its troubleshooting questions before moving on; see the Scenario & Troubleshooting Focus section below for how to do that systematically.
Key Concepts
A handful of ideas recur across nearly every topic above. If you can explain each of these precisely — not just recognize the term — you're in good shape for most of what follows:
- Reconciliation, not imperative execution. Controllers don't "do" things once; they continuously compare desired state (spec) against observed state (status) and act on the difference. This explains why a
kubectl applythat "did nothing visible" might still be working correctly, and why manually patching a live object is fragile — the next reconciliation loop can silently revert it. - Requests vs. limits are different mechanisms with different consequences. Requests drive scheduling decisions and (for CPU) HPA math; limits drive OOM kills and CPU throttling. Confusing the two is behind a surprising fraction of "why is this failing" questions.
- A Service routes to Endpoints, not to pods directly. Endpoints are populated by matching the Service's selector against pod labels, filtered by readiness. A perfectly healthy pod with a mismatched label produces a Service with zero Endpoints — and no error anywhere.
PendingandTerminatingboth mean "waiting on a precondition," not "broken." A Pending pod is waiting on a scheduling decision (capacity, taints, affinity, volume binding); a stuck Terminating pod is usually waiting on a finalizer or a node that can't confirm it's actually gone. Both require identifying what specifically is being waited on, not guessing.- Cluster-level control plane health (etcd, the API server, admission webhooks) is a distinct failure domain from workload health. A perfectly correct Deployment can fail for reasons that have nothing to do with its own spec — a slow admission webhook, an etcd disk-latency problem, an expired webhook certificate. Staff/principal-level questions concentrate here precisely because these failures are rarer but far more consequential.
Interview Focus
What's actually being evaluated shifts meaningfully by level, even when the surface topic (say, "tell me about Services") looks the same:
- Junior DevOps / DevOps Engineer — do you know what the core objects are and what they're for, and can you use
kubectlto actually investigate a straightforward problem (a Pending pod, a CrashLoopBackOff)? Interviewers are checking for real hands-on fluency, not memorized definitions. - Senior DevOps — can you reason about why something is happening, not just recognize the symptom? This is where "the Service has no Endpoints" needs to become "here's how I'd confirm that, and here's the three most likely reasons it happened."
- SRE — expect more weight on observability, incident response process, and designing for failure (PodDisruptionBudgets, autoscaling headroom, alerting on the right signals) rather than pure object mechanics.
- DevSecOps — RBAC, admission control, and cluster hardening questions carry more weight; expect to be asked not just "how does this work" but "what's the actual blast radius if this is misconfigured."
- Staff / Principal — architecture and system-design questions dominate: control plane scaling, multi-tenancy trade-offs, and judgment calls with no single correct answer, where the interviewer is evaluating how you reason under ambiguity, not whether you land on their preferred answer.
Across every level, the strongest answers name the actual kubectl command or field you'd check, not just the general idea — "I'd look at the Service's Endpoints" is weaker than "I'd run kubectl get endpoints <name> and compare the pod IPs it returns against kubectl get pods -o wide."
Scenario & Troubleshooting Focus
Troubleshooting is where Kubernetes interviews most often separate candidates who've operated real clusters from candidates who've only read about them — because the investigation process matters as much as the eventual fix. A strong answer to "a pod is stuck Pending" doesn't jump straight to a guess; it walks a narrowing sequence: what does kubectl describe pod actually say in its Events, is this a scheduling problem or an admission problem, what's the specific resource/taint/volume constraint involved.
The full set of troubleshooting-typed Kubernetes questions is available at Kubernetes Troubleshooting Questions — 44 questions spanning scheduling, networking, storage, autoscaling, RBAC, and admission control, each built around a real production symptom rather than a textbook definition. The scenario-style questions (open-ended, judgment-heavy situations rather than single-symptom debugging) are at Kubernetes Scenario Questions.
When practicing these, resist the urge to jump straight to the "Resolution" section of a question's answer. Read the symptom, write down (even mentally) what you'd actually check first and why, and only then compare against the question's own Investigation Steps — that rehearsal is what actually transfers to a live interview, since being handed the answer and recognizing it as correct is a different skill from producing it yourself under pressure.
Common Mistakes
- Treating
kubectlcommand memorization as the goal. Interviewers can tell the difference between someone who's memorizedkubectl get pods -o wideand someone who understands why that specific command answers the question at hand. Understand the reasoning, and the commands follow naturally. - Answering at the wrong altitude for the role. A junior-level answer that's all mechanics ("here's the YAML") reads as thin at senior level; a senior-level answer that's all trade-off discussion with no concrete command or field reads as vague at junior level. Calibrate to what Interview Focus above says your target level is actually evaluating.
- Skipping networking and storage because they feel less "interesting" than autoscaling or CRDs. These are consistently under-practiced relative to how often they actually come up, precisely because their failure modes are often silent rather than loudly erroring.
- Preparing troubleshooting and architecture as separate, unrelated skills. The best system-design answers are informed by having actually debugged the failure modes being designed around — practice them together, not sequentially in isolation.
- Not distinguishing "I don't know" from "let me reason through what I'd check." Kubernetes has enough surface area that no one knows every corner. Interviewers consistently rate "I haven't hit that specifically, but here's how I'd investigate it" far higher than a confident wrong guess.
Recommended Preparation Path
If you have limited time before an interview, prioritize in this order:
- Start here: Work through the Must Practice list above — it's deliberately spread across troubleshooting, networking, scheduling, storage, RBAC, resource management, and admission control, giving you a representative sample of the whole bank in under an hour.
- Next: Go back to the Learning / Interview Path and spend real time on whichever numbered stage felt weakest while doing the Must Practice set — don't move forward until that gap is closed, since later stages assume it.
- Then: Work through the troubleshooting questions for your two or three weakest topic areas specifically (linked above), since troubleshooting fluency is both the largest single category in this bank and the area interviewers weight most heavily as a real signal.
- Advanced / final pass: If you're targeting Senior, SRE, DevSecOps, or Staff/Principal specifically, use the "Prepare by Interview Level" links below to pull the question set scoped to your actual target level, and focus your remaining time there rather than continuing to work broadly.
References
Practice Questions
The list below is generated live from the current Kubernetes question bank — nothing here is a fixed list that goes stale as new questions are added. Start with Must Practice for a fast, representative cross-section, then use the difficulty and interview-level groupings to focus your remaining time on where you actually need it.
Must Practice
Short on time? This is a deliberately small, representative cross-section — work through these first.
- A pod goes into CrashLoopBackOff immediately after you roll out a ConfigMap change, but only in one namespace. How do you investigate it?Intermediate
- You deploy a new version of a Kubernetes Deployment. The pods show Running and pass their readiness checks, but the Service in front of them stops routing traffic entirely. Where do you look?Intermediate
- A pod stays Pending with node(s) had untolerated taint — how do you diagnose it and decide toleration vs. removing the taint?Intermediate
- A pod is stuck Pending with an event about its PVC failing to bind — how do you diagnose why?Intermediate
- A ServiceAccount that worked fine before a deployment suddenly gets Forbidden errors on the Kubernetes API — how do you diagnose it?Intermediate
- Your container runs fine locally but repeatedly gets OOMKilled after you deploy it to Kubernetes. How would you investigate it?Intermediate
- Users report increasing latency under load, but the HPA isn't scaling the Deployment at all — how do you figure out why?Advanced
- Every pod creation cluster-wide suddenly starts failing with an admission webhook TLS error — what happened, and how do you recover quickly?Expert
By Subcategory
What each area of Kubernetes interview prep actually covers, and how much of it exists in the bank.
- A CronJob has been silently creating thousands of failed Jobs over several days — how did this happen, and how would you prevent it?
- How would you decide between a Deployment, a StatefulSet, and a DaemonSet for three different real services (a stateless API, a database, a node agent)?
- A Job is supposed to run to completion exactly once, but it created multiple pods — why, and is that actually a bug?
- A pod stays Pending with node(s) had untolerated taint — how do you diagnose it and decide toleration vs. removing the taint?
- How would you dedicate a set of nodes exclusively to one team's workloads, while still letting that team's pods run elsewhere too?
- How would you design pod topology spread constraints to keep a Deployment's replicas evenly distributed across availability zones?
- Your container runs fine locally but repeatedly gets OOMKilled after you deploy it to Kubernetes. How would you investigate it?
- What's the difference in behavior between a container hitting its own memory limit versus the underlying node running out of memory overall?
- How would you distinguish a genuine memory leak from a legitimately growing in-memory cache, using only Kubernetes-level metrics?
- Users report increasing latency under load, but the HPA isn't scaling the Deployment at all — how do you figure out why?
- A team wants to run HPA and VPA on the same Deployment for both CPU and memory — what breaks if you're not careful, and how do you combine them safely?
- What's the difference between HPA scaling on CPU utilization versus a custom metric like queue depth, and when is CPU actually the wrong signal?
- You deploy a new version of a Kubernetes Deployment. The pods show Running and pass their readiness checks, but the Service in front of them stops routing traffic entirely. Where do you look?
- How would you migrate a cluster from one CNI plugin to another without a full cluster rebuild — what's actually risky about it?
- How does a Service-not-routing-traffic failure mode differ between a plain ClusterIP Service and one fronted by an Ingress controller?
Storage
11- A pod is stuck Pending with an event about its PVC failing to bind — how do you diagnose why?
- A CSI driver upgrade causes new attach operations to fail while already-mounted volumes keep working — how do you investigate, and how would you roll this out more safely next time?
- How would you back up and restore persistent volume data for a stateful app, given kubectl alone doesn't capture volume contents?
- An app reads an env var from a Secret, but after rotating the Secret's value, the running pod still uses the old one — why?
- A Secret manifest with real credentials was committed to a public repo — how does remediation differ from a generic leaked-secret response?
- How would you audit which pods across a cluster consume a specific Secret, before rotating it, to know what needs restarting?
- A ServiceAccount that worked fine before a deployment suddenly gets Forbidden errors on the Kubernetes API — how do you diagnose it?
- A security team rejects a pod spec requesting privileged: true — what SecurityContext alternatives would you propose to meet the actual requirement?
- How would you audit an entire cluster to find ServiceAccounts with effectively cluster-admin permissions before a security review?
- A security scan found the kubelet's API port reachable without authentication on some nodes — what can an attacker actually do with that, and how do you fix it?
- Runtime security tooling alerts that a specific pod is exhibiting behavior consistent with compromise — walk through your immediate containment response.
- A security scan flags the API server's anonymous authentication as enabled — what does that actually expose, and how would you harden it safely?
- Every pod creation cluster-wide suddenly starts failing with an admission webhook TLS error — what happened, and how do you recover quickly?
- How would you design and roll out a policy blocking :latest image tags cluster-wide, without breaking every existing deployment on day one?
- An admission webhook's failurePolicy is set to Fail — what happens if the webhook itself becomes unavailable, and why might that be the wrong default?
- A custom resource is stuck in Terminating status indefinitely after being deleted — what's a finalizer, and how does it cause this?
- A CRD needs a breaking schema change, but existing custom resources and consumers depend on the old shape — how do you version a CRD safely?
- Why should a CRD's status be a separate subresource from spec, and what belongs in status versus spec?
- The API server responds slowly to all requests — how do you determine whether etcd, the API server, or something else is the bottleneck?
- How would you design multi-cluster architecture — when does an org actually need multiple clusters instead of namespaces?
- How would you safely drain and remove a node without disrupting running workloads?
- A pod goes into CrashLoopBackOff immediately after you roll out a ConfigMap change, but only in one namespace. How do you investigate it?
- How would you set up alerting to catch a CrashLoopBackOff-class issue before it reaches production traffic, rather than discovering it via a user-facing outage?
- How would your investigation differ if a Pod entered ImagePullBackOff instead of CrashLoopBackOff?
Prepare by Interview Level
These aren't just counts — each one is a curated preparation path for that specific role. Pick the level you're actually interviewing for.
By Difficulty
Work upward if you're building from fundamentals, or jump straight to the difficulty an interview at your level will actually probe.
By Question Type
Practice one specific skill at a time — troubleshooting, architecture trade-offs, head-to-head comparisons, and more.
Related Guides
If your interview also covers AWS, the AWS DevOps Interview Guide is a natural companion — many teams running Kubernetes run it on EKS. If your interview also covers infrastructure provisioning, the Terraform Interview Guide is a companion too — Terraform-managed Kubernetes clusters raise their own state-and-module questions worth preparing alongside these.
Last updated August 23, 2026 · Last reviewed August 23, 2026