DevOps Engineer Roadmap
A sequenced path from DevOps foundations to production-ready engineering skills, built from this bank's 552 most job-relevant questions.
Introduction
This roadmap sequences 552 of this bank's 590 technical interview questions into a single path: what to learn, in what order, to go from DevOps fundamentals to job-ready. It doesn't answer interview questions for you — the DevOps Interview Guide and category pages already do that. What this roadmap adds is sequencing: which of the bank's 16 preparation domains to tackle first, which depend on which, and when you've actually covered enough of a stage to move on.
Twelve stages, each scoped to a live-derived slice of the question bank, each with a small checkpoint set and a short representative practice sample — never the full stage dumped onto one page.
Who This Roadmap Is For
This targets the DevOps Engineer interview level specifically — the bulk of a working DevOps engineer's day-to-day tooling, troubleshooting, and judgment calls, not a first-role primer and not a staff-level design review. If you're earlier in your career, the sequence still applies, but expect some stages (Cloud, GitOps, Security) to assume more working familiarity than a from-scratch tutorial would.
Reality check on the source material: true beginner-difficulty content in this bank is thin (50 of 590 questions, concentrated almost entirely in Foundations, Cloud Fundamentals, Docker, and Git). Past those, the bank assumes you already have working familiarity and tests judgment, not vocabulary. This roadmap's job is sequencing and checkpointing that reality, not manufacturing tutorial content that doesn't exist here — where a stage's own coverage is thin, "What to Learn" points to the general skill, and the checkpoint questions are there to test it, not teach it from zero.
Deferred Areas
Four categories are deliberately not part of this roadmap, each for a specific, data-grounded reason:
sre— zero questions at the devops-engineer level in this bank; every one of them assumes senior/staff framing (error budgets, on-call design). Reserved for a future SRE roadmap.platform-engineering— same pattern: organizational/strategic content (golden paths, internal developer platforms) that presupposes the tooling this roadmap teaches, rather than teaching it. Reserved for a future Platform Engineer roadmap.databases— 8 of its 10 questions are advanced or expert difficulty (replication split-brain, sync-vs-async trade-offs); this category skews senior in this bank, not devops-engineer. Reserved for a future SRE roadmap, alongsidesre.system-design— 100% senior/staff-level, zero devops-engineer coverage. Reserved for a future Senior DevOps roadmap.
This isn't a content gap being hidden — it's the same category-level data used to sequence the 12 stages below, applied honestly to decide what doesn't belong in a DevOps Engineer's path yet.
Stages
Foundations
52 questions in this stage's category pool
Linux, Bash, Python, and YAML — the vocabulary every later stage assumes without re-teaching it. This is also where this bank's beginner-difficulty content concentrates most heavily (21 of 52 questions), making it the one stage genuinely safe to start from zero.
Why It Matters
Nearly every later stage's questions lean on this vocabulary silently — a CI/CD question about a failing pipeline step assumes you can read the Bash it's running; a Kubernetes manifest question assumes YAML fluency. Skipping this stage doesn't remove that assumption, it just means later stages read as harder than they are.
What to Learn
- Navigate and troubleshoot a Linux filesystem, process tree, and disk usage from the command line.
- Write and debug Bash scripts, including error handling (
set -e, exit codes, quoting) — not just happy-path one-liners. - Read and write YAML correctly, including the specific gotchas (anchors, implicit type coercion) that break configs silently.
- Explain core DevOps vocabulary and practices well enough to use them precisely in later stages, not just recognize them.
Resources
Checkpoint Questions
- A backup script runs perfectly when you execute it by hand, but fails silently every night when cron runs it. How would you debug it?
- A Kubernetes manifest that looked correct in the editor fails to apply with a cryptic YAML parsing error. How do you systematically debug indentation-related YAML errors?
- df says a server's disk is 100% full, but running du -sh on every top-level directory only adds up to a fraction of that. Where did the rest of the space go?
Representative Practice
- A script that processes filenames from a directory listing works fine in testing, but breaks (or silently does the wrong thing) the moment a filename has a space in it. Why, and how do you fix it?
- A script needs cloud CLI credentials (like AWS/Azure/gcloud auth) that are normally provided by an interactive session's environment. How do you handle that for a cron-scheduled script?
- A Python automation script uses subprocess.run(f'kubectl get pod {pod_name}', shell=True) where pod_name comes from user input. What's actually wrong with this, and how do you fix it?
- A script processes 10,000 files one at a time in a for loop and takes hours. Someone suggests using xargs -P to parallelize it. What are the actual trade-offs and risks of doing that?
- A provisioning script failed halfway through, and re-running it from the top caused errors because some resources already existed. How do you make it safely re-runnable?
Ready to Move On When...
- [ ] Can explain why a script with
set -estill let a failing command through - [ ] Can diagnose disk usage that doesn't match what
du/dfreport inside a container - [ ] Can identify a YAML indentation or type-coercion bug before it reaches a linter
- [ ] Comfortable reading unfamiliar Bash and Python scripts without line-by-line hand-holding
Source Control & Collaboration
30 questions in this stage's category pool
Git internals and platform-level governance (GitHub, GitLab) — sequenced right after Foundations because the next stage, CI/CD, is triggered by and reasons about Git events.
Prerequisites: Foundations
Why It Matters
You can't reason about a stuck or misfiring CI/CD pipeline without first understanding what triggered it — a branch protection rule, a CODEOWNERS mismatch, a force-push that rewrote history a teammate was building on. This stage builds that vocabulary before CI/CD needs it.
What to Learn
- Recover from common Git incidents: a bad force-push, a botched merge-conflict resolution, a history rewrite that needs undoing.
- Understand branch protection rules and why a required status check can get stuck in a state that looks wrong but isn't.
- Understand platform-level governance mechanics (CODEOWNERS, team permissions, API rate limits) beyond plain
gitcommands.
Checkpoint Questions
- A teammate's work was lost after a force-push overwrote their commits on a shared branch. How would you recover it?
- A required status check on a GitHub branch protection rule is permanently stuck showing "Expected — Waiting for status to be reported", blocking every PR from merging. What's going on?
- A teammate resolved a merge conflict by keeping 'their' version of every conflicting hunk without actually reading the other side's changes. Why is that dangerous, and how should conflicts actually be resolved?
Representative Practice
- An internal tool making GitHub API calls started getting 403 errors, but checking the rate limit endpoint shows plenty of requests remaining. What else could be causing this?
- Your org requires branch protection on main everywhere, but a critical tool has just one maintainer who finds required review genuinely slows down urgent fixes. How do you balance this?
- A bug was introduced somewhere in the last 200 commits, but nobody knows exactly which one, and the bug doesn't reproduce reliably enough to eyeball the diffs. How would you find the exact commit using git bisect?
- A repository has grown to several gigabytes because it stores large binary assets (design files, ML model weights) directly, making every clone painfully slow. How would you fix this?
- A secret was committed several commits ago and has since been rotated, but it's still sitting in the repository's Git history. How do you actually remove it, not just delete it in a new commit?
Ready to Move On When...
- [ ] Can recover a teammate's lost work after a force-push overwrite
- [ ] Can diagnose why a required GitHub status check is stuck rather than failed
- [ ] Understands the difference between a Git problem and a platform-governance problem
Containers
29 questions in this stage's category pool
Docker and container runtime internals — deliberately before Orchestration, since this bank's Kubernetes questions routinely assume you already understand what a container actually is at the OS level.
Prerequisites: Foundations
Why It Matters
Kubernetes doesn't re-explain namespaces, cgroups, or image layers — it assumes them. A candidate who jumps straight to Kubernetes without this stage can usually operate kubectl but struggles the moment a question asks why a container behaves a certain way, not just what command fixes it.
What to Learn
- Explain what actually isolates a container (namespaces, cgroups) and where that isolation is thinner than people assume.
- Diagnose container networking and volume issues — permission errors, disk usage that doesn't add up, connectivity between containers.
- Distinguish an OOM-killed container from other failure modes, and explain why the distinction matters for the fix.
Checkpoint Questions
- Two containers on the same Docker host can't talk to each other over the network, even though both are running. How would you debug the connection?
- A container keeps getting killed, and both the container's own memory limit and the host's overall available memory look potentially responsible. How do you determine which one actually caused it?
- A container fails to write to a bind-mounted directory with permission denied, even though the host directory has permissive permissions — why?
Representative Practice
- Your containerized application's process count keeps growing over time, even though nothing appears to be leaking connections or memory. What's likely happening, and why is it specific to running in a container?
- How would you explain the security trade-off between containers and VMs to someone deciding whether to run untrusted third-party code?
- How does Docker layer caching behave differently on CI runners without a persistent Docker daemon between builds, and how would you mitigate the resulting slowdown?
- How would you design volume mounts for a containerized production database, considering both data persistence and backup requirements?
- What's actually the difference between a container and a virtual machine, at the operating system level — not the marketing-slide version?
Ready to Move On When...
- [ ] Can explain the difference between a container being OOM-killed and crashing for another reason
- [ ] Can diagnose a bind-mount permission error
- [ ] Can explain why anonymous volumes accumulate disk space over time
- [ ] Comfortable with basic image, network, and volume troubleshooting without looking up every flag
CI/CD
44 questions in this stage's category pool
Pipeline design and the platform-specific mechanics of GitHub Actions, GitLab CI, Jenkins, and Azure Pipelines — building on Source Control (you can't reason about a stuck pipeline without understanding what triggered it) and Containers (build stages assume runtime literacy).
Prerequisites: Source Control & Collaboration, Containers
Why It Matters
This is where "I can write a YAML pipeline" turns into "I can diagnose why this specific pipeline is stuck, flaky, or silently wrong" — concurrency control, cache-vs-artifact confusion, and parallel-stage failures are the actual day-to-day of a DevOps engineer's CI/CD work, not pipeline syntax.
What to Learn
- Design pipelines with correct concurrency control, so simultaneous merges don't race each other.
- Understand the real difference between cache and artifacts, and why conflating them causes subtle failures.
- Diagnose parallel-stage and self-hosted-runner failures across at least one major CI/CD platform in depth.
Resources
Checkpoint Questions
- Two people merged to main within a minute of each other, and now two deployment workflow runs are executing at the same time, racing to deploy. How do you prevent this?
- A teammate used 'cache' to pass compiled build output from the build job to the deploy job, and it intermittently fails to find the files. What did they get wrong?
- Your Jenkins pipeline runs three test suites in parallel stages. One fails after 2 minutes, but the pipeline keeps running the other two for another 20 minutes before reporting failure. How do you fix this?
Representative Practice
- An Azure Pipelines YAML job runs fine on Microsoft-hosted agents but fails only on self-hosted agents with 'command not found' for a tool the job needs. Why, and how do you fix it?
- How would you keep a fleet of self-hosted Azure Pipelines agents consistent over time as tool requirements evolve, instead of them slowly drifting apart?
- How would you design a multi-stage Azure Pipelines YAML pipeline so production deployment requires manual approval, while dev and staging deploy automatically?
- A single Azure service connection with subscription-wide Contributor access is used by every pipeline, including ones that only read a storage account. What's wrong here?
- Beyond tool availability, what are the actual trade-offs of self-hosted versus Microsoft-hosted Azure Pipelines agents — cost, control, security, and startup latency?
Ready to Move On When...
- [ ] Can explain why two merges within a minute of each other caused a pipeline race
- [ ] Can diagnose a
cachevsartifactsmisuse in GitLab CI - [ ] Can debug a parallel Jenkins stage failure
- [ ] Comfortable reading pipeline YAML across more than one CI/CD platform
Orchestration
142 questions in this stage's category pool
Kubernetes and Helm — by a wide margin the largest stage in this roadmap (142 of 552 questions, over a quarter of the roadmap's total scope). Go deep here; it's the dominant real-world DevOps Engineer skill in this bank. The dedicated Kubernetes Interview Guide covers this stage's full 128-question Kubernetes slice in depth.
Prerequisites: Containers
Why It Matters
This is the single heaviest-weighted skill area in the entire bank. A DevOps Engineer interview that touches Kubernetes at all will go deep, and this bank's own weighting reflects that reality — budget your largest single block of preparation time here.
What to Learn
- Understand workloads, scheduling, and resource management (requests/limits) well enough to explain why the scheduler made a given decision.
- Diagnose autoscaling (HPA) behavior, including the specific ways misconfigured requests/limits silently break it.
- Understand RBAC, cluster security, and admission control at a working level, not just "apply this YAML."
- Understand Helm release lifecycle — hooks, rollbacks, values precedence — since it's the packaging layer most real clusters run on top of.
Checkpoint Questions
- Why does setting only limits (no requests) on a container break HPA's CPU-based scaling calculation?
- A Helm pre-upgrade hook Job doesn't run before the Deployment update it's supposed to precede — why, and how do you fix the ordering?
- Two mutating webhooks both touch a pod's containers field, and the final result isn't what either webhook intended alone — how do you diagnose and resolve this?
Representative Practice
- A helm rollback succeeds but the application still behaves like the newer version — why might rollback not fully revert the deployed state?
- A new mutating webhook accidentally intercepted kube-system pod creation and broke core cluster components — how would you design its scoping to prevent this?
- How would you share common templates (labels, resource boilerplate) across many microservice charts without copy-pasting them into every chart?
- An admission webhook's failurePolicy is set to Fail — what happens if the webhook itself becomes unavailable, and why might that be the wrong default?
- What's the difference between a helm test resource and a Helm lifecycle hook, and when would you use each?
Ready to Move On When...
- [ ] Can explain why an HPA doesn't scale despite rising load (and the requests/limits root cause)
- [ ] Can diagnose a Helm hook that isn't running in the expected order
- [ ] Understands RBAC and admission control well enough to reason about a rejected request
- [ ] Comfortable navigating workloads, scheduling, and storage without leaning on a cheat sheet
Infrastructure as Code
40 questions in this stage's category pool
Terraform, tool-agnostic IaC practice, and Ansible — sequenced right after Orchestration, since a meaningful share of this bank's Terraform questions provision and manage the exact kind of infrastructure the previous stage just covered. The dedicated Terraform Interview Guide covers this stage's Terraform slice in depth.
Prerequisites: Orchestration
Why It Matters
Provisioning is a different discipline from operating what's provisioned — state management, drift, and idempotency are their own category of failure mode, distinct from anything in the Orchestration stage, and this bank tests them accordingly.
What to Learn
- Understand Terraform state mechanics well enough to diagnose a stuck lock or an unexpected plan diff.
- Recognize configuration drift (a manual console change diverging from code) and know how to reconcile it safely.
- Understand what idempotency actually guarantees in Ansible playbooks, and why "ran successfully" doesn't always mean "didn't change anything it shouldn't have."
Resources
Checkpoint Questions
- terraform plan fails with Error acquiring the state lock — how do you diagnose whether it's a genuine concurrent run or a stale lock?
- Someone manually changed a cloud resource in the console that's managed by your IaC. What actually happens the next time the pipeline runs, and how would you handle the drift?
- An Ansible playbook reports "changed" on the same tasks every single run, even when nothing about the target host actually changed. Why, and how do you fix it?
Representative Practice
- terraform plan shows changes to a resource you didn't touch, and the diff traces back to a data source — what's actually happening?
- How would you convince a team to invest time in an idempotency retrofit when the playbook "already works"?
- How would you audit an existing large Ansible playbook to find all the places check mode's coverage is actually incomplete?
- How would you design scheduled Terraform drift detection so it alerts the right team without becoming noise nobody reads?
- A database password is passed as a resource argument in Terraform — does marking the variable sensitive actually protect it in the state file?
Ready to Move On When...
- [ ] Can diagnose a Terraform state lock acquisition failure
- [ ] Can explain how to detect and reconcile infrastructure drift
- [ ] Can explain why an Ansible playbook reported changes on a run that should have been a no-op
- [ ] Comfortable reading a Terraform plan diff and explaining what's about to happen before running apply
Cloud Platforms
120 questions in this stage's category pool
AWS, GCP, Azure, cloud fundamentals, and cross-cloud architecture. Lands here deliberately — by this point you already understand provisioning and orchestration, so cloud-specific content (IAM models, managed Kubernetes, serverless) is an application of those concepts to a vendor, not a disconnected topic. The dedicated AWS DevOps Interview Guide, Azure DevOps Interview Guide, and GCP Interview Guide cover this stage's three single-provider slices. Pick one provider to go deep on rather than splitting attention evenly across all three. AWS is this bank's deepest single-provider path (39 questions, a dedicated Guide, and a clear IAM/Lambda/S3 structure). Azure is a close second (34 questions, its own dedicated Guide covering Identity & Networking, Compute, Storage, and AKS) — a strong choice if AWS isn't your target, especially if AKS is central to the role. GCP is comparably deep (27 questions, its own dedicated Guide covering IAM, Storage, and Cloud Functions). All three now have a dedicated Guide sequencing the material for you — pick based on which provider your target role actually uses, not on depth alone.
Prerequisites: Infrastructure as Code
Why It Matters
Nearly every DevOps Engineer role is anchored to a specific cloud provider in practice — depth on one provider reads far stronger in an interview than shallow familiarity with three.
What to Learn
- Go deep on IAM, storage, and compute (serverless or managed containers) for one primary provider.
- Understand the security failure modes specific to that provider (public bucket exposure, over-permissioned roles) well enough to both cause and fix them in a discussion.
- Understand managed Kubernetes autoscaling behavior on at least one provider, connecting back to the Orchestration stage.
Resources
Checkpoint Questions
- A security scanner just flagged one of your production S3 buckets as publicly readable. Walk through how you'd respond in the first hour and prevent a repeat.
- Your AKS cluster is under CPU pressure and pods are stuck Pending, but the cluster autoscaler isn't adding nodes. How would you troubleshoot it?
- A Lambda function times out for about 2% of invocations, seemingly at random, but works fine when you test it manually. How would you track down the cause?
Representative Practice
- How would you detect, from CloudWatch metrics alone, whether a Lambda function's tail latency problem is cold-start-driven versus something else?
- How would you make the case for the cost of a separate AWS account, if leadership pushes back on the added complexity?
- How would you audit whether an ECS task role or Lambda execution role is actually scoped tightly, versus just copy-pasted from a broader existing role?
- Even after mitigating cold starts, a small amount of irreducible tail latency remains. How would you design the caller's retry behavior to handle that remaining tail?
- You inherit an EC2 workload that authenticates to AWS using an IAM user with AdministratorAccess. How would you migrate it to least-privilege access without causing an outage?
Ready to Move On When...
- [ ] Can explain how to determine what changed after an accidental public storage exposure
- [ ] Can diagnose a managed-Kubernetes autoscaler that isn't scaling under load
- [ ] Can debug a serverless function timeout or cold-start latency issue
- [ ] Has picked one primary provider and can go deep on it, not just describe all three at a survey level
Security Fundamentals
24 questions in this stage's category pool
Security and DevSecOps — explicitly scoped as baseline DevOps security competency, not full DevSecOps specialization (that's a future, dedicated roadmap). Placed after CI/CD, Containers, and Cloud, since most of this stage's questions assume you already understand the surfaces being secured.
Prerequisites: CI/CD, Containers, Cloud Platforms
Why It Matters
A DevOps Engineer interview expects you to reason about IAM sprawl, supply-chain risk, and access-control failures as part of normal operations — not as a separate specialist's job. This stage is the floor every DevOps Engineer needs, not the ceiling a DevSecOps Engineer eventually reaches.
What to Learn
- Diagnose IAM and access-control failures: service-account sprawl, shared admin accounts, SSO lockouts.
- Understand supply-chain risk at a working level — dependency confusion, third-party CI/CD action risk — without needing to design a full mitigation program.
- Recognize when a security question is asking for a specific mechanism versus a general awareness answer, and calibrate accordingly.
Checkpoint Questions
- A critical production system is accessed via one shared 'admin' account used by six engineers, with no individual audit trail. How do you fix this?
- Your build just pulled in a package from the public npm registry instead of your internal package with the same name, and it wasn't the version your team published. What's happening, and how do you respond?
- After a change to your SSO/identity provider configuration, nobody — including admins — can log into any connected system. How do you get back in, and how do you diagnose the actual cause?
Representative Practice
- A widely-used open-source dependency your organization relies on is publicly disclosed as compromised — a malicious backdoor was found in a recent release. How do you respond?
- How would you handle a pre-existing backlog of medium-severity security findings that predates your new scanning rollout, without blocking every team's work on day one?
- How would you measure whether engineers are actually acting on security scan findings, versus just dismissing them to unblock their PR?
- How would you design artifact signing into your CI/CD pipeline so a deployed container image or binary can be verified as genuinely coming from your build, not tampered with in transit?
- A developer just committed a live database password directly into a public GitHub repository. It's been merged and pushed. What do you do, in order?
Ready to Move On When...
- [ ] Can diagnose why a security audit found hundreds of untraceable service accounts
- [ ] Can explain dependency confusion and why it's a real supply-chain risk, not a theoretical one
- [ ] Can reason about an SSO lockout without immediately reaching for "just reset the password"
Observability Fundamentals
24 questions in this stage's category pool
Monitoring and observability — you need something running (every stage before this one) before "how do you know it's healthy" is a meaningful question. The deeper SRE-specific content (error budgets, on-call design) is deliberately excluded here; see Deferred Areas.
Prerequisites: Cloud Platforms
Why It Matters
Cardinality explosions, alert fatigue, and SLO-setting-without-history are the practical, day-to-day observability failures a DevOps Engineer actually hits — distinct from the org-level SRE practice this roadmap defers.
What to Learn
- Understand metric cardinality well enough to diagnose (and prevent) a cardinality explosion before it takes down monitoring itself.
- Set an initial SLO reasonably when there's no historical data to base it on.
- Recognize alert fatigue as a design problem with the alerting rules, not a tooling problem.
Checkpoint Questions
- A team added a 'user_id' label to a metric to debug a specific customer's issue, and now your Prometheus instance is running out of memory and queries are timing out. What happened?
- A Prometheus instance's memory usage keeps growing and query performance is degrading — how do you diagnose a cardinality explosion?
- How would you set an initial SLO for a service that has no historical performance data to base it on at all?
Representative Practice
- How would you handle an SLO breach caused by a shared dependency affecting multiple services at once?
- Your on-call team is getting paged 40+ times a week and starting to ignore alerts. How would you redesign the alerting strategy to fix that without missing real incidents?
- Your monitoring currently only checks whether an external health-check endpoint returns 200 OK. What's missing, and how do black-box and white-box monitoring actually complement each other?
- How does multi-window burn-rate alerting actually work mathematically — what does "burn rate" mean precisely?
- Your team has 40 Grafana dashboards, and during the last incident nobody knew which one to actually look at first. How would you fix this dashboard sprawl?
Ready to Move On When...
- [ ] Can explain why adding a high-cardinality label to a metric caused a monitoring outage
- [ ] Can propose a reasonable initial SLO with no historical data
- [ ] Can diagnose noisy alert thresholds causing fatigue, and propose a fix
GitOps & Modern Delivery
23 questions in this stage's category pool
GitOps and Argo CD — a synthesis of Source Control, CI/CD, and Orchestration, sequenced after all three because it presupposes them. Genuinely more advanced content in this bank (only a third of its questions sit at the devops-engineer level), appropriately placed late in the path.
Prerequisites: Source Control & Collaboration, CI/CD, Orchestration
Why It Matters
Git-as-source-of-truth deployment is increasingly how modern Kubernetes delivery actually works in production — this stage connects everything learned so far into the deployment model a growing share of DevOps Engineer roles now run on.
What to Learn
- Diagnose an Argo CD application stuck OutOfSync despite matching pods.
- Understand Argo CD sync waves and hook ordering well enough to control migration and cleanup behavior deliberately.
- Reason about what belongs in Git versus what a mutating webhook should handle at apply time.
Checkpoint Questions
- An Argo CD application shows OutOfSync even though the pods running in the cluster look correct and healthy. What would you check?
- How would you clean up old, hash-named Argo CD migration Jobs so they don't accumulate indefinitely in the cluster?
- A mutating webhook injects configuration into your resources. When should that injected config actually be tracked in Git instead of ignored via ignoreDifferences?
Representative Practice
- How would you distinguish 'the migration itself is broken' from 'this was a transient failure worth retrying,' in terms of alerting design?
- How would you handle a migration that needs to run exactly once across an entire fleet of clusters, not just once per cluster?
- GitOps means Git is the source of truth for everything deployed, but you obviously can't commit plaintext secrets to Git. How do you actually reconcile this?
- How do Argo CD's resource hooks compare to Helm's pre-install/pre-upgrade hooks, given Argo CD can also deploy Helm charts directly?
- How do you control the order Argo CD applies resources within a single Application, e.g. making sure a database migration Job completes before the Deployment that depends on it rolls out?
Ready to Move On When...
- [ ] Can diagnose an Argo CD application stuck OutOfSync
- [ ] Can explain how sync waves control apply order and why that matters for migrations
- [ ] Can reason about what configuration belongs in Git versus a runtime admission mechanism
Networking Essentials
10 questions in this stage's category pool
DNS, load balancing, and the connectivity failures that show up in production — small (10 questions) but consistently referenced by earlier stages (Docker networking, Kubernetes networking, cloud VPCs) without ever being taught standalone. This is the dedicated, focused pass.
Prerequisites: Containers, Cloud Platforms
Why It Matters
Networking failures are disproportionately represented in real production incidents relative to how much standalone attention they get in most preparation — this stage exists because every earlier stage assumed just enough networking to function, without ever teaching it directly.
What to Learn
- Diagnose DNS resolution failures, including the specific client-side vs. server-side distinction.
- Diagnose load-balancer health-check flapping and connection-draining issues during deployments.
- Understand the practical trade-offs load balancers make (sticky sessions, MTU mismatches) well enough to reason about them, not just name them.
Resources
Checkpoint Questions
- An application intermittently fails to resolve a hostname — most requests succeed, but a small percentage fail with a DNS error. How would you track this down?
- Your load balancer keeps marking healthy backend servers as unhealthy and removing them from rotation, causing capacity to drop and requests to concentrate on fewer servers. How do you diagnose and fix this?
- During every deployment, a handful of in-flight requests get dropped with connection-reset errors right as old servers are terminated. How do you fix this?
Representative Practice
- How would you differentiate a client-side DNS resolution problem from the authoritative or upstream DNS server itself being unreliable?
- An application currently relies on sticky sessions (a user's requests always route to the same backend) to work correctly. Why is this considered an anti-pattern, and how would you actually remove the dependency?
- What's the actual difference between a Layer 4 and a Layer 7 load balancer, and how does that difference affect what routing decisions each can make?
- How does NodeLocal DNSCache in Kubernetes actually reduce DNS-related failures, mechanically?
- Why does DNS primarily use UDP instead of TCP, and what does that choice trade off for reliability?
Ready to Move On When...
- [ ] Can differentiate a client-side from a server-side DNS resolution problem
- [ ] Can diagnose a load balancer flapping healthy backends as unhealthy
- [ ] Can explain what happens to in-flight requests during a deployment without connection draining
Production Readiness
14 questions in this stage's category pool
Scenarios and troubleshooting methodology — the capstone. Only 14 questions live directly in these two categories, but this stage's real material is the cross-cutting scenario (92 questions) and troubleshooting (141 questions) question types spread across every stage above. This is the synthesis stage, not a new topic.
A note on precision, since this trips people up: /troubleshooting (this stage's own category) is 4 general-methodology questions. /type/troubleshooting is 141 questions — every troubleshooting-tagged question across the entire bank, spanning every category above. They are not the same thing, and this roadmap never adds them together.
Prerequisites: Security Fundamentals, Observability Fundamentals, GitOps & Modern Delivery, Networking Essentials
Why It Matters
By this stage you've covered the mechanics of every domain — this is where judgment gets tested: triage under ambiguity, deciding when to escalate, and reasoning through legacy systems and imperfect trade-offs with no single correct answer. It's also the most realistic simulation of what an actual interview's hardest questions look like.
What to Learn
- Triage an ambiguous incident ("the app is slow," no other context) methodically rather than guessing.
- Decide when to escalate versus continue investigating alone, and articulate why.
- Reason through legacy-system scenarios — refactor vs. rewrite, inheriting undocumented systems — as genuine trade-offs, not solvable-with-one-right-answer puzzles.
- Practice broadly across the cross-cutting troubleshooting question type and scenario question type, not just this stage's own 14 questions.
Checkpoint Questions
- You get paged with just "the app is slow" and no other context. Walk through your actual troubleshooting methodology before you touch anything.
- How would you decide whether to gradually refactor a legacy pipeline versus rewriting it from scratch?
- You've proposed a robust, well-architected solution to a problem, but your manager wants a quick, hacky fix instead because the problem is genuinely minor. Are they wrong, or are you over-engineering?
Representative Practice
- You find a piece of a legacy pipeline that seems actively dangerous — overly broad credentials, say — but nobody can explain why it's configured that way. What do you do?
- How do you decide when to escalate or pull in another team during an incident, versus continuing to investigate solo?
- You inherit a legacy CI/CD pipeline with zero documentation, and the person who built it left the company months ago. How do you approach understanding it and safely making your first change?
- How do you balance the time investment in fully understanding a legacy system against the pressure to just make the requested change quickly?
- You inherit ownership of a system whose fundamental architecture you think was the wrong call — not a small detail, a core design decision. How do you handle this, given the system is already in production?
Ready to Move On When...
- [ ] Can lay out a methodical triage plan for a vague "app is slow" page with no other context
- [ ] Can articulate a clear, reasoned escalation threshold
- [ ] Can argue both sides of a refactor-vs-rewrite decision on a legacy system
- [ ] Has practiced enough of the cross-cutting troubleshooting/scenario question types to be comfortable with unfamiliar categories, not just the ones covered directly above
Related Guides
Three stages in this roadmap have a dedicated deep-reference Guide: Kubernetes Interview Guide for the full Orchestration stage, Terraform Interview Guide for the Infrastructure as Code stage's Terraform slice, and the AWS DevOps Interview Guide, Azure DevOps Interview Guide, and GCP Interview Guide for the Cloud Platforms stage's three single-provider paths. For the roadmap's full scope in one place — all 16 domains, not just the 12 stages sequenced here — see the umbrella DevOps Interview Guide, which this roadmap draws its structure from but sequences differently: the Guide is organized for reference (jump to any domain), this roadmap is organized for progression (work through in order).
Last updated August 23, 2026 · Last reviewed August 23, 2026