>DevOps Interview KB

Comparison Interview Questions — All Categories

127 questions across 34 categories

Head-to-head trade-off questions — the kind that reveal whether a candidate has real judgment or just memorized facts.

How would you decide between a Deployment, a StatefulSet, and a DaemonSet for three different real services (a stateless API, a database, a node agent)?

IntermediateKubernetes6 min

What's the difference between a CronJob's concurrencyPolicy: Forbid and Replace, and what production incident does picking the wrong one cause?

IntermediateKubernetes5 min

What's actually different at the API level between kubectl rollout restart and deleting all of a Deployment's pods manually?

BeginnerKubernetes5 min

Why does a StatefulSet's rolling update behave completely differently from a Deployment's, and why can that ordering guarantee become a problem mid-incident?

IntermediateKubernetes6 min

How does debugging a systemd timer failure differ from debugging the same issue in cron?

IntermediateLinuxSystemd7 min

Your monitoring currently only checks whether an external health-check endpoint returns 200 OK. What's missing, and how do black-box and white-box monitoring actually complement each other?

IntermediateMonitoring6 min

You need to track request latency percentiles across your fleet. When would you use a Prometheus Histogram versus a Summary metric type, and why can't you just average a Summary's percentiles across instances?

AdvancedPrometheus7 min

A dashboard using rate() on a counter metric shows a value that doesn't match what you'd expect from doing the math manually. What's actually going on, and when should you use rate() versus increase()?

IntermediatePrometheus6 min

What's the actual difference between a Layer 4 and a Layer 7 load balancer, and how does that difference affect what routing decisions each can make?

IntermediateNetworking6 min

Should TLS terminate at the load balancer, or should encrypted traffic pass all the way through to the backend servers? What's the actual security and operational trade-off?

IntermediateNetworkingSecurity7 min

What's the difference between the RED and USE monitoring methodologies, and which resources should each actually be applied to?

IntermediatePrometheus6 min

What's the difference between synthetic monitoring and real user monitoring, and why would you need both rather than just one?

IntermediatePrometheus6 min

How does trace sampling — head-based versus tail-based — affect whether distributed tracing actually catches the incidents you care about?

AdvancedOpentelemetry7 min

You already have logs and metrics for your services. When does it actually become worth investing in distributed tracing, and what problem does it solve that the other two don't?

IntermediateOpentelemetry7 min

Should platform engineers be organized as one centralized team, or embedded within product teams? What actually determines the right structure?

IntermediatePlatform Engineering6 min

Do we actually need a full internal developer portal (like Backstage), or would genuinely good documentation and a couple of CLI tools solve most of the same problem?

IntermediatePlatform Engineering6 min

You want to speed up a Python automation script that processes many files. Should you use threading or multiprocessing? What does the GIL actually have to do with this decision?

IntermediatePython7 min

What's the difference between a Docker container being OOM-killed by its own cgroup memory limit versus the host machine itself running out of memory?

IntermediateDocker6 min

When would you actually choose manual credential rotation over fully automated rotation, given that automation is generally considered the more secure default?

IntermediateSecurity6 min

How would your incident response to a hardcoded secret committed to a repository differ if the repository were private rather than public?

IntermediateGit6 min

What's the trade-off between rewriting Git history to remove a committed secret versus simply rotating it and leaving the now-worthless value in history?

IntermediateGit6 min

How would you choose the rolling window length for an error budget — 7 days, 30 days, or 90 days — and what does that choice actually trade off?

IntermediateSRE6 min

What would make you choose a private Terraform module registry over Git-tag-based module sourcing, or vice versa?

IntermediateTerraform6 min

Should you use Terraform workspaces or separate directories/state files to manage dev, staging, and production — what's the actual trade-off?

IntermediateTerraform7 min