>DevOps Interview KB

Conceptual Interview Questions — All Categories

166 questions across 34 categories

The underlying mechanisms and trade-offs interviewers expect you to actually understand, not just recite.

What's the difference in behavior between a container hitting its own memory limit versus the underlying node running out of memory overall?

IntermediateKubernetes6 min

How do requests.memory and limits.memory affect Pod scheduling and node-level eviction behavior differently from each other?

IntermediateKubernetes6 min

How does the Kubernetes scheduler actually decide which node to place a pod on — walk through filtering and scoring?

IntermediateKubernetes6 min

Why might mounting the same ReadWriteOnce PVC work for two pods on some clusters but fail on others?

AdvancedKubernetes6 min

What does volumeBindingMode: WaitForFirstConsumer actually solve, and what breaks in a multi-zone cluster if you don't set it?

AdvancedKubernetes6 min

How would your investigation differ if a Pod entered ImagePullBackOff instead of CrashLoopBackOff?

BeginnerKubernetes6 min

How do liveness and readiness probes interact with a Pod that's already crash-looping on startup?

IntermediateKubernetes6 min

A Job is supposed to run to completion exactly once, but it created multiple pods — why, and is that actually a bug?

IntermediateKubernetes6 min

What's actually different at the API level between kubectl rollout restart and deleting all of a Deployment's pods manually?

BeginnerKubernetes5 min

Why does Unix allow you to delete a file that's still open by a running process, instead of blocking the delete like some other operating systems do?

IntermediateLinux6 min

chmod 755 vs chmod 644 — what do these numbers actually mean, and how would you figure out the right mode for a new file without guessing?

BeginnerLinux6 min

A server's load average shows 8.0 on a 4-core machine, but CPU usage in top only shows 20%. How can load be double the core count while CPU usage looks low?

IntermediateLinux7 min

How does debugging a systemd timer failure differ from debugging the same issue in cron?

IntermediateLinuxSystemd7 min

How does multi-window burn-rate alerting actually work mathematically — what does "burn rate" mean precisely?

AdvancedSREPrometheus8 min

Your monitoring currently only checks whether an external health-check endpoint returns 200 OK. What's missing, and how do black-box and white-box monitoring actually complement each other?

IntermediateMonitoring6 min

Your team has 40 Grafana dashboards, and during the last incident nobody knew which one to actually look at first. How would you fix this dashboard sprawl?

IntermediateMonitoringGrafana6 min

How does NodeLocal DNSCache in Kubernetes actually reduce DNS-related failures, mechanically?

AdvancedKubernetesDNS7 min

Why does DNS primarily use UDP instead of TCP, and what does that choice trade off for reliability?

IntermediateDNS6 min

A team wants to log every value they'd normally track as a metric, reasoning logs give more detail. What breaks when metrics and logs' roles get confused?

IntermediatePrometheus6 min

You already have logs and metrics for your services. When does it actually become worth investing in distributed tracing, and what problem does it solve that the other two don't?

IntermediateOpentelemetry7 min

What's the relationship between distributed traces and structured logs — should trace IDs actually be injected into log lines, and why?

IntermediateOpentelemetry6 min

How do you decide what belongs on a golden path / internal developer platform versus what teams should just be free to do themselves?

AdvancedPlatform Engineering8 min

Finance is asking your platform team to justify its headcount with concrete ROI numbers. How would you actually measure and communicate the value a platform team provides?

IntermediatePlatform Engineering7 min

What does it actually mean to run an internal platform 'as a product,' and how is that concretely different from just building infrastructure tooling?

IntermediatePlatform Engineering6 min