>DevOps Interview KB

Practical Interview Questions — All Categories

131 questions across 30 categories

Hands-on, command-driven questions testing whether you can actually operate these tools, not just discuss them.

How would you set up alerting to catch a CrashLoopBackOff-class issue before it reaches production traffic, rather than discovering it via a user-facing outage?

IntermediateKubernetesPrometheus7 min

How would you build automated alerting specifically for the df/du divergence pattern, rather than relying on someone noticing it during an incident?

IntermediateLinuxPrometheus6 min

chmod 755 vs chmod 644 — what do these numbers actually mean, and how would you figure out the right mode for a new file without guessing?

BeginnerLinux6 min

A production process appears to be running (it's in the process list, consuming no CPU) but isn't responding to any requests. How would you use strace to figure out what it's actually doing?

AdvancedLinux8 min

A misbehaving service is stuck in a rapid restart loop, consuming CPU and flooding logs, because systemd keeps restarting it immediately after every crash. How would you configure this correctly?

IntermediateLinuxSystemd6 min

How would you set an initial SLO for a service that has no historical performance data to base it on at all?

IntermediateSRE6 min

Your on-call team is getting paged 40+ times a week and starting to ignore alerts. How would you redesign the alerting strategy to fix that without missing real incidents?

AdvancedPrometheusMonitoring10 min

Your team has 40 Grafana dashboards, and during the last incident nobody knew which one to actually look at first. How would you fix this dashboard sprawl?

IntermediateMonitoringGrafana6 min

A dashboard using rate() on a counter metric shows a value that doesn't match what you'd expect from doing the math manually. What's actually going on, and when should you use rate() versus increase()?

IntermediatePrometheus6 min

A team has started ignoring alerts because most of them turn out to be noise — how do you diagnose and fix alert fatigue systematically?

AdvancedPrometheus8 min

During an incident spanning multiple services, how would you quickly correlate metrics across them to find where the problem actually originates?

AdvancedPrometheusGrafana8 min

A service's dashboard has 40 panels, and during a real incident nobody can find the signal that actually explains what's wrong. How would you redesign it?

IntermediatePrometheusGrafana6 min

How would you retrofit distributed tracing into a system with no existing context propagation, without a disruptive big-bang migration?

AdvancedOpentelemetry8 min

How do you handle a team that has a legitimate reason to deviate from a golden path — what does a good exception process actually look like?

IntermediatePlatform Engineering6 min

How would you measure whether a golden path is actually succeeding, versus just being nominally adopted because it's mandated?

IntermediatePlatform Engineering6 min

Usage data shows one of your platform's supported deployment strategies is used by exactly two teams out of sixty, but maintaining it still consumes real platform-team time. How do you handle deprecating it?

IntermediatePlatform Engineering6 min

Finance is asking your platform team to justify its headcount with concrete ROI numbers. How would you actually measure and communicate the value a platform team provides?

IntermediatePlatform Engineering7 min

A Python CLI tool used in CI pipelines always exits with code 0, even when it detects and reports a real failure. What breaks because of this, and how do you fix it?

IntermediatePython6 min

An automation script opens a database connection, does some work, and closes it at the end — but an exception partway through leaves the connection open. Why does 'with' fix this, and how does it actually work?

IntermediatePython6 min

A script calling a flaky API retries immediately with a fixed 1-second delay. During a real outage, this made things worse. Why, and how should retry logic actually be designed?

IntermediatePython7 min

How would you handle a data transformation that genuinely needs to see the whole dataset at once, like a global sort, when streaming isn't an option?

AdvancedPython8 min

How would you design a recurring privileged-access review that catches stale access at scale without becoming a rubber-stamp exercise nobody takes seriously?

AdvancedSecurity7 min

Your team resolves most incidents quickly because a few senior engineers just 'know' what to do. What's actually wrong with that, and how would you convert that knowledge into runbooks without slowing the seniors down?

IntermediateSRE6 min

Your SRE team spends most of its time on manual, repetitive operational work and has no time left for the reliability engineering they were hired to do. How would you identify and reduce this 'toil' systematically?

IntermediateSRE7 min