Senior DevOps Interview Questions
516 questions across 39 categories
Deeper trade-off and design questions — why a specific approach was chosen, what breaks at scale, and how to reason about a system you didn't build.
A rolling update to a Deployment is stuck at 50% — how do you determine whether it's a bad readiness probe, insufficient capacity, or a PodDisruptionBudget blocking it?
AdvancedKubernetes8 min
How would you safely roll out a breaking change to a DaemonSet running a critical node-level agent across a large production cluster?
AdvancedKubernetes7 min
A StatefulSet pod is deleted but isn't recreated with the same identity fast enough — what's actually blocking it?
AdvancedKubernetes7 min
Why does a StatefulSet's rolling update behave completely differently from a Deployment's, and why can that ordering guarantee become a problem mid-incident?
IntermediateKubernetes6 min
How would you build automated alerting specifically for the df/du divergence pattern, rather than relying on someone noticing it during an incident?
IntermediateLinuxPrometheus6 min
df says a server's disk is 100% full, but running du -sh on every top-level directory only adds up to a fraction of that. Where did the rest of the space go?
IntermediateLinux8 min
How would you find the same df/du phantom-disk-usage issue on a containerized workload, where the culprit process might be in a different mount namespace?
AdvancedLinuxDocker8 min
Why does Unix allow you to delete a file that's still open by a running process, instead of blocking the delete like some other operating systems do?
IntermediateLinux6 min
A containerized process shows only 40% average CPU usage, well under its configured limit, but application metrics show frequent latency spikes correlating with CPU throttling events. How is this possible?
AdvancedLinuxKubernetes8 min
A server's load average shows 8.0 on a 4-core machine, but CPU usage in top only shows 20%. How can load be double the core count while CPU usage looks low?
IntermediateLinux7 min
A service takes 25 seconds to shut down gracefully, but your orchestrator sends SIGKILL after a 10-second grace period, causing corrupted in-progress writes on every deploy. How do you actually fix this?
IntermediateLinux7 min
A production process appears to be running (it's in the process list, consuming no CPU) but isn't responding to any requests. How would you use strace to figure out what it's actually doing?
AdvancedLinux8 min
A misbehaving service is stuck in a rapid restart loop, consuming CPU and flooding logs, because systemd keeps restarting it immediately after every crash. How would you configure this correctly?
IntermediateLinuxSystemd6 min
A service running fine for months suddenly throws 'Too many open files' errors under increased load. How do you fix it correctly, rather than just raising the limit blindly?
IntermediateLinux7 min
How does debugging a systemd timer failure differ from debugging the same issue in cron?
IntermediateLinuxSystemd7 min
How does multi-window burn-rate alerting actually work mathematically — what does "burn rate" mean precisely?
AdvancedSREPrometheus8 min
How would you set an initial SLO for a service that has no historical performance data to base it on at all?
IntermediateSRE6 min
How would you handle an SLO breach caused by a shared dependency affecting multiple services at once?
AdvancedSRE7 min
Your on-call team is getting paged 40+ times a week and starting to ignore alerts. How would you redesign the alerting strategy to fix that without missing real incidents?
AdvancedPrometheusMonitoring10 min
Your monitoring currently only checks whether an external health-check endpoint returns 200 OK. What's missing, and how do black-box and white-box monitoring actually complement each other?
IntermediateMonitoring6 min
A team added a 'user_id' label to a metric to debug a specific customer's issue, and now your Prometheus instance is running out of memory and queries are timing out. What happened?
AdvancedPrometheus8 min
Your team has 40 Grafana dashboards, and during the last incident nobody knew which one to actually look at first. How would you fix this dashboard sprawl?
IntermediateMonitoringGrafana6 min
You need to track request latency percentiles across your fleet. When would you use a Prometheus Histogram versus a Summary metric type, and why can't you just average a Summary's percentiles across instances?
AdvancedPrometheus7 min
A dashboard using rate() on a counter metric shows a value that doesn't match what you'd expect from doing the math manually. What's actually going on, and when should you use rate() versus increase()?
IntermediatePrometheus6 min