DevOps Engineer Interview Questions
372 questions across 38 categories
The bulk of a working DevOps engineer's day-to-day: hands-on tooling, common troubleshooting, and the practical judgment calls that come up once you're actually operating systems.
Why does Unix allow you to delete a file that's still open by a running process, instead of blocking the delete like some other operating systems do?
IntermediateLinux6 min
chmod 755 vs chmod 644 — what do these numbers actually mean, and how would you figure out the right mode for a new file without guessing?
BeginnerLinux6 min
A server's load average shows 8.0 on a 4-core machine, but CPU usage in top only shows 20%. How can load be double the core count while CPU usage looks low?
IntermediateLinux7 min
A service takes 25 seconds to shut down gracefully, but your orchestrator sends SIGKILL after a 10-second grace period, causing corrupted in-progress writes on every deploy. How do you actually fix this?
IntermediateLinux7 min
A misbehaving service is stuck in a rapid restart loop, consuming CPU and flooding logs, because systemd keeps restarting it immediately after every crash. How would you configure this correctly?
IntermediateLinuxSystemd6 min
A service running fine for months suddenly throws 'Too many open files' errors under increased load. How do you fix it correctly, rather than just raising the limit blindly?
IntermediateLinux7 min
How does debugging a systemd timer failure differ from debugging the same issue in cron?
IntermediateLinuxSystemd7 min
How would you set an initial SLO for a service that has no historical performance data to base it on at all?
IntermediateSRE6 min
Your monitoring currently only checks whether an external health-check endpoint returns 200 OK. What's missing, and how do black-box and white-box monitoring actually complement each other?
IntermediateMonitoring6 min
Your team has 40 Grafana dashboards, and during the last incident nobody knew which one to actually look at first. How would you fix this dashboard sprawl?
IntermediateMonitoringGrafana6 min
A dashboard using rate() on a counter metric shows a value that doesn't match what you'd expect from doing the math manually. What's actually going on, and when should you use rate() versus increase()?
IntermediatePrometheus6 min
How would you differentiate a client-side DNS resolution problem from the authoritative or upstream DNS server itself being unreliable?
IntermediateDNS7 min
An application intermittently fails to resolve a hostname — most requests succeed, but a small percentage fail with a DNS error. How would you track this down?
IntermediateDNSLinux8 min
Why does DNS primarily use UDP instead of TCP, and what does that choice trade off for reliability?
IntermediateDNS6 min
During every deployment, a handful of in-flight requests get dropped with connection-reset errors right as old servers are terminated. How do you fix this?
IntermediateNetworking6 min
What's the actual difference between a Layer 4 and a Layer 7 load balancer, and how does that difference affect what routing decisions each can make?
IntermediateNetworking6 min
An application currently relies on sticky sessions (a user's requests always route to the same backend) to work correctly. Why is this considered an anti-pattern, and how would you actually remove the dependency?
IntermediateNetworking7 min
Should TLS terminate at the load balancer, or should encrypted traffic pass all the way through to the backend servers? What's the actual security and operational trade-off?
IntermediateNetworkingSecurity7 min
A service's dashboard has 40 panels, and during a real incident nobody can find the signal that actually explains what's wrong. How would you redesign it?
IntermediatePrometheusGrafana6 min
How would you design a structured logging convention for a service, so logs are actually usable for debugging production incidents?
IntermediateLogging7 min
A team wants to log every value they'd normally track as a metric, reasoning logs give more detail. What breaks when metrics and logs' roles get confused?
IntermediatePrometheus6 min
What's the difference between the RED and USE monitoring methodologies, and which resources should each actually be applied to?
IntermediatePrometheus6 min
What's the difference between synthetic monitoring and real user monitoring, and why would you need both rather than just one?
IntermediatePrometheus6 min
You already have logs and metrics for your services. When does it actually become worth investing in distributed tracing, and what problem does it solve that the other two don't?
IntermediateOpentelemetry7 min