Senior DevOps Interview Questions
516 questions across 39 categories
Deeper trade-off and design questions — why a specific approach was chosen, what breaks at scale, and how to reason about a system you didn't build.
Your Prometheus storage costs are growing unsustainably from keeping raw-resolution metrics for a year. How would you design a retention and downsampling strategy to control costs?
AdvancedPrometheus7 min
How would you differentiate a client-side DNS resolution problem from the authoritative or upstream DNS server itself being unreliable?
IntermediateDNS7 min
An application intermittently fails to resolve a hostname — most requests succeed, but a small percentage fail with a DNS error. How would you track this down?
IntermediateDNSLinux8 min
How does NodeLocal DNSCache in Kubernetes actually reduce DNS-related failures, mechanically?
AdvancedKubernetesDNS7 min
Why does DNS primarily use UDP instead of TCP, and what does that choice trade off for reliability?
IntermediateDNS6 min
During every deployment, a handful of in-flight requests get dropped with connection-reset errors right as old servers are terminated. How do you fix this?
IntermediateNetworking6 min
Your load balancer keeps marking healthy backend servers as unhealthy and removing them from rotation, causing capacity to drop and requests to concentrate on fewer servers. How do you diagnose and fix this?
AdvancedNetworking8 min
What's the actual difference between a Layer 4 and a Layer 7 load balancer, and how does that difference affect what routing decisions each can make?
IntermediateNetworking6 min
Small requests to a service work fine, but any request with a larger payload just hangs and eventually times out, with no error on either side. What's the likely cause?
AdvancedNetworking8 min
An application currently relies on sticky sessions (a user's requests always route to the same backend) to work correctly. Why is this considered an anti-pattern, and how would you actually remove the dependency?
IntermediateNetworking7 min
Should TLS terminate at the load balancer, or should encrypted traffic pass all the way through to the backend servers? What's the actual security and operational trade-off?
IntermediateNetworkingSecurity7 min
A team has started ignoring alerts because most of them turn out to be noise — how do you diagnose and fix alert fatigue systematically?
AdvancedPrometheus8 min
How would you design a centralized log aggregation pipeline for a fleet of services, from collection through to searchable storage?
AdvancedLogging8 min
How would you choose the right SLIs for a new service, rather than defaulting to generic latency and error rate for everything?
AdvancedPrometheus7 min
During an incident spanning multiple services, how would you quickly correlate metrics across them to find where the problem actually originates?
AdvancedPrometheusGrafana8 min
A service's dashboard has 40 panels, and during a real incident nobody can find the signal that actually explains what's wrong. How would you redesign it?
IntermediatePrometheusGrafana6 min
How would you design a structured logging convention for a service, so logs are actually usable for debugging production incidents?
IntermediateLogging7 min
A team wants to log every value they'd normally track as a metric, reasoning logs give more detail. What breaks when metrics and logs' roles get confused?
IntermediatePrometheus6 min
A Prometheus instance's memory usage keeps growing and query performance is degrading — how do you diagnose a cardinality explosion?
AdvancedPrometheus8 min
What's the difference between the RED and USE monitoring methodologies, and which resources should each actually be applied to?
IntermediatePrometheus6 min
What's the difference between synthetic monitoring and real user monitoring, and why would you need both rather than just one?
IntermediatePrometheus6 min
How does trace sampling — head-based versus tail-based — affect whether distributed tracing actually catches the incidents you care about?
AdvancedOpentelemetry7 min
You already have logs and metrics for your services. When does it actually become worth investing in distributed tracing, and what problem does it solve that the other two don't?
IntermediateOpentelemetry7 min
How would you retrofit distributed tracing into a system with no existing context propagation, without a disruptive big-bang migration?
AdvancedOpentelemetry8 min