SRE Interview Questions
34 questions across 3 categories
Site Reliability Engineering questions on SLOs, error budgets, incident response, and the observability practices that define the discipline.
How does multi-window burn-rate alerting actually work mathematically — what does "burn rate" mean precisely?
AdvancedSREPrometheus8 min
How would you set an initial SLO for a service that has no historical performance data to base it on at all?
IntermediateSRE6 min
How would you handle an SLO breach caused by a shared dependency affecting multiple services at once?
AdvancedSRE7 min
Your on-call team is getting paged 40+ times a week and starting to ignore alerts. How would you redesign the alerting strategy to fix that without missing real incidents?
AdvancedPrometheusMonitoring10 min
Your monitoring currently only checks whether an external health-check endpoint returns 200 OK. What's missing, and how do black-box and white-box monitoring actually complement each other?
IntermediateMonitoring6 min
A team added a 'user_id' label to a metric to debug a specific customer's issue, and now your Prometheus instance is running out of memory and queries are timing out. What happened?
AdvancedPrometheus8 min
Your team has 40 Grafana dashboards, and during the last incident nobody knew which one to actually look at first. How would you fix this dashboard sprawl?
IntermediateMonitoringGrafana6 min
You need to track request latency percentiles across your fleet. When would you use a Prometheus Histogram versus a Summary metric type, and why can't you just average a Summary's percentiles across instances?
AdvancedPrometheus7 min
A dashboard using rate() on a counter metric shows a value that doesn't match what you'd expect from doing the math manually. What's actually going on, and when should you use rate() versus increase()?
IntermediatePrometheus6 min
Your Prometheus storage costs are growing unsustainably from keeping raw-resolution metrics for a year. How would you design a retention and downsampling strategy to control costs?
AdvancedPrometheus7 min
A team has started ignoring alerts because most of them turn out to be noise — how do you diagnose and fix alert fatigue systematically?
AdvancedPrometheus8 min
How would you design a centralized log aggregation pipeline for a fleet of services, from collection through to searchable storage?
AdvancedLogging8 min
How would you choose the right SLIs for a new service, rather than defaulting to generic latency and error rate for everything?
AdvancedPrometheus7 min
During an incident spanning multiple services, how would you quickly correlate metrics across them to find where the problem actually originates?
AdvancedPrometheusGrafana8 min
A service's dashboard has 40 panels, and during a real incident nobody can find the signal that actually explains what's wrong. How would you redesign it?
IntermediatePrometheusGrafana6 min
How would you design a structured logging convention for a service, so logs are actually usable for debugging production incidents?
IntermediateLogging7 min
A team wants to log every value they'd normally track as a metric, reasoning logs give more detail. What breaks when metrics and logs' roles get confused?
IntermediatePrometheus6 min
A Prometheus instance's memory usage keeps growing and query performance is degrading — how do you diagnose a cardinality explosion?
AdvancedPrometheus8 min
What's the difference between the RED and USE monitoring methodologies, and which resources should each actually be applied to?
IntermediatePrometheus6 min
What's the difference between synthetic monitoring and real user monitoring, and why would you need both rather than just one?
IntermediatePrometheus6 min
How does trace sampling — head-based versus tail-based — affect whether distributed tracing actually catches the incidents you care about?
AdvancedOpentelemetry7 min
You already have logs and metrics for your services. When does it actually become worth investing in distributed tracing, and what problem does it solve that the other two don't?
IntermediateOpentelemetry7 min
How would you retrofit distributed tracing into a system with no existing context propagation, without a disruptive big-bang migration?
AdvancedOpentelemetry8 min
What's the relationship between distributed traces and structured logs — should trace IDs actually be injected into log lines, and why?
IntermediateOpentelemetry6 min