>DevOps Interview KB

Observability Interview Questions

14 questions

Observability interview questions on metrics, logs, and traces as complementary signals — instrumentation decisions, cardinality trade-offs, and building systems you can actually debug in production rather than just watch.

A team has started ignoring alerts because most of them turn out to be noise — how do you diagnose and fix alert fatigue systematically?

AdvancedPrometheus8 min

How would you design a centralized log aggregation pipeline for a fleet of services, from collection through to searchable storage?

AdvancedLogging8 min

How would you choose the right SLIs for a new service, rather than defaulting to generic latency and error rate for everything?

AdvancedPrometheus7 min

During an incident spanning multiple services, how would you quickly correlate metrics across them to find where the problem actually originates?

AdvancedPrometheusGrafana8 min

A service's dashboard has 40 panels, and during a real incident nobody can find the signal that actually explains what's wrong. How would you redesign it?

IntermediatePrometheusGrafana6 min

How would you design a structured logging convention for a service, so logs are actually usable for debugging production incidents?

IntermediateLogging7 min

A team wants to log every value they'd normally track as a metric, reasoning logs give more detail. What breaks when metrics and logs' roles get confused?

IntermediatePrometheus6 min

A Prometheus instance's memory usage keeps growing and query performance is degrading — how do you diagnose a cardinality explosion?

AdvancedPrometheus8 min

What's the difference between the RED and USE monitoring methodologies, and which resources should each actually be applied to?

IntermediatePrometheus6 min

What's the difference between synthetic monitoring and real user monitoring, and why would you need both rather than just one?

IntermediatePrometheus6 min

How does trace sampling — head-based versus tail-based — affect whether distributed tracing actually catches the incidents you care about?

AdvancedOpentelemetry7 min

You already have logs and metrics for your services. When does it actually become worth investing in distributed tracing, and what problem does it solve that the other two don't?

IntermediateOpentelemetry7 min

How would you retrofit distributed tracing into a system with no existing context propagation, without a disruptive big-bang migration?

AdvancedOpentelemetry8 min

What's the relationship between distributed traces and structured logs — should trace IDs actually be injected into log lines, and why?

IntermediateOpentelemetry6 min