>DevOps Interview KB

Advanced DevOps Interview Questions

194 questions

How would you safely roll out a breaking change to a DaemonSet running a critical node-level agent across a large production cluster?

AdvancedKubernetes7 min

A StatefulSet pod is deleted but isn't recreated with the same identity fast enough — what's actually blocking it?

AdvancedKubernetes7 min

How would you find the same df/du phantom-disk-usage issue on a containerized workload, where the culprit process might be in a different mount namespace?

AdvancedLinuxDocker8 min

A containerized process shows only 40% average CPU usage, well under its configured limit, but application metrics show frequent latency spikes correlating with CPU throttling events. How is this possible?

AdvancedLinuxKubernetes8 min

A production process appears to be running (it's in the process list, consuming no CPU) but isn't responding to any requests. How would you use strace to figure out what it's actually doing?

AdvancedLinux8 min

How does multi-window burn-rate alerting actually work mathematically — what does "burn rate" mean precisely?

AdvancedSREPrometheus8 min

How would you handle an SLO breach caused by a shared dependency affecting multiple services at once?

AdvancedSRE7 min

Your on-call team is getting paged 40+ times a week and starting to ignore alerts. How would you redesign the alerting strategy to fix that without missing real incidents?

AdvancedPrometheusMonitoring10 min

A team added a 'user_id' label to a metric to debug a specific customer's issue, and now your Prometheus instance is running out of memory and queries are timing out. What happened?

AdvancedPrometheus8 min

You need to track request latency percentiles across your fleet. When would you use a Prometheus Histogram versus a Summary metric type, and why can't you just average a Summary's percentiles across instances?

AdvancedPrometheus7 min

Your Prometheus storage costs are growing unsustainably from keeping raw-resolution metrics for a year. How would you design a retention and downsampling strategy to control costs?

AdvancedPrometheus7 min

How does NodeLocal DNSCache in Kubernetes actually reduce DNS-related failures, mechanically?

AdvancedKubernetesDNS7 min

Your load balancer keeps marking healthy backend servers as unhealthy and removing them from rotation, causing capacity to drop and requests to concentrate on fewer servers. How do you diagnose and fix this?

AdvancedNetworking8 min

Small requests to a service work fine, but any request with a larger payload just hangs and eventually times out, with no error on either side. What's the likely cause?

AdvancedNetworking8 min

A team has started ignoring alerts because most of them turn out to be noise — how do you diagnose and fix alert fatigue systematically?

AdvancedPrometheus8 min

How would you design a centralized log aggregation pipeline for a fleet of services, from collection through to searchable storage?

AdvancedLogging8 min

How would you choose the right SLIs for a new service, rather than defaulting to generic latency and error rate for everything?

AdvancedPrometheus7 min

During an incident spanning multiple services, how would you quickly correlate metrics across them to find where the problem actually originates?

AdvancedPrometheusGrafana8 min

A Prometheus instance's memory usage keeps growing and query performance is degrading — how do you diagnose a cardinality explosion?

AdvancedPrometheus8 min

How does trace sampling — head-based versus tail-based — affect whether distributed tracing actually catches the incidents you care about?

AdvancedOpentelemetry7 min

How would you retrofit distributed tracing into a system with no existing context propagation, without a disruptive big-bang migration?

AdvancedOpentelemetry8 min

How would you sunset or retire a golden path that's no longer the right default as the organization's needs have evolved?

AdvancedPlatform Engineering7 min

How do you decide what belongs on a golden path / internal developer platform versus what teams should just be free to do themselves?

AdvancedPlatform Engineering8 min

A Python automation script uses subprocess.run(f'kubectl get pod {pod_name}', shell=True) where pod_name comes from user input. What's actually wrong with this, and how do you fix it?

AdvancedPython7 min