>DevOps Interview KB

Intermediate DevOps Interview Questions

322 questions

A Job is supposed to run to completion exactly once, but it created multiple pods — why, and is that actually a bug?

IntermediateKubernetes6 min

Why does a StatefulSet's rolling update behave completely differently from a Deployment's, and why can that ordering guarantee become a problem mid-incident?

IntermediateKubernetes6 min

How would you build automated alerting specifically for the df/du divergence pattern, rather than relying on someone noticing it during an incident?

IntermediateLinuxPrometheus6 min

df says a server's disk is 100% full, but running du -sh on every top-level directory only adds up to a fraction of that. Where did the rest of the space go?

IntermediateLinux8 min

Why does Unix allow you to delete a file that's still open by a running process, instead of blocking the delete like some other operating systems do?

IntermediateLinux6 min

A server's load average shows 8.0 on a 4-core machine, but CPU usage in top only shows 20%. How can load be double the core count while CPU usage looks low?

IntermediateLinux7 min

A service takes 25 seconds to shut down gracefully, but your orchestrator sends SIGKILL after a 10-second grace period, causing corrupted in-progress writes on every deploy. How do you actually fix this?

IntermediateLinux7 min

A misbehaving service is stuck in a rapid restart loop, consuming CPU and flooding logs, because systemd keeps restarting it immediately after every crash. How would you configure this correctly?

IntermediateLinuxSystemd6 min

A service running fine for months suddenly throws 'Too many open files' errors under increased load. How do you fix it correctly, rather than just raising the limit blindly?

IntermediateLinux7 min

How does debugging a systemd timer failure differ from debugging the same issue in cron?

IntermediateLinuxSystemd7 min

How would you set an initial SLO for a service that has no historical performance data to base it on at all?

IntermediateSRE6 min

Your monitoring currently only checks whether an external health-check endpoint returns 200 OK. What's missing, and how do black-box and white-box monitoring actually complement each other?

IntermediateMonitoring6 min

Your team has 40 Grafana dashboards, and during the last incident nobody knew which one to actually look at first. How would you fix this dashboard sprawl?

IntermediateMonitoringGrafana6 min

A dashboard using rate() on a counter metric shows a value that doesn't match what you'd expect from doing the math manually. What's actually going on, and when should you use rate() versus increase()?

IntermediatePrometheus6 min

How would you differentiate a client-side DNS resolution problem from the authoritative or upstream DNS server itself being unreliable?

IntermediateDNS7 min

An application intermittently fails to resolve a hostname — most requests succeed, but a small percentage fail with a DNS error. How would you track this down?

IntermediateDNSLinux8 min

Why does DNS primarily use UDP instead of TCP, and what does that choice trade off for reliability?

IntermediateDNS6 min

During every deployment, a handful of in-flight requests get dropped with connection-reset errors right as old servers are terminated. How do you fix this?

IntermediateNetworking6 min

What's the actual difference between a Layer 4 and a Layer 7 load balancer, and how does that difference affect what routing decisions each can make?

IntermediateNetworking6 min

An application currently relies on sticky sessions (a user's requests always route to the same backend) to work correctly. Why is this considered an anti-pattern, and how would you actually remove the dependency?

IntermediateNetworking7 min

Should TLS terminate at the load balancer, or should encrypted traffic pass all the way through to the backend servers? What's the actual security and operational trade-off?

IntermediateNetworkingSecurity7 min

A service's dashboard has 40 panels, and during a real incident nobody can find the signal that actually explains what's wrong. How would you redesign it?

IntermediatePrometheusGrafana6 min

How would you design a structured logging convention for a service, so logs are actually usable for debugging production incidents?

IntermediateLogging7 min

A team wants to log every value they'd normally track as a metric, reasoning logs give more detail. What breaks when metrics and logs' roles get confused?

IntermediatePrometheus6 min