>DevOps Interview KB

Troubleshooting Interview Questions — All Categories

141 questions across 32 categories

Real production incidents, spanning Kubernetes, AWS, Docker, Terraform, CI/CD, networking, and databases — diagnosing a specific symptom back to its actual cause, not textbook definitions.

Small requests to a service work fine, but any request with a larger payload just hangs and eventually times out, with no error on either side. What's the likely cause?

AdvancedNetworking8 min

A Prometheus instance's memory usage keeps growing and query performance is degrading — how do you diagnose a cardinality explosion?

AdvancedPrometheus8 min

Your platform team started as an enabler, but teams now complain that every new capability request sits in the platform team's backlog for months. How do you diagnose and fix this?

IntermediatePlatform Engineering7 min

A script calling a flaky API retries immediately with a fixed 1-second delay. During a real outage, this made things worse. Why, and how should retry logic actually be designed?

IntermediatePython7 min

A log-processing script uses json.load() to parse a multi-gigabyte newline-delimited JSON log file and gets killed with an out-of-memory error. What's wrong, and how do you fix it?

IntermediatePython7 min

A Python automation script uses subprocess.run(f'kubectl get pod {pod_name}', shell=True) where pod_name comes from user input. What's actually wrong with this, and how do you fix it?

AdvancedPython7 min

A Python script that processes a large file works fine on your laptop but gets OOM-killed when run in a container with a memory limit. What's likely wrong, and how do you fix it?

IntermediatePythonDocker8 min

A security audit finds 340 service accounts, and nobody can say for certain which ones are still in use. How do you find and safely remove the dead ones without breaking production?

AdvancedSecurity8 min

A critical production system is accessed via one shared 'admin' account used by six engineers, with no individual audit trail. How do you fix this?

IntermediateSecurity7 min

After a change to your SSO/identity provider configuration, nobody — including admins — can log into any connected system. How do you get back in, and how do you diagnose the actual cause?

AdvancedSecurity8 min

A developer just committed a live database password directly into a public GitHub repository. It's been merged and pushed. What do you do, in order?

AdvancedSecurityGitHub10 min

Your on-call engineers are getting paged 15+ times a week, and they've started treating every page as probably-not-urgent before even looking. How do you fix this?

AdvancedSRE8 min

terraform plan shows changes to a resource you didn't touch, and the diff traces back to a data source — what's actually happening?

AdvancedTerraform7 min

Converting a resource block from count to for_each caused Terraform to want to destroy and recreate every instance — why, and how do you migrate safely?

AdvancedTerraform7 min

terraform init fails with a provider checksum mismatch after a teammate on a different OS committed the lock file — what's happening?

IntermediateTerraform6 min

terraform plan fails with Error acquiring the state lock — how do you diagnose whether it's a genuine concurrent run or a stale lock?

IntermediateTerraform6 min

A Terraform pipeline that worked yesterday suddenly fails today with no code changes — how does an unpinned provider version cause this?

IntermediateTerraform6 min

terraform plan shows a production RDS instance will be destroyed and recreated after a change you thought was trivial. What would you investigate before running apply?

AdvancedTerraformAWS10 min

You get paged with just "the app is slow" and no other context. Walk through your actual troubleshooting methodology before you touch anything.

IntermediateObservability8 min

A Kubernetes manifest that looked correct in the editor fails to apply with a cryptic YAML parsing error. How do you systematically debug indentation-related YAML errors?

BeginnerYAML6 min

A CI config field intended as the string "no" silently breaks everything downstream. Why does this happen with YAML specifically, and how do you avoid this whole class of bug?

IntermediateYAML6 min