>DevOps Interview KB

DevOps Engineer Interview Questions

372 questions across 38 categories

The bulk of a working DevOps engineer's day-to-day: hands-on tooling, common troubleshooting, and the practical judgment calls that come up once you're actually operating systems.

A critical production system is accessed via one shared 'admin' account used by six engineers, with no individual audit trail. How do you fix this?

IntermediateSecurity7 min

A third-party vendor's contract is up for renewal, and their integration still has the broad access it was granted two years ago during initial setup. How do you review and right-size it before renewing?

IntermediateSecurity7 min

How would your incident response to a hardcoded secret committed to a repository differ if the repository were private rather than public?

IntermediateGit6 min

What's the trade-off between rewriting Git history to remove a committed secret versus simply rotating it and leaving the now-worthless value in history?

IntermediateGit6 min

How would you choose the rolling window length for an error budget — 7 days, 30 days, or 90 days — and what does that choice actually trade off?

IntermediateSRE6 min

How would you go about setting an SLO and error budget for a service that's never had one, and what should actually happen once that error budget runs out?

IntermediateSRE9 min

How can a postmortem process be genuinely blameless while still holding people accountable for mistakes? Doesn't 'blameless' just mean nobody's responsible for anything?

IntermediateSRE7 min

How would you design an incident commander role/process for a company that currently has no formal structure — incidents are just 'whoever notices it first tries to fix it'?

IntermediateSRE7 min

One engineer on your team has quietly been covering far more on-call shifts than everyone else because they're 'just better at it' and others avoid volunteering. How do you fix this before it causes burnout?

IntermediateSRE6 min

Your team resolves most incidents quickly because a few senior engineers just 'know' what to do. What's actually wrong with that, and how would you convert that knowledge into runbooks without slowing the seniors down?

IntermediateSRE6 min

Your SRE team spends most of its time on manual, repetitive operational work and has no time left for the reliability engineering they were hired to do. How would you identify and reduce this 'toil' systematically?

IntermediateSRE7 min

What would make you choose a private Terraform module registry over Git-tag-based module sourcing, or vice versa?

IntermediateTerraform6 min

Your team has copy-pasted the same VPC Terraform configuration into six different repositories. Design a module structure and versioning strategy to fix that.

IntermediateTerraform10 min

A production database was created manually through the AWS console years ago — how would you bring it under Terraform management without recreating it?

IntermediateTerraform7 min

terraform init fails with a provider checksum mismatch after a teammate on a different OS committed the lock file — what's happening?

IntermediateTerraform6 min

How would you deploy the same Terraform resource type to multiple AWS regions within a single configuration, using provider aliases?

IntermediateTerraform6 min

Should you use Terraform workspaces or separate directories/state files to manage dev, staging, and production — what's the actual trade-off?

IntermediateTerraform7 min

How would you make it structurally difficult for someone to accidentally run terraform destroy against a production state?

IntermediateTerraform6 min

terraform plan fails with Error acquiring the state lock — how do you diagnose whether it's a genuine concurrent run or a stale lock?

IntermediateTerraform6 min

A Terraform pipeline that worked yesterday suddenly fails today with no code changes — how does an unpinned provider version cause this?

IntermediateTerraform6 min

What's the difference between state drift and a genuine configuration change in Terraform, and how does -refresh-only help distinguish them?

IntermediateTerraform6 min

What's the difference between terraform state rm and just deleting a resource block from configuration — when would you use state rm specifically?

IntermediateTerraform6 min

How do you decide when to escalate or pull in another team during an incident, versus continuing to investigate solo?

IntermediateDevops6 min

How do you balance "mitigate first" against the risk that a rollback or failover masks the real problem, letting it recur later?

IntermediateObservability6 min