Advanced DevOps Interview Questions
194 questions
How would you handle a data transformation that genuinely needs to see the whole dataset at once, like a global sort, when streaming isn't an option?
AdvancedPython8 min
You find a piece of a legacy pipeline that seems actively dangerous — overly broad credentials, say — but nobody can explain why it's configured that way. What do you do?
AdvancedSecurity7 min
How would you decide whether to gradually refactor a legacy pipeline versus rewriting it from scratch?
AdvancedCI/CD7 min
You inherit ownership of a system whose fundamental architecture you think was the wrong call — not a small detail, a core design decision. How do you handle this, given the system is already in production?
AdvancedDevops7 min
You have two weeks to ship a solution, and neither option available to you is actually good — one is fast but creates real tech debt, the other is more correct but won't make the deadline. How do you approach this?
AdvancedDevops7 min
How would you design a recurring privileged-access review that catches stale access at scale without becoming a rubber-stamp exercise nobody takes seriously?
AdvancedSecurity7 min
You're designing the IAM/role structure for a shared platform used by 12 different teams. How do you avoid both 'everyone is admin' and a role-request bottleneck that blocks every team on you?
AdvancedSecurity8 min
A security audit finds 340 service accounts, and nobody can say for certain which ones are still in use. How do you find and safely remove the dead ones without breaking production?
AdvancedSecurity8 min
After a change to your SSO/identity provider configuration, nobody — including admins — can log into any connected system. How do you get back in, and how do you diagnose the actual cause?
AdvancedSecurity8 min
How would you design credential architecture so a hardcoded secret, if it happens again, has a much smaller blast radius?
AdvancedSecurity8 min
A developer just committed a live database password directly into a public GitHub repository. It's been merged and pushed. What do you do, in order?
AdvancedSecurityGitHub10 min
A single catastrophic incident burns an entire quarter's error budget in one day. Does the error budget policy still apply the same way?
AdvancedSRE7 min
How do you handle a service with multiple different user-facing SLIs — latency, availability, correctness — that might have conflicting error budget states at the same time?
AdvancedSRE7 min
Your on-call engineers are getting paged 15+ times a week, and they've started treating every page as probably-not-urgent before even looking. How do you fix this?
AdvancedSRE8 min
How would you measure whether a self-service CI/CD platform is actually succeeding, beyond just "teams are using it"?
AdvancedCI/CD7 min
How would you structure a shared Terraform module differently if its consuming repos were owned by teams with different release cadences and risk tolerances?
AdvancedTerraform7 min
How would you handle a breaking change to a shared Terraform module that genuinely needs to reach all consumers within a specific timeframe, like a security fix?
AdvancedTerraform7 min
terraform plan shows changes to a resource you didn't touch, and the diff traces back to a data source — what's actually happening?
AdvancedTerraform7 min
Converting a resource block from count to for_each caused Terraform to want to destroy and recreate every instance — why, and how do you migrate safely?
AdvancedTerraform7 min
A corrupted or accidentally-overwritten Terraform state file threatens to make you lose track of an entire environment's real infrastructure — how would you design for recoverability?
AdvancedTerraform7 min
How would you design a CI/CD pipeline to automatically block a Terraform apply that would destroy a production database?
AdvancedTerraform8 min
How does Terraform's create_before_destroy interact with resources that have unique naming constraints, like a fixed identifier?
AdvancedTerraform7 min
A single Terraform state file managing an entire environment's infrastructure takes 10 minutes to plan and any change risks touching everything — how would you split it?
AdvancedTerraform8 min
Using terraform_remote_state to reference another team's outputs works, but creates a tight coupling that breaks when they change their state — what's the alternative?
AdvancedTerraform7 min