Senior DevOps Interview Questions
516 questions across 39 categories
Deeper trade-off and design questions — why a specific approach was chosen, what breaks at scale, and how to reason about a system you didn't build.
You've proposed a robust, well-architected solution to a problem, but your manager wants a quick, hacky fix instead because the problem is genuinely minor. Are they wrong, or are you over-engineering?
IntermediateDevops6 min
You inherit ownership of a system whose fundamental architecture you think was the wrong call — not a small detail, a core design decision. How do you handle this, given the system is already in production?
AdvancedDevops7 min
You've just joined a team responsible for critical infrastructure, and there's essentially no documentation — the previous engineers who understood it have all left. How do you approach getting up to speed safely?
IntermediateDevops6 min
Your organization depends heavily on an internal tool maintained by exactly one engineer, who's showing clear signs of burnout and has hinted they might leave. How do you handle this before it becomes a crisis?
IntermediateDevops6 min
Product leadership keeps deprioritizing infrastructure tech debt work in favor of features, and the debt is now visibly slowing down every new feature. How do you actually change this dynamic?
IntermediateDevops7 min
You have two weeks to ship a solution, and neither option available to you is actually good — one is fast but creates real tech debt, the other is more correct but won't make the deadline. How do you approach this?
AdvancedDevops7 min
How would you design a recurring privileged-access review that catches stale access at scale without becoming a rubber-stamp exercise nobody takes seriously?
AdvancedSecurity7 min
You're designing the IAM/role structure for a shared platform used by 12 different teams. How do you avoid both 'everyone is admin' and a role-request bottleneck that blocks every team on you?
AdvancedSecurity8 min
When would you actually choose manual credential rotation over fully automated rotation, given that automation is generally considered the more secure default?
IntermediateSecurity6 min
A security audit finds 340 service accounts, and nobody can say for certain which ones are still in use. How do you find and safely remove the dead ones without breaking production?
AdvancedSecurity8 min
A critical production system is accessed via one shared 'admin' account used by six engineers, with no individual audit trail. How do you fix this?
IntermediateSecurity7 min
After a change to your SSO/identity provider configuration, nobody — including admins — can log into any connected system. How do you get back in, and how do you diagnose the actual cause?
AdvancedSecurity8 min
A third-party vendor's contract is up for renewal, and their integration still has the broad access it was granted two years ago during initial setup. How do you review and right-size it before renewing?
IntermediateSecurity7 min
How would you design credential architecture so a hardcoded secret, if it happens again, has a much smaller blast radius?
AdvancedSecurity8 min
A developer just committed a live database password directly into a public GitHub repository. It's been merged and pushed. What do you do, in order?
AdvancedSecurityGitHub10 min
How would your incident response to a hardcoded secret committed to a repository differ if the repository were private rather than public?
IntermediateGit6 min
What's the trade-off between rewriting Git history to remove a committed secret versus simply rotating it and leaving the now-worthless value in history?
IntermediateGit6 min
A single catastrophic incident burns an entire quarter's error budget in one day. Does the error budget policy still apply the same way?
AdvancedSRE7 min
How would you choose the rolling window length for an error budget — 7 days, 30 days, or 90 days — and what does that choice actually trade off?
IntermediateSRE6 min
How would you go about setting an SLO and error budget for a service that's never had one, and what should actually happen once that error budget runs out?
IntermediateSRE9 min
How do you handle a service with multiple different user-facing SLIs — latency, availability, correctness — that might have conflicting error budget states at the same time?
AdvancedSRE7 min
Your on-call engineers are getting paged 15+ times a week, and they've started treating every page as probably-not-urgent before even looking. How do you fix this?
AdvancedSRE8 min
How can a postmortem process be genuinely blameless while still holding people accountable for mistakes? Doesn't 'blameless' just mean nobody's responsible for anything?
IntermediateSRE7 min
How would you design an incident commander role/process for a company that currently has no formal structure — incidents are just 'whoever notices it first tries to fix it'?
IntermediateSRE7 min