Expert DevOps Interview Questions
34 questions
A CSI driver upgrade causes new attach operations to fail while already-mounted volumes keep working — how do you investigate, and how would you roll this out more safely next time?
ExpertKubernetes8 min
How would you design a Job for a task that must never run twice, even if a pod fails partway through (e.g., a billing charge)?
ExpertKubernetes8 min
How would you design a 'break-glass' emergency access process that lets an engineer bypass normal approval during a critical incident, without that becoming a permanent backdoor around your access controls?
ExpertSecurity8 min
How would a self-service CI/CD platform handle a team that's on a fundamentally different tech stack the golden path doesn't cover well?
ExpertCI/CD8 min
How would the self-service CI/CD platform design change if the organization had 1,000 teams instead of 100?
ExpertCI/CD8 min
Design an alerting and on-call paging system for a company running services across 3 regions, where the paging system itself must not go down along with the region it's monitoring.
ExpertObservabilitySRE14 min
Design a centralized logging platform for an organization running roughly 500 microservices across multiple Kubernetes clusters, where engineers currently can't find logs during incidents.
ExpertKubernetesObservability14 min
Design a centralized secrets management system for an org currently scattering credentials across env vars, config files, and CI/CD tool stores, with no consistent rotation or audit trail.
ExpertSecurityDevSecOps14 min
Design a deployment orchestration system that lets any of your organization's 200 services safely adopt canary or blue-green deployments, without every team building their own rollout automation from scratch.
ExpertKubernetesPlatform Engineering14 min
Design a self-service CI/CD platform for an engineering org with roughly 100 teams, each owning multiple services, without a central platform team becoming a bottleneck.
ExpertCI/CDKubernetesPlatform Engineering15 min