Expert DevOps Interview Questions
34 questions
How would you handle a migration that needs to run exactly once across an entire fleet of clusters, not just once per cluster?
ExpertArgo CDKubernetes8 min
How does the workload-identity comparison extend to EKS, where pod identity is yet another mechanism (IRSA or Pod Identity)?
ExpertAWSEKSKubernetes8 min
How would you scale S3 public-exposure alerting for an organization with hundreds of AWS accounts, where per-account Config rules alone don't scale operationally?
ExpertAWS8 min
How would you design secret rotation for a live, high-traffic application backed by Key Vault, so rotating a credential never causes a production outage?
ExpertAzureKey Vault8 min
How would you design the data layer differently for an active-active multi-region architecture versus an active-passive (failover) one?
ExpertAWS10 min
What would you do differently for a large-table backfill migration if the table were actively receiving high write throughput during the migration window?
ExpertPostgresql8 min
During a network partition, both your old primary and a newly promoted replica briefly accept writes at the same time. How does automatic failover cause this, and how do you design against it?
ExpertDatabases8 min
A widely-used open-source dependency your organization relies on is publicly disclosed as compromised — a malicious backdoor was found in a recent release. How do you respond?
ExpertDevSecOps9 min
A function migrated from Gen 1 to Gen 2 started returning occasionally-wrong results under load — what changed, and how do you fix it?
ExpertGCPCloud Functions8 min
How would you design a fast, event-driven alerting system to catch an accidentally-public Cloud Storage bucket within minutes, GCP-native?
ExpertGCPCloud Storage8 min
How would the OIDC-based deploy design change for a monorepo where multiple independent deploy targets live in one repository?
ExpertGitHub ActionsAWS8 min
How would you design a deploy pipeline so a killed CI job can never leave a Helm release ambiguously stuck?
ExpertHelmKubernetes8 min
Two mutating webhooks both touch a pod's containers field, and the final result isn't what either webhook intended alone — how do you diagnose and resolve this?
ExpertKubernetes8 min
Every pod creation cluster-wide suddenly starts failing with an admission webhook TLS error — what happened, and how do you recover quickly?
ExpertKubernetes8 min
The API server responds slowly to all requests — how do you determine whether etcd, the API server, or something else is the bottleneck?
ExpertKubernetes9 min
A team wants to run HPA and VPA on the same Deployment for both CPU and memory — what breaks if you're not careful, and how do you combine them safely?
ExpertKubernetes8 min
A Deployment's pods get evicted during scale-up because new nodes take too long to become ready — how would you close that gap?
ExpertKubernetes8 min
Runtime security tooling alerts that a specific pod is exhibiting behavior consistent with compromise — walk through your immediate containment response.
ExpertKubernetes8 min
How would you design a policy requiring every image deployed to a cluster be cryptographically signed, and what does that actually protect against?
ExpertKubernetes8 min
A security scan found the kubelet's API port reachable without authentication on some nodes — what can an attacker actually do with that, and how do you fix it?
ExpertKubernetes8 min
A CRD needs a breaking schema change, but existing custom resources and consumers depend on the old shape — how do you version a CRD safely?
ExpertKubernetes8 min
How would you migrate a cluster from one CNI plugin to another without a full cluster rebuild — what's actually risky about it?
ExpertKubernetes8 min
A pod's ServiceAccount token was found in a public repo — what's your incident response, and how do you reduce blast radius for next time?
ExpertKubernetes8 min
How would you troubleshoot a pod stuck Pending even though kubectl describe pod shows no scheduling errors at all?
ExpertKubernetes8 min