>DevOps Interview KB

Staff / Principal Interview Questions

228 questions across 34 categories

Expert-level and system-design questions — the scope, trade-offs, and organizational reasoning expected at the most senior technical levels.

Why is it specifically dangerous to use self-hosted GitHub Actions runners on a public repository, in a way that doesn't apply to a private repository?

AdvancedGitHub Actions7 min

A team wants to auto-merge every Dependabot PR that passes CI, to reduce the toil of manually reviewing hundreds of dependency bumps. What's the actual risk, and how would you design this safely?

AdvancedGitHubDevSecOps8 min

You have a monorepo where a single commit might touch 1 service or 15. How would you use GitLab's dynamic child pipelines so you only run CI for the services that actually changed?

AdvancedGitLab CI/CD8 min

A compliance requirement mandates that every commit merged into a regulated project be cryptographically signed and traceable to a verified author. How would you enforce this in GitLab?

AdvancedGitLab CI/CD7 min

How does Argo CD's ignoreDifferences configuration interact with its automated self-healing (selfHeal) feature?

AdvancedArgo CDKubernetes6 min

A mutating webhook injects configuration into your resources. When should that injected config actually be tracked in Git instead of ignored via ignoreDifferences?

AdvancedArgo CDKubernetes7 min

GitOps means Git is the source of truth for everything deployed, but you obviously can't commit plaintext secrets to Git. How do you actually reconcile this?

AdvancedGitOpsSecurity8 min

How does Helm's --atomic flag change the stuck-release failure mode, and what window of risk does it not actually cover?

AdvancedHelm6 min

How would you share common templates (labels, resource boilerplate) across many microservice charts without copy-pasting them into every chart?

AdvancedHelm7 min

How would you design a deploy pipeline so a killed CI job can never leave a Helm release ambiguously stuck?

ExpertHelmKubernetes8 min

A Helm pre-upgrade hook Job doesn't run before the Deployment update it's supposed to precede — why, and how do you fix the ordering?

AdvancedHelm7 min

A helm rollback succeeds but the application still behaves like the newer version — why might rollback not fully revert the deployed state?

AdvancedHelm7 min

How would you design scheduled Terraform drift detection so it alerts the right team without becoming noise nobody reads?

AdvancedTerraform8 min

You need to make a breaking change to a shared infrastructure module used by 30 different projects. How do you version and roll this out without breaking everyone simultaneously?

AdvancedTerraform8 min

You've extracted common pipeline logic into a Jenkins shared library used by 40+ Jenkinsfiles. How do you version it so you can improve the library without breaking every pipeline that depends on it?

AdvancedJenkins7 min

How would Kaniko or rootless BuildKit avoid the Docker socket mounting problem entirely, and what do you give up by switching to them?

AdvancedJenkinsDocker7 min

Two mutating webhooks both touch a pod's containers field, and the final result isn't what either webhook intended alone — how do you diagnose and resolve this?

ExpertKubernetes8 min

A new mutating webhook accidentally intercepted kube-system pod creation and broke core cluster components — how would you design its scoping to prevent this?

AdvancedKubernetes7 min

Every pod creation cluster-wide suddenly starts failing with an admission webhook TLS error — what happened, and how do you recover quickly?

ExpertKubernetes8 min

An admission webhook's failurePolicy is set to Fail — what happens if the webhook itself becomes unavailable, and why might that be the wrong default?

AdvancedKubernetes7 min

A validating webhook is rejecting pod creations that look completely valid — how do you diagnose what it's actually objecting to?

AdvancedKubernetes7 min

kubectl apply --dry-run=server behaves unexpectedly for requests a webhook handles — what does a webhook's sideEffects field have to do with it?

AdvancedKubernetes6 min

The API server responds slowly to all requests — how do you determine whether etcd, the API server, or something else is the bottleneck?

ExpertKubernetes9 min

How would you design backup/DR for etcd, and what's actually recoverable from a snapshot versus what isn't?

AdvancedKubernetesEtcd8 min