Staff / Principal Interview Questions
228 questions across 34 categories
Expert-level and system-design questions — the scope, trade-offs, and organizational reasoning expected at the most senior technical levels.
How would you audit an existing large Ansible playbook to find all the places check mode's coverage is actually incomplete?
AdvancedAnsible7 min
How would you build confidence in a playbook's check-mode output when it has to use command/shell for something with no equivalent module?
AdvancedAnsible7 min
How would you retrofit idempotency checks into an existing large Ansible playbook full of command/shell tasks, without rewriting every task at once?
AdvancedAnsible8 min
How would you verify a purpose-built module replacement produces the exact same end state as the shell command it replaced, in a repeatable way?
AdvancedAnsible7 min
How do Argo CD's resource hooks compare to Helm's pre-install/pre-upgrade hooks, given Argo CD can also deploy Helm charts directly?
AdvancedArgo CDHelm7 min
How would you distinguish 'the migration itself is broken' from 'this was a transient failure worth retrying,' in terms of alerting design?
AdvancedArgo CD7 min
How do you control the order Argo CD applies resources within a single Application, e.g. making sure a database migration Job completes before the Deployment that depends on it rolls out?
AdvancedArgo CDKubernetes8 min
What would break if a Helm chart's hooks specifically relied on helm rollback semantics, and how would that manifest when deployed via Argo CD instead?
AdvancedArgo CDHelm7 min
How would you handle a migration that needs to run exactly once across an entire fleet of clusters, not just once per cluster?
ExpertArgo CDKubernetes8 min
How would you handle an Argo CD migration Job that should run on every sync versus one that should only run when the migration itself actually changed?
AdvancedArgo CDKubernetes7 min
Why might Argo CD's sync-wave model be considered more expressive than Helm's own hook-weight system for controlling ordering?
AdvancedArgo CDHelm6 min
What happens to a PreSync hook Job that succeeded on a previous sync, if the overall sync is retried after a later hook fails?
AdvancedArgo CDKubernetes6 min
How would you verify a Helm chart's hooks behave correctly under Argo CD before migrating a production chart from helm install to GitOps?
AdvancedArgo CDHelm7 min
How would you audit whether an ECS task role or Lambda execution role is actually scoped tightly, versus just copy-pasted from a broader existing role?
AdvancedAWS7 min
What CloudTrail-based alerting would you specifically set up for a narrowly-scoped static-key IAM user, and how would you tune it to avoid false positives?
AdvancedAWS7 min
How would you design the exception process for an SCP blocking IAM user creation, so legitimate cases aren't blocked indefinitely by bureaucracy?
AdvancedAWS7 min
How does the workload-identity comparison extend to EKS, where pod identity is yet another mechanism (IRSA or Pod Identity)?
ExpertAWSEKSKubernetes8 min
What's the mechanism difference between how EC2's IMDS delivers credentials versus how Lambda delivers them to a function's environment?
AdvancedAWSLambda6 min
You inherit an EC2 workload that authenticates to AWS using an IAM user with AdministratorAccess. How would you migrate it to least-privilege access without causing an outage?
AdvancedAWSIAMEC212 min
How would you prevent a new workload from ever being built directly on a static IAM user again, at an organizational level rather than case by case?
AdvancedAWS8 min
A detective scan finds dozens of pre-existing IAM users with active keys across many accounts. How would you prioritize remediation?
AdvancedAWS7 min
How would you design automated credential rotation to handle the overlap window safely, so the application never experiences an auth failure?
AdvancedAWS7 min
A third-party application only supports static AWS access keys and can't use an instance profile or role. How do you handle this without abandoning least privilege entirely?
AdvancedAWS8 min
Even after mitigating cold starts, a small amount of irreducible tail latency remains. How would you design the caller's retry behavior to handle that remaining tail?
AdvancedAWSLambda8 min