>DevOps Interview KB

Staff / Principal Interview Questions

228 questions across 34 categories

Expert-level and system-design questions — the scope, trade-offs, and organizational reasoning expected at the most senior technical levels.

A lifecycle policy quietly moved a set of blobs to Archive tier months ago, and now someone urgently needs one back within the hour. What are your actual options?

AdvancedAzureBlob Storage7 min

A storage account's firewall was locked down to specific VNets for security, and now diagnostic logging and a few other first-party Azure integrations silently stopped working. What's actually happening?

AdvancedAzureBlob Storage6 min

A compliance requirement says certain records must be provably unmodifiable for seven years. How would you design this with Blob Storage's immutability features, and what breaks if you choose the wrong policy type?

AdvancedAzureBlob Storage7 min

A script has 'set -e' at the top, but a failing command inside a function doesn't stop the script the way you'd expect. Why not, and how do you fix it?

AdvancedBash7 min

A team wants zero-downtime deploys with fast rollback for a high-traffic payments API. Would you recommend blue-green or canary deployment, and why?

AdvancedCI/CDKubernetes10 min

How would you design the metrics and thresholds that gate automatic canary promotion versus automatic rollback?

AdvancedKubernetes8 min

How does a database schema migration change your deployment strategy for either blue-green or canary — what breaks if you don't account for it?

AdvancedDatabases8 min

Your architecture serves a large volume of data out to end users and other services, and data transfer/egress costs have become one of your largest line items. What architectural decisions actually reduce this?

AdvancedAWS8 min

Leadership wants to adopt a multi-cloud strategy specifically to reduce vendor lock-in and negotiate better pricing. What are the real cost trade-offs they might not be accounting for?

AdvancedAWSAzureGCP7 min

How would you design the data layer differently for an active-active multi-region architecture versus an active-passive (failover) one?

ExpertAWS10 min

In a real system with dozens of services, how would you decide which specific ones actually need multi-region treatment versus which can safely stay single-region?

AdvancedAWS8 min

What's the actual difference between designing for high availability across Availability Zones versus across Regions, and when do you actually need multi-region instead of just multi-AZ?

AdvancedAWS8 min

What actually changed between cgroups v1 and v2, and why did some container resource limit behavior change when a host migrated to a v2-only kernel/distro?

AdvancedContainersLinux7 min

A container keeps getting killed, and both the container's own memory limit and the host's overall available memory look potentially responsible. How do you determine which one actually caused it?

AdvancedContainersLinux7 min

Your containerized application's process count keeps growing over time, even though nothing appears to be leaking connections or memory. What's likely happening, and why is it specific to running in a container?

AdvancedContainersLinux7 min

What would you do differently for a large-table backfill migration if the table were actively receiving high write throughput during the migration window?

ExpertPostgresql8 min

How would you monitor a batched backfill in production to know it's progressing safely, not falling behind, and not causing replication lag?

AdvancedPostgresql7 min

How would the safe-migration approach for adding a NOT NULL column to a large table differ on MySQL versus PostgreSQL, given their different online-DDL capabilities?

AdvancedMysqlPostgresql8 min

How do you safely add a NOT NULL column to a large production table without locking it for the duration of a slow backfill?

AdvancedPostgresql9 min

A user updates their profile, immediately reloads the page, and sees their old data — but only sometimes. What's causing this, and how do you fix it?

AdvancedDatabases8 min

How would you design the actual mechanism that decides whether a given database query goes to the primary or a read replica, for an application that wasn't originally built with this split in mind?

AdvancedDatabases8 min

During a network partition, both your old primary and a newly promoted replica briefly accept writes at the same time. How does automatic failover cause this, and how do you design against it?

ExpertDatabases8 min

What's the actual trade-off between synchronous and asynchronous database replication, and how would you decide which to use for a payments system versus an analytics dashboard?

AdvancedDatabases7 min

How would you design artifact signing into your CI/CD pipeline so a deployed container image or binary can be verified as genuinely coming from your build, not tampered with in transit?

AdvancedDevSecOps7 min