>DevOps Interview KB

Architecture Interview Questions — All Categories

98 questions across 27 categories

Designing systems under real constraints: requirements, trade-offs, and the reasoning behind a specific design choice.

A production workload needs persistent storage, and its pods may be rescheduled to different nodes — how do you design storage so data survives that?

BeginnerKubernetes6 min

How would you set up alerting to catch a CrashLoopBackOff-class issue before it reaches production traffic, rather than discovering it via a user-facing outage?

IntermediateKubernetesPrometheus7 min

How would you decide between a Deployment, a StatefulSet, and a DaemonSet for three different real services (a stateless API, a database, a node agent)?

IntermediateKubernetes6 min

How would you design a Job for a task that must never run twice, even if a pod fails partway through (e.g., a billing charge)?

ExpertKubernetes8 min

How would you safely roll out a breaking change to a DaemonSet running a critical node-level agent across a large production cluster?

AdvancedKubernetes7 min

Your on-call team is getting paged 40+ times a week and starting to ignore alerts. How would you redesign the alerting strategy to fix that without missing real incidents?

AdvancedPrometheusMonitoring10 min

Your Prometheus storage costs are growing unsustainably from keeping raw-resolution metrics for a year. How would you design a retention and downsampling strategy to control costs?

AdvancedPrometheus7 min

An application currently relies on sticky sessions (a user's requests always route to the same backend) to work correctly. Why is this considered an anti-pattern, and how would you actually remove the dependency?

IntermediateNetworking7 min

How would you design a centralized log aggregation pipeline for a fleet of services, from collection through to searchable storage?

AdvancedLogging8 min

How would you choose the right SLIs for a new service, rather than defaulting to generic latency and error rate for everything?

AdvancedPrometheus7 min

How would you design a structured logging convention for a service, so logs are actually usable for debugging production incidents?

IntermediateLogging7 min

How would you design a 'break-glass' emergency access process that lets an engineer bypass normal approval during a critical incident, without that becoming a permanent backdoor around your access controls?

ExpertSecurity8 min

You're designing the IAM/role structure for a shared platform used by 12 different teams. How do you avoid both 'everyone is admin' and a role-request bottleneck that blocks every team on you?

AdvancedSecurity8 min

How would you design credential architecture so a hardcoded secret, if it happens again, has a much smaller blast radius?

AdvancedSecurity8 min

How would you design an incident commander role/process for a company that currently has no formal structure — incidents are just 'whoever notices it first tries to fix it'?

IntermediateSRE7 min

How would a self-service CI/CD platform handle a team that's on a fundamentally different tech stack the golden path doesn't cover well?

ExpertCI/CD8 min

How would the self-service CI/CD platform design change if the organization had 1,000 teams instead of 100?

ExpertCI/CD8 min

Design an alerting and on-call paging system for a company running services across 3 regions, where the paging system itself must not go down along with the region it's monitoring.

ExpertObservabilitySRE14 min

Design a centralized logging platform for an organization running roughly 500 microservices across multiple Kubernetes clusters, where engineers currently can't find logs during incidents.

ExpertKubernetesObservability14 min

Design a centralized secrets management system for an org currently scattering credentials across env vars, config files, and CI/CD tool stores, with no consistent rotation or audit trail.

ExpertSecurityDevSecOps14 min

Design a deployment orchestration system that lets any of your organization's 200 services safely adopt canary or blue-green deployments, without every team building their own rollout automation from scratch.

ExpertKubernetesPlatform Engineering14 min

Design a self-service CI/CD platform for an engineering org with roughly 100 teams, each owning multiple services, without a central platform team becoming a bottleneck.

ExpertCI/CDKubernetesPlatform Engineering15 min

Your team has copy-pasted the same VPC Terraform configuration into six different repositories. Design a module structure and versioning strategy to fix that.

IntermediateTerraform10 min

A corrupted or accidentally-overwritten Terraform state file threatens to make you lose track of an entire environment's real infrastructure — how would you design for recoverability?

AdvancedTerraform7 min