How would you prevent a new workload from ever being built directly on a static IAM user again, at an organizational level rather than case by case?
Short Answer
Combine a preventive guardrail (a Service Control Policy or IAM permissions boundary that blocks creating new IAM users with programmatic access outside an explicit exception process) with a detective control (scheduled scanning for IAM users with active access keys, alerting when one appears) and a paved-road default (make the role-based path the easy, documented default for provisioning new workload access, not an extra step people skip under deadline pressure) — prevention alone tends to get bypassed under pressure without a genuinely easier default path available.
Detailed Explanation
Fixing one instance of a workload built on a static IAM user solves that instance; it doesn't address why it happened, which is usually some combination of "the role-based path wasn't the obvious default" and "there was no guardrail actively stopping the shortcut." A durable fix addresses both.
Preventive control: an AWS Organizations Service Control Policy (SCP) can deny iam:CreateAccessKey (or iam:CreateUser entirely) account-wide, with an explicit exception mechanism (e.g. a specific tag or a separate break-glass account) for the rare legitimate case — like the third-party-vendor scenario requiring static keys — where it's genuinely unavoidable. This makes the shortcut structurally unavailable by default rather than merely discouraged in a wiki page nobody reads under deadline pressure.
Detective control: even with a preventive SCP, drift and exceptions happen — a scheduled scan (AWS Config rule, or a custom Lambda on a schedule) checking for IAM users with active access keys, reporting findings to a channel the platform/security team actually watches, catches anything that slips through or predates the guardrail. This is the safety net for the cases prevention alone doesn't cover.
Paved-road default: guardrails alone tend to just create friction if there's no easier alternative offered alongside them — if provisioning a proper IAM role for a new workload is a slower, less-documented process than "just create a user with an access key," people will find a way around the guardrail (or petition for an exception) under time pressure. Pairing the restriction with genuinely easy self-service role provisioning (a Terraform module, a service-catalog-style internal tool, clear documentation with copy-pasteable examples) removes the incentive to route around the guardrail in the first place, since the correct path is no longer the harder one.
The combination matters because each piece covers a different failure mode: prevention stops the common case outright, detection catches what prevention misses (legitimate exceptions, older resources, policy gaps), and a good default removes the pressure that causes people to seek workarounds for the prevention in the first place.
Interview Follow-Up Questions
- How would you design the exception process for the SCP so legitimate cases (like a vendor requiring static keys) aren't blocked indefinitely by bureaucracy?
- What would you do if the detective scan found dozens of pre-existing IAM users with active keys across many accounts — how would you prioritize remediation?
- How would you measure whether this governance change actually worked, six months later?
Key Takeaways
- Fixing one instance doesn't prevent recurrence — that requires addressing why it happened structurally, not just this one case.
- A preventive SCP denying IAM user/access-key creation by default (with an explicit exception path) makes the shortcut structurally unavailable.
- A detective scan for IAM users with active access keys catches what prevention misses or predates.
- Guardrails without an easier paved-road alternative just create pressure to bypass them — pair restriction with genuinely easy self-service role provisioning.
References
Related Questions
- You inherit an EC2 workload that authenticates to AWS using an IAM user with AdministratorAccess. How would you migrate it to least-privilege access without causing an outage?Advanced
- How would you make the case for the cost of a separate AWS account, if leadership pushes back on the added complexity?SuggestedIntermediate
- How does the approach to workload identity and least privilege differ if a workload runs on ECS or Lambda instead of EC2?SuggestedIntermediate
Last updated August 21, 2026 · Last reviewed August 21, 2026