PRACTICE TRACK / 148 QUESTIONS

AWS Architect
Think it through.

Solutions architecture, IAM, VPC, compute, and resilience.

Choose a question, explain your approach, then reveal the supplied answer. Difficulty labels come from the existing question library.

148 questions

Answers stay closed until you choose to reveal them.

QUESTION 01AWS ArchitectEasy

What is the difference between an IAM role and an IAM user?

#
Reveal answer guidance

A user is a long-lived identity with permanent credentials, meant for a person or legacy app. A role has no permanent credentials — it's assumed temporarily, handing back short-lived STS credentials. Roles are the secure default for services (EC2, Lambda), cross-account access, and federated/OIDC logins because nothing long-lived is stored.

QUESTION 02AWS ArchitectEasy

What is the difference between a security group and a NACL?

#
Reveal answer guidance

A security group is a stateful firewall attached to an ENI/instance — return traffic is automatically allowed, and it has allow rules only. A NACL is a stateless firewall at the subnet level — you must allow both directions explicitly, it supports allow and deny rules, and rules are evaluated in numbered order. SGs are your primary control; NACLs are a coarse subnet-wide backstop.

QUESTION 03AWS ArchitectMedium

How do you give an EC2 instance access to S3 without hardcoding credentials?

#
Reveal answer guidance

Attach an IAM role to the instance via an instance profile, with a least-privilege policy scoped to the specific bucket/prefix. The SDK automatically retrieves temporary credentials from the instance metadata service (prefer IMDSv2). No keys are stored on disk and credentials rotate automatically.

QUESTION 04AWS ArchitectMedium

Explain how a request is routed in a typical public/private VPC setup.

#
Reveal answer guidance

Public subnets have a route to an Internet Gateway; resources there (e.g. an ALB or NAT gateway) get public IPs. Private subnets route outbound internet traffic through a NAT gateway in a public subnet, so instances can reach the internet but aren't directly reachable. Route tables per subnet, plus security groups and NACLs, define what can talk to what. App/DB tiers live in private subnets; only the load balancer is public.

QUESTION 05AWS ArchitectHard

How does an IAM policy evaluation decide allow vs deny?

#
Reveal answer guidance

Default is implicit deny. AWS gathers all applicable policies (identity, resource, permission boundaries, SCPs, session policies). Any explicit Deny anywhere wins immediately. Otherwise the request is allowed only if there's an explicit Allow AND it isn't blocked by a permission boundary or SCP. So effective permissions are the intersection of identity policy ∩ boundary ∩ SCP, minus any explicit deny.

QUESTION 06AWS ArchitectHard

An S3 bucket triggers a Lambda function on s3:ObjectCreated:*. The Lambda writes a message to SQS. How do you guarantee exactly-once processing when S3 event notifications are at-least-once and Lambda may retry?

#
Reveal answer guidance

S3 event notifications are inherently at-least-once — the same event can be delivered multiple times due to retries or replication events. Lambda may also retry on failure, compounding duplicates. To achieve effective exactly-once, you must implement idempotency in the consumer. The Lambda should extract the S3 object's ETag (or a combination of bucket+key+eventTime) and use it as a deduplication ID. If using SQS FIFO, set MessageDeduplicationId to the event's unique hash — SQS FIFO will reject duplicate messages within the 5-minute deduplication window. However, S3 Event Notifications do not support FIFO natively, so the Lambda must handle dedup by writing the event hash to DynamoDB with a TTL (e.g., 1 hour) using a conditional write (ConditionExpression: attribute_not_exists(PK)). If the conditional write fails, the Lambda skips processing. This pattern catches duplicates from both S3 retries and Lambda retries. Without deduplication, downstream systems may double-count inventory, charge customers twice, or create duplicate database records.

QUESTION 07AWS ArchitectHard

You have a DynamoDB table with a partition key of userId and a sort key of timestamp. Under heavy write load, you see throttling on a few partitions. Diagnose the root cause and design a mitigation strategy without full table rebuild.

#
Reveal answer guidance

This is a "hot partition" problem. The userId partition key causes all writes for a popular user to hit a single partition, exceeding the 1000 WCU/partition limit. Even if the table has 10000 WCU provisioned, throughput is divided across partitions — a single popular user can only consume 1000 WCU. Mitigations: (1) Add a suffix to the partition key — append a random digit (0–9) to userId during writes, and use a scatter-read pattern (Query with BETWEEN for all suffixes) during reads. This spreads writes across 10 partitions but increases read cost. (2) Use write sharding with a configurable shard count stored in a config table — clients compute userId + "-" + (hash(userId) % shardCount). (3) If the access pattern supports it, use DynamoDB Accelerator (DAX) to absorb read traffic, freeing read capacity for write-heavy partitions. (4) Switch to on-demand capacity, which scales per-partition throughput dynamically but may cost more. The core insight: partition key design must account for access distribution, not just uniqueness.

QUESTION 08AWS ArchitectMedium

What is the difference between an S3 Event Notification sent directly to a Lambda function versus routing through SQS first? When would you insert SQS into the pipeline?

#
Reveal answer guidance

Direct S3-to-Lambda invocation invokes the function asynchronously — S3 publishes the event, and Lambda queues it internally. However, Lambda has a concurrency limit (1000 by default per account) and an internal event queue with a 6-hour TTL. If the Lambda throttles, S3 retries but the retry window is limited (eventual retries up to 2 times). SQS decouples the producer from the consumer: S3 sends to SQS, and the Lambda polls from SQS. This provides: (1) a durable buffer (up to 14 days) so no events are lost during Lambda cold starts or throttles, (2) batch processing (up to 10 messages per invocation), (3) DLQ support for poison messages, and (4) decoupled retry policies — SQS retries based on a configurable redrivePolicy before discarding. Use direct invocation when latency is critical and the processing is lightweight. Insert SQS when the downstream is variable-throughput, needs batching, or requires guaranteed delivery with custom retry semantics.

QUESTION 09AWS ArchitectHard

A corporate network uses a Transit Gateway with multiple VPCs attached. VPC A (spoke) needs to reach VPC B's private RDS instance, but VPC B must not initiate connections back to VPC A. How do you configure route tables and security groups to enforce this?

#
Reveal answer guidance

Transit Gateway route tables control inter-VPC traffic. To enforce one-way communication, create separate route tables: associate VPC A with Route Table RT_A, and VPC B with Route Table RT_B. In RT_A, add a static route for VPC B's CIDR pointing to the VPC B attachment. Do NOT add a return route in RT_B for VPC A's CIDR — VPC B will not have a route back, so return traffic from RDS (which is a response to an established connection) still flows because AWS tracks connection state at the network interface level. However, RDS itself does not initiate connections. For defense in depth, add a security group on the RDS instance that allows inbound from VPC A's CIDR or the VPC A security group, and set egress to deny all in VPC A's security group. If VPC B has other resources that could be compromised and used as a pivot, also configure a network ACL on VPC B's subnets to deny return traffic to VPC A. Transit Gateway also supports cross-account attachments if VPCs are in different accounts.

QUESTION 10AWS ArchitectMedium

Lambda functions running behind an Application Load Balancer show high P99 latency during cold starts, but P50 is normal. What factors contribute to cold start latency, and how can you mitigate without always-on provisioned concurrency?

#
Reveal answer guidance

Cold start latency is the sum of: (1) sandbox allocation (container/download of code), (2) runtime initialization (Node.js parsing, Python import, JVM class loading), (3) handler code initialization (DB connections, config loading inside the handler closure). For ALB-triggered Lambdas, the ALB adds a 750ms idle timeout before recycling the sandbox, making cold starts more frequent under sparse traffic. Mitigations without provisioned concurrency: (1) Use SnapStart for Java/Python — Lambda snapshots the initialized runtime and resumes from snapshot (reducing cold start from seconds to <200ms). (2) Keep functions warm with a CloudWatch Events rule triggering the function every 5 minutes with a dummy event; branch on the event shape to skip real work. (3) Reduce deployment package size — use AWS SDK v3 with tree-shaking, avoid bundling large native modules. (4) Use the arm64 architecture (Graviton2) — lower cost and often faster initialization. (5) Move initialization code outside the handler for connection reuse across warm invocations. For ALB specifically, increase the idle timeout to 60s to keep the connection alive and reduce re-init.

QUESTION 11AWS ArchitectHard

An existing t3.medium EC2 instance must be migrated to a different subnet in the same VPC without changing its public IP, private IP, or elastic IP assignment. How do you achieve a zero-downtime migration?

#
Reveal answer guidance

In AWS, changing a subnet requires stopping the instance, detaching the ENI, and attaching a new ENI in the target subnet — which changes the private IP. To keep the same private IP, you must create a secondary ENI in the target subnet with the desired IP, then swap the primary ENI. However, Elastic IPs can be reassigned: detach the EIP from the old ENI and associate it with the new ENI. For zero-downtime, use a multi-step approach: (1) Provision a second ENI in the target subnet with the same private IP (possible only if the IP is not already in use), (2) Attach the second ENI as eth1, (3) Update the OS routing table to prefer the new ENI, (4) Reassociate the EIP to the new ENI, (5) Remove the old ENI. If the same private IP is impossible (IP belongs to the old subnet's CIDR range), you must use a different private IP and update DNS or use an internal ALB/NLB to abstract the IP. The better approach is to use an Application Load Balancer or NLB in front of the instance, making the instance's IP irrelevant for clients.

QUESTION 12AWS ArchitectHard

You have a Multi-AZ RDS PostgreSQL instance and a read replica in a different region. On the primary, a DROP TABLE is executed accidentally. Can the read replica be promoted to recover the table? What is the recovery strategy?

#
Reveal answer guidance

A cross-region read replica uses logical replication (streaming WAL), which applies all DDL/DML from the primary — including DROP TABLE. By the time you notice, the DROP has already been replayed on the replica, so promoting it gives you a database without the table. For point-in-time recovery (PITR), RDS automatically takes automated backups with transaction log retention. You can restore the primary instance to a point in time just before the DROP (usually within the last 35 days, default 7). Restore to a new instance, extract the table via pg_dump, then import it back to the primary. If the DROP was recent, AWS also supports backtracking for Aurora (not standard RDS). To prevent recurrence use: (1) sql_safe_updates in PostgreSQL via rds.force_autovacuum_logging_level, (2) IAM policy deny rds:DeleteDBInstance on the replica, (3) RDS event notifications for DDL changes, (4) logical replication slots with a secondary consumer that logs incoming WAL for replay audit. The read replica is for read scaling and disaster recovery, not a backup mechanism for DDL mistakes.

QUESTION 13AWS ArchitectMedium

CloudFront serves content from an S3 origin with RestrictBucketAccess enabled (OAI). But a user is still able to access the S3 object directly using the S3 URL if they guess the object key. How do you block direct S3 access without breaking CloudFront?

#
Reveal answer guidance

The OAI allows only CloudFront to access S3 via the bucket policy, but if the OAI is not properly configured, or the bucket policy permits s3:GetObject from Principal: "*", direct access works. To block direct S3 access: (1) Configure the bucket policy to deny all principals except the CloudFront OAI (or OAC for v2). Example: "Condition": {"StringEquals": {"AWS:SourceArn": "arn:aws:cloudfront::ACCOUNT:distribution/DIST_ID"}} with Principal: "*", Action: "s3:GetObject", and specifying the OAI ARN. (2) Ensure S3 Block Public Access settings are enabled. (3) Use an Origin Access Control (OAC) instead of OAI — OAC supports aws:SourceArn condition keys for stronger verification. (4) If users still access directly, they are likely getting the objects through a cached response — set Cache-Control: private or add a Referer header check via a CloudFront Lambda@Edge that signs a custom header checked by S3 via bucket policy. (5) Use pre-signed URLs for private content and set CloudFront to sign using trusted key groups.

QUESTION 14AWS ArchitectHard

An ECS service using Fargate launch type runs a task that mounts an EFS filesystem. Multiple tasks across AZs write to the same file concurrently, causing data corruption. How do you design for safe concurrent writes without rewriting the application?

#
Reveal answer guidance

EFS supports the NFSv4.1 protocol with file locking, but the application must use flock() or fcntl() advisory locking, and all tasks must run the same locking protocol. If the application does not use locks, you have three options: (1) Use EFS's GeneralPurpose performance mode with Bursting throughput, which does not help with locking — the app must be fixed. (2) Insert a sidecar container running a distributed lock service like Redis (ElastiCache) or ZooKeeper — the application obtains a lock from the sidecar before writing. This requires code changes. (3) Restructure the workload: instead of all tasks writing to the same file, partition the data by task ID and write to separate files, then merge at read time. This avoids concurrent writes entirely. (4) Use an S3 bucket as the shared store with S3's strong read-after-write consistency (since December 2020) — the application writes to S3 instead of EFS, using conditional writes (If-None-Match: * for new objects) to prevent overwrites. EFS is not designed for high-concurrency same-file writes — it provides NFS semantics, not distributed consensus. If the app cannot be modified, the simplest fix is to serialize writes through an SQS queue consumed by a single task.

QUESTION 15AWS ArchitectMedium

You encrypt an S3 bucket with SSE-KMS. When you enable S3 Inventory, the inventory reports show "Failed" for some objects. What is causing this, and how do you ensure inventory succeeds for KMS-encrypted objects?

#
Reveal answer guidance

S3 Inventory reads object metadata to generate reports. For SSE-KMS objects, the inventory service needs kms:Decrypt permission on the KMS key used to encrypt each object. If the KMS key policy or the IAM role executing inventory lacks kms:Decrypt or kms:GenerateDataKey, the inventory cannot read the object's metadata (it tries to decrypt the object's encryption context for the report). Additionally, if the objects are encrypted with different KMS keys (e.g., per-object customer keys), the inventory service must have permissions for every key. Solutions: (1) Grant the S3 Inventory service principal (s3-inventory.amazonaws.com) kms:Decrypt in the KMS key policy using the aws:SourceAccount and aws:SourceArn condition keys. (2) Use a single KMS key for the entire bucket, reducing permission complexity. (3) Verify the inventory is configured with the correct optional fields — requesting SSE-KMS-related fields may require additional permissions. (4) If the objects are in OUTPOSTS, inventory does not support SSE-KMS objects at all — you must use SSE-S3 or SSE-C.

QUESTION 16AWS ArchitectMedium

Your EC2 instance is running, but the application is inaccessible. What would you check first?

#
Reveal answer guidance

Evaluate the issue layer-by-layer: (1) Security Group & Network ACLs: Verify inbound rules allow traffic on the application port (e.g., 80/443) from the client's CIDR block. (2) Route Tables: Check if the subnet is associated with an Internet Gateway (IGW) route (0.0.0.0/0 -> igw-xxxx) for public ingress, or a NAT Gateway for private egress. (3) Local OS Listener: SSH in and run ss -ltnp or netstat to confirm the application service is listening on the correct interface (e.g., 0.0.0.0 vs 127.0.0.1) and port. (4) Systemd logs: Run journalctl -u app.service to check if the app crashed or failed startup diagnostics. (5) ALB target health: Verify the load balancer target status if it sits behind an ALB.

QUESTION 17AWS ArchitectMedium

Why would you choose an Auto Scaling Group over a single EC2 instance?

#
Reveal answer guidance

An ASG provides several key production benefits over a standalone instance: (1) High Availability & Self-Healing: If an instance becomes unhealthy or fails status checks, the ASG terminates it and launches a new one automatically across Availability Zones. (2) Dynamic Scaling: Horizontally scales capacity out or in based on traffic metrics (CPU, request count, target tracking) to handle sudden spikes. (3) Cost Optimization: Automatically matches compute capacity to demand, scaling down to a minimum count during off-peak hours and utilizing mixed instance policies (Spot and On-Demand). (4) Automated Rollouts: Integrates with Launch Templates to execute rolling updates (InstanceRefresh) with zero downtime.

QUESTION 18AWS ArchitectMedium

What happens when an EC2 instance fails a health check?

#
Reveal answer guidance

The response depends on the configuration: (1) Behind ALB: The target group marks the instance unhealthy and stops routing new connections to it, draining existing ones. (2) Inside ASG: The Auto Scaling Group waits for the specified HealthCheckGracePeriod (to allow boot/initialization), and if the instance fails its EC2 status check or ELB target health check, it terminates the instance and launches a replacement based on the current launch template. (3) Troubleshooting: To prevent loss of debug state on termination, you can use ASG Lifecycle Hooks to pause termination, or inspect system logs before the instance is deleted.

QUESTION 19AWS ArchitectMedium

How would you securely access a private EC2 instance?

#
Reveal answer guidance

Avoid using jump servers with open port 22. Instead, use: (1) AWS Systems Manager (SSM) Session Manager: Recommended standard. Requires the SSM Agent running on the instance and an IAM Role with AmazonSSMManagedInstanceCore permissions. Access is authenticated via IAM and logged, with no inbound port 22 open. (2) EC2 Instance Connect Endpoint (EICE): Allows SSH tunneling using the AWS CLI or Console via a managed VPC endpoint. (3) Bastion Jump Host: Hardened instance in a public subnet, restricted by security groups to allow SSH from a specific developer IP address only.

QUESTION 20AWS ArchitectMedium

EBS vs EFS — when would you use each?

#
Reveal answer guidance

(1) EBS (Elastic Block Store): Block storage attached to a single EC2 instance (ReadWriteOnce). Best for performance-sensitive, transactional, or low-latency workloads like database storage (gp3/io2 Block Express for databases). (2) EFS (Elastic File System): Serverless POSIX-compliant shared filesystem (ReadWriteMany) that can be mounted concurrently by thousands of compute instances (EC2, ECS, EKS) across multiple Availability Zones. Best for shared application assets, configuration directories, and parallel pipelines needing concurrent file access.

QUESTION 21AWS ArchitectEasy

Why should an application use an IAM Role instead of Access Keys?

#
Reveal answer guidance

IAM Roles provide superior security: (1) Temporary Credentials: Roles use the Security Token Service (STS) to generate credentials that rotate automatically every 1 to 12 hours, eliminating the risk of compromised long-lived keys. (2) No Stored Secrets: Applications do not need access key secrets stored on disk or environment variables. (3) Scoped Governance: AWS dynamically provisions metadata credentials to EC2 instances, ECS task execution, or Lambda execution environments. (4) Audit trails: Role assumptions are recorded in CloudTrail.

QUESTION 22AWS ArchitectMedium

A user cannot access an S3 bucket. How would you troubleshoot it?

#
Reveal answer guidance

Evaluate the authorization layers in sequence: (1) S3 Block Public Access: Verify if global settings block public reads. (2) Bucket Policy: Check for explicit Denys or IP/VPC restrictions in the bucket's resource policy. (3) Identity-Based IAM Policy: Ensure the user's IAM role/user has appropriate permissions (s3:GetObject / s3:ListBucket). (4) Permission Boundary / SCP: Check for restrictions at the Organizations level. (5) KMS Key Policy: If the bucket uses SSE-KMS, verify the IAM user has kms:Decrypt and kms:GenerateDataKey on the key policy.

QUESTION 23AWS ArchitectEasy

What is the difference between an S3 Bucket Policy and an IAM Policy?

#
Reveal answer guidance

(1) S3 Bucket Policy: A resource-based policy written in JSON and attached directly to the bucket. It specifies *who* (Principals) can perform actions on the bucket, which is ideal for configuring cross-account access. (2) IAM Policy: An identity-based policy attached directly to a user, group, or role, defining *what* actions that specific identity can perform across any resources. Both policies are evaluated together, and any explicit Deny in either overrides all Allows.

QUESTION 24AWS ArchitectEasy

How would you recover a deleted file from S3?

#
Reveal answer guidance

(1) Versioning Enabled: S3 does not delete the actual object on a simple Delete call; it writes a "Delete Marker" as the current version. To recover, toggle "Show Versions" in the console or CLI, find the Delete Marker, and delete it. The previous version becomes active again. (2) Versioning Disabled: The deletion is permanent. Recovery is impossible unless you have replication configured (CRR/SRR) or must restore from an S3 backup/snapshot.

QUESTION 25AWS ArchitectEasy

What is the purpose of S3 Versioning?

#
Reveal answer guidance

S3 Versioning protects against accidental deletions and overwrites by retaining multiple iterations of an object in the same bucket. It allows you to roll back files to previous historical versions and recover deleted files. It is also a prerequisite for S3 Replication (Cross-Region/Same-Region) and S3 Object Lock. Pair with S3 Lifecycle rules to transition older versions to cheaper storage (Glacier) to control costs.

QUESTION 26AWS ArchitectEasy

Public Subnet vs Private Subnet?

#
Reveal answer guidance

(1) Public Subnet: The route table has an explicit route pointing to an Internet Gateway (0.0.0.0/0 -> igw-xxxx). Resources inside get public IP addresses and can be accessed directly from the internet. Best for ALBs, public NAT Gateways, and bastions. (2) Private Subnet: The route table does not point to an Internet Gateway. Resources only have private IPs and route outbound traffic through a NAT Gateway in a public subnet. Best for applications, databases, and private backends.

QUESTION 27AWS ArchitectEasy

Why do we need a NAT Gateway?

#
Reveal answer guidance

A NAT (Network Address Translation) Gateway resides in a public subnet and enables resources in private subnets (like database EC2s or Lambda functions) to securely connect outbound to the internet (for package updates, third-party API calls, or patches) while preventing external hosts on the internet from initiating inbound connections back to those private resources.

QUESTION 28AWS ArchitectEasy

Security Group vs NACL?

#
Reveal answer guidance

(1) Security Group: Stateful firewall at the instance/ENI level. Inbound allowed traffic automatically allows return outbound traffic. Evaluates all rules, supports "Allow" only. (2) Network ACL (NACL): Stateless firewall at the subnet boundary. You must explicitly configure both inbound and outbound rules. Evaluates rules sequentially in numbered order, supports both "Allow" and "Deny" rules.

QUESTION 29AWS ArchitectMedium

What happens if a route table is misconfigured?

#
Reveal answer guidance

Misconfigured route tables drop packets, resulting in "Connection Timeout" or routing failures. Common errors include: (1) Missing internet egress (0.0.0.0/0 route) in a public subnet, (2) Private subnets routing to an IGW instead of a NAT, (3) Missing peer/Transit Gateway routes for VPC-to-VPC traffic. Diagnose using VPC Reachability Analyzer or by checking route status in the console.

QUESTION 30AWS ArchitectMedium

How would you design a highly available VPC?

#
Reveal answer guidance

(1) Multi-AZ: Spread resources across at least two (ideally three) Availability Zones. (2) Tiered Subnets: Deploy public subnets (external load balancers), private application subnets (EC2, ECS, EKS), and isolated data subnets (databases, caches) in each AZ. (3) Redundant NATs: Deploy one NAT Gateway in each AZ's public subnet to ensure that an AZ failure does not disrupt outbound connectivity for the other zones. (4) Route Tables: Separate route tables per subnet type and AZ.

QUESTION 31AWS ArchitectMedium

ALB vs NLB — when would you use each?

#
Reveal answer guidance

(1) Application Load Balancer (ALB): Operates at Layer 7 (Application). Supports HTTP/HTTPS, path-based routing, host-based routing, SSL termination, WAF, gRPC, and user authentication. Best for web applications and REST APIs. (2) Network Load Balancer (NLB): Operates at Layer 4 (Transport). Can handle millions of connections per second with ultra-low latency, supports static/Elastic IPs, TCP/UDP sockets, TLS pass-through, and is best for gaming, TCP services, or high-throughput stream processing.

QUESTION 32AWS ArchitectMedium

Why would you place an Application Load Balancer in public subnets?

#
Reveal answer guidance

An internet-facing ALB must be placed in public subnets so it can accept public incoming traffic, map public IP addresses to its listeners, and route requests. It then terminates SSL and forwards traffic internally to the application backend servers (EC2 instances or containers) residing securely in private subnets, acting as a buffer that prevents exposing private compute directly.

QUESTION 33AWS ArchitectMedium

How does Route 53 help with high availability?

#
Reveal answer guidance

Route 53 provides high availability through DNS routing strategies: (1) DNS Failover: Uses health checks to monitor backend endpoints and automatically routes traffic away from failed resources to healthy targets or static backup pages. (2) Active-Active/Active-Passive routing, (3) Latency and Geolocation routing: Directs users to the closest healthy AWS region. (4) Weighted routing: Shifts traffic gradually during canary updates.

QUESTION 34AWS ArchitectEasy

What happens when a DNS record is updated?

#
Reveal answer guidance

When a DNS record is updated in Route 53, the changes are propagated globally to all Route 53 name servers in seconds. However, clients and local ISP resolvers cache the old record based on the Time-to-Live (TTL) value. Until the TTL expires, clients will continue routing traffic to the old destination. For migrations, you should lower the record TTL (e.g., to 60s) before executing the cutover.

QUESTION 35AWS ArchitectEasy

CloudWatch vs CloudTrail?

#
Reveal answer guidance

(1) CloudWatch: Observability and monitoring. Captures performance metrics (CPU, memory, disk, network), application logs, and generates alerts/alarms based on threshold metrics. Focuses on *resource performance*. (2) CloudTrail: Auditing and compliance. Records every API call made in the AWS account, including the user, time, source IP, and action taken. Focuses on *account activity and security*.

QUESTION 36AWS ArchitectMedium

What AWS services would you use to monitor application health?

#
Reveal answer guidance

A comprehensive monitoring stack includes: (1) CloudWatch Metrics & Alarms for system and application status, (2) Route 53 Health Checks for external endpoint uptime, (3) CloudWatch Synthetics (canaries) to run automated browser scripts simulating user login and checkout flows, (4) AWS X-Ray for distributed tracing to identify API dependency latency and bottlenecks.

QUESTION 37AWS ArchitectMedium

How would you investigate a sudden spike in AWS costs?

#
Reveal answer guidance

Use a structured cost audit: (1) AWS Cost Explorer: Group cost by service, region, or usage type to pinpoint the spike. (2) Cost Anomaly Detection: Identify the exact resource that deviated. (3) Cost Allocation Tags: Track which project or team owns the resource. (4) AWS Budgets: Set alerts to notify on forecasts. (5) Athena + CUR (Cost & Usage Report): Run SQL queries on CUR logs for granular billing metadata (e.g., NAT Gateway data transfer, unattached EBS volumes).

QUESTION 38AWS ArchitectEasy

What metrics do you monitor for EC2 instances?

#
Reveal answer guidance

Monitor both hypervisor and OS metrics: (1) Default Metrics: CPU utilization, Network In/Out, Disk Reads/Writes (Ops/Bytes), and Status Checks (StatusCheckFailed_System / StatusCheckFailed_Instance). (2) Custom Agent Metrics: Memory utilization and local disk space (requires the CloudWatch Agent). (3) Burst Metrics: CPU Credit Balance for burstable instance classes (t3/t4g) to prevent unexpected performance throttling.

QUESTION 39AWS ArchitectEasy

What is the purpose of an Auto Scaling Policy?

#
Reveal answer guidance

An Auto Scaling Policy defines how an Auto Scaling Group adjusts its instance count. (1) Target Tracking: Keeps a metric stable (e.g., average CPU at 60%). (2) Step Scaling: Responds to alarms with stepped instance changes (e.g., if CPU is 80% add 2, if 95% add 4). (3) Scheduled Scaling: Scales capacity based on known timing patterns (e.g., scaling up on business mornings, scaling down at night).

QUESTION 40AWS ArchitectMedium

How would you secure secrets and passwords in AWS?

#
Reveal answer guidance

Implement a zero-secret policy: (1) AWS Secrets Manager: Best for database keys. Supports automatic rotation, KMS encryption at rest, and secret versioning. (2) Systems Manager (SSM) Parameter Store: SecureString parameters integrate with KMS for lightweight config tokens. (3) Runtime Injection: Inject secrets as environment variables into ECS/Fargate task definitions or Lambda configs at runtime instead of hardcoding them in container images or source code.

QUESTION 41AWS ArchitectMedium

What would happen if an RDS instance becomes unavailable?

#
Reveal answer guidance

(1) Single-AZ Setup: The database goes offline, causing application downtime (high RTO) and potential data loss (RPO). You must restore from snapshots. (2) Multi-AZ Setup: AWS automatically fails over to a standby replica in a different AZ. The DNS CNAME record of the database endpoint is updated to point to the new primary. The failover takes 60–120 seconds, ensuring high availability with no data loss.

QUESTION 42AWS ArchitectMedium

Multi-AZ vs Read Replica?

#
Reveal answer guidance

(1) Multi-AZ: High availability and disaster recovery. Uses synchronous replication to a standby instance in another AZ. Standby cannot serve traffic; failover is automatic. (2) Read Replicas: Scalability. Uses asynchronous replication to read-only database nodes (can be cross-region). Replicas serve read queries to offload the primary, but failover is manual (requires promoting the replica).

QUESTION 43AWS ArchitectHard

How would you perform a zero-downtime deployment in AWS?

#
Reveal answer guidance

Use one of these three deployment strategies: (1) Blue/Green: Spin up a duplicate environment (Green), deploy new code, test it, then swap DNS or ALB target group weights to shift users to Green. (2) Canary: Deploy the new code to a small percentage of instances, routing 5% of traffic, monitor error rates/logs, then scale up. (3) Rolling: ASG Instance Refresh terminates and replaces instances in batches (e.g., 25% at a time), keeping the minimum healthy instances active.

QUESTION 44AWS ArchitectMedium

A deployment succeeds, but users cannot access the application. What would you check?

#
Reveal answer guidance

Check the network path and target state: (1) Target Group Health: Verify if targets are passing health checks. If they fail, inspect the application startup logs. (2) DNS Routing: Ensure Route 53 records resolve to the correct load balancer. (3) Security Groups: Check if the deploy updated security group rules, blocking communication between the ALB and instances. (4) CloudWatch Metrics: Inspect ALB request count, 5XX errors, and target connection errors.

QUESTION 45AWS ArchitectHard

Explain the AWS architecture you worked on and why it was designed that way.

#
Reveal answer guidance

Designed a highly resilient, 3-tier microservice architecture: (1) Edge/Ingress: Route 53 with latency-based routing to CloudFront for static assets and public ALBs. (2) Compute: EKS (Elastic Kubernetes Service) with Fargate profiles in private subnets, avoiding node management and autoscaling based on CPU/Memory and SQS queue depth. (3) Database: Amazon Aurora PostgreSQL Serverless Multi-AZ for transactional data, combined with DynamoDB for session state and S3 for static assets. (4) Security: IAM Roles for Service Accounts (IRSA) for least privilege, SSM Session Manager for secure console access, and KMS for full-stack data encryption at rest.

QUESTION 46AWS ArchitectHard

Your API Gateway behind CloudFront shows intermittent 503 errors. CloudFront logs show OriginConnectError. ALB target health checks pass. Walk through the complete debugging process.

#
Reveal answer guidance

CloudFront's OriginConnectError means CloudFront cannot establish a TCP connection to your custom origin (API Gateway or ALB). Causes and checks: (1) CloudFront origin protocol mismatch — if the origin is set to HTTPS but API Gateway only accepts HTTPS, verify the origin protocol matches. (2) API Gateway endpoint type — if using EDGE optimized, CloudFront in front of an Edge-optimized API Gateway creates a circular dependency; use REGIONAL endpoint type behind CloudFront. (3) Custom domain name — API Gateway requires the Host header to match its configured domain; CloudFront may send the wrong Host. Set the Origin Custom Header to Host: <api-gateway-host>. (4) ALB idle timeout — if the origin is an ALB, ensure the ALB idle timeout (default 60s) exceeds CloudFront's origin response timeout (default 30s, max 120s for custom origins). (5) Security group — ALB security group must allow inbound from CloudFront's IP ranges (published in https://d7uri8nf7uskq.cloudfront.net/tools/list-cloudfront-ips), not from 0.0.0.0/0. (6) WAF rate limiting — if API Gateway has WAF with rate-based rules, CloudFront requests from shared IPs can trigger blocks. Enable CloudFront-Forwarded-Headers or use managed rule groups. Check CloudWatch logs for 503s with IntegrationErrorMessage to pinpoint the cause.

QUESTION 47AWS ArchitectHard

A Lambda function processes SQS messages in batches of 10. Under high load, some messages succeed while others in the same batch fail. The failed messages retry, but the successful ones are reprocessed, causing duplicates. How do you ensure at-most-once or exactly-once processing?

#
Reveal answer guidance

SQS + Lambda uses partial batch responses (since Nov 2020): your function returns {"batchItemFailures": [{"itemIdentifier": "<messageId>"}]}, telling Lambda which messages failed. Successful messages are removed from the queue; only failed ones are retried. Implementation: in your handler, process each message in a loop; if a message fails, catch the error, record its messageId in a list, and continue processing the rest. After the loop, return { batchItemFailures: failedIds.map(id => ({ itemIdentifier: id })) }. Lambda then deletes the successful messages from the queue and leaves the failed ones for retry. This achieves at-least-once for successful messages (exactly-once is impossible without dedup). For stronger guarantees: use SQS FIFO with MessageDeduplicationId based on a business-level unique key — SQS FIFO deduplicates within 5 minutes. However, Lambda doesn't support FIFO partial batch responses in all regions yet. Combine DynamoDB conditional writes for idempotency as a fallback: store processed message IDs with a TTL.

QUESTION 48AWS ArchitectMedium

Your team accidentally deleted an S3 bucket with 5 years of customer data. Versioning was not enabled. MFA Delete was disabled. What recovery options exist and what do you do in the first 15 minutes?

#
Reveal answer guidance

Without versioning, S3 object deletions are permanent within minutes. Recovery options depend on timing: (1) AWS Support case — open an urgent ticket with AWS Support; they may be able to recover deleted buckets within a window (typically a few hours to days) via internal S3 garbage collection processes. This is not guaranteed. (2) Cross-region replication — if CRR was configured, the replica bucket may still have the data. (3) Backup — restore from any external backup (AWS Backup, Velero, database dumps, ETL pipelines). (4) Glacier/Vault — if a Vault Lock or Glacier Archive was created, data may be recoverable. First 15 minutes: (1) Do NOT make any further API calls to the bucket — no writes, no list operations. (2) Open AWS Support urgent case, provide bucket name and approximate deletion time. (3) Check CloudTrail to identify who/what performed the delete — review s3:DeleteBucket events. (4) Check if the bucket name is still reserved (try aws s3api head-bucket). (5) If the bucket name is available, create it immediately with the same name to prevent someone else from registering it. (6) Contact the account team if you have Enterprise Support. Prevention: S3 Versioning + MFA Delete + AWS Backup + S3 Object Lock (write-once-read-many) for critical data. Also configure deletion prevention via SCPs that deny s3:DeleteBucket without MFA.

QUESTION 49AWS ArchitectHard

AWS Architect incident: One AZ impairment causes user-facing errors even though resources exist in two AZs. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: ALB target health, subnet/AZ mapping, route tables, RDS events, and CloudWatch metrics. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Place load-balanced stateless tiers across AZs, verify health checks per AZ, use managed Multi-AZ data services, and test AZ evacuation rather than assuming redundancy works.

QUESTION 50AWS ArchitectHard

AWS Architect architecture scenario: Design a highly available web tier with database failover. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Place load-balanced stateless tiers across AZs, verify health checks per AZ, use managed Multi-AZ data services, and test AZ evacuation rather than assuming redundancy works.

QUESTION 51AWS ArchitectMedium

AWS Architect security scenario: Security groups allow broad east-west traffic between tiers. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Place load-balanced stateless tiers across AZs, verify health checks per AZ, use managed Multi-AZ data services, and test AZ evacuation rather than assuming redundancy works.

QUESTION 52AWS ArchitectHard

AWS Architect release scenario: Move a single-AZ service to multi-AZ without downtime. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Place load-balanced stateless tiers across AZs, verify health checks per AZ, use managed Multi-AZ data services, and test AZ evacuation rather than assuming redundancy works.

QUESTION 53AWS ArchitectMedium

AWS Architect reliability/cost scenario: Cross-AZ data transfer and idle capacity costs are rising. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Place load-balanced stateless tiers across AZs, verify health checks per AZ, use managed Multi-AZ data services, and test AZ evacuation rather than assuming redundancy works.

QUESTION 54AWS ArchitectHard

AWS Architect incident: A Lambda gets AccessDenied in production but works with admin permissions. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: CloudTrail denied events, IAM policy simulator, Access Analyzer, and service-specific condition keys. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Start from observed actions, scope resources and conditions, use roles not users, add permission boundaries where needed, and keep break-glass access separate.

QUESTION 55AWS ArchitectHard

AWS Architect architecture scenario: Design least-privilege IAM for compute workloads. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Start from observed actions, scope resources and conditions, use roles not users, add permission boundaries where needed, and keep break-glass access separate.

QUESTION 56AWS ArchitectMedium

AWS Architect security scenario: Wildcard IAM policies allow privilege escalation. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Start from observed actions, scope resources and conditions, use roles not users, add permission boundaries where needed, and keep break-glass access separate.

QUESTION 57AWS ArchitectHard

AWS Architect release scenario: Replace broad IAM policies with scoped permissions. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Start from observed actions, scope resources and conditions, use roles not users, add permission boundaries where needed, and keep break-glass access separate.

QUESTION 58AWS ArchitectMedium

AWS Architect reliability/cost scenario: Security review blocks production release. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Start from observed actions, scope resources and conditions, use roles not users, add permission boundaries where needed, and keep break-glass access separate.

QUESTION 59AWS ArchitectHard

AWS Architect incident: Private instances cannot reach a managed service endpoint. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: route tables, security groups, NACLs, VPC flow logs, endpoint policies, and DNS resolution. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Trace source subnet route, destination DNS, endpoint policy, security group, NACL, and return path. Use gateway/interface endpoints for supported services and keep NAT for true internet egress.

QUESTION 60AWS ArchitectHard

AWS Architect architecture scenario: Design private networking for app, data, and shared services. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Trace source subnet route, destination DNS, endpoint policy, security group, NACL, and return path. Use gateway/interface endpoints for supported services and keep NAT for true internet egress.

QUESTION 61AWS ArchitectMedium

AWS Architect security scenario: Public IPs are added to fix connectivity quickly. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Trace source subnet route, destination DNS, endpoint policy, security group, NACL, and return path. Use gateway/interface endpoints for supported services and keep NAT for true internet egress.

QUESTION 62AWS ArchitectHard

AWS Architect release scenario: Move outbound traffic from internet paths to VPC endpoints. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Trace source subnet route, destination DNS, endpoint policy, security group, NACL, and return path. Use gateway/interface endpoints for supported services and keep NAT for true internet egress.

QUESTION 63AWS ArchitectMedium

AWS Architect reliability/cost scenario: NAT gateway costs are unexpectedly high. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Trace source subnet route, destination DNS, endpoint policy, security group, NACL, and return path. Use gateway/interface endpoints for supported services and keep NAT for true internet egress.

QUESTION 64AWS ArchitectHard

AWS Architect incident: A private object is reachable through an unexpected URL. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: bucket policy, public access block, CloudTrail data events, CloudFront origin config, and access logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Enable Block Public Access, use OAC/OAI correctly, restrict bucket policy by distribution, log object access, and use lifecycle/caching for cost control.

QUESTION 65AWS ArchitectHard

AWS Architect architecture scenario: Design S3 access through CloudFront only. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Enable Block Public Access, use OAC/OAI correctly, restrict bucket policy by distribution, log object access, and use lifecycle/caching for cost control.

QUESTION 66AWS ArchitectMedium

AWS Architect security scenario: Bucket policy uses public principals or weak conditions. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Enable Block Public Access, use OAC/OAI correctly, restrict bucket policy by distribution, log object access, and use lifecycle/caching for cost control.

QUESTION 67AWS ArchitectHard

AWS Architect release scenario: Migrate from public S3 hosting to private origin access. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Enable Block Public Access, use OAC/OAI correctly, restrict bucket policy by distribution, log object access, and use lifecycle/caching for cost control.

QUESTION 68AWS ArchitectMedium

AWS Architect reliability/cost scenario: Data transfer and request costs spike after hotlinking. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Enable Block Public Access, use OAC/OAI correctly, restrict bucket policy by distribution, log object access, and use lifecycle/caching for cost control.

QUESTION 69AWS ArchitectHard

AWS Architect incident: A bad migration corrupts production data. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: RDS events, automated backup status, PITR window, snapshots, parameter changes, and database logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use PITR, final snapshots, tested restore drills, migration approvals, least-privilege DB users, and rollback plans that include application compatibility.

QUESTION 70AWS ArchitectHard

AWS Architect architecture scenario: Design backup and restore for relational databases. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use PITR, final snapshots, tested restore drills, migration approvals, least-privilege DB users, and rollback plans that include application compatibility.

QUESTION 71AWS ArchitectMedium

AWS Architect security scenario: Application roles have destructive database privileges. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use PITR, final snapshots, tested restore drills, migration approvals, least-privilege DB users, and rollback plans that include application compatibility.

QUESTION 72AWS ArchitectHard

AWS Architect release scenario: Introduce blue/green database migration safety. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use PITR, final snapshots, tested restore drills, migration approvals, least-privilege DB users, and rollback plans that include application compatibility.

QUESTION 73AWS ArchitectMedium

AWS Architect reliability/cost scenario: Storage and backup retention costs are increasing. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use PITR, final snapshots, tested restore drills, migration approvals, least-privilege DB users, and rollback plans that include application compatibility.

QUESTION 74AWS ArchitectHard

AWS Architect incident: A few keys throttle while table-level capacity looks healthy. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: CloudWatch throttles, Contributor Insights, partition key distribution, and access pattern review. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Design from access patterns, spread hot writes with sharding where needed, avoid scans, use GSIs carefully, and validate cost under realistic traffic.

QUESTION 75AWS ArchitectHard

AWS Architect architecture scenario: Design a table for high-write multi-tenant traffic. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Design from access patterns, spread hot writes with sharding where needed, avoid scans, use GSIs carefully, and validate cost under realistic traffic.

QUESTION 76AWS ArchitectMedium

AWS Architect security scenario: IAM allows full table scans from application roles. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Design from access patterns, spread hot writes with sharding where needed, avoid scans, use GSIs carefully, and validate cost under realistic traffic.

QUESTION 77AWS ArchitectHard

AWS Architect release scenario: Change partition strategy without a big-bang migration. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Design from access patterns, spread hot writes with sharding where needed, avoid scans, use GSIs carefully, and validate cost under realistic traffic.

QUESTION 78AWS ArchitectMedium

AWS Architect reliability/cost scenario: On-demand capacity costs are climbing. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Design from access patterns, spread hot writes with sharding where needed, avoid scans, use GSIs carefully, and validate cost under realistic traffic.

QUESTION 79AWS ArchitectHard

AWS Architect incident: One event source consumes all Lambda concurrency and starves APIs. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: Lambda concurrency metrics, throttles, event source mapping settings, DLQ depth, and CloudWatch logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Set reserved concurrency for critical functions, tune batch size and retry behavior, use DLQs or destinations, and make handlers idempotent.

QUESTION 80AWS ArchitectHard

AWS Architect architecture scenario: Design serverless concurrency isolation. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Set reserved concurrency for critical functions, tune batch size and retry behavior, use DLQs or destinations, and make handlers idempotent.

QUESTION 81AWS ArchitectMedium

AWS Architect security scenario: Functions have broad permissions across event sources. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Set reserved concurrency for critical functions, tune batch size and retry behavior, use DLQs or destinations, and make handlers idempotent.

QUESTION 82AWS ArchitectHard

AWS Architect release scenario: Introduce reserved concurrency and DLQs. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Set reserved concurrency for critical functions, tune batch size and retry behavior, use DLQs or destinations, and make handlers idempotent.

QUESTION 83AWS ArchitectMedium

AWS Architect reliability/cost scenario: Retries and duplicate work increase cost. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Set reserved concurrency for critical functions, tune batch size and retry behavior, use DLQs or destinations, and make handlers idempotent.

QUESTION 84AWS ArchitectHard

AWS Architect incident: Tasks restart during deployment and the service drops traffic. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: ECS service events, target group health, task logs, CPU/memory metrics, and deployment config. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Tune health checks, min/max healthy percentages, task sizing, autoscaling, IAM roles, and log visibility before migrating production traffic.

QUESTION 85AWS ArchitectHard

AWS Architect architecture scenario: Design a resilient container service on ECS. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Tune health checks, min/max healthy percentages, task sizing, autoscaling, IAM roles, and log visibility before migrating production traffic.

QUESTION 86AWS ArchitectMedium

AWS Architect security scenario: Task roles and execution roles are over-permissive. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Tune health checks, min/max healthy percentages, task sizing, autoscaling, IAM roles, and log visibility before migrating production traffic.

QUESTION 87AWS ArchitectHard

AWS Architect release scenario: Move EC2-hosted containers to Fargate. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Tune health checks, min/max healthy percentages, task sizing, autoscaling, IAM roles, and log visibility before migrating production traffic.

QUESTION 88AWS ArchitectMedium

AWS Architect reliability/cost scenario: Fargate costs are higher than expected. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Tune health checks, min/max healthy percentages, task sizing, autoscaling, IAM roles, and log visibility before migrating production traffic.

QUESTION 89AWS ArchitectHard

AWS Architect incident: A node group upgrade evicts too many pods at once. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: EKS events, node drain logs, PDBs, pod scheduling, IAM access entries, and autoscaler logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use managed node groups carefully, enforce PDBs, isolate namespaces, use IRSA/workload identity, and test upgrades in a staging cluster first.

QUESTION 90AWS ArchitectHard

AWS Architect architecture scenario: Design an EKS platform for multiple teams. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use managed node groups carefully, enforce PDBs, isolate namespaces, use IRSA/workload identity, and test upgrades in a staging cluster first.

QUESTION 91AWS ArchitectMedium

AWS Architect security scenario: Cluster admin access is granted broadly. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use managed node groups carefully, enforce PDBs, isolate namespaces, use IRSA/workload identity, and test upgrades in a staging cluster first.

QUESTION 92AWS ArchitectHard

AWS Architect release scenario: Upgrade Kubernetes version with minimal downtime. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use managed node groups carefully, enforce PDBs, isolate namespaces, use IRSA/workload identity, and test upgrades in a staging cluster first.

QUESTION 93AWS ArchitectMedium

AWS Architect reliability/cost scenario: Cluster autoscaler overprovisions nodes. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use managed node groups carefully, enforce PDBs, isolate namespaces, use IRSA/workload identity, and test upgrades in a staging cluster first.

QUESTION 94AWS ArchitectHard

AWS Architect incident: Users see stale or wrong content after a deploy. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: CloudFront cache headers, access logs, origin logs, invalidation history, and cache policy config. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use versioned assets, correct Cache-Control, targeted invalidations, separate cache behaviors, and explicit private content controls.

QUESTION 95AWS ArchitectHard

AWS Architect architecture scenario: Design global static and dynamic content delivery. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use versioned assets, correct Cache-Control, targeted invalidations, separate cache behaviors, and explicit private content controls.

QUESTION 96AWS ArchitectMedium

AWS Architect security scenario: Signed URL or header logic leaks private content. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use versioned assets, correct Cache-Control, targeted invalidations, separate cache behaviors, and explicit private content controls.

QUESTION 97AWS ArchitectHard

AWS Architect release scenario: Change cache policies and origin request policies safely. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use versioned assets, correct Cache-Control, targeted invalidations, separate cache behaviors, and explicit private content controls.

QUESTION 98AWS ArchitectMedium

AWS Architect reliability/cost scenario: Origin load and CloudFront invalidation costs are high. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use versioned assets, correct Cache-Control, targeted invalidations, separate cache behaviors, and explicit private content controls.

QUESTION 99AWS ArchitectHard

AWS Architect incident: A poison message blocks useful work and retries repeatedly. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: queue depth, age of oldest message, receive count, DLQ metrics, and consumer logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use visibility timeout based on processing time, DLQs, idempotency keys, batch failure handling, and backpressure-aware consumers.

QUESTION 100AWS ArchitectHard

AWS Architect architecture scenario: Design asynchronous processing with backpressure. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use visibility timeout based on processing time, DLQs, idempotency keys, batch failure handling, and backpressure-aware consumers.

QUESTION 101AWS ArchitectMedium

AWS Architect security scenario: Messages contain sensitive payloads without encryption controls. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use visibility timeout based on processing time, DLQs, idempotency keys, batch failure handling, and backpressure-aware consumers.

QUESTION 102AWS ArchitectHard

AWS Architect release scenario: Introduce DLQ and redrive policies. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use visibility timeout based on processing time, DLQs, idempotency keys, batch failure handling, and backpressure-aware consumers.

QUESTION 103AWS ArchitectMedium

AWS Architect reliability/cost scenario: Retry storms increase Lambda and downstream costs. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use visibility timeout based on processing time, DLQs, idempotency keys, batch failure handling, and backpressure-aware consumers.

QUESTION 104AWS ArchitectHard

AWS Architect incident: A service cannot decrypt data after a key policy change. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: CloudTrail KMS events, key policy, grants, encryption context, and service role permissions. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Separate admin and usage permissions, use grants for AWS services when appropriate, test key policy changes, and cache data keys where services support it.

QUESTION 105AWS ArchitectHard

AWS Architect architecture scenario: Design encryption key ownership across accounts. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Separate admin and usage permissions, use grants for AWS services when appropriate, test key policy changes, and cache data keys where services support it.

QUESTION 106AWS ArchitectMedium

AWS Architect security scenario: Key administrators can also decrypt production data. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Separate admin and usage permissions, use grants for AWS services when appropriate, test key policy changes, and cache data keys where services support it.

QUESTION 107AWS ArchitectHard

AWS Architect release scenario: Rotate customer-managed keys. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Separate admin and usage permissions, use grants for AWS services when appropriate, test key policy changes, and cache data keys where services support it.

QUESTION 108AWS ArchitectMedium

AWS Architect reliability/cost scenario: KMS request costs grow with high-volume workloads. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Separate admin and usage permissions, use grants for AWS services when appropriate, test key policy changes, and cache data keys where services support it.

QUESTION 109AWS ArchitectHard

AWS Architect incident: A deployment role works in dev but cannot assume the prod role. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: STS assume-role errors, trust policy, SCPs, permission boundaries, and CloudTrail. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use Organizations, SCP guardrails, explicit trust, external IDs where needed, centralized identity, and account vending with standard baselines.

QUESTION 110AWS ArchitectHard

AWS Architect architecture scenario: Design multi-account landing zone access. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use Organizations, SCP guardrails, explicit trust, external IDs where needed, centralized identity, and account vending with standard baselines.

QUESTION 111AWS ArchitectMedium

AWS Architect security scenario: Trust policies allow unknown principals. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use Organizations, SCP guardrails, explicit trust, external IDs where needed, centralized identity, and account vending with standard baselines.

QUESTION 112AWS ArchitectHard

AWS Architect release scenario: Move workloads into separate AWS accounts. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use Organizations, SCP guardrails, explicit trust, external IDs where needed, centralized identity, and account vending with standard baselines.

QUESTION 113AWS ArchitectMedium

AWS Architect reliability/cost scenario: Account sprawl creates operational overhead. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use Organizations, SCP guardrails, explicit trust, external IDs where needed, centralized identity, and account vending with standard baselines.

QUESTION 114AWS ArchitectHard

AWS Architect incident: An outage occurs but CloudWatch alarms did not fire. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: CloudWatch metrics, logs, X-Ray traces, alarm history, synthetic checks, and dashboard gaps. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Alert on symptoms and saturation, keep debug logs sampled, protect sensitive fields, and align alarms to user impact rather than every low-level metric.

QUESTION 115AWS ArchitectHard

AWS Architect architecture scenario: Design meaningful AWS monitoring for user journeys. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Alert on symptoms and saturation, keep debug logs sampled, protect sensitive fields, and align alarms to user impact rather than every low-level metric.

QUESTION 116AWS ArchitectMedium

AWS Architect security scenario: Logs expose secrets or customer data. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Alert on symptoms and saturation, keep debug logs sampled, protect sensitive fields, and align alarms to user impact rather than every low-level metric.

QUESTION 117AWS ArchitectHard

AWS Architect release scenario: Introduce SLO-based alarms across services. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Alert on symptoms and saturation, keep debug logs sampled, protect sensitive fields, and align alarms to user impact rather than every low-level metric.

QUESTION 118AWS ArchitectMedium

AWS Architect reliability/cost scenario: Log ingestion costs are too high. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Alert on symptoms and saturation, keep debug logs sampled, protect sensitive fields, and align alarms to user impact rather than every low-level metric.

QUESTION 119AWS ArchitectHard

AWS Architect incident: DNS failover did not move users during a regional incident. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: Route 53 health checks, resolver output, TTLs, regional ALB health, and synthetic tests. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Understand DNS TTL limits, use external health checks, test failover regularly, and decide active-active versus active-passive based on RTO/RPO and cost.

QUESTION 120AWS ArchitectHard

AWS Architect architecture scenario: Design multi-region traffic failover. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Understand DNS TTL limits, use external health checks, test failover regularly, and decide active-active versus active-passive based on RTO/RPO and cost.

QUESTION 121AWS ArchitectMedium

AWS Architect security scenario: Health checks expose internal endpoints. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Understand DNS TTL limits, use external health checks, test failover regularly, and decide active-active versus active-passive based on RTO/RPO and cost.

QUESTION 122AWS ArchitectHard

AWS Architect release scenario: Introduce weighted routing before active-active. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Understand DNS TTL limits, use external health checks, test failover regularly, and decide active-active versus active-passive based on RTO/RPO and cost.

QUESTION 123AWS ArchitectMedium

AWS Architect reliability/cost scenario: Multi-region duplication costs are high. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Understand DNS TTL limits, use external health checks, test failover regularly, and decide active-active versus active-passive based on RTO/RPO and cost.

QUESTION 124AWS ArchitectHard

AWS Architect incident: Monthly AWS spend doubles after a platform rollout. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: Cost Explorer, CUR, tags, Trusted Advisor, Compute Optimizer, and service-level usage metrics. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Find top drivers, tag ownership, rightsizing compute, optimize data transfer/NAT/logs, and create budgets with accountable owners.

QUESTION 125AWS ArchitectHard

AWS Architect architecture scenario: Design cost-aware cloud architecture. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Find top drivers, tag ownership, rightsizing compute, optimize data transfer/NAT/logs, and create budgets with accountable owners.

QUESTION 126AWS ArchitectMedium

AWS Architect security scenario: Cost visibility roles expose account-wide metadata broadly. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Find top drivers, tag ownership, rightsizing compute, optimize data transfer/NAT/logs, and create budgets with accountable owners.

QUESTION 127AWS ArchitectHard

AWS Architect release scenario: Introduce budgets and cost allocation tags. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Find top drivers, tag ownership, rightsizing compute, optimize data transfer/NAT/logs, and create budgets with accountable owners.

QUESTION 128AWS ArchitectMedium

AWS Architect reliability/cost scenario: Leadership asks for immediate savings without downtime. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Find top drivers, tag ownership, rightsizing compute, optimize data transfer/NAT/logs, and create budgets with accountable owners.

QUESTION 129AWS ArchitectHard

AWS Architect incident: A restore takes longer than the stated RTO. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: AWS Backup jobs, restore test logs, snapshot age, cross-region copy status, and IAM delete permissions. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Define RTO/RPO per workload, test restores, isolate backup delete rights, copy critical backups cross-account or cross-region, and tier retention.

QUESTION 130AWS ArchitectHard

AWS Architect architecture scenario: Design disaster recovery across AWS services. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Define RTO/RPO per workload, test restores, isolate backup delete rights, copy critical backups cross-account or cross-region, and tier retention.

QUESTION 131AWS ArchitectMedium

AWS Architect security scenario: Backup delete permissions are too broad. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Define RTO/RPO per workload, test restores, isolate backup delete rights, copy critical backups cross-account or cross-region, and tier retention.

QUESTION 132AWS ArchitectHard

AWS Architect release scenario: Add AWS Backup policies and restore drills. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Define RTO/RPO per workload, test restores, isolate backup delete rights, copy critical backups cross-account or cross-region, and tier retention.

QUESTION 133AWS ArchitectMedium

AWS Architect reliability/cost scenario: Long retention increases storage cost. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Define RTO/RPO per workload, test restores, isolate backup delete rights, copy critical backups cross-account or cross-region, and tier retention.

QUESTION 134AWS ArchitectHard

AWS Architect incident: API latency spikes but backend Lambda duration is normal. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: API Gateway access logs, integration latency, authorizer logs, throttling metrics, and WAF logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Separate gateway latency from backend duration, add throttling/WAF/auth, cache safe responses, and log request IDs end to end.

QUESTION 135AWS ArchitectHard

AWS Architect architecture scenario: Design managed API ingress with throttling and auth. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Separate gateway latency from backend duration, add throttling/WAF/auth, cache safe responses, and log request IDs end to end.

QUESTION 136AWS ArchitectMedium

AWS Architect security scenario: Authorizer misconfiguration allows unintended access. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Separate gateway latency from backend duration, add throttling/WAF/auth, cache safe responses, and log request IDs end to end.

QUESTION 137AWS ArchitectHard

AWS Architect release scenario: Move public APIs behind API Gateway. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Separate gateway latency from backend duration, add throttling/WAF/auth, cache safe responses, and log request IDs end to end.

QUESTION 138AWS ArchitectMedium

AWS Architect reliability/cost scenario: Request charges grow due to abusive clients. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Separate gateway latency from backend duration, add throttling/WAF/auth, cache safe responses, and log request IDs end to end.

QUESTION 139AWS ArchitectHard

AWS Architect incident: Long-running connections drop during deployments. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: target health, deregistration delay, access logs, listener rules, and backend connection metrics. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use ALB for HTTP routing and NLB for TCP/static IP needs, tune draining and idle timeouts, and prevent clients from bypassing the load balancer.

QUESTION 140AWS ArchitectHard

AWS Architect architecture scenario: Choose ALB or NLB for mixed HTTP and TCP workloads. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use ALB for HTTP routing and NLB for TCP/static IP needs, tune draining and idle timeouts, and prevent clients from bypassing the load balancer.

QUESTION 141AWS ArchitectMedium

AWS Architect security scenario: Security groups allow direct instance bypass. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use ALB for HTTP routing and NLB for TCP/static IP needs, tune draining and idle timeouts, and prevent clients from bypassing the load balancer.

QUESTION 142AWS ArchitectHard

AWS Architect release scenario: Change target groups and listener rules safely. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use ALB for HTTP routing and NLB for TCP/static IP needs, tune draining and idle timeouts, and prevent clients from bypassing the load balancer.

QUESTION 143AWS ArchitectMedium

AWS Architect reliability/cost scenario: Load balancer hours and LCU costs are increasing. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use ALB for HTTP routing and NLB for TCP/static IP needs, tune draining and idle timeouts, and prevent clients from bypassing the load balancer.

QUESTION 144AWS ArchitectHard

AWS Architect incident: A rotated secret breaks applications using cached credentials. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and recent change: rotation Lambda logs, secret versions, application auth failures, and IAM access logs. Check logs, metrics, health checks, dependency errors, and config drift before changing anything. Recover with the smallest reversible action, then document the root cause. Prevention: Use staged labels correctly, make apps refresh secrets safely, avoid logging values, cache with TTL, and test rotation with rollback.

QUESTION 145AWS ArchitectHard

AWS Architect architecture scenario: Design safe secret rotation. What design would you choose, and what tradeoffs would you call out in an interview?

#
Reveal answer guidance

Design for failure domains, rollback, observability, and least privilege first. Validate capacity, limits, network paths, and operational ownership. The practical answer is not one service or command; it is the architecture plus the runbook. For this scenario: Use staged labels correctly, make apps refresh secrets safely, avoid logging values, cache with TTL, and test rotation with rollback.

QUESTION 146AWS ArchitectMedium

AWS Architect security scenario: Secrets are copied into environment variables and logs. How do you harden it without breaking production?

#
Reveal answer guidance

Baseline current behavior, add guardrails in report-only or staged mode where possible, and test the highest-risk paths first. Roll out with logs, alerts, and a rollback plan. Use least privilege, explicit ownership, and automated checks. For this scenario: Use staged labels correctly, make apps refresh secrets safely, avoid logging values, cache with TTL, and test rotation with rollback.

QUESTION 147AWS ArchitectHard

AWS Architect release scenario: Move static secrets to Secrets Manager rotation. How do you ship the change safely?

#
Reveal answer guidance

Separate build, deploy, validation, and cutover. Use canary or blue/green where possible, keep the old path available until health checks pass, and define rollback before starting. Watch saturation, errors, latency, and user-facing checks. For this scenario: Use staged labels correctly, make apps refresh secrets safely, avoid logging values, cache with TTL, and test rotation with rollback.

QUESTION 148AWS ArchitectMedium

AWS Architect reliability/cost scenario: Secret version sprawl and API calls raise cost. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Look at utilization, error rate, latency, queue depth, throttling, quota, retry volume, and recent deployment history. Optimize the bottleneck rather than guessing. Prefer rightsizing, caching, batching, and lifecycle policies before broad rewrites. For this scenario: Use staged labels correctly, make apps refresh secrets safely, avoid logging values, cache with TTL, and test rotation with rollback.

CONTINUE PRACTICING

Try another perspective.