Service models, regions and AZs, shared responsibility, IAM, VPC networking, load balancing, autoscaling, storage and databases, messaging, serverless and Java cold starts, twelve-factor, containers, cost, HA/DR, observability and cloud incident scenarios (AWS-focused).
Examples use AWS terminology; the principles apply to Azure and GCP with different product names.
Theory
Q1
What are IaaS, PaaS, SaaS and FaaS?
basic
They differ in how much of the stack the provider manages. IaaS gives virtual machines, disks and networks; PaaS gives a managed runtime for your code; SaaS gives a finished application; FaaS runs individual functions on demand.
IaaS: EC2, EBS, VPC. You patch the OS and runtime.
PaaS: Elastic Beanstalk, App Runner, RDS (managed platform for a database).
SaaS: Gmail, Salesforce. You only configure and use it.
FaaS: AWS Lambda. You supply a function; scaling, patching and capacity are the provider's job.
⚠ Follow-up traps
Is a managed database PaaS? Yes, it is a platform service: the provider runs the engine and OS, you manage schema and data.
Is FaaS the same as serverless? FaaS is one kind of serverless; S3, DynamoDB and SQS are serverless too.
#service-models#iaas#paas#saas#faas
Q2
What is a region and what is an Availability Zone?
basic
A region is a geographic area (for example ap-south-1) made of several isolated Availability Zones. An AZ is one or more data centers with independent power, cooling and networking, connected to its siblings by low-latency links.
Deploy across at least two AZs for high availability inside a region.
Regions are isolated from each other; data does not leave a region unless you replicate it.
⚠ Follow-up traps
Is us-east-1a the same physical AZ in every account? No, AWS shuffles the name-to-AZ mapping per account; use AZ IDs (use1-az1) to coordinate across accounts.
Is cross-AZ traffic free? No, it is billed per GB, which matters for chatty services.
#regions#availability-zones
Q3
How do you choose a region?
basic
Weigh data residency and compliance first, then latency to users, then service availability and price.
Some services and instance types launch in a few regions only.
Prices vary per region.
Disaster recovery regions should be far enough to avoid shared regional events.
⚠ Follow-up traps
Does a CDN remove the need to pick a region near users? It helps for cacheable reads; writes and dynamic API calls still go to the origin region.
#regions#latency#compliance
Q4
Explain the shared responsibility model.
basic
The provider is responsible for security "of" the cloud (hardware, hypervisor, facilities, managed-service internals); the customer is responsible for security "in" the cloud (data, IAM, network rules, OS patching on IaaS, application code).
The split shifts with the service model: on EC2 you patch the OS; on Lambda AWS patches the runtime but you still own code, dependencies and permissions.
Encryption configuration, public bucket exposure and leaked keys are customer-side failures.
⚠ Follow-up traps
Who patches the guest OS of an RDS instance? AWS, via maintenance windows; you control the window and engine version upgrades.
Who is at fault for a public S3 bucket? The customer; the service works as configured.
#security#shared-responsibility
Q5
What is IAM and what are its main building blocks?
basic
IAM controls who (principal) can do what (action) on which resource under which conditions. Building blocks: users, groups, roles, and policies.
A role has no long-term credentials; principals assume it and receive temporary credentials from STS.
Policies are JSON documents attached to identities or resources.
Explicit Deny always wins over Allow; the default is implicit deny.
⚠ Follow-up traps
Is a group a principal? No, you cannot reference a group in a resource policy principal.
What is the root user for? A few account-level tasks only; lock it with MFA and do not use it daily.
#iam#security
Q6
What is the principle of least privilege and how do you apply it?
basic
Grant only the actions and resources a workload needs, for as long as it needs them. Start narrow and widen from evidence.
Use specific actions and resource ARNs instead of *.
Use IAM Access Analyzer and last-accessed data to prune unused permissions.
Add conditions (source VPC, tags, MFA) and permission boundaries.
Is s3:* on one bucket least privilege? No, it includes delete and policy changes.
Does an AWS managed AdministratorAccess policy for a CI job pass review? Rarely; scope it to the stacks and services it deploys.
#iam#least-privilege
Q7
IAM user with access keys versus IAM role: which should an application use?
intermediate
Use a role. Compute services (EC2 instance profile, ECS task role, Lambda execution role, EKS IRSA/Pod Identity) deliver short-lived credentials automatically, so there is nothing to store or rotate.
Access keys are long-lived and leak through repos, logs and images.
The AWS SDK default credentials provider chain picks up role credentials with no code change.
⚠ Follow-up traps
How does an on-prem job call AWS without keys? IAM Roles Anywhere, or federation via OIDC/SAML.
How does GitHub Actions deploy without stored keys? OIDC federation: it assumes a role using a short-lived token.
#iam#roles#credentials
Q8
Identity-based versus resource-based policies, and how are they evaluated together?
intermediate
Identity-based policies attach to users and roles; resource-based policies attach to the resource (S3 bucket, SQS queue, KMS key). Within one account either one allowing is enough; across accounts both sides must allow.
Evaluation: explicit deny anywhere wins, then SCPs and permission boundaries must allow, then an allow must exist.
Session policies can only further restrict.
⚠ Follow-up traps
Can an SCP grant permissions? No, it only sets the maximum available.
Do SCPs affect the management account? They do not restrict the management account.
#iam#policies#evaluation
Q9
What is a VPC and what are its main components?
basic
A VPC is your logically isolated virtual network with a CIDR range, spanning all AZs in a region. Inside it you create subnets (each in one AZ), route tables, gateways and security controls.
Internet Gateway (IGW) for public internet access.
NAT Gateway for outbound-only access from private subnets.
VPC endpoints for private access to AWS services.
⚠ Follow-up traps
Can a subnet span two AZs? No, a subnet lives in exactly one AZ.
Can you change a VPC's primary CIDR? No, but you can add secondary CIDRs.
#vpc#networking
Q10
What makes a subnet public or private?
basic
A subnet is public if its route table has a route 0.0.0.0/0 to an Internet Gateway; otherwise it is private. Instances also need a public or Elastic IP to be reachable from the internet.
Typical layout: load balancers in public subnets; app and database tiers in private subnets across 2-3 AZs.
Private instances reach the internet outbound through a NAT Gateway in a public subnet.
⚠ Follow-up traps
Does a public subnet make instances reachable? Not without a public IP and permissive security group.
Is one NAT Gateway enough? It is AZ-scoped; for HA use one per AZ.
#vpc#subnets#routing
Q11
Security groups versus network ACLs.
intermediate
Security groups are stateful, attach to network interfaces, and only have allow rules. NACLs are stateless, attach to subnets, have numbered allow and deny rules evaluated in order.
Stateful: return traffic is automatically allowed.
NACLs need explicit ephemeral-port rules for responses.
Security groups can reference other security groups, which is the preferred way to model tiers.
⚠ Follow-up traps
Can a security group deny a specific IP? No, use a NACL or WAF.
What is the default for a new security group's outbound rules? Allow all outbound, no inbound.
#security-groups#nacl#vpc
Q12
What are VPC endpoints and why use them?
intermediate
They let resources reach AWS services without traversing the internet or a NAT Gateway. Gateway endpoints (S3, DynamoDB) are route-table entries and free; interface endpoints (PrivateLink) are ENIs with private IPs and are billed hourly plus per GB.
Reduces NAT data-processing cost and exposure.
Endpoint policies can restrict which buckets or actions are reachable.
⚠ Follow-up traps
Is an S3 gateway endpoint billed? No.
Does the endpoint work across regions? No, it reaches services in the same region.
#vpc#endpoints#privatelink
Q13
How do you connect two VPCs or a VPC to on-premises?
intermediate
VPC peering is a point-to-point, non-transitive link. Transit Gateway is a hub for many VPCs and on-premises links. Site-to-Site VPN runs over the internet; Direct Connect is a dedicated private circuit.
Peering requires non-overlapping CIDRs.
PrivateLink exposes a single service instead of the whole network.
⚠ Follow-up traps
If A peers with B and B with C, can A reach C? No, peering is not transitive.
Does Direct Connect encrypt traffic? Not by default; add a VPN or MACsec.
#vpc#peering#transit-gateway#vpn
Q14
What types of load balancers does AWS offer?
basic
Application Load Balancer (L7, HTTP/HTTPS, path and host routing), Network Load Balancer (L4, TCP/UDP/TLS, very high throughput, static IPs), Gateway Load Balancer (L3 for virtual appliances).
NLB preserves source IP and handles millions of requests per second with low latency.
⚠ Follow-up traps
Which supports static Elastic IPs? NLB (ALB IPs change; use Global Accelerator for static IPs in front).
Which terminates TLS and inspects paths? ALB.
#load-balancer#alb#nlb
Q15
How do load balancer health checks and target groups work?
intermediate
A target group holds targets (instances, IPs, Lambdas). The LB probes each target on a path and port; targets failing a threshold of consecutive checks are removed from rotation until they pass again.
Make the health endpoint cheap and reflect the ability to serve (readiness), not full dependency health.
Tune interval, thresholds and deregistration delay (connection draining).
⚠ Follow-up traps
Should /health call the database? A deep check can cascade an outage: DB blips mark all targets unhealthy. Prefer shallow readiness plus separate alerts.
What if all targets are unhealthy? ALB fails open and sends to all targets.
#load-balancer#health-checks
Q16
What is cross-zone load balancing?
intermediate
With it enabled, each LB node distributes requests across targets in all AZs; disabled, each node only sends to targets in its own AZ.
Always on for ALB; off by default for NLB (cross-AZ data charges apply when enabled).
Without it, uneven target counts per AZ cause uneven load.
⚠ Follow-up traps
Any cost impact? For NLB, inter-AZ transfer is charged.
#load-balancer#availability-zones
Q17
What is autoscaling and what scaling policies exist?
basic
Autoscaling adjusts capacity to demand. EC2 Auto Scaling groups support target tracking (keep a metric near a value), step scaling, scheduled scaling and predictive scaling.
Target tracking on average CPU or RequestCountPerTarget is the simplest start.
Set min, desired and max; spread across AZs.
Cooldowns and warm-up prevent flapping.
⚠ Follow-up traps
Is CPU the right metric for a Java service? Often request count or queue depth is better; JVM warm-up and GC distort CPU.
Does max capacity protect the bill? Yes, it is the cap on both scale-out and cost.
#autoscaling#ec2
Q18
Horizontal versus vertical scaling.
basic
Vertical scaling makes one machine bigger; horizontal scaling adds more machines. Horizontal needs stateless services and gives fault tolerance; vertical is simple but capped and usually needs a restart.
⚠ Follow-up traps
Which scales a relational primary write path? Mostly vertical; reads scale horizontally via replicas.
Can you scale horizontally with in-memory sessions? Only with sticky sessions or an external session store; prefer the store.
#scaling#horizontal#vertical
Q19
Compare object, block and file storage.
basic
Object (S3): flat key-addressed blobs over HTTP, virtually unlimited, eventually durable at 11 nines. Block (EBS): raw volumes attached to one instance, low latency, for databases and boot disks. File (EFS/FSx): shared POSIX or SMB file systems mounted by many hosts.
⚠ Follow-up traps
Can EBS be shared by many instances? Generally one instance at a time (io2 Multi-Attach is a narrow exception).
Is S3 a file system? No; there are no real directories and no in-place partial updates.
#storage#s3#ebs#efs
Q20
What are S3 consistency, durability and storage classes?
intermediate
S3 offers strong read-after-write consistency for all operations since December 2020. Standard stores data across at least three AZs with 99.999999999% durability.
Classes: Standard, Intelligent-Tiering, Standard-IA, One Zone-IA, Glacier Instant/Flexible/Deep Archive.
Lifecycle rules transition or expire objects automatically.
Versioning and Object Lock protect against deletes and overwrites.
⚠ Follow-up traps
Is durability the same as availability? No; Standard is 99.99% available, durability is about not losing data.
Is One Zone-IA resilient to AZ loss? No.
#s3#durability#storage-classes
Q21
How do you secure an S3 bucket?
intermediate
Enable Block Public Access at account level, use bucket policies with least privilege, enforce TLS, encrypt with SSE-S3 or SSE-KMS, and enable access logging and versioning.
Share content with presigned URLs or CloudFront Origin Access Control instead of making objects public.
Use VPC endpoint policies to constrain access paths.
⚠ Follow-up traps
Is a presigned URL revocable? Not individually; it expires or you revoke the signer's permissions.
Is S3 encrypted by default? Yes, SSE-S3 is applied to new objects by default since January 2023.
#s3#security
Q22
EBS volume types and when to pick each.
intermediate
gp3 is the general default with baseline 3,000 IOPS and 125 MB/s adjustable independent of size; io2 is for latency-sensitive databases with provisioned IOPS; st1/sc1 are HDD for throughput and cold data.
EBS is AZ-bound; snapshots go to S3 and can restore in another AZ or region.
⚠ Follow-up traps
Is gp3 cheaper than gp2? Usually, about 20% per GB, and IOPS are decoupled from size.
Does instance type limit EBS throughput? Yes, each instance has its own EBS bandwidth cap.
#ebs#storage#performance
Q23
What is a managed database service and what does it do for you?
basic
A managed database (RDS, Aurora, DynamoDB) automates provisioning, patching, backups, failover and monitoring. You still own schema design, indexing, queries, capacity choices and access control.
RDS supports PostgreSQL, MySQL, MariaDB, Oracle, SQL Server; Aurora is a cloud-native MySQL/PostgreSQL-compatible engine.
Limitations: no OS access, restricted superuser, constrained parameters.
⚠ Follow-up traps
Does managed mean no tuning? No, bad queries still hurt.
Can you SSH into RDS? No.
#rds#managed-database
Q24
How do RDS Multi-AZ and read replicas differ?
intermediate
Multi-AZ is for availability: a synchronous standby in another AZ with automatic failover, not readable (in the classic single-standby setup). Read replicas are for read scaling: asynchronous, readable, can be cross-region, and promotable manually.
Failover typically takes 60-120 seconds and the endpoint DNS flips.
Replica lag means stale reads.
⚠ Follow-up traps
Does Multi-AZ increase read throughput? Not in single-standby mode.
Does the app need a new connection string after failover? No, but it must reconnect and handle DNS caching (short JVM DNS TTL).
#rds#multi-az#read-replica
Q25
How does Aurora differ from standard RDS?
intermediate
Aurora separates compute from a distributed storage layer that keeps six copies across three AZs, auto-grows to 128 TiB, and supports up to 15 low-lag replicas that share the same storage.
Faster failover (typically under 30 seconds) as replicas share storage.
Aurora Serverless v2 scales compute in fine increments.
Aurora Global Database replicates across regions with typically sub-second lag.
⚠ Follow-up traps
Is Aurora always cheaper? No; I/O charges and instance cost can exceed RDS for steady small workloads.
#aurora#rds
Q26
When pick DynamoDB over a relational database?
intermediate
Pick DynamoDB for known access patterns, massive scale, single-digit-millisecond latency and minimal operations. Pick relational for ad hoc queries, joins and complex transactions.
Design keys from access patterns; a poor partition key causes hot partitions.
On-demand mode bills per request; provisioned mode with autoscaling is cheaper for steady load.
Global tables give multi-region active-active replication.
⚠ Follow-up traps
Can DynamoDB do joins? No; denormalize or use multiple queries.
Does Scan scale? It reads the whole table and costs accordingly; avoid in hot paths.
#dynamodb#nosql
Q27
SQS versus SNS versus Kinesis (or MSK).
intermediate
SQS is a queue: one consumer group pulls and deletes each message. SNS is pub/sub push to many subscribers. Kinesis/MSK are ordered, replayable streams where consumers track their position.
Fan-out pattern: SNS topic to several SQS queues.
Streams support replay and multiple independent readers; queues do not retain after delete.
⚠ Follow-up traps
Is SQS Standard ordered? No, it is best-effort ordering and at-least-once; FIFO queues give ordering per message group.
Can SQS replay? No.
#sqs#sns#kinesis#messaging
Q28
How does SQS visibility timeout and a dead-letter queue work?
intermediate
When a consumer receives a message it becomes invisible for the visibility timeout. If not deleted before it expires, the message reappears. After maxReceiveCount failed receives it moves to the DLQ.
Set the timeout above the processing time (Lambda: at least 6x function timeout).
Alarm on DLQ depth; redrive after fixing the bug.
⚠ Follow-up traps
What if processing outlasts the timeout? Duplicate processing; make consumers idempotent.
Does SQS guarantee exactly-once? Only FIFO with deduplication within a 5-minute window; design for idempotency anyway.
#sqs#dlq#visibility-timeout
Q29
What is serverless computing and what are its trade-offs?
basic
Serverless means no server management, automatic scaling (including to zero) and pay-per-use pricing. Trade-offs: cold starts, execution limits (Lambda 15 minutes), vendor coupling, harder local debugging, and cost surprises at steady high volume.
⚠ Follow-up traps
Are there servers? Yes, you just do not manage them.
Is serverless always cheaper? No; a constantly busy service is usually cheaper on containers or instances.
#serverless#lambda
Q30
What is a Lambda cold start?
basic
A cold start is the extra latency when Lambda must create a new execution environment: download code, start the runtime, and run initialization before the handler. Later invocations reuse the warm environment.
Phases: environment creation, runtime init, your static/init code, then the handler.
Cold start impact is rare in steady traffic but visible after idle periods or scale-out bursts.
⚠ Follow-up traps
Does every request see a cold start? No, only the first one per new environment.
Is the init phase billed? Yes, init duration is now billed for managed runtimes too.
#lambda#cold-start
Q31
Why are Java Lambda cold starts slow and how do you reduce them?
intermediate
JVM startup, class loading, JIT warm-up and framework initialization (reflection, classpath scanning, DI) dominate. Typical Spring Boot cold starts are several seconds.
Use lightweight frameworks or functional approaches; avoid classpath scanning.
Shrink the jar; lazy-initialize SDK clients but build them outside the handler.
Use SnapStart, provisioned concurrency, or GraalVM native image.
More memory means more CPU and faster init.
Prefer UrlConnection/CRT HTTP clients over heavy defaults.
⚠ Follow-up traps
Does more memory cost more? Per ms yes, but shorter runtime can make it cheaper overall; measure with power tuning.
#lambda#java#cold-start
Q32
What is Lambda SnapStart?
intermediate
SnapStart takes an encrypted snapshot of the initialized execution environment (memory and disk state) when you publish a version, and resumes new environments from that snapshot instead of initializing from scratch. It supports Java 11+ managed runtimes and cuts cold starts often by up to 10x.
Works on published versions or aliases, not $LATEST.
Uniqueness issues: random seeds, unique IDs, and network connections created during init may be duplicated or stale after restore.
CRaC runtime hooks (beforeCheckpoint, afterRestore) let you refresh state.
⚠ Follow-up traps
Is it compatible with provisioned concurrency? No, you choose one.
Does it help EFS or ephemeral storage over 512 MB? Not supported with those configurations.
#lambda#snapstart#java
Q33
Provisioned concurrency versus reserved concurrency.
intermediate
Provisioned concurrency keeps N environments initialized and warm (costs money continuously, removes cold starts for that capacity). Reserved concurrency carves out a guaranteed and maximum concurrency for a function from the account pool (free).
Reserved acts as both a guarantee and a throttle cap, useful to protect a downstream database.
Account default concurrency limit is 1,000 per region (can be raised).
⚠ Follow-up traps
Does reserved concurrency avoid cold starts? No.
What happens beyond provisioned capacity? It spills to on-demand environments with cold starts.
#lambda#concurrency
Q34
How does Lambda scale and what limits matter?
intermediate
Each concurrent request gets its own environment; Lambda adds environments quickly (burst of up to 1,000 per 10 seconds per function) until the account or reserved limit.
Throttled sync calls return 429; async ones are retried by Lambda.
⚠ Follow-up traps
What does Lambda do to a database at scale? It can exhaust connections; use RDS Proxy or limit concurrency.
Is the 15-minute limit configurable? No, it is a hard cap.
#lambda#scaling#limits
Q35
What are the twelve-factor app principles most relevant to cloud services?
basic
Twelve-factor is a methodology for portable, scalable apps. Key factors: one codebase per app, explicit dependencies, config in the environment, backing services as attached resources, strict build/release/run separation, stateless processes, port binding, concurrency via processes, disposability, dev/prod parity, logs as event streams, admin tasks as one-off processes.
⚠ Follow-up traps
Where should config live? In environment variables or a config service, never in the artifact.
Can a twelve-factor app write local files? Only as scratch; durable state goes to backing services.
#twelve-factor
Q36
Why must cloud services be stateless, and where does state go?
basic
Instances are created, replaced and terminated at any time, so no instance may hold unique state. Put sessions in Redis/ElastiCache or use signed tokens, files in S3, data in a managed database, and work items in queues.
⚠ Follow-up traps
Are sticky sessions acceptable? They are a workaround; scale-in or failure still loses the session.
Is a local cache stateful? Only if correctness depends on it; a read-through cache is fine.
#stateless#twelve-factor
Q37
Containers versus virtual machines.
basic
VMs virtualize hardware and each run a full guest OS under a hypervisor; containers share the host kernel and isolate processes with namespaces and cgroups. Containers start in seconds, are lighter and denser; VMs give stronger isolation.
Firecracker microVMs (Lambda, Fargate) combine VM isolation with container-like startup.
Containers must match the host kernel family (Linux containers need a Linux kernel).
⚠ Follow-up traps
Are containers secure sandboxes? A kernel exploit can escape; untrusted multi-tenant code needs microVMs or gVisor.
Does a container hold an OS? Only user-space files, not a kernel.
#containers#vm
Q38
ECS, EKS, Fargate and EC2: how do they relate?
intermediate
ECS and EKS are orchestrators (AWS-native and managed Kubernetes). EC2 and Fargate are the compute they run on: EC2 means you manage the nodes, Fargate is serverless per-task compute.
Fargate: no node patching, per-vCPU/GB-second billing, slower scaling than warm nodes, no privileged containers or host access.
EKS adds portability and the Kubernetes ecosystem at the cost of operational complexity and a control-plane fee.
⚠ Follow-up traps
Is Fargate cheaper than EC2? Not at steady high utilization; convenience has a premium.
#ecs#eks#fargate#containers
Q39
How should a Java application be sized and tuned for containers?
intermediate
Modern JVMs (Java 10+, 8u191+) are container-aware and size heap from the cgroup limit. Set -XX:MaxRAMPercentage (for example 70-75) to leave room for metaspace, thread stacks and native memory.
Memory limit must cover heap plus non-heap; otherwise the OOM killer terminates the container (exit code 137).
CPU limits below 1 core make the JVM pick the serial GC and 1 compiler thread.
Use layered or jlink-trimmed images and graceful shutdown to handle SIGTERM.
⚠ Follow-up traps
Is -Xmx equal to container memory? No, that guarantees OOM kills.
What is the default max heap in a container? 25% of the container memory limit.
#java#containers#jvm
Q40
What are RTO and RPO?
basic
RTO (Recovery Time Objective) is the maximum acceptable time to restore service. RPO (Recovery Point Objective) is the maximum acceptable data loss measured in time.
RPO of 5 minutes means backups or replication must be no more than 5 minutes behind.
Tighter targets cost more.
⚠ Follow-up traps
Does Multi-AZ give RPO 0? For synchronous replication within a region, effectively yes; it does not cover regional loss or logical corruption.
Are RTO and RPO technical-only decisions? No, the business sets them from impact and cost.
#disaster-recovery#rto#rpo
Q41
Describe the four AWS disaster recovery strategies.
intermediate
From cheapest and slowest to most expensive and fastest: backup and restore (hours), pilot light (core data replicated, compute off, tens of minutes), warm standby (scaled-down full stack running, minutes), multi-site active/active (near zero).
⚠ Follow-up traps
Which is pilot light versus warm standby? Pilot light has only data and minimal core services; warm standby serves traffic at reduced capacity.
Is a backup that has never been restored a plan? No; test restores regularly.
#disaster-recovery#backup-restore#pilot-light
Q42
What is high availability and how do you compute it?
basic
HA means the system keeps serving despite component failure, achieved by removing single points of failure through redundancy across AZs. Availability = uptime / total time; 99.9% allows about 43 minutes of downtime a month, 99.99% about 4.3 minutes.
Serial dependencies multiply availability: two 99.9% services in series give about 99.8%.
Parallel redundancy raises it.
⚠ Follow-up traps
Does a 99.99% provider SLA make your app 99.99%? No, your own design and dependencies bound it.
Is an SLA the same as an SLO? An SLA is a contract with penalties; an SLO is an internal target.
#high-availability#sla
Q43
When is a multi-region architecture justified?
advanced
For regional-outage tolerance beyond what AZs give, strict RTO/RPO, data-residency partitioning, or global low latency. It adds major cost and complexity: data replication, conflict resolution, routing and operations.
Patterns: active/passive with Route 53 failover; active/active with latency routing and global tables or Aurora Global Database.
Active/active needs a strategy for write conflicts (last-writer-wins, regional ownership of data).
⚠ Follow-up traps
Is multi-AZ enough for most apps? Yes, regional failures are rare; many teams do not need multi-region.
Does Route 53 failover work if the control plane is down? Rely on data-plane health checks and pre-provisioned records, not on making changes during the event.
#multi-region#high-availability
Q44
How does Route 53 route traffic?
intermediate
Route 53 supports simple, weighted, latency-based, geolocation, geoproximity, failover and multivalue routing, driven by optional health checks.
DNS-based failover is bounded by TTL and client caching.
Alias records point to AWS resources at the zone apex with no extra query charge.
⚠ Follow-up traps
Can a CNAME exist at the zone apex? No, use an alias record.
Does the JVM respect DNS TTL? By default it caches successful lookups per networkaddress.cache.ttl (30 seconds without a security manager); set it low for failover-sensitive clients.
#route53#dns#routing
Q45
What is a CDN and when do you use CloudFront?
basic
A CDN caches content at edge locations close to users, cutting latency and origin load. CloudFront fronts S3, ALB or custom origins, terminates TLS at the edge and integrates with WAF and Shield.
Control caching with Cache-Control headers and cache policies; invalidate sparingly.
Versioned file names avoid invalidation.
⚠ Follow-up traps
Can a CDN accelerate dynamic APIs? Yes, via persistent connections and the AWS backbone, even without caching.
#cdn#cloudfront#caching
Q46
What are the three pillars of observability?
basic
Metrics (aggregated numbers over time), logs (discrete events) and traces (a request's path across services). Together they answer what is wrong, why, and where.
AWS: CloudWatch Metrics/Logs/Alarms, X-Ray or OpenTelemetry (ADOT), and CloudTrail for API audit.
Emit structured JSON logs with a correlation or trace ID.
⚠ Follow-up traps
Is monitoring the same as observability? Monitoring checks known failure modes; observability lets you ask new questions of the data.
Is CloudTrail an application log? No, it records AWS API calls.
#observability#logs#metrics#traces
Q47
What should you alert on, and what are the golden signals?
intermediate
Alert on symptoms that affect users, not every cause. The four golden signals: latency, traffic, errors and saturation.
Alert on SLO burn rate rather than a raw CPU threshold.
Use percentiles (p95, p99), not averages, for latency.
Every alert needs an owner and a runbook.
⚠ Follow-up traps
Why not alert on average latency? Averages hide tail latency that real users feel.
Is high CPU always an incident? No; if latency and errors are fine it is a capacity signal, not a page.
#observability#alerting#slo
Q48
How should application configuration and secrets be managed?
intermediate
Keep config outside the artifact (environment variables, SSM Parameter Store, AppConfig) and secrets in a secret store (AWS Secrets Manager) encrypted by KMS, fetched at runtime via IAM roles.
Secrets Manager supports automatic rotation (RDS integration); Parameter Store is cheaper and fine for plain config and SecureString.
Never commit secrets, bake them into images, or print them in logs.
⚠ Follow-up traps
Are environment variables a safe place for secrets? Visible in process listings, crash dumps and console config; a store with audit logging is better.
How does a long-running Java app pick up a rotated secret? Cache with a short TTL or refresh on authentication failure.
#config#secrets#secrets-manager
Q49
What is KMS and how does envelope encryption work?
intermediate
KMS manages encryption keys in HSM-backed storage. Envelope encryption generates a data key per object, encrypts the data locally with it, then encrypts the data key with a KMS key and stores both together.
Avoids sending large data to KMS (which caps direct encrypt at 4 KB).
Key policies and grants control use; every use is logged in CloudTrail.
⚠ Follow-up traps
Does KMS see your plaintext data? No, only data keys.
What if the KMS key is deleted? Data is permanently unrecoverable after the waiting period (7-30 days).
#kms#encryption
Q50
What is Infrastructure as Code and why does it matter?
basic
IaC defines infrastructure in versioned declarative files (CloudFormation, CDK, Terraform) so environments are reproducible, reviewable and auditable.
Prevents configuration drift and manual console changes.
State handling: Terraform state must be stored remotely with locking; CloudFormation tracks state in stacks.
⚠ Follow-up traps
What is drift? Actual resources diverging from the template; detect it and reconcile.
Should you commit Terraform state? No, it can hold secrets; use a remote encrypted backend.
#iac#cloudformation#terraform
Q51
How do you approach cloud cost optimization?
intermediate
Make cost visible, remove waste, right-size, then commit. Tag resources, use Cost Explorer and budgets, delete idle resources, pick right instance sizes, then buy Savings Plans or Reserved Instances for stable baseline.
Spot instances give up to ~90% discount for interruptible work.
Use S3 lifecycle and Intelligent-Tiering, gp3 volumes, Graviton (arm64) instances for better price-performance.
Reduce data transfer: VPC endpoints, same-AZ traffic, CDN.
⚠ Follow-up traps
Do Savings Plans reduce usage? No, they only discount committed spend; waste still wastes.
Is Spot suitable for a stateful primary DB? No.
#cost-optimization#finops
Q52
On-demand, Reserved/Savings Plans and Spot: how do you mix them?
intermediate
Cover the steady baseline with Savings Plans or Reserved Instances, handle variable load on-demand, and run fault-tolerant batch or stateless workers on Spot.
Spot gets a two-minute interruption notice; handle it with checkpointing and diversified instance pools.
Compute Savings Plans apply across EC2, Fargate and Lambda.
⚠ Follow-up traps
Can a web tier use Spot? Yes, mixed with on-demand in an ASG and behind a load balancer.
#cost-optimization#spot#savings-plans
Q53
What are the main sources of hidden cloud costs?
intermediate
Data transfer (egress to internet, cross-AZ, cross-region), NAT Gateway per-GB processing, unattached EBS volumes and old snapshots, idle load balancers, over-verbose CloudWatch logs, and API request charges.
⚠ Follow-up traps
Is inbound data transfer charged? Generally free; egress is charged.
Why can NAT cost exceed the compute? All outbound traffic including S3 and ECR pulls passes through it per GB unless endpoints exist.
#cost-optimization#data-transfer
Q54
What is the Well-Architected Framework?
basic
AWS's guidance organized in six pillars: operational excellence, security, reliability, performance efficiency, cost optimization and sustainability, used for design reviews and trade-off discussions.
⚠ Follow-up traps
Can you maximize all pillars? No, they trade off (for example reliability versus cost).
#well-architected#design
Q55
What are API Gateway and its role in serverless APIs?
intermediate
API Gateway is a managed front door that handles routing, auth (IAM, Cognito, JWT, Lambda authorizers), throttling, request validation and caching for REST, HTTP and WebSocket APIs.
HTTP APIs are cheaper and faster than REST APIs but have fewer features.
Synchronous integration timeout is 29 seconds by default for REST APIs.
⚠ Follow-up traps
Can a Lambda behind API Gateway run 5 minutes? Not synchronously; use async patterns.
#api-gateway#serverless
Q56
What is the difference between synchronous and asynchronous integration, and why prefer async in the cloud?
intermediate
Synchronous calls couple caller availability and latency to the callee; asynchronous messaging buffers work in a queue, absorbs spikes and lets each side scale and fail independently.
Needs idempotent consumers, retry with backoff and DLQs.
Trade-off: eventual consistency and harder end-to-end tracing.
⚠ Follow-up traps
Does a queue make a slow consumer faster? No, it only protects the producer and smooths load.
#async#decoupling#resilience
Q57
What are retries, timeouts, backoff and circuit breakers for in cloud calls?
intermediate
Cloud calls fail transiently. Use bounded timeouts, retry only idempotent operations with exponential backoff plus jitter, and trip a circuit breaker to stop hammering a failing dependency.
Retries multiply load (retry storms); cap attempts and budget retries.
AWS SDK v2 has built-in retry policies with backoff.
⚠ Follow-up traps
Why jitter? Without it, clients retry in synchronized waves.
Can you retry a POST payment safely? Only with an idempotency key.
#resilience#retry#circuit-breaker
Q58
What are the common deployment strategies in the cloud?
intermediate
Rolling (replace instances gradually), blue/green (a full new environment, switch traffic), canary (a small percentage first, then ramp), and feature flags to decouple deploy from release.
Blue/green gives instant rollback but doubles capacity temporarily.
Database changes must be backward compatible (expand/contract) so old and new code coexist.
⚠ Follow-up traps
What breaks rollback? Irreversible schema migrations.
#deployment#blue-green#canary
Q59
What is the difference between elasticity and scalability?
basic
Scalability is the ability to handle more load by adding resources; elasticity is doing so automatically and also releasing them when demand drops, so you pay only for what you use.
⚠ Follow-up traps
Is an over-provisioned fixed fleet elastic? No.
#elasticity#scalability
Q60
What is multi-tenancy and what isolation options exist?
advanced
Multi-tenancy serves many customers on shared infrastructure. Models: silo (account or stack per tenant), pool (shared resources with a tenant ID), or bridge (mix).
Pool is cheapest but needs strict row-level access control and noisy-neighbor controls such as per-tenant throttling and quotas.
Silo gives stronger blast-radius and compliance isolation at higher cost.
⚠ Follow-up traps
Is a tenant ID filter in the query enough? One missed filter leaks data; enforce at a lower layer (row-level security, scoped credentials).
#multi-tenancy#isolation
Q61
Why use multiple AWS accounts and what is AWS Organizations?
intermediate
An account is the strongest isolation boundary for permissions, quotas and billing. Organizations groups accounts, gives consolidated billing, and applies Service Control Policies as guardrails.
Limits and service quotas are per account, which also contains blast radius.
⚠ Follow-up traps
Do SCPs grant permissions? No, they only cap them.
#organizations#accounts#governance
Scenarios
Q62
One Availability Zone has an outage. What happens to your service and what should you check?
intermediate
If the service runs in at least two AZs behind a load balancer, health checks drain the failed AZ's targets and traffic shifts to survivors; the Auto Scaling group launches replacements in healthy AZs. A single-AZ deployment goes down.
Check: capacity headroom (survivors must absorb the extra load, so run N+1 AZ capacity), RDS Multi-AZ failover, single-AZ resources (NAT Gateway, EBS volumes, One Zone caches).
Do not fail over manually before verifying health checks and metrics.
⚠ Follow-up traps
With 2 AZs at 60% CPU each, are you safe? No, one AZ then needs 120%; keep each below ~50% or use 3 AZs.
Does the ASG rebalance back automatically? Yes, after recovery it rebalances, which launches and terminates instances.
#outage#availability-zones#high-availability
Q63
The AZ failure is real but the database stays in the dead AZ. Why might that happen?
advanced
The database was deployed Single-AZ, or Multi-AZ failover is delayed or blocked. Single-AZ RDS has no standby, so recovery needs a restore from snapshot in another AZ.
Multi-AZ failover takes a minute or two and the app must reconnect; stale pooled connections and cached DNS prolong the outage.
Mitigate with Multi-AZ enabled, connection validation and retry in the pool, and a short DNS TTL in the JVM.
⚠ Follow-up traps
Does the endpoint change after failover? The DNS name stays the same but resolves to the new primary.
Why do connections hang for minutes? Dead TCP connections need socket timeouts and pool validation to be detected.
#outage#rds#multi-az
Q64
Your monthly AWS bill tripled overnight. How do you investigate?
intermediate
Use Cost Explorer grouped by service, then by usage type, account, region and tag, comparing day over day to find the spike. Then drill down in CloudTrail and resource listings to identify who created what.
Common causes: leaked access keys mining crypto in an unused region, NAT/data transfer from a loop, runaway Lambda recursion, unbounded logs, forgotten large instances or snapshots.
Contain first: disable the keys, stop resources; then fix and add guardrails.
⚠ Follow-up traps
Why check other regions? Attackers launch resources in regions you never use.
Does Cost Explorer show real-time data? It lags by up to a day; use Budgets alerts and Cost Anomaly Detection.
#runaway-bill#cost#incident
Q65
A Lambda triggered by S3 writes back to the same bucket and the bill explodes. What happened?
intermediate
An infinite trigger loop: each output object fires a new invocation which writes another object. Concurrency and cost grow until throttled.
Fix: write to a different bucket or prefix and use prefix/suffix filters on the notification.
Set reserved concurrency to cap blast radius; Lambda also detects some recursive SQS/SNS/Lambda loops and stops them.
Add a billing alarm and anomaly detection.
⚠ Follow-up traps
Is a prefix filter always enough? Only if the output prefix does not overlap the trigger prefix.
#runaway-bill#lambda#s3#recursion
Q66
A leaked AWS access key was found in a public repository. What do you do?
intermediate
Deactivate the key immediately (not just delete, so you can audit), review CloudTrail for its activity, revoke active sessions via a deny policy keyed on aws:TokenIssueTime, remove unauthorized resources, then rotate and move the workload to roles.
Removing the commit from git history is not enough; assume it was scraped.
Enable secret scanning and pre-commit hooks.
⚠ Follow-up traps
Is rotating the key enough? No, check for persistence such as new users, roles or access keys created by the attacker.
Do role sessions issued with the stolen key die with it? No, they stay valid until expiry unless revoked.
#security#incident#credentials
Q67
A noisy neighbor slows your service on a shared instance. How do you detect and fix it?
intermediate
Symptoms: latency spikes without matching load, high CPU steal time, or throttled I/O. Check %steal in top, burstable credit balance, and EBS queue length.
Fixes: dedicated hosts or the largest size of a family, Nitro-based non-burstable types, provisioned IOPS (io2/gp3), container CPU and memory requests and limits.
For multi-tenant apps, apply per-tenant rate limits and bulkheads.
⚠ Follow-up traps
Is it always a neighbor? Often it is your own T-instance running out of CPU credits.
Does a bigger instance size help? The largest size of a family uses the whole host and removes neighbors.
#noisy-neighbor#performance#ec2
Q68
Your T3 instances are fast for hours, then become slow. Why?
basic
Burstable instances earn CPU credits while below baseline and spend them when above. When credits run out, the CPU is throttled to the baseline (for example 20% of a vCPU on a t3.medium in standard mode).
Check the CPUCreditBalance metric.
Fix: switch to M/C/R families, enable unlimited mode (extra charge), or right-size.
⚠ Follow-up traps
Does unlimited mode cost nothing? Surplus credits above the average baseline are billed.
#ec2#burstable#cpu-credits
Q69
Traffic spikes 10x in two minutes (flash sale). Autoscaling is too slow. What do you do?
advanced
Reactive scaling takes minutes (metric lag, instance boot, JVM warm-up). Use scheduled or predictive scaling before the known event, keep headroom, and absorb bursts with queues, caching and rate limiting.
Use warm pools, smaller faster-booting images, or containers.
Scale on a leading indicator (request rate, queue depth) rather than CPU.
Contact AWS for extreme known spikes and check service quotas.
⚠ Follow-up traps
Does a new JVM instance serve at full speed immediately? No, JIT warm-up means high latency at first; use slow-start in the target group.
Do you check account quotas? Yes, EC2 vCPU limits can block scale-out silently.
#autoscaling#spikes#capacity
Q70
The ASG keeps launching and terminating instances (flapping). Why?
intermediate
Common causes: health check grace period shorter than app startup time, so instances are marked unhealthy and killed before ready; scale-in and scale-out thresholds too close; or a crashing app.
Set the grace period above JVM and Spring startup time, use target tracking, and separate cooldowns.
Check the ALB health check path, port and security group reachability.
⚠ Follow-up traps
Which health check type should the ASG use with an ALB? ELB health checks in addition to EC2 status checks.
#autoscaling#health-checks#troubleshooting
Q71
Users get random 502/504 errors from an ALB. How do you debug?
intermediate
502 means the target sent an invalid response or closed the connection; 504 means the target did not respond within the ALB idle timeout (default 60 seconds).
Classic 502 cause: the backend's keep-alive timeout is shorter than the ALB idle timeout, so the backend closes a reused connection. Set the app keep-alive higher than the ALB idle timeout (for example Tomcat keepAliveTimeout above 60 s).
504: slow queries or downstream calls; add timeouts and check target response time.
Review ALB access logs and target_status_code.
⚠ Follow-up traps
Is 503 the same? 503 usually means no healthy targets.
#load-balancer#troubleshooting#timeouts
Q72
After deploy, the ALB shows all targets unhealthy but the app runs. What's wrong?
basic
Typical causes: the target's security group does not allow inbound from the load balancer's security group on the health check port; wrong health check path or port; app returns 401/302 on the path; or the app has not finished starting.
Fix by referencing the ALB security group as the source.
Make the health path unauthenticated and return 200.
⚠ Follow-up traps
Should you open the instance to 0.0.0.0/0 to test? No, allow the ALB's security group only.
#load-balancer#security-groups#troubleshooting
Q73
A private-subnet instance cannot reach the internet or S3. Walk through the checks.
intermediate
Verify the route table: a 0.0.0.0/0 route to a NAT Gateway (which itself sits in a public subnet with an IGW route and an Elastic IP). Then check security group egress, NACLs (ephemeral ports), DNS resolution and, for S3, a gateway endpoint plus endpoint and bucket policies.
Use VPC Reachability Analyzer and flow logs to see where packets drop.
⚠ Follow-up traps
Can a NAT Gateway be in a private subnet? Not for internet access; it must be in a public subnet.
Do flow logs show which rule rejected? No, only ACCEPT/REJECT; use Reachability Analyzer.
#vpc#nat#troubleshooting
Q74
Lambda in a VPC cannot call an external API. Why?
intermediate
A VPC-attached Lambda has only private IPs, so it needs a route through a NAT Gateway in a public subnet to reach the internet. Placing it in a public subnet does not help since its ENI has no public IP.
For AWS services use VPC endpoints to avoid NAT cost.
Only attach a Lambda to a VPC if it needs VPC resources like RDS.
⚠ Follow-up traps
Does VPC attachment still add a large cold-start delay? No, since Hyperplane ENIs (2019) the extra delay is small.
#lambda#vpc#nat
Q75
Your Spring Boot Lambda takes 8 seconds to cold start. Give a prioritized plan.
advanced
Measure first (init versus handler duration in the REPORT line), then apply the cheapest wins.
Enable SnapStart on a published version/alias; typically cuts to a few hundred ms.
Raise memory (CPU scales with it) and use arm64.
Trim dependencies, lazy beans, avoid classpath scanning, or move to Spring Cloud Function, Micronaut or Quarkus.
Provisioned concurrency for latency-critical paths.
GraalVM native image for the best cold start with build complexity.
⚠ Follow-up traps
Why is the first SnapStart restore still slower? Restore pulls snapshot pages on demand; a warm-up before the checkpoint helps.
Is a keep-warm ping the best fix? It only warms one environment and fails under concurrency; use provisioned concurrency.
#lambda#java#cold-start#spring
Q76
After enabling SnapStart, two invocations return the same "random" token. Why?
advanced
Random or ID-generator state was initialized during init and captured in the snapshot, so restored environments start from the same state.
Generate randomness and unique IDs inside the handler, or reseed in a CRaC afterRestore hook.
Re-create network connections and refresh credentials or time-bound caches after restore.
⚠ Follow-up traps
Does the runtime reseed everything? It handles its own default SecureRandom; custom or cached state is your responsibility.
Where should DB connections be opened? Lazily at first use or in afterRestore.
#lambda#snapstart#uniqueness
Q77
A Lambda that works locally times out when talking to RDS under load. What's happening?
advanced
Each concurrent environment opens its own connection, so a burst of 500 invocations exhausts max_connections and new connections hang or fail.
Use RDS Proxy to pool connections, cap reserved concurrency, keep pool size at 1 per environment, and reuse connections across invocations.
Alternatives: DynamoDB, or a queue in front to throttle consumers.
⚠ Follow-up traps
Is HikariCP with 10 connections fine in Lambda? No, it multiplies by concurrency; use a pool of 1.
#lambda#rds#connections#rds-proxy
Q78
An SQS-triggered Lambda processes the same message twice. Is that a bug?
intermediate
No. SQS Standard is at-least-once. Duplicates occur on visibility timeout expiry, partial batch failures or retries, so consumers must be idempotent.
Use a deduplication store (DynamoDB conditional write, idempotency key) or naturally idempotent operations (upsert).
Report partial batch failures with batchItemFailures so successful messages are not re-run.
⚠ Follow-up traps
If one message in a batch of 10 fails? Without partial batch response all 10 return to the queue.
Does FIFO fully eliminate duplicates? Only within the dedupe window; keep idempotent handlers.
#sqs#lambda#idempotency
Q79
A poison message keeps failing and blocks the queue. How do you handle it?
intermediate
Configure a dead-letter queue with a maxReceiveCount (for example 3-5) so repeated failures move aside, alarm on DLQ depth, inspect and redrive after a fix.
On FIFO queues a stuck message blocks its message group, so a DLQ is essential.
Distinguish transient errors (retry with backoff) from permanent ones (send to DLQ immediately).
⚠ Follow-up traps
DLQ retention shorter than the source queue? Messages can expire before you look; set the DLQ retention to the maximum (14 days).
#sqs#dlq#poison-message
Q80
A consumer cannot keep up and the queue grows. What are your options?
intermediate
Scale consumers on ApproximateNumberOfMessagesVisible (or backlog per instance), increase batch size, optimize processing, and for downstream limits apply backpressure or rate limiting rather than more consumers.
Track age of the oldest message as the key SLO signal.
If ordering is required per key, FIFO message groups allow parallelism across groups only.
⚠ Follow-up traps
Does adding consumers always help? Not if the bottleneck is the database or a shared lock.
#sqs#scaling#backpressure
Q81
You must choose between Kinesis Data Streams and SQS for an event pipeline. How do you decide?
advanced
Pick Kinesis (or Kafka/MSK) when you need ordered replay, multiple independent consumers of the same data, or time-window processing. Pick SQS for simple work distribution with per-message acknowledgement and easy scaling.
Kinesis ordering is per shard; hot partition keys limit throughput (1 MB/s or 1,000 records/s per shard write).
A failing record blocks the shard unless handled with retries, bisect on error or skip.
⚠ Follow-up traps
Can a new consumer read old messages from SQS? No, from Kinesis yes within retention (24 hours default, up to 365 days).
#kinesis#sqs#streams
Q82
You need exactly-once effect for payments consumed from a queue. How?
advanced
Exactly-once delivery is not realistic; build exactly-once effect with at-least-once delivery plus idempotent processing. Store an idempotency key with the result in the same transaction as the side effect.
Use unique constraints or a conditional write (attribute_not_exists) keyed by message or business ID.
For cross-service flows use the outbox pattern and sagas.
⚠ Follow-up traps
Does FIFO plus dedup ID solve it? Only for 5 minutes at the queue; your handler can still be retried after partial success.
#idempotency#messaging#consistency
Q83
Aurora/RDS primary CPU is 95% on a read-heavy workload. What do you do?
intermediate
Find the top queries with Performance Insights and fix indexes first. Then offload reads to replicas (reader endpoint), add a cache (ElastiCache) for hot data, and only then scale the instance up.
Replica reads are eventually consistent; route read-your-writes paths to the primary.
Pool connections with RDS Proxy.
⚠ Follow-up traps
Does a read replica help write-heavy load? No, writes still hit the primary.
#rds#read-replica#caching
Q84
How do you run a regional DR plan for an Aurora-backed service with RPO 1 minute and RTO 15 minutes?
advanced
Use Aurora Global Database (replication lag typically under 1 second) with a warm standby stack in the secondary region. On disaster, detach and promote the secondary cluster, then shift Route 53 to the secondary region.
Pre-provision compute, secrets, KMS keys and config in both regions; Secrets Manager supports multi-region replication.
Automate and rehearse the runbook; measure real RTO.
⚠ Follow-up traps
Does promotion lose data? Possibly the unreplicated tail; unplanned failover can lose a few seconds.
What about KMS keys? Encrypted cross-region replicas need a key in the target region (multi-Region keys).
#disaster-recovery#aurora#multi-region
Q85
S3 data was deleted by a faulty release. How could you have prevented the loss?
intermediate
Enable S3 versioning (and MFA Delete or Object Lock for compliance), restrict s3:DeleteObject* permissions, and replicate to another account or region.
With versioning, a delete adds a delete marker; remove the marker to restore.
Lifecycle rules should expire noncurrent versions after a retention period to control cost.
⚠ Follow-up traps
Does cross-region replication protect against bad changes? Not alone; replicas can receive the bad overwrite, so keep versioning and backups too.
#s3#backup#versioning
Q86
A deploy introduced a schema change and rollback now fails. How should it have been done?
advanced
Use expand/contract: first deploy a backward-compatible schema (add a nullable column), deploy code that writes both, backfill, switch reads, and only later drop the old column. Old and new app versions must both work against the intermediate schema.
Blue/green and rolling deployments always have a window where both versions run.
⚠ Follow-up traps
Is renaming a column in one release safe? No, the old version breaks immediately.
#deployment#database#migration
Q87
A container keeps restarting with exit code 137. What does it mean?
intermediate
Exit 137 is SIGKILL (128 + 9), most often the kernel OOM killer after the container exceeded its memory limit. In Java the cause is usually heap plus metaspace, threads and direct buffers exceeding the limit, not a Java OutOfMemoryError.
Set -XX:MaxRAMPercentage=70, cap -XX:MaxMetaspaceSize and -Xss if needed, and raise the limit or reduce load.
Enable -XX:+HeapDumpOnOutOfMemoryError to a mounted path; check OOMKilled in task or pod status.
⚠ Follow-up traps
If heap shows only 50% used, can it still be OOM-killed? Yes, native memory is outside the heap.
Why does the pod restart in a loop? The restart policy relaunches it; the cause stays until fixed.
#containers#oom#java
Q88
Rolling deploy drops in-flight requests. Why and how to fix?
intermediate
Old instances receive SIGTERM and exit before finishing requests, or the load balancer keeps sending traffic while they stop. Deregister first, then drain, then stop.
Enable Spring Boot server.shutdown=graceful and a spring.lifecycle.timeout-per-shutdown-phase.
Set the ALB deregistration delay and the ECS stopTimeout (or Kubernetes terminationGracePeriodSeconds plus a preStop sleep) to cover the drain.
Ensure PID 1 receives signals (exec form entrypoint).
⚠ Follow-up traps
Why a preStop sleep in Kubernetes? Endpoint removal and SIGTERM race; the sleep lets the endpoint propagate.
What is the ECS default stop timeout? 30 seconds before SIGKILL.
#deployment#graceful-shutdown#containers
Q89
Different environments behave differently after "identical" deploys. What causes it?
intermediate
Configuration drift: manual console changes, per-environment config baked into builds, unpinned dependencies or mutable latest image tags.
Build once, promote the same immutable artifact (digest-pinned image) through environments with only injected config differing.
Provision everything with IaC and detect drift.
⚠ Follow-up traps
Is latest acceptable in production? No, it is mutable and breaks reproducibility and rollback.
#twelve-factor#config#drift
Q90
A secret rotation causes a production outage. What went wrong?
intermediate
The app read the secret once at startup, so after rotation its credentials became invalid. Or the rotation was not two-phase and the old password stopped working while consumers still used it.
Fetch with a short TTL cache and refetch on auth failure.
Use alternating-user rotation (Secrets Manager multi-user strategy) so old and new credentials overlap.
Test rotation in staging.
⚠ Follow-up traps
Does the Hikari pool pick up a new password? Existing connections stay valid, new ones fail; rebuild the pool or refresh the datasource.
#secrets#rotation#resilience
Q91
A developer needs temporary prod access for debugging. How do you grant it safely?
intermediate
Use federated SSO with a time-boxed role (IAM Identity Center permission set), read-only by default, with MFA, an approval workflow and CloudTrail logging. Prefer session-based access (SSM Session Manager) over opening SSH ports.
No shared accounts or long-lived keys.
Break-glass roles with alerts when used.
⚠ Follow-up traps
Is opening port 22 to your IP fine? IPs change and ports remain a risk; Session Manager needs no inbound ports.
#iam#access#security
Q92
A bucket policy and an IAM policy disagree (one allows, one denies). Who wins?
basic
An explicit deny in any applicable policy wins. In the same account, an allow in either the identity or the resource policy is sufficient when there is no deny.
⚠ Follow-up traps
Cross-account with only a bucket policy allow? Denied, the caller's identity policy must also allow.
Does a missing allow equal a deny? Yes, implicit deny.
#iam#policy-evaluation#s3
Q93
`AccessDenied` calling `s3:GetObject` even though the role policy allows it. What do you check?
intermediate
Check for explicit denies in the bucket policy, SCPs, permission boundaries or session policies; the object's KMS key policy (the role also needs kms:Decrypt); the correct resource ARN (bucket/* versus bucket); VPC endpoint policy; and the object owner for cross-account writes.
Use the IAM policy simulator and CloudTrail error details.
⚠ Follow-up traps
Why does ListBucket need the bucket ARN but GetObject the object ARN? They act on different resource types.
Is the error from KMS visible as S3 AccessDenied? Yes, it often surfaces as a generic 403.
#iam#troubleshooting#s3#kms
Q94
Costs from NAT Gateway are 40% of the bill. How do you cut them?
intermediate
Find the top talkers with VPC Flow Logs, then add S3 and DynamoDB gateway endpoints (free), interface endpoints for heavy AWS services (ECR, CloudWatch, STS), and keep inter-service traffic inside the VPC.
Pull container images via ECR endpoints; cache dependencies.
A NAT per AZ avoids cross-AZ charges but costs hourly each; weigh volume.
⚠ Follow-up traps
Are interface endpoints always cheaper? No, each costs hourly per AZ; they pay off only with enough traffic.
#cost-optimization#nat#vpc-endpoints
Q95
A data transfer bill spikes with cross-AZ traffic between microservices. What are the options?
advanced
Cross-AZ traffic costs about $0.01/GB each direction. Reduce chattiness (batching, compression), co-locate heavily communicating services per AZ with zone-aware routing (topology-aware routing in Kubernetes, ALB zonal affinity), and cache.
Do not drop to a single AZ to save money; you lose resilience.
⚠ Follow-up traps
Is traffic through an ALB across AZs billed? ALB-to-target cross-AZ traffic is not charged, but client-to-service calls between instances are.
Dev and test environments cost more than production. What do you do?
basic
Schedule shutdowns outside working hours, use smaller instance sizes, Spot, and ephemeral per-PR environments torn down automatically; apply TTL tags with cleanup jobs and budgets per account.
⚠ Follow-up traps
What else survives a stop? EBS volumes, Elastic IPs and snapshots still bill.
#cost-optimization#environments
Q97
You must choose between Lambda and Fargate for an API with steady 200 requests per second. Which?
advanced
For steady, sustained load containers are usually cheaper and avoid cold starts and Lambda concurrency limits. Lambda wins for spiky, low-volume or idle-heavy workloads and for minimal operations.
Rough check: Lambda cost scales with requests x duration x memory; at constant 200 rps with 100 ms and 512 MB, a few Fargate tasks cost less.
Consider team skills, latency requirements and deployment model.
⚠ Follow-up traps
Does Lambda scale to zero cost? Yes, which makes it better for rare calls.
Can you mix? Yes, containers for the core API, Lambda for event glue and bursty jobs.
#lambda#fargate#cost#trade-offs
Q98
A Lambda function times out on large file processing. What are the alternatives?
intermediate
Lambda caps at 15 minutes and 10 GB memory. Break work into chunks via SQS or Step Functions, stream the file rather than loading it, or move to Fargate/AWS Batch for long jobs.
Step Functions orchestrates fan-out with the Map state.
Use S3 byte-range reads and increase /tmp ephemeral storage (up to 10 GB) if required.
⚠ Follow-up traps
Can you raise the timeout above 15 minutes? No.
#lambda#limits#architecture
Q99
Latency to users in Asia is 400 ms while your region is in Virginia. What do you do?
basic
First cache static and cacheable API responses at the edge with CloudFront. For dynamic traffic, use CloudFront or Global Accelerator for backbone routing, then deploy a regional stack closer to users with latency-based routing.
The hard part is data: read replicas or global tables near users, with a single-writer or conflict strategy for writes.
⚠ Follow-up traps
Will a bigger instance fix it? No, distance is the bottleneck.
#latency#cdn#multi-region
Q100
Active/active across two regions causes lost updates. Why?
advanced
Both regions accept writes to the same record and replicate asynchronously; DynamoDB global tables resolve conflicts by last-writer-wins on timestamp, so one write silently overwrites the other.
Mitigate: route each entity to a home region (partition ownership), use CRDTs or merge logic, or make writes append-only/commutative.
Aurora Global Database is single-writer by design.
⚠ Follow-up traps
Does global tables replication guarantee order? Eventual, not strict global order.
#multi-region#consistency#dynamodb
Q101
After a region failover, users report missing recent orders. Explain.
intermediate
Replication was asynchronous, so writes committed in the failed region but not yet replicated were lost; this is the actual RPO, which is nonzero for cross-region async replication.
Reconcile later from the primary's logs or event stream if recoverable, and revisit the RPO decision.
Synchronous cross-region replication is slow and rarely used.
⚠ Follow-up traps
Can backups fix it? Not recent data; backups only restore to their point in time.
#disaster-recovery#rpo#replication
Q102
CloudWatch Logs cost is huge. How do you reduce it?
basic
Ingestion is the main cost. Lower log levels in production (no DEBUG), drop noisy health-check logs, sample verbose events, set retention periods, and archive to S3 via subscription or Firehose for long-term storage.
Use the Infrequent Access log class for rarely queried logs.
Metric filters and Embedded Metric Format avoid logging just to count.
⚠ Follow-up traps
Is storage or ingestion costlier? Ingestion, per GB, usually dominates.
What is the default log retention? Never expire.
#observability#cost-optimization#logging
Q103
You cannot find the root cause of slow requests across five microservices. What do you add?
intermediate
Distributed tracing: propagate a trace context (W3C traceparent) across calls and export spans with OpenTelemetry (ADOT collector to X-Ray or another backend). Put the trace ID in logs for correlation.
Spring Boot 3 integrates Micrometer Tracing and the OpenTelemetry bridge.
Look at span waterfalls for the slowest hop (often a DB or an external call).
⚠ Follow-up traps
Why sample traces? Volume and cost; use tail-based sampling to keep errors and slow traces.
Does tracing show payloads? Not by default; avoid recording PII.
#observability#tracing#opentelemetry
Q104
The pager fires constantly but users are fine. What do you change?
basic
Alert fatigue: replace cause-based thresholds (CPU > 80%) with symptom-based SLO alerts (error rate, latency burn rate), add runbooks, tune or delete noisy alarms, and route non-urgent signals to tickets.
⚠ Follow-up traps
Should every alarm page? No, only actionable and urgent ones.
#alerting#observability#slo
Q105
A Java service on EC2 slows after hours; CPU is low but latency high. Where do you look?
advanced
Low CPU with high latency points to waiting: connection pool exhaustion, thread pool saturation, GC pauses, slow downstream calls or EBS/network throttling.
Does low CPU mean healthy? No, blocked threads use no CPU.
#troubleshooting#jvm#observability
Q106
A stateful session is lost whenever the ASG scales in. What is the fix?
basic
Sessions were kept in instance memory. Externalize to ElastiCache (Redis) via Spring Session, or use stateless JWT, and enable connection draining so in-flight requests finish.
⚠ Follow-up traps
Do sticky sessions fix the loss on scale-in? No, the instance is still terminated.
#stateless#sessions#autoscaling
Q107
You must migrate a monolith to the cloud quickly with minimal change. Which strategy?
intermediate
Rehost (lift and shift) onto EC2 is the fastest; replatform (move DB to RDS, files to S3, add ASG) gives early gains with modest change; refactor into cloud-native services has the highest payoff and risk. Choose by deadline and value, and iterate (the "6 Rs": rehost, replatform, repurchase, refactor, retire, retain).
⚠ Follow-up traps
Does lift and shift save money automatically? Often not; right-size and use managed services afterward.
#migration#rehost#replatform
Q108
You want zero-downtime failover of an RDS database but your app errors for two minutes. Why?
advanced
Failover swaps DNS to the standby. The JVM may cache the old IP, the pool holds dead connections, and no socket or connect timeouts mean threads block until TCP timeouts.
Set networkaddress.cache.ttl to a low value (such as 5-30), configure Hikari connectionTimeout, validationTimeout and maxLifetime, set driver socketTimeout and use the AWS Advanced JDBC Wrapper's failover plugin to react faster.
Retry idempotent transactions.
⚠ Follow-up traps
Is Multi-AZ failover instant? No, typically 60-120 seconds for classic RDS.
#rds#failover#jvm#dns
Q109
A bucket receives 5,000 PUT/s to one prefix and sees 503 SlowDown. What now?
advanced
S3 supports at least 3,500 PUT/COPY/POST/DELETE and 5,500 GET/HEAD requests per second per prefix, and scales by prefix. Spread keys across more prefixes (a hash or date-sharded prefix), retry with exponential backoff, and batch small objects.
⚠ Follow-up traps
Do you need random prefixes for performance? Not for sequential keys anymore but spreading load across prefixes raises the limit.
#s3#performance#throttling
Q110
A tenant's heavy usage degrades others in a shared-pool SaaS. How do you contain it?
advanced
Apply per-tenant rate limits and quotas at the gateway (API Gateway usage plans or custom token buckets), bulkhead resources (separate thread pools or queues per tier), shard heavy tenants to dedicated cells, and monitor per-tenant metrics.
Use SQS fair queuing or per-tenant message groups to stop one tenant starving others.
⚠ Follow-up traps
Is autoscaling a substitute? No, it raises everyone's cost and may overload shared databases.
#noisy-neighbor#multi-tenancy#throttling
Q111
How do you design a cell-based architecture and why?
advanced
Split the service into independent, identical cells (full stacks with their own data) serving a subset of customers, with a thin router in front. A failure or bad deploy affects only one cell, limiting blast radius, and capacity grows by adding cells.
Deploy changes to one cell first as a natural canary.
Cost: operational overhead and rebalancing tenants between cells.
⚠ Follow-up traps
Is the router a single point of failure? Keep it minimal, stateless and highly redundant.
#cell-based#blast-radius#high-availability
Q112
Your CI/CD pipeline role has admin rights and an attacker abuses a pull request. How do you harden it?
advanced
Use OIDC federation with roles trusting only specific repositories and branches via sub claim conditions, scope permissions to the deployment, separate roles per environment, require approvals for production, and prevent pull requests from forks accessing deploy roles.
⚠ Follow-up traps
Why not store keys as CI secrets? They are long-lived and exfiltratable by malicious build steps.
Does a permission boundary help? Yes, it caps what the pipeline can grant itself.
#security#ci-cd#iam
Q113
Compliance requires encrypting data and proving who accessed it. What do you enable?
intermediate
Encrypt at rest with KMS (customer-managed keys for control and audit) and in transit with TLS; enable CloudTrail (data events for S3 where needed) in all regions with log file validation, send logs to a locked-down separate account, and use AWS Config and Security Hub for continuous compliance.
⚠ Follow-up traps
Does enabling encryption at rest protect against a compromised IAM role? No, authorized principals get decrypted data.
#compliance#encryption#auditing
Q114
An engineer proposes running the whole system in a single region and single AZ "to save money". How do you respond?
intermediate
Quantify instead of refusing: compare the extra cost of a second AZ (a second small instance set, cross-AZ transfer, Multi-AZ DB at about 2x DB cost) with the revenue and reputation loss of downtime against the SLO. Most production systems justify multi-AZ; non-critical internal tools may not.
⚠ Follow-up traps
Can you get HA with one instance? Auto-recovery restores it, but you still have downtime.
Is multi-region the next step? Only if the business RTO/RPO demands it.
#trade-offs#high-availability#cost
Q115
A deploy to production passes in staging but fails with throttling on a cloud API. What happened?
advanced
Production scale hit service quotas or API rate limits (for example EC2 API calls, KMS requests per second, Lambda concurrency, or SES sending limits) that staging never reached.
Monitor Service Quotas, request increases ahead of launches, cache KMS data keys, and use SDK retries with backoff and jitter.
Load test at production-like scale.
⚠ Follow-up traps
Are all quotas adjustable? No, some are hard limits.
Is it safe to retry immediately? No, it worsens throttling; use backoff.