Courseiva

CCNA Design High-Performing Architectures Questions

75 of 215 questions · Page 1/3 · Design High-Performing Architectures · Answers revealed

1
MCQmedium

A read-heavy document portal repeatedly queries the same product catalogue data from DynamoDB with millisecond latency requirements. Which service can reduce read latency and table load? The architecture review board prefers a managed AWS-native control.

A.Amazon Kinesis Data Firehose
B.S3 Transfer Acceleration
C.DynamoDB Accelerator (DAX)
D.AWS Glue Data Catalog
AnswerC

DynamoDB Accelerator (DAX) is an in-memory cache that sits in front of DynamoDB, providing microsecond latency for repeated reads while maintaining eventual consistency by default. It automatically intercepts read and write operations, populates the cache on read misses, and supports TTL-based expiration, making it ideal for read-heavy workloads with repetitive access patterns. Because the portal repeatedly queries the same product data, DAX directly addresses the latency issue without changing application code beyond the DAX endpoint.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache for DynamoDB that delivers up to 10x read performance improvement, reducing read latency to microseconds for repeated queries. It offloads read traffic from the DynamoDB table, lowering consumed read capacity units and table load, making it ideal for read-heavy workloads with millisecond latency requirements. As a fully managed, AWS-native service, DAX aligns with the architecture review board's preference for managed controls.

Exam trap

The trap here is that candidates may confuse DAX with ElastiCache (which is also a caching service but not DynamoDB-native) or assume that any AWS caching service works interchangeably, but DAX is the only managed, DynamoDB-specific cache that integrates directly with the DynamoDB API without application code changes.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Firehose is a streaming data ingestion service for loading data into data stores and analytics tools, not a caching layer for DynamoDB reads; it cannot reduce read latency or table load for repeated queries. Option B is wrong because S3 Transfer Acceleration speeds up uploads and downloads to/from S3 over long distances using AWS edge locations, but it does not cache DynamoDB data or reduce read latency for DynamoDB queries. Option D is wrong because AWS Glue Data Catalog is a metadata repository for ETL jobs and data lake schemas, not a caching service for DynamoDB reads; it has no impact on DynamoDB read latency or table load.

2
MCQeasy

A team runs a stateless web app on Amazon EC2 behind an Application Load Balancer. During traffic spikes, new EC2 instances take several minutes to finish bootstrapping before they can receive traffic. Which Auto Scaling configuration most directly reduces the time until additional capacity is available?

A.Increase the ALB target group deregistration delay.
B.Use an Auto Scaling warm pool so pre-initialized instances are ready to enter service.
C.Reduce the Auto Scaling group minimum size to one instance.
D.Replace the Application Load Balancer with a Network Load Balancer.
AnswerB

A warm pool keeps pre-initialised instances in a stopped or running state, so scaling out moves already-bootstrapped capacity into service instead of waiting several minutes for bootstrap, directly cutting the delay before new instances receive traffic.

Why this answer

An Auto Scaling warm pool allows you to maintain a pool of pre-initialized instances that are ready to quickly enter the target group and start serving traffic. Instead of waiting for new instances to boot and configure during a scale-out event, the warm pool provides instances that have already completed bootstrapping, drastically reducing the time to additional capacity.

Exam trap

The trap here is that candidates may confuse the deregistration delay (which handles graceful connection draining) with a mechanism to speed up instance readiness, or they may incorrectly assume that reducing the minimum size or switching to a Network Load Balancer will improve scaling speed, when neither addresses the root cause of slow bootstrapping.

Why the other options are wrong

A

Increasing the deregistration delay only keeps existing connections alive longer; it does not speed up the bootstrapping of new instances, so it does not reduce the time until additional capacity is available.

C

Reducing the minimum size to one instance does not address the bootstrapping delay; it only lowers the baseline capacity, potentially worsening performance during traffic spikes.

D

Replacing the ALB with a Network Load Balancer does not address the bootstrapping delay of EC2 instances; NLB operates at layer 4 and does not affect instance initialization time.

3
MCQmedium

A analytics dashboard uses RDS MySQL and receives many read-only reporting queries that slow down the primary database. What should the architect add?

A.S3 lifecycle policy
B.RDS read replica and route reporting queries to it
C.Multi-AZ standby and route reads to the standby
D.A larger NAT gateway
AnswerB

An RDS read replica is a separate MySQL instance that uses asynchronous replication from the primary database. Routing reporting queries to the replica's endpoint shifts the read IOPS and CPU burden away from the primary, which continues to handle writes and can scale to support higher read throughput. This is the standard AWS pattern for offloading read-only queries in a MySQL environment, though it requires updating application data-source configuration to direct those queries appropriately.

Why this answer

Adding an RDS read replica offloads read-heavy reporting queries from the primary MySQL instance, preserving write performance. The read replica asynchronously replicates data using MySQL's native binlog replication, and routing reporting queries to its endpoint reduces contention on the primary.

Exam trap

The trap here is confusing a Multi-AZ standby (which is for high availability only and cannot serve reads) with a read replica (which is explicitly designed to offload read traffic).

How to eliminate wrong answers

Option A is wrong because S3 lifecycle policies manage object transitions and expirations in S3, not database query offloading. Option C is wrong because a Multi-AZ standby is a synchronous replica used only for failover; it does not serve read traffic (RDS does not allow direct reads from the standby). Option D is wrong because a larger NAT gateway increases outbound internet bandwidth for private subnets, which does not address database read query performance.

4
MCQeasy

Based on the exhibit, what change best reduces Lambda cold-start impact for a predictable user-upload workflow?

A.Set a reserved concurrency limit for the function to protect it from throttling.
B.Enable provisioned concurrency for the function.
C.Increase the function timeout to give more time for initialization.
D.Move the function to a larger memory setting only to eliminate all initialization time.
AnswerB

Provisioned concurrency keeps a pre-initialized pool of Lambda execution environments ready to respond immediately. The exhibit shows long init duration after inactivity, which is the classic symptom of cold starts affecting user experience. Because the traffic pattern is predictable during launches, provisioned concurrency is the most direct way to reduce startup latency and smooth response times.

Why this answer

Provisioned concurrency pre-warms a specified number of execution environments so that when a user upload triggers the Lambda function, there is no cold-start latency. This is the most direct way to eliminate initialization time for a predictable workload, as it keeps instances ready to handle requests immediately.

Exam trap

The trap here is that candidates often confuse reserved concurrency (which limits concurrency) with provisioned concurrency (which pre-warms instances), or they assume that increasing memory or timeout will solve cold starts, when in fact only provisioned concurrency directly addresses initialization latency for predictable workloads.

How to eliminate wrong answers

Option A is wrong because reserved concurrency only caps the maximum number of concurrent executions to prevent throttling; it does not pre-warm instances or reduce cold-start impact. Option C is wrong because increasing the function timeout does not affect initialization time; it only extends the maximum duration a function can run, which does not address cold starts. Option D is wrong because moving to a larger memory setting can reduce initialization time by providing more CPU and resources, but it does not eliminate all initialization time, and it is not as targeted or effective as provisioned concurrency for predictable workloads.

5
Multi-Selectmedium

A company is designing a high-performance database architecture for an e-commerce platform that experiences rapid spikes in read traffic during flash sales. The database must handle millions of reads per second with sub-millisecond latency. The data is key-value in nature, with a small number of attributes per item. Which three options should be included in the architecture? (Choose three.)

Select 3 answers
.Amazon DynamoDB as the primary database.
.Amazon RDS for MySQL with Multi-AZ and Read Replicas.
.DynamoDB Accelerator (DAX) as an in-memory cache.
.Amazon ElastiCache for Redis with cluster mode enabled.
.Amazon S3 as a primary data store accessed via Select and Range queries.
.Amazon Redshift with auto-scaling for real-time reads.

Why this answer

Amazon DynamoDB is a fully managed NoSQL key-value database that delivers single-digit millisecond latency at any scale, making it ideal for high-traffic e-commerce platforms with key-value data. DynamoDB Accelerator (DAX) is an in-memory cache that sits in front of DynamoDB, reducing read latency to microseconds for millions of reads per second. Amazon ElastiCache for Redis with cluster mode enabled provides a distributed in-memory cache that can offload read traffic from the primary database, further reducing latency and handling spikes during flash sales.

Exam trap

The trap here is that candidates often choose Amazon RDS with Read Replicas for read scaling, but they fail to recognize that relational databases cannot achieve sub-millisecond latency for millions of reads per second, and that DynamoDB with caching layers is the correct high-performance key-value solution.

6
MCQeasy

A development team is building a new application that stores session state in a relational database. The application experiences unpredictable read traffic, and the team wants a fully managed database that can scale read capacity automatically and provide a reader endpoint that distributes connections across multiple replicas. Which AWS service should a solutions architect recommend?

A.Amazon Aurora with Aurora Replicas and the reader endpoint.
B.Amazon Redshift with concurrency scaling enabled.
C.Amazon DynamoDB with on-demand capacity mode.
D.Amazon RDS for SQL Server with Multi-AZ deployment.
AnswerA

Aurora is a fully managed relational database compatible with MySQL and PostgreSQL, and Aurora Replicas can be added and scaled to handle read traffic. The cluster reader endpoint automatically distributes connections across the available Aurora Replicas, which is exactly the behavior described. It provides automatic storage scaling and managed replication, making it the correct relational, read-scalable choice for this scenario.

Why this answer

Aurora is a managed relational database whose Aurora Replicas serve read traffic, and the cluster reader endpoint spreads incoming read connections across those replicas automatically. That directly matches the need for automatic read scaling with a reader endpoint. DynamoDB is non-relational, RDS Multi-AZ standbys do not serve reads, and Redshift is an analytics warehouse rather than an OLTP database.

Exam trap

The trap here is assuming a Multi-AZ standby in RDS can absorb read traffic, when it exists only for failover and does not serve application reads.

7
MCQhard

A Lambda-based retail API has unpredictable traffic spikes and users see latency caused by cold starts. The function must respond consistently during expected campaign windows. What should be configured? The design must avoid adding custom operational scripts.

A.A larger deployment package
B.Reserved concurrency only
C.Provisioned concurrency during campaign windows
D.CloudTrail data events
AnswerC

Provisioned concurrency is the correct choice because it instructs Lambda to pre-initialize a specified number of execution environments before traffic arrives, eliminating cold-start latency for those warm instances. During known campaign windows, you can adjust the provisioned concurrency level to align with expected demand. This allows the retail API to serve requests instantly during a traffic surge instead of waiting for environment creation. You pay for the pre-warmed capacity even when idle, but it directly solves the stated low-latency requirement.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, eliminating cold starts during campaign windows. This ensures consistent response times for the Lambda-based retail API under unpredictable traffic spikes without requiring custom scripts.

Exam trap

The trap here is confusing reserved concurrency (which prevents throttling but does not address cold starts) with provisioned concurrency (which eliminates cold starts by keeping environments warm).

How to eliminate wrong answers

Option A is wrong because a larger deployment package increases cold start duration, making latency worse. Option B is wrong because reserved concurrency only caps the maximum number of concurrent executions to prevent throttling, but does not pre-warm environments to avoid cold starts. Option D is wrong because CloudTrail data events record API activity for auditing, not performance optimization.

8
MCQmedium

A team serves static assets from an S3 origin through CloudFront. Cache hit ratio is low. Analytics show that requests include an Authorization header (even though the assets are public) and the cache key currently varies on that header, causing CloudFront to treat the same asset as different cache entries. What is the best change to improve cache hit ratio without breaking access controls?

A.Keep Authorization in the CloudFront cache key, but increase the origin response minimum TTL to 1 day.
B.Modify the CloudFront cache policy so the cache key does not include the Authorization header.
C.Switch the S3 origin from the current bucket to a website endpoint to enable automatic caching headers.
D.Enable CloudFront to forward all headers to S3 so origin can decide caching behavior per request.
AnswerB

CloudFront cache hit ratio depends on what constitutes a unique cache key. If Authorization is included, identical public assets requested with different Authorization values will map to different cache objects and reduce reuse. Removing Authorization from the cache key makes those requests share the same edge cache entry, improving hit ratio and reducing origin traffic. Because the scenario states the assets are public, removing Authorization from the cache key does not break access controls (access is not controlled by Authorization at the origin).

Why this answer

The low cache hit ratio is caused by the Authorization header being included in the CloudFront cache key, which creates separate cache entries for the same object even though the assets are public. By modifying the cache policy to exclude the Authorization header, CloudFront will treat all requests for the same asset as identical, dramatically improving the cache hit ratio without affecting access controls because the assets are already public.

Exam trap

The trap here is that candidates may think increasing TTL or changing the origin type will fix caching, when the real issue is the cache key composition—specifically, the Authorization header fragmenting the cache.

How to eliminate wrong answers

Option A is wrong because increasing the minimum TTL does not address the root cause—the cache key still varies on the Authorization header, so separate cache entries will persist and the cache hit ratio will remain low. Option C is wrong because switching to an S3 website endpoint does not change how CloudFront caches based on headers; the cache key is still controlled by the CloudFront cache policy, not the origin type. Option D is wrong because forwarding all headers to S3 would include the Authorization header in the cache key, making the problem worse by further fragmenting the cache.

9
MCQmedium

A read-heavy media archive repeatedly queries the same product catalogue data from DynamoDB with millisecond latency requirements. Which service can reduce read latency and table load? The architecture review board prefers a managed AWS-native control.

A.DynamoDB Accelerator (DAX)
B.Amazon Kinesis Data Firehose
C.AWS Glue Data Catalog
D.S3 Transfer Acceleration
AnswerA

DynamoDB Accelerator is an in-memory cache that sits in front of DynamoDB and returns items for repeated identical queries at single-digit-millisecond latency. Because this media archive repeatedly queries the same product, DAX caches those hot items and queries, reducing read latency and decreasing the read capacity units consumed by DynamoDB. It is purpose-built to accelerate DynamoDB reads while maintaining API compatibility.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache for DynamoDB that delivers microsecond read latency, directly addressing the millisecond requirement. By caching frequently accessed product catalogue data, DAX offloads read requests from the DynamoDB table, reducing table load and read capacity unit consumption. As a fully managed, AWS-native service, it aligns with the architecture review board's preference for managed controls.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration (which optimizes uploads to S3) with a caching solution for DynamoDB, or mistakenly think Glue Data Catalog or Kinesis Firehose can cache database queries, when only DAX provides in-memory acceleration for DynamoDB reads.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Firehose is a streaming data ingestion service for loading data into data lakes or analytics tools, not a caching or read-latency reduction solution for DynamoDB. Option C is wrong because AWS Glue Data Catalog is a metadata repository for ETL and data discovery, not a cache that accelerates DynamoDB read queries. Option D is wrong because S3 Transfer Acceleration uses AWS edge locations to speed up uploads to S3 over long distances, but it does not cache DynamoDB data or reduce read latency for repeated queries.

10
MCQmedium

A global video platform serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most? The team wants the control to be enforceable during normal operations.

A.A larger S3 bucket
B.Amazon CloudFront distribution with the S3 bucket as origin
C.RDS read replicas
D.An EC2 Auto Scaling group in one Region
AnswerB

Amazon CloudFront is a global content delivery network (CDN) that caches static objects from the S3 origin at edge locations geographically close to viewers. When a user requests an image or 3D asset, CloudFront serves it from the nearest edge cache, avoiding a round trip to the S3 bucket's origin Region, which dramatically reduces latency and improves throughput. CloudFront also supports features like origin shielding, signed URLs, and compression, making it the recommended service for accelerating delivery of static S3-hosted content to a worldwide audience.

Why this answer

Amazon CloudFront is a content delivery network (CDN) that caches static content (images, JavaScript) at edge locations worldwide, reducing latency for users in distant countries. By using the S3 bucket as an origin, CloudFront serves cached copies from the nearest edge, drastically improving load times. This solution is enforceable during normal operations because CloudFront provides cache control headers and invalidation APIs to manage content freshness.

Exam trap

The trap here is that candidates may confuse scaling compute (EC2 Auto Scaling) or database (RDS read replicas) with content delivery, failing to recognize that static content performance is solved by a CDN like CloudFront, not by scaling backend resources.

How to eliminate wrong answers

Option A is wrong because increasing the S3 bucket size does not reduce latency; S3 is a regional service and does not cache content globally. Option C is wrong because RDS read replicas are for database read scaling, not for serving static files from S3. Option D is wrong because an EC2 Auto Scaling group in one Region does not address global latency; it only scales compute capacity in a single geographic area, leaving distant users unaffected.

11
MCQmedium

A financial services company runs a high-traffic REST API on Amazon EC2 instances behind an Application Load Balancer. The API retrieves user session data from an Amazon DynamoDB table for every request. During peak hours, DynamoDB read capacity is exhausted, causing throttling and increased latency. The workload is read-heavy and the session data is accessed frequently but changes infrequently. The solutions architect needs to reduce DynamoDB read load and improve API response times with minimal application changes. Which solution meets these requirements?

A.Migrate the session data to Amazon ElastiCache for Redis and modify the application to read from Redis instead of DynamoDB.
B.Create a read replica of the DynamoDB table in another region and direct all reads to the replica.
C.Increase the DynamoDB read capacity to a higher provisioned level and enable auto scaling.
D.Enable DynamoDB Accelerator (DAX) for the table and modify the application to use the DAX client for session reads.
AnswerD

DAX is an in-memory cache for DynamoDB that provides microsecond read latency and reduces the number of read capacity units consumed. It requires minimal code changes: the application uses the DAX client instead of the standard DynamoDB client. Since session data is read frequently and changes infrequently, DAX is ideal. It offloads read traffic from the table, preventing throttling and improving API response times.

Why this answer

DynamoDB Accelerator (DAX) is a fully managed, highly available in-memory cache for DynamoDB that delivers up to a 10x performance improvement, from milliseconds to microseconds. It is API-compatible with DynamoDB, so the application only needs to use the DAX client. For read-heavy workloads with infrequent updates, DAX reduces read capacity consumption and improves response times without requiring a major rewrite or data migration.

Exam trap

The trap here is assuming that simply increasing read capacity or adding a read replica will reduce latency, when the real need is to offload repeated reads from the table with an in-memory cache.

12
MCQmedium

A global video platform serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most? The design must avoid adding custom operational scripts.

A.A larger S3 bucket
B.Amazon CloudFront distribution with the S3 bucket as origin
C.RDS read replicas
D.An EC2 Auto Scaling group in one Region
AnswerB

Amazon CloudFront distribution with the S3 bucket as origin is the correct solution because CloudFront caches static images and video at edge locations around the world, dramatically reducing the distance data must travel to reach users. When a user requests content, CloudFront serves it from the nearest edge cache, and only fetches from the S3 origin on a cache miss, which also offloads throughput from S3 and lowers recurring costs. CloudFront integrates securely with S3 via Origin Access Control (OAC), ensuring that the bucket stays private while the CDN delivers content globally.

Why this answer

Amazon CloudFront is a content delivery network (CDN) that caches static content (images, JavaScript files) at edge locations worldwide. By distributing content closer to users, it reduces latency and improves load times for distant countries without requiring any custom operational scripts or changes to the S3 bucket.

Exam trap

The trap here is that candidates may think increasing S3 bucket size or adding compute resources (EC2, RDS) can solve latency issues, but the correct solution is a CDN like CloudFront that brings content physically closer to users.

How to eliminate wrong answers

Option A is wrong because increasing the size of an S3 bucket does not improve performance; S3 bucket size has no impact on latency or throughput for static content delivery. Option C is wrong because RDS read replicas are designed to offload read traffic from a relational database, not to accelerate delivery of static files stored in S3. Option D is wrong because an EC2 Auto Scaling group in a single Region does not reduce latency for users in distant countries; it only provides compute scaling within one geographic area, not global edge caching.

13
MCQhard

Based on the exhibit, a batch-processing service runs on Amazon EC2. The workload is Linux-based, can run on ARM64, and is CPU-bound during its nightly processing window. The team wants the best throughput per dollar without changing the application logic. Which EC2 instance family should the solutions architect recommend?

A.C7g instances based on AWS Graviton processors
B.R7i instances because more memory will improve CPU-bound job throughput.
C.M7a instances because general-purpose families are always the safest performance choice.
D.T3 instances because burstable instances can handle occasional nighttime spikes at lower cost.
AnswerA

C7g instances are compute optimized and use Graviton processors, which often deliver strong price-performance for CPU-bound Linux workloads that can run on ARM64. The exhibit shows the application is compatible and even benchmarks faster on ARM.

Why this answer

The C7g instances are based on AWS Graviton processors (ARM64 architecture), which offer up to 25% better performance per dollar compared to x86-based instances for CPU-bound workloads. Since the workload is Linux-based, can run on ARM64, and is CPU-bound, the C7g family provides the best throughput per dollar without requiring any application logic changes.

Exam trap

The trap here is that candidates may choose memory-optimized or general-purpose instances (like R7i or M7a) thinking they are safer, or burstable instances (T3) assuming they handle spikes cheaply, without recognizing that compute-optimized ARM64 instances (C7g) provide the best throughput per dollar for CPU-bound, ARM64-compatible workloads.

How to eliminate wrong answers

Option B is wrong because R7i instances are memory-optimized, designed for workloads that require large amounts of memory, not for CPU-bound jobs where additional memory does not improve throughput. Option C is wrong because M7a instances are general-purpose and balance compute, memory, and networking, but they are not optimized for CPU-bound workloads and use x86 architecture, which typically offers lower performance per dollar compared to ARM64-based instances for this specific scenario. Option D is wrong because T3 instances are burstable and designed for workloads with low baseline CPU usage and occasional spikes, but they are not suitable for sustained CPU-bound processing during a nightly window, as they would exhaust CPU credits and incur performance throttling or additional costs.

14
MCQmedium

A media analytics company ingests a continuous stream of JSON clickstream events, roughly 20,000 records per second, into an Amazon Kinesis Data Streams stream with 32 shards. Downstream consumers must be able to re-read the same records up to 7 days later to rebuild a reporting index. Which combination of settings should the team use to maximize the number of records each consumer can read per second while preserving this replay capability?

A.Switch the stream to on-demand capacity mode, keep the retention period at 24 hours, and have consumers read with the Kinesis Client Library because it partitions reads across shards with no per-shard limit.
B.Increase the shard count to 64, set the data retention period to 168 hours, and have each consumer use enhanced fan-out with its own dedicated 2 MiB/s read throughput per shard.
C.Increase the shard count to 64, set the retention period to 168 hours, and rely on the shared GetRecords polling model because it automatically scales to 10 MiB/s per shard when many consumers register.
D.Keep 32 shards, set the retention period to 24 hours, and have consumers poll with the GetRecords API using a single shared throughput budget of 2 MiB/s per shard.
AnswerB

Adding shards raises the aggregate write and read ceiling, the 168-hour retention period keeps records replayable for a full week, and enhanced fan-out gives each registered consumer a dedicated 2 MiB/s per-shard pipe rather than sharing the 2 MiB/s per-shard limit. This directly satisfies both the throughput and the 7-day replay requirements.

Why this answer

The scenario needs higher aggregate read throughput plus a full week of replayable data. Enhanced fan-out is the only mechanism that gives each consumer a dedicated 2 MiB/s per-shard read pipe instead of sharing the 2 MiB/s per-shard polling budget, and the 168-hour retention setting preserves records for the required rebuild window. Adding shards raises the aggregate ceiling further.

Exam trap

The trap here is assuming that adding more registered consumers to a standard Kinesis Data Streams stream increases the shared 2 MiB/s per-shard read throughput limit.

15
Multi-Selectmedium

An Aurora PostgreSQL application has an OLTP writer and a reporting dashboard that issues many read-only queries. The writer is healthy, but read latency rises noticeably during reporting windows. Which two changes should you make? Select two.

Select 2 answers
A.Add Aurora Replicas to scale out the read workload.
B.Send read-only application traffic to the reader endpoint.
C.Scale up only the writer instance and keep all queries on it.
D.Replace the cluster with a single-AZ RDS instance to reduce replication overhead.
E.Move the dashboard to DynamoDB without changing the query model.
AnswersA, B

Aurora Replicas are independent compute instances in the same Aurora cluster that share the underlying storage volume. By adding one or more replicas, you create additional read endpoints that can absorb dashboard queries and other read-only traffic, directly offloading the writer instance. Because Aurora's storage is distributed and replicated separately, adding replicas does not cause significant write overhead, making horizontal read scaling the most efficient and cost-effective solution for an OLTP workload with heavy reads.

Why this answer

Adding Aurora Replicas (Option A) is correct because Aurora Replicas are dedicated read-only instances that share the same underlying storage volume as the writer, allowing you to scale read capacity linearly without impacting write performance. Sending read-only traffic to the reader endpoint (Option B) is correct because the reader endpoint automatically load-balances connections across all available Aurora Replicas, ensuring that dashboard queries are distributed and do not overload a single instance.

Exam trap

The trap here is that candidates may think scaling up the writer instance (Option C) is sufficient, but they overlook that read-heavy workloads require horizontal read scaling via replicas, not just vertical scaling of the writer.

Why the other options are wrong

C

Scaling up the writer instance does not offload read queries; the writer still handles all traffic, so read latency remains high during reporting windows. Aurora Replicas are needed to distribute read-only queries.

E

Moving the dashboard to DynamoDB without changing the query model is wrong because DynamoDB is a NoSQL database with a different query model (key-value and document), so existing SQL queries from the reporting dashboard would not work without significant application changes.

16
MCQhard

Based on the exhibit, a media company serves versioned JavaScript and CSS files from an Amazon S3 origin through CloudFront. After a frontend release, the cache hit ratio dropped sharply even though the file names are versioned. The application team says the browser requests include the same Authorization header on every asset request because the frontend and API share one domain. What should the solutions architect do to improve CloudFront cache hit ratio without changing the application authentication model for the API?

A.Enable S3 Transfer Acceleration on the bucket so CloudFront fetches objects faster from the origin.
B.Create a CloudFront cache policy that excludes Authorization, cookies, and unnecessary query strings from the cache key.
C.Switch the origin from S3 to an Application Load Balancer so CloudFront can cache dynamic responses more effectively.
D.Configure CloudFront to forward every viewer header to the origin so the origin can decide whether the content is cacheable.
AnswerB

This reduces cache fragmentation because CloudFront can reuse the same cached object for many viewers. Since the assets are immutable and versioned, the Authorization header is not needed to vary the cache for these files. Keeping API authentication separate preserves the application model while improving hit ratio.

Why this answer

The sharp drop in cache hit ratio is caused by the Authorization header being included in the cache key, which makes each request unique even though the file names are versioned. By creating a CloudFront cache policy that excludes the Authorization header (and unnecessary cookies/query strings) from the cache key, CloudFront can serve cached responses to requests with different Authorization headers, restoring the cache hit ratio without altering the application's authentication model for the API.

Exam trap

The trap here is that candidates may think the Authorization header is required for caching or that forwarding all headers is safe, but in reality, including it in the cache key destroys cache efficiency for static assets, and the correct solution is to exclude it via a cache policy.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration improves upload/download speed over long distances but does not affect CloudFront's cache key or hit ratio. Option C is wrong because switching to an Application Load Balancer would not solve the cache key issue; ALB is for dynamic content and would not improve caching for static versioned files served from S3. Option D is wrong because forwarding every viewer header to the origin would include the Authorization header in the cache key, making each request unique and further reducing the cache hit ratio, which is the opposite of what is needed.

17
MCQeasy

A DynamoDB-backed multi-tenant app experiences throttling during a promotion. Most writes and reads target tenant "ACME" and use the same partition key value, causing a hot partition. Which design change most directly improves performance?

A.Add a "shard" component to the partition key (for example, tenantId + hashed bucket) to spread traffic across partitions
B.Increase the table’s read capacity without changing the partition key
C.Switch all reads to strongly consistent reads to guarantee faster results
D.Store ACME data in S3 and query it directly to avoid DynamoDB throttling
AnswerA

DynamoDB throughput is distributed across physical partitions. If one partition key value receives most traffic, that partition throttles. Adding a shard component to the partition key increases the number of partition key values being used, spreading requests across more partitions and reducing hot-partition throttling.

Why this answer

Adding a shard component to the partition key (e.g., appending a random or hash-based suffix to the tenant ID) distributes writes and reads for the same tenant across multiple physical partitions. This directly alleviates the hot partition caused by all ACME traffic hitting a single partition key value, allowing DynamoDB to utilize its full provisioned throughput across partitions.

Exam trap

The trap here is that candidates may think increasing total table capacity (Option B) solves throttling, but they overlook that DynamoDB throttles at the partition level, not the table level, so a single hot partition remains constrained regardless of total capacity.

How to eliminate wrong answers

Option B is wrong because increasing the table’s read capacity does not fix the hot partition issue—DynamoDB distributes throughput evenly across partitions, so a single partition can still throttle even if total table capacity is high. Option C is wrong because strongly consistent reads do not improve performance; they are slower and consume more read capacity units than eventually consistent reads, and they do not spread traffic across partitions. Option D is wrong because storing ACME data in S3 and querying it directly bypasses DynamoDB’s low-latency access patterns and introduces additional complexity (e.g., S3 eventual consistency, lack of native querying), making it an inefficient and indirect solution for a hot partition problem.

18
MCQmedium

A production application writes to an Amazon Aurora PostgreSQL cluster. Users report that during business-hour reporting runs, write latency increases. The application team wants to keep the writer focused on OLTP writes while still providing low-latency reads for reporting queries. What architectural approach should the solutions architect recommend?

A.Create Aurora read replicas and direct reporting read-only connections to the cluster reader endpoint.
B.Resize the writer instance to a larger class so it can handle both writes and reads with fewer slowdowns.
C.Enable cross-region replication for the entire cluster so reporting always runs in the secondary Region.
D.Disable read replicas and use caching only in the application layer, keeping all queries connected to the writer endpoint.
AnswerA

Aurora read replicas are separate DB instances that share the same underlying storage volume, so they serve read-only traffic without adding load to the writer. The cluster reader endpoint automatically load-balances connections across all replicas, allowing reporting queries to run in parallel with production writes. This decouples read and write workloads, reducing contention on the writer and improving overall responsiveness. It also lets you scale read capacity independently by adding or resizing replicas.

Why this answer

A is correct because creating Aurora read replicas and directing reporting read-only connections to the cluster reader endpoint offloads read traffic from the writer instance. This allows the writer to focus on OLTP writes, while the reader endpoint load-balances read-only queries across replicas, providing low-latency reads for reporting without impacting write performance.

Exam trap

The trap here is that candidates may think resizing the writer instance (Option B) is sufficient, but the exam tests the architectural principle of separating read and write workloads to avoid resource contention, not just scaling vertically.

Why the other options are wrong

B

Resizing the writer instance to a larger class does not offload read traffic from the writer; reporting queries still compete with OLTP writes on the same instance, failing to isolate workloads and reduce write latency.

C

Cross-region replication does not reduce latency for reporting queries in the primary region; it creates a separate cluster in another region, which would not help with low-latency reads for local reporting and introduces additional cost and complexity.

D

Disabling read replicas and using only application-layer caching forces all reporting queries through the writer endpoint, increasing write latency during reporting runs. This contradicts the goal of offloading reads to keep the writer focused on OLTP writes.

19
MCQmedium

A telemetry pipeline uses RDS MySQL and receives many read-only reporting queries that slow down the primary database. What should the architect add? The architecture review board prefers a managed AWS-native control.

A.Multi-AZ standby and route reads to the standby
B.RDS read replica and route reporting queries to it
C.S3 lifecycle policy
D.A larger NAT gateway
AnswerB

Creating an RDS read replica is the correct approach because it provisions a separate, read-only MySQL instance that asynchronously replicates all changes from the primary database. Reporting and analytics queries can be pointed to the read replica's own endpoint, offloading read-heavy telemetry workloads from the primary and freeing its compute/IO capacity for writes. This pattern is purpose-built for scaling read throughput and matches the stated requirement to support read-only reporting traffic.

Why this answer

RDS read replicas are designed specifically to offload read-heavy workloads like reporting queries from the primary database. They provide an asynchronous read-only copy of the database that can handle SELECT statements without impacting the primary's write performance. This is a fully managed AWS-native solution that aligns with the architecture review board's preference.

Exam trap

The trap here is confusing Multi-AZ standby (which is for failover only) with read replicas (which are for read scaling), leading candidates to incorrectly choose Option A.

How to eliminate wrong answers

Option A is wrong because a Multi-AZ standby is a synchronous replica used for high availability and failover, not for read traffic; it cannot serve read queries directly. Option C is wrong because an S3 lifecycle policy manages object storage transitions and expiration, which is unrelated to offloading database read queries. Option D is wrong because a larger NAT gateway increases outbound internet capacity for private subnets, which does not address read query load on an RDS database.

20
MCQeasy

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application experiences variable traffic patterns, with sudden spikes during marketing campaigns. The operations team wants to ensure that the application can scale out quickly to handle the spikes and scale in when traffic decreases, while minimizing costs. Which solution should a solutions architect recommend?

A.Configure an Auto Scaling group with a target tracking scaling policy based on average CPU utilization.
B.Deploy the application on AWS Lambda behind an API Gateway and remove the EC2 instances.
C.Use a scheduled scaling policy that adds instances at the start of each marketing campaign.
D.Enable sticky sessions on the ALB and increase the instance size.
AnswerA

A target tracking scaling policy automatically adjusts the number of instances to maintain a specified metric, such as average CPU utilization, at a target value. This is ideal for variable traffic because it scales out when CPU rises and scales in when it falls. It helps handle sudden spikes and reduces costs during low traffic by terminating unnecessary instances.

Why this answer

A target tracking scaling policy with an Auto Scaling group is the most effective way to handle variable traffic. It automatically adjusts capacity based on CPU utilization, scaling out during spikes and scaling in during low traffic, which optimizes both performance and cost. Other options are either static, do not address scaling, or require major architectural changes.

Exam trap

The trap here is assuming that scheduled scaling is always sufficient for variable traffic, when in fact unpredictable spikes require dynamic scaling policies.

21
MCQeasy

Based on the exhibit, which EBS volume type should the team use to meet the performance need at lower cost than overprovisioning capacity?

A.Use gp3 and provision the needed IOPS independently of volume size.
B.Use sc1 because it is optimized for infrequent access and large objects.
C.Use st1 because it provides high throughput for streaming data.
D.Use standard magnetic storage because it is compatible with all EC2 instances.
AnswerA

gp3 is the best fit because it lets you provision IOPS and throughput separately from volume size. The exhibit shows the workload needs around 10,000 IOPS and experiences queue buildup on gp2. With gp3, the team can raise performance without unnecessarily increasing storage capacity, which is usually more cost-effective for this kind of database workload.

Why this answer

The team needs to meet performance requirements at lower cost than overprovisioning capacity. gp3 allows you to provision baseline performance of 3,000 IOPS and 125 MiB/s throughput for any volume size, and you can independently increase IOPS up to 16,000 and throughput up to 1,000 MiB/s without needing to increase volume size. This avoids the cost of overprovisioning large gp2 volumes to achieve higher IOPS, which are tied to volume size (3 IOPS per GiB).

Exam trap

The trap here is that candidates assume all EBS volume types require overprovisioning capacity to achieve higher IOPS, forgetting that gp3 decouples performance from size, making it the most cost-effective choice for workloads needing specific IOPS without large storage.

Why the other options are wrong

B

The question asks for lower cost than overprovisioning capacity, but sc1 is a cold HDD volume that cannot meet performance needs requiring IOPS independent of volume size; it has low IOPS and throughput limits unsuitable for consistent performance.

C

The question asks for a lower-cost solution than overprovisioning capacity, but st1 is a throughput-optimized HDD volume that requires provisioning storage to achieve baseline throughput, which can lead to overprovisioning. Additionally, the performance need likely involves IOPS, not just throughput, and st1 has low IOPS.

D

Standard magnetic storage (previous generation) does not offer the ability to provision IOPS independently of volume size, so it cannot meet the performance need at lower cost than overprovisioning capacity. It also lacks the performance and cost efficiency of gp3 for this use case.

22
MCQmedium

Your company needs a high-throughput, low-latency TCP service using a custom binary protocol. Requirements: preserve the original client source IP for rate limiting, keep latency minimal, and use TCP health checks. The current setup uses an Application Load Balancer and performance is inconsistent. Which load balancer choice best meets these requirements?

A.Keep the Application Load Balancer (ALB), because ALBs also preserve client source IP for TCP protocols.
B.Use a Network Load Balancer (NLB) with TCP listeners so traffic stays at Layer 4 and the original source IP is preserved.
C.Use Amazon API Gateway because it preserves client source IP and provides TCP health checks for all protocols.
D.Use Amazon CloudFront with an S3 origin, because CloudFront reduces latency for TCP-based protocols.
AnswerB

NLB is designed for Layer 4 TCP/UDP traffic with very low latency and high throughput. It supports TCP health checks and preserves the original client source IP by default, which enables accurate client-IP-based rate limiting for a custom TCP protocol.

Why this answer

A Network Load Balancer (NLB) operates at Layer 4 and preserves the original client source IP by default, which is essential for accurate rate limiting. Its TCP listeners provide low-latency, high-throughput handling of custom binary protocols, and it supports TCP health checks natively. This directly addresses the performance inconsistency seen with the Application Load Balancer, which operates at Layer 7 and introduces additional processing overhead.

Exam trap

The trap here is that candidates often assume Application Load Balancers preserve client source IP for all protocols, but they only do so for HTTP/HTTPS traffic via the X-Forwarded-For header, not for raw TCP traffic, and they introduce higher latency due to Layer 7 processing.

How to eliminate wrong answers

Option A is wrong because an Application Load Balancer operates at Layer 7 (HTTP/HTTPS) and does not preserve the original client source IP for TCP traffic; it terminates the client connection and re-establishes a new one, so the source IP seen by the backend is the ALB's private IP. Option C is wrong because Amazon API Gateway is a fully managed service for creating RESTful and WebSocket APIs, not a load balancer; it does not support TCP listeners or TCP health checks, and it operates at Layer 7. Option D is wrong because Amazon CloudFront is a content delivery network (CDN) that caches content at edge locations, but it does not support TCP-based custom binary protocols (it works with HTTP/HTTPS and WebSocket) and cannot use an S3 origin for a TCP service; it also does not preserve the original client source IP for TCP traffic.

23
Matchinghard

A company runs a stateless application tier behind an Application Load Balancer. Match each observed scaling pattern on the left to the best Auto Scaling strategy or metric on the right.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Scale the Auto Scaling group on ALB RequestCountPerTarget.

Scale on SQS queue depth using a custom CloudWatch metric.

Use scheduled scaling to add capacity before the recurring surge.

Use target tracking on EC2 CPUUtilization.

Why these pairings

Steady increase is best handled by step scaling for gradual adjustments; sudden spikes use simple scaling for immediate action; cyclical patterns benefit from scheduled scaling; consistent low traffic may not need scaling; unpredictable bursts are managed by target tracking to maintain a metric; gradual decrease uses simple scaling to reduce capacity.

24
MCQeasy

A company is deploying a new web application on AWS. The application will serve static content (HTML, CSS, JavaScript, images) and dynamic API requests. The company expects a global user base and wants to minimize latency for all users. The static content is stored in an Amazon S3 bucket, and the dynamic APIs are hosted on Amazon EC2 instances behind an Application Load Balancer. Which service should the company use to accelerate both static and dynamic content delivery?

A.Amazon Route 53 with latency-based routing to multiple Regions
B.Amazon CloudFront with the S3 bucket and ALB as origins
C.AWS Global Accelerator
D.AWS Direct Connect with a public virtual interface
AnswerB

Amazon CloudFront is a content delivery network that caches static content at edge locations and can also forward dynamic requests to the ALB origin. By configuring both the S3 bucket and the ALB as origins, CloudFront can serve static content from cache and dynamically route API calls, reducing latency for global users. This is the most effective solution for mixed content.

Why this answer

Amazon CloudFront is a global content delivery network that caches static content at edge locations and can also forward dynamic requests to an Application Load Balancer origin. By using CloudFront with both the S3 bucket and the ALB as origins, the company can accelerate delivery of static and dynamic content for a global user base, reducing latency and improving performance.

Exam trap

The trap here is confusing AWS Global Accelerator with a CDN; Global Accelerator does not cache content at the edge, so it is not the best choice for accelerating static content delivery.

25
Multi-Selecthard

A genomics research team stores about 400 TB of compressed sequence files in Amazon S3 and runs a distributed analysis on Amazon EC2 instances in the same Region. The analysis reads each file sequentially and writes intermediate results to local instance storage. The team reports that the S3 GET requests are a bottleneck and wants to improve read throughput while keeping data durable. (Choose two.)

Select 2 answers
A.Prefix the object keys with a hash or random value to spread requests across multiple S3 partitions.
B.Use S3 byte-range fetches to retrieve each object in parallel parts across multiple threads or instances.
C.Mount the bucket with an S3 File Gateway and read the files over NFS from the EC2 fleet.
D.Enable S3 Versioning on the bucket so concurrent readers can access older object versions.
E.Configure the bucket for S3 Standard-Infrequent Access to lower the cost of frequent GET requests.
AnswersA, B

S3 scales request rates by partitioning an index by key prefix, and a hot prefix can throttle GET throughput. Distributing keys with a high-cardinality prefix spreads the load across many partitions, raising the request rate the bucket can sustain. This complements byte-range fetches and keeps objects durable in S3, directly improving the reported GET bottleneck.

Why this answer

S3 request performance is bounded by how well requests spread across the index partitions and by per-object concurrency. Randomizing key prefixes distributes GETs across many partitions so no single prefix throttles, and byte-range fetches let many threads or instances read one large object in parallel. Together they raise aggregate read throughput for the 400 TB analysis while the objects stay durable in S3.

Exam trap

The trap here is reaching for storage-class changes or gateway products to fix throughput, when the real levers are request distribution across partitions and parallel byte-range reads.

26
MCQmedium

A high-volume telemetry pipeline writes streaming click events that must be processed by multiple independent consumers. Which service is most appropriate?

A.Amazon Kinesis Data Streams
B.AWS DataSync
C.Amazon EBS
D.Amazon Route 53
AnswerA

Kinesis Data Streams retains ordered records for a configurable period, letting multiple independent consumers read the same stream at their own pace via separate iterators. This satisfies the stem's requirement for parallel, independent processing of high-volume click events.

Why this answer

Amazon Kinesis Data Streams is the most appropriate service because it is designed for real-time streaming data ingestion and processing. It can capture and store terabytes of data per hour from hundreds of thousands of sources, such as click events, and allows multiple independent consumers to read and process the same stream concurrently using the Kinesis Client Library (KCL) or enhanced fan-out with dedicated throughput.

Exam trap

The trap here is that candidates may confuse Kinesis Data Streams with simpler messaging services like SQS or SNS, but the key differentiator is that Kinesis supports multiple independent consumers processing the same stream in real-time with replay capability, whereas SQS is designed for point-to-point message delivery and SNS for pub/sub with push-based fan-out.

How to eliminate wrong answers

Option B (AWS DataSync) is wrong because it is a data transfer service for moving large datasets between on-premises storage and AWS services, not for real-time streaming or multiple consumer processing. Option C (Amazon EBS) is wrong because it provides block-level storage volumes for EC2 instances, not a streaming data ingestion or processing capability. Option D (Amazon Route 53) is wrong because it is a DNS web service for domain name resolution and routing, not for handling streaming telemetry data.

27
MCQmedium

A read-heavy document portal repeatedly queries the same product catalogue data from DynamoDB with millisecond latency requirements. Which service can reduce read latency and table load?

A.Amazon Kinesis Data Firehose
B.S3 Transfer Acceleration
C.DynamoDB Accelerator (DAX)
D.AWS Glue Data Catalog
AnswerC

DynamoDB Accelerator (DAX) is a fully managed, in-memory cache designed specifically for Amazon DynamoDB, delivering microsecond response times for repeated reads. When the document portal queries the same product key frequently, DAX serves subsequent identical queries from memory, drastically reducing read latency and offloading read capacity units from the table. It integrates seamlessly with the DynamoDB API, requiring only an endpoint change, and is the correct solution for this read-heavy, hot-key access pattern.

Why this answer

DynamoDB Accelerator (DAX) is a fully managed, in-memory cache for Amazon DynamoDB that delivers up to 10x read performance improvement by caching frequently accessed data. For a read-heavy workload querying the same product catalogue data, DAX reduces read latency to microseconds and offloads read requests from the DynamoDB table, lowering consumed read capacity units and table load.

Exam trap

The trap here is that candidates often confuse caching services (DAX) with data transfer acceleration (S3 Transfer Acceleration) or data ingestion (Kinesis Data Firehose), failing to recognize that the core requirement is to reduce DynamoDB read latency and table load, which only a dedicated in-memory cache like DAX can achieve.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Firehose is a streaming data ingestion service for loading data into data stores or analytics tools, not a caching layer for DynamoDB read operations. Option B is wrong because S3 Transfer Acceleration speeds up uploads and downloads to Amazon S3 over long distances using AWS edge locations, but it does not cache DynamoDB data or reduce read latency for DynamoDB queries. Option D is wrong because AWS Glue Data Catalog is a metadata repository for data assets used in ETL and analytics, not a cache for DynamoDB read requests.

28
MCQhard

Based on the exhibit, which change best reduces latency during peak traffic without overprovisioning the fleet?

A.Replace the instances with a larger instance family so each server has more headroom.
B.Change the Auto Scaling policy to target tracking on ALB RequestCountPerTarget.
C.Use scheduled scaling to add instances only during the business hours peak window.
D.Replace the ALB with a Network Load Balancer to reduce request latency.
AnswerB

RequestCountPerTarget matches the actual demand reaching each instance and scales capacity before the thread pool saturates. Because CPU is still low, CPU-based scaling would react too late or not at all. Target tracking on request count helps keep queue depth and latency down while avoiding unnecessary overprovisioning during quieter periods.

Why this answer

Using a target tracking scaling policy on ALB RequestCountPerTarget dynamically adjusts the fleet size based on the actual load per instance, ensuring that capacity scales with demand during peak traffic without manual intervention or overprovisioning. This approach directly addresses latency caused by high request rates per instance by maintaining a target request count, which reduces response time without adding unnecessary instances during off-peak periods.

Exam trap

The trap here is that candidates confuse 'reducing latency' with 'improving network throughput' (Option D) or 'static capacity increases' (Option A), missing that dynamic scaling based on per-target request count directly addresses the latency caused by overloaded instances during peak traffic.

How to eliminate wrong answers

Option A is wrong because simply replacing instances with a larger family increases per-instance capacity but does not automatically scale the fleet; it leads to overprovisioning during low traffic and fails to adapt to variable peak loads, wasting cost and not reducing latency efficiently. Option C is wrong because scheduled scaling adds instances only during a fixed business hours window, which cannot handle unpredictable peak traffic spikes outside that window, leaving the fleet either under-provisioned or over-provisioned. Option D is wrong because replacing the ALB with a Network Load Balancer (NLB) reduces transport-layer latency but does not address the root cause of latency—high request load per target—and NLB lacks the application-layer metrics (like RequestCountPerTarget) needed for intelligent Auto Scaling based on request volume.

29
MCQmedium

A global video platform serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most? The architecture review board prefers a managed AWS-native control.

A.A larger S3 bucket
B.Amazon CloudFront distribution with the S3 bucket as origin
C.RDS read replicas
D.An EC2 Auto Scaling group in one Region
AnswerB

CloudFront caches the static images and JavaScript at edge locations worldwide, so distant users retrieve objects from nearby points of presence instead of the S3 origin region. This directly cuts latency, and it is a managed AWS-native service, satisfying the architecture review board's stated preference.

Why this answer

Amazon CloudFront is a global content delivery network (CDN) that caches static content (images, JavaScript) at edge locations closer to users, drastically reducing latency for distant countries. By using the S3 bucket as the origin, CloudFront offloads requests from S3 and accelerates delivery via HTTP/2, TCP optimizations, and persistent connections. This is the most effective managed AWS-native solution for improving global load times for static assets.

Exam trap

The trap here is that candidates may confuse scaling storage (larger bucket) or compute (Auto Scaling) with performance improvement, overlooking that latency for static content is primarily a network distance problem solved by a CDN like CloudFront.

How to eliminate wrong answers

Option A is wrong because increasing the S3 bucket size does not improve data transfer speed or reduce latency; S3 performance is independent of bucket size and is limited by the bucket's regional location. Option C is wrong because RDS read replicas are designed for scaling database read traffic, not for serving static files or accelerating HTTP content delivery. Option D is wrong because an EC2 Auto Scaling group in a single Region does not reduce latency for users in distant countries; it only provides regional scalability and fault tolerance, not global edge caching.

30
MCQmedium

A DynamoDB table stores device status items. The partition key is deviceId, and the partition distribution is healthy (no single partition dominates). However, during peak periods the application experiences high read latency because many clients repeatedly request the latest status for the same devices. Which action best improves read latency without changing the DynamoDB partitioning model?

A.Add Amazon DAX as a caching layer in front of DynamoDB and route repeated read operations through DAX.
B.Change the partition key to a random value for each request to eliminate hot partitions.
C.Increase write capacity only, because writes generally determine read latency in DynamoDB.
D.Create an additional Global Secondary Index (GSI) and read exclusively from the index to accelerate reads.
AnswerA

Amazon DAX is an in-memory caching layer for DynamoDB that accelerates repeated reads. When many clients request the same items (for example, “latest status” point reads by deviceId), DAX can serve cached responses directly, reducing round trips to DynamoDB and lowering read latency during peak periods.

Why this answer

Amazon DAX is a fully managed, in-memory cache for DynamoDB that provides microsecond read latency. By caching the results of repeated GetItem and Query requests for the same device status items, DAX offloads read traffic from the underlying DynamoDB table, reducing the number of read capacity units consumed and eliminating the latency caused by repeated fetches from disk. This directly addresses the high read latency during peak periods without altering the existing partition key or partitioning model.

Exam trap

The trap here is that candidates may think a GSI can magically speed up reads, but GSIs do not provide caching and still read from the same storage layer, so they do not reduce latency for repeated identical queries.

Why the other options are wrong

C

Increasing write capacity does not reduce read latency; read latency is affected by read capacity and throttling, not write capacity. The problem is high read demand on the same items, which write capacity cannot address.

D

Creating a GSI does not reduce read latency for repeated requests to the same items; it only provides an alternate query pattern. The hot partition issue is caused by high read frequency on specific items, which a GSI does not alleviate.

31
MCQmedium

A containerized service fleet running on EC2 instances needs to share user-uploaded files and access them with low latency. The workload is bursty: sometimes dozens of instances concurrently read the same directory for short periods, and then traffic drops. Which Amazon EFS configuration best matches these performance needs?

A.Use Amazon EFS General Purpose performance mode and Throughput mode set to Bursting.
B.Use Amazon EFS Max I/O performance mode with Throughput mode set to Provisioned.
C.Use Amazon EFS General Purpose performance mode with Throughput mode set to Provisioned.
D.Use Amazon EFS Max I/O performance mode with Throughput mode set to Bursting.
AnswerA

EFS General Purpose performance mode is designed for latency-sensitive use cases with a broad range of I/O sizes, including typical file-sharing and web-content workloads. Throughput mode Bursting provides baseline throughput and allows throughput to scale up during demand spikes, which matches the pattern of short read bursts from many instances. When traffic drops, the system returns to baseline without requiring you to provision peak throughput for all time.

Why this answer

The workload is bursty with concurrent reads of the same directory, which favors the General Purpose performance mode for its strong consistency and lower latency per operation. The Bursting Throughput mode is ideal for bursty traffic as it allows the file system to accumulate burst credits during idle periods and consume them during high-demand spikes, matching the described pattern without incurring additional costs.

Exam trap

The trap here is that candidates often assume Max I/O is always better for high concurrency, but they overlook that General Purpose mode provides lower latency and stronger consistency for directory-heavy bursty reads, which is the actual requirement.

How to eliminate wrong answers

Option B is wrong because Max I/O performance mode is designed for highly parallelized workloads (e.g., thousands of instances) but sacrifices consistency and can introduce higher per-operation latency, which is not suitable for low-latency access to the same directory. Option C is wrong because Provisioned Throughput mode is intended for steady-state throughput requirements and would waste cost on a bursty workload that could use Bursting mode's credit-based model. Option D is wrong because Max I/O performance mode is not optimal for low-latency, directory-heavy access patterns, and while Bursting mode fits the bursty nature, the combination with Max I/O undermines the low-latency requirement.

32
MCQeasy

A startup runs a read-heavy product catalogue API on Amazon DynamoDB. The table uses on-demand capacity mode, and the team notices that repeated queries for the same popular items return consistently, causing high read request charges. The application can tolerate data that is a few seconds stale. Which change reduces read costs while keeping latency low?

A.Create a global secondary index on the queried attributes and read from the index instead of the table.
B.Switch the table to provisioned capacity mode and purchase reserved capacity.
C.Increase the table's read capacity and enable auto scaling to handle the repeated queries.
D.Enable DynamoDB Accelerator (DAX) for the table and route the read queries through the DAX cluster.
AnswerD

DAX is an in-memory cache purpose-built for DynamoDB that serves repeated reads of the same items in microseconds and absorbs that traffic before it reaches the table. Because the application tolerates slightly stale data, cached reads are acceptable, and the reduction in table read requests directly lowers cost while keeping response times low.

Why this answer

A DAX cluster caches item reads in memory and serves repeated requests for the same popular items without touching the table, which cuts read request charges and keeps latency in the microsecond range. Because the application can accept data that is a few seconds stale, the cache's item TTL behavior is acceptable, making DAX the direct fit for this read-heavy, repeated-key pattern.

Exam trap

The trap here is reaching for capacity or index changes, which alter how reads are served by DynamoDB but never remove the repeated read requests that drive the cost.

33
MCQmedium

A global mobile game backend serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most? The design must avoid adding custom operational scripts.

A.RDS read replicas
B.Amazon CloudFront distribution with the S3 bucket as origin
C.A larger S3 bucket
D.An EC2 Auto Scaling group in one Region
AnswerB

An Amazon CloudFront distribution with the S3 bucket as origin places your static images and other objects at a global network of edge locations, so users receive data from the nearest point of presence rather than traversing the internet to the origin bucket. CloudFront caches objects based on cache-control headers or default TTLs, dramatically reducing latency and origin load, and it integrates with S3 via Origin Access Control to keep the bucket private while still serving public content. This directly solves the global latency problem described in the scenario.

Why this answer

Amazon CloudFront is a content delivery network (CDN) that caches static content at edge locations worldwide, reducing latency for users in distant countries by serving files from the nearest edge. Using the S3 bucket as the origin, CloudFront distributes the content globally without requiring any custom operational scripts, directly addressing the slow load times for static assets.

Exam trap

The trap here is that candidates may confuse database read replicas (Option A) with content delivery, not realizing that static asset acceleration requires a CDN like CloudFront, not a database scaling solution.

How to eliminate wrong answers

Option A is wrong because RDS read replicas are designed to offload read traffic from a relational database, not to accelerate delivery of static files stored in S3. Option C is wrong because increasing the S3 bucket size does not improve performance; S3 performance is independent of bucket size and does not reduce latency for distant users. Option D is wrong because an EC2 Auto Scaling group in a single Region does not provide global edge caching; it only scales compute capacity within one geographic area, failing to reduce latency for users in other regions.

34
MCQhard

A DynamoDB table for a travel booking site has a partition key based only on the current date. Write throttling occurs during business hours. What is the best design change? The design must avoid adding custom operational scripts.

A.Create a global secondary index with the same date key
B.Move the table to S3 Glacier Instant Retrieval
C.Reduce the table's write capacity
D.Use a higher-cardinality partition key that distributes writes across partitions
AnswerD

A date-only partition key funnels every write into one partition, exhausting its throughput ceiling. Switching to a high-cardinality key spreads writes across many partitions, removing the hot-partition throttling without needing custom scripts or operational workarounds.

Why this answer

Using a low-cardinality partition key like the current date causes all writes to land on a single partition, leading to throttling. By designing a higher-cardinality key (e.g., combining date with a random suffix or user ID), writes are distributed evenly across partitions, fully utilizing the provisioned write capacity without custom scripts.

Exam trap

The trap here is that candidates may think adding a GSI (Option A) solves the issue, but GSIs inherit the same partition key design flaws and can also throttle independently.

How to eliminate wrong answers

Option A is wrong because a global secondary index (GSI) with the same date key would still concentrate writes on a single partition in the index, replicating the throttling issue. Option B is wrong because S3 Glacier Instant Retrieval is a storage class for infrequently accessed objects, not a replacement for DynamoDB's transactional write throughput, and moving the table would break the application's access pattern. Option C is wrong because reducing write capacity would worsen throttling during business hours, not solve the underlying partition hot-spotting problem.

35
MCQeasy

A travel booking site uses EC2 instances behind an ALB. CPU is consistently high during peak traffic, and request latency rises. What should be configured? The architecture review board prefers a managed AWS-native control.

A.A VPC endpoint for CloudWatch only
B.Auto Scaling policy based on an appropriate CloudWatch metric
C.S3 Object Lock
D.Disable health checks
AnswerB

An Auto Scaling policy based on an appropriate CloudWatch metric—most commonly average CPUUtilization across the Auto Scaling group—directly addresses the high CPU condition. With target tracking, Auto Scaling dynamically adjusts desired capacity to keep the metric near a defined value, such as 60%. When CPU rises above the target, it launches additional EC2 instances to spread the load; when it falls, it terminates instances to reduce cost. This is the standard, correct solution for a workload whose compute demand varies and is currently saturating existing instances.

Why this answer

An Auto Scaling policy based on an appropriate CloudWatch metric (such as CPUUtilization or request latency) dynamically adds or removes EC2 instances to match demand. This managed AWS-native control directly addresses high CPU and rising latency during peak traffic by scaling out capacity, which is the preferred approach per the architecture review board's requirement for a managed solution.

Exam trap

The trap here is that candidates may confuse monitoring (VPC endpoints) or data protection (S3 Object Lock) with performance scaling, overlooking that Auto Scaling is the direct AWS-native solution for handling variable load and high CPU.

How to eliminate wrong answers

Option A is wrong because a VPC endpoint for CloudWatch only provides private connectivity to CloudWatch APIs, not scaling or performance improvement; it does not reduce CPU load or latency. Option C is wrong because S3 Object Lock is a data protection feature for preventing object deletion or overwrites in S3, unrelated to EC2 performance or scaling. Option D is wrong because disabling health checks would cause the ALB to route traffic to unhealthy instances, worsening latency and availability, not solving the high CPU issue.

36
MCQeasy

An application repeatedly reads the same DynamoDB items with very low latency requirements. The application can tolerate slightly stale data (for example, within a few seconds). You want to improve read latency without changing the existing DynamoDB table schema. Which service is the best choice?

A.Amazon DAX
B.Amazon S3 Transfer Acceleration
C.Amazon EFS
D.AWS CloudTrail for data plane reads
AnswerA

Amazon DAX is an in-memory cache specifically designed for DynamoDB reads. It can significantly reduce read latency for frequently accessed items. Because the application can tolerate brief staleness, DAX’s caching behavior is appropriate and does not require a DynamoDB schema change.

Why this answer

Amazon DAX (DynamoDB Accelerator) is an in-memory cache that sits between your application and DynamoDB, providing microsecond read latency for frequently accessed items. Since the application can tolerate slightly stale data (within seconds), DAX's default TTL-based caching is ideal because it reduces read pressure on DynamoDB while serving cached results with significantly lower latency than direct DynamoDB reads.

Exam trap

The trap here is confusing a caching layer (DAX) with unrelated acceleration or storage services (S3 Transfer Acceleration, EFS) or with auditing tools (CloudTrail), leading candidates to pick options that don't address DynamoDB read latency at all.

How to eliminate wrong answers

Option B (Amazon S3 Transfer Acceleration) is wrong because it accelerates uploads to S3 over long distances using edge locations, not DynamoDB reads, and does not address DynamoDB latency. Option C (Amazon EFS) is wrong because it is a file storage service for EC2 instances, not a cache for DynamoDB items, and introduces network filesystem overhead incompatible with sub-millisecond read requirements. Option D (AWS CloudTrail for data plane reads) is wrong because CloudTrail records API activity for auditing, not caching or accelerating data reads, and enabling it would add latency and cost without improving read performance.

37
MCQmedium

A DynamoDB-backed multi-tenant app experiences throttling. Most write traffic for tenant 'ACME' targets a single logical stream of events (you write items for ACME in near-real time). The table currently uses partition key = tenantId and sort key = eventTimestamp. CloudWatch shows partition-level throttling concentrated in the ACME partition. What design change most directly improves write throughput for the hottest tenant while still enabling efficient queries for recent events for that tenant?

A.Add a Global Secondary Index (GSI) with the same partition key (tenantId) and eventTimestamp, and rely on the GSI to spread load.
B.Mitigate the hotspot by changing the partition key to include a shard value (for example, tenantId + '#' + shardId) and write using shardId. Query recent events by fanning out across ACME shards and merging results by eventTimestamp.
C.Increase the table’s write capacity (or on-demand baseline) without changing the partition key, because DynamoDB will automatically balance hotspots.
D.Switch the sort key to a random value to prevent writes from landing on the same physical partition.
AnswerB

In DynamoDB, the partition key controls which physical partitions receive traffic for that key value. By adding shardId into the partition key, ACME writes are distributed across multiple partitions, increasing aggregate write capacity and reducing partition-level throttling. Efficient recent-event queries are still possible by querying each ACME shard for the relevant time range (using eventTimestamp as the sort key) and merging the ordered results.

Why this answer

It directly addresses the partition-level throttling by introducing a shard key (e.g., tenantId + '#' + shardId) as the partition key, which distributes ACME's write load across multiple physical partitions. To query recent events for ACME, the application must fan out queries across all shards and merge results by eventTimestamp, which is efficient because each shard holds a subset of the data and the sort key remains eventTimestamp for ordering.

Exam trap

The trap here is that candidates often think increasing provisioned capacity or switching to on-demand mode will automatically resolve a hot partition, but DynamoDB's per-partition throughput limit (3,000 RCU or 1,000 WCU) is a hard ceiling that cannot be overcome without redistributing the partition key.

How to eliminate wrong answers

Option A is wrong because adding a GSI with the same partition key (tenantId) does not spread the write load; the base table still experiences the same hotspot, and the GSI inherits the same partition-level throttling. Option C is wrong because DynamoDB does not automatically balance hotspots caused by a single partition key; increasing write capacity on a table with a skewed partition key only raises the per-partition limit but does not distribute the load across more partitions. Option D is wrong because changing the sort key to a random value would prevent efficient queries for recent events (since sort key ordering is lost) and does not change the partition key, so writes still target the same physical partition.

38
MCQhard

Based on the exhibit, a serverless API on AWS Lambda experiences a predictable cold-start penalty every weekday at 09:00 UTC when a marketing campaign begins. The team wants the first requests to stay fast while minimizing extra cost during quiet periods. What is the best approach?

A.Enable provisioned concurrency on the published version and schedule it to scale up shortly before the spike.
B.Increase the Lambda timeout so cold starts have more time to complete.
C.Move the function behind an Application Load Balancer to improve warm-up behavior.
D.Increase the function memory to the maximum value and leave concurrency unchanged.
AnswerA

Provisioned concurrency pre-initializes Lambda execution environments for a specific version or alias, eliminating the cold-start delay for the first request. By configuring a scheduled scaling action (via Application Auto Scaling) to raise provisioned concurrency just before the spike, the environments are already warm and ready to handle the burst. This approach directly targets the initialization penalty shown in the exhibit, and because provisioned concurrency is billed only while it is enabled, scheduling it around the known spike avoids paying for idle resources the rest of the day.

Why this answer

Provisioned concurrency pre-warms a specified number of Lambda execution environments so that incoming requests do not incur a cold start. By scheduling the provisioned concurrency to scale up just before the 09:00 UTC spike and scale down afterward, the team eliminates the cold-start penalty during the campaign while minimizing cost during quiet periods. This directly addresses the predictable, time-bound traffic pattern without requiring code changes or over-provisioning.

Exam trap

The trap here is that candidates confuse increasing memory or timeout with solving cold starts, or they mistakenly think an ALB can pre-warm Lambda, when in fact only provisioned concurrency guarantees warm containers for the first requests in a predictable traffic spike.

How to eliminate wrong answers

Option B is wrong because increasing the Lambda timeout does not prevent cold starts; it only extends the maximum execution duration, which has no effect on the initialization latency of a new execution environment. Option C is wrong because placing the function behind an Application Load Balancer does not warm up the Lambda; ALB is a request router and does not maintain warm containers or alter Lambda's scaling behavior. Option D is wrong because increasing memory (which also increases CPU allocation) can reduce cold-start duration but does not eliminate it, and setting memory to the maximum value (10,240 MB) would significantly increase cost without guaranteeing zero cold starts for the first requests.

39
MCQeasy

A data processing application runs on a single EC2 instance and needs persistent block storage with sustained low-latency random read/write performance (high IOPS). Which storage choice is most appropriate?

A.EBS io2 provisioned IOPS SSD
B.Amazon S3 Standard
C.Amazon EFS for POSIX file sharing between multiple instances
D.EBS Throughput Optimized HDD (st1) storage
AnswerA

Amazon EBS io2 is a block-level storage volume designed for low-latency, high-IOPS workloads. It provides provisioned IOPS with consistent performance, making it ideal for a single EC2 instance running data-processing applications that require high random I/O and low latency. Unlike other options, io2 volumes deliver 99.999% durability and are directly attached to the instance, ensuring minimal network overhead.

Why this answer

Amazon EBS io2 Provisioned IOPS SSD volumes are designed for I/O-intensive workloads that require sustained, low-latency random read/write performance with high IOPS. They provide consistent performance by allowing you to provision a specific IOPS level (up to 256,000 IOPS per volume) and offer a 99.999% durability guarantee, making them ideal for a single EC2 instance needing persistent block storage.

Exam trap

The trap here is that candidates often confuse throughput-optimized HDD (st1) with IOPS-optimized SSD (io2) because both are EBS volume types, but st1 is designed for sequential, not random, I/O workloads.

How to eliminate wrong answers

Option B is wrong because Amazon S3 is an object storage service accessed via HTTP/HTTPS, not a block storage device, and cannot be attached directly to an EC2 instance for low-latency random read/write operations. Option C is wrong because Amazon EFS is a file-level, NFS-based storage service designed for shared access across multiple EC2 instances, not for providing persistent block storage with sustained low-latency random I/O to a single instance. Option D is wrong because EBS Throughput Optimized HDD (st1) volumes are optimized for large, sequential workloads (e.g., big data, log processing) and cannot deliver the sustained low-latency random I/O performance required for high IOPS workloads.

40
Multi-Selectmedium

A retail company is deploying a read-heavy product catalog on Amazon Aurora MySQL. The primary instance is heavily loaded with read traffic, and the team wants to offload reads to Aurora Replicas while keeping the application resilient to replica failures. The application connects using a single endpoint. Which two actions should the team take to meet these requirements? (Choose two.)

Select 2 answers
A.Configure the application to connect directly to each Aurora Replica by its individual instance endpoint and round-robin between them in code.
B.Enable Aurora Auto Scaling for the cluster so that Aurora Replicas are added or removed based on read workload.
C.Create an Amazon ElastiCache for Memcached cluster and cache all product catalog queries there.
D.Convert the Aurora cluster to a multi-master cluster so that all instances accept writes and reads.
E.Create one or more Aurora Replicas and connect the application to the reader endpoint of the Aurora cluster.
AnswersB, E

Aurora Auto Scaling monitors read workload and automatically adds or removes Aurora Replicas within a defined range. Combined with the reader endpoint, this keeps read capacity aligned with demand and improves resilience because the endpoint always includes healthy replicas. It reduces manual intervention and helps maintain performance during traffic spikes.

Why this answer

Aurora separates the writer endpoint from the reader endpoint, and the reader endpoint automatically distributes connections across all Aurora Replicas. Pairing that endpoint with Aurora Auto Scaling lets the cluster add or remove replicas based on read demand, so the application gains both read offloading and resilience without custom failover logic. Direct instance endpoints, caching, and multi-master do not fulfill the specific goal of routing reads through replicas.

Exam trap

The trap here is assuming that connecting directly to each replica provides better control, when it actually bypasses the built-in load balancing and failover of the reader endpoint.

41
Multi-Selecthard

A latency-sensitive video platform uploads large files to S3 from users around the world. Which two features can improve upload performance?

Select 2 answers
A.S3 Object Lock
B.S3 Transfer Acceleration
C.S3 multipart upload
D.S3 Inventory
AnswersB, C

S3 Transfer Acceleration routes uploads through AWS's globally distributed edge locations, which then forward data to the bucket over the private, high-bandwidth AWS backbone instead of the public internet. This shortens the effective network path and reduces the impact of long-distance latency and packet loss, making it ideal for large, latency-sensitive videos uploaded from many global locations. While it incurs a per-GB cost, it directly addresses the throughput and latency bottlenecks typical for cross-continental uploads.

Why this answer

S3 Transfer Acceleration (B) uses AWS edge locations to route uploads over optimized network paths, reducing latency and packet loss for global users. S3 multipart upload (C) allows large files to be uploaded in parallel parts, improving throughput and enabling retries of individual parts without restarting the entire upload.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration with CloudFront or think multipart upload is only for resumability, when in fact both features directly address latency and throughput for global, large-file uploads.

42
Multi-Selecthard

A retail analytics table stores events in Amazon DynamoDB with partition key tenantId and sort key eventTime. During a promotion, one tenant generates most writes and repeatedly polls the same latest-status items, causing throttling on a single partition key and high latency on reads. The business can tolerate read results that are a few seconds stale. Which two changes will most effectively reduce throttling and latency? Select two.

Select 2 answers
A.Introduce write sharding by adding a bounded random suffix to the hot tenant partition key and fan out reads across the shards.
B.Add DynamoDB Accelerator (DAX) in front of the table for the repeated status reads.
C.Keep the same key design and increase only the table’s provisioned RCUs and WCUs.
D.Replace the table reads with a Scan operation to distribute the load across all partitions.
E.Move the table to another Availability Zone so the hot tenant uses a different storage node.
AnswersA, B

Sharding spreads the hot tenant’s traffic across multiple partitions so DynamoDB is no longer forced to serve all writes through one physical partition. Querying across the shard set restores access to the tenant’s data while reducing throttling. This is the standard fix when a single partition key becomes a hot spot.

Why this answer

Write sharding distributes the hot tenant's writes across multiple partitions by appending a bounded random suffix to the partition key, preventing a single partition from throttling. Reads then fan out across all shards and aggregate results, which is acceptable since the business tolerates a few seconds of staleness. This directly addresses the single-partition bottleneck without changing the overall data model.

Exam trap

The trap here is that candidates often assume DAX alone can fix both read and write throttling, but DAX only caches reads and does not address the write-side partition bottleneck that causes throttling in the first place.

43
Multi-Selecthard

A media company serves versioned JavaScript and CSS files from Amazon S3 through CloudFront. After each release, the cache hit ratio drops sharply because the same distribution also fronts a personalized API path, and the current cache policy forwards cookies, all query strings, and several headers to every origin request. The static assets already use content-hashed filenames. Which two changes will most directly improve cache hit ratio for the static assets without changing the application behavior? Select two.

Select 2 answers
A.Create a dedicated cache behavior for the static asset path that excludes cookies, query strings, and unneeded headers from the cache key.
B.Keep the content-hashed filenames and send long Cache-Control max-age and immutable headers for the versioned objects.
C.Increase the size of the S3 bucket’s underlying storage to absorb more origin traffic.
D.Add Lambda@Edge logic to append a timestamp to every asset request so updates are always fetched immediately.
E.Disable compression so CloudFront can treat each object as a separate cache entry.
AnswersA, B

Separating the static asset behavior lets CloudFront cache those objects independently from the personalized API. Excluding cookies, query strings, and unnecessary headers prevents cache fragmentation, so many viewers can reuse the same cached object. This is the most direct way to raise hit ratio without altering how the application serves assets.

Why this answer

Creating a dedicated cache behavior for the static asset path (e.g., /static/*) allows you to configure a cache policy that excludes cookies, query strings, and unneeded headers from the cache key. Since the static assets use content-hashed filenames, they are immutable and do not vary by user-specific attributes. By removing these variables from the cache key, CloudFront can serve the same cached object to all users, drastically improving the cache hit ratio.

Exam trap

The trap here is that candidates may think that content-hashed filenames alone guarantee high cache hit ratios, but they overlook that the shared cache policy forwarding cookies and query strings creates many unique cache keys for the same static file, negating the benefit of hashed filenames.

44
MCQmedium

A high-volume analytics dashboard writes streaming click events that must be processed by multiple independent consumers. Which service is most appropriate? The design must avoid adding custom operational scripts.

A.Amazon Route 53
B.Amazon EBS
C.Amazon Kinesis Data Streams
D.AWS DataSync
AnswerC

Kinesis Data Streams supports high-throughput event ingestion with multiple consumers reading from the stream.

Why this answer

Amazon Kinesis Data Streams is the correct choice because it is designed for real-time streaming data ingestion and processing by multiple independent consumers. Each consumer can read from the stream at its own pace using its own shard iterator, enabling parallel processing of click events without custom scripts. This aligns with the requirement for a high-volume analytics dashboard where multiple downstream applications need to consume the same stream independently.

Exam trap

The trap here is that candidates may confuse Kinesis Data Streams with Kinesis Data Firehose, but Firehose delivers data to a single destination and does not support multiple independent consumers natively.

How to eliminate wrong answers

Option A is wrong because Amazon Route 53 is a DNS and traffic management service, not designed for streaming data ingestion or processing. Option B is wrong because Amazon EBS provides block-level storage volumes for EC2 instances, not a streaming data platform for multiple consumers. Option D is wrong because AWS DataSync is a data transfer service for moving large datasets between on-premises storage and AWS services, not for real-time streaming or multi-consumer processing.

45
MCQhard

A document portal needs low-latency full-text search across product descriptions and filtered attributes. Which managed service is most suitable?

A.Amazon OpenSearch Service
B.AWS Config
C.Amazon EFS
D.Amazon SQS
AnswerA

Amazon OpenSearch Service provides a distributed, Lucene-based search engine that creates inverted indexes over document fields, enabling fast full-text queries with tokenization, stemming, and relevance scoring. For a document portal, it can ingest content via APIs or streaming integrations and serve sub-second searches across large corpora. Its built-in features like multi-tenancy, searchable snapshots, and Kibana dashboards make it the purpose-built choice for low-latency text search.

Why this answer

Amazon OpenSearch Service is purpose-built for full-text search, log analytics, and real-time application monitoring. It provides low-latency indexing and querying of unstructured and semi-structured data, making it ideal for searching product descriptions and filtering attributes. The service uses a RESTful API and supports advanced query DSL for complex search operations.

Exam trap

The trap here is that candidates may confuse a storage service (EFS) or a messaging service (SQS) with a search service, or mistakenly think AWS Config can be used for text search because it stores configuration data in a queryable format.

How to eliminate wrong answers

Option B (AWS Config) is wrong because it is a resource inventory and compliance auditing service, not a search engine; it cannot perform full-text search across product descriptions. Option C (Amazon EFS) is wrong because it is a scalable file storage service for Linux-based workloads, not a search or indexing service; it provides shared file access but no search capabilities. Option D (Amazon SQS) is wrong because it is a fully managed message queuing service for decoupling application components, not a search service; it cannot index or query text data.

46
MCQmedium

An API team runs an AWS Lambda function behind an Application Load Balancer (ALB). During predictable hourly traffic spikes, p95 response latency increases due to occasional cold starts. The team wants stable latency during those spikes without permanently overprovisioning resources for all functions. Which configuration is the most appropriate way to reduce cold starts for this Lambda function?

A.Publish a version of the function and configure provisioned concurrency on an alias, using autoscaling for the alias.
B.Increase the function memory size and rely on faster initialization to reduce cold starts.
C.Set reserved concurrency equal to the expected peak requests per second for the function.
D.Use an event source mapping with a higher batch size so Lambda triggers earlier and keeps the runtime warm.
AnswerA

Provisioned concurrency pre-initializes execution environments for a specific published function version. By attaching provisioned concurrency to an alias, you can control warm capacity and (with the right settings) autoscale the provisioned capacity for predictable spike patterns, reducing cold-start-driven latency increases.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, keeping them warm and ready to handle requests without cold start latency. By configuring provisioned concurrency on an alias with autoscaling, the team can dynamically adjust the number of pre-warmed environments to match predictable traffic spikes, avoiding permanent overprovisioning while ensuring stable p95 latency.

Exam trap

The trap here is confusing reserved concurrency (which limits concurrency but does not prevent cold starts) with provisioned concurrency (which pre-warms environments), leading candidates to select Option C as a cost-saving measure that fails to address latency.

Why the other options are wrong

B

Increasing memory size improves CPU speed and can reduce initialization time, but it does not eliminate cold starts; it only shortens them. The question asks to reduce cold starts, not just mitigate their duration.

C

Reserved concurrency limits the maximum number of concurrent executions but does not pre-warm instances; cold starts still occur during traffic spikes when new execution environments are needed.

D

Event source mappings (e.g., SQS, DynamoDB Streams) are not used with ALB triggers; ALB invokes Lambda synchronously via a function URL or alias ARN. Increasing batch size does not apply to ALB-triggered functions and does not prevent cold starts during predictable spikes.

47
MCQeasy

Your web application runs on EC2 instances behind an Application Load Balancer (ALB). During traffic spikes, p95 response time increases, but average CPU utilization remains below 40%. The current Auto Scaling policy scales based on average CPU%. What should you change to improve performance during spikes?

A.Keep scaling on CPU% to avoid over-scaling
B.Scale on a request-driven metric such as ALB RequestCount per target (or target-group request rate)
C.Disable scaling and manually increase capacity during business hours
D.Scale only when network packet drops fall below a threshold
AnswerB

A request-driven metric correlates directly with incoming workload pressure. Scaling on request rate helps ensure enough capacity is added before request queues build up, which can reduce p95 response time even when CPU remains low.

Why this answer

The p95 response time is increasing during traffic spikes while CPU utilization remains low, indicating that the bottleneck is not compute capacity but rather request handling or connection overhead. By scaling on ALB RequestCountPerTarget, you directly target the metric causing latency—each target's request load—rather than an indirect metric like CPU. This ensures that new instances are launched precisely when individual targets are overwhelmed by requests, reducing queueing delays and improving response times.

Exam trap

The trap here is that candidates assume high latency always means high CPU, but AWS tests the understanding that p95 latency can spike due to request queueing even when CPU is idle, making request-based scaling the correct choice over CPU-based scaling.

How to eliminate wrong answers

Option A is wrong because continuing to scale on CPU% ignores the actual symptom (high p95 latency with low CPU), leading to under-provisioning during request bursts. Option C is wrong because manual scaling during business hours is not elastic and cannot react to unpredictable traffic spikes, violating the principle of auto scaling for performance. Option D is wrong because scaling on network packet drops is irrelevant to the described issue (low CPU, high latency) and packet drops typically indicate network congestion or buffer exhaustion, not request overload on the application layer.

48
MCQhard

A media archive needs low-latency full-text search across product descriptions and filtered attributes. Which managed service is most suitable? The design must avoid adding custom operational scripts.

A.AWS Config
B.Amazon OpenSearch Service
C.Amazon EFS
D.Amazon SQS
AnswerB

Amazon OpenSearch Service provides managed full-text search with low-latency querying across document fields, plus filter clauses for structured attributes. It satisfies the no-custom-operational-scripts constraint because AWS handles cluster provisioning, patching and scaling, unlike self-managed search engines requiring bespoke maintenance automation.

Why this answer

Amazon OpenSearch Service is the correct choice because it provides managed, low-latency full-text search capabilities with support for filtering on structured attributes (e.g., product categories, price ranges). It indexes JSON documents and exposes a RESTful API for search queries, eliminating the need for custom operational scripts while meeting the media archive's requirements.

Exam trap

The trap here is that candidates might confuse AWS Config's resource tracking or EFS's file storage with search capabilities, overlooking that OpenSearch Service is the only managed option purpose-built for full-text search and filtering.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for auditing and evaluating resource configurations against compliance rules, not for full-text search or indexing product descriptions. Option C is wrong because Amazon EFS is a scalable NFS file system for shared storage, not a search engine; it cannot perform low-latency full-text queries across text content. Option D is wrong because Amazon SQS is a managed message queue for decoupling application components, not a search or indexing service, and it does not support querying stored data.

49
Multi-Selecthard

A solutions architect is designing a high-performance computing (HPC) workload that requires a shared file system with high throughput and low latency for thousands of compute instances. The workload also requires a caching layer to accelerate repeated reads of the same data. Which two AWS services should be combined to meet these requirements? (Choose two.)

Select 2 answers
A.Amazon FSx for Lustre
B.Amazon ElastiCache for Redis
C.Amazon S3 Glacier
D.Amazon EBS Multi-Attach enabled io1 volumes
E.AWS Storage Gateway File Gateway
AnswersA, B

Amazon FSx for Lustre is a high-performance file system optimized for HPC workloads. It provides sub-millisecond latencies, millions of IOPS, and hundreds of gigabytes per second of throughput. It integrates natively with Amazon S3, allowing data to be lazily loaded from S3 and written back. FSx for Lustre is designed for compute-intensive workloads and can be linked to an S3 bucket as a data repository. This makes it ideal for the shared file system requirement.

Why this answer

Amazon FSx for Lustre provides the high-performance shared file system needed for HPC, with low latency and high throughput. Amazon ElastiCache for Redis adds an in-memory caching layer to accelerate repeated reads, reducing load on the file system. Together, they meet the requirements for a scalable, high-performance HPC storage and caching solution.

Other options either have high latency (Glacier), are for hybrid scenarios (File Gateway), or lack the necessary scalability (EBS Multi-Attach).

Exam trap

The trap here is assuming that any shared storage service can handle HPC scale; services like EBS Multi-Attach or Storage Gateway are not designed for thousands of instances or low-latency HPC.

50
Multi-Selecthard

A latency-sensitive video platform uploads large files to S3 from users around the world. Which two features can improve upload performance? The architecture review board prefers a managed AWS-native control.

Select 2 answers
A.S3 Object Lock
B.S3 Transfer Acceleration
C.S3 multipart upload
D.S3 Inventory
AnswersB, C

Transfer Acceleration routes uploads through AWS edge locations to the nearest S3 endpoint over AWS's optimised backbone, cutting latency for globally distributed users. It is fully AWS-managed, meeting the review board's AWS-native constraint while improving worldwide upload performance.

Why this answer

S3 Transfer Acceleration (B) uses AWS edge locations to accelerate uploads over long distances by routing traffic through the AWS global network, reducing latency and packet loss compared to the public internet. Multipart upload (C) improves performance by splitting large files into smaller parts that can be uploaded in parallel, increasing throughput and allowing retries of individual parts without restarting the entire upload.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration with CloudFront or think multipart upload is only for reliability, not performance, while overlooking that both features are managed AWS-native controls that directly address latency and throughput for large file uploads.

51
MCQeasy

A startup runs a stateless web application on Amazon EC2 instances behind an Application Load Balancer. Traffic is steady during the day but drops to almost zero overnight, and the team wants to reduce compute cost without manual intervention or a service interruption. Which action should the team take?

A.Replace the Auto Scaling group with a single large EC2 instance and use Elastic Load Balancing health checks.
B.Create a scheduled scaling action on an Auto Scaling group that reduces the desired capacity at night and restores it each morning.
C.Attach an Amazon EBS volume with higher IOPS to each instance to reduce the number of instances needed.
D.Switch to Spot Instances for the entire fleet and rely on capacity-optimized allocation.
AnswerB

Scheduled scaling adjusts the desired capacity of an Auto Scaling group at defined times, so the fleet can shrink overnight when traffic is near zero and grow before the morning peak. Because the group manages instance lifecycle and the ALB drains connections, there is no service interruption. This matches a predictable daily pattern and removes manual intervention.

Why this answer

The workload has a predictable daily pattern, which is exactly what scheduled scaling on an Auto Scaling group is designed for. Reducing desired capacity overnight and restoring it in the morning lowers compute spend while the group and load balancer keep the application available. Connection draining on the ALB ensures in-flight requests finish before instances terminate, so users see no interruption.

Exam trap

The trap here is choosing Spot Instances purely for cost when the scenario's real requirement is matching capacity to a known daily traffic curve without risking reclamation.

52
MCQeasy

A compute workload uses temporary scratch space for intermediate results (reproducible), and it can tolerate data loss if the instance is terminated. The workload benefits from very high local I/O throughput. Which storage option is the best fit for the scratch data?

A.Amazon EBS General Purpose (gp3) volumes to persist intermediate results across reboots.
B.Amazon EFS for a shared file system between multiple instances.
C.Instance store for local temporary files that can be lost when the instance stops.
D.Amazon S3 for scratch data so it is always durable and accessible from anywhere.
AnswerC

Instance store provides physically attached NVMe SSDs delivering the highest local I/O throughput and lowest latency, unlike EBS network storage. Because the intermediate results are reproducible and loss on termination is acceptable, the ephemeral, non-persistent nature of instance store matches the workload's tolerance exactly.

Why this answer

Instance store volumes provide very high local I/O throughput because they are physically attached to the host server, making them ideal for temporary scratch data that is reproducible and can tolerate loss. Since the workload explicitly accepts data loss on instance termination and does not require persistence across reboots, instance store is the best fit for this use case.

Exam trap

The trap here is that candidates often choose EBS gp3 (Option A) because they assume all block storage is persistent and high-performance, overlooking the fact that instance store offers even higher local throughput and is explicitly designed for temporary, loss-tolerant workloads.

Why the other options are wrong

B

Amazon EFS provides a shared file system, but the question specifies scratch data for a single instance that benefits from very high local I/O throughput. EFS is network-attached and has higher latency than local storage, making it unsuitable for high local I/O needs.

D

Amazon S3 is designed for durable, highly available object storage with high latency, not for high local I/O throughput scratch space. It cannot provide the very high local I/O performance required for temporary scratch data.

53
MCQmedium

A global mobile game backend serves mostly static images and JavaScript files from an S3 origin. Users in distant countries report slow load times. What should improve performance most? The architecture review board prefers a managed AWS-native control.

A.RDS read replicas
B.Amazon CloudFront distribution with the S3 bucket as origin
C.A larger S3 bucket
D.An EC2 Auto Scaling group in one Region
AnswerB

CloudFront caches the static images and JavaScript at edge locations close to distant users, cutting latency versus direct S3 retrieval. It is fully managed and AWS-native, satisfying the review board's preference, and integrates directly with the S3 bucket as origin.

Why this answer

Amazon CloudFront is a global content delivery network (CDN) that caches static content (images, JavaScript) at edge locations close to users, drastically reducing latency. By using the S3 bucket as the origin, CloudFront offloads requests from S3 and serves cached objects from the nearest edge, which directly addresses slow load times for distant users. This is a managed AWS-native service that aligns with the architecture review board's preference.

Exam trap

The trap here is that candidates may think increasing S3 bucket size or using RDS replicas can improve static content delivery, but the core issue is geographic latency, which only a CDN like CloudFront can solve by caching content at edge locations.

How to eliminate wrong answers

Option A is wrong because RDS read replicas are designed to offload read traffic from a relational database, not to accelerate delivery of static files stored in S3; they have no effect on S3 latency. Option C is wrong because increasing the S3 bucket size does not improve data transfer speed or reduce latency; S3 performance is independent of bucket size and is limited by regional endpoints. Option D is wrong because an EC2 Auto Scaling group in a single Region does not provide geographic distribution; users in distant countries would still experience high latency connecting to that single Region, and it adds unnecessary compute overhead for serving static content.

54
MCQmedium

A logistics company runs a REST API on Amazon ECS using the Fargate launch type behind an Application Load Balancer. The API's response times are acceptable, but the operations team wants to reduce the number of database calls per request by caching frequently accessed reference data in memory inside the tasks. The data changes infrequently and slight staleness is acceptable. Which approach best meets these requirements?

A.Store the reference data in an Amazon ElastiCache for Redis cluster and have tasks query it on each request.
B.Implement an in-process cache in the application code that loads reference data at task startup and refreshes it periodically.
C.Enable Amazon ElastiCache for Memcached and configure automatic discovery of cache nodes.
D.Use the ECS task definition to mount an Amazon EFS file system containing the reference data.
AnswerB

An in-process cache keeps the reference data in the task's own memory, so requests are served without any database call or network hop. Since the data changes infrequently and slight staleness is acceptable, a periodic refresh is sufficient and simple. This directly satisfies the goal of reducing database calls per request while using the Fargate tasks' memory, with no additional managed service required.

Why this answer

Caching reference data in the application's own process memory means each request is served locally with no database or network call, which is the most direct way to reduce per-request database calls. Because the data is small, changes rarely, and tolerates slight staleness, periodic refresh is simple and effective. External caches like Redis or Memcached still require a network call, and EFS is storage rather than memory.

Exam trap

The trap here is treating any cache, such as ElastiCache, as equivalent to in-process caching, when external caches still require a network round trip per request.

55
MCQhard

Based on the exhibit, which change will most improve the CloudFront cache hit ratio for the static assets while still serving the same files to all users?

A.Create a custom cache policy that includes only the v query string and excludes cookies.
B.Enable Origin Shield and keep the current cache behavior unchanged.
C.Move the static assets to individual presigned URLs for each viewer.
D.Increase the CloudFront default TTL to 24 hours while continuing to forward all cookies and query strings.
AnswerA

This removes unnecessary cache-key fragmentation. Since all users receive identical static files, forwarding user-specific cookies and irrelevant query strings destroys cache reuse. Keeping only the version parameter preserves correct object variation while allowing many more requests to hit the same cached object at the edge.

Why this answer

The CloudFront cache hit ratio for static assets is reduced when query strings and cookies are forwarded to the origin, because each unique combination creates a separate cache entry. By creating a custom cache policy that includes only the 'v' query string (used for versioning) and excludes cookies, CloudFront can cache a single object for all users regardless of other query parameters or cookie values, maximizing cache hits while still serving the same file.

Exam trap

The trap here is that candidates assume increasing TTL or enabling Origin Shield will fix a low cache hit ratio, when the real issue is an overly broad cache key caused by forwarding all query strings and cookies.

How to eliminate wrong answers

Option B is wrong because enabling Origin Shield reduces load on the origin and improves cache fill efficiency, but it does not address the root cause of low cache hit ratio—forwarding all query strings and cookies still creates many unique cache keys. Option C is wrong because moving static assets to individual presigned URLs for each viewer would force CloudFront to treat each URL as a distinct object, drastically reducing the cache hit ratio and defeating the purpose of caching. Option D is wrong because increasing the default TTL to 24 hours while continuing to forward all cookies and query strings does not reduce the number of unique cache keys; CloudFront will still cache separate copies for each cookie and query string combination, so the cache hit ratio remains low.

56
MCQmedium

A high-volume analytics dashboard writes streaming click events that must be processed by multiple independent consumers. Which service is most appropriate?

A.Amazon Route 53
B.Amazon EBS
C.Amazon Kinesis Data Streams
D.AWS DataSync
AnswerC

Kinesis Data Streams retains an ordered, replayable record sequence for a configurable period, and each consumer reads independently via its own iterator, so multiple analytics applications process the same click events without competing for messages.

Why this answer

Amazon Kinesis Data Streams is the most appropriate service because it is designed for real-time streaming data ingestion and can be consumed by multiple independent consumers in parallel. Each shard within a Kinesis stream supports up to 5 read transactions per second and a total data read rate of 2 MB per second, allowing multiple consumer applications to process the same stream of click events concurrently without interfering with each other.

Exam trap

The trap here is that candidates often confuse Amazon Kinesis Data Streams with Amazon SQS or Amazon SNS, but SQS is a message queue for decoupled point-to-point communication and SNS is a pub/sub notification service, neither of which natively supports multiple independent consumers processing the same stream of data with replay capability.

How to eliminate wrong answers

Option A is wrong because Amazon Route 53 is a DNS web service that translates domain names to IP addresses and does not ingest or process streaming data. Option B is wrong because Amazon EBS provides block-level storage volumes for EC2 instances and cannot natively support multiple independent consumers reading a continuous stream of events. Option D is wrong because AWS DataSync is a data transfer service for moving large datasets between on-premises storage and AWS services, not for real-time streaming event processing.

57
MCQeasy

A web application uses an Amazon Aurora DB cluster. The workload is becoming read-heavy, and the application team wants to increase read throughput without changing the database schema. They can adjust the application to route reads differently. What should they do?

A.Add Aurora read replicas and route read queries to the cluster reader endpoint
B.Switch the cluster to Multi-AZ with a longer failover target clock
C.Move all reads to the writer endpoint to reduce connection overhead
D.Disable automated backups to reduce storage overhead and speed reads
AnswerA

Aurora read replicas scale read throughput by creating up to 15 independent instances that share the same distributed storage volume. The cluster reader endpoint automatically routes new connections to any available replica, allowing SELECT queries to run in parallel across multiple instances while the writer instance focuses on update operations. Because replicas require no data copy and remain fully in sync with minimal replica lag, this approach directly addresses a read-heavy workload without schema or application refactoring.

Why this answer

Adding Aurora read replicas and routing read queries to the cluster reader endpoint is the correct approach because Aurora replicas share the same underlying storage volume as the primary instance, so they can serve read traffic with minimal replication lag. The reader endpoint automatically load-balances connections across all available replicas, increasing aggregate read throughput without requiring any schema changes.

Exam trap

The trap here is confusing Multi-AZ with read replicas: candidates often think Multi-AZ improves read performance, but in standard RDS Multi-AZ the standby is passive and cannot serve reads, whereas Aurora's architecture allows all replicas to actively handle read traffic.

How to eliminate wrong answers

Option B is wrong because Multi-AZ with a longer failover target clock does not increase read throughput; it only provides high availability by maintaining a standby in another Availability Zone, and the standby cannot serve reads. Option C is wrong because moving all reads to the writer endpoint would increase load on the single writer instance, reducing overall read throughput and potentially impacting write performance. Option D is wrong because disabling automated backups does not increase read throughput; backups are stored separately and do not affect the performance of read operations on the cluster.

58
MCQmedium

Your company currently uses an Application Load Balancer (ALB) in front of a service that receives a large number of TCP and UDP packets (including UDP-based telemetry). During load tests, you need to support both TCP and UDP traffic at high throughput while keeping stable IP endpoints for a downstream firewall allowlist. Which change best meets these requirements?

A.Switch to a Network Load Balancer (NLB) configured for TCP/UDP, and use Elastic IPs to provide stable endpoint IP addresses for allowlisting.
B.Keep the ALB and add an AWS WAF Web ACL to improve throughput and add static IP support.
C.Replace the ALB with an API Gateway REST API to support UDP because API Gateway can forward UDP packets.
D.Use an Auto Scaling group with multiple EC2 instances and no load balancer to avoid any networking bottlenecks.
AnswerA

NLB operates at Layer 4 and supports both TCP and UDP. For stable IP allowlists, you can associate Elastic IP addresses with the NLB so the load balancer exposes consistent IPs (as opposed to relying on dynamic addresses). This combination directly satisfies protocol support and stable endpoint requirements.

Why this answer

A Network Load Balancer (NLB) operates at Layer 4 and can handle both TCP and UDP traffic natively, unlike an ALB which only supports HTTP/HTTPS and cannot forward UDP packets. By assigning Elastic IPs to the NLB, you provide stable, static IP endpoints that can be added to a downstream firewall allowlist, meeting both the protocol and throughput requirements.

Exam trap

The trap here is that candidates assume an ALB can handle all traffic types because it is the most commonly used load balancer, but they forget that ALB is strictly Layer 7 and cannot process UDP packets, making the NLB the only correct choice for mixed TCP/UDP workloads requiring static IPs.

How to eliminate wrong answers

Option B is wrong because an ALB cannot handle UDP traffic (it only supports HTTP/HTTPS and WebSocket), and AWS WAF does not add static IP support or improve throughput for Layer 4 traffic. Option C is wrong because API Gateway REST APIs do not support UDP traffic; they only handle HTTP/HTTPS and WebSocket protocols. Option D is wrong because removing the load balancer eliminates the stable IP endpoint required for the firewall allowlist and introduces a single point of failure, while also not addressing the need for high-throughput TCP/UDP handling with a consistent front-end IP.

59
MCQhard

A media company serves video thumbnails from an Amazon S3 bucket in us-east-1 to viewers across Europe and Asia. The thumbnails are immutable after upload and are requested repeatedly by the same users. The company wants to reduce latency for the global audience and reduce data transfer costs, and it does not want to modify application code. Which solution meets these requirements with the LEAST operational effort?

A.Replicate the bucket to eu-west-1 and ap-southeast-1 using S3 Cross-Region Replication, and have the application select the nearest bucket.
B.Create an Amazon CloudFront distribution with the S3 bucket as origin, enable caching, and configure an origin access control.
C.Attach an S3 bucket policy that allows public read and place an Application Load Balancer in front of the bucket.
D.Enable S3 Transfer Acceleration on the bucket and update the application to use the accelerated endpoint.
AnswerB

CloudFront caches immutable thumbnails at edge locations close to European and Asian viewers, so repeat requests are served from the edge with lower latency and fewer origin fetches, which reduces S3 data transfer cost. Origin access control restricts direct bucket access while keeping the same object URLs behind the distribution. This requires no application code change and is the least-effort global acceleration option.

Why this answer

CloudFront is the AWS content delivery network that caches objects at edge locations worldwide, so repeated thumbnail requests from Europe and Asia are served close to viewers instead of from us-east-1. Because the objects are immutable, cache hit ratios stay high, cutting origin fetches and S3 egress charges. Origin access control keeps the bucket private, and the application continues to use the distribution domain without code changes.

Exam trap

The trap here is confusing S3 Transfer Acceleration, which accelerates individual transfers to a bucket, with edge caching, which is what actually serves repeat reads quickly to a global audience.

60
MCQeasy

A retail company runs a read-heavy product catalog on Amazon RDS for MySQL. During flash sales, read replicas lag behind the primary and the application serves stale prices. The team wants to scale read traffic while ensuring the application reads the most current data for price lookups. Which solution should a solutions architect recommend?

A.Enable Multi-AZ on the RDS for MySQL instance and point all reads to the standby.
B.Increase the size of the read replicas and enable automatic scaling of replicas.
C.Configure the application to retry reads on a replica until the price matches the primary.
D.Route price lookups to the primary DB instance and other reads to read replicas.
AnswerD

Read replicas are asynchronous, so they can lag during heavy write periods. Directing price lookups, which require the latest committed data, to the primary instance guarantees current values, while offloading less sensitive reads to replicas preserves read scaling. This matches the requirement for fresh prices without abandoning replica-based scaling.

Why this answer

RDS read replicas use asynchronous replication, so they can lag under write-heavy flash-sale conditions. Price lookups need the latest committed value, so sending them to the primary guarantees freshness, while other reads continue to use replicas for scale. Multi-AZ standbys cannot serve reads, and simply adding replica capacity does not eliminate lag, so routing consistency-sensitive queries to the primary is the correct design.

Exam trap

The trap here is assuming that scaling read replicas also guarantees fresh data, when asynchronous replication means replicas can lag regardless of their size.

61
MCQeasy

A retail analytics app uses Amazon RDS for PostgreSQL. Read traffic is growing, and the database CPU spikes mainly due to SELECT-heavy workloads. Writes are less frequent, and the app can tolerate eventually consistent reads for the reports. What is the most appropriate AWS-native way to improve read performance with minimal application changes?

A.Create an RDS read replica and point the reporting queries to the replica endpoint.
B.Switch the cluster to DynamoDB without redesigning the data model.
C.Enable S3 event notifications to trigger a Lambda function after each write to the database.
D.Replace the RDS instance class with a smaller size to reduce cost and improve performance.
AnswerA

Amazon RDS read replicas use asynchronous replication from the primary DB instance to one or more read-only copies, typically within the same region or cross-region. By pointing reporting and analytics queries to a replica's DNS endpoint, you offload SELECT-heavy traffic from the primary, reducing CPU and I/O contention on the source instance. This lets the primary focus on OLTP writes while analysts query near-real-time data from the replica, requiring no application schema changes. For a retail analytics app with read pressure, this is the minimal-risk, AWS-native fix.

Why this answer

Creating an RDS read replica is the most appropriate AWS-native solution because it offloads SELECT-heavy workloads from the primary database instance to a read-only copy, reducing CPU spikes on the primary. The application can tolerate eventually consistent reads for reports, which is exactly the consistency model of RDS read replicas (typically sub-second replication lag). This requires minimal application changes—only updating the reporting queries to point to the replica endpoint—and fully leverages PostgreSQL's built-in replication capabilities.

Exam trap

The trap here is that candidates might assume read replicas require application changes to handle eventual consistency, but the question explicitly states the app can tolerate eventually consistent reads, making the replica endpoint swap a minimal-change solution.

Why the other options are wrong

B

Switching to DynamoDB without redesigning the data model is not feasible because RDS PostgreSQL and DynamoDB have fundamentally different data models (relational vs. NoSQL), requiring significant application changes to adapt queries, schema, and consistency models.

C

This option does not directly improve read performance for SELECT-heavy workloads on RDS PostgreSQL. It introduces asynchronous S3 event notifications and Lambda processing, which adds latency and complexity without offloading read queries from the primary database.

D

Replacing the RDS instance with a smaller size would reduce CPU capacity, worsening performance under SELECT-heavy workloads, not improving it. The question asks for improved read performance, not cost reduction.

62
MCQeasy

A service performs many repeated read requests for the same DynamoDB items. The reads are latency-sensitive, but the application can tolerate slightly stale data. Which AWS service is the best fit to reduce read latency?

A.Amazon DAX (DynamoDB Accelerator)
B.Amazon S3 Select
C.Amazon SQS FIFO queue
D.AWS Lambda provisioned concurrency
AnswerA

Amazon DAX is an in-memory cache for DynamoDB. It reduces latency for repeated reads by caching results and serving subsequent read requests from the DAX cluster rather than repeatedly calling DynamoDB. Because it provides cached reads that may be slightly stale, it matches the scenario’s tolerance.

Why this answer

Amazon DAX (DynamoDB Accelerator) is an in-memory cache specifically designed for DynamoDB. It reduces read latency from single-digit milliseconds to microseconds by caching frequently accessed items, and it supports eventually consistent reads, which aligns with the application's tolerance for slightly stale data. DAX handles repeated read requests without additional DynamoDB read capacity unit consumption, making it the optimal choice for this latency-sensitive workload.

Exam trap

The trap here is that candidates often confuse caching services (DAX) with data retrieval services (S3 Select) or assume that a queue (SQS) or compute optimization (Lambda provisioned concurrency) can solve read latency issues, when only a purpose-built in-memory cache like DAX directly addresses repeated DynamoDB reads with stale data tolerance.

How to eliminate wrong answers

Option B (Amazon S3 Select) is wrong because it retrieves subsets of data from objects stored in S3 using SQL-like queries, not from DynamoDB items, and it does not provide a caching layer to reduce read latency for repeated DynamoDB reads. Option C (Amazon SQS FIFO queue) is wrong because it is a message queuing service for decoupling and ordering messages, not a caching or read-acceleration service for DynamoDB; it adds latency rather than reducing it for repeated reads. Option D (AWS Lambda provisioned concurrency) is wrong because it pre-warms Lambda execution environments to reduce cold starts, but it does not cache DynamoDB items or reduce read latency for repeated database queries.

63
MCQmedium

A Lambda function behind an API needs consistent low latency. Traffic normally drops to near zero, then spikes several times per hour. During spikes, the p95 latency often spikes above 800 ms due to cold starts. The team wants to keep using Lambda (no containers) but minimize cold start impact during predictable spikes. What is the best AWS configuration to meet this goal?

A.Enable Lambda provisioned concurrency on a published function alias and set the minimum provisioned instances to the baseline expected during spikes.
B.Increase the function memory size to the maximum and rely on the larger memory to eliminate cold starts.
C.Configure an ALB with target group health checks to keep Lambda warm by sending periodic requests.
D.Turn on AWS CloudTrail data events to monitor cold start frequency and tune the runtime accordingly.
AnswerA

Provisioned concurrency pre-creates and initializes Lambda execution environments for a specific published alias or version, so requests are served immediately without a cold start. Setting the provisioned minimum to your baseline expected during spikes ensures that the required capacity is already warm and ready, maintaining consistent low latency under load. This is the correct, managed mechanism designed by AWS for this exact problem.

Why this answer

Provisioned concurrency initializes a specified number of execution environments in advance, keeping them warm and ready to handle requests instantly. By setting the minimum provisioned instances to the baseline expected during spikes, the function avoids cold starts for those requests, ensuring p95 latency stays low even when traffic surges from near zero.

Exam trap

The trap here is that candidates may confuse provisioned concurrency with reserved concurrency, or assume that increasing memory or using health checks can eliminate cold starts, when only provisioned concurrency guarantees pre-warmed environments for predictable spikes.

Why the other options are wrong

B

Increasing memory size can reduce cold start duration but does not eliminate cold starts entirely; it only shortens the initialization time. The question requires minimizing cold start impact during predictable spikes, which provisioned concurrency achieves by keeping instances pre-warmed.

C

ALB health checks send requests to a target, but Lambda functions behind an ALB are invoked only when health checks are configured to hit a specific endpoint. However, health checks are typically sent at intervals (e.g., every 30 seconds), which may not keep the function warm during the unpredictable spikes described. Additionally, health checks can cause unnecessary invocations and costs without guaranteeing that all needed concurrent executions are warm.

64
MCQhard

A gaming company uses Amazon DynamoDB to store player session data. The table has a partition key of PlayerID and a sort key of SessionStartTime. The company needs to retrieve all sessions for a specific player within a date range, sorted by session start time. The table is large and the company wants to minimize read latency. Which approach should they use?

A.Use DynamoDB Accelerator (DAX) to cache the results of a Scan operation.
B.Perform a Query operation on the table using the partition key and a condition on the sort key.
C.Perform a Scan operation with a FilterExpression on PlayerID and SessionStartTime.
D.Create a global secondary index with PlayerID as the partition key and SessionStartTime as the sort key, then Query the index.
AnswerB

A Query operation directly accesses items with a specific partition key and allows a condition on the sort key to retrieve a range of session start times. This is efficient and leverages the table's primary key design, providing low-latency retrieval of sorted results. It is the optimal approach for this access pattern.

Why this answer

The table's primary key already supports the required access pattern: partition key PlayerID and sort key SessionStartTime. A Query operation with a condition on the sort key retrieves exactly the items needed, sorted by session start time, with minimal latency. Scans, GSIs, and DAX do not provide a more efficient solution for this specific query.

Exam trap

The trap here is thinking that a GSI is needed to query by PlayerID and SessionStartTime, but the base table already has that key schema.

65
Multi-Selecthard

A distributed analytics engine runs 12 EC2 instances in one Availability Zone. The nodes exchange thousands of tiny messages per second and must keep jitter as low as possible. The current design launches the instances across multiple placement groups and uses general-purpose burstable instances. Which two changes will most directly lower east-west network latency and variability? Select two.

Select 2 answers
A.Move all instances into a cluster placement group.
B.Use instance families that provide high network bandwidth and support enhanced networking.
C.Spread the instances across three Availability Zones for better fault tolerance.
D.Front the nodes with an Application Load Balancer to balance the internal messages.
E.Store the messages on EBS volumes so the nodes avoid network communication.
AnswersA, B

Cluster placement groups pack instances closely together in a single Availability Zone, which minimizes network distance and improves latency consistency. This is the best placement strategy when the workload is highly chatty and needs very low jitter between nodes. It directly targets east-west performance.

Why this answer

A cluster placement group provides a low-latency, high-bandwidth network connection by placing instances in a single Availability Zone within the same logical rack or cluster. This minimizes the physical distance and network hops between instances, directly reducing east-west latency and jitter for the thousands of tiny messages per second.

Exam trap

The trap here is that candidates often confuse 'fault tolerance' (spreading across AZs) with 'performance' (cluster placement group), or they mistakenly think a load balancer can optimize internal node-to-node traffic, when in fact it adds latency and is designed for client-facing traffic.

66
MCQmedium

A media processing service runs ECS tasks in multiple Availability Zones. Each task must read and write the same shared filesystem with low latency because tasks stream intermediate artifacts to other tasks. The team currently mounts an EBS volume per task, and cross-AZ tasks frequently cannot see each other’s files. Which option best resolves the shared filesystem requirement while supporting high-performing access?

A.Keep using EBS, but attach the same EBS volume to tasks in multiple Availability Zones using EBS multi-attach so all tasks share the filesystem.
B.Use Amazon EFS with mount targets in each Availability Zone so all tasks mount a common NFS filesystem over the AWS network.
C.Use Amazon S3 for the intermediate artifacts and rely on S3 event notifications to emulate POSIX file operations.
D.Switch to instance store on each task and use SQS messages between tasks to copy intermediate artifacts.
AnswerB

EFS is designed for shared, NFS-like file storage that can be mounted concurrently from compute resources across multiple Availability Zones. By creating mount targets in each AZ used by the ECS tasks, you enable low-latency network access patterns so tasks can read and write the same shared filesystem reliably.

Why this answer

Amazon EFS provides a fully managed, shared NFS filesystem that can be mounted concurrently by ECS tasks across multiple Availability Zones with low latency. It supports POSIX file operations, making it ideal for streaming intermediate artifacts between tasks. EFS mount targets in each AZ ensure local access, meeting the requirement for high-performing shared storage.

Exam trap

The trap here is that candidates may assume EBS multi-attach works across Availability Zones, but it is strictly limited to a single AZ and requires specific instance types, making it unsuitable for multi-AZ shared filesystem requirements.

Why the other options are wrong

A

EBS multi-attach does not support attaching a single volume to instances across different Availability Zones; it only works within a single AZ. Therefore, cross-AZ tasks cannot share the same EBS volume.

C

S3 does not provide a POSIX-compliant shared filesystem with low-latency file locking and immediate consistency needed for streaming intermediate artifacts between tasks; it is an object store, not a filesystem.

D

Instance store is ephemeral and not shared across tasks, so tasks in different AZs cannot access a common filesystem. SQS message copying adds latency and complexity, failing the low-latency shared filesystem requirement.

67
MCQhard

A genomics company stores 400 TB of compressed reference data in Amazon S3. Researchers in an on-premises lab must run high-throughput reads of this data over a 10 Gbps AWS Direct Connect connection. The team observes that reads are slower than expected and wants to maximize throughput per S3 request while minimizing request costs. Which S3 feature should they implement?

A.S3 Cross-Region Replication to a bucket in the lab's nearest Region.
B.S3 byte-range fetches using parallel GET requests against the same object.
C.S3 Transfer Acceleration on the bucket, because it speeds up long-distance transfers.
D.S3 Intelligent-Tiering to automatically move frequently accessed data to a faster tier.
AnswerB

Byte-range fetches let a client issue multiple concurrent GET requests for different portions of one large object, which multiplies achievable throughput and uses the available Direct Connect bandwidth more fully. This is the standard S3 pattern for high-throughput reads of large objects and can reduce total request cost by retrieving full object data efficiently.

Why this answer

Large-object throughput in S3 scales with concurrent connections, so issuing parallel byte-range GETs against each object saturates the Direct Connect link far better than a single sequential GET. Transfer Acceleration targets internet paths, replication moves copies without changing read mechanics, and Intelligent-Tiering is a cost tool rather than a throughput tool.

Exam trap

The trap here is reaching for Transfer Acceleration whenever S3 feels slow, even when the workload already uses private Direct Connect connectivity.

68
MCQhard

Based on the exhibit, a DynamoDB-backed event processing system is throttling during a promotion. The table uses tenantId as the partition key and eventTime as the sort key. One tenant accounts for most of the write traffic, and the application must preserve fast lookups for that tenant without relying on a single hot partition. What change is the best fix?

A.Add a sharding suffix to the partition key, such as tenantId#shardId, and query across the tenant's shards.
B.Enable DynamoDB Streams so the table can process writes more quickly.
C.Switch the table to on-demand capacity mode and keep the same key design.
D.Add a global secondary index on eventTime and query the index instead of the base table.
AnswerA

Sharding the partition key spreads ACME traffic across multiple partitions, which removes the hot key problem. Because the application still needs tenant-scoped time-range queries, it can fan out across the shard values and merge results.

Why this answer

Adding a sharding suffix (e.g., tenantId#shardId) to the partition key distributes write traffic for the hot tenant across multiple partitions, eliminating the single-partition bottleneck while preserving fast lookups by querying across all shards for that tenant. DynamoDB's partition key determines physical storage; without sharding, all writes for the hot tenant land on one partition, causing throttling even if the table has sufficient total capacity.

Exam trap

The trap here is that candidates often assume on-demand mode (Option C) eliminates all throttling, but it does not resolve the physical partition limit—a single hot partition still caps at 1,000 WCU/3,000 RCU, so throttling persists regardless of capacity mode.

How to eliminate wrong answers

Option B is wrong because enabling DynamoDB Streams does not increase write throughput; it captures item-level changes asynchronously and does not alleviate throttling caused by a hot partition. Option C is wrong because switching to on-demand capacity mode only removes the need to provision capacity manually, but it does not solve the underlying hot partition issue—DynamoDB still throttles if a single partition exceeds 1,000 WCU or 3,000 RCU, regardless of capacity mode. Option D is wrong because adding a GSI on eventTime does not distribute write load; the base table's partition key remains tenantId, so the hot tenant still causes throttling on the base table, and the GSI inherits the same write patterns.

69
Multi-Selectmedium

A solutions architect is designing a high-performance architecture for a real-time analytics application. The application ingests a continuous stream of data from thousands of IoT devices. The data must be processed in near real-time, and the results must be stored in a durable, scalable data store for later analysis. The architect needs to choose AWS services that can handle the ingestion and processing of the stream. (Choose two.)

Select 2 answers
A.Use Amazon Kinesis Data Streams to ingest the data.
B.Use AWS Glue to process the stream in real time.
C.Use Amazon Redshift to ingest the data directly from the devices.
D.Use Amazon SQS FIFO queues to ingest the data.
E.Use AWS Lambda to process the stream in real time.
AnswersA, E

Amazon Kinesis Data Streams is designed for real-time streaming data ingestion at scale. It can handle thousands of data sources and provides durable, ordered storage of records for up to 7 days. It integrates with AWS Lambda, Kinesis Data Analytics, and other services for processing. This makes it ideal for ingesting the continuous stream from IoT devices, ensuring high throughput and low latency for the analytics application.

Why this answer

Amazon Kinesis Data Streams provides a scalable, durable ingestion layer for real-time streaming data, capable of handling thousands of IoT devices. AWS Lambda can be integrated as a consumer to process records in near real-time, automatically scaling with the stream. Together, they form a serverless pipeline that ingests and processes data with low latency, and can write results to a durable store like S3 or DynamoDB.

Exam trap

The trap here is assuming that any queue or ETL service can handle real-time streaming; SQS FIFO is for message decoupling with ordering, and Glue is for batch ETL, not continuous stream processing.

70
MCQmedium

A media archive requires consistent high IOPS for a transactional database on EC2. Which EBS volume type is most suitable? The architecture review board prefers a managed AWS-native control.

A.Provisioned IOPS SSD such as io2
B.st1 Throughput Optimized HDD
C.Instance store only
D.sc1 Cold HDD
AnswerA

Provisioned IOPS SSD volumes such as io2 deliver a guaranteed, consistent level of IOPS that is ideal for media archive databases or applications requiring predictable high random I/O performance. io2 also offers 99.999% durability and, with Block Express, can scale up to 256,000 IOPS per volume, making it the only option that truly meets the stated requirement for sustained high IOPS without compromising persistence.

Why this answer

The io2 Provisioned IOPS SSD volume type is designed for latency-sensitive transactional workloads that require consistent, high IOPS. It provides a service-level agreement (SLA) of 99.999% durability and supports up to 256,000 IOPS per volume, making it ideal for a database on EC2 that demands predictable performance. As a managed AWS-native EBS volume, it aligns with the architecture review board's preference for a fully AWS-controlled storage solution.

Exam trap

The trap here is that candidates often confuse throughput-optimized HDD (st1) with IOPS requirements, mistakenly thinking high throughput equals high IOPS, but transactional databases need random I/O performance, not sequential throughput.

How to eliminate wrong answers

Option B (st1 Throughput Optimized HDD) is wrong because it is optimized for sequential throughput, not random IOPS, and cannot deliver the consistent high IOPS required by a transactional database. Option C (Instance store only) is wrong because instance store volumes are ephemeral and data is lost on instance stop or termination, making them unsuitable for a persistent database. Option D (sc1 Cold HDD) is wrong because it is designed for infrequently accessed data with the lowest cost and lowest IOPS, far below the needs of a transactional database.

71
MCQeasy

A media company uses CloudFront in front of an S3 bucket origin for video thumbnails. They want to prevent users from bypassing CloudFront and accessing the S3 bucket directly, while still allowing CloudFront to fetch objects. What is the best option?

A.Keep the bucket public and rely on signed cookies for all thumbnail requests.
B.Use CloudFront Origin Access Control (OAC) or Origin Access Identity (OAI) and update the bucket policy to allow only CloudFront.
C.Enable S3 static website hosting so users access thumbnails directly from the S3 website endpoint.
D.Set S3 bucket permissions to allow all IAM users and block access only by using a WAF rule at CloudFront.
AnswerB

OAC or OAI makes CloudFront sign origin requests with a service principal, so the bucket policy can grant read access solely to that identity. This blocks direct S3 URLs, satisfying the requirement that users cannot bypass CloudFront while CloudFront still fetches objects.

Why this answer

CloudFront Origin Access Control (OAC) or Origin Access Identity (OAI) allows you to restrict direct access to an S3 bucket by configuring the bucket policy to grant read permissions only to the CloudFront distribution's service principal. This ensures that users can only retrieve thumbnails through CloudFront, leveraging its caching and security features, while blocking any direct S3 requests.

Exam trap

The trap here is that candidates often think signed cookies or URLs alone are sufficient to secure direct S3 access, but they forget that those mechanisms only control access through CloudFront and do not restrict the S3 bucket's public endpoint unless the bucket policy explicitly denies direct access.

Why the other options are wrong

A

Keeping the bucket public allows anyone with the S3 URL to access thumbnails directly, bypassing CloudFront. Signed cookies control access via CloudFront but do not prevent direct S3 access, so this fails to meet the requirement.

C

Enabling S3 static website hosting exposes the S3 bucket via its website endpoint, allowing users to bypass CloudFront and access thumbnails directly, which violates the requirement to prevent direct access.

D

WAF rules at CloudFront can block certain requests but do not prevent direct access to the S3 bucket; users could still bypass CloudFront and access the bucket directly if the bucket policy allows it.

72
MCQeasy

Your application uses ElastiCache Redis as a cache for user profiles stored in DynamoDB. You must ensure that when a profile is updated, subsequent reads see the latest value quickly. Which cache strategy is generally the best fit for this requirement?

A.Write to DynamoDB only, and never update or invalidate the Redis cache.
B.Use a cache-aside approach with TTL plus explicit invalidation after writes.
C.Cache only for reads, and do not fetch from DynamoDB when a key is missing.
D.Rely on eventual consistency of Redis replication to propagate updates to all nodes.
AnswerB

A cache-aside (lazy loading) pattern reads from cache first; if missing/expired, it fetches from the source of truth. After an update, explicitly invalidating or updating the cached entry ensures subsequent reads quickly reflect changes. TTL provides protection against missed invalidations while invalidation accelerates correctness after writes.

Why this answer

A cache-aside (lazy loading) strategy with TTL and explicit invalidation ensures that after a write to DynamoDB, the stale Redis entry is removed, forcing the next read to fetch the fresh profile from DynamoDB and repopulate the cache. This combination minimizes the window of stale reads while maintaining high read performance, which is critical for user profile caches where consistency matters.

Exam trap

The trap here is that candidates often confuse eventual consistency within Redis replication (which only applies to Redis-to-Redis sync) with the need to synchronize the cache with the authoritative data store (DynamoDB), leading them to pick option D, which does not address the core requirement of reflecting DynamoDB updates in the cache.

How to eliminate wrong answers

Option A is wrong because never updating or invalidating the Redis cache means stale data persists indefinitely, violating the requirement that subsequent reads see the latest value quickly. Option C is wrong because caching only for reads and not fetching from DynamoDB when a key is missing would result in cache misses returning no data, effectively breaking the application's ability to serve user profiles. Option D is wrong because relying on eventual consistency of Redis replication does not guarantee that updates to DynamoDB are reflected in the cache; Redis replication only synchronizes data between Redis nodes, not between DynamoDB and Redis, and does not address cache invalidation after writes.

73
MCQeasy

An ECS service runs on EC2 capacity. During peak traffic, tasks frequently wait for available container instances. The team wants faster scale-out for the underlying EC2 capacity when tasks increase. What is the best first architectural step?

A.Tune the container health check settings so tasks stop failing and stay running.
B.Use an ECS capacity provider (or Auto Scaling integration) to scale the EC2 instances based on ECS demand.
C.Pin all tasks to a single Availability Zone to reduce placement overhead.
D.Switch the tasks to run only on Fargate so EC2 scaling is no longer relevant.
AnswerB

When ECS tasks need compute, capacity must scale at the EC2 layer so there are enough container instances to place tasks. Integrating ECS with an Auto Scaling capacity provider allows the cluster to scale out in response to pending tasks. This reduces waiting time and improves responsiveness under load.

Why this answer

An ECS capacity provider (or Auto Scaling integration) directly links ECS task-level demand to EC2 instance scaling. When tasks are pending due to insufficient container instances, the capacity provider triggers a scale-out event on the Auto Scaling group, adding EC2 instances to accommodate the workload. This is the most efficient architectural step to reduce placement delays during peak traffic.

Exam trap

The trap here is that candidates may confuse task-level scaling (e.g., Service Auto Scaling) with infrastructure-level scaling, and incorrectly assume that tuning health checks or placement strategies will resolve a capacity shortage caused by insufficient EC2 instances.

Why the other options are wrong

A

Tuning health check settings does not address the root cause of tasks waiting for EC2 capacity; it only affects task stability, not instance availability.

C

Pinning tasks to a single Availability Zone does not address the root cause of insufficient EC2 capacity; it actually reduces fault tolerance and may increase placement constraints, making scaling slower.

D

Switching to Fargate eliminates EC2 scaling concerns but does not address the existing EC2 capacity scaling issue; the question specifically asks for faster scale-out of underlying EC2 capacity, not a migration to a different compute type.

74
MCQeasy

An application uses DynamoDB to store order status. Reads happen extremely frequently for the same few keys (for example, the most recent orders), and the team wants lower read latency without changing the table’s partition key design. Which AWS service best fits this requirement?

A.Amazon DAX (DynamoDB Accelerator) to cache frequently read items
B.Provision AWS WAF rules to reduce DynamoDB read latency caused by bots
C.Enable multi-region writes in DynamoDB Global Tables to speed up reads locally
D.Add more read capacity units to DynamoDB and avoid caching entirely
AnswerA

DAX is an in-memory caching layer specifically built for DynamoDB. It reduces read latency for hot keys by serving cached responses quickly while still reading from DynamoDB when a key is not cached (or when the cached entry expires). This avoids the need to redesign partition keys.

Why this answer

Amazon DAX (DynamoDB Accelerator) is an in-memory cache that sits between your application and DynamoDB, providing microsecond read latency for frequently accessed items. Because the workload involves extremely frequent reads of the same few keys (hot keys), DAX reduces the load on the DynamoDB table and delivers faster responses without requiring any changes to the partition key design.

Exam trap

The trap here is that candidates often confuse throughput scaling (adding RCUs) with latency optimization, or they mistakenly think Global Tables improve read latency within a single region, when in fact DAX is the only option that directly caches hot keys to reduce read latency without altering the table design.

Why the other options are wrong

B

AWS WAF is a web application firewall that protects against web exploits and bots, but it does not reduce read latency for DynamoDB queries. The latency issue here is due to high read frequency on a few keys, not bot traffic.

C

Multi-region writes in DynamoDB Global Tables improve write availability and disaster recovery, but they do not reduce read latency for frequently accessed hot keys because reads still go to the local region's DynamoDB table without caching. The question specifically asks for lower read latency without changing the partition key design, which DAX addresses via caching.

D

Adding more read capacity units increases throughput but does not reduce latency for hot keys; it only prevents throttling. The question specifically asks for lower read latency, which caching addresses directly.

75
MCQeasy

A travel booking site uses EC2 instances behind an ALB. CPU is consistently high during peak traffic, and request latency rises. What should be configured? The design must avoid adding custom operational scripts.

A.A VPC endpoint for CloudWatch only
B.Auto Scaling policy based on an appropriate CloudWatch metric
C.S3 Object Lock
D.Disable health checks
AnswerB

Auto Scaling adds capacity when load increases and removes it when load falls.

Why this answer

An Auto Scaling policy based on a CloudWatch metric like CPUUtilization or request latency directly addresses the high CPU and rising latency by automatically adding EC2 instances during peak traffic. This eliminates the need for custom scripts and ensures the application scales horizontally to maintain performance.

Exam trap

The trap here is that candidates might think a VPC endpoint (Option A) is needed for CloudWatch metrics, but CloudWatch metrics are already available without a VPC endpoint, and scaling requires an Auto Scaling policy, not just metric access.

How to eliminate wrong answers

Option A is wrong because a VPC endpoint for CloudWatch only enables private connectivity to CloudWatch, but does not provide any scaling or performance improvement for EC2 instances. Option C is wrong because S3 Object Lock is used for data retention and compliance, not for scaling compute resources or reducing latency. Option D is wrong because disabling health checks would cause the ALB to route traffic to unhealthy instances, worsening latency and availability issues.

Page 1 of 3 · 215 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Design High-Performing Architectures questions.

CCNA Design High-Performing Architectures Questions — Page 1 of 3 | Courseiva