Courseiva

CCNA Data Store Management Questions

75 of 358 questions · Page 1/5 · Data Store Management topic · Answers revealed

1
MCQeasy

A company stores application logs in Amazon S3 and wants to analyze them using Amazon Athena. The logs are in JSON format and are compressed with gzip. The data engineer needs to create an Athena table that can query these logs efficiently. The logs are stored in an S3 bucket with the prefix logs/year=2023/month=10/day=15/. The engineer wants to minimize query costs and improve performance. Which action should the engineer take?

A.Create an external table with the JSON SerDe, specifying the S3 location as s3://bucket/logs/ and using the OpenCSVSerDe for better performance.
B.Create an external table with the CSV SerDe, specifying the S3 location as s3://bucket/logs/ and defining partitions for year, month, and day.
C.Create an external table with the JSON SerDe, specifying the S3 location as s3://bucket/logs/year=2023/month=10/day=15/ and no partitions.
D.Create an external table with the JSON SerDe, specifying the S3 location as s3://bucket/logs/ and defining partitions for year, month, and day.
AnswerD

Creating an external table with the JSON SerDe and defining partitions allows Athena to read the JSON logs and use partition pruning to scan only relevant data. Specifying the root S3 location and adding partitions (either manually or via MSCK REPAIR) enables efficient queries and reduces cost by limiting data scanned.

Why this answer

To query JSON logs efficiently in Athena, the table must use the JSON SerDe to parse the data correctly. Defining partitions on year, month, and day allows Athena to prune partitions based on query filters, reducing the amount of data scanned and lowering costs. The table location should point to the root prefix so that all partitions are included.

This combination provides accurate parsing and cost-effective queries.

Exam trap

The trap here is selecting a SerDe based on compression or performance assumptions rather than matching it to the actual data format, which is JSON.

2
Multi-Selectmedium

A data engineer is optimizing an Amazon S3 data lake for cost and performance. The data lake contains large volumes of CSV files that are queried by Amazon Athena. The engineer wants to reduce query costs and improve query performance. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Partition the data by a frequently filtered column, such as date.
B.Enable S3 server access logging on the bucket.
C.Enable S3 Transfer Acceleration on the bucket.
D.Convert the CSV files to Apache Parquet format.
E.Increase the number of Athena workgroups.
AnswersA, D

Partitioning by a commonly filtered column allows Athena to scan only the relevant partitions, reducing the amount of data read. This lowers query cost and improves performance by skipping irrelevant data. It is a best practice for large datasets in S3 when queries filter on that column. Combined with Parquet, it maximizes savings.

Why this answer

Converting CSV to Parquet and partitioning the data by a frequently filtered column are the two most effective actions to reduce Athena query costs and improve performance. Parquet reduces scanned data through columnar compression and predicate pushdown, while partitioning limits the data scanned to relevant partitions. The other actions do not affect query execution or data scanned.

Exam trap

The trap here is confusing S3 bucket-level features like Transfer Acceleration or access logging with query optimization, when the real gains come from data format and partitioning.

3
MCQhard

A company stores sensitive financial data in an Amazon S3 bucket. The data engineering team must ensure that all data is encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys are rotated annually. The team also needs to audit key usage. Which solution meets these requirements?

A.Enable default encryption on the S3 bucket using SSE-S3, and configure S3 server access logging to track key usage.
B.Use client-side encryption with a customer provided key stored in AWS Secrets Manager, and enable AWS CloudTrail to log S3 API calls.
C.Enable default encryption on the S3 bucket using SSE-KMS with an AWS managed key, and configure S3 Inventory to report on encryption status.
D.Enable default encryption on the S3 bucket using SSE-KMS with a customer managed key, enable automatic key rotation for the KMS key, and enable AWS CloudTrail to log KMS API calls.
AnswerD

SSE-KMS with a customer managed key satisfies the encryption requirement. Enabling automatic key rotation on the KMS key rotates the key annually. AWS CloudTrail logs all KMS API calls, providing an audit trail of key usage. This combination meets all specified requirements: encryption with customer managed keys, annual rotation, and auditing.

Why this answer

The correct solution uses SSE-KMS with a customer managed key to encrypt data at rest, enables automatic key rotation for that key to meet the annual rotation requirement, and enables AWS CloudTrail to log KMS API calls for auditing key usage. Other options either use incorrect key types (SSE-S3 or AWS managed keys) or do not provide the necessary auditing capabilities.

Exam trap

The trap here is confusing AWS managed keys with customer managed keys and assuming that S3 server access logging or S3 Inventory can audit KMS key usage.

4
MCQmedium

A company uses Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The data is JSON and each record is about 2 KB. The delivery stream is configured to buffer incoming data to 5 MB or 60 seconds, whichever comes first. The data engineering team notices that the S3 bucket contains many small files (average 2 MB), which makes subsequent processing inefficient. They need to reduce the number of small files without increasing the latency beyond 5 minutes. Which solution should they implement?

A.Enable compression (GZIP) on the delivery stream.
B.Increase the buffer size to 50 MB and the buffer interval to 300 seconds.
C.Use a Lambda function to merge small files after delivery.
D.Decrease the buffer size to 1 MB and the buffer interval to 60 seconds.
AnswerB

Raising the buffer to 50 MB lets Firehose accumulate far more records before each S3 PUT, directly cutting the small-file count. The 300-second interval still satisfies the stem's five-minute latency ceiling, so larger objects arrive without breaching the stated constraint.

Why this answer

Increasing the buffer size to 50 MB and the buffer interval to 300 seconds directly addresses the root cause: the current 5 MB buffer size triggers a flush too frequently, producing many 2 MB files. By raising the buffer size to 50 MB, each flush will contain more data, resulting in larger S3 objects (up to ~50 MB uncompressed), while the 300-second interval ensures latency stays within the 5-minute requirement. This reduces the number of small files without requiring additional services or post-processing.

Exam trap

The trap here is that candidates often assume compression (Option A) reduces file count, but compression only reduces file size, not the number of files; the real issue is the flush frequency controlled by buffer size and interval.

How to eliminate wrong answers

Option A is wrong because enabling GZIP compression reduces the storage size of each file but does not change the buffer size or flush behavior; the delivery stream will still flush at 5 MB or 60 seconds, producing the same number of small files (just compressed). Option C is wrong because using a Lambda function to merge small files after delivery adds complexity, cost, and latency (Lambda invocation delays, S3 event processing), and does not address the root cause of premature flushes; it also violates the goal of not increasing latency beyond 5 minutes due to the merging overhead. Option D is wrong because decreasing the buffer size to 1 MB and keeping the buffer interval at 60 seconds would make the problem worse, producing even smaller files (average ~1 MB) and more frequent flushes, increasing the number of small files.

5
MCQmedium

A data engineer is building an Amazon Redshift data warehouse. The cluster will ingest data from Amazon S3 using the COPY command. The engineer needs to ensure that the data is loaded in a way that maximizes query performance for future complex analytical queries. The data is currently stored as uncompressed CSV files in S3. Which action should the engineer take to optimize the load and subsequent query performance?

A.Use the COPY command with the GZIP option to load compressed files, which automatically distributes the data evenly across all nodes.
B.Load the data into a table with no compression and then run a VACUUM command to reorganize and compress the data.
C.Load the data into a single large staging table and then use CREATE TABLE AS (CTAS) to redistribute it into multiple tables.
D.Use the COPY command with the COMPUPDATE ON option to automatically apply compression, and define an appropriate distribution key and sort key on the target table.
AnswerD

Using COPY with COMPUPDATE ON allows Redshift to analyze the data and apply optimal compression encodings automatically, which reduces storage and improves I/O performance. Defining a distribution key ensures even data distribution across nodes, and a sort key improves query performance by enabling efficient range scans and minimizing data movement during joins. This combination directly addresses the need for optimized load and query performance.

Why this answer

The correct approach is to use COPY with COMPUPDATE ON to automatically apply compression, and to define an appropriate distribution key and sort key. This optimizes storage, reduces I/O, and improves query performance by ensuring even data distribution and efficient data access patterns. Other options either do not address compression, distribution, or sort keys adequately, or rely on commands that do not achieve the desired optimization.

Exam trap

The trap here is assuming that loading compressed files or running VACUUM will automatically optimize the table's physical design for query performance.

6
Multi-Selecteasy

Which TWO features of Amazon S3 help protect data from accidental deletion or modification? (Choose two.)

Select 2 answers
A.Lifecycle policies
B.Default encryption
C.S3 MFA Delete
D.S3 Object Versioning
E.Cross-Region Replication
AnswersC, D

S3 MFA Delete enforces a second authentication factor on delete operations, requiring the bucket owner's MFA token to permanently remove object versions or change versioning state. This directly satisfies the accidental-deletion constraint by blocking unauthorised or mistaken deletions.

Why this answer

C is correct because S3 MFA Delete requires multi-factor authentication for permanent deletion of objects or suspension of versioning, adding a critical layer of protection against accidental or malicious deletions. D is correct because S3 Object Versioning preserves every version of an object, allowing recovery from accidental overwrites or deletions by restoring a previous version.

Exam trap

The trap here is that candidates often confuse data protection features (like encryption or replication) with deletion prevention, mistakenly selecting Lifecycle policies or Cross-Region Replication because they think 'protection' includes backup or security, when the question specifically asks about preventing accidental deletion or modification.

7
MCQmedium

A data engineer manages an Amazon S3 data lake that ingests millions of small JSON files daily from an IoT fleet. Query performance in Amazon Athena has degraded significantly, and each query scans far more data than expected. The engineer wants to reduce per-query cost and improve performance without changing the raw data. Which solution should the engineer implement?

A.Convert the JSON files to Apache Parquet, partition the data, and compact small files with AWS Glue ETL jobs.
B.Move the data into an Amazon Redshift cluster and query it with Redshift Spectrum.
C.Increase the Athena workgroup data usage control limit and rerun the queries.
D.Enable S3 Transfer Acceleration on the ingestion bucket.
AnswerA

Athena charges by bytes scanned, so columnar Parquet with compression and partition pruning dramatically reduces data read. Compacting millions of tiny files into larger row groups removes per-object overhead and improves parallel scan throughput. Partitioning by a common filter column such as device date or region lets Athena skip irrelevant prefixes entirely, cutting both latency and cost.

Why this answer

Athena performance and cost are driven by how many bytes a query scans, so converting row-oriented JSON to columnar Parquet, partitioning on common filter columns, and compacting small files directly reduce scanned data. These three changes work together: columnar storage reads only referenced columns, partitions prune entire prefixes, and compaction removes per-object overhead. The result is faster queries at lower cost without altering the source data in S3.

Exam trap

The trap here is assuming that a networking or workgroup setting such as S3 Transfer Acceleration or a higher data usage limit can fix slow Athena queries, when the real driver is bytes scanned and file layout.

8
MCQmedium

A data engineering team stores clickstream events in an Amazon S3 bucket under the prefix s3://analytics/raw/. New objects arrive continuously, and the team wants Amazon Athena queries to scan only the events for the current day without scanning the entire prefix. The events are written as JSON files partitioned by year/month/day, but queries still scan all partitions because the partition metadata is not registered. Which action should the data engineer take to enable partition pruning in Athena?

A.Run MSCK REPAIR TABLE on the Athena table to add the year/month/day partitions to the AWS Glue Data Catalog.
B.Create an Amazon CloudFront distribution in front of the S3 bucket and point Athena at the distribution.
C.Enable S3 Transfer Acceleration on the bucket so Athena can retrieve the daily events faster.
D.Convert the JSON files to Apache Parquet and rely on columnar storage alone to limit the scan to one day.
AnswerA

MSCK REPAIR TABLE scans the S3 prefix, discovers Hive-style partition folders such as year=2024/month=03/day=15, and registers each partition in the Data Catalog so Athena can prune non-matching partitions. This directly addresses the missing partition metadata that causes full-prefix scans in this scenario.

Why this answer

Partition pruning requires the query engine to know which partition values exist and where their data lives. Registering the year/month/day folders as partitions in the Data Catalog lets Athena build a plan that reads only the matching day, cutting scanned bytes and cost. Physical optimizations such as acceleration or columnar formats do not supply that partition metadata.

Exam trap

The trap here is assuming that a columnar file format or faster data transfer alone limits how much data Athena scans, when partition pruning depends on registered partition metadata.

9
Multi-Selectmedium

Which TWO of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose TWO.)

Select 2 answers
A.Improves write throughput
B.Provides microsecond read latency
C.Offloads read traffic from the DynamoDB table
D.Provides data durability across Availability Zones
E.Reduces storage costs
AnswersB, C

DAX is an in-memory write-through cache sitting in front of DynamoDB, serving eventually consistent reads from RAM in microseconds rather than the single-digit millisecond latency of direct DynamoDB access. This satisfies the microsecond read latency benefit.

Why this answer

Option B is correct because DAX is an in-memory write-through cache placed in front of DynamoDB that serves eventually consistent reads from RAM, delivering microsecond read latency for cached items instead of the single-digit millisecond latency of direct DynamoDB reads. Option C is correct because read requests that hit the DAX cache are answered by DAX itself, so they never reach the DynamoDB table, which offloads read traffic and reduces the read capacity units consumed on the underlying table. Option A is not correct because DAX does not accelerate writes: write operations are still written through to DynamoDB synchronously, so write throughput and write latency remain governed by the table.

Option D is not correct because durability across Availability Zones is provided by DynamoDB's own multi-AZ replication of table data, not by DAX, which is a cache and can lose cached data. Option E is not correct because DAX does not reduce the storage consumed by the DynamoDB table; it is an additional caching layer that itself incurs cost.

Exam trap

The trap here is that candidates often confuse DAX's read acceleration with write performance improvements, or assume that a cache provides durability guarantees similar to the underlying database.

10
MCQhard

A financial services company stores trade records in Amazon DynamoDB. An application performs many reads per second for individual trades by tradeId and also runs a nightly analytics job that must read every trade for a given trading day. The table uses tradeId as the partition key only. The nightly job currently performs a full table scan and is slow and expensive. The company wants to optimize the nightly access pattern with minimal application changes. Which change should the data engineer make?

A.Add a local secondary index on the base table with tradeId as the sort key and query it for each day.
B.Create a global secondary index with a partition key of tradingDay and a sort key of tradeId, then query that index for each day.
C.Increase the table's provisioned read capacity units and schedule the full table scan during off-peak hours.
D.Enable DynamoDB Streams on the table and have the nightly job consume the stream to reconstruct the day's trades.
AnswerB

A global secondary index with tradingDay as partition key groups all trades for a day into one logical partition, so the nightly job can issue efficient Query operations instead of scanning the whole table. Keeping tradeId as the sort key allows ordering and range access. The base table partition key and application writes remain unchanged, meeting the minimal-change requirement.

Why this answer

DynamoDB access is most efficient when queries target a specific partition key value. Because the base table partitions by tradeId, retrieving all trades for a day requires a Scan. A global secondary index partitioned by tradingDay gives each day its own logical partition, so the nightly job can use Query to read only that day's items.

The base table schema and write path stay the same, so application changes are minimal.

Exam trap

The trap here is thinking that a local secondary index can change the partition grouping, when it must reuse the base table partition key and therefore cannot group by trading day.

11
MCQmedium

A data engineer notices that an Amazon Redshift cluster is experiencing slow query performance. The engineer suspects that tables are not properly sorted. Which diagnostic query should the engineer run to identify unsorted rows?

A.SELECT * FROM SVV_TABLE_INFO ORDER BY unsorted DESC;
B.SELECT * FROM PG_CATALOG;
C.SELECT * FROM STV_TBL_PERM;
D.SELECT * FROM STL_LOAD_ERRORS;
AnswerA

SVV_TABLE_INFO exposes per-table statistics including the unsorted percentage, so ordering by unsorted DESC surfaces the tables with the most unsorted rows. That directly identifies where a VACUUM SORT is needed to address the slow query performance.

Why this answer

The `SVV_TABLE_INFO` system view in Amazon Redshift provides metadata about each table, including the `unsorted` column which shows the percentage of unsorted rows. By ordering by `unsorted DESC`, the engineer can quickly identify tables with the highest proportion of unsorted data, which directly impacts query performance due to inefficient zone maps and scan pruning.

Exam trap

The trap here is that candidates may confuse `SVV_TABLE_INFO` with `STV_TBL_PERM` (which shows block counts) or `STL_LOAD_ERRORS` (which is for load debugging), missing that only `SVV_TABLE_INFO` exposes the `unsorted` column specifically designed for sort health analysis.

How to eliminate wrong answers

Option B is wrong because `PG_CATALOG` is a system schema containing PostgreSQL catalog tables (e.g., `pg_class`, `pg_attribute`), not a diagnostic view for unsorted rows; it lacks the `unsorted` metric. Option C is wrong because `STV_TBL_PERM` provides block-level storage information (e.g., number of blocks per slice) but does not include a column for unsorted row percentage. Option D is wrong because `STL_LOAD_ERRORS` logs errors from COPY and INSERT operations, such as data type mismatches or malformed CSV rows, and has no relevance to sort key efficiency.

12
MCQeasy

A company uses Amazon DynamoDB for a gaming application. They need to store player session data that expires after 24 hours. Which DynamoDB feature should they use to automatically delete expired items?

A.Time to Live (TTL)
B.DynamoDB auto scaling
C.DynamoDB Streams
D.Point-in-time recovery
AnswerA

Time to Live (TTL) lets you define an attribute holding an expiry timestamp; DynamoDB then deletes each item automatically once that timestamp passes, with no write capacity consumed. This directly satisfies the 24-hour automatic expiry requirement for player session data, eliminating manual cleanup jobs.

Why this answer

DynamoDB Time to Live (TTL) is the correct feature because it allows you to define a per-item timestamp attribute (e.g., `expireAt`) that DynamoDB automatically deletes once that timestamp is reached. This is ideal for expiring session data after 24 hours without requiring custom code or scheduled jobs to scan and delete items, reducing cost and operational overhead.

Exam trap

The trap here is that candidates may confuse TTL with DynamoDB Streams, thinking streams can automatically delete items, but streams only notify of changes and require separate logic to perform deletions.

How to eliminate wrong answers

Option B (DynamoDB auto scaling) is wrong because it manages throughput capacity (read/write units) based on traffic, not item expiration or deletion. Option C (DynamoDB Streams) is wrong because it captures item-level changes (inserts, updates, deletes) in near real-time for downstream processing, but does not automatically delete items. Option D (Point-in-time recovery) is wrong because it provides continuous backups to restore a table to any point within the last 35 days, but does not handle automatic deletion of expired data.

13
MCQhard

A company is using Amazon Redshift for data warehousing. The data engineer notices that the STL_ALERT_EVENT_LOG table shows many 'missing statistics' alerts. What is the best course of action to address this issue?

A.Increase the WLM concurrency slots.
B.Run VACUUM on the tables.
C.Enable compression on the tables.
D.Run ANALYZE on the tables.
AnswerD

ANALYZE refreshes table statistics that the query planner uses to build optimal execution plans. Missing statistics force suboptimal plans, degrading performance and generating STL_ALERT_EVENT_LOG entries. Running ANALYZE directly resolves the alerts by restoring accurate cardinality estimates.

Why this answer

The STL_ALERT_EVENT_LOG table records alerts about query performance issues, including 'missing statistics' alerts. This indicates that the query optimizer lacks up-to-date table statistics, leading to suboptimal query plans. Running the ANALYZE command updates table statistics, enabling the optimizer to generate efficient execution plans.

Therefore, option D is the correct course of action.

Exam trap

The trap here is that candidates often confuse VACUUM (which reorganizes data) with ANALYZE (which updates statistics), assuming both are needed for query performance, but only ANALYZE directly resolves 'missing statistics' alerts.

How to eliminate wrong answers

Option A is wrong because increasing WLM concurrency slots does not address missing statistics; it only allows more queries to run simultaneously, which could worsen performance if statistics are outdated. Option B is wrong because VACUUM reclaims disk space and sorts rows but does not update table statistics; it is used for managing data storage, not query optimization. Option C is wrong because enabling compression reduces storage and I/O but does not provide the optimizer with the statistical metadata needed for efficient query planning.

14
MCQmedium

A company is using Amazon RDS for MySQL and needs to reduce read latency for a global user base. Which AWS feature should be implemented?

A.Multi-AZ deployment
B.Aurora Auto Scaling
C.Read Replicas
D.Cross-Region Replication
AnswerC

Read Replicas allow offloading read queries to reduce latency.

Why this answer

Amazon RDS Read Replicas allow you to offload read traffic from the primary DB instance to one or more read-only copies, which can be placed in different AWS Regions to reduce read latency for a global user base. Unlike Multi-AZ, which is designed for high availability, Read Replicas directly address read performance and latency by distributing read queries closer to users. Cross-Region Replication for RDS MySQL is achieved through Read Replicas, making option C the correct choice for reducing read latency globally.

Exam trap

The trap here is that candidates often confuse Multi-AZ (high availability) with Read Replicas (read scaling), or they assume Cross-Region Replication is a separate feature when it is actually implemented via Read Replicas in RDS MySQL.

How to eliminate wrong answers

Option A is wrong because Multi-AZ deployment provides high availability and automatic failover by maintaining a standby replica in a different Availability Zone, but it does not offload read traffic or reduce read latency for a global user base. Option B is wrong because Aurora Auto Scaling automatically adjusts the number of Aurora Replicas based on workload, but it is a feature of Amazon Aurora, not Amazon RDS for MySQL, and the question specifies RDS for MySQL. Option D is wrong because Cross-Region Replication is not a standalone feature for RDS MySQL; it is implemented using Read Replicas in a different region, so it is a subset of the correct answer, not a separate feature.

15
MCQmedium

A company runs a multi-AZ Amazon RDS for PostgreSQL instance. They need to run a one-time analytical query that will take several hours and consume significant I/O. The query should not impact the primary workload. What should the data engineer do?

A.Create a read replica of the RDS instance and run the query on the replica.
B.Run the query directly on the primary instance during off-peak hours.
C.Increase the instance size to handle the load.
D.Enable Multi-AZ and run the query on the standby instance.
AnswerA

A read replica offloads the analytical query onto a separate database instance, isolating its heavy I/O from the primary. Because replication is asynchronous, the primary workload keeps running unaffected, satisfying the requirement that the query must not impact production. Note that a multi-AZ standby cannot serve read traffic, so it would not work here.

Why this answer

Creating a read replica of the RDS for PostgreSQL instance allows the analytical query to run on a separate database engine without affecting the primary workload. Read replicas in Amazon RDS use asynchronous replication from the source instance, so the replica can handle heavy I/O and long-running queries independently. This ensures the primary instance remains available for the production workload without performance degradation.

Exam trap

The trap here is that candidates often confuse the Multi-AZ standby instance with a read replica, assuming the standby can be used for queries, but in Amazon RDS, the standby is only for high availability and is not accessible for read operations.

How to eliminate wrong answers

Option B is wrong because running the query directly on the primary instance, even during off-peak hours, still consumes significant I/O and CPU resources on that instance, which can impact the primary workload and potentially cause performance issues or increased latency. Option C is wrong because increasing the instance size only adds more resources to the same single instance; the analytical query would still compete with the primary workload for I/O and memory, and scaling up does not isolate the workload. Option D is wrong because the standby instance in a Multi-AZ deployment is not directly accessible for read or write operations; it is a synchronous replica used only for automatic failover, and Amazon RDS does not allow connecting to the standby for queries.

16
MCQhard

A data engineer is designing a multi-region disaster recovery plan for an Amazon DynamoDB table. The table stores critical user profile data and must have a Recovery Point Objective (RPO) of less than 1 minute and a Recovery Time Objective (RTO) of less than 5 minutes. Which solution meets these requirements?

A.Configure DynamoDB Streams and a Lambda function to replicate data to another region.
B.Use DynamoDB on-demand backup and restore to another region.
C.Use DynamoDB global tables to replicate data to another region.
D.Enable point-in-time recovery (PITR) and restore to another region.
AnswerC

DynamoDB global tables use multi-region, active-active replication that propagates writes across regions, typically within one second. This satisfies the sub-one-minute RPO, while the replica's continuous availability enables failover well inside the five-minute RTO without restore operations.

Why this answer

DynamoDB global tables provide active-active multi-region replication with sub-second latency, ensuring that data written in one region is automatically replicated to other regions within seconds. This meets the RPO of less than 1 minute and RTO of less than 5 minutes because the table is already available in the secondary region for immediate reads and writes, with no manual restore or failover steps required.

Exam trap

The trap here is that candidates confuse point-in-time recovery (PITR) or on-demand backups with multi-region replication, not realizing that restore operations are manual and time-consuming, whereas global tables provide automatic, near-real-time replication.

How to eliminate wrong answers

Option A is wrong because DynamoDB Streams with a Lambda function introduces asynchronous replication that can have variable latency and potential data loss if the Lambda fails or is throttled, making it unreliable for a sub-1-minute RPO. Option B is wrong because on-demand backup and restore is a manual process that takes minutes to hours to complete, far exceeding the 5-minute RTO and 1-minute RPO. Option D is wrong because point-in-time recovery (PITR) only allows restoring to a point within the last 35 days in the same region, and restoring to another region requires exporting and re-importing data, which cannot achieve sub-5-minute RTO.

17
MCQmedium

A company has an Amazon RDS for PostgreSQL DB instance with a large table that is frequently updated. The data engineer needs to reduce storage costs by archiving old records that are no longer accessed. The archived records must be retained for 7 years due to compliance requirements. Which solution is MOST cost-effective?

A.Use RDS native backup and restore to keep a separate backup.
B.Export old records using pg_dump and store in S3 Glacier Deep Archive.
C.Enable storage autoscaling on the RDS instance.
D.Move old records to a separate table in the same RDS instance.
AnswerB

pg_dump extracts the old records, and S3 Glacier Deep Archive provides the lowest-cost long-term retention tier suitable for seven-year compliance archiving. This removes the data from expensive RDS storage while preserving it durably and cheaply.

Why this answer

Exporting old records via pg_dump and storing them in S3 Glacier Deep Archive provides the lowest-cost storage for data that must be retained for 7 years but is never accessed. S3 Glacier Deep Archive offers retrieval times of 12–48 hours at a storage cost of approximately $0.00099/GB/month, far cheaper than any RDS storage tier. This approach removes the archived data from the RDS instance, reducing provisioned storage costs while meeting compliance requirements.

Exam trap

The trap here is that candidates confuse 'archiving' with 'backup' or 'storage autoscaling,' failing to recognize that only moving data out of the RDS instance to a low-cost storage class like S3 Glacier Deep Archive actually reduces ongoing storage costs.

How to eliminate wrong answers

Option A is wrong because RDS native backup and restore creates full or snapshot backups of the entire DB instance, not just the old records, and storing these backups for 7 years would incur high costs for redundant data and storage fees. Option C is wrong because enabling storage autoscaling only increases storage capacity when thresholds are reached, which does not reduce costs or archive old records—it merely prevents out-of-space errors. Option D is wrong because moving old records to a separate table in the same RDS instance does not reduce storage costs; the data still occupies the same provisioned storage, and the table remains part of the instance's billed capacity.

18
Multi-Selecthard

A data engineering team is building a data lake on Amazon S3. They need to catalog data and make it queryable by Amazon Athena and Amazon Redshift Spectrum. The data arrives in multiple formats and the schema evolves frequently. Which TWO actions should the team take to support schema evolution and efficient querying? (Choose two.)

Select 2 answers
A.Define tables in the AWS Glue Data Catalog and use AWS Glue crawlers to infer and update schemas as new data arrives.
B.Store all data as uncompressed CSV to maximize compatibility with all query engines.
C.Create a separate S3 bucket for each schema version and require analysts to query the correct bucket manually.
D.Use open columnar formats such as Apache Parquet or ORC with partition prefixes, and register partitions in the Data Catalog.
E.Enable S3 Object Lock on the data lake bucket to preserve schema versions.
AnswersA, D

The AWS Glue Data Catalog is the central metastore used by Athena and Redshift Spectrum. Crawlers inspect data, infer schemas, and update table definitions, so new columns and partitions are reflected automatically. This supports frequent schema evolution without manual DDL, and both query engines can immediately use the updated metadata.

Why this answer

The AWS Glue Data Catalog with crawlers keeps table and partition metadata current as schemas evolve, and both Athena and Redshift Spectrum read from it. Using open columnar formats such as Parquet or ORC with partitioned prefixes enables column projection and partition pruning, which cut scanned bytes and cost. Together, these two actions support evolving schemas and efficient querying across both engines.

Exam trap

The trap here is assuming that a single format or manual bucket separation solves schema evolution, when the Data Catalog plus columnar, partitioned storage is what enables both evolution and efficient querying.

19
MCQhard

A company uses Amazon DynamoDB as the primary data store for a gaming application. The application experiences sudden spikes in traffic. The data engineer notices that write requests are throttled during peak times. The partition keys are well-distributed. What should the data engineer do to reduce throttling?

A.Use DynamoDB global tables to distribute writes across regions.
B.Configure DynamoDB auto scaling to adjust write capacity automatically.
C.Increase the number of partition keys to improve write distribution.
D.Enable DynamoDB Accelerator (DAX) to cache write operations.
AnswerB

Auto scaling adjusts provisioned write capacity in response to consumed capacity, directly addressing the sudden traffic spikes causing throttling. Since partition keys are already well-distributed, the bottleneck is provisioned throughput rather than hot partitions, so scaling write capacity automatically absorbs peak demand without manual intervention.

Why this answer

DynamoDB auto scaling allows the table to automatically adjust its provisioned write capacity based on actual traffic patterns, preventing throttling during sudden spikes without manual intervention. Since the partition keys are already well-distributed, throttling is likely due to insufficient write capacity units, which auto scaling can dynamically increase.

Exam trap

The trap here is that candidates may confuse throttling due to hot partitions (uneven key distribution) with throttling due to insufficient overall capacity, leading them to incorrectly choose option C even when the question explicitly states partition keys are well-distributed.

How to eliminate wrong answers

Option A is wrong because DynamoDB global tables replicate data across regions for disaster recovery and low-latency reads, but they do not increase write capacity within a single region; writes are still subject to the same per-table capacity limits. Option C is wrong because the partition keys are already well-distributed, so adding more partition keys would not resolve throttling caused by insufficient provisioned write capacity; throttling occurs when write requests exceed the table's write capacity units, not due to partition key distribution. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for read operations only, not writes; it cannot reduce write throttling.

20
MCQmedium

A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The migration must have minimal downtime. Which service should be used to perform the migration?

A.AWS Glue
B.AWS DataSync
C.AWS Database Migration Service (DMS)
D.Amazon S3 Transfer Acceleration
AnswerC

AWS Database Migration Service performs continuous change data capture replication from MongoDB to Amazon DocumentDB, keeping the target synchronised until cutover. This satisfies the minimal-downtime constraint, since the source stays live during the initial full load and ongoing replication, unlike offline dump-and-restore approaches that require stopping writes.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it supports continuous replication from MongoDB to Amazon DocumentDB using change data capture (CDC), enabling near-zero downtime migrations. DMS can perform a full load of existing data and then apply ongoing changes from the source MongoDB oplog, keeping the target DocumentDB synchronized until the cutover.

Exam trap

The trap here is that candidates may confuse AWS DMS with AWS DataSync or AWS Glue, assuming any 'migration' or 'data transfer' service can handle live database replication, but only DMS provides the necessary CDC engine for heterogeneous database migrations with minimal downtime.

How to eliminate wrong answers

Option A is wrong because AWS Glue is a serverless data integration service for ETL (extract, transform, load) jobs, not designed for live database migration with minimal downtime; it lacks native CDC support for MongoDB to DocumentDB replication. Option B is wrong because AWS DataSync is optimized for moving large volumes of file data (e.g., NFS, SMB) to AWS storage services like S3 or EFS, not for heterogeneous database migrations or ongoing replication. Option D is wrong because Amazon S3 Transfer Acceleration is a feature that speeds up uploads to S3 buckets over long distances using edge locations; it has no capability to migrate or replicate a MongoDB database to DocumentDB.

21
MCQeasy

A company is storing large amounts of log data in Amazon S3. The data is accessed frequently for the first 30 days, then rarely after that. The company wants to automatically transition the data to a lower-cost storage class after 30 days. Which S3 feature should the data engineer use?

A.S3 Intelligent-Tiering
B.S3 Lifecycle policies
C.S3 Cross-Region Replication
D.S3 Batch Operations
AnswerB

Lifecycle policies can transition objects after a specified number of days.

Why this answer

S3 Lifecycle policies allow you to define rules that automatically transition objects between storage classes based on age or other criteria. In this scenario, a lifecycle rule can be configured to transition objects from S3 Standard to a lower-cost class like S3 Glacier Deep Archive after 30 days, directly meeting the requirement for automated cost optimization.

Exam trap

The trap here is that candidates confuse S3 Intelligent-Tiering's automatic cost optimization with the ability to enforce a fixed time-based transition, when in fact Intelligent-Tiering monitors access patterns and may not align with a strict 30-day policy.

How to eliminate wrong answers

Option A is wrong because S3 Intelligent-Tiering automatically moves data between access tiers based on changing access patterns, but it does not allow you to set a fixed 30-day transition rule; it monitors usage and may not transition data that is rarely accessed after exactly 30 days. Option C is wrong because S3 Cross-Region Replication is used to copy objects to a different AWS region for disaster recovery or compliance, not to transition objects to a lower-cost storage class within the same region. Option D is wrong because S3 Batch Operations is designed for bulk actions like copying, tagging, or restoring objects, not for automating storage class transitions based on time.

22
MCQeasy

A data engineer needs to store large volumes of infrequently accessed compliance data in Amazon S3 for 10 years. The data must be retrievable within 12 hours if required for audits. The engineer wants the most cost-effective storage solution. Which S3 storage class should be used?

A.S3 Glacier Instant Retrieval
B.S3 Standard-Infrequent Access (S3 Standard-IA)
C.S3 Glacier Deep Archive
D.S3 Glacier Flexible Retrieval
AnswerC

S3 Glacier Deep Archive is the lowest-cost storage class in Amazon S3, designed for long-term retention of data that is rarely accessed. It provides standard retrieval within 12 hours, which meets the audit requirement. For compliance data retained for 10 years with infrequent access, it offers the most cost-effective solution.

Why this answer

The scenario requires long-term retention (10 years) with infrequent access and retrieval within 12 hours. S3 Glacier Deep Archive is specifically designed for this use case, offering the lowest storage cost among S3 classes while providing standard retrieval within 12 hours. It is the most cost-effective choice for compliance data that is rarely accessed.

Exam trap

The trap here is assuming that any Glacier class would be equally cost-effective, but S3 Glacier Deep Archive is the cheapest and still meets the 12-hour retrieval requirement.

23
MCQmedium

A data engineer is responsible for an Amazon Redshift cluster that ingests data continuously from Amazon Kinesis Data Streams. The engineer needs to ensure that the raw streaming data is immediately queryable in Redshift with minimal latency. Which approach should the engineer take?

A.Use the Amazon Kinesis Data Firehose delivery stream to load data into Redshift.
B.Use Amazon Redshift Spectrum to query the Kinesis stream directly.
C.Use Amazon Redshift federated query to query Kinesis Data Streams.
D.Use AWS Database Migration Service (AWS DMS) to replicate from Kinesis to Redshift.
AnswerA

Kinesis Data Firehose can deliver streaming data directly to Amazon Redshift, using an intermediate S3 bucket and issuing COPY commands. This provides near-real-time ingestion with minimal latency, making the data immediately queryable. Firehose handles buffering, compression, and retry logic, simplifying the pipeline.

Why this answer

Kinesis Data Firehose is the managed service for loading streaming data into Amazon Redshift. It buffers, compresses, and delivers data to an intermediate S3 bucket, then issues COPY commands to load into Redshift. This provides low-latency, near-real-time ingestion, making the data immediately available for querying.

Other options either do not support Kinesis as a source or are not optimized for this scenario.

Exam trap

The trap here is assuming that Redshift Spectrum or federated query can directly access streaming data, but they are designed for querying external data in S3 or relational databases, not Kinesis streams.

24
MCQeasy

A data engineer is configuring an Amazon S3 lifecycle policy to transition objects to S3 Glacier Deep Archive after 90 days. The bucket receives new objects daily. The engineer wants to ensure that objects are not deleted before 90 days. Which lifecycle action should be used?

A.Expiration
B.Transition
C.NoncurrentVersionTransition
D.AbortIncompleteMultipartUpload
AnswerB

Transition moves objects to another storage class without deleting them, so objects remain retrievable in Glacier Deep Archive after 90 days. Expiration would delete them, and the engineer explicitly wants no deletion before 90 days, making Transition the appropriate lifecycle action.

Why this answer

(Transition) is correct because the S3 Lifecycle Transition action moves objects between storage classes over time. To ensure objects are moved to S3 Glacier Deep Archive after 90 days without deletion, a Transition rule is configured to specify the target storage class and the number of days from object creation.

Exam trap

The trap here is confusing Expiration (which deletes objects) with Transition (which moves objects to another storage class), leading candidates to select Expiration when the goal is to retain objects for a minimum period before moving them to archival storage.

How to eliminate wrong answers

Option A (Expiration) is wrong because it permanently deletes objects after a specified number of days, which would remove them before they could be transitioned to Glacier Deep Archive. Option C (NoncurrentVersionTransition) is wrong because it applies only to noncurrent versions of versioned objects, not to current objects in a non-versioned or versioned bucket. Option D (AbortIncompleteMultipartUpload) is wrong because it only aborts incomplete multipart uploads after a specified number of days, not transitioning or deleting complete objects.

25
MCQeasy

A company needs to store application log files for 90 days for compliance. The logs are generated continuously and are rarely accessed after 30 days. The data engineer must minimize storage costs. Which storage solution should the engineer choose?

A.Amazon CloudWatch Logs with a retention policy of 90 days
B.Amazon S3 Glacier Deep Archive
C.Amazon EBS gp3 volumes attached to an EC2 instance
D.Amazon S3 Standard with a lifecycle policy to transition to S3 Standard-IA after 30 days and expire after 90 days
AnswerD

S3 Standard with a lifecycle policy transitioning to Standard-IA at 30 days and expiring at 90 days matches the access pattern: frequent early access, rare later access, and deletion at the compliance boundary. This minimises cost while meeting the 90-day retention requirement.

Why this answer

Amazon S3 Standard with a lifecycle policy to transition to S3 Standard-IA after 30 days and expire after 90 days is correct because it aligns with the access pattern: logs are frequently accessed only in the first 30 days, then rarely accessed for the remaining 60 days. S3 Standard-IA offers lower storage costs for infrequently accessed data while still providing millisecond retrieval, and the lifecycle policy automates the transition and eventual deletion, minimizing costs without sacrificing availability.

Exam trap

The trap here is that candidates often choose CloudWatch Logs (Option A) because it is a familiar logging service, but they overlook that its cost model (per GB ingested, per GB stored, and per GB archived) can be significantly higher than S3 for long-term retention of large log volumes, and it lacks the automated tiering to lower-cost storage classes.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch Logs is designed for real-time monitoring and log ingestion, not for long-term, cost-optimized archival storage; its retention policy only controls deletion, not tiered storage transitions, and costs can be higher than S3 for large volumes of rarely accessed logs. Option B is wrong because S3 Glacier Deep Archive is intended for data that is accessed at most once or twice a year and has retrieval times of 12 hours or more, making it unsuitable for logs that may need occasional access within 90 days; it also incurs minimum storage charges that make it cost-ineffective for short retention periods. Option C is wrong because EBS gp3 volumes attached to an EC2 instance incur compute costs even when idle, and managing log storage on block storage requires manual lifecycle management, leading to higher operational overhead and cost compared to a fully managed object storage solution.

26
Multi-Selectmedium

Which TWO actions are recommended for securing data at rest in Amazon S3? (Choose two.)

Select 2 answers
A.Enable default encryption on the S3 bucket using SSE-S3 or SSE-KMS.
B.Use S3 Bucket Key to reduce KMS request costs.
C.Enable S3 Versioning to protect against accidental deletions.
D.Apply a bucket policy that denies PutObject requests without the x-amz-server-side-encryption header.
E.Configure cross-region replication to replicate data to another bucket.
AnswersA, D

Enabling default encryption ensures every object written to the bucket is encrypted at rest automatically, satisfying the data-at-rest requirement without relying on per-request headers. SSE-S3 provides AES-256 with Amazon-managed keys, while SSE-KMS adds customer-managed key control and audit trails through CloudTrail, meeting compliance needs.

Why this answer

Option A is correct because enabling default bucket encryption with SSE-S3 or SSE-KMS ensures that every object written to the bucket is automatically encrypted at rest, satisfying the core requirement for data-at-rest protection in Amazon S3. Option D is correct because a bucket policy that denies PutObject requests lacking the x-amz-server-side-encryption header enforces encryption on upload, preventing unencrypted objects from being stored even if default encryption is bypassed or misconfigured. Option B is not a security control; S3 Bucket Key reduces AWS KMS request costs and CloudTrail event volume but does not itself secure data at rest.

Option C addresses availability and data durability against accidental deletion through versioning, not encryption at rest. Option E provides geographic redundancy and disaster recovery via cross-region replication, but replication does not encrypt data at rest by itself.

Exam trap

The trap here is that candidates often confuse data protection features like Versioning or replication with encryption controls, but the question specifically asks for securing data at rest, which requires encryption mechanisms such as default encryption or policy-enforced encryption headers.

27
MCQmedium

A company runs a data warehouse on Amazon Redshift. Queries are slow, and the team suspects data distribution is skewed. Which approach would best help identify distribution skew?

A.Check the STL_LOAD_ERRORS table for load failures
B.Query the SVV_TABLE_INFO table to see table size
C.Query the SVV_DISKUSAGE table to examine data distribution across slices
D.Review the WLM configuration in the parameter group
AnswerC

SVV_DISKUSAGE reports per-slice row counts and disk usage, directly exposing uneven distribution across slices that causes skew. Unlike planner-oriented views such as SVV_TABLE_INFO, it shows physical storage per slice, satisfying the need to confirm skew empirically before choosing a distribution key.

Why this answer

The SVV_DISKUSAGE table provides per-slice data distribution information, allowing you to identify skew by comparing the number of blocks allocated to each slice for a given table. In Amazon Redshift, data is distributed across slices based on the distribution key, and significant variation in block counts across slices indicates distribution skew, which can cause query performance degradation due to uneven workload distribution.

Exam trap

The trap here is that candidates confuse table-level metadata (SVV_TABLE_INFO) with slice-level distribution data (SVV_DISKUSAGE), assuming overall table size alone can reveal skew, when in fact only per-slice block counts expose uneven data distribution.

How to eliminate wrong answers

Option A is wrong because STL_LOAD_ERRORS records errors during COPY or INSERT operations, such as data type mismatches or malformed data, and has no relation to data distribution skew. Option B is wrong because SVV_TABLE_INFO shows overall table size, row count, and compression ratios, but it does not provide per-slice data distribution details needed to identify skew. Option D is wrong because WLM configuration in the parameter group manages query concurrency and memory allocation, not data distribution or skew detection.

28
MCQhard

A data engineer is troubleshooting an Amazon DynamoDB table that has frequent throttling exceptions for write requests. The table has auto scaling enabled. What is the most likely cause?

A.The partition key is causing a hot partition
B.The table's read capacity is set too low
C.The table's auto scaling is disabled
D.The table is using global tables without conflict resolution
AnswerA

Auto scaling adjusts capacity for uniform load, but a hot partition concentrates writes on one partition key value, exceeding that partition's throughput ceiling regardless of table-level capacity. This is the most likely cause of persistent write throttling.

Why this answer

Auto scaling adjusts capacity based on utilization, but it cannot prevent throttling caused by a hot partition. If a single partition key value receives a disproportionate share of write traffic, that partition's throughput limit (3,000 WCU or 10 MB per partition) is exceeded, triggering ProvisionedThroughputExceededException. Auto scaling operates at the table level, not per partition, so it cannot resolve this imbalance.

Exam trap

The trap here is that candidates assume auto scaling automatically prevents all throttling, but it only adjusts table-level capacity and cannot fix uneven data access patterns like a hot partition.

How to eliminate wrong answers

Option B is wrong because write throttling is unrelated to read capacity; the question specifies write request throttling, so read capacity settings are irrelevant. Option C is wrong because the question states auto scaling is enabled, so this option describes a scenario that does not match the given condition. Option D is wrong because global tables with conflict resolution handle eventual consistency and replication conflicts, not throughput throttling on write requests.

29
MCQmedium

A company is running a data warehouse on Amazon Redshift. The data engineering team notices that query performance has degraded over time. They suspect that data distribution is causing excessive data movement between nodes. The table is joined frequently on the customer_id column. Which column should be chosen as the distribution key to optimize join performance?

A.AUTO distribution
B.customer_id
C.order_date
D.EVEN distribution
AnswerB

Choosing customer_id as the distribution key colocates rows with identical customer_id values on the same node, so frequent joins on that column execute locally without network redistribution. This directly addresses the excessive data movement degrading query performance.

Why this answer

(customer_id) because Redshift distributes data across nodes based on the distribution key. When two tables are joined on customer_id, using it as the distribution key ensures that matching rows from both tables are co-located on the same node, eliminating the need for data redistribution (broadcast or shuffle) during the join. This minimizes network traffic and reduces query latency, directly addressing the performance degradation caused by excessive data movement.

Exam trap

The trap here is that candidates may choose EVEN distribution (D) thinking it balances data evenly, but they overlook that it causes maximum data movement for joins, while AUTO distribution (A) seems safe but does not guarantee co-location for the specific join column.

How to eliminate wrong answers

Option A (AUTO distribution) is wrong because AUTO lets Redshift choose the distribution style based on table size and usage patterns, but it may not guarantee co-location for frequent joins on customer_id, potentially still causing data movement. Option C (order_date) is wrong because it is not the join column; using it as the distribution key would scatter customer_id values across nodes, forcing redistribution for every join on customer_id. Option D (EVEN distribution) is wrong because it distributes rows round-robin across nodes without considering join keys, which maximizes data movement during joins on customer_id and degrades performance.

30
MCQhard

A data engineer created the IAM policy shown in the exhibit. The engineer then attempts to upload an object to 'my-bucket' using the AWS CLI with the command: aws s3 cp file.txt s3://my-bucket/ --sse aws:kms. The upload fails with an 'AccessDenied' error. What is the most likely cause?

A.The policy resource is incorrect
B.The policy requires SSE-S3 (AES256), but the command uses SSE-KMS
C.The policy does not allow the s3:PutObject action
D.The command is missing the --sse-customer-algorithm parameter
AnswerB

The policy's condition permits only AES256 (SSE-S3) encryption, but the CLI command requests aws:kms, so the encryption header mismatches the allowed value and S3 denies the upload. Aligning the command with SSE-S3 or amending the policy resolves it.

Why this answer

The IAM policy in the exhibit requires the `s3:x-amz-server-side-encryption` header to be set to `AES256`, which corresponds to SSE-S3. The AWS CLI command uses `--sse aws:kms`, which sets the header to `aws:kms` for SSE-KMS. This mismatch causes the request to fail the `s3:PutObject` condition check in the policy, resulting in an 'AccessDenied' error.

Exam trap

The trap here is that candidates may overlook the condition key in the policy and assume the error is due to a missing action or incorrect resource, rather than recognizing that the encryption header value must exactly match the policy's requirement.

How to eliminate wrong answers

Option A is wrong because the policy resource `arn:aws:s3:::my-bucket/*` correctly specifies the bucket and its objects, so the resource is not the issue. Option B is wrong because the policy explicitly requires SSE-S3 (AES256), but the command uses SSE-KMS, which is the direct cause of the failure. Option C is wrong because the policy does allow `s3:PutObject` via the `Effect: Allow` statement; the failure is due to the condition key mismatch, not a missing action.

Option D is wrong because `--sse-customer-algorithm` is used for SSE-C, not SSE-KMS or SSE-S3, and the command already specifies `--sse aws:kms` correctly for SSE-KMS.

31
MCQeasy

A data engineer is designing a data lake on AWS using Amazon S3. The data consists of CSV files generated by IoT devices. The data is accessed by multiple analytics jobs, and the engineer needs to ensure that new files are immediately visible to all consumers after writing. What S3 consistency model applies?

A.Consistent reads require S3 Object Lock.
B.Strong consistency for all operations.
C.Eventual consistency for all operations.
D.Read-after-write consistency for new object PUTS.
AnswerB

Amazon S3 now provides strong read-after-write consistency for all PUT and DELETE operations across all Regions, so newly written IoT CSV objects are immediately visible to every analytics job. This satisfies the stem's requirement that new files be instantly visible to all consumers.

Why this answer

Amazon S3 now provides strong consistency for all operations. After a successful write of a new object (or overwrite of an existing object), any subsequent read request immediately receives the latest version of the object, and list operations are also strongly consistent. Therefore, for new CSV files written to S3, the applicable model is strong consistency for all operations.

Exam trap

Candidates may incorrectly choose 'Read-after-write consistency for new object PUTS' (option D) because the question mentions new files after writing. While read-after-write behavior for new object PUTs is true, the current S3 consistency model is strong consistency for all operations. Another pitfall is selecting 'Eventual consistency for all operations' (option C) due to outdated knowledge of S3's previous eventual consistency model.

How to eliminate wrong answers

Option A is wrong because S3 Object Lock is a feature for preventing object deletion or overwrites for compliance or retention, not for ensuring consistency. Option B is wrong because while S3 now offers strong consistency for all operations (including overwrites and deletes), this was not always the case; historically S3 offered eventual consistency for overwrites, and the question's phrasing about 'new files' specifically tests the read-after-write consistency model for new PUTS. Option C is wrong because S3 no longer provides eventual consistency for new object PUTS; it guarantees strong read-after-write consistency for new objects since December 2020.

32
MCQmedium

A data engineer reviewed the S3 lifecycle policy shown in the exhibit. The engineer notices that objects under the 'logs/' prefix are being deleted after 365 days. The business requirement is to retain logs for at least 5 years. What should the engineer change in the lifecycle policy?

A.Change the prefix to 'logs/archive/'
B.Set the expiration days to 1825
C.Change the transition to GLACIER on day 365
D.Remove the expiration action
AnswerB

The lifecycle rule expires objects after 365 days, but the business needs five years of retention. Setting expiration to 1825 days (5 × 365) aligns deletion with the retention requirement, correcting the premature removal of logs under the 'logs/' prefix.

Why this answer

The business requirement is to retain logs for at least 5 years, which is 1,825 days (5 × 365). The current lifecycle policy sets expiration to 365 days, causing premature deletion. By setting the expiration days to 1,825, the S3 lifecycle policy will delete objects under the 'logs/' prefix only after 5 years, meeting the retention requirement.

Exam trap

The trap here is that candidates may confuse transition actions (which change storage class) with expiration actions (which delete objects), or incorrectly assume that changing the prefix or removing expiration will meet the retention requirement without adjusting the day count.

How to eliminate wrong answers

Option A is wrong because changing the prefix to 'logs/archive/' would only apply the lifecycle rules to a different subset of objects, not fix the retention period for the original 'logs/' prefix. Option C is wrong because transitioning to GLACIER on day 365 only changes the storage class for cost optimization; it does not extend the deletion timeline, so objects would still be deleted after 365 days. Option D is wrong because removing the expiration action entirely would mean objects are never automatically deleted, which may lead to indefinite storage and increased costs, not a 5-year retention.

33
MCQmedium

A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a nested folder structure. The engineer wants to automatically discover the schema and update the AWS Glue Data Catalog as new files are added. Which solution should the engineer use?

A.Run an AWS Glue ETL job that reads the CSV files and writes to a new location.
B.Create an AWS Glue crawler that points to the S3 bucket and schedule it to run periodically.
C.Use AWS Lake Formation to define a data lake and register the S3 location.
D.Manually define tables in the AWS Glue Data Catalog using the AWS Management Console.
AnswerB

AWS Glue crawlers automatically scan data sources, infer schemas, and populate the Data Catalog. They can be scheduled to run periodically to detect new files and update table definitions. This meets the requirement for automatic schema discovery and catalog updates with minimal manual effort. The crawler can handle nested folder structures and CSV format.

Why this answer

AWS Glue crawlers are designed to automatically scan data stores, infer schemas, and create or update tables in the Data Catalog. Scheduling a crawler ensures that new files are detected and the catalog stays current. This is the standard, low-effort solution for the described scenario.

Other options either require manual intervention or do not provide automatic schema discovery.

Exam trap

The trap here is confusing AWS Lake Formation with AWS Glue crawlers; Lake Formation manages permissions and governance but does not automatically discover schemas.

34
Multi-Selecteasy

A company is building a data pipeline that ingests streaming data from IoT devices. The data must be stored in a durable, scalable, and cost-effective manner for batch processing. Which TWO AWS services should be used together?

Select 2 answers
A.Amazon ElastiCache
B.Amazon Kinesis Data Streams
C.Amazon Redshift
D.Amazon DynamoDB
E.Amazon S3
AnswersB, E

Amazon Kinesis Data Streams ingests the IoT telemetry durably and scales by shard, satisfying the streaming ingestion requirement. It decouples producers from consumers, buffering records for the companion storage service that handles batch processing, so ingestion spikes never overwhelm downstream analytics.

Why this answer

Amazon Kinesis Data Streams (B) is the correct ingestion service for streaming IoT data because it provides a durable, scalable, and real-time data streaming platform that can capture and store data records for up to 365 days. Amazon S3 (E) is the correct storage service for batch processing because it offers virtually unlimited durability (99.999999999%), cost-effective tiered storage, and native integration with batch processing frameworks like Amazon EMR and AWS Glue. Together, they form a classic streaming-to-batch pipeline: Kinesis ingests and buffers the streaming data, which is then persisted in S3 for downstream batch analytics.

Exam trap

The trap here is that candidates often confuse Amazon Kinesis Data Streams with Amazon Kinesis Data Firehose (which directly writes to S3) or mistakenly choose Amazon Redshift for storage, overlooking that S3 is the correct durable and cost-effective storage layer for raw streaming data before any warehousing.

35
MCQmedium

A company is using Amazon S3 to store large amounts of archival data. The data is accessed infrequently but must be immediately retrievable when needed. Which storage class is the most cost-effective choice?

A.S3 Standard
B.S3 Standard-IA
C.S3 Glacier Deep Archive
D.S3 Intelligent-Tiering
AnswerB

S3 Standard-IA matches the stem's dual constraint: infrequent access plus immediate retrieval. Its lower per-GB storage price than S3 Standard suits archival data, while millisecond retrieval preserves instant availability. Glacier classes would fail the immediate-retrieval requirement despite lower cost.

Why this answer

S3 Standard-IA (Infrequent Access) is the most cost-effective choice because it offers lower storage costs than S3 Standard while still providing millisecond first-byte latency for immediate retrieval. The data is accessed infrequently but requires instant availability, which matches the IA use case exactly.

Exam trap

The trap here is that candidates confuse 'immediately retrievable' with 'lowest cost' and choose Glacier Deep Archive, overlooking the critical requirement for instant access versus the 12-48 hour retrieval time of Deep Archive.

How to eliminate wrong answers

Option A is wrong because S3 Standard is designed for frequently accessed data and has higher storage costs than Standard-IA, making it less cost-effective for archival data with infrequent access. Option C is wrong because S3 Glacier Deep Archive has the lowest storage cost but retrieval times range from 12 to 48 hours, failing the 'immediately retrievable' requirement. Option D is wrong because S3 Intelligent-Tiering automatically moves data between tiers based on access patterns but incurs a monthly monitoring and automation fee per object, making it less cost-effective than Standard-IA for a predictable infrequent access pattern.

36
Multi-Selectmedium

A company is designing a data lake on Amazon S3. Which TWO strategies improve query performance for Amazon Athena?

Select 2 answers
A.Enable S3 Versioning on the bucket.
B.Use server-side encryption with AWS KMS (SSE-KMS).
C.Partition the data by frequently queried columns such as date or region.
D.Use columnar file formats like Parquet or ORC.
E.Store data in CSV format with header rows.
AnswersC, D

Partitioning by frequently filtered columns lets Athena prune entire S3 prefixes using partition metadata, so it scans only relevant data rather than the whole table. This directly cuts bytes scanned and query runtime for date- or region-filtered workloads.

Why this answer

Option C is correct because partitioning the S3 data by frequently filtered columns such as date or region lets Athena prune partitions and scan only the relevant prefixes, drastically reducing the amount of data read and thus improving query speed and lowering cost. Option D is correct because columnar formats like Parquet or ORC allow Athena to read only the columns referenced in the query and benefit from compression and predicate pushdown, minimizing I/O compared to row-based formats. Options A and B do not improve query performance: S3 Versioning only retains object versions for recovery, and SSE-KMS only encrypts data at rest, adding no scan optimization (and KMS calls can even add latency).

Option E is incorrect because CSV is a row-based, uncompressed text format that forces Athena to scan entire rows and cannot skip unneeded columns, making queries slower and more expensive than Parquet/ORC.

Exam trap

The trap here is that candidates often confuse data management features (like versioning or encryption) with performance optimizations, or assume that simpler formats like CSV are sufficient for analytics, ignoring the significant performance benefits of partitioning and columnar storage.

37
Multi-Selecthard

A company is using Amazon Redshift for its data warehouse. The data engineering team needs to improve query performance for a large fact table that is frequently joined with multiple dimension tables. Which THREE strategies should be considered?

Select 3 answers
A.Define sort keys on columns used in WHERE clauses.
B.Use DISTSTYLE EVEN to distribute data evenly.
C.Increase the number of nodes in the cluster.
D.Choose an appropriate distribution key based on join columns.
E.Apply columnar compression to reduce storage and I/O.
AnswersA, D, E

Sort keys physically order rows on disk by the chosen columns, so range and equality predicates in WHERE clauses skip irrelevant blocks via zone maps. This reduces the rows scanned for the large fact table, directly improving the filtered queries feeding the joins.

Why this answer

Option A is correct because defining sort keys on columns frequently used in WHERE clauses allows Redshift to use zone maps and block-level metadata to skip irrelevant blocks, dramatically reducing I/O for filtered queries on the large fact table. Option D is correct because choosing a distribution key based on the fact table's join columns co-locates matching rows on the same node slice, enabling collocated joins that avoid costly data redistribution (broadcast or shuffle) across the cluster. Option E is correct because columnar compression reduces the amount of data read from disk and improves I/O efficiency, which is especially beneficial for large fact tables scanned by analytical queries.

Option B is not ideal here because DISTSTYLE EVEN spreads rows uniformly but does not co-locate join keys, so joins with dimension tables require network redistribution, hurting performance for frequent joins. Option C is not a targeted strategy for join performance; adding nodes increases compute and storage capacity but does not by itself address sort keys, distribution keys, or compression, and may not resolve join-related bottlenecks.

Exam trap

The trap here is that candidates often assume DISTSTYLE EVEN is always the best choice for performance, but for frequently joined fact tables, a distribution key aligned with the join columns is critical to avoid network-heavy data shuffling.

38
Multi-Selecthard

A company stores sensitive data in Amazon S3. The security team requires encryption at rest and that the encryption keys are managed by the company using AWS KMS. The data is frequently accessed by multiple AWS services. Which THREE steps should be taken to meet these requirements?

Select 3 answers
A.Use client-side encryption with the KMS key before uploading to S3
B.Configure the KMS key policy to allow the necessary AWS services to use the key for decryption
C.Enable default encryption on the S3 bucket using SSE-S3
D.Create a bucket policy that denies s3:PutObject if the object is not encrypted with SSE-KMS
E.Enable default encryption on the S3 bucket using SSE-KMS
AnswersB, D, E

Services must have decrypt permissions to access the encrypted objects.

Why this answer

The security team requires that encryption keys be managed by the company using AWS KMS, and that multiple AWS services can access the data. To allow those services to decrypt objects encrypted with a customer-managed KMS key, the KMS key policy must explicitly grant the necessary AWS services (e.g., AWS Lambda, Amazon Athena) permission to use the key for decryption (kms:Decrypt). Without this policy, even if the bucket is configured for SSE-KMS, the services will fail to read the encrypted objects.

Exam trap

AWS often tests the distinction between enforcing encryption (bucket policy) and enabling access to encrypted data (KMS key policy), leading candidates to overlook the KMS key policy step when multiple services need to decrypt objects.

39
MCQeasy

A company uses Amazon RDS for PostgreSQL. The data engineer needs to ensure that the database is automatically backed up and that backups are retained for 35 days. What is the simplest way to achieve this?

A.Use AWS Backup to schedule daily backups with a 35-day retention.
B.Enable automated backups with a retention period of 35 days in the RDS instance configuration.
C.Create a manual snapshot every day and delete them after 35 days using a script.
D.Enable automatic export of transaction logs to Amazon S3 and use S3 lifecycle policies.
AnswerB

Configuring automated backups on the RDS instance with a 35-day retention period uses RDS's native backup mechanism, which takes daily snapshots and retains them for the specified window. This is the simplest approach, requiring no custom scripting or external tooling.

Why this answer

Amazon RDS for PostgreSQL allows you to enable automated backups directly in the instance configuration. By setting the backup retention period to 35 days, RDS automatically performs daily snapshots and retains transaction logs for point-in-time recovery within that window. This is the simplest method because it requires no external services or custom scripting.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing AWS Backup (Option A) or manual scripting (Option C), not realizing that RDS native automated backups already provide the simplest, fully managed way to achieve the required retention period.

How to eliminate wrong answers

Option A is wrong because AWS Backup is an additional service that adds complexity and cost; RDS native automated backups already support retention up to 35 days without needing AWS Backup. Option C is wrong because manual snapshots require custom scripting to create and delete daily, which is not the simplest approach and does not provide automated point-in-time recovery. Option D is wrong because automatic export of transaction logs to S3 is not a native RDS feature for PostgreSQL; RDS handles transaction logs internally for point-in-time recovery, and using S3 lifecycle policies would not replace the need for automated backups.

40
MCQmedium

A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow due to many small inserts. Which technique would improve write performance?

A.Use the COPY command to load data from Amazon S3.
B.Define DISTKEY and SORTKEY on the table.
C.Increase the number of nodes in the cluster.
D.Configure workload management (WLM) queues.
AnswerA

Row-by-row inserts force Redshift to commit many small transactions, which is inefficient. COPY loads large batches in parallel from Amazon S3 directly into the cluster, dramatically improving write throughput and reducing the overhead of many small inserts.

Why this answer

The COPY command is the recommended way to load data into Amazon Redshift because it performs bulk inserts in parallel across all nodes, leveraging the cluster's distributed architecture. Small individual INSERT statements cause high overhead due to transaction logging and commit processing, leading to slow write performance. By loading data from Amazon S3 using COPY, you bypass these per-row overheads and achieve optimal throughput.

Exam trap

The trap here is that candidates often confuse performance tuning for reads (DISTKEY/SORTKEY) or general scaling (adding nodes) with the specific write performance bottleneck caused by many small inserts, overlooking the COPY command as the primary solution for bulk data loading.

How to eliminate wrong answers

Option B is wrong because defining DISTKEY and SORTKEY improves query read performance by optimizing data distribution and sort order, but does not directly address the write performance issue caused by many small inserts. Option C is wrong because increasing the number of nodes adds compute and storage capacity, but does not solve the fundamental problem of per-insert overhead; small inserts will still be slow on a larger cluster. Option D is wrong because configuring workload management (WLM) queues manages concurrency and prioritizes queries, but does not reduce the overhead of individual small INSERT statements.

41
MCQeasy

A company uses Amazon DynamoDB as its primary data store for a web application. The application experiences high latency during peak hours. The data engineer notices that the table has a large number of items with the same partition key. Which DynamoDB feature should the engineer use to improve performance?

A.Redesign the partition key to use a composite key that includes a timestamp or random suffix.
B.Enable DynamoDB Accelerator (DAX) to cache read requests.
C.Create a global table to replicate data across multiple Regions.
D.Enable auto scaling on the table to increase write capacity.
AnswerA

Many items sharing one partition key concentrate reads on a single partition, creating a hot partition that throttles throughput and raises latency. A composite key adding a timestamp or random suffix spreads items across partitions, distributing load and restoring performance.

Why this answer

The high latency is caused by a hot partition, where many items share the same partition key, overwhelming a single DynamoDB partition. Redesigning the partition key to include a timestamp or random suffix distributes the workload evenly across partitions, improving throughput and reducing latency. This directly addresses the root cause of the performance issue.

Exam trap

The trap here is that candidates often confuse caching solutions (DAX) or scaling mechanisms (auto scaling) with the need to fix the data model itself, which is the only way to resolve a hot partition caused by a skewed partition key.

How to eliminate wrong answers

Option B is wrong because DynamoDB Accelerator (DAX) caches read requests, which can reduce read latency but does not solve the underlying hot partition issue caused by skewed write or read traffic on a single partition key. Option C is wrong because creating a global table replicates data across multiple Regions for disaster recovery or low-latency global access, but it does not distribute load within a single table's partitions. Option D is wrong because enabling auto scaling increases the table's provisioned capacity, but if the workload is concentrated on one partition, the partition's throughput limit (3000 RCU or 1000 WCU) will still be exceeded, causing throttling and high latency.

42
MCQeasy

A company needs to store JSON documents that are accessed by a key-value pattern. The data is 500 GB and requires single-digit millisecond latency. Which AWS database is most suitable?

A.Amazon Redshift
B.Amazon DynamoDB
C.Amazon Neptune
D.Amazon RDS for MySQL
AnswerB

DynamoDB is a fully managed key-value and document store delivering consistent single-digit millisecond latency at any scale, comfortably handling 500 GB of JSON documents accessed by primary key. Relational or object stores cannot meet that latency and access pattern.

Why this answer

Amazon DynamoDB is the most suitable choice because it is a fully managed NoSQL key-value and document database that delivers single-digit millisecond latency at any scale, making it ideal for storing and retrieving JSON documents via a key-value access pattern. It supports document data types natively and can handle 500 GB of data efficiently with consistent low-latency performance.

Exam trap

The trap here is that candidates may choose Amazon RDS for MySQL because they associate JSON documents with relational databases, overlooking that DynamoDB is purpose-built for key-value and document workloads with guaranteed single-digit millisecond latency, while RDS introduces schema rigidity and higher latency for this pattern.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a petabyte-scale data warehouse optimized for complex analytical queries using SQL, not for low-latency key-value lookups on JSON documents; it incurs higher latency and is not designed for single-digit millisecond access patterns. Option C is wrong because Amazon Neptune is a graph database optimized for highly connected data and graph queries (e.g., using Gremlin or SPARQL), not for simple key-value access to JSON documents; it would add unnecessary complexity and cost. Option D is wrong because Amazon RDS for MySQL is a relational database that requires predefined schemas and is not optimized for key-value access patterns on JSON documents; while it can store JSON, it lacks the native partitioning and low-latency throughput of DynamoDB for this use case.

43
Multi-Selecteasy

Which TWO methods can be used to enforce least-privilege access to an Amazon S3 bucket? (Choose two.)

Select 2 answers
A.Use IAM policies to grant specific permissions to users and roles.
B.Set bucket ACLs to allow full control to the bucket owner only.
C.Use an S3 bucket policy that explicitly denies actions not required.
D.Configure a VPC endpoint to restrict access to the bucket.
E.Generate pre-signed URLs for all access.
AnswersA, C

IAM policies allow granular permissions.

Why this answer

IAM policies allow you to grant granular, specific permissions to individual users and roles, adhering to the principle of least privilege by explicitly allowing only the actions required. This avoids granting broad or default permissions, ensuring that each identity has only the access necessary for its function.

Exam trap

The trap here is that candidates often confuse network-level controls (like VPC endpoints) with identity-based access controls, or they mistakenly think that granting full control to the owner is a form of least privilege, when in fact it violates the principle by providing excessive permissions.

44
Multi-Selectmedium

Which TWO actions can help improve query performance in Amazon Redshift? (Choose two.)

Select 2 answers
A.Use appropriate sort keys for tables.
B.Disable SSL encryption for connections.
C.Use VARCHAR instead of CHAR for fixed-length strings.
D.Apply compression encodings to columns.
E.Increase the number of nodes in the cluster.
AnswersA, D

Sort keys help the query optimizer scan less data.

Why this answer

Defining appropriate sort keys in Amazon Redshift enables the query optimizer to use zone maps to skip irrelevant data blocks during table scans, significantly reducing the amount of data read from disk. Sort keys also improve the effectiveness of merge joins and the performance of range-restricted queries by physically co-locating rows with similar sort key values on disk.

Exam trap

The trap here is that candidates often assume scaling out (adding nodes) always speeds up individual queries, but in Redshift, query performance is more dependent on data layout (sort keys, distribution, compression) than on cluster size, and adding nodes primarily benefits concurrent workloads rather than single-query latency.

45
MCQhard

A data engineering team is designing a data lake on Amazon S3. They need to store raw data in its original format and transformed data in Parquet. The data is accessed by multiple analytics services, including Amazon Athena and Amazon Redshift Spectrum. Compliance requirements mandate that all data be encrypted at rest with AWS KMS and that the encryption keys be rotated every 90 days. Which S3 bucket configuration meets these requirements?

A.Use SSE-KMS with a customer-managed KMS key that has automatic key rotation enabled.
B.Use SSE-C with client-managed keys and rotate them manually.
C.Use a bucket policy to enforce encryption and rely on default S3 encryption.
D.Use SSE-S3 with default encryption enabled.
AnswerA

SSE-KMS with automatic rotation meets compliance requirements.

Why this answer

SSE-KMS with a customer-managed KMS key allows you to implement custom key rotation, such as every 90 days, by creating new keys and updating the bucket policy or key alias. AWS KMS automatic key rotation for customer-managed keys occurs yearly, not every 90 days, but you can achieve a 90-day rotation schedule manually or through automation (e.g., AWS Lambda). SSE-C requires manual key management and does not integrate with AWS services like Amazon Athena.

SSE-S3 does not support configurable rotation, and the default encryption option (C) does not meet compliance if rotation is required.

Exam trap

The trap is that candidates assume SSE-S3 or default encryption meets the 90-day rotation requirement because AWS rotates keys automatically, but they overlook that SSE-S3 key rotation is not configurable—it follows AWS-managed rotation, which is not guaranteed every 90 days. SSE-KMS with a customer-managed key is the only option that allows a custom rotation schedule, even though it may require additional automation beyond the automatic yearly rotation.

How to eliminate wrong answers

Option B is wrong because SSE-C requires you to manage and rotate encryption keys client-side, which adds operational overhead and does not integrate with AWS KMS for automated rotation; manual rotation every 90 days is possible but not automated, and it violates the requirement to use AWS KMS. Option C is wrong because relying on default S3 encryption (SSE-S3) uses S3-managed keys that cannot be rotated on a 90-day schedule; AWS rotates SSE-S3 keys annually, but you have no control over the rotation frequency. Option D is wrong because SSE-S3 does not support customer-controlled key rotation; it uses S3-managed keys with automatic rotation by AWS, but the rotation period is not configurable and does not meet the 90-day requirement.

46
MCQmedium

A data engineer is designing a data lake on Amazon S3. The team wants to optimize query performance and reduce storage costs for a large dataset of JSON logs that are queried frequently by Amazon Athena. The logs are currently stored as uncompressed JSON files, each around 1 GB, in a single prefix. The engineer needs to improve query performance and reduce costs without changing the data format. Which action should the engineer take?

A.Compress the JSON files with gzip and split them into smaller files of approximately 128 MB.
B.Convert the JSON files to Apache Parquet and partition the data by date.
C.Move the data to Amazon Redshift and query it using Redshift Spectrum.
D.Enable S3 Transfer Acceleration on the bucket to speed up data retrieval.
AnswerA

Compressing with gzip reduces storage costs and the amount of data scanned by Athena. Splitting large files into smaller ones (around 128 MB) enables parallel processing and improves query performance. This approach maintains the JSON format and directly addresses the requirements without altering the data structure.

Why this answer

Compressing the JSON files with gzip reduces the storage footprint and the amount of data scanned by Athena, lowering costs. Splitting the large files into smaller, more manageable sizes allows Athena to process them in parallel, significantly improving query performance. This solution respects the requirement to keep the data in JSON format and directly addresses both performance and cost concerns.

Exam trap

The trap here is assuming that converting to a columnar format like Parquet is always the best optimization, but the scenario explicitly requires keeping the data in its original JSON format.

47
MCQhard

A company runs a real-time analytics platform on Amazon ECS that ingests streaming data from Amazon Kinesis Data Streams, processes it, and stores results in Amazon DynamoDB. The data volume spikes unpredictably, causing DynamoDB to throttle write requests. The application uses on-demand capacity mode. The data engineer notices that the throttling occurs on a specific partition due to a hot key. The hot key is a customer ID that receives a disproportionate number of writes. The application cannot change the partition key design immediately. The engineer needs to reduce throttling while maintaining low latency. Which solution is most effective?

A.Switch to provisioned capacity with auto scaling and increase the write capacity units.
B.Implement a write buffer using Amazon SQS, and have consumers write to DynamoDB at a controlled rate.
C.Enable DynamoDB Accelerator (DAX) to cache the hot key writes.
D.Use DynamoDB Streams to trigger a Lambda function that retries throttled writes.
AnswerB

An SQS write buffer decouples ingestion from DynamoDB, letting consumers write at a controlled rate so the hot partition key is no longer overwhelmed. This absorbs unpredictable spikes while preserving low latency, and works without redesigning the partition key, which the stem forbids.

Why this answer

Buffering writes through Amazon SQS decouples the ingestion rate from DynamoDB's capacity, allowing consumers to write at a controlled pace. This directly mitigates throttling on the hot key without requiring a partition key redesign, and SQS provides low-latency, durable buffering suitable for real-time analytics.

Exam trap

The trap here is that candidates often assume on-demand capacity eliminates all throttling, but it does not protect against hot key skew; they may also confuse DAX's read caching with write buffering, or think retrying throttled writes is a viable solution rather than a reactive fix that increases latency.

How to eliminate wrong answers

Option A is wrong because switching to provisioned capacity with auto scaling does not solve the hot key issue; throttling occurs on a specific partition regardless of total capacity, and increasing write capacity units would not prevent a single partition from exceeding its 1,000 WCU limit. Option C is wrong because DAX is a caching layer for reads, not writes; it cannot buffer or absorb write throttling on a hot key. Option D is wrong because using DynamoDB Streams to retry throttled writes introduces latency and does not prevent throttling; it only retries failed writes, which can lead to backlog and increased latency, not a controlled rate.

48
MCQeasy

A company uses Amazon DynamoDB as the primary data store for a web application. The application experiences occasional throttling on write requests. The data engineer needs to implement a solution that handles throttling gracefully without losing data. Which approach should the engineer use?

A.Increase the provisioned write capacity to a higher value
B.Use an Amazon SQS queue to buffer write requests before sending to DynamoDB
C.Implement exponential backoff in the application's write retry logic
D.Enable DynamoDB Accelerator (DAX) to cache writes
AnswerC

Exponential backoff retries throttled writes with progressively longer, randomised delays, absorbing transient capacity bursts without dropping requests. DynamoDB returns ProvisionedThroughputExceededException for throttled writes; retrying with backoff lets the request succeed once capacity frees, satisfying the no-data-loss constraint. It handles throttling gracefully rather than preventing it.

Why this answer

Implementing exponential backoff in the application's write retry logic is the standard AWS-recommended approach for handling DynamoDB throttling (ProvisionedThroughputExceededException). Exponential backoff gradually increases the wait time between retries, reducing the retry rate and allowing the throttling condition to subside, while ensuring no write data is lost as long as the retries eventually succeed. This approach is lightweight, requires no additional AWS services, and aligns with best practices for building resilient applications against DynamoDB throttling.

Exam trap

The trap here is that candidates often confuse DAX as a write cache or assume SQS is the only way to buffer writes, but the question specifically asks for handling throttling gracefully without losing data, and exponential backoff is the direct, built-in mechanism for retrying throttled requests in DynamoDB.

How to eliminate wrong answers

Option A is wrong because simply increasing provisioned write capacity may reduce throttling but does not handle throttling gracefully when it occurs; it also incurs higher costs and does not address the root cause of occasional spikes. Option B is wrong because using an SQS queue to buffer write requests introduces eventual consistency and potential data loss if the queue messages expire or are not processed before the DynamoDB write; it also adds complexity and latency, and is not the standard pattern for handling DynamoDB throttling directly. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for reads only, not writes; it cannot cache write requests or mitigate write throttling.

49
MCQeasy

A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for later analysis using Amazon Athena. The logs are generated continuously, and the total volume is about 1 TB per day. The data must be queryable within minutes of arrival. Which storage solution is most appropriate?

A.Amazon DynamoDB table with JSON attribute
B.Amazon RDS for PostgreSQL table with JSON column
C.Amazon S3 bucket with partitioned folders
D.Amazon Redshift cluster with JSON ingestion
AnswerC

Amazon S3 stores JSON at low cost per terabyte and integrates natively with Athena, which queries data in place. Partitioning folders by date or service prunes scanned data, cutting query cost and latency, so logs become queryable within minutes of arrival.

Why this answer

Amazon S3 with partitioned folders is the most appropriate solution because it provides a cost-effective, scalable storage layer for semi-structured JSON logs, and integrates natively with Amazon Athena for serverless querying. By partitioning the data by time (e.g., year/month/day/hour), Athena can use partition pruning to minimize scanned data, enabling queries within minutes of arrival. S3's low cost per GB and lifecycle policies further optimize storage for the 1 TB/day volume.

Exam trap

AWS often tests the misconception that a data warehouse (Redshift) or a NoSQL database (DynamoDB) is required for analytical queries on semi-structured data, when in fact S3 with Athena is the most cost-effective and scalable solution for serverless ad-hoc analysis on raw logs.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is optimized for key-value and document access patterns with low-latency reads/writes, not for ad-hoc analytical queries on large volumes of JSON logs; scanning 1 TB/day would be prohibitively expensive and slow, and it lacks native integration with Athena. Option B is wrong because Amazon RDS for PostgreSQL is a relational database designed for transactional workloads, not for storing and analyzing 1 TB/day of semi-structured logs; it would require manual partitioning, incur high storage costs, and cannot scale to petabyte-scale analytics efficiently. Option D is wrong because Amazon Redshift is a petabyte-scale data warehouse optimized for complex analytical queries, but it is overkill and more expensive than S3 for raw log storage; ingesting 1 TB/day of JSON logs into Redshift requires an ETL pipeline (e.g., COPY from S3) and incurs compute costs even when idle, whereas S3 with Athena is serverless and pay-per-query.

50
MCQmedium

A data engineer manages an Amazon S3 data lake where analytics queries run through Amazon Athena. Monthly partition folders hold Parquet files, and each partition contains tens of thousands of small files averaging 40 KB. Athena queries that scan a single month take much longer than expected and consume far more bytes scanned than the actual data volume. The engineer must improve query performance without changing the table schema or the folder layout. What should the engineer do?

A.Enable S3 Transfer Acceleration on the bucket so that Athena can retrieve the small objects with lower latency across edge locations.
B.Run an AWS Glue ETL job that reads each partition and rewrites it as fewer, larger Parquet files of roughly 128 MB, then update the partitions in the Data Catalog.
C.Attach an S3 Lifecycle policy that transitions the Parquet objects to S3 Glacier Instant Retrieval after 30 days to improve read throughput.
D.Increase the number of partitions by splitting each monthly folder into daily folders and repointing the table at the new prefixes.
AnswerB

Consolidating many tiny files into fewer large Parquet files drastically reduces the per-file overhead that Athena pays when listing and opening objects, so scans finish faster and less metadata work is repeated. The rewrite keeps the same partition structure and schema, satisfying the constraint that neither the table definition nor the folder layout may change.

Why this answer

Athena performance degrades when a partition holds a huge number of very small objects because each object requires a separate request and contributes metadata overhead, inflating both runtime and reported bytes scanned. Compacting the files with an AWS Glue job into larger Parquet objects preserves the schema and partition layout while cutting that overhead, which is the standard remedy for this pattern.

Exam trap

The trap here is assuming that a storage-class or network-acceleration change can fix a small-file problem, when the real fix is rewriting the objects into fewer larger files.

51
MCQeasy

A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data is immutable and must be stored cost-effectively. Which AWS service should the engineer use?

A.Amazon DynamoDB with on-demand capacity
B.Amazon EBS with gp3 volume
C.Amazon S3 with S3 Standard-IA storage class
D.Amazon RDS for PostgreSQL with JSONB data type
AnswerC

S3 Standard-IA suits infrequent access with millisecond retrieval, satisfying the minutes-based constraint. Its lower storage cost than S3 Standard meets the cost-effectiveness requirement, while object immutability is preserved through versioning and Object Lock. Unlike Glacier tiers, no retrieval job or restore delay is needed, keeping access immediate.

Why this answer

Amazon S3 Standard-IA (Infrequent Access) is designed for data that is accessed less frequently but requires rapid retrieval when needed, with retrieval times in milliseconds. It offers lower storage costs than S3 Standard while maintaining high durability and availability, making it ideal for storing immutable semi-structured JSON files that must be retrievable within minutes. The service is cost-effective for infrequently accessed data because it charges a retrieval fee per GB, but the storage price is significantly lower than standard tiers.

Exam trap

The trap here is that candidates often confuse 'infrequently accessed' with 'archival' and choose Glacier or Deep Archive, but the requirement for retrieval within minutes eliminates those options, while DynamoDB or RDS seem plausible for JSON but are not cost-effective for immutable, infrequently accessed data.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB with on-demand capacity is a NoSQL database optimized for high-frequency, low-latency queries and is not cost-effective for infrequently accessed, immutable JSON files; it charges per read/write request unit and storage, which would be wasteful for archival-like data. Option B is wrong because Amazon EBS with gp3 volume is a block storage service designed for EC2 instances and requires an attached compute instance to access data, adding unnecessary cost and complexity; it is not a standalone object storage solution for infrequently accessed files. Option D is wrong because Amazon RDS for PostgreSQL with JSONB data type is a relational database service that incurs ongoing compute and storage costs, even when idle, and is overkill for storing immutable JSON files that are only occasionally retrieved; it is designed for transactional workloads and complex queries, not cost-effective archival storage.

52
MCQmedium

A data engineer manages an Amazon Redshift cluster that stores sales data in a table with a sort key on the sale_date column. The table is growing rapidly, and queries that filter by sale_date are becoming slower. The engineer notices that the table has a high percentage of unsorted rows. What should the engineer do to improve query performance with the least effort?

A.Run the VACUUM REINDEX command on the table.
B.Run the ANALYZE command on the table.
C.Run the VACUUM SORT ONLY command on the table.
D.Run the VACUUM FULL command on the table.
AnswerC

VACUUM SORT ONLY sorts the rows in the table according to the sort key without reclaiming space. This directly addresses the high percentage of unsorted rows and improves the efficiency of range-restricted scans on sale_date. It is a straightforward maintenance operation that requires minimal effort and targets the specific issue of unsorted data.

Why this answer

VACUUM SORT ONLY sorts the rows in a table according to the defined sort key, which directly improves the performance of queries that filter on that key. It is the least intrusive vacuum operation that addresses unsorted rows. ANALYZE only updates statistics, while VACUUM FULL and VACUUM REINDEX are more resource-intensive and not required for the described issue.

Exam trap

The trap here is confusing VACUUM SORT ONLY with VACUUM FULL; the former sorts without reclaiming space, while the latter does both and is overkill for just sorting.

53
Multi-Selectmedium

Which THREE of the following are valid storage classes in Amazon S3? (Choose THREE.)

Select 3 answers
A.S3 Standard
B.S3 Archive
C.S3 Intelligent-Tiering
D.S3 Cold
E.S3 One Zone-IA
AnswersA, C, E

S3 Standard is a genuine storage class, designed for frequently accessed data with low latency and high throughput, replicated across a minimum of three Availability Zones. It satisfies the question's requirement for a valid Amazon S3 storage class.

Why this answer

S3 Standard (A) is a valid storage class designed for frequently accessed data with low latency and high throughput, and it is the default class for S3 objects. S3 Intelligent-Tiering (C) is a valid class that automatically moves objects between access tiers based on changing access patterns, charging a small monitoring and automation fee. S3 One Zone-IA (E) is a valid class that stores data in a single Availability Zone at lower cost, intended for infrequently accessed data that can be easily recreated.

The other options are not real S3 storage classes: S3 Archive (B) does not exist — the archival class is S3 Glacier, and S3 Cold (D) is not a class either, as the closest real classes are S3 Glacier Instant Retrieval and S3 Glacier Flexible Retrieval.

Exam trap

AWS often tests the distinction between valid S3 storage classes and fabricated names like 'S3 Archive' or 'S3 Cold', expecting candidates to recall the exact naming conventions (e.g., S3 Glacier, S3 Glacier Deep Archive) rather than generic terms.

54
Multi-Selectmedium

A data engineer is designing a disaster recovery strategy for an Amazon RDS for PostgreSQL database. The primary database is in us-east-1. Which TWO approaches provide cross-region disaster recovery?

Select 2 answers
A.Configure cross-region automated backups to copy to us-west-2.
B.Take a manual snapshot and copy it to us-west-2 daily.
C.Use Amazon S3 cross-region replication for the database export.
D.Enable Multi-AZ in us-east-1.
E.Create a cross-region read replica in us-west-2.
AnswersA, E

Backups are automatically copied and can be restored.

Why this answer

Amazon RDS supports cross-region automated backups, which automatically copy backup data (snapshots and transaction logs) from the primary region (us-east-1) to a secondary region (us-west-2). This provides a fully managed, automated disaster recovery solution that allows point-in-time recovery in the secondary region without manual intervention.

Exam trap

The trap here is that candidates often confuse Multi-AZ (which provides in-region high availability) with cross-region disaster recovery, or they assume manual snapshot copying is equivalent to automated cross-region backups, not realizing the significant difference in RPO and operational overhead.

55
MCQhard

A company is using Amazon ElastiCache for Redis to cache frequently accessed data. The cache hit ratio is low, and the engineering team suspects that the eviction policy is causing important data to be removed. Which eviction policy should be used to minimize eviction of the most frequently accessed keys?

A.allkeys-lru
B.allkeys-lfu
C.noeviction
D.volatile-lru
AnswerB

allkeys-lfu evicts the least frequently used keys across the entire keyspace, so frequently accessed items survive even when newer keys arrive. This directly raises the hit ratio by retaining hot keys, unlike LRU or volatile-ttl policies that discard them based on recency or expiry.

Why this answer

The allkeys-lfu (Least Frequently Used) eviction policy is the correct choice because it explicitly tracks and retains keys that are accessed most frequently across the entire keyspace. Since the cache hit ratio is low due to eviction of important data, LFU ensures that frequently accessed keys are evicted last, directly addressing the problem of important data being removed.

Exam trap

The trap here is that candidates often confuse recency (LRU) with frequency (LFU), assuming that 'least recently used' also implies 'least frequently used,' but LRU can evict a frequently accessed key that hasn't been touched recently, which is exactly the problem described.

How to eliminate wrong answers

Option A is wrong because allkeys-lru (Least Recently Used) evicts keys based on recency of access, not frequency, so a frequently accessed key that hasn't been used recently could be evicted. Option C is wrong because noeviction returns errors for write operations when memory is full, which would cause application failures rather than solving the low hit ratio. Option D is wrong because volatile-lru only applies to keys with a TTL set, leaving keys without TTLs unprotected and potentially evicting important data that lacks an expiration.

56
MCQhard

A data engineer is designing an Amazon DynamoDB table for an order-processing application. The table uses a partition key of order_id and a sort key of order_date. The application needs to retrieve all orders for a specific customer within a date range, and the queries must be efficient at scale. The engineer must choose a design that supports these access patterns without full table scans. What should the engineer do?

A.Enable DynamoDB Streams and use a Lambda function to write customer-based query results to a separate table.
B.Create a local secondary index with customer_id as the partition key and order_date as the sort key.
C.Create a global secondary index with customer_id as the partition key and order_date as the sort key.
D.Use a Scan operation with a FilterExpression on customer_id and order_date.
AnswerC

A global secondary index with customer_id as the partition key and order_date as the sort key allows the application to query by customer and date range directly. DynamoDB can then use the index to retrieve only the relevant items, avoiding a full table scan. This design matches the required access pattern and scales horizontally with the table.

Why this answer

The application needs to query orders by customer and date range efficiently. A global secondary index can use a different partition key and sort key from the base table, so customer_id as the partition key and order_date as the sort key provides the required access path. Local secondary indexes must share the base table partition key, scans are inefficient, and stream-based copies do not replace a queryable index.

Exam trap

The trap here is confusing local secondary indexes with global secondary indexes, because a local secondary index cannot use a different partition key from the base table.

57
MCQhard

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query loads. The cluster uses a dc2.large node type with 2 nodes. Analysis shows that the workload involves frequent large table scans and complex joins. The engineer wants to improve query performance without changing the overall data volume. Which action should the engineer take?

A.Change the distribution style of large tables to KEY on the join columns and enable sort keys on frequently filtered columns.
B.Increase the number of nodes in the cluster by resizing to a larger node type with more memory.
C.Enable concurrency scaling to automatically add transient clusters during peak loads.
D.Enable short query acceleration (SQA) to prioritize short-running queries.
AnswerA

Choosing an appropriate distribution key ensures that join data is collocated on the same nodes, reducing data movement during joins. Sort keys on filtered columns enable efficient range scans. Together, these optimizations significantly improve performance for large scans and complex joins on a small cluster.

Why this answer

Proper distribution and sort keys are fundamental Redshift tuning techniques. Distributing large tables on join keys collocates data, minimizing network traffic during joins. Sort keys on filter columns allow the query engine to skip blocks.

These changes directly address the performance bottlenecks without adding hardware, making them the most effective and cost-efficient solution.

Exam trap

The trap here is assuming that adding more nodes or enabling concurrency scaling will automatically fix performance issues, when the root cause is suboptimal data distribution and sorting.

58
Multi-Selecteasy

Which THREE actions can help improve read performance in Amazon DynamoDB? (Choose THREE.)

Select 3 answers
A.Use DynamoDB global tables to replicate data.
B.Use parallel scans to distribute read load across partitions.
C.Use strongly consistent reads for all queries.
D.Enable DynamoDB Accelerator (DAX) to cache reads.
E.Increase the read capacity units (RCU) for the table.
AnswersB, D, E

Parallel scans can improve scan performance.

Why this answer

Parallel scans in DynamoDB can improve read performance by dividing a scan operation into multiple segments that are processed concurrently across partitions. This reduces the overall latency of the scan by leveraging the distributed nature of DynamoDB's storage, though it consumes more read capacity units (RCUs) due to the parallel execution.

Exam trap

The DEA-C01 exam often tests the misconception that strongly consistent reads always improve performance, when in fact they increase latency and RCU consumption, making eventually consistent reads the better choice for read-heavy workloads.

59
MCQmedium

A data engineer is building a data lake on Amazon S3. The raw data arrives as JSON files, but the analytics team needs to query the data using standard SQL in Amazon Athena with optimal performance and minimal cost. The engineer wants to convert the JSON to a columnar format that supports predicate pushdown and is natively supported by Athena. Which storage format should the engineer choose?

A.Apache Parquet
B.CSV
D.Apache Avro
AnswerA

Apache Parquet is a columnar format that stores data by column, enabling Athena to read only the columns referenced in a query. This reduces I/O and cost. Parquet also supports predicate pushdown and is natively supported by Athena, making it ideal for this scenario where the goal is optimal query performance and minimal cost.

Why this answer

Apache Parquet is the best choice because it is a columnar format that allows Athena to read only the necessary columns, reducing data scanned and cost. It also supports predicate pushdown, which filters data at the storage level. Other formats like Avro, JSON, and CSV are row-based and less efficient for analytical queries in Athena.

Exam trap

The trap here is assuming that any format supported by Athena is equally performant, but columnar formats like Parquet provide significant advantages for query performance and cost.

60
MCQmedium

A data engineer manages an Amazon S3 data lake with millions of small JSON files under prefixes partitioned by year/month/day. Amazon Athena queries scan far more data than expected and return slowly. The engineer wants to reduce bytes scanned and improve query performance while keeping files in S3 and queryable by Athena. Which solution meets these requirements with the LEAST operational overhead?

A.Enable S3 Transfer Acceleration on the bucket and configure Athena workgroups to use it for all queries.
B.Use AWS Glue to convert the JSON files to Apache Parquet, write them to partitioned prefixes, and register the resulting tables in the AWS Glue Data Catalog.
C.Attach an S3 Lifecycle policy that transitions the JSON objects to S3 Glacier Instant Retrieval after 30 days and query them through Athena.
D.Create an Amazon S3 Inventory report for each partition and use it to rewrite the Athena queries so they reference only the inventory files.
AnswerB

Converting to columnar Parquet with partitioning lets Athena read only needed columns and partitions, cutting bytes scanned. Registering the tables in the Data Catalog makes them immediately queryable without managing a separate metastore, and Glue handles the conversion job, keeping operational overhead low compared with self-managed ETL.

Why this answer

Athena charges and performs based on bytes scanned, and row-oriented JSON forces reading entire objects. Converting to columnar Parquet lets the engine read only referenced columns, while partitioning by year/month/day enables partition pruning so irrelevant dates are skipped. Registering the converted data in the AWS Glue Data Catalog makes it available to Athena with minimal management, directly reducing scanned bytes and latency.

Exam trap

The trap here is assuming that changing storage class or accelerating transfer improves Athena query cost and speed, when the real driver is data format and partition pruning.

61
MCQmedium

A data engineer maintains an Amazon S3 data lake with millions of small JSON objects. The engineer needs to improve query performance by reducing the number of objects and compressing them into a columnar format that Amazon Athena can query efficiently. The data must remain partitioned by date. Which solution should the engineer use?

A.Enable S3 Transfer Acceleration on the bucket and run Amazon Athena queries with the existing JSON objects.
B.Use AWS Glue ETL to read the JSON objects, repartition by date, write them as Parquet with Snappy compression, and register the resulting tables in the AWS Glue Data Catalog.
C.Use Amazon S3 Lifecycle policies to transition the JSON objects to S3 Glacier Instant Retrieval and query them with Athena.
D.Create an Amazon EMR cluster and run a MapReduce job that merges the JSON files into larger JSON files without changing the format.
AnswerB

AWS Glue ETL can transform JSON to Parquet, repartition by date, and update the Data Catalog, which Athena uses for schema and partition metadata. Parquet with Snappy reduces storage and scan size, and fewer larger files improve query performance. This directly addresses the small-file problem and columnar requirement while preserving date partitioning.

Why this answer

Converting small JSON files to Parquet with Snappy compression and repartitioning by date reduces storage footprint and the amount of data scanned by Athena. AWS Glue ETL can perform this transformation and update the Data Catalog so Athena queries the new tables. This is the standard pattern for optimizing S3 data lakes for analytics.

Exam trap

The trap here is assuming that simply merging files or changing storage classes solves the small-file problem, when the key is converting to a columnar format and repartitioning.

62
MCQeasy

A data engineer needs to store semi-structured JSON event logs in a data lake on Amazon S3 and query them with Amazon Athena using SQL, including filtering on individual JSON attributes. The team wants to avoid transforming the files before querying. Which approach should the engineer use?

A.Store the logs as CSV and use Athena to parse the columns
B.Use Amazon CloudWatch Logs Insights to query the JSON logs in place
C.Load the JSON into Amazon Redshift using COPY JSON and query it there
D.Define an Athena table with a struct or map column over the JSON and query nested fields
AnswerD

Athena, built on the AWS Glue Data Catalog, supports JSON SerDe and complex types such as struct, array, and map. Defining columns with these types lets SQL reference nested attributes directly, for example using dot notation on a struct field. No preprocessing is needed, which satisfies the requirement to query raw JSON in place while still allowing filters on individual attributes.

Why this answer

Athena queries data in place on S3 using table definitions in the AWS Glue Data Catalog. For semi-structured JSON, defining columns as struct, array, or map with the appropriate JSON SerDe lets SQL access nested attributes directly without preprocessing. Storing as CSV would require flattening, loading into Redshift adds a transformation step, and CloudWatch Logs Insights is scoped to log groups rather than the S3 data lake.

Exam trap

The trap here is assuming JSON must be flattened or loaded into a warehouse before SQL can filter individual attributes, when Athena supports complex types that expose nested fields directly.

63
Multi-Selecteasy

A company is designing a data lake on Amazon S3. The data includes CSV files, Parquet files, and images. The data engineering team needs to catalog the metadata and enable SQL queries. Which TWO AWS services should be used together?

Select 2 answers
A.Amazon EMR
B.Amazon Redshift Spectrum
C.Amazon QuickSight
D.Amazon Athena
E.AWS Glue
AnswersD, E

Amazon Athena queries S3 data directly using standard SQL, reading the table and partition metadata held in the AWS Glue Data Catalog. It satisfies the SQL query requirement without loading data, complementing Glue's cataloguing role for the CSV, Parquet and image lake.

Why this answer

Amazon Athena is correct because it is a serverless interactive query service that can directly query data stored in Amazon S3 using standard SQL, without needing to load or transform data. AWS Glue is correct because it provides a fully managed data catalog (AWS Glue Data Catalog) that stores metadata about the data lake's schema, partitions, and locations, which Athena can use to discover and query the data efficiently.

Exam trap

The trap here is that candidates often confuse Amazon Redshift Spectrum (which requires a Redshift cluster) with Athena (which is serverless), or they think Amazon EMR is needed for SQL queries on S3, not realizing Athena provides a simpler, cluster-free solution.

64
MCQmedium

An e-commerce company uses Amazon DynamoDB as the primary data store for its product catalog. The table has a simple primary key (ProductID) and handles 10,000 writes per second during peak hours. Recently, the engineering team noticed increased write latency and throttled requests during peak times. The table's provisioned write capacity is set to 12,000 WCU. What is the most likely cause of the throttling?

A.The table has reached the maximum number of partitions
B.DynamoDB Accelerator (DAX) is not configured
C.Write traffic is unevenly distributed across partitions
D.A global secondary index is consuming write capacity
AnswerC

Unevenly distributed write traffic concentrates load on a subset of partitions, so individual partitions exceed their 1,000 WCU per-partition limit even though total provisioned capacity of 12,000 WCU is not exhausted. DynamoDB throttles at partition level, making hot partitions the likely cause of the latency and throttling during peak hours.

Why this answer

DynamoDB partitions data by the primary key's hash value. If write traffic is unevenly distributed across partitions (e.g., a few ProductIDs receive most writes), those hot partitions can exceed their individual throughput limits (3,000 WCU per partition for provisioned tables), causing throttling even when the table's total provisioned WCU of 12,000 is not fully utilized.

Exam trap

The trap here is that candidates assume throttling only occurs when total provisioned capacity is exceeded, overlooking the per-partition throughput limits that cause throttling on hot partitions even when the table's overall WCU is underutilized.

How to eliminate wrong answers

Option A is wrong because DynamoDB tables do not have a maximum number of partitions; partitions are automatically added or removed based on storage and throughput needs. Option B is wrong because DAX is an in-memory cache for reads, not writes; it does not affect write capacity or throttling. Option D is wrong because while a global secondary index (GSI) does consume write capacity from the table's WCU pool, the question states the table has 12,000 WCU provisioned, and throttling occurs during peak writes of 10,000 writes per second, so the GSI would only contribute to throttling if its own provisioned WCU were insufficient, but the scenario does not indicate that.

65
MCQhard

A media company ingests millions of small JSON files per day into an Amazon S3 bucket. Analysts run Amazon Athena queries over this data and report that each query scans far more data than the files matching their filters, resulting in high cost and slow performance. The files are partitioned by year/month/day in S3. What should a data engineer do to reduce the data scanned per query?

A.Move the files into an Amazon Redshift cluster using COPY and query them with Redshift Spectrum.
B.Create an Amazon CloudFront distribution in front of the S3 bucket and run Athena against the distribution.
C.Enable S3 Transfer Acceleration on the bucket and increase the Athena query timeout.
D.Convert the JSON files to Apache Parquet, store them in the existing partition structure, and define the table in the AWS Glue Data Catalog with the correct partition columns.
AnswerD

Athena charges by data scanned, and columnar Parquet with predicate pushdown lets the engine read only the columns and row groups needed. Combined with partition pruning on year/month/day, queries skip irrelevant prefixes entirely. This directly reduces bytes scanned, lowering cost and improving latency without changing the query interface.

Why this answer

Athena's cost model is based on bytes scanned, so the most effective optimization is to store data in a columnar format such as Parquet and rely on partition pruning. Parquet enables column projection and predicate pushdown, while year/month/day partitions let Athena skip entire prefixes. Together they sharply reduce scanned bytes, cutting cost and improving query speed for the analyst workload.

Exam trap

The trap here is focusing on ingestion or delivery speedups such as Transfer Acceleration or CloudFront, which do not change how many bytes Athena reads during a query.

66
MCQhard

A company uses Amazon Redshift for its data warehouse. The data engineer notices that queries are slow on a large table that is frequently filtered on a column 'transaction_date'. Which optimization technique best improves query performance?

A.Apply compression encoding to 'transaction_date'.
B.Set the sort key to 'transaction_date'.
C.Set the distribution key to 'transaction_date'.
D.Run VACUUM on the table.
AnswerB

Sort keys physically order rows on disk by transaction_date, so Redshift's zone maps skip irrelevant blocks when filtering that column, cutting scanned data. This targets the stem's frequently filtered column directly, unlike distribution keys, which address join and data-skew costs instead.

Why this answer

Setting the sort key to 'transaction_date' organizes the table data physically by that column, which allows Redshift to use zone maps to skip blocks that don't match query filters. This dramatically reduces the amount of data scanned for range-restricted queries on 'transaction_date', improving query performance.

Exam trap

The trap here is that candidates confuse distribution keys (which optimize joins) with sort keys (which optimize filtering and range scans), leading them to pick distribution key as the answer for a single-table filter performance issue.

How to eliminate wrong answers

Option A is wrong because compression encoding reduces storage size and I/O but does not directly optimize query filtering on a column; it can even slow down scans if the column is frequently used in predicates. Option C is wrong because setting the distribution key to 'transaction_date' distributes rows across nodes based on that column, which can help with joins but does not improve the efficiency of range-restricted scans on a single table. Option D is wrong because VACUUM reclaims space and re-sorts data but does not improve query performance unless the table is already sorted on a key; without a sort key on 'transaction_date', VACUUM has no effect on filter performance.

67
MCQmedium

A company stores log files in Amazon S3. They want to automatically move logs older than 90 days to S3 Glacier Deep Archive to reduce costs. Which S3 feature should be used?

A.S3 Intelligent-Tiering
B.S3 Lifecycle configuration
C.S3 Replication
D.S3 Object Lock
AnswerB

S3 Lifecycle configuration applies transition rules that automatically move objects to Glacier Deep Archive once they cross the specified age threshold, here 90 days. It satisfies the stem's requirement for automatic, age-based archival without custom code, and transitions are billed per request rather than requiring retrieval or manual intervention.

Why this answer

S3 Lifecycle configuration allows you to define rules that automatically transition objects to colder storage classes, such as S3 Glacier Deep Archive, based on age. By setting a rule to move objects older than 90 days to S3 Glacier Deep Archive, you reduce storage costs without manual intervention. This is the correct feature for automating tier-based data lifecycle management.

Exam trap

The trap here is that candidates may confuse S3 Intelligent-Tiering with lifecycle policies, but Intelligent-Tiering does not support age-based transitions to Glacier Deep Archive and is designed for unpredictable access patterns, not fixed retention schedules.

How to eliminate wrong answers

Option A is wrong because S3 Intelligent-Tiering automatically moves data between access tiers based on changing access patterns, not on a fixed age-based schedule, and it does not support direct transition to S3 Glacier Deep Archive. Option C is wrong because S3 Replication is used to copy objects across buckets or regions for redundancy or compliance, not to transition objects to colder storage classes. Option D is wrong because S3 Object Lock is designed to prevent object deletion or overwrites for a specified retention period, not to manage storage tier transitions.

68
MCQeasy

A data engineer needs to store semi-structured JSON data that is accessed infrequently but requires immediate retrieval when needed. The data must be durable and cost-effective. Which Amazon S3 storage class should be used?

A.S3 Standard-IA
B.S3 Glacier
C.S3 Standard
D.S3 One Zone-IA
AnswerA

S3 Standard-IA suits infrequent access with immediate retrieval, satisfying the low-frequency, on-demand constraint. Its lower storage cost than S3 Standard meets cost-effectiveness, while replicating across a minimum of three Availability Zones preserves durability. Unlike Glacier classes, no retrieval delay applies, so JSON remains instantly available.

Why this answer

S3 Standard-IA is the correct choice because it offers the same durability and low-latency retrieval as S3 Standard but at a lower storage cost, making it ideal for infrequently accessed data that still needs immediate retrieval when requested. The scenario specifies 'infrequently accessed' and 'immediate retrieval,' which aligns with Standard-IA's design for data accessed less than once a month but with millisecond first-byte latency.

Exam trap

The trap here is that candidates often confuse 'infrequently accessed' with 'archival' and choose S3 Glacier, overlooking the 'immediate retrieval' requirement that rules out Glacier's multi-minute or multi-hour retrieval times.

How to eliminate wrong answers

Option B (S3 Glacier) is wrong because it is designed for archival data with retrieval times ranging from minutes to hours, not immediate retrieval. Option C (S3 Standard) is wrong because it is optimized for frequently accessed data and would be less cost-effective for infrequently accessed data, incurring higher storage costs without benefit. Option D (S3 One Zone-IA) is wrong because it stores data in a single Availability Zone, which does not meet the durability requirement of the scenario (data must be durable, implying multi-AZ resilience).

69
MCQmedium

A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. The table uses a partition key of deviceId and a sort key of timestamp, with a provisioned read capacity of 100 RCUs. During a sudden spike in traffic, the engineer observes throttling on read operations even though the consumed read capacity is well below the provisioned limit. What is the MOST likely cause of the throttling?

A.The table uses eventually consistent reads, which are throttled more aggressively than strongly consistent reads.
B.The provisioned read capacity is set too low for the table's data volume.
C.The table lacks a global secondary index (GSI) on the timestamp attribute.
D.The table's partition key is a low-cardinality attribute, causing a hot partition.
AnswerD

A low-cardinality partition key, such as deviceId with few unique values, concentrates read traffic on a small number of partitions. DynamoDB limits each partition to 3,000 RCUs, so even if total consumed capacity is under 100 RCUs, a single partition can exceed its local limit and throttle requests. This matches the symptom of throttling despite low aggregate usage.

Why this answer

DynamoDB throttling can occur even when total consumed capacity is below the provisioned level if reads are concentrated on a few partitions. Each partition has a maximum throughput of 3,000 RCUs. A low-cardinality partition key creates hot partitions, causing localized throttling.

The solution is to use a higher-cardinality partition key or add a write sharding strategy.

Exam trap

The trap here is assuming that throttling always means total provisioned capacity is insufficient, overlooking partition-level limits.

70
MCQhard

Refer to the exhibit. An IAM policy is attached to an IAM role used by an application. The application needs to decrypt objects in an S3 bucket using a customer managed KMS key. What is the effect of this policy?

A.The application cannot perform any KMS operations.
B.The application can decrypt objects from any service.
C.The application can decrypt objects only when accessing them through S3.
D.The application can encrypt but not decrypt objects.
AnswerC

The policy grants kms:Decrypt but conditions it on the request arriving via the S3 service, so the principal can decrypt only through S3 GetObject calls. Direct KMS decrypt requests, or access via other services, are denied by the condition.

Why this answer

The IAM policy grants the `kms:Decrypt` permission with a `kms:ViaService` condition key set to `s3.amazonaws.com`. This condition restricts the decryption operation to only when the request is made through the S3 service. Therefore, the application can decrypt objects only when accessing them through S3, not via direct KMS API calls or other services.

Exam trap

AWS often tests the `kms:ViaService` condition key to trap candidates who assume that granting `kms:Decrypt` alone allows decryption from any source, ignoring the service-specific restriction.

How to eliminate wrong answers

Option A is wrong because the policy explicitly allows `kms:Decrypt` under the condition, so the application can perform KMS decryption operations when invoked via S3. Option B is wrong because the `kms:ViaService` condition restricts decryption to S3 only, preventing decryption from any other service or direct KMS API calls. Option D is wrong because the policy grants `kms:Decrypt` permission, not `kms:Encrypt`, so the application can decrypt but not encrypt objects.

71
Multi-Selectmedium

A data engineer needs to store event data from IoT devices that arrives in bursts. The data is key-value and requires single-digit millisecond read and write latency. The engineer also needs to run complex analytical queries on the data for reporting. Which TWO services should be used together? (Choose TWO.)

Select 2 answers
A.Amazon DynamoDB
B.Amazon ElastiCache for Redis
C.Amazon Redshift
D.Amazon S3
E.Amazon RDS for MySQL
AnswersA, C

Amazon DynamoDB delivers consistent single-digit millisecond latency for key-value workloads, satisfying the burst IoT ingestion requirement. Its partition-based architecture scales horizontally without manual sharding. However, DynamoDB alone cannot run complex analytical queries, so it must pair with a separate analytics service to meet the reporting constraint.

Why this answer

Amazon DynamoDB (A) is correct because it is a fully managed key-value and document NoSQL database that delivers consistent single-digit millisecond read and write latency at any scale, making it ideal for bursty IoT event ingestion. Amazon Redshift (C) is correct because it is a petabyte-scale, columnar data warehouse designed for complex analytical queries and reporting, so it complements DynamoDB by handling the analytics workload. Together, DynamoDB captures the high-throughput key-value events while Redshift serves the reporting and complex query needs.

Amazon ElastiCache for Redis (B) is an in-memory cache, not a durable primary store for event data, and it does not natively run complex analytical SQL queries. Amazon S3 (D) is object storage with higher latency and no native single-digit millisecond key-value access or complex query engine. Amazon RDS for MySQL (E) is a relational database that generally cannot match DynamoDB's single-digit millisecond latency at burst scale and is not optimized for large-scale complex analytical reporting like Redshift.

Exam trap

The trap here is that candidates often choose ElastiCache for Redis because of its low latency, forgetting that it is not a durable data store for analytical queries, or they pick S3 thinking it can serve as a primary database, ignoring its lack of single-digit millisecond latency for key-value access.

72
MCQhard

A media company ingests thousands of small JSON files per hour into an Amazon S3 bucket. A data engineer needs to convert these files into a compact, columnar format for efficient querying with Amazon Athena. The engineer wants to minimize storage costs and improve query performance. Which approach should the engineer take?

A.Use AWS Glue ETL to read the JSON files, convert them to Apache Parquet, and write the output to a new S3 prefix partitioned by date.
B.Use Amazon Athena to create a new table as Parquet using CREATE TABLE AS SELECT (CTAS) from the JSON table.
C.Use Amazon Kinesis Data Firehose to convert JSON to Parquet and deliver to S3.
D.Use AWS Lambda to read each JSON file, convert to Parquet, and write back to S3.
AnswerA

AWS Glue ETL can read JSON, transform to Parquet, and write partitioned data to S3. Parquet is columnar, reducing storage and improving Athena query performance. Partitioning by date further reduces data scanned. This directly addresses the need for compact columnar format and cost efficiency.

Why this answer

AWS Glue ETL is a fully managed extract, transform, and load service that can efficiently process large volumes of data. It can read JSON from S3, convert to Parquet, and write partitioned output. This reduces storage costs due to Parquet's compression and columnar format, and improves Athena query performance by reducing data scanned.

Partitioning by date further optimizes queries.

Exam trap

The trap here is assuming that streaming services like Kinesis Data Firehose or serverless functions like Lambda are the best fit for batch conversion of existing small files, when a managed ETL service like AWS Glue is more appropriate.

73
Multi-Selectmedium

Which TWO options are valid ways to reduce storage costs for an Amazon S3 data lake that stores historical data rarely accessed after 30 days? (Choose TWO.)

Select 2 answers
A.Enable S3 Transfer Acceleration for all uploads.
B.Create a lifecycle policy to transition objects to S3 Standard-IA after 30 days.
C.Create a lifecycle policy to delete objects after 30 days.
D.Create a lifecycle policy to transition objects to S3 Glacier Deep Archive after 90 days.
E.Enable S3 Versioning to preserve all object versions.
AnswersB, D

Standard-IA reduces storage cost for infrequent access.

Why this answer

S3 Standard-IA (Infrequent Access) is designed for data accessed less frequently but requires rapid access when needed. Transitioning objects to Standard-IA after 30 days reduces storage costs compared to S3 Standard while maintaining low-latency retrieval, making it ideal for a data lake where historical data is rarely accessed after the first month.

Exam trap

The trap here is that candidates often confuse data protection features (like Versioning) with cost optimization, or they assume that deleting data is the only way to reduce costs, overlooking lifecycle transitions to lower-cost storage classes that retain data accessibility.

74
MCQmedium

A data engineer is designing a data store for a real-time analytics application that requires low-latency reads and writes at scale. The data model includes time-series data with high ingest rates and queries that aggregate data over sliding time windows. The engineer needs a fully managed AWS service that supports automatic scaling and can handle millions of writes per second. Which service should the engineer choose?

A.Amazon S3
B.Amazon Redshift
C.Amazon RDS for PostgreSQL
D.Amazon DynamoDB
AnswerD

DynamoDB is a fully managed NoSQL database that supports high-throughput reads and writes with automatic scaling. It can handle millions of requests per second and provides low-latency performance. With features like time-to-live (TTL) and on-demand capacity mode, it is well-suited for time-series data and real-time analytics. Its ability to scale horizontally without downtime makes it the best fit for this scenario.

Why this answer

Amazon DynamoDB is a fully managed NoSQL database that provides low-latency performance at any scale, with automatic scaling and support for high-throughput workloads. It can handle millions of writes per second and is ideal for time-series data due to features like TTL and on-demand capacity. Its ability to scale horizontally without downtime makes it the best choice for real-time analytics applications requiring low-latency reads and writes.

Exam trap

The trap here is assuming that a relational database or data warehouse can handle the same scale and latency requirements as a purpose-built NoSQL service like DynamoDB.

75
MCQeasy

A company's analytics team needs a petabyte-scale, fully managed data warehouse that supports standard SQL, columnar storage, and massively parallel query execution, and it must integrate with existing business intelligence tools with minimal operational effort. Which AWS service should the data engineer choose?

A.Amazon Redshift
B.AWS Database Migration Service
C.Amazon ElastiCache for Redis
D.Amazon DynamoDB
AnswerA

Amazon Redshift is a fully managed, petabyte-scale data warehouse that uses columnar storage on a cluster of nodes with massively parallel processing, and it speaks standard SQL that BI tools connect to through JDBC and ODBC drivers. It matches every stated requirement, including minimal operational effort, because AWS handles provisioning, patching, and backups of the cluster.

Why this answer

The workload calls for a managed analytical database with columnar storage, massively parallel execution, and standard SQL connectivity for BI tools. Amazon Redshift delivers all three as a managed service, so the team avoids building and operating its own warehouse. The other services target transactional NoSQL access, in-memory caching, or data movement rather than petabyte-scale SQL analytics.

Exam trap

The trap here is being drawn to any highly scalable AWS data store and overlooking that only a columnar, MPP SQL warehouse satisfies the analytics and BI requirements.

Page 1 of 5 · 358 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Store Management questions.