Courseiva

CCNA Data Store Management Questions

75 of 358 questions · Page 2/5 · Data Store Management topic · Answers revealed

76
MCQmedium

A data engineer needs to migrate an on-premises Apache Hadoop cluster to AWS. The cluster stores data in HDFS and runs MapReduce jobs. The company wants to minimize operational overhead and leverage serverless technologies where possible. Which AWS service should the data engineer use to replace HDFS storage?

A.Amazon EBS
B.Amazon EMR
C.Amazon S3
D.Amazon Redshift
AnswerC

Amazon S3 provides durable, serverless object storage that replaces HDFS without cluster management, satisfying the minimise-operational-overhead constraint. MapReduce workloads migrate to Amazon EMR or Athena, while S3 becomes the decoupled storage layer, eliminating NameNode and DataNode administration entirely.

Why this answer

Amazon S3 is the correct replacement for HDFS because it provides highly durable, scalable, and serverless object storage that can be used as the primary storage layer for Amazon EMR. Unlike HDFS, S3 decouples storage from compute, eliminating the need to manage cluster storage and allowing jobs to run on ephemeral clusters, which minimizes operational overhead. S3 integrates with EMR via the EMR File System (EMRFS), enabling MapReduce jobs to read/write data directly from S3 as if it were HDFS.

Exam trap

The trap here is that candidates confuse Amazon EMR (a compute service) with a storage service, assuming it replaces HDFS, when in fact EMR can use either HDFS or S3 for storage, and the question explicitly asks for the storage replacement.

How to eliminate wrong answers

Option A is wrong because Amazon EBS provides block-level storage volumes attached to EC2 instances, which is not serverless and requires manual management of volume size, snapshots, and replication; it also ties storage to a specific compute instance, defeating the purpose of decoupling storage from compute for a Hadoop migration. Option B is wrong because Amazon EMR is a managed big data platform that runs MapReduce jobs, not a storage service; it can use HDFS or S3 for storage, but the question specifically asks for a replacement of HDFS storage, not the compute framework. Option D is wrong because Amazon Redshift is a fully managed data warehouse optimized for SQL-based analytics and structured data, not a general-purpose distributed file system for Hadoop workloads; it does not support HDFS semantics or MapReduce jobs natively.

77
MCQmedium

A data engineer stores application logs in an Amazon S3 bucket. Compliance requires that log objects be retained for seven years and that they cannot be deleted or overwritten by any user, including the account root user, during that period. The engineer must configure the bucket to enforce this. Which combination of settings should the engineer apply?

A.Enable S3 Object Lock in governance mode and grant the s3:BypassGovernanceRetention permission only to the root user.
B.Enable S3 Versioning and attach a bucket policy that denies s3:DeleteObject to all principals except the root user.
C.Enable S3 Object Lock in compliance mode with a default retention period of seven years on the bucket.
D.Configure an S3 Lifecycle rule to transition objects to S3 Glacier Deep Archive after one day and expire them after seven years.
AnswerC

S3 Object Lock in compliance mode prevents any user, including the root user, from overwriting or deleting a protected object version until the retention period expires. Setting a default retention of seven years on the bucket applies the rule automatically to new objects, satisfying the immutability and retention requirements.

Why this answer

S3 Object Lock provides write-once-read-many protection at the object version level. Compliance mode is the strictest setting: no principal, including the account root user, can delete or overwrite a locked version before retention expires. Applying a bucket default retention of seven years automatically protects newly uploaded log objects without per-object calls, meeting the compliance requirement.

Exam trap

The trap here is believing that a bucket policy or governance mode can stop the root user, when only compliance-mode Object Lock removes all override capability.

78
MCQhard

A company uses Amazon S3 to store sensitive financial data. The security team requires that all objects be encrypted at rest using AWS KMS with a customer-managed key. Additionally, they want to audit all KMS decrypt calls for compliance. Which configuration should be used to meet these requirements?

A.Enable default encryption on the bucket with SSE-KMS using an AWS managed key.
B.Use SSE-S3 with a bucket policy that denies uploads without encryption.
C.Use SSE-KMS with a customer-managed KMS key and enable CloudTrail data events for the key.
D.Use SSE-C with client-managed keys and log S3 API calls.
AnswerC

SSE-KMS with a customer-managed key satisfies the encryption-at-rest requirement, while CloudTrail data events record KMS Decrypt API calls against that key, delivering the required audit trail. Default CloudTrail management events alone would not capture per-object decrypt activity.

Why this answer

SSE-KMS with a customer-managed key allows the company to control the encryption key lifecycle and meet the requirement for customer-managed keys. Enabling CloudTrail data events for the KMS key captures all decrypt API calls, providing the necessary audit trail for compliance.

Exam trap

The trap here is that candidates may confuse enabling default encryption on the bucket (which can use SSE-KMS) with the need for a customer-managed key and CloudTrail data events, or they may think SSE-S3 or SSE-C can satisfy the audit requirement without KMS-specific logging.

How to eliminate wrong answers

Option A is wrong because it uses an AWS managed key, not a customer-managed key, so the security team cannot control key rotation or access policies. Option B is wrong because SSE-S3 uses server-side encryption with S3-managed keys, which does not provide customer-managed key control, and the bucket policy only enforces encryption, not auditing of decrypt calls. Option D is wrong because SSE-C requires the client to manage the encryption keys, which does not meet the requirement for AWS KMS, and logging S3 API calls alone does not capture KMS decrypt events.

79
MCQeasy

A startup is building a mobile application that requires a database to store user profiles and preferences. The database must scale automatically with minimal administration. Which AWS service should they use?

A.Amazon Redshift
B.Amazon Aurora
C.Amazon DynamoDB
D.Amazon RDS for PostgreSQL
AnswerC

DynamoDB is fully managed and serverless, scaling throughput automatically via on-demand capacity with no server provisioning or patching. This directly satisfies the stem's constraints of automatic scaling and minimal administration for storing user profiles and preferences.

Why this answer

Amazon DynamoDB is a fully managed NoSQL key-value and document database that delivers single-digit millisecond performance at any scale. It supports automatic scaling of throughput capacity and storage with no downtime, making it ideal for a mobile application that requires minimal administrative overhead. The serverless, pay-per-request billing model aligns perfectly with the startup's need for automatic scaling and low operational burden.

Exam trap

The trap here is that candidates often choose a relational database like Aurora or RDS because they assume user profiles require complex joins or ACID transactions, but DynamoDB's single-table design and conditional updates can handle most mobile app patterns with simpler, more scalable operations.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a petabyte-scale data warehouse optimized for complex analytical queries, not for transactional user profile storage, and it requires manual scaling and cluster management. Option B is wrong because Amazon Aurora is a relational database that, while offering some auto-scaling for storage, still requires manual provisioning of compute resources and is not as fully serverless as DynamoDB for this use case. Option D is wrong because Amazon RDS for PostgreSQL is a managed relational database but requires manual scaling of instance size and storage, and does not offer the same level of automatic, seamless scaling as DynamoDB for a mobile app's unpredictable workload.

80
Multi-Selecthard

Which TWO of the following are best practices for Amazon Redshift table design? (Choose TWO.)

Select 2 answers
A.Choose sort keys based on query patterns
B.Use INSERT statements for large data loads
C.Avoid compression encoding to reduce CPU overhead
D.Specify distribution keys to minimize data movement
E.Set distribution style to ALL for all tables
AnswersA, D

Sort keys determine the order in which rows are stored on disk, so choosing them from actual query filter and join patterns lets Redshift skip irrelevant blocks via zone maps, reducing I/O. This directly satisfies the best-practise requirement for table design.

Why this answer

Option A is correct because choosing sort keys based on query patterns allows Amazon Redshift to use zone maps and skip scanning irrelevant blocks, dramatically reducing I/O for range-filtered and ordered queries. Option D is correct because specifying an appropriate distribution key colocates matching rows on the same compute node slice, so joins and aggregations on that key avoid costly data redistribution across the network. Option B is wrong because large data loads should use COPY (ideally from Amazon S3) rather than INSERT, which is slow and row-by-row.

Option C is wrong because compression encoding is a best practice that reduces storage and I/O; Redshift uses columnar compression with minimal CPU penalty, so avoiding it is counterproductive. Option E is wrong because distribution style ALL replicates the entire table to every node, which is only suitable for small dimension tables and causes excessive storage and load overhead on large tables.

Exam trap

DEA-C01 often tests whether candidates confuse distribution styles (ALL vs. KEY vs. EVEN) and mistakenly believe that ALL is a safe default, when it actually causes severe storage and performance penalties for large tables.

81
Multi-Selecthard

A company is migrating a large Oracle data warehouse to Amazon Redshift. Which THREE considerations are important for optimizing the Redshift cluster?

Select 3 answers
A.Purchasing reserved instances for the cluster.
B.Using columnar storage format.
C.Defining appropriate sort keys for the tables.
D.Applying compression encoding to columns.
E.Choosing the right distribution style (KEY, ALL, EVEN).
AnswersC, D, E

Improves query performance by reducing scans.

Why this answer

Sort keys in Amazon Redshift determine the physical order of data on disk, which directly impacts the efficiency of range-restricted queries and compression. By defining appropriate sort keys (compound or interleaved), the query optimizer can use zone maps to skip large blocks of data that don't match the filter criteria, significantly reducing the number of blocks scanned and improving query performance.

Exam trap

The trap here is that candidates confuse cost-saving measures (reserved instances) or inherent architecture features (columnar storage) with active optimization choices, when in fact only sort keys, distribution styles, and compression encoding are configurable settings that directly impact query performance in Redshift.

82
MCQmedium

A data engineer is building a data lake on Amazon S3 and needs to store JSON logs that will be queried by Amazon Athena. The engineer wants to minimize query cost and improve performance by reducing the amount of data scanned. The logs are approximately 1 KB each and arrive continuously. Which solution should the engineer implement?

A.Convert the JSON logs to Apache Parquet, partition the data by date, and store it in Amazon S3. Query with Athena.
B.Store the JSON logs in Amazon S3 Standard-Infrequent Access (S3 Standard-IA) and query them with Athena.
C.Enable S3 Transfer Acceleration on the bucket and query the JSON logs with Athena.
D.Store the JSON logs in Amazon S3 Glacier Instant Retrieval and query them with Athena.
AnswerA

Converting to columnar Parquet enables Athena to read only the columns referenced in a query, and partitioning by date allows Athena to skip irrelevant partitions. Together, these reduce the amount of data scanned, lowering query cost and improving performance, which directly meets the requirement.

Why this answer

Athena charges based on the amount of data scanned per query. Converting JSON to a columnar format like Parquet allows column pruning, so only the needed columns are read, and partitioning by date enables partition pruning, so only relevant folders are scanned. This combination minimizes data scanned, reducing cost and improving query speed.

Exam trap

The trap here is assuming that changing the S3 storage class reduces Athena query costs, when Athena cost is driven by data scanned, not by where the data is stored.

83
MCQmedium

A financial analytics company stores daily transaction records in Amazon S3 as Apache Parquet files, partitioned by year/month/day. The data engineering team queries these files with Amazon Athena. To reduce query runtime and cost, they want to apply fine-grained access control and column-level filtering without changing the files. Which solution should they use?

A.Define an AWS Lake Formation table with column-level permissions and use Lake Formation to manage access for Athena users.
B.Create an AWS Glue Data Catalog table with partition projection and use Amazon S3 Access Points for authorization.
C.Enable Amazon S3 Block Public Access and use IAM policies with condition keys to filter columns.
D.Store the data in Amazon Redshift Spectrum and use Redshift database roles to restrict column access.
AnswerA

AWS Lake Formation provides fine-grained access control at the table, column, and row level for data in the AWS Glue Data Catalog. Athena integrates with Lake Formation to enforce these permissions during queries without modifying underlying data. This directly satisfies the requirement for column-level filtering and access control.

Why this answer

AWS Lake Formation is designed to provide granular access control for data lakes built on Amazon S3 and the AWS Glue Data Catalog. It allows administrators to grant or revoke permissions at the database, table, column, and row level. Athena respects these permissions, enabling column-level filtering without data duplication or transformation.

This meets the requirement to reduce runtime and cost by limiting data access.

Exam trap

The trap here is assuming that S3 Access Points or IAM policies can enforce column-level access control for Athena queries, when in fact only Lake Formation provides that granularity.

84
Multi-Selecthard

A data engineer is building a near-real-time ingestion pipeline into Amazon S3. Small JSON files arrive continuously from thousands of devices, and the engineer must optimize the data lake for downstream Amazon Athena queries while minimizing storage cost and query latency. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Store the raw JSON files in S3 Glacier Instant Retrieval to reduce storage cost immediately.
B.Configure an S3 Lifecycle rule to abort incomplete multipart uploads after 7 days.
C.Use AWS Glue ETL jobs to compact small JSON files into larger Parquet files partitioned by ingestion date.
D.Use Amazon Kinesis Data Firehose to buffer incoming records and deliver larger aggregated files to S3.
E.Enable S3 Transfer Acceleration on the ingestion bucket to speed up uploads from devices.
AnswersC, D

Compacting many small JSON files into larger Parquet files reduces the number of S3 objects and takes advantage of columnar storage, which lowers Athena scan costs and improves query latency. Partitioning by ingestion date further enables partition pruning. This directly addresses both the small-file problem and the need for efficient downstream analytics.

Why this answer

The pipeline suffers from many small JSON files, which increase Athena query overhead and cost. Compacting files into larger Parquet objects with AWS Glue ETL and using Kinesis Data Firehose to buffer and aggregate records both reduce object count and improve columnar query performance. Transfer Acceleration, Glacier Instant Retrieval, and aborting multipart uploads do not address file size or format for analytics.

Exam trap

The trap here is focusing on upload speed or storage class alone, when the real issue is the number and format of files that Athena must scan.

85
MCQhard

An IAM role 'DataLakeRole' has the above S3 bucket policy attached to an S3 bucket. The role is assumed by an AWS Glue job. The Glue job is failing with 'Access Denied' errors when trying to list objects in the bucket. Which action should be added to the policy to fix the issue?

A.Add s3:ListObjects action for the bucket ARN.
B.Add s3:ListBucket action for the bucket ARN (arn:aws:s3:::my-data-lake).
C.Add s3:GetObjectVersion action for the object ARN.
D.Add s3:ListBucket action for the object ARN (arn:aws:s3:::my-data-lake/*).
AnswerB

Adding `s3:ListBucket` on the bucket ARN (`arn:aws:s3:::my-data-lake`) satisfies the failing list operation, because bucket-level actions such as `ListBucket` are evaluated against the bucket resource itself, not the objects within it. Object ARNs (`arn:aws:s3:::my-data-lake/*`) only govern object-level actions like `GetObject`.

Why this answer

The Glue job is failing with 'Access Denied' when trying to list objects, which requires the s3:ListBucket permission on the bucket itself (not on objects). Option B correctly adds s3:ListBucket for the bucket ARN (arn:aws:s3:::my-data-lake), which grants permission to list the contents of the bucket. Without this action, even if other permissions exist, the ListObjectsV2 API call used by AWS Glue to enumerate objects will be denied.

Exam trap

The trap here is that candidates confuse s3:ListBucket (which applies to the bucket itself) with s3:GetObject or s3:ListObjects (which are often misapplied to object ARNs), leading them to pick Option D or A, not realizing that listing requires the bucket-level permission and a bucket ARN, not an object ARN.

How to eliminate wrong answers

Option A is wrong because s3:ListObjects is an alias for s3:ListBucket and must be applied to the bucket ARN, not the bucket ARN with a trailing slash or object path; however, the key issue is that the action name itself is correct but the ARN in the answer is unspecified, and the question asks for the action to add, not the ARN—but more critically, s3:ListObjects is a legacy action name and the exam expects s3:ListBucket for consistency with the S3 API. Option C is wrong because s3:GetObjectVersion is used to retrieve a specific version of an object, not to list objects, and does not address the 'list objects' failure. Option D is wrong because s3:ListBucket must be applied to the bucket ARN (arn:aws:s3:::my-data-lake), not to an object ARN (arn:aws:s3:::my-data-lake/*); applying it to an object ARN would be invalid and would not grant the permission to list the bucket's contents.

86
Multi-Selecthard

A company stores sensitive financial data in an Amazon Redshift cluster. The data engineer must ensure that all queries are logged for audit purposes and that the logs are stored in Amazon S3 with server-side encryption. Which THREE steps should the data engineer take to meet these requirements?

Select 3 answers
A.Configure audit logs to be stored in an Amazon S3 bucket.
B.Enable encryption on the Redshift cluster.
C.Enable AWS CloudTrail to log Redshift queries.
D.Enable audit logging on the Redshift cluster.
E.Enable default encryption on the S3 bucket using SSE-S3 or SSE-KMS.
AnswersA, D, E

Audit logs can be delivered to an S3 bucket.

Why this answer

Amazon Redshift audit logs can be configured to be stored directly in an Amazon S3 bucket, which is a native feature for exporting connection logs, user logs, and query logs. This satisfies the requirement to log all queries for audit purposes without relying on external services.

Exam trap

The trap here is confusing AWS CloudTrail (which logs control-plane API calls) with Redshift's native audit logging (which logs data-plane SQL queries), leading candidates to incorrectly select CloudTrail as a solution for query auditing.

87
MCQhard

A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow and the system is experiencing high disk usage. The engineer suspects that the distribution style is suboptimal. Which action should the engineer take to improve query performance?

A.Convert all tables to use SORTKEY on the most frequently filtered column.
B.Increase the number of nodes in the cluster to distribute data across more slices.
C.Use the DISTSTYLE AUTO setting and analyze query patterns to let Redshift choose.
D.Set all tables to DISTSTYLE EVEN to distribute data evenly.
AnswerC

AUTO adapts distribution based on workload.

Why this answer

DISTSTYLE AUTO allows Amazon Redshift to automatically assign distribution styles (KEY, EVEN, or ALL) based on query patterns and table size, optimizing data distribution for improved query performance. This is particularly effective when the engineer suspects suboptimal distribution but lacks detailed knowledge of the ideal key, as Redshift analyzes workload patterns to reduce data movement and disk usage.

Exam trap

The trap here is that candidates often confuse distribution style with sort key or node scaling, leading them to choose options that address symptoms (e.g., disk usage via node count) rather than the root cause of suboptimal data distribution.

How to eliminate wrong answers

Option A is wrong because SORTKEY improves query performance by reducing the amount of data scanned via block-level filtering, but it does not address distribution style or high disk usage caused by data skew. Option B is wrong because increasing the number of nodes distributes data across more slices but does not fix the root cause of suboptimal distribution; it may even exacerbate disk usage if data is already skewed. Option D is wrong because setting all tables to DISTSTYLE EVEN distributes rows evenly across slices, which can eliminate data skew but may cause excessive data movement (broadcast or redistribution) during joins, degrading query performance for tables that are frequently joined on specific columns.

88
MCQhard

A data engineer is designing a multi-Region disaster recovery solution for an Amazon DynamoDB table. The table must be available in a secondary Region with minimal data loss and automatic failover. Which feature should be used?

A.DynamoDB on-demand backup and restore in the secondary Region
B.DynamoDB global tables
C.DynamoDB point-in-time recovery (PITR)
D.DynamoDB cross-Region snapshot export to S3
AnswerB

DynamoDB global tables replicate data across Regions with active-active writes and automatic failover, meeting the secondary-Region availability and minimal-data-loss requirements. Single-Region backups or point-in-time recovery cannot provide an operational secondary Region, so they fail the stated disaster recovery constraint.

Why this answer

DynamoDB global tables provide a fully managed, multi-Region, multi-active database solution that replicates data automatically across selected AWS Regions. This ensures automatic failover with eventual consistency and minimal data loss, meeting the disaster recovery requirements for high availability and automatic failover without manual intervention.

Exam trap

The trap here is that candidates often confuse point-in-time recovery (PITR) with cross-Region disaster recovery, but PITR is a single-Region feature that does not provide automatic failover or multi-Region replication.

How to eliminate wrong answers

Option A is wrong because on-demand backup and restore is a manual process that requires user intervention to initiate a restore in the secondary Region, not providing automatic failover or minimal data loss in real time. Option C is wrong because point-in-time recovery (PITR) protects against accidental writes or deletes within a single Region by restoring to a point in time, but it does not replicate data across Regions or enable automatic failover. Option D is wrong because cross-Region snapshot export to S3 is a manual, batch-oriented process that exports table data to Amazon S3 in another Region, requiring manual import and setup for failover, and does not provide automatic, continuous replication or failover.

89
MCQmedium

A company is using an Amazon RDS for MySQL database for its e-commerce platform. During a recent flash sale, the database experienced high read traffic, causing slow query performance. The company needs a solution that offloads read traffic with minimal application changes. Which action should be taken?

A.Enable DynamoDB Accelerator (DAX) on the RDS instance.
B.Migrate the database to Amazon Aurora and enable Aurora Global Database.
C.Implement Amazon ElastiCache for Redis to cache database queries.
D.Create an Amazon RDS read replica in the same region.
AnswerD

An RDS read replica uses MySQL's native asynchronous replication to serve read-only queries from a separate endpoint, diverting SELECT traffic from the primary. This offloads read load with minimal application change, satisfying the flash-sale read-traffic constraint.

Why this answer

Creating an Amazon RDS read replica in the same region offloads read traffic from the primary DB instance by directing read queries to a read-only copy. This requires minimal application changes—only modifying the database connection string to point read queries to the replica endpoint. RDS read replicas use MySQL's native asynchronous replication, making them ideal for scaling read-heavy workloads like flash sales.

Exam trap

The trap here is that candidates may choose ElastiCache (Option C) because it is a caching solution, but they overlook the explicit requirement for minimal application changes, which caching typically does not satisfy without code modifications.

How to eliminate wrong answers

Option A is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for Amazon DynamoDB, not for RDS for MySQL; it cannot be enabled on an RDS instance. Option B is wrong because migrating to Aurora and enabling Aurora Global Database is designed for cross-region disaster recovery and global reads, not for offloading read traffic within a single region, and it requires significant application and migration effort. Option C is wrong because while ElastiCache for Redis can cache query results, it requires application code changes to implement caching logic (e.g., cache-aside pattern), which contradicts the requirement for minimal application changes.

90
Multi-Selectmedium

Which TWO actions can help optimize Amazon S3 storage costs for a data lake? (Choose two.)

Select 2 answers
A.Enable S3 Replication to another region
B.Use S3 Intelligent-Tiering for unpredictable access patterns
C.Use S3 Select to retrieve only needed data
D.Enable S3 Transfer Acceleration
E.Implement S3 Lifecycle policies to transition objects to Glacier
AnswersB, E

S3 Intelligent-Tiering automatically moves objects between frequent and infrequent access tiers based on observed access patterns, charging a small monitoring fee. This satisfies the stem's unpredictable access constraint, since lifecycle policies require known patterns to schedule transitions reliably.

Why this answer

Option B is correct because S3 Intelligent-Tiering automatically moves objects between frequent and infrequent access tiers based on changing access patterns, eliminating retrieval fees and avoiding the cost of manual tiering when access is unpredictable. Option E is correct because S3 Lifecycle policies can transition aging objects to lower-cost storage classes such as S3 Glacier Instant Retrieval, Glacier Flexible Retrieval, or Glacier Deep Archive, directly reducing storage costs for cold data. Option A is not a cost optimization for a data lake because S3 Replication to another region incurs additional storage, request, and inter-region data transfer charges.

Option C reduces data scanned and compute cost for querying, but it does not lower S3 storage costs. Option D speeds up uploads over long distances via AWS edge locations but adds a premium charge, so it increases rather than optimizes storage cost.

Exam trap

The trap here is that candidates confuse cost optimization for storage (reducing stored data cost) with cost optimization for data transfer or retrieval, leading them to select options like S3 Select or Transfer Acceleration that address different cost dimensions.

91
MCQeasy

A company uses Amazon S3 to store customer documents. The data engineer needs to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted with a customer-managed AWS KMS key. What should the data engineer do?

A.Use pre-signed URLs for all uploads that include encryption parameters.
B.Create a bucket policy that denies uploads without encryption.
C.Enable S3 Versioning on the bucket.
D.Set default encryption on the bucket to use SSE-KMS with the customer-managed key.
AnswerD

Setting default bucket encryption to SSE-KMS with the customer-managed key enforces encryption at upload without changing client behaviour, satisfying the requirement that every object be encrypted automatically. S3 applies the default to objects lacking an explicit encryption header, and SSE-KMS provides the customer-managed key control the stem demands.

Why this answer

Setting default encryption on the S3 bucket to SSE-KMS with the customer-managed key ensures that all objects uploaded without explicit encryption headers are automatically encrypted using that KMS key. This satisfies the requirement without relying on client-side behavior, as S3 applies the encryption server-side at the time of write.

Exam trap

The trap here is that candidates often confuse bucket policies that deny unencrypted uploads (which only reject non-compliant requests) with default encryption (which automatically encrypts objects), leading them to choose Option B instead of D.

How to eliminate wrong answers

Option A is wrong because pre-signed URLs only grant temporary access to upload or download objects; they do not enforce encryption on the uploaded data, and including encryption parameters in the URL is optional and client-dependent. Option B is wrong because a bucket policy that denies uploads without encryption can enforce that clients must include encryption headers, but it does not automatically encrypt objects; if the client fails to include the header, the upload is denied rather than encrypted. Option C is wrong because S3 Versioning preserves multiple versions of an object but has no effect on encryption; it does not encrypt objects or enforce encryption policies.

92
MCQhard

A data engineer is troubleshooting an Amazon Redshift cluster that is running out of disk space. The engineer runs STV_PARTITIONS and notices that some slices have significantly more data than others. What is the most likely cause and solution?

A.Poorly chosen sort keys; redefine sort keys
B.Data distribution skew due to uneven distribution style; change distribution style to EVEN or correct KEY
C.Some nodes are underutilized; add more nodes
D.Concurrency scaling is disabled; enable concurrency scaling
AnswerB

Uneven slice storage indicates distribution skew: rows cluster on certain slices under a poor KEY distribution style. Redistributing with EVEN, or choosing a higher-cardinality KEY column, spreads data evenly across slices and resolves the disk-space imbalance.

Why this answer

B is correct because STV_PARTITIONS shows per-slice disk usage, and significant variation indicates data distribution skew. Uneven distribution causes some slices to fill faster, leading to premature disk-full errors. Changing the distribution style to EVEN (for tables without join keys) or correcting the KEY distribution style (using a high-cardinality, evenly distributed column) rebalances data across slices.

Exam trap

The trap here is that candidates confuse sort keys (which improve query performance via zone maps) with distribution keys (which control data placement across slices), leading them to incorrectly select sort key redefinition as the fix for disk space skew.

How to eliminate wrong answers

Option A is wrong because sort keys affect query performance (min/max zone maps and block pruning), not how data is distributed across slices; disk space skew is a distribution issue, not a sort key issue. Option C is wrong because adding nodes increases total cluster capacity but does not fix existing data skew; the problem is uneven data placement, not insufficient total nodes. Option D is wrong because concurrency scaling handles workload bursts by adding transient compute capacity, not disk space; it does not affect how data is stored on existing slices.

93
MCQeasy

A company uses Amazon DynamoDB to store user session data. The table has a partition key of UserID and a sort key of SessionStartTime. The application frequently queries for all sessions of a specific user within a date range. The table is provisioned with 1000 RCUs and 1000 WCUs. During peak hours, the application experiences throttling on read requests. Which action should a data engineer take to resolve the throttling with minimal changes?

A.Increase the provisioned read capacity units (RCUs) for the table.
B.Create a global secondary index (GSI) on the SessionStartTime attribute.
C.Enable DynamoDB Accelerator (DAX) for the table.
D.Change the table's capacity mode to on-demand.
AnswerA

Throttling on read requests indicates insufficient read capacity. Increasing RCUs directly addresses the issue by allowing more read operations per second. Since the table is already provisioned, adjusting RCUs is a minimal change that can be done without application modifications or schema changes.

Why this answer

Read throttling occurs when the consumed read capacity exceeds the provisioned RCUs. Increasing RCUs is the most direct and minimal solution, as it provides more read capacity without requiring application changes or architectural modifications. It addresses the root cause of the throttling.

Exam trap

The trap here is assuming that a more complex solution like DAX or on-demand mode is needed, when simply increasing provisioned capacity resolves the immediate throttling issue.

94
MCQmedium

A data engineer manages a large Amazon S3 data lake with millions of small JSON files ingested daily. Amazon Athena queries against this data lake are slow and costly due to high per-query data scanned. The engineer wants to optimize the storage layout to improve query performance and reduce cost, while keeping the data queryable in place. Which solution should the engineer implement?

A.Use AWS Glue ETL to compact the small JSON files into larger Parquet files partitioned by commonly filtered columns, then update the AWS Glue Data Catalog.
B.Enable S3 Transfer Acceleration on the bucket to speed up data retrieval by Athena.
C.Convert the data to CSV format and add more columns to the AWS Glue Data Catalog table.
D.Move the data to Amazon Redshift and query it using Redshift Spectrum.
AnswerA

Compacting small files into larger Parquet files reduces the number of S3 GET requests and leverages columnar storage, which minimizes data scanned by Athena. Partitioning by frequently filtered columns enables partition pruning, further reducing data scanned. Updating the Data Catalog ensures Athena queries use the new schema and partitions. This directly addresses performance and cost without moving data out of S3.

Why this answer

Compacting small files into larger Parquet files and partitioning by frequently filtered columns directly reduces the amount of data scanned by Athena, improving performance and lowering cost. Parquet's columnar format allows Athena to read only the columns needed, and partitioning enables partition pruning. Updating the Data Catalog ensures queries use the optimized layout.

The other options do not address the core issues of small files and inefficient format.

Exam trap

The trap here is assuming that S3 Transfer Acceleration or format changes alone improve Athena query performance, when the key is reducing data scanned through compaction and partitioning.

95
MCQmedium

A company stores sensitive data in an S3 bucket. To meet compliance requirements, they must ensure that all objects are encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement should be applied to deny uploads that do not use the required encryption?

A.{"Effect":"Deny","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucketname/*","Condition":{"StringNotEquals":{"s3:x-amz-server-side-encryption":"aws:kms"}}}
B.{"Effect":"Deny","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucketname/*","Condition":{"StringNotEquals":{"s3:x-amz-server-side-encryption":"AES256"}}}
C.{"Effect":"Deny","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucketname/*","Condition":{"StringNotEquals":{"s3:x-amz-server-side-encryption-aws-kms-key-id":"arn:aws:kms:us-east-1:123456789012:key/abc123"}}}
D.{"Effect":"Deny","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucketname/*","Condition":{"Null":{"s3:x-amz-server-side-encryption":"true"}}}
AnswerA

The `StringNotEquals` condition on `s3:x-amz-server-side-encryption` denies any `PutObject` request whose encryption header is not `aws:kms`, directly enforcing the compliance requirement that all objects use SSE-KMS. Unlike `StringNotEqualsIfExists`, this operator also blocks uploads that omit the header entirely, closing the unencrypted-upload gap.

Why this answer

It uses the `s3:x-amz-server-side-encryption` condition key with `StringNotEquals` set to `aws:kms`, which denies any `s3:PutObject` request where the encryption header does not specify `aws:kms`. This ensures all uploaded objects are encrypted at rest using server-side encryption with AWS KMS, meeting the compliance requirement.

Exam trap

The trap here is that candidates often confuse the condition keys for encryption type (`s3:x-amz-server-side-encryption`) with the specific KMS key ID (`s3:x-amz-server-side-encryption-aws-kms-key-id`), or mistakenly use `Null` to check for the presence of encryption instead of enforcing the correct encryption algorithm.

How to eliminate wrong answers

Option B is wrong because it checks for `AES256`, which corresponds to SSE-S3 (Amazon S3-managed keys), not SSE-KMS (AWS KMS keys), so it would allow objects encrypted with SSE-S3 instead of enforcing KMS encryption. Option C is wrong because it uses the `s3:x-amz-server-side-encryption-aws-kms-key-id` condition key to require a specific KMS key ID, but the requirement is only to use KMS encryption, not a particular key; this would deny uploads using any other KMS key, even if they use KMS encryption. Option D is wrong because it uses the `Null` condition to deny uploads where the `s3:x-amz-server-side-encryption` header is not present (i.e., null), but it does not enforce that the encryption must be `aws:kms`; it would also allow uploads with `AES256` or other encryption values, failing to meet the KMS-specific requirement.

96
Multi-Selectmedium

A company is using Amazon S3 to store sensitive data. They need to ensure that all objects are encrypted at rest. Which combination of actions should be taken? (Choose TWO.)

Select 2 answers
A.Enable S3 Versioning on the bucket.
B.Enable MFA Delete on the bucket.
C.Configure S3 Access Points with network policies.
D.Use a bucket policy to deny PutObject requests that do not include the x-amz-server-side-encryption header.
E.Enable default encryption on the S3 bucket.
AnswersD, E

The bucket policy enforces encryption at upload time by rejecting any PutObject lacking the x-amz-server-side-encryption header, so unencrypted writes cannot succeed. This satisfies the requirement that all objects be encrypted at rest by preventing plaintext objects from ever being stored.

Why this answer

A bucket policy that denies PutObject requests lacking the `x-amz-server-side-encryption` header enforces encryption at the time of upload, ensuring that any object written without explicit encryption headers is rejected. Option E is correct because enabling default encryption on the S3 bucket automatically applies server-side encryption (SSE-S3 or SSE-KMS) to any object uploaded without specifying encryption headers, providing a fallback that covers all objects. Together, these actions ensure that every object stored in the bucket is encrypted at rest, either by explicit client request or by default bucket settings.

Exam trap

The trap here is that candidates often confuse data protection features like Versioning or MFA Delete with encryption controls, or assume that network policies (Access Points) somehow enforce encryption, when in reality only explicit bucket policies and default encryption settings directly ensure objects are encrypted at rest.

97
MCQeasy

A data engineer needs to store semi-structured JSON logs from multiple sources in a centralized data store for querying using SQL. The logs are immutable and need to be retained for 90 days. Which AWS service should be used?

A.Amazon RDS for MySQL.
B.Amazon DynamoDB.
C.Amazon S3 with Amazon Athena.
D.Amazon ElastiCache for Redis.
AnswerC

Amazon S3 stores immutable JSON objects durably at low cost, while Athena queries them in place using standard SQL through the Glue Data Catalog. This satisfies the semi-structured, SQL-queryable, 90-day retention constraints without loading data into a separate database.

Why this answer

Amazon S3 with Amazon Athena is the correct choice because S3 provides durable, cost-effective storage for immutable semi-structured JSON logs, and Athena enables serverless SQL querying directly against the data in S3 without needing to load or transform it. This combination meets the 90-day retention requirement and supports querying semi-structured data using standard SQL via Athena's built-in JSON SerDe.

Exam trap

The trap here is that candidates may choose DynamoDB for its JSON support and querying flexibility, overlooking that it is not designed for cost-effective long-term retention of immutable logs and lacks native SQL querying, while S3 with Athena directly addresses both requirements.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for MySQL is a relational database designed for structured data with predefined schemas, not optimized for storing large volumes of immutable semi-structured JSON logs, and it incurs higher costs for long-term retention. Option B is wrong because Amazon DynamoDB is a NoSQL key-value and document database that can store JSON, but it is not cost-effective for 90-day retention of immutable logs due to per-request pricing and storage costs, and it lacks native SQL querying capabilities without additional services like DynamoDB Accelerator or PartiQL. Option D is wrong because Amazon ElastiCache for Redis is an in-memory cache designed for low-latency access to transient data, not for durable, long-term storage of immutable logs, and it does not support SQL querying.

98
MCQmedium

A company has an Amazon S3 bucket with versioning enabled. They want to automatically delete noncurrent versions of objects after 30 days. Which lifecycle rule action should be used?

A.Expiration
B.NoncurrentVersionExpiration
C.NoncurrentVersionTransition
D.AbortIncompleteMultipartUpload
AnswerB

NoncurrentVersionExpiration targets noncurrent object versions specifically, permanently deleting them once the specified retention period elapses. This satisfies the stem's 30-day constraint on noncurrent versions, unlike CurrentVersionExpiration, which acts only on the active version. Versioning must be enabled, which the scenario confirms.

Why this answer

The NoncurrentVersionExpiration lifecycle action is specifically designed to remove noncurrent object versions after a specified number of days. Since versioning is enabled and the requirement is to delete noncurrent versions after 30 days, this action directly meets the goal without affecting current versions or other lifecycle aspects.

Exam trap

The trap here is confusing NoncurrentVersionExpiration with Expiration, as candidates often mistakenly apply the standard Expiration action to delete old versions, not realizing it only affects the current version.

How to eliminate wrong answers

Option A is wrong because Expiration deletes the current version of an object (or marks it for deletion in non-versioned buckets), not noncurrent versions. Option C is wrong because NoncurrentVersionTransition moves noncurrent versions to a different storage class (e.g., S3 Glacier), but does not delete them. Option D is wrong because AbortIncompleteMultipartUpload only aborts incomplete multipart uploads that are older than a specified number of days, and has no effect on existing object versions.

99
MCQeasy

A company wants to store data from thousands of IoT devices with varying data rates. The data must be stored in a schema-on-read fashion and support SQL queries. Which AWS service should be used?

A.Amazon RDS for MySQL
B.Amazon S3 with Amazon Athena
C.Amazon DynamoDB
D.Amazon Redshift
AnswerB

S3's schema-on-read storage decouples ingestion from structure, letting thousands of IoT devices write at varying rates without transformation. Athena then queries that data in place using standard SQL, satisfying both the schema-on-read and SQL query constraints without managing servers or loading into a warehouse.

Why this answer

Amazon S3 stores data in its native format (e.g., JSON, Parquet) without requiring a predefined schema, enabling schema-on-read. Amazon Athena uses Presto-based SQL to query data directly from S3, making it ideal for IoT data with varying rates and ad-hoc SQL analysis without provisioning servers.

Exam trap

The trap here is that candidates confuse schema-on-read with schema-on-write, assuming DynamoDB's flexible schema or Redshift's SQL support fits, but they miss that DynamoDB lacks native SQL and Redshift requires upfront table definitions, while Athena directly queries raw files in S3 with SQL.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for MySQL requires a fixed schema defined before writing data, which contradicts the schema-on-read requirement and cannot handle the high write throughput of thousands of IoT devices without scaling limitations. Option C is wrong because Amazon DynamoDB is a NoSQL key-value and document database that does not support SQL queries natively (it uses PartiQL with limited SQL compatibility) and is not designed for schema-on-read. Option D is wrong because Amazon Redshift is a columnar data warehouse that requires schema-on-write (tables must be defined before loading data) and is optimized for structured, batch-loaded analytics rather than streaming IoT ingestion with varying data rates.

100
Multi-Selecteasy

Which TWO of the following are features of Amazon RDS Multi-AZ deployments? (Choose 2.)

Select 2 answers
A.Read replicas in the same region for offloading read traffic.
B.Automatic failover to a standby instance in case of an AZ failure.
C.A standby instance that is not accessible for reads or writes.
D.Automatic storage scaling based on usage.
E.Synchronous replication across AWS Regions.
AnswersB, C

Multi-AZ automatically fails over to the standby in another AZ.

Why this answer

Amazon RDS Multi-AZ deployments automatically handle failover to a standby instance in a different Availability Zone when the primary instance fails or the AZ becomes unavailable. This synchronous replication ensures zero data loss and minimal downtime, with the standby instance automatically promoted to primary without manual intervention.

Exam trap

The DEA-C01 exam often tests the distinction between Multi-AZ (high availability with a passive standby) and read replicas (scaling reads with active replicas), leading candidates to incorrectly associate read offloading with Multi-AZ deployments.

101
MCQmedium

A data engineer is designing a data store for a real-time leaderboard application that requires sub-millisecond read and write latency. The leaderboard stores scores for millions of users and needs to be sorted by score. Which AWS service should the engineer use?

A.Amazon RDS for PostgreSQL with an index on score
B.Amazon DynamoDB with a global secondary index on score
C.Amazon ElastiCache for Redis with a sorted set
D.Amazon Neptune with a graph model
AnswerC

Amazon ElastiCache for Redis sorted sets store members ordered by score, giving O(log N) inserts and rank queries entirely in memory, which delivers the sub-millisecond latency the leaderboard demands. This satisfies the stem's requirement for millions of users sorted by score, unlike disk-based stores such as DynamoDB.

Why this answer

Amazon ElastiCache for Redis provides a sorted set data structure (ZADD/ZRANGE commands) that maintains elements ordered by a numeric score with O(log N) complexity for both writes and reads, enabling sub-millisecond latency for real-time leaderboard updates and queries. This is the only option purpose-built for in-memory, sorted, real-time leaderboards at scale.

Exam trap

The trap here is that candidates often choose DynamoDB (Option B) because it is a common NoSQL choice for high-performance applications, but they overlook that DynamoDB lacks a native sorted data structure and requires costly scan operations to retrieve a globally sorted leaderboard, whereas Redis sorted sets are purpose-built for this exact use case.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for PostgreSQL, even with an index on score, is a disk-based relational database that cannot guarantee sub-millisecond read/write latency under high concurrency due to disk I/O and transaction overhead. Option B is wrong because Amazon DynamoDB with a global secondary index on score does not natively maintain a globally sorted order; it requires expensive scan operations to retrieve the top scores, and write latency can exceed sub-millisecond under heavy load due to throughput limits and index propagation. Option D is wrong because Amazon Neptune is a graph database designed for traversing relationships, not for sorted, real-time score retrieval, and its query latency is not optimized for sub-millisecond leaderboard operations.

102
Multi-Selecthard

Which THREE steps are recommended for migrating an on-premises Oracle database to Amazon RDS for Oracle with minimal downtime? (Choose 3.)

Select 3 answers
A.Set up a VPN or Direct Connect between on-premises and AWS
B.Disable archiving on the source database
C.Use AWS Schema Conversion Tool (SCT) to convert the schema
D.Perform a full load migration without change data capture
E.Use AWS Database Migration Service (DMS) for ongoing replication
AnswersA, C, E

Secure connectivity is essential.

Why this answer

Establishing a VPN or Direct Connect provides a secure, private, and low-latency network connection between the on-premises environment and AWS. This is essential for minimizing downtime during a migration, as it ensures reliable and fast data transfer for both the initial full load and ongoing replication, reducing the risk of network interruptions that could extend the migration window.

Exam trap

The trap here is that candidates often think disabling archiving simplifies the migration, but they miss that CDC requires archived logs for minimal downtime, and they may also assume a full load alone is sufficient without realizing it forces a longer outage to ensure data consistency.

103
MCQmedium

A data engineer applies the bucket policy shown in the exhibit to an S3 bucket. The bucket contains sensitive data that must be encrypted at rest and accessed only over HTTPS. Which of the following statements is true?

A.The policy allows both HTTP and HTTPS access.
B.The policy allows anonymous access to list objects in the bucket.
C.The policy enforces that all PutObject requests must include the x-amz-server-side-encryption header with value AES256.
D.The policy requires the use of AWS KMS for server-side encryption.
AnswerC

The policy's `StringNotEquals` condition on `s3:x-amz-server-side-encryption` with value `AES256` denies any PutObject lacking that header, so uploads without SSE-AES256 encryption are rejected. This directly satisfies the stem's encryption-at-rest constraint, though it does not address the HTTPS-only requirement.

Why this answer

The bucket policy includes a condition that denies PutObject requests unless the `s3:x-amz-server-side-encryption` header is present and set to `AES256`. This enforces server-side encryption with S3-managed keys (SSE-S3) for all uploads, ensuring data at rest is encrypted.

Exam trap

AWS often tests the distinction between SSE-S3 (`AES256`) and SSE-KMS (`aws:kms`) in bucket policy conditions, and candidates may mistakenly think the policy requires KMS when it actually specifies AES256.

How to eliminate wrong answers

Option A is wrong because the policy includes a `Deny` statement that blocks requests when `aws:SecureTransport` is `false`, which effectively denies HTTP access and allows only HTTPS. Option B is wrong because the policy does not grant any `s3:ListBucket` permission to anonymous principals; it only denies requests that fail encryption or transport conditions, but does not allow anonymous listing. Option D is wrong because the policy requires the `x-amz-server-side-encryption` header with value `AES256`, which corresponds to SSE-S3, not AWS KMS (which would require `aws:kms`).

104
MCQmedium

A data engineer needs to keep a near-real-time copy of an Amazon DynamoDB table in Amazon S3 for analytics, capturing every item-level change with the before and after images and no impact on table write latency. Which approach meets these requirements with the LEAST operational effort?

A.Enable DynamoDB point-in-time recovery and export the recovery window to S3 on a schedule.
B.Schedule an AWS Glue job every five minutes to run a full table scan and overwrite the S3 objects.
C.Enable DynamoDB Streams on the table and write an AWS Lambda function to batch stream records into S3.
D.Use AWS Database Migration Service with change data capture from DynamoDB to Amazon S3.
AnswerC

DynamoDB Streams captures an ordered log of item-level modifications, and configuring the stream view type to include both new and old images provides the before and after values. A Lambda function subscribed to the stream can batch records into S3 with no servers to manage and no added latency to table writes, since streams are written asynchronously.

Why this answer

DynamoDB Streams records every item-level change in order and can be configured to include both the old and new images of each item. A Lambda function consuming the stream writes batched records to S3 asynchronously, so table write latency is unaffected, and the serverless subscription keeps operational effort low. The alternative approaches are periodic, snapshot-based, or heavier to operate.

Exam trap

The trap here is treating point-in-time recovery or scheduled scans as change data capture, when only a stream carries every item modification with before and after images in near real time.

105
MCQeasy

A company stores its application logs in an Amazon S3 bucket. The logs are accessed frequently for the first 30 days, after which they are rarely accessed but must be retained for 7 years for compliance. The company wants to optimize storage costs while maintaining immediate retrieval availability for the first 30 days and the ability to retrieve logs within 12 hours after that. Which lifecycle policy should the data engineer configure?

A.Delete objects after 30 days to minimize storage costs.
B.Transition objects to S3 Standard-IA after 30 days and then to S3 Glacier Deep Archive after 1 year.
C.Transition objects to S3 One Zone-IA after 30 days and delete after 7 years.
D.Transition objects to S3 Glacier Flexible Retrieval after 30 days and delete after 7 years.
AnswerB

Standard-IA provides immediate retrieval for the first 30 days, then Deep Archive for cost-effective long-term retention.

Why this answer

It uses S3 Standard-IA for the first 30 days (frequent access, immediate retrieval) and then transitions to S3 Glacier Deep Archive after 1 year, which provides retrieval within 12 hours at the lowest cost for long-term retention. This meets the compliance requirement of 7-year retention while optimizing costs by moving data to progressively cheaper storage classes based on access patterns.

Exam trap

The DEA-C01 exam often tests the misconception that S3 Glacier Flexible Retrieval is the cheapest option for long-term archival, but S3 Glacier Deep Archive is significantly cheaper for data that is rarely accessed and can tolerate a 12-hour retrieval time.

How to eliminate wrong answers

Option A is wrong because deleting objects after 30 days violates the 7-year compliance retention requirement. Option C is wrong because S3 One Zone-IA does not provide the durability (99.999999999% vs 99.9999999999%) or availability needed for compliance data, and it lacks the 12-hour retrieval capability required after 30 days. Option D is wrong because S3 Glacier Flexible Retrieval has a retrieval time of minutes to hours (typically 1-5 minutes for expedited, 3-5 hours for standard), but the requirement is for retrieval within 12 hours, which is met; however, transitioning directly to Glacier Flexible Retrieval after 30 days is more expensive than using Standard-IA first, and the option does not include a transition to Deep Archive for further cost optimization over 7 years.

106
MCQhard

A data engineer is designing a real-time analytics solution using Amazon DynamoDB. The workload requires capturing all changes to a DynamoDB table and processing them in near-real-time to update a materialized view in Amazon Redshift. Which approach should the engineer use to capture and process the changes?

A.Enable DynamoDB global tables and use Amazon Kinesis Data Streams to replicate changes to Redshift.
B.Use AWS Database Migration Service (AWS DMS) with ongoing replication from DynamoDB to Redshift.
C.Use DynamoDB Accelerator (DAX) to cache changes and periodically export to Redshift.
D.Enable DynamoDB Streams and use AWS Lambda to process the stream and write to Redshift.
AnswerD

DynamoDB Streams captures item-level changes in near-real-time. AWS Lambda can be triggered by the stream to process records and write to Redshift. This serverless approach is scalable and requires minimal operational overhead. It is the recommended pattern for real-time change data capture from DynamoDB.

Why this answer

DynamoDB Streams captures item-level changes in near-real-time, and AWS Lambda can process these changes and write to Amazon Redshift. This serverless architecture is scalable, cost-effective, and requires minimal operational overhead. It is the standard pattern for change data capture from DynamoDB to other data stores.

Exam trap

The trap here is confusing DynamoDB Streams with other features like global tables or DAX, which serve different purposes such as multi-region replication or caching, and do not provide change data capture for real-time processing.

107
MCQmedium

A financial services company stores transaction records in an Amazon DynamoDB table. An audit requires that all data older than 7 years be automatically and permanently deleted. The data engineering team must implement this with minimal operational overhead and no application code changes. What should the team do?

A.Configure a TTL attribute on the table with an expiry timestamp set to 7 years from the transaction date.
B.Enable point-in-time recovery and restore the table to a point before the 7-year window, then delete the original table.
C.Create a scheduled Amazon EventBridge rule that invokes an AWS Step Functions state machine to scan and delete old items.
D.Enable DynamoDB Streams and write an AWS Lambda function that deletes items older than 7 years.
AnswerA

DynamoDB Time to Live (TTL) lets you define an attribute holding an expiration timestamp; DynamoDB automatically deletes expired items at no extra cost, with no code or infrastructure to manage. Setting the attribute to transaction date plus 7 years satisfies the audit requirement and requires only a one-time schema and write-path change, not ongoing operations.

Why this answer

DynamoDB TTL is the native, serverless mechanism for automatic item expiration. By storing an epoch timestamp attribute set to the transaction date plus seven years, DynamoDB deletes expired items in the background without consuming write capacity or requiring custom code. It directly satisfies the audit's permanent deletion requirement with minimal operational effort, unlike stream-based, step-function, or restore-based approaches.

Exam trap

The trap here is assuming that DynamoDB Streams or point-in-time recovery can enforce time-based retention, when only TTL provides automatic, attribute-driven item expiration.

108
Multi-Selectmedium

A data engineer is building an Amazon Redshift data warehouse that ingests large staged files from Amazon S3 using the COPY command. The team wants to maximize load performance and minimize the time spent on ingestion. Which TWO practices should the engineer apply? (Choose two.)

Select 2 answers
A.Load compressed columnar files such as Parquet in a single COPY statement
B.Run VACUUM SORT ONLY on the target table immediately after each COPY
C.Use a single large compressed file to reduce the number of S3 GET requests
D.Split the input data into multiple files sized roughly equal and load them in parallel
E.Add a sort key to the target table on the timestamp column before loading
AnswersA, D

COPY supports columnar formats like Parquet and ORC, and columnar compression reduces the bytes transferred from S3 while allowing column-level pruning during load. Loading multiple Parquet files in one COPY statement lets Redshift distribute them across slices in parallel. Combining columnar compression with parallel file distribution reduces I/O and CPU work, directly improving ingestion throughput and shortening load windows.

Why this answer

COPY performance in Redshift depends on parallelism and data volume. Splitting input into multiple similarly sized files lets every slice read concurrently, and using compressed columnar formats such as Parquet reduces bytes moved and enables column pruning. Together these practices shorten load time.

Sort keys, single-file loads, and post-load VACUUM operations affect query performance or add overhead rather than accelerating ingestion, so they do not belong in this optimization.

Exam trap

The trap here is equating query-performance tuning such as sort keys and VACUUM with load-performance tuning, when COPY speed is driven primarily by parallel file distribution and compressed columnar formats.

109
MCQeasy

A data engineer is migrating an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB in size. The engineer needs to minimize downtime. Which AWS service should be used for the migration?

A.AWS Data Pipeline
B.AWS Database Migration Service (DMS)
C.AWS Snowball
D.Amazon S3
AnswerB

DMS supports continuous replication with minimal downtime.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it supports continuous replication (change data capture) from an on-premises PostgreSQL source to Amazon RDS for PostgreSQL, enabling near-zero downtime migration. DMS can handle a 2 TB database by using a large replication instance and tuning task settings, and it automatically converts the source schema to the target RDS engine.

Exam trap

The trap here is that candidates often choose AWS Snowball for large databases, mistakenly thinking physical transfer is faster, but they overlook that Snowball requires stopping writes to the source database during the export and shipping process, causing unacceptable downtime for a live migration.

How to eliminate wrong answers

Option A is wrong because AWS Data Pipeline is a batch-oriented workflow orchestration service for moving and transforming data between AWS services, but it does not support live, ongoing replication or schema conversion for database migrations, making it unsuitable for minimizing downtime. Option C is wrong because AWS Snowball is a physical data transfer device designed for large-scale offline data movement (e.g., petabyte-scale), but it introduces significant downtime due to shipping and manual transfer, and it cannot perform continuous replication for a live migration. Option D is wrong because Amazon S3 is an object storage service and cannot directly migrate a live PostgreSQL database to RDS; it would require an intermediate export/import process that causes extended downtime and lacks native change data capture.

110
Multi-Selecthard

A data engineer is optimizing an Amazon Redshift cluster for a workload that includes frequent complex queries with multiple joins and aggregations. The engineer wants to improve query performance by using appropriate distribution styles and sort keys. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Create a materialized view for each complex query.
B.Set the distribution style of small dimension tables to ALL.
C.Use EVEN distribution for all tables to ensure uniform data distribution.
D.Apply a sort key on the column used in the WHERE clause of frequent queries.
E.Set the distribution style of large fact tables to KEY on the join column.
AnswersB, E

ALL distribution replicates small dimension tables to every node, eliminating the need to redistribute data during joins. This is effective for small tables that are frequently joined with large fact tables. It reduces network overhead and improves join performance, but should only be used for tables that are small enough to fit comfortably in node memory and storage.

Why this answer

Setting the distribution style of large fact tables to KEY on the join column collocates matching rows, reducing data movement during joins. Setting small dimension tables to ALL replicates them to all nodes, eliminating redistribution. Together, these actions minimize network traffic and accelerate complex join queries.

They are fundamental best practices for optimizing Amazon Redshift performance in join-heavy workloads.

Exam trap

The trap here is assuming that sort keys or materialized views are the primary optimizations for join performance, when distribution styles have a more direct impact on data movement during joins.

111
MCQhard

A company has an Amazon Redshift cluster with a mix of frequently accessed hot data and rarely accessed cold data. They want to reduce storage costs without affecting query performance for the hot data. Which strategy is MOST effective?

A.Use RA3 nodes with managed storage to automatically offload cold data to Amazon S3.
B.Reduce the number of nodes and increase the number of slices.
C.Create external tables in Redshift Spectrum to query cold data in S3.
D.Use Dense Compute nodes and unload cold data to Amazon S3 manually.
AnswerA

RA3 nodes separate compute from Redshift Managed Storage, automatically moving cold blocks to Amazon S3 while hot data stays on local SSD. This lowers storage cost without degrading hot-data query performance, and needs no manual data movement.

Why this answer

RA3 nodes with managed storage automatically separate compute and storage, offloading cold data to Amazon S3 while keeping hot data on local SSD for fast queries. This reduces storage costs without manual intervention or affecting hot data performance.

Exam trap

The trap here is that candidates may choose Redshift Spectrum (Option C) thinking it automatically offloads cold data, but Spectrum requires manual external table creation and does not integrate with the cluster's automatic storage tiering.

How to eliminate wrong answers

Option B is wrong because reducing nodes and increasing slices does not address cold data storage; it changes cluster configuration without reducing storage costs for cold data. Option C is wrong because creating external tables in Redshift Spectrum allows querying cold data in S3 but does not automatically offload cold data from the cluster; it requires manual data movement and schema management. Option D is wrong because Dense Compute nodes are compute-optimized and do not support managed storage offloading; manually unloading cold data to S3 adds operational overhead and does not leverage automatic tiering.

112
MCQhard

A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The workload includes frequent queries that retrieve all sessions for a user within a specific time range. The engineer notices that some queries are slow and wants to optimize the table design. Which action should the engineer take to improve query performance?

A.Ensure that queries use the partition key and sort key condition to retrieve data efficiently, and consider enabling adaptive capacity.
B.Increase the provisioned read capacity units (RCUs) for the table to handle the query load.
C.Create a global secondary index (GSI) with session_start_time as the partition key and user_id as the sort key.
D.Enable DynamoDB Streams to capture changes and use AWS Lambda to maintain a separate queryable table.
AnswerA

The base table's key schema (user_id as partition key, session_start_time as sort key) is ideal for the query pattern. Using both keys in queries allows DynamoDB to efficiently locate the relevant items. Enabling adaptive capacity helps handle uneven access patterns by automatically adjusting throughput. This approach directly addresses performance by leveraging the existing design and DynamoDB's built-in optimizations.

Why this answer

The correct action is to ensure queries use the partition key and sort key condition to retrieve data efficiently, and to consider enabling adaptive capacity. The table's key schema is already optimized for the query pattern, so leveraging it properly is key. Other options introduce unnecessary complexity, do not address the query pattern, or simply scale capacity without fixing efficiency.

Exam trap

The trap here is assuming that adding a GSI or increasing capacity will solve slow queries, when the base table design is already correct and the issue may be query patterns.

113
MCQhard

A company uses Amazon Redshift for a data warehouse. They notice that queries are slow due to heavy data skew. Which optimization technique should be applied first?

A.Configure workload management (WLM) queues
B.Define sort keys on frequently filtered columns
C.Set an appropriate distribution style
D.Apply compression encodings to columns
AnswerC

Data skew concentrates rows on some slices, so those slices do redundant work during joins and aggregations. Choosing a distribution style that spreads rows evenly, such as KEY on a high-cardinality column or ALL for small tables, addresses the root cause first.

Why this answer

Data skew occurs when rows are distributed unevenly across Redshift slices, causing some nodes to process far more data than others. Setting an appropriate distribution style (e.g., KEY, EVEN, or ALL) redistributes the data to balance the workload, directly addressing the root cause of the slowness. This is the first optimization to apply because skew is a fundamental distribution issue that other tuning steps cannot fix.

Exam trap

The trap here is that candidates often confuse distribution skew with sort key optimization or compression, mistakenly believing that improving data organization on disk (sort keys) or reducing I/O (compression) will fix uneven data distribution across nodes.

How to eliminate wrong answers

Option A is wrong because WLM queues manage concurrency and memory allocation for query slots, not the physical distribution of data across nodes; they cannot fix performance degradation caused by data skew. Option B is wrong because sort keys optimize the order of data on disk to improve range-restricted scans and merge joins, but they do not redistribute data or alleviate skew across slices. Option D is wrong because compression encodings reduce storage footprint and I/O by compressing column data, but they have no effect on how rows are distributed across nodes or on query parallelism.

114
MCQmedium

A data engineer is designing a data lake on Amazon S3 to store JSON logs from an application. The logs are written once and never modified. The engineer needs to query the data using Amazon Athena with the best performance and lowest cost. The engineer wants to partition the data by year, month, and day based on the log timestamp. Which approach should the engineer use to organize the S3 objects?

A.Enable S3 Inventory to generate a manifest of objects and use it to create a partitioned table in Athena.
B.Use a Hive-style prefix structure such as s3://bucket/year=YYYY/month=MM/day=DD/ and define partitions in the AWS Glue Data Catalog.
C.Store objects in a flat structure and create an AWS Glue crawler to automatically add partitions based on file metadata.
D.Store all objects in a single prefix and create a table with a partition projection based on the timestamp column.
AnswerB

This approach creates a hierarchical prefix that Athena and AWS Glue recognize as partitions. When queries filter on year, month, or day, Athena prunes irrelevant partitions, scanning less data and lowering cost. It is the standard, efficient way to organize time-series data in S3 for query engines like Athena and Redshift Spectrum.

Why this answer

Organizing data in a Hive-style prefix hierarchy allows Athena and AWS Glue to recognize partitions. When queries filter on partition columns, only relevant prefixes are scanned, reducing data scanned and cost. This is the recommended practice for time-series data in S3.

Other options either do not create real partitions or fail to provide efficient pruning.

Exam trap

The trap here is assuming that any S3 organization can be partitioned by simply defining a table schema, without the physical prefix structure that enables partition pruning.

115
MCQmedium

A data engineer needs to transfer 10 TB of data from an on-premises Hadoop cluster to Amazon S3. The network bandwidth is limited to 100 Mbps, and the transfer must be completed within 48 hours. Which solution meets the requirements?

A.Use AWS DataSync to transfer data online
B.Use AWS Snowball Edge device to transfer data offline
C.Use S3 Transfer Acceleration over the internet
D.Set up AWS Direct Connect to increase bandwidth
AnswerB

At 100 Mbps, transferring 10 TB over the network would take roughly ten days, far exceeding the 48-hour deadline. AWS Snowball Edge ships the data physically, sidestepping the bandwidth constraint entirely and meeting the required completion window.

Why this answer

The on-premises Hadoop cluster has 10 TB of data to transfer, but the network bandwidth is only 100 Mbps. At 100 Mbps, the theoretical maximum transfer rate is about 12.5 MB/s, which would take approximately 10 TB / 12.5 MB/s ≈ 800,000 seconds ≈ 222 hours — far exceeding the 48-hour window. AWS Snowball Edge is an offline, physical device that bypasses network constraints entirely, allowing you to transfer the data by shipping the device, which completes within days regardless of bandwidth.

Exam trap

The trap here is that candidates may assume S3 Transfer Acceleration or Direct Connect can magically overcome a hard bandwidth cap, but neither increases the last-mile bandwidth; the only way to transfer 10 TB in under 48 hours with a 100 Mbps link is to use an offline physical device like Snowball Edge.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is an online data transfer service that still relies on network bandwidth; at 100 Mbps, it cannot transfer 10 TB within 48 hours due to the same bandwidth limitation. Option C is wrong because S3 Transfer Acceleration only optimizes routing over the internet using AWS edge locations, but it does not increase the underlying 100 Mbps bandwidth; the transfer would still take far longer than 48 hours. Option D is wrong because AWS Direct Connect provides a dedicated network connection, but it does not inherently increase bandwidth beyond the 100 Mbps limit unless you provision a higher-capacity circuit, which is not specified and would still require time to set up; the question assumes the bandwidth is fixed at 100 Mbps.

116
Multi-Selecthard

Which THREE of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose three.)

Select 3 answers
A.Offloads read traffic from the DynamoDB table.
B.Improves write throughput by batching writes.
C.Reduces read latency from single-digit milliseconds to microseconds.
D.Supports write-through caching to improve write performance.
E.Provides in-memory caching for DynamoDB tables.
AnswersA, C, E

DAX handles read requests, reducing load on the table.

Why this answer

DAX acts as a read-through cache that offloads read traffic from the DynamoDB table, reducing the number of read requests that hit the underlying table and thus lowering the consumed read capacity units (RCUs). This allows the table to handle more concurrent reads without scaling up provisioned capacity.

Exam trap

The trap here is that candidates often assume DAX improves write performance or supports write-through caching, but DAX is strictly a read cache and does not accelerate or batch writes.

117
MCQeasy

A company wants to migrate its on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which AWS service should be used for the migration?

A.AWS Database Migration Service (DMS)
B.AWS Schema Conversion Tool (SCT)
C.AWS DataSync
D.AWS Direct Connect
AnswerA

AWS Database Migration Service performs ongoing replication from the on-premises MySQL source to Amazon RDS for MySQL, keeping both in sync until cutover. This continuous change data capture satisfies the minimal-downtime constraint, unlike a one-off dump and restore, which would require halting writes for the full load duration.

Why this answer

AWS Database Migration Service (DMS) is purpose-built for migrating databases to AWS with minimal downtime by using ongoing replication (change data capture, CDC) from the source MySQL database to the target Amazon RDS for MySQL instance. This allows the source to remain fully operational during the migration, meeting the minimal-downtime requirement.

Exam trap

The trap here is that candidates confuse AWS DMS with AWS DataSync or SCT, assuming any data transfer tool works for database migration, but DMS is the only service that supports ongoing replication for minimal downtime database migrations.

How to eliminate wrong answers

Option B (AWS Schema Conversion Tool) is wrong because SCT is used for converting database schemas from one engine to another (e.g., Oracle to Aurora), not for migrating data with minimal downtime; it does not handle ongoing replication. Option C (AWS DataSync) is wrong because DataSync is designed for moving large volumes of file data (e.g., NFS, SMB) to Amazon S3 or EFS, not for database migrations or CDC replication. Option D (AWS Direct Connect) is wrong because Direct Connect establishes a dedicated network connection between on-premises and AWS, but it is a connectivity service, not a migration tool; it does not perform data migration or replication.

118
MCQhard

A data engineer is migrating an on-premises Apache HBase workload to Amazon DynamoDB. The HBase table has a row key with composite structure: customer_id (10 chars) + timestamp (10 digits). The access pattern is to query by customer_id and retrieve the latest entries. How should the DynamoDB table be designed to optimize performance?

A.Create a table with partition key = customer_id and sort key = timestamp.
B.Use Amazon S3 with customer_id as prefix and timestamp as object name.
C.Create a table with partition key = concatenated customer_id and timestamp.
D.Create a table with partition key = timestamp and sort key = customer_id.
AnswerA

Splitting the composite row key into partition key customer_id and sort key timestamp lets DynamoDB distribute items across partitions by customer, while the sort key orders entries chronologically. Querying a single customer_id with ScanIndexForward=false returns the latest entries efficiently, satisfying the query-by-customer, retrieve-latest access pattern without full table scans.

Why this answer

DynamoDB's partition key (customer_id) evenly distributes data across partitions, while the sort key (timestamp) enables efficient range queries using Query with ScanIndexForward=false to retrieve the latest entries. This design directly maps the HBase composite row key pattern to DynamoDB's primary key structure, optimizing for the described access pattern.

Exam trap

The trap here is that candidates may think concatenating the row key into a single partition key (Option C) preserves the query pattern, but DynamoDB requires the partition key to be known exactly for queries, making it impossible to query by customer_id alone without a full scan.

How to eliminate wrong answers

Option B is wrong because Amazon S3 is an object store, not a low-latency NoSQL database; it lacks native support for range queries and cannot efficiently retrieve the latest entries by timestamp without scanning all objects. Option C is wrong because using a concatenated partition key (customer_id + timestamp) prevents querying by customer_id alone, as DynamoDB requires the exact partition key value for queries, forcing a full scan. Option D is wrong because using timestamp as the partition key leads to hot partitions (e.g., all writes for the same second hit one partition) and does not allow efficient retrieval by customer_id without a scan.

119
MCQhard

A data engineer is reviewing an IAM policy that controls access to an S3 bucket. The policy is attached to a user group. The policy includes a condition that explicitly requires server-side encryption with SSE-S3 for all GetObject requests. The engineer notices that users are unable to download objects from the bucket. What is the likely cause?

A.The policy is attached to a user group instead of an IAM role.
B.The policy does not specify the correct bucket ARN.
C.The policy does not allow the s3:GetObject action.
D.The objects are encrypted using SSE-KMS, not SSE-S3.
AnswerD

The condition requires SSE-S3 (AES256), so SSE-KMS objects are denied.

Why this answer

The IAM policy includes a condition that allows downloads only if the object is encrypted with SSE-S3. However, the objects in the bucket are encrypted using SSE-KMS, which does not satisfy the condition. As a result, the s3:GetObject request is denied.

This is a common scenario where a specific encryption condition in the policy blocks access when the actual encryption method differs. The other options are less likely: attaching the policy to a user group is valid and does not cause download issues; an incorrect bucket ARN would affect all operations, not just downloads; and missing s3:GetObject would be a straightforward policy error that would be easily identified.

Exam trap

The trap here is that candidates often focus only on S3 actions (like `s3:GetObject`) and overlook the required KMS permissions when SSE-KMS is involved, assuming SSE-S3 or no encryption is the default.

How to eliminate wrong answers

Option A is wrong because attaching a policy to a user group is a valid and common practice for granting permissions to multiple users; the issue is not about the attachment target but the permissions themselves. Option B is wrong because an incorrect bucket ARN would typically cause all actions to fail, not just downloads, and the question implies other operations might work. Option C is wrong because if the policy did not allow `s3:GetObject`, users would likely receive an Access Denied error for any read operation, but the question specifically mentions download failures, which can occur even with `s3:GetObject` allowed if KMS decrypt is missing.

120
MCQhard

A company has an Amazon Redshift cluster that stores petabytes of data. Queries are experiencing high disk usage due to large intermediate results. The data engineer needs to improve query performance without adding more nodes. Which action should the engineer take?

A.Set appropriate distribution keys to minimize data movement.
B.Configure workload management (WLM) queues to limit concurrency.
C.Apply column compression encoding to reduce data size.
D.Define sort keys on all columns used in WHERE clauses.
AnswerA

Distribution keys control which node stores each row, so choosing them well colocates joined rows and reduces data redistribution across nodes. That cuts the disk spill from large intermediate results without adding nodes, directly addressing the stated constraint.

Why this answer

In Amazon Redshift, distribution keys determine how table rows are spread across compute nodes. When joined tables share the same distribution key (or use ALL distribution for small dimension tables), matching rows are co-located on the same node slice, eliminating the broadcast or redistribution steps that spill large intermediate result sets to disk. Since the question specifies high disk usage from large intermediate results and prohibits adding nodes, fixing distribution keys directly reduces the data movement that generates those spills.

Exam trap

The trap here is conflating storage/scan optimizations (compression, sort keys) with join-time data movement; candidates pick compression or sort keys because they sound like general performance fixes, but only distribution keys address the intermediate-result disk spill caused by row redistribution.

How to eliminate wrong answers

Option B is wrong because WLM queues only manage concurrency and memory allocation for query slots; they throttle or queue queries but do not reduce the intermediate result size or disk spill caused by poor data co-location. Option C is wrong because column compression encoding reduces on-disk storage footprint and I/O for scans, but it does not address the redistribution/broadcast of rows during joins that produces large intermediate results. Option D is wrong because sort keys optimize range-restricted scans and merge joins by enabling zone-map block skipping, but defining sort keys on every WHERE column is impractical (Redshift allows only one sort key per table) and does not fix the data-movement problem driving disk usage.

121
MCQhard

A company runs a real-time analytics platform using Amazon Kinesis Data Streams with a shard count of 10. The data is consumed by an AWS Lambda function that writes to an Amazon DynamoDB table. The DynamoDB table has a partition key of 'user_id' and a sort key of 'timestamp'. The table is provisioned with 5000 RCUs and 5000 WCUs. Recently, the application experienced increased write latency and throttling errors (ProvisionedThroughputExceededException) on the DynamoDB table. The CloudWatch metrics show that ConsumedWriteCapacityUnits averages 4500 with occasional spikes to 6000. The Lambda function’s concurrency is set to 1000. The data engineer suspects the issue is due to hot partitions. Upon investigation, the engineer finds that a small number of users generate a disproportionately large amount of data. Which course of action would best resolve the throttling while minimizing cost?

A.Enable DynamoDB adaptive capacity and implement write sharding by adding a suffix to the partition key for high-volume users
B.Increase the provisioned WCUs to 10000 to handle spikes
C.Switch the DynamoDB table to on-demand capacity mode
D.Reduce the Lambda function concurrency to 100 to limit write requests
AnswerA

Write sharding splits a hot user's items across multiple synthetic partitions, letting DynamoDB spread the load rather than concentrating it on one partition key. Adaptive capacity alone only isolates frequently accessed items; it cannot fix a single partition key exceeding its per-partition write limit. Together they resolve the ProvisionedThroughputExceededException spikes without raising the 5000 WCUs, minimising cost.

Why this answer

The root cause is hot partitions caused by a small number of high-volume users. Enabling DynamoDB adaptive capacity allows the table to automatically adjust throughput to accommodate uneven access patterns, but the key fix is write sharding — adding a random or calculated suffix to the partition key for those high-volume users. This distributes writes across multiple physical partitions, eliminating the hot spot without requiring a global increase in provisioned capacity, thus resolving throttling while minimizing cost.

Exam trap

AWS often tests the misconception that throttling is always solved by increasing total provisioned capacity or switching to on-demand, when the real issue is partition-level hot spots that require key design changes like write sharding.

How to eliminate wrong answers

Option B is wrong because simply increasing provisioned WCUs to 10000 does not address the hot partition issue; the throttling is due to uneven distribution of writes across partitions, not a lack of total capacity, so this would waste money without fixing the root cause. Option C is wrong because switching to on-demand capacity mode would handle spikes but at a significantly higher cost for sustained high write volumes, and it still does not solve the hot partition problem — on-demand tables can still throttle individual partitions if a single partition exceeds 1,000 WCUs (the per-partition throughput limit). Option D is wrong because reducing Lambda concurrency to 100 would limit the total write throughput, potentially causing data backlogs in Kinesis, and it does not address the uneven distribution of writes across DynamoDB partitions; the hot partition would still be throttled even with fewer concurrent writers.

122
MCQeasy

A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB in size. The migration must have minimal downtime and must be completed within a week. Which AWS service should the data engineer use to perform the migration?

A.Amazon S3 Transfer Acceleration.
B.AWS Snowball Edge.
C.AWS DataSync.
D.AWS Database Migration Service (DMS).
AnswerD

AWS DMS performs continuous replication from the on-premises MySQL source to RDS for MySQL, keeping the target current until cutover. This satisfies the minimal-downtime constraint while completing the 500 GB migration well within a week.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it is specifically designed for migrating databases to AWS with minimal downtime. DMS can perform a full load of the 500 GB MySQL database and then continuously replicate ongoing changes from the on-premises source to the Amazon RDS target using change data capture (CDC), allowing the migration to complete within a week with near-zero downtime.

Exam trap

The trap here is that candidates may confuse AWS DataSync (a file-transfer service) with database migration, or assume Snowball Edge is required for any migration over a few hundred gigabytes, when in fact DMS with CDC is the appropriate service for minimizing downtime in database migrations even with moderate data sizes.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Transfer Acceleration is a feature for speeding up uploads to S3 buckets over long distances using optimized network paths and edge locations; it is not a database migration service and cannot handle schema conversion or ongoing replication. Option B is wrong because AWS Snowball Edge is a physical data transport device intended for large-scale data transfers (typically tens of TB or more) when network bandwidth is limited; for a 500 GB database with a one-week timeline and minimal downtime requirement, the network transfer is feasible and Snowball Edge would introduce unnecessary latency and downtime for shipping. Option C is wrong because AWS DataSync is designed for moving large amounts of file data (e.g., NFS, SMB shares) to or from Amazon S3, Amazon EFS, or Amazon FSx; it does not support database engines like MySQL and cannot perform schema conversion or ongoing replication of transactional changes.

123
MCQmedium

A company runs a critical application on Amazon RDS for MySQL. To ensure high availability and automatic failover, the database is deployed as a Multi-AZ DB instance. The application uses read-heavy workloads. Which additional configuration should be used to offload read traffic without impacting write performance?

A.Use the Multi-AZ standby instance for read queries.
B.Create one or more Read Replicas in different AZs.
C.Use Amazon ElastiCache to cache read queries.
D.Change to a Single-AZ deployment with a larger instance size.
AnswerB

Read Replicas asynchronously replicate from the primary and serve read-only traffic, offloading SELECT queries while writes continue unaffected on the Multi-AZ primary. Placing them in different AZs also improves resilience, complementing the existing automatic failover configuration.

Why this answer

Amazon RDS Read Replicas are designed to offload read traffic from the primary DB instance without affecting write performance. Unlike the Multi-AZ standby instance, which is not accessible for reads, Read Replicas can be placed in different Availability Zones and serve read queries independently, improving read scalability while the primary handles writes.

Exam trap

The trap here is that candidates often assume the Multi-AZ standby instance can be used for read queries, but AWS explicitly prevents this to ensure failover consistency, making Read Replicas the correct choice for offloading reads.

How to eliminate wrong answers

Option A is wrong because the Multi-AZ standby instance is a passive replica that cannot serve read traffic; it only provides automatic failover and is not accessible for read queries. Option C is wrong because while Amazon ElastiCache can cache read queries, it does not offload read traffic from the database itself—it only reduces the number of queries hitting the database by caching results, but the question specifically asks for offloading read traffic without impacting write performance, and Read Replicas directly address this by providing additional database endpoints for reads. Option D is wrong because switching to a Single-AZ deployment with a larger instance size eliminates high availability and does not offload read traffic; it only increases capacity for both reads and writes, which can still impact write performance under heavy read load.

124
Multi-Selecthard

A data engineer manages an Amazon Redshift provisioned cluster that serves a nightly ELT workload. The cluster's largest fact table is loaded with new rows each night, and queries frequently filter on a date column and join to a customer dimension. The engineer wants to improve query performance and reduce the time spent vacuuming. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Enable automatic compression on all columns and disable the sort key to reduce storage footprint.
B.Convert the fact table to a view over the raw staging table to avoid storing duplicate data.
C.Set the customer join column as the distribution key so matching rows of the fact and dimension tables co-locate on the same slice.
D.Increase the cluster's number of nodes without changing the table design or sort keys.
E.Define the date column as the sort key so range-filtered queries skip irrelevant blocks during scans.
AnswersC, E

Distributing both the fact table and the customer dimension on the join column places matching rows on the same slice, enabling collocated joins that avoid broadcasting or shuffling data across the network. This lowers join cost and improves overall query performance for the frequent fact-to-dimension joins in this workload.

Why this answer

Performance and vacuum efficiency in Redshift are driven primarily by table design. Sorting on the frequently filtered date column enables zone-map pruning and makes VACUUM's work easier, while distributing the fact table and customer dimension on the join column enables collocated joins that avoid network-heavy shuffles. Together they reduce scanned blocks and join cost without adding hardware.

Exam trap

The trap here is believing that adding cluster capacity or compression alone fixes slow filtered joins, when the decisive levers are the sort key and distribution key choices.

125
MCQhard

Refer to the exhibit. A data engineer ran the CLI command to check the configuration of an RDS instance named 'mydb'. Which statement accurately describes the current configuration?

A.The database is in the 'stopped' state
B.The database is a Single-AZ deployment and is not a read replica
C.The database is a read replica of another instance
D.The database is a Multi-AZ deployment
AnswerB

The CLI output shows no MultiAZ or ReadReplica fields, confirming a standalone Single-AZ instance. This satisfies the stem's requirement to describe the current configuration accurately: the database runs in one Availability Zone with no read replica, so it offers no automatic failover or read scaling.

Why this answer

The CLI command output shows 'DBInstanceStatus: available' and 'MultiAZ: False', indicating the database is running as a Single-AZ deployment. Additionally, the absence of a 'SourceDBInstanceIdentifier' field confirms it is not a read replica. The 'ReadReplicaSourceDBInstanceIdentifier' is not present, which would be required if it were a read replica.

Exam trap

AWS often tests the misconception that a database with 'available' status must be a read replica or Multi-AZ, but the absence of the 'ReadReplicaSourceDBInstanceIdentifier' and 'MultiAZ: False' clearly identify it as a standalone Single-AZ instance.

How to eliminate wrong answers

Option A is wrong because the CLI output shows 'DBInstanceStatus: available', not 'stopped', so the database is running, not stopped. Option C is wrong because the output lacks a 'ReadReplicaSourceDBInstanceIdentifier' field, which is mandatory for a read replica; without it, the instance is a primary database, not a replica. Option D is wrong because the output explicitly shows 'MultiAZ: False', which means it is a Single-AZ deployment, not Multi-AZ.

126
MCQeasy

A data engineer needs to store a large volume of time-series data from IoT sensors. The data will be queried by timestamp and sensor ID, and the engineer wants to use a managed AWS database that can handle high write throughput and provide fast queries on recent data. The engineer also wants to automatically expire old data after 90 days to reduce storage costs. Which AWS service should the engineer use?

A.Amazon DynamoDB with TTL
B.Amazon Timestream
C.Amazon RDS for PostgreSQL with table partitioning
D.Amazon S3 with lifecycle policies
AnswerB

Amazon Timestream is a purpose-built time-series database that automatically scales to handle high write throughput and provides fast queries on recent data. It has built-in time-series functions and supports automatic data retention policies to expire data after a specified period, such as 90 days. This directly matches the requirements for IoT sensor data with timestamp and sensor ID queries and automatic data expiration.

Why this answer

Amazon Timestream is a fully managed time-series database that handles high write throughput, provides fast queries on recent data, and supports automatic data retention policies. It is designed specifically for time-series data like IoT sensor readings, with built-in functions for time-based queries. Other options either lack native time-series optimization or require manual management of retention and querying.

Exam trap

The trap here is assuming that DynamoDB with TTL is sufficient for time-series data, but Timestream offers purpose-built time-series capabilities and automatic retention that simplify management.

127
Multi-Selectmedium

A data engineering team is designing a data lake on Amazon S3 for storing sensor data from IoT devices. The data is written in near real-time and needs to be queried using Amazon Athena. Which TWO configurations should the team implement to optimize query performance and minimize costs?

Select 2 answers
A.Compress the data using GZIP.
B.Use S3 Standard-IA storage class.
C.Store the data in Apache Parquet format.
D.Partition the data by date and sensor ID.
E.Enable Requester Pays on the S3 bucket.
AnswersC, D

Parquet is columnar, so Athena reads only the columns referenced in each query and scans far less data than row-based CSV or JSON. This directly satisfies the stem's requirement to optimise query performance and minimise cost, since Athena charges per terabyte scanned.

Why this answer

Option C is correct because storing the data in Apache Parquet, a columnar format, allows Athena to read only the columns referenced in a query and leverages columnar compression and predicate pushdown, which dramatically reduces the amount of data scanned and therefore query cost and latency. Option D is correct because partitioning the data by date and sensor ID lets Athena prune irrelevant partitions using the partition metadata in the AWS Glue Data Catalog, so queries only scan the specific date/sensor combinations needed instead of the entire dataset. Together, Parquet plus partitioning is the standard best practice for cost-efficient, high-performance Athena queries over S3 data lakes.

Option A is not correct because GZIP is a row-based compression format that cannot be split for parallel reads and does not provide column pruning, so it is inferior to Parquet's built-in columnar compression for this use case. Option B is not correct because S3 Standard-IA is a storage-class cost optimization for infrequently accessed data, not a query-performance optimization, and it can even add retrieval costs for frequently queried near real-time data. Option E is not correct because Requester Pays shifts data-transfer costs to the requester and does nothing to improve Athena query performance or reduce the team's own query scanning costs.

Exam trap

AWS often tests the misconception that any compression (like GZIP alone) is sufficient for Athena optimization, but the trap is that without a columnar format like Parquet or ORC, compression alone does not enable column pruning or predicate pushdown, leading to higher scan costs and slower queries.

128
MCQmedium

A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a Redshift table. The data is in Parquet format and is partitioned by date in S3. The engineer wants to load only the data for the last 7 days to reduce load time and cost. Which Redshift command should the engineer use?

A.COPY command with a manifest file that lists only the files for the last 7 days.
B.UNLOAD command with the PARTITION BY clause.
C.INSERT INTO ... SELECT FROM spectrum ... WHERE date >= ...
D.COPY command with the PARTITION clause.
AnswerA

The COPY command can use a manifest file to specify exactly which files to load. By creating a manifest that lists only the Parquet files for the last 7 days, the engineer can load only the required data. This reduces load time and cost by avoiding unnecessary data transfer.

Why this answer

The COPY command with a manifest file allows precise selection of files to load, enabling the engineer to load only the Parquet files for the last 7 days. This minimizes data transfer and load time. Other options either do not exist (PARTITION clause in COPY) or are for unloading data (UNLOAD) or are less efficient for bulk loading (INSERT from Spectrum).

Exam trap

The trap here is assuming that the COPY command can filter by partition automatically, but it requires a manifest or specific prefix to limit the files loaded.

129
MCQeasy

An e-commerce application uses Amazon ElastiCache for Redis to cache product catalog data. The cache currently uses lazy loading. The team wants to ensure that frequently accessed product data is always fresh. Which caching strategy should they implement?

A.Write-through caching
B.Set a TTL of 5 minutes for all cached items
C.Use database read replicas to serve data
D.Lazy loading with TTL
AnswerA

Write-through caching updates the Redis cache synchronously on every database write, so cached product data never goes stale. This directly satisfies the freshness constraint that lazy loading cannot guarantee, since lazy loading only populates entries on a miss and leaves existing values stale until eviction or expiry.

Why this answer

Write-through caching ensures that data is written to the cache simultaneously with the database, guaranteeing that frequently accessed product data is always fresh. This strategy eliminates stale reads by synchronously updating the cache on every write, which directly addresses the requirement for freshness without relying on expiration or lazy population.

Exam trap

The trap here is that candidates often assume lazy loading with a short TTL is sufficient for freshness, but the exam tests the understanding that only write-through (or write-behind) strategies guarantee synchronous cache updates without relying on expiration windows.

How to eliminate wrong answers

Option B is wrong because setting a TTL of 5 minutes does not guarantee freshness; data can still become stale within the TTL window, and frequently accessed items may be served from the cache even after they have been updated in the database. Option C is wrong because database read replicas serve stale data asynchronously and do not cache product data in ElastiCache, failing to meet the caching freshness requirement. Option D is wrong because lazy loading with TTL still allows stale data to be served until the TTL expires or a cache miss triggers a refresh, which does not ensure that frequently accessed data is always fresh.

130
MCQhard

A company is using Amazon S3 to store sensitive customer data. The security team requires that all data be encrypted in transit and at rest. Additionally, they want to prevent any accidental public access. Which combination of actions should the data engineer take?

A.Enable default encryption with SSE-S3, enforce HTTPS only via bucket policy, and enable S3 Block Public Access.
B.Enable default encryption with SSE-KMS, allow both HTTP and HTTPS, and set bucket ACLs to private.
C.Use client-side encryption, enforce HTTPS via bucket policy, and enable S3 Block Public Access.
D.Enable default encryption with SSE-S3, allow HTTP and HTTPS, and use bucket ACLs to block public access.
AnswerA

SSE-S3 default encryption protects data at rest, a bucket policy denying non-TLS requests enforces encryption in transit, and S3 Block Public Access prevents accidental exposure. Together these satisfy all three stated requirements: encryption at rest, encryption in transit, and no public access.

Why this answer

Enabling default encryption with SSE-S3 ensures data is encrypted at rest automatically, enforcing HTTPS only via bucket policy ensures encryption in transit by rejecting HTTP requests, and enabling S3 Block Public Access prevents any accidental public exposure regardless of bucket policies or ACLs. This combination satisfies all security requirements: encryption in transit, encryption at rest, and prevention of public access.

Exam trap

The trap here is that candidates may think bucket ACLs or client-side encryption alone satisfy the requirements, but the exam tests that S3 Block Public Access is needed to fully prevent accidental public access and that HTTPS enforcement is mandatory for encryption in transit.

How to eliminate wrong answers

Option B is wrong because allowing both HTTP and HTTPS violates the encryption-in-transit requirement; HTTPS-only must be enforced. Option C is wrong because client-side encryption does not guarantee encryption at rest on the server side (the data may be decrypted before upload), and the security team requires encryption at rest managed by AWS. Option D is wrong because allowing HTTP and HTTPS fails encryption-in-transit, and bucket ACLs alone are insufficient to block all public access (e.g., bucket policies can still grant public access).

131
MCQhard

A data engineer is configuring an Amazon S3 bucket for a data lake. The bucket must store sensitive financial data and comply with a regulation that requires all data to be encrypted at rest with keys that are automatically rotated every year. The engineer also needs to audit key usage and control access to the keys separately from other AWS services. Which encryption option should the engineer choose?

A.Server-side encryption with Amazon S3 managed keys (SSE-S3)
B.Server-side encryption with AWS KMS keys (SSE-KMS)
C.Server-side encryption with customer-provided keys (SSE-C)
D.Client-side encryption with an AWS KMS customer managed key
AnswerB

SSE-KMS uses AWS KMS customer managed keys, which can be configured for automatic annual rotation. KMS provides separate access control through key policies and IAM, and it logs key usage in AWS CloudTrail for auditing. This meets the compliance requirements for encryption, rotation, auditing, and separate key management. SSE-KMS is the appropriate choice for sensitive data with regulatory mandates.

Why this answer

SSE-KMS with customer managed keys provides automatic key rotation, separate access control via key policies and IAM, and detailed auditing through CloudTrail. These features directly address the compliance requirements for encryption at rest, annual key rotation, and auditing key usage. SSE-S3 and SSE-C lack the necessary auditing and separate access control, while client-side encryption shifts management burden to the application.

Exam trap

The trap here is assuming that any encryption option with automatic rotation meets all compliance needs, ignoring the requirement for separate key access control and auditing.

132
MCQeasy

A data engineer needs to store large volumes of semi-structured JSON data in Amazon S3 and query it using Amazon Athena. The data is generated continuously and appended to S3 in small files. The engineer wants to optimize query performance and reduce costs. Which action should the engineer take?

A.Convert the JSON data to Apache Parquet and partition it by date.
B.Enable S3 Transfer Acceleration for faster data uploads.
C.Use Amazon S3 Select to query the JSON data directly.
D.Store the JSON data in a single large file per day.
AnswerA

Converting JSON to Parquet provides columnar storage, enabling Athena to read only needed columns and reduce scan volume. Partitioning by date allows Athena to prune partitions based on query filters, further reducing data scanned. This combination significantly improves query performance and lowers costs, making it the optimal solution for large-scale semi-structured data queried by Athena.

Why this answer

Converting JSON to Parquet and partitioning by date leverages columnar storage and partition pruning, which are key optimizations for Athena. Parquet reduces the amount of data scanned by reading only required columns, while partitioning allows Athena to skip irrelevant partitions based on query filters. Together, they minimize cost and improve performance for large-scale semi-structured data.

Exam trap

The trap here is focusing on file size or upload speed instead of the format and partitioning strategy that directly impact Athena query cost and performance.

133
MCQhard

A large e-commerce company uses Amazon DynamoDB to store shopping cart data. The table has a partition key of 'user_id' and a sort key of 'item_id'. The application performs frequent updates to the 'quantity' attribute for items in a user's cart. Recently, the operations team noticed that write requests are being throttled during peak shopping hours. The table is provisioned with 10,000 write capacity units (WCUs) and uses DynamoDB Accelerator (DAX) for read caching. The data engineer suspects that the throttling is due to hot partitions. The application uses a single AWS SDK client configured with retries. After reviewing the Amazon CloudWatch metrics, the engineer sees that the WriteThrottleEvents metric spikes for a few partition keys. The table has a high number of partitions. What should the data engineer do to resolve the throttling issue with minimal application changes?

A.Increase the provisioned write capacity to 20,000 WCUs permanently.
B.Enable DynamoDB Global Tables to distribute writes across regions.
C.Add more nodes to the DAX cluster to offload write traffic.
D.Configure DynamoDB Auto Scaling with a maximum WCU setting of 20,000 and a target utilization of 70%.
AnswerD

Auto Scaling dynamically adjusts capacity based on traffic, reducing throttling without permanent overprovisioning.

Why this answer

DynamoDB Auto Scaling can dynamically adjust write capacity in response to traffic patterns, reducing throttling on hot partitions without requiring application changes. Option A is incorrect because permanently increasing WCUs does not adapt to variable demand and may lead to over-provisioning. Option B (Global Tables) replicates data across regions but does not increase write capacity for a single table, so it does not resolve hot partition throttling.

Option C (DAX) is a read cache and does not offload write traffic; it only improves read performance.

134
MCQhard

A financial services company stores sensitive transaction data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys be automatically rotated every year. A data engineer needs to configure the S3 bucket to meet these requirements with minimal ongoing operational effort. Which solution should the engineer implement?

A.Enable default encryption on the S3 bucket using SSE-S3 (AES-256).
B.Enable default encryption on the S3 bucket using SSE-KMS with a customer managed key that has automatic key rotation enabled.
C.Use S3 client-side encryption with a customer-provided key stored in AWS Secrets Manager, and rotate the key annually using a Lambda function.
D.Enable default encryption on the S3 bucket using SSE-KMS with an AWS managed key (aws/s3).
AnswerB

SSE-KMS with a customer managed key allows the use of AWS KMS keys that the customer controls. Enabling automatic key rotation on the customer managed key rotates the key material every year by default, meeting the annual rotation requirement. This configuration requires minimal ongoing effort because S3 automatically encrypts new objects with the specified key.

Why this answer

To meet the requirements, the engineer should use server-side encryption with AWS KMS customer managed keys (SSE-KMS) and enable automatic key rotation on the key. This ensures data is encrypted at rest with a key the customer controls, and rotation occurs annually without manual intervention. The bucket default encryption setting ensures all new objects are encrypted automatically.

Exam trap

The trap here is assuming that SSE-S3 or AWS managed keys provide customer control over rotation, when only customer managed keys allow configuring annual automatic rotation.

135
MCQhard

Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket. What is the effect of this policy?

A.It enforces server-side encryption for all objects written to the bucket.
B.It allows the DataLakeRole to read and write objects, but only over HTTPS.
C.It allows anonymous access to the bucket for HTTPS requests.
D.It denies all access to the bucket except for requests from the DataLakeRole.
AnswerB

The policy grants DataLakeRole s3:GetObject and s3:PutObject permissions, but the condition restricts access to requests made over TLS. Plain HTTP requests are denied, so the role can read and write objects only when the connection is encrypted via HTTPS.

Why this answer

The bucket policy uses a condition key `aws:SecureTransport` set to `true`, which restricts access to HTTPS (TLS) connections only. The `Principal` is `DataLakeRole`, and the `Action` includes `s3:GetObject` and `s3:PutObject`, so the policy allows that role to read and write objects exclusively over HTTPS, enforcing encrypted data in transit.

Exam trap

AWS often tests the distinction between encryption in transit (HTTPS/TLS) and encryption at rest (SSE), leading candidates to confuse the `aws:SecureTransport` condition with server-side encryption requirements.

How to eliminate wrong answers

Option A is wrong because the policy does not reference `s3:x-amz-server-side-encryption` or any condition enforcing server-side encryption (SSE) at rest; it only enforces encryption in transit via `aws:SecureTransport`. Option C is wrong because the `Principal` is explicitly set to `DataLakeRole` (an IAM role ARN), not `"*"` or `{"AWS": "*"}`, so anonymous access is not granted. Option D is wrong because the policy includes an `Allow` effect for `DataLakeRole` under the HTTPS condition, but it does not contain a `Deny` statement for other principals or conditions; without an explicit `Deny`, other access may still be allowed by other policies (e.g., bucket ACLs or IAM policies), so it does not deny all other access.

136
MCQeasy

A company uses Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed key that is rotated annually. Which encryption option should be used?

A.SSE-KMS (Server-Side Encryption with AWS KMS).
B.SSE-S3 (Server-Side Encryption with S3-managed keys).
C.Client-side encryption.
D.SSE-C (Server-Side Encryption with Customer-Provided keys).
AnswerA

SSE-KMS encrypts objects at rest using keys held in AWS KMS, and supports customer-managed keys with automatic annual rotation. This satisfies the requirement for a customer-managed, annually rotated key, unlike SSE-S3, which uses AWS-managed keys that cannot be rotated on demand.

Why this answer

SSE-KMS is the correct choice because it allows you to use a customer-managed key (CMK) in AWS KMS, which you can configure to rotate automatically on an annual schedule. This satisfies the security team's requirement for encryption at rest with a key you control and rotate yearly, while still leveraging server-side encryption that integrates with S3's existing infrastructure.

Exam trap

The trap here is that candidates often confuse SSE-C with customer-managed keys, but SSE-C requires you to supply the key on every operation and does not support AWS-managed rotation, making it unsuitable for the 'rotated annually' requirement.

How to eliminate wrong answers

Option B (SSE-S3) is wrong because it uses S3-managed keys that are automatically rotated by AWS, not customer-managed keys, so you cannot control the rotation schedule or manage the key yourself. Option C (Client-side encryption) is wrong because it encrypts data before it reaches S3, which does not meet the requirement for server-side encryption at rest managed by AWS; it also places the key management burden entirely on the client, not the customer-managed key service. Option D (SSE-C) is wrong because it requires you to provide your own encryption key with each request, and AWS does not manage or rotate the key—you must handle key storage and rotation entirely outside of AWS, which contradicts the requirement for a customer-managed key that is rotated annually within AWS.

137
MCQhard

Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket named data-lake-bucket. The engineer wants to allow only GET requests from the corporate network (10.0.0.0/16) over HTTPS. However, users report that they cannot access objects even when connected to the corporate network. What is the issue?

A.The Deny statement should include a condition on the source IP.
B.The Allow statement should include a condition for SecureTransport.
C.The Allow statement should specify s3:GetObject instead of s3:GetObject.
D.The Deny statement blocks all requests that are not using HTTPS, including those from the corporate network.
AnswerD

Deny overrides Allow when condition is met.

Why this answer

The Deny statement with `aws:SecureTransport` set to `false` blocks all HTTP requests. Since the Allow statement only permits GET requests from the corporate network (10.0.0.0/16) but does not require HTTPS, any request from that network that uses HTTP is denied by the explicit Deny. The Deny statement overrides the Allow, so even legitimate corporate users are blocked if they use HTTP.

Exam trap

AWS often tests the principle that an explicit Deny overrides any Allow, leading candidates to focus on fixing the Allow statement rather than recognizing that the Deny unconditionally blocks HTTP traffic from all sources, including the corporate network.

How to eliminate wrong answers

Option A is wrong because the Deny statement already includes a condition on `aws:SecureTransport`, not on source IP; adding a source IP condition would not fix the HTTPS enforcement issue. Option B is wrong because the Allow statement already includes a condition for `aws:SecureTransport` equal to `true` in the Deny, but the Allow itself lacks a SecureTransport condition, so it permits both HTTP and HTTPS; adding SecureTransport to the Allow would not resolve the Deny blocking HTTP. Option C is wrong because `s3:GetObject` is the correct action for GET requests; the typo 's3:GetObject' in the question is a red herring, and the actual policy uses the correct action.

138
MCQmedium

A company is using Amazon RDS for MySQL and needs to automate backups with a retention period of 35 days. They also want to be able to restore to any point within the retention period. Which configuration should be used?

A.Enable manual snapshots daily and retain for 35 days.
B.Set the backup retention period to 35 days and enable automatic backups.
C.Set the backup retention period to 7 days and create daily manual snapshots.
D.Disable automated backups and rely on Multi-AZ for recovery.
AnswerB

Automatic backups capture daily snapshots plus continuous transaction logs, enabling point-in-time recovery to any second within the retention window. Setting retention to 35 days satisfies the stated requirement exactly, since RDS permits 0–35 days for MySQL. Manual snapshots alone cannot deliver point-in-time restore.

Why this answer

Amazon RDS for MySQL supports automated backups with a configurable retention period of up to 35 days. By setting the backup retention period to 35 days and enabling automatic backups, RDS automatically performs daily snapshots and transaction log backups, enabling point-in-time recovery (PITR) to any second within the retention window. This meets the requirement for both a 35-day retention and full PITR capability without manual intervention.

Exam trap

The trap here is that candidates often confuse manual snapshots (which are retained indefinitely but do not support PITR) with automated backups (which support PITR but have a maximum retention of 35 days), leading them to choose Option A or C, thinking manual snapshots can extend the PITR window.

How to eliminate wrong answers

Option A is wrong because manual snapshots are not automatically taken daily and do not support point-in-time recovery; they only provide a single point-in-time restore, not continuous PITR. Option C is wrong because setting the backup retention period to 7 days limits automated backups and PITR to only 7 days, and adding daily manual snapshots does not extend the PITR window beyond 7 days. Option D is wrong because disabling automated backups eliminates both automated snapshots and transaction log backups, making PITR impossible; Multi-AZ provides high availability but does not create backups or enable recovery to any point in time.

139
Multi-Selectmedium

A company is designing a data store for IoT sensor data that is written once and never updated. The data must be stored with high durability and low cost. Which TWO AWS storage services are most suitable? (Choose TWO.)

Select 2 answers
A.Amazon ElastiCache
B.Amazon EBS
C.Amazon S3
D.Amazon DynamoDB
E.Amazon S3 Glacier Deep Archive
AnswersC, E

S3 provides 99.999999999% durability and low cost for infrequently accessed data.

Why this answer

Amazon S3 is correct because it provides 99.999999999% (11 9's) durability, is designed for write-once-read-many (WORM) workloads, and offers low-cost storage tiers suitable for IoT sensor data that is never updated. S3's object storage model and lifecycle policies allow automatic transition to colder storage, making it ideal for immutable data at scale.

Exam trap

The trap here is that candidates often choose DynamoDB (D) for its scalability and low latency, overlooking that the question emphasizes low cost and write-once immutability, where S3 and Glacier Deep Archive are orders of magnitude cheaper per GB stored.

140
MCQeasy

A company needs to migrate an on-premises 10 TB PostgreSQL database to Amazon RDS for PostgreSQL with minimal downtime. Which AWS service should be used for the migration?

A.AWS Storage Gateway
B.AWS Snowball Edge
C.AWS DataSync
D.AWS Database Migration Service (DMS)
AnswerD

DMS supports continuous replication.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it is specifically designed for migrating databases to AWS with minimal downtime. It supports continuous replication from an on-premises PostgreSQL source to Amazon RDS for PostgreSQL using change data capture (CDC), allowing the source database to remain operational during the migration.

Exam trap

The trap here is that candidates may confuse data transfer services (like Snowball Edge or DataSync) with database migration tools, overlooking that DMS is the only AWS service that supports live, ongoing replication and schema conversion for relational databases like PostgreSQL.

How to eliminate wrong answers

Option A is wrong because AWS Storage Gateway is a hybrid storage service for on-premises access to cloud storage, not a database migration tool; it cannot perform schema conversion or ongoing replication for a PostgreSQL database. Option B is wrong because AWS Snowball Edge is a physical data transport device for large-scale data transfers, but it does not support live database replication or CDC, making it unsuitable for minimal-downtime migrations of a live 10 TB database. Option C is wrong because AWS DataSync is designed for moving large amounts of file data over the network, not for database-level migrations; it lacks the ability to handle PostgreSQL-specific objects, transactions, or ongoing replication.

141
MCQeasy

A company stores its application logs in Amazon S3. The logs are generated daily and need to be retained for 3 years for compliance. The logs are accessed frequently for the first 30 days, occasionally for the next 6 months, and rarely after that. The data engineering team wants to minimize storage costs while ensuring that logs are available for retrieval within 12 hours for the first 6 months and within 48 hours after that. The team also wants to automatically delete logs after 3 years. Which lifecycle policy should the team implement?

A.Transition to S3 One Zone-IA after 30 days, delete after 6 months.
B.Transition to S3 Standard-IA after 30 days, delete after 3 years.
C.Transition to S3 Standard-IA after 30 days, to S3 Glacier after 6 months, delete after 3 years.
D.Transition to S3 Standard-IA after 30 days, to S3 Glacier Deep Archive after 6 months, delete after 3 years.
AnswerD

Meets cost and retrieval time requirements.

Why this answer

It aligns with the access patterns and retrieval requirements: S3 Standard-IA after 30 days for occasional access with immediate retrieval, then S3 Glacier Deep Archive after 6 months for rare access with a 12-hour retrieval time (via expedited or standard retrieval), and deletion after 3 years. This minimizes storage costs while meeting the 48-hour retrieval window for older logs.

Exam trap

The trap here is that candidates may choose Option C (S3 Glacier) thinking it is the cheapest cold storage, but S3 Glacier Deep Archive is actually the lowest-cost option for data that is rarely accessed and can tolerate a 12-hour retrieval time, which still satisfies the 48-hour requirement.

How to eliminate wrong answers

Option A is wrong because it transitions to S3 One Zone-IA after 30 days, which does not provide the durability or availability needed for compliance logs, and it deletes after 6 months instead of 3 years. Option B is wrong because it keeps logs in S3 Standard-IA for the entire 3 years, which is more expensive than transitioning to a colder storage class after 6 months, and it does not meet the cost-minimization goal. Option C is wrong because it transitions to S3 Glacier after 6 months, which has a retrieval time of 1-5 minutes for expedited or 3-5 hours for standard, exceeding the 48-hour requirement but not being the most cost-effective option; S3 Glacier Deep Archive is cheaper and still meets the 48-hour retrieval window.

142
MCQhard

A data engineer is using Amazon Redshift and needs to improve query performance for a workload that involves frequent joins between a large fact table and a small dimension table. The dimension table is updated daily with new records. The engineer wants to minimize data movement during joins. Which Redshift distribution style should the engineer use for the dimension table?

A.AUTO
B.EVEN
C.KEY
D.ALL
AnswerD

ALL distribution replicates the entire dimension table to every compute node. This ensures that joins with the fact table can be performed locally on each node without data movement, significantly improving join performance. It is ideal for small, slowly changing dimension tables.

Why this answer

ALL distribution replicates the entire dimension table to every node, eliminating the need to redistribute data during joins with the fact table. This minimizes network traffic and speeds up join operations. It is best suited for small dimension tables that are frequently joined with larger tables.

Exam trap

The trap here is assuming that KEY distribution on the join column always eliminates data movement, but it does not when the fact table is distributed differently.

143
MCQmedium

A data engineer needs to store semi-structured JSON data from IoT devices. The data is written once, read rarely, but must be queryable using SQL. The storage cost must be minimized. Which storage solution should the engineer choose?

A.Store JSON in Amazon Redshift as SUPER data type
B.Store JSON in an Amazon RDS for MySQL table
C.Store JSON documents in Amazon DynamoDB and use PartiQL for queries
D.Store JSON files in Amazon S3 and use Amazon Athena for queries
AnswerD

S3 provides the lowest-cost durable storage for write-once, rarely read JSON, and Athena queries that data in place using standard SQL over the AWS Glue Data Catalog, avoiding the cost of a continuously running database or cluster.

Why this answer

Amazon S3 provides the lowest-cost storage for data that is written once and rarely read, while Amazon Athena enables serverless SQL querying directly on JSON files stored in S3. This combination minimizes storage costs because S3 charges only for the data stored and retrieval, with no minimum fees or provisioning required, and Athena charges only for the data scanned per query. The workload's write-once, read-rarely pattern aligns perfectly with S3's durability and lifecycle policies, making it the most cost-effective choice.

Exam trap

The trap here is that candidates often choose DynamoDB (Option C) because it supports JSON natively and PartiQL provides SQL-like queries, but they overlook that DynamoDB's provisioned throughput and storage costs are significantly higher than S3 for write-once, read-rarely workloads, and that Athena on S3 is the serverless, cost-optimized solution for ad-hoc SQL queries on infrequently accessed data.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a petabyte-scale data warehouse designed for high-performance analytics on structured and semi-structured data, but it incurs significant costs for provisioned clusters even when data is rarely queried, making it unsuitable for minimizing storage costs. Option B is wrong because Amazon RDS for MySQL is a relational database that requires provisioning and paying for a database instance 24/7, and storing JSON in a MySQL table incurs overhead for indexing and transactions that are unnecessary for write-once, read-rarely data. Option C is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for low-latency, high-throughput workloads, but its storage costs are higher than S3 for rarely accessed data, and PartiQL queries on DynamoDB still consume read capacity units, leading to ongoing costs that exceed S3+Athena for infrequent queries.

144
Multi-Selectmedium

A data engineer is configuring an Amazon Redshift cluster and needs to optimize query performance for complex analytical queries that involve large joins. The engineer wants to reduce the amount of data movement during query execution. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Increase the number of nodes in the cluster to add more compute resources.
B.Define sort keys on columns frequently used in join and filter conditions.
C.Enable concurrency scaling to handle concurrent queries.
D.Choose a distribution style that colocates joined tables on the same node slices.
E.Use columnar storage for all tables.
AnswersB, D

Sort keys determine the order in which data is stored on disk. When sort keys are defined on columns used in joins and filters, Redshift can use zone maps to skip irrelevant blocks and perform efficient merge joins. This reduces I/O and data movement, enhancing query performance for complex analytical workloads.

Why this answer

To reduce data movement during joins in Amazon Redshift, the engineer should choose a distribution style that colocates joined tables on the same node slices and define sort keys on columns used in joins and filters. These actions minimize network traffic and enable efficient merge joins. Increasing nodes, enabling concurrency scaling, or using columnar storage do not directly address data movement within a query.

Exam trap

The trap here is assuming that adding more nodes or enabling concurrency scaling will reduce data movement during joins, but those actions address capacity and concurrency, not join efficiency.

145
MCQeasy

A company is using an Amazon RDS for PostgreSQL database to store application data. The data engineering team needs to run complex analytical queries that join multiple large tables. These queries are causing performance degradation on the production database. The team wants to offload the analytical workload to a separate system that can handle large-scale data processing. Which AWS service should the team use?

A.Amazon Redshift
B.Amazon ElastiCache for Redis
C.Amazon RDS for MySQL
D.Amazon DynamoDB
AnswerA

Amazon Redshift is a fully managed, petabyte-scale data warehouse designed for analytical queries. It uses columnar storage and massively parallel processing to handle complex joins and aggregations efficiently. Offloading analytical workloads to Redshift frees up the RDS production database and provides better performance for large-scale data processing.

Why this answer

Amazon Redshift is purpose-built for analytical workloads, offering columnar storage, massively parallel processing, and advanced query optimization. By moving the analytical queries to Redshift, the team can achieve faster performance without impacting the production RDS database. This separation of transactional and analytical workloads is a common best practice.

Exam trap

The trap here is assuming that any database can handle analytical queries, but transactional databases like RDS are not optimized for large-scale joins and aggregations.

146
MCQmedium

A company is using Amazon DynamoDB to store session data for a web application. The data engineer needs to ensure that the data is encrypted at rest. Which action should the data engineer take?

A.Enable encryption at rest on the DynamoDB Accelerator (DAX) cluster.
B.Use client-side encryption before writing to DynamoDB.
C.Ensure encryption at rest is enabled on the DynamoDB table (default).
D.Enable DynamoDB Time to Live (TTL) to encrypt data.
AnswerC

DynamoDB encrypts all data at rest by default using AWS-owned keys, so no action is needed to enable it. Verifying that default encryption is active on the table satisfies the requirement without configuring client-side or KMS key options.

Why this answer

DynamoDB tables are encrypted at rest by default using AWS Key Management Service (KMS) with an AWS owned key. The data engineer does not need to take any additional action to enable encryption at rest, as it is automatically enabled for all new DynamoDB tables. This ensures that all data stored on disk, including the table's primary key, local secondary indexes, and global secondary indexes, is encrypted before being written to SSDs in AWS data centers.

Exam trap

The trap here is that candidates may assume encryption at rest must be explicitly enabled or configured, when in fact DynamoDB enables it by default for all tables, leading them to incorrectly select client-side encryption or other unrelated options.

How to eliminate wrong answers

Option A is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that does not store data at rest; it only caches data in memory, and encryption at rest is not a configurable feature for DAX clusters. Option B is wrong because client-side encryption is an additional security measure that encrypts data before sending it to DynamoDB, but it is not required to meet the requirement of encryption at rest, which is already provided by DynamoDB's default server-side encryption. Option D is wrong because DynamoDB Time to Live (TTL) is a feature that automatically deletes expired items from a table to manage storage costs; it does not provide any encryption functionality.

147
MCQhard

A data engineer is designing a data lake on Amazon S3 and needs to ensure that data is encrypted at rest. The company requires that encryption keys be managed by AWS and automatically rotated annually. The engineer also wants to audit key usage. Which S3 encryption option should the engineer choose?

A.Server-side encryption with Amazon S3 managed keys (SSE-S3).
B.Server-side encryption with customer-provided keys (SSE-C).
C.Client-side encryption with a custom key management solution.
D.Server-side encryption with AWS KMS managed keys (SSE-KMS).
AnswerD

SSE-KMS uses AWS KMS to manage encryption keys, which can be automatically rotated annually. AWS KMS integrates with AWS CloudTrail to log key usage, providing the required auditability. This option meets both the encryption and auditing requirements.

Why this answer

SSE-KMS uses AWS KMS to manage encryption keys, supports automatic annual key rotation, and logs key usage in AWS CloudTrail. This provides both the required encryption and auditability. Other options either do not support AWS-managed rotation or lack auditing capabilities.

Exam trap

The trap here is assuming that SSE-S3 provides auditing because it is AWS-managed, but only KMS-backed keys provide CloudTrail logs of key usage.

148
MCQeasy

A company stores application logs in Amazon S3 and needs to query them using standard SQL. The logs are in JSON format and are updated daily. The data engineering team wants a serverless solution that requires minimal management and can automatically discover the schema. Which AWS service should they use?

A.Amazon EMR with Apache Hive.
B.Amazon Athena with an AWS Glue Data Catalog table.
C.Amazon RDS for PostgreSQL with the aws_s3 extension.
D.Amazon Redshift Spectrum.
AnswerB

Amazon Athena is a serverless query service that allows SQL queries directly on S3 data. By using an AWS Glue Data Catalog table, the schema can be automatically discovered and updated. This combination requires no infrastructure management and is ideal for querying JSON logs stored in S3 with minimal overhead.

Why this answer

Amazon Athena is a serverless query service that integrates with the AWS Glue Data Catalog to automatically discover schemas and query data directly in S3. It requires no infrastructure management and supports standard SQL, making it ideal for querying JSON logs with minimal operational overhead.

Exam trap

The trap here is confusing serverless querying with managed cluster services like Redshift Spectrum or EMR, which require more management and are not fully serverless.

149
MCQeasy

A data engineer needs to store transaction data that requires strong consistency, ACID transactions, and complex join queries. Which AWS service is most appropriate?

A.Amazon DynamoDB
B.Amazon RDS for PostgreSQL
C.Amazon S3
D.Amazon Redshift
AnswerB

Amazon RDS for PostgreSQL provides ACID transactions, strong consistency and full SQL joins, matching all three stated requirements. Its relational engine supports complex multi-table joins natively, unlike purpose-built NoSQL stores that trade joins and cross-item transactions for scale.

Why this answer

Amazon RDS for PostgreSQL is the most appropriate choice because it provides full ACID transaction support, strong consistency, and the ability to perform complex join queries using standard SQL. Unlike NoSQL or data warehouse solutions, PostgreSQL is a relational database that excels at enforcing referential integrity and supporting multi-table joins with advanced indexing.

Exam trap

The trap here is that candidates often confuse DynamoDB's 'eventually consistent reads' with strong consistency, or assume its limited transaction API can replace full ACID relational databases, but the question explicitly requires complex joins and ACID transactions, which only a relational database like PostgreSQL can provide.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database that does not support complex join queries or ACID transactions across multiple items (it only offers limited transactional APIs with restrictions). Option C is wrong because Amazon S3 is an object storage service with no support for ACID transactions, complex joins, or relational query capabilities. Option D is wrong because Amazon Redshift is a columnar data warehouse optimized for analytical queries on large datasets, not for transactional workloads requiring ACID compliance and complex joins at the row level.

150
MCQeasy

A data engineer applies the above IAM policy to a user. The user attempts to upload an object to the bucket 'my-data-lake' without specifying server-side encryption. What will happen?

A.The upload fails only if the bucket has a default encryption setting
B.The upload succeeds if the bucket policy allows unencrypted uploads
C.The upload succeeds because the policy allows s3:PutObject
D.The upload fails because the condition requires encryption
AnswerD

The policy's `s3:PutObject` statement carries a `StringNotEquals` condition on `s3:x-amz-server-side-encryption`, so a request omitting that header evaluates as non-compliant and is denied. Because the user supplied no encryption parameter, the condition fails and the upload is rejected with AccessDenied, satisfying the stem's unencrypted-upload constraint.

Why this answer

The IAM policy includes a condition that requires the `s3:x-amz-server-side-encryption` header to be present with a value of `AES256`. When the user attempts to upload an object without specifying server-side encryption, the condition is not satisfied, so the `s3:PutObject` permission is denied. This causes the upload to fail, regardless of any bucket default encryption settings or bucket policies.

Exam trap

The trap here is that candidates assume bucket default encryption automatically satisfies an IAM condition requiring the encryption header, but the condition checks the request headers, not the bucket's configuration.

How to eliminate wrong answers

Option A is wrong because the failure is due to the IAM policy condition, not the bucket's default encryption setting; even if the bucket has no default encryption, the IAM policy still denies the upload. Option B is wrong because the bucket policy is irrelevant here—the IAM policy explicitly denies unencrypted uploads via a condition, and bucket policies cannot override an explicit IAM deny. Option C is wrong because while the policy allows `s3:PutObject` in general, the condition `s3:x-amz-server-side-encryption` must be met; without it, the permission is effectively denied.

← PreviousPage 2 of 5 · 358 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Store Management questions.