Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 751–825

1321 questions total · 18pages · All types, answers revealed

Page 10

Page 11 of 18

Page 12
751
MCQmedium

A company is ingesting log files from multiple EC2 instances into Amazon S3 using the CloudWatch agent. The logs are delivered to a CloudWatch Logs group, and a subscription filter sends them to a Lambda function for transformation, then to Firehose. The Firehose stream is configured with a buffer interval of 60 seconds and buffer size of 5 MB. The logs are critical and must be available in S3 within 5 minutes. What is the most cost-effective way to reduce the delivery latency?

A.Replace Firehose with Amazon Kinesis Data Streams
B.Increase the buffer size to 10 MB
C.Increase the buffer interval to 120 seconds
D.Decrease the buffer interval to 10 seconds
AnswerD

Lower buffer interval reduces delivery latency.

Why this answer

Decreasing the Firehose buffer interval to 10 seconds directly reduces the maximum time data waits in the buffer before being delivered to S3, ensuring logs reach S3 within the required 5-minute window. Since the current 60-second buffer interval is the primary contributor to latency, lowering it to 10 seconds minimizes delivery delay without incurring additional costs, as Firehose charges are based on data volume, not buffer frequency.

Exam trap

The trap here is that candidates may think increasing buffer size or interval improves throughput, but the question asks for reduced latency, so decreasing the buffer interval is the direct and cost-effective solution.

How to eliminate wrong answers

Option A is wrong because replacing Firehose with Kinesis Data Streams would require additional components (e.g., a consumer to write to S3) and increase cost and complexity, not reduce latency cost-effectively. Option B is wrong because increasing the buffer size to 10 MB would allow more data to accumulate before delivery, potentially increasing latency, not reducing it. Option C is wrong because increasing the buffer interval to 120 seconds would double the maximum buffering time, worsening delivery latency.

752
MCQmedium

A data engineering team notices that an AWS Glue ETL job, which processes hourly data from an S3 bucket, is taking progressively longer to run. The job reads Parquet files partitioned by date and hour. Which action is MOST likely to improve the job's performance?

A.Enable pushdown predicate filtering on the job's data source.
B.Convert Parquet files to CSV to improve read performance.
C.Increase the number of DPUs for the job.
D.Switch from Spark to Python shell for simpler processing.
AnswerA

Pushdown predicates filter partitions and row groups at the S3/Parquet source, so Glue reads only matching date and hour data instead of scanning the full dataset. This cuts I/O and shuffle volume, directly addressing the growing runtime.

Why this answer

Pushdown predicate filtering allows AWS Glue to push filter conditions down to the data source, so only the relevant partitions and rows are read from S3. Since the data is partitioned by date and hour, predicate pushdown minimizes the amount of data scanned, reducing I/O and improving job performance. This is especially effective for Parquet, which supports columnar pruning and predicate pushdown.

Exam trap

DEA-C01 often tests the misconception that adding more DPUs is always the answer to performance issues, but the exam expects candidates to identify data pruning and predicate pushdown as the primary optimization for partitioned data.

How to eliminate wrong answers

Option B is wrong because converting Parquet to CSV would remove the columnar storage benefits and increase file size, making reads slower and more expensive. Option C is wrong because increasing DPUs adds more compute resources but does not address the root cause of reading unnecessary data; it may temporarily speed up the job but is not the most efficient or cost-effective solution. Option D is wrong because Python shell jobs are limited in scalability and are not suitable for complex ETL transformations; switching from Spark to Python shell would likely degrade performance for large datasets.

753
MCQhard

A company runs a nightly batch ETL job using AWS Glue to transform data from Amazon RDS for MySQL to Amazon S3. The job reads 100 tables and writes Parquet files partitioned by date. Recently, the job started failing with 'ThrottlingException' from the RDS database. The data volume has increased, and the Glue job is reading large tables without any filtering. The job uses a single Glue job with multiple Spark executors. The engineer needs to reduce the load on the RDS database while maintaining the same processing time. What should the engineer do?

A.Use AWS DMS to continuously replicate data to S3 and then run Glue on the S3 data.
B.Change the job to read all data from RDS into a staging table in S3 first, then transform.
C.Increase the number of DPUs to process data faster.
D.Modify the Glue job to use a JDBC connection with a WHERE clause to read only the latest partition.
AnswerD

Reading all rows from every table drives concurrent JDBC connections that overwhelm RDS, triggering ThrottlingException. Adding a WHERE clause restricts each read to the latest partition, sharply reducing rows scanned and database load while preserving the same overall processing time.

Why this answer

The ThrottlingException originates from RDS because the Glue job issues full-table JDBC reads across 100 tables with no predicate pushdown. Adding a WHERE clause to read only the latest partition (D) reduces the number of rows scanned and transferred per run, directly lowering the query load on RDS while keeping the same nightly processing window. This is the least disruptive change that preserves the existing architecture and processing time.

Exam trap

The trap is assuming 'more DPUs = faster and less load'; in JDBC-based Glue jobs, more parallelism increases concurrent connections and worsens source-side throttling, so the correct answer is to reduce rows read, not add compute.

How to eliminate wrong answers

Option A is wrong because introducing AWS DMS adds a continuous replication pipeline — a larger architectural change than needed, and it doesn't address the immediate throttling without new infrastructure and cutover effort. Option B is wrong because staging all data in S3 first still requires the same full-table reads from RDS, so it does not reduce load on the source database. Option C is wrong because adding DPUs increases parallelism and therefore increases concurrent JDBC connections and query pressure on RDS, making throttling worse rather than better.

754
MCQhard

Refer to the exhibit. A data engineer applies this bucket policy to an S3 bucket. A user within the 10.0.0.0/24 IP range attempts to upload an object to the bucket using an HTTP (non-HTTPS) request. What is the outcome?

A.The upload succeeds because the Allow statement grants permission.
B.The upload succeeds because the user's IP is allowed.
C.The upload fails because the user's IP is not in the allowed range for PutObject.
D.The upload fails because the request is not using HTTPS.
AnswerD

The bucket policy includes a Deny statement conditioned on aws:SecureTransport being false, which rejects any non-HTTPS request regardless of source IP. The 10.0.0.0/24 user's HTTP upload therefore fails, satisfying the explicit deny that overrides the IP-based allowance.

Why this answer

The bucket policy includes a condition `aws:SecureTransport` set to `false`, which explicitly denies any request that does not use HTTPS. Since the user is making an HTTP (non-HTTPS) request, the Deny statement overrides any Allow statement, causing the upload to fail. The correct answer is D.

Exam trap

The DEA-C01 exam often tests the precedence of explicit Deny over Allow in IAM policies, and the trap here is that candidates focus on the IP range in the Allow statement and overlook the Deny condition that blocks non-HTTPS requests.

How to eliminate wrong answers

Option A is wrong because the Allow statement is overridden by the explicit Deny when the condition `aws:SecureTransport` equals `false`; the upload does not succeed. Option B is wrong because even though the user's IP is within the allowed range (10.0.0.0/24), the Deny statement for non-HTTPS requests takes precedence and blocks the upload. Option C is wrong because the user's IP is actually within the allowed range for PutObject; the failure is due to the lack of HTTPS, not the IP range.

755
MCQhard

A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 bucket containing nested JSON, flattens arrays using Relationalize, and writes Parquet to another S3 bucket. The job must run only when new objects land in the source bucket. The engineer wants to avoid unnecessary job runs and minimize cost. Which approach should the engineer use?

A.Configure an Amazon EventBridge rule that matches S3 Object Created events on the source prefix and targets the Glue job as the rule's target.
B.Create an AWS Lambda function that polls the S3 bucket using the ListObjectsV2 API on a one-minute schedule and starts the Glue job when new keys appear.
C.Schedule the Glue job with a cron trigger every five minutes and rely on Glue job bookmarks to skip already-processed files.
D.Enable S3 Event Notifications to publish to an Amazon SNS topic and subscribe the data engineer's email address to the topic.
AnswerA

EventBridge can match S3 Object Created events filtered by bucket and key prefix, then invoke a Glue job directly as a target. This event-driven pattern starts the job exactly when new objects arrive, avoids idle scheduled runs, and reduces cost while meeting the requirement to run only on new data.

Why this answer

S3 can emit Object Created events that EventBridge matches by bucket and prefix, and EventBridge supports an AWS Glue job as a direct target. This is the canonical event-driven pattern: the job starts only when a new object lands, eliminating idle scheduled runs. Polling and SNS notifications either add latency or require manual steps, and a fixed cron schedule still burns DPUs on empty runs.

Exam trap

The trap here is choosing a scheduled cron trigger with bookmarks, which still launches the job on every interval even when no new objects exist.

756
MCQhard

A data engineer maintains an Amazon Redshift cluster where a nightly COPY job loads data into a large fact table. After the load, analysts run queries that filter on a `sale_date` column and join to a small dimension table. Query performance degrades over time as the fact table grows. The engineer wants to improve performance for these recurring queries without changing the query text. Which combination of actions should the engineer take?

A.Enable concurrency scaling on the cluster and increase the number of slices per node.
B.Add an interleaved sort key on every column of the fact table and distribute the dimension table as EVEN.
C.Apply a compound sort key on `sale_date` and enable automatic vacuum and analyze, then distribute the dimension table as ALL.
D.Convert the fact table to a view over the raw staging table and rely on Redshift Spectrum for all reads.
AnswerC

A compound sort key on `sale_date` lets Redshift skip blocks outside the filtered range, reducing scanned data for date-filtered queries. Distributing the small dimension table as ALL replicates it to every node, eliminating data movement during joins. Enabling automatic vacuum and analyze keeps statistics current and reclaims space after the nightly loads, so the query planner continues choosing efficient plans as the table grows.

Why this answer

The recurring workload filters on `sale_date` and joins to a small dimension. A compound sort key on the filter column enables zone-map block skipping, while an ALL distribution on the small dimension removes network redistribution during joins. Automatic vacuum and analyze keep the physical layout and statistics healthy as nightly loads accumulate, so the planner keeps producing efficient plans without query changes.

Exam trap

The trap here is assuming that adding concurrency scaling or more slices will fix a single query's cost, when scan and join efficiency depend on sort and distribution design.

757
Multi-Selectmedium

A data engineer is using AWS Glue DataBrew to prepare a dataset stored in Amazon S3. The dataset contains missing values, inconsistent date formats, and duplicate rows. The engineer needs to clean the data and produce a transformed output for downstream analytics. (Choose two.)

Select 2 answers
A.Schedule a DataBrew recipe job to run on a cron expression so the cleaning steps execute periodically.
B.Create a recipe step using the Fill missing values transformation to replace nulls with a specified value or strategy.
C.Use the DataBrew profile job to generate statistics about missing values and duplicates in the dataset.
D.Configure a DataBrew dataset to use the S3 bucket as a source and set the file type to CSV with header detection enabled.
E.Apply a recipe step that uses the Remove duplicates transformation to eliminate duplicate rows.
AnswersB, E

The Fill missing values transformation lets you replace null or missing entries with a constant, a calculated value, or a strategy such as most frequent. It directly resolves the missing values problem in the dataset. Because it is a recipe step, it remains part of the reusable transformation pipeline applied to the S3 data.

Why this answer

Cleaning the dataset requires transformation steps that actually modify data. Removing duplicates eliminates redundant rows, and filling missing values addresses nulls. Together these recipe steps resolve two of the stated data quality issues.

Date format normalization would need an additional step. Source configuration, profiling, and scheduling support the workflow but do not themselves clean the data.

Exam trap

The trap here is selecting operational or diagnostic actions like profiling or scheduling instead of the transformation steps that actually change the data.

758
MCQhard

A company runs a data ingestion pipeline that uses AWS Glue to read 500 GB of JSON files from an S3 bucket (s3://raw-data/) every hour. The Glue ETL job transforms the data and writes Parquet files to another S3 bucket (s3://processed-data/). The job is triggered by a time-based CloudWatch Events rule. Recently, the job has started taking over 2 hours to complete, causing delays in downstream processes. The data volume has been consistent, and no changes have been made to the job code or infrastructure. The S3 bucket 's3://raw-data/' receives new files continuously, but the Glue job reads all files in the bucket each run (no incremental processing). The engineer suspects that the job is reprocessing old data. Which action should the engineer take FIRST to reduce the job duration?

A.Enable Glue job bookmarking and configure the job to process only new data.
B.Increase the parallelism of the Spark job by repartitioning the data.
C.Add partition pruning by modifying the S3 path to include date-based partitions.
D.Increase the number of DPUs for the Glue job to 100.
AnswerA

Without job bookmarks, every run rereads the entire s3://raw-data/ prefix, so the 500 GB backlog is reprocessed hourly and runtime balloons. Enabling bookmarks tracks previously processed objects, letting Glue read only newly arrived files, which directly shortens the job.

Why this answer

The Glue job is reprocessing all 500 GB of data every hour because it lacks job bookmarking, which tracks previously processed data. Enabling job bookmarking allows Glue to persist state about which files have already been processed, so subsequent runs only read new files. This directly reduces the input volume and thus the job duration without changing the job logic or infrastructure.

Since the data volume is consistent and no code changes were made, the issue is purely due to reprocessing old data, making bookmarking the correct first step.

Exam trap

DEA-C01 often tests the misconception that scaling resources (more DPUs) or optimizing data layout (partitioning) is the first step to improve job performance, when the root cause is reprocessing old data due to missing job bookmarks.

How to eliminate wrong answers

Option B is wrong because repartitioning the data increases parallelism but does not reduce the total volume of data processed; the job would still read all 500 GB each run, so the duration would not improve significantly. Option C is wrong because partition pruning requires the data to be organized in a partitioned structure (e.g., by date) and the job to be written to leverage partitions; the scenario does not indicate that the data is partitioned, and adding partition pruning would require restructuring the data and modifying the job, which is not the first step. Option D is wrong because increasing DPUs scales resources but does not address the root cause of reprocessing old data; it would increase cost without solving the inefficiency, and the job would still read all data each run.

759
MCQmedium

A data engineer is designing a data ingestion pipeline to load data from an on-premises Oracle database to Amazon S3. The pipeline should capture changes in near real-time (within minutes) and minimize impact on the source database. The source table has a 'last_modified' timestamp column. Which service combination would meet these requirements?

A.AWS DMS with a replication task in CDC mode, writing to S3 in Parquet format.
B.Amazon Kinesis Data Firehose with a Lambda function that queries Oracle.
C.AWS Data Pipeline with a periodic SQL query activity to copy full table snapshots.
D.AWS Glue with a JDBC connection to Oracle, running a crawler every 5 minutes.
AnswerA

DMS change data capture reads the Oracle redo logs continuously, applying minimal query load on the source, and delivers near-real-time changes to S3. Writing Parquet reduces storage and query cost, satisfying the minutes-level latency and low source impact constraints.

Why this answer

AWS DMS with a replication task in CDC (Change Data Capture) mode is the correct choice because it continuously reads the Oracle redo logs to capture near real-time changes (within seconds to minutes) with minimal impact on the source database. It can directly write to S3 in Parquet format, meeting the requirement for low-latency ingestion without full table scans.

Exam trap

The trap here is that candidates assume a 'last_modified' timestamp column enables easy CDC via polling (options B, D), but the question tests that true near real-time CDC with minimal source impact requires reading database redo logs, not querying the table, and that services like Kinesis or Glue cannot perform log-based CDC without additional custom logic.

How to eliminate wrong answers

Option B is wrong because Kinesis Data Firehose with a Lambda function that queries Oracle would require periodic polling of the source table, which either misses changes between polls or causes high load on the database if polled too frequently, and it does not natively support CDC from Oracle redo logs. Option C is wrong because AWS Data Pipeline with a periodic SQL query activity to copy full table snapshots performs full table scans, which impacts the source database and cannot capture changes in near real-time (only at scheduled intervals). Option D is wrong because AWS Glue with a JDBC connection running a crawler every 5 minutes performs full table scans or incremental queries based on the 'last_modified' column, but this still causes repeated query load on Oracle and cannot achieve true near real-time CDC without reading redo logs.

760
MCQhard

A data engineer is designing a data ingestion pipeline for clickstream data from a mobile app. The data volume varies, with occasional spikes up to 10 MB/s. The pipeline must persist the raw data in Amazon S3 and make it available for near-real-time analytics via Amazon Athena. Which combination of services minimizes cost and operational overhead?

A.Amazon Kinesis Data Streams with Amazon Kinesis Data Analytics, then Amazon S3
B.Amazon SQS with an Auto Scaling group of EC2 instances writing to Amazon S3
C.Amazon Kinesis Data Streams with AWS Lambda for transformation, then Amazon S3
D.Amazon Kinesis Data Firehose with direct delivery to Amazon S3, then Amazon Athena
AnswerD

Firehose is fully managed, scales automatically, and delivers to S3.

Why this answer

Amazon Kinesis Data Firehose is the most cost-effective and low-overhead solution for ingesting variable-volume clickstream data (up to 10 MB/s) into Amazon S3 because it is a fully managed service that automatically scales, buffers, and compresses data before delivery. It integrates directly with S3 without requiring custom code or infrastructure management, and the data is immediately queryable by Amazon Athena with no additional transformation steps.

Exam trap

The trap here is that candidates often choose Amazon Kinesis Data Streams with Lambda (Option C) because they think it provides more control, but they overlook Lambda's concurrency limits and the operational burden of managing stream shards, making Firehose the simpler and cheaper choice for raw data ingestion to S3.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Analytics adds unnecessary cost and complexity for a pipeline that only needs to persist raw data to S3, and it requires a separate stream consumer. Option B is wrong because Amazon SQS with Auto Scaling EC2 instances introduces significant operational overhead for managing instance scaling, SQS polling, and potential data loss or duplication, and it is not optimized for near-real-time streaming at 10 MB/s. Option C is wrong because AWS Lambda has a maximum invocation duration of 15 minutes and a concurrency limit that can cause throttling during spikes, making it unsuitable for sustained 10 MB/s throughput without complex sharding and retry logic.

761
MCQmedium

A data engineer manages an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to an Amazon Redshift table. The job runs daily and recently started failing with the error 'Unable to find a suitable security group for the connection'. The Glue connection is configured with a VPC, subnet, and security group. The engineer verifies that the IAM role has the necessary permissions and the S3 bucket is accessible. What is the most likely cause of this error?

A.The security group attached to the Glue connection does not allow outbound traffic to the Redshift cluster on the required port.
B.The IAM role used by the Glue job lacks permissions to describe security groups in the VPC.
C.The Glue connection is associated with a subnet that does not have a route to the Redshift cluster.
D.The security group specified in the Glue connection does not exist or is not in the same VPC as the subnet.
AnswerD

Correct. AWS Glue requires that the security group and subnet belong to the same VPC. If the security group is deleted, renamed, or belongs to a different VPC, Glue cannot use it and throws this error. This is a common misconfiguration when VPC settings are changed or when the connection is created with incorrect parameters.

Why this answer

The error 'Unable to find a suitable security group for the connection' occurs when the security group specified in the Glue connection is invalid for the subnet. This typically happens if the security group does not exist, is in a different VPC, or is not associated with the subnet's VPC. Ensuring the security group and subnet are in the same VPC and that the security group exists resolves the issue.

Exam trap

The trap here is assuming the error is about network routing or IAM permissions, when it specifically points to a security group mismatch with the subnet's VPC.

762
MCQhard

A company is using Amazon EMR to process data stored in Amazon S3. The S3 bucket is configured with a bucket policy that denies access unless the request includes a specific tag. The EMR cluster's IAM role has s3:GetObject permission. However, the EMR job fails to read data from S3. What is the most likely cause?

A.The bucket policy is not attached to the EMR role.
B.The EMR cluster is not in the same account as the S3 bucket.
C.The IAM role does not have a condition that matches the required tag.
D.The EMR role does not have s3:GetObject permission.
AnswerC

The bucket policy requires a tag, and the role must have a matching condition.

Why this answer

The bucket policy denies access unless the request includes a specific tag. Even though the EMR cluster's IAM role has s3:GetObject permission, the IAM role does not have a condition key (e.g., aws:RequestTag) that matches the required tag. Therefore, the request is denied by the bucket policy, causing the EMR job to fail.

Exam trap

AWS often tests the interaction between IAM policies and S3 bucket policies, specifically that a bucket policy with a deny condition can override IAM permissions, and candidates mistakenly think the issue is missing IAM permissions rather than a missing condition in the request.

How to eliminate wrong answers

Option A is wrong because bucket policies are attached to the S3 bucket, not to IAM roles; the policy is already configured on the bucket. Option B is wrong because cross-account access is possible with proper permissions, and the question does not indicate a different account; the failure is due to the tag condition, not account mismatch. Option D is wrong because the question explicitly states the IAM role has s3:GetObject permission, so the failure is not due to missing permission.

763
MCQhard

A data engineer maintains an Amazon DynamoDB table that stores IoT telemetry. The table uses a partition key of deviceId and a sort key of timestamp, with on-demand capacity. A few very active devices generate millions of writes per hour while thousands of other devices write sporadically. The engineer observes throttling on writes for the active devices and wants to reduce it with the least application change. Which action should the engineer take?

A.Enable DynamoDB Streams and process writes asynchronously
B.Switch the table to provisioned capacity with a high write capacity unit setting
C.Add a random suffix to the partition key to spread writes across partitions
D.Create a global secondary index on timestamp and write through the index
AnswerC

Write sharding by appending a random or calculated suffix to the partition key distributes a hot device's writes across multiple partitions, each with its own throughput budget. This directly relieves the per-partition write limit that causes throttling for high-volume devices. Queries must then fan out across the shards and aggregate results, but the change is confined to the key composition and read logic, keeping application change moderate and targeted.

Why this answer

Throttling from a few high-volume partition keys is a hot-partition problem, not a capacity problem. DynamoDB partitions have a fixed write ceiling regardless of table-level capacity mode. Write sharding by adding a suffix to the partition key spreads a single device's writes across many partitions, each with independent throughput.

Provisioned capacity, streams, and GSIs do not raise the per-partition write limit, so they fail to address the observed throttling.

Exam trap

The trap here is thinking that switching to provisioned capacity or raising WCUs fixes throttling, when the real constraint is the per-partition write limit on a single hot partition key.

764
MCQhard

A company is using Amazon EMR with Kerberos authentication. They want to ensure that data in transit between EMR cluster nodes is encrypted. Which configuration should be applied?

A.Use VPC peering to connect the cluster nodes.
B.Configure the EMR cluster to use in-transit encryption.
C.Enable S3 server-side encryption for the cluster's output data.
D.Enable EBS encryption on the cluster instances.
AnswerB

EMR security configurations enable in-transit encryption using TLS for inter-node communication, plus encryption of data within the cluster. This satisfies the requirement to encrypt data moving between EMR cluster nodes, which at-rest encryption and Kerberos alone do not provide.

Why this answer

In-transit encryption on Amazon EMR is enabled through a security configuration that specifies a PEM certificate and, optionally, a TLS certificate provider, which encrypts traffic between cluster nodes and to the EMR service. This is the only option that directly addresses data in transit between nodes.

Exam trap

The trap is confusing encryption at rest (EBS, S3 SSE) with encryption in transit; candidates often pick EBS or S3 encryption because they sound like security controls but do not secure node-to-node traffic.

How to eliminate wrong answers

Option A is wrong because VPC peering only provides network connectivity between VPCs; it does not encrypt traffic. Option C is wrong because S3 server-side encryption protects data at rest in S3, not data moving between EMR nodes. Option D is wrong because EBS encryption protects volumes at rest on the cluster instances, not network traffic.

765
MCQmedium

A data engineer is migrating an on-premises MongoDB database to Amazon DocumentDB. Which migration strategy minimizes downtime?

A.Take a snapshot of the MongoDB database and restore it to DocumentDB.
B.Use AWS Database Migration Service (AWS DMS) with full load only.
C.Export data using mongodump and import using mongorestore.
D.Use AWS DMS with full load and ongoing replication from MongoDB to DocumentDB.
AnswerD

AWS DMS performs a full load then applies ongoing change data capture replication from MongoDB to DocumentDB, keeping the target continuously synchronised until cutover. This continuous replication is what minimises downtime compared with one-off dump-and-restore migrations.

Why this answer

AWS DMS with full load and ongoing replication (change data capture) minimizes downtime by continuously synchronizing changes from the source MongoDB to the target DocumentDB after the initial full load, allowing a cutover with only a brief pause. This is the only option that supports near-zero downtime migration for live databases.

Exam trap

The trap here is that candidates assume any AWS DMS migration automatically minimizes downtime, but only the full load plus ongoing replication (CDC) option achieves near-zero downtime, while full load only still requires a write stop.

How to eliminate wrong answers

Option A is wrong because taking a snapshot and restoring it captures only a point-in-time copy, requiring the source database to be offline or read-only during the snapshot, causing downtime. Option B is wrong because AWS DMS full load only transfers the current data once, without capturing ongoing changes, so any writes during the migration are lost and downtime is needed to stop writes before cutover. Option C is wrong because mongodump and mongorestore are offline tools that require the source MongoDB to stop accepting writes during the export, resulting in significant downtime.

766
MCQeasy

A data engineer needs to ingest streaming data from thousands of IoT devices and immediately process each record with minimal latency. Which AWS service should be used as the ingestion point?

A.AWS Lambda
B.Amazon S3
C.Amazon Kinesis Data Streams
D.AWS Glue
AnswerC

Amazon Kinesis Data Streams ingests high-volume device telemetry with millisecond-level latency and preserves record order per shard, letting consumers process each record immediately. It satisfies the minimal-latency, per-record processing constraint, unlike Amazon S3 or Kinesis Data Firehose, which batch and buffer before delivery.

Why this answer

Amazon Kinesis Data Streams is designed as a massively scalable, low-latency ingestion service for streaming data, capable of handling thousands of producers (IoT devices) and delivering records to consumers within milliseconds. It durably stores data in shards for up to 365 days and integrates natively with Lambda, Kinesis Data Analytics, and Firehose for immediate processing. This makes it the canonical ingestion point for real-time IoT pipelines on AWS.

Exam trap

DEA-C01 often tests the distinction between ingestion services (Kinesis, MSK, IoT Core) and processing services (Lambda, Glue, EMR), so candidates mistakenly pick Lambda as the 'streaming' answer.

How to eliminate wrong answers

Option A is wrong because AWS Lambda is a compute service for processing events, not a durable ingestion buffer for thousands of concurrent producers; it can be a consumer of Kinesis but not the ingestion endpoint itself. Option B is wrong because Amazon S3 is object storage optimized for batch/throughput, not sub-second streaming ingestion, and it lacks native record-level streaming semantics. Option D is wrong because AWS Glue is a serverless ETL/catalog service for batch and micro-batch jobs, not a real-time streaming ingestion layer.

767
MCQhard

A company uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data is transformed using an AWS Lambda function. Recently, the transformation errors have increased due to Lambda timeouts. The data engineer needs to diagnose and resolve the issue without losing data. What should the engineer do?

A.Increase the Lambda function timeout and ensure that failed records are sent to a backup S3 bucket
B.Enable Amazon CloudWatch Logs for the Lambda function to capture errors and store failed records in CloudWatch
C.Configure the Lambda function to write failed records to an Amazon SQS queue for later reprocessing
D.Modify the Lambda function to store failed records in Amazon S3 before processing
AnswerA

Increasing timeout reduces failures, and configuring a backup bucket prevents data loss.

Why this answer

Increasing the Lambda function timeout directly addresses the root cause of transformation errors (timeouts), and configuring a backup S3 bucket for failed records ensures no data loss. Kinesis Data Firehose can be configured to send failed records to a separate S3 bucket as a dead-letter queue, which preserves the data for later reprocessing while the primary transformation pipeline is fixed.

Exam trap

The trap here is that candidates may confuse logging (CloudWatch Logs) with actual data preservation, or assume that SQS is a native Firehose failure destination, when in fact Firehose only supports S3 or Redshift as backup destinations for failed records.

How to eliminate wrong answers

Option B is wrong because enabling CloudWatch Logs captures error logs but does not store the actual failed records; it only provides visibility into errors without preventing data loss. Option C is wrong because Kinesis Data Firehose does not natively support sending failed records to an SQS queue; the Lambda function would need custom code to write to SQS, and this does not address the timeout issue. Option D is wrong because storing failed records in S3 before processing would require modifying the Lambda function to write to S3 first, which adds complexity and does not resolve the timeout; the records are already in the Firehose stream and need to be processed or redirected after failure.

768
MCQhard

A data engineer is using AWS Glue to catalog data stored in Amazon S3. The data is in Apache Parquet format and partitioned by year, month, and day. The engineer notices that AWS Glue crawlers are taking a long time to run and are not correctly identifying new partitions. The engineer needs to improve the crawler performance and ensure new partitions are added automatically. Which action should the engineer take?

A.Use AWS Glue partition projection instead of crawling for partitions.
B.Enable the crawler option to update the Data Catalog with new partitions only, and set a partition index.
C.Increase the number of data processing units (DPUs) for the crawler.
D.Configure the crawler to use a custom classifier for Parquet.
AnswerA

AWS Glue partition projection allows you to define partition patterns and ranges without running crawlers. This eliminates the need for crawlers to discover partitions, significantly improving performance and ensuring new partitions are automatically available. It is ideal for data with a known, regular partitioning scheme like year/month/day.

Why this answer

Partition projection in AWS Glue allows you to specify the partitioning scheme directly in the table properties, so the crawler does not need to discover partitions. This reduces crawler runtime and ensures that new partitions are immediately available for query engines like Athena. It is the most efficient solution for regularly partitioned data.

Exam trap

The trap here is assuming that increasing crawler resources or using custom classifiers will fix partition detection, but the real solution is to bypass crawling altogether with partition projection.

769
MCQmedium

A data engineering team is building an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another. The security team mandates that both read and write operations use a customer-managed AWS KMS key so they can audit key usage. Which configuration should the data engineer apply to the Glue job to meet this requirement?

A.Enable default encryption on the S3 bucket with the customer-managed KMS key and rely on Glue to use that default for all operations.
B.Modify the S3 bucket policy to deny any requests that do not include the aws:SecureTransport condition and require the KMS key in the request headers.
C.Add a job parameter --encryption-type sse-kms and specify the KMS key ARN in the job script, then set the Security configuration to use the same key.
D.Create an AWS Glue Security configuration that enables S3 encryption with the customer-managed KMS key for both reads and writes, and attach it to the job.
AnswerD

A Glue Security configuration allows you to specify a KMS key for S3 encryption, which Glue uses when reading from and writing to S3. By attaching this configuration to the job, all data access uses the specified customer-managed key, satisfying the audit requirement. This is the intended mechanism for controlling encryption in Glue jobs.

Why this answer

The correct approach is to use an AWS Glue Security configuration that specifies the customer-managed KMS key for S3 encryption. This configuration is applied at the job level and ensures that Glue uses the key for both reading and writing data. It provides a centralized way to enforce encryption and enables auditing of key usage.

Other methods like bucket policies or default encryption do not guarantee that the Glue job will use the specified key for all operations.

Exam trap

The trap here is assuming that S3 bucket default encryption or bucket policies alone will force AWS Glue to use a specific customer-managed KMS key for both reads and writes.

770
MCQmedium

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing sensitive customer records. The job must remove all columns flagged as PII before writing to a target S3 location. The PII columns are not known in advance and vary by file. Which approach should the engineer use to ensure the PII is removed dynamically?

A.Use AWS Glue's `DetectPII` transform to automatically identify and redact PII columns based on built-in patterns.
B.Use the AWS Glue DynamicFrame's `drop_fields` method with a list of PII column names derived from the Glue Data Catalog metadata.
C.Use AWS Glue DataBrew to create a recipe that identifies PII columns using pattern matching and apply it as a transformation step in the Glue job.
D.Use a custom PySpark script within the Glue job that scans column names and data samples against a regex pattern for PII, then drops matching columns.
AnswerD

A custom PySpark script can dynamically inspect each DataFrame's schema and sample data to apply regex patterns that identify PII-like content (e.g., email, SSN). It can then drop those columns before writing. This approach is flexible, scales with Glue's distributed processing, and handles varying PII columns per file, satisfying the requirement without relying on non-existent built-in transforms.

Why this answer

The requirement is to dynamically remove PII columns that vary by file. AWS Glue does not have a native PII detection transform, and the Data Catalog does not store PII classifications unless manually added. A custom PySpark script within the Glue job can inspect schema and data, apply regex patterns, and drop matching columns.

This provides the necessary flexibility and scalability.

Exam trap

The trap here is assuming AWS Glue has a built-in PII detection transform or that the Data Catalog automatically identifies PII columns.

771
MCQeasy

A company needs to ingest data from a MySQL database into Amazon S3 in near real-time. The database is running on EC2. The data engineer wants to minimize the impact on the source database. Which service should be used?

A.AWS Database Migration Service (DMS) with ongoing replication
B.AWS Glue ETL job with a JDBC connection
C.Amazon RDS for MySQL with read replica
D.AWS Schema Conversion Tool (SCT)
AnswerA

DMS with ongoing replication performs change data capture from the MySQL binary log, streaming inserts, updates and deletes to Amazon S3 continuously. Reading the log rather than polling tables keeps load on the source EC2 database minimal while achieving near real-time ingestion.

Why this answer

AWS DMS with ongoing replication (change data capture) is the correct choice because it can continuously replicate changes from a MySQL source database to Amazon S3 with minimal performance impact. DMS uses a transactional log-based approach (MySQL binlog) to capture changes as they occur, avoiding heavy SELECT queries on the source. This enables near real-time ingestion without adding significant load to the production database.

Exam trap

The trap here is that candidates often confuse AWS Glue's batch JDBC capabilities with streaming ingestion, or assume that a read replica can directly feed data into S3 without an intermediary service like DMS or Kinesis.

How to eliminate wrong answers

Option B is wrong because AWS Glue ETL jobs with JDBC connections run batch queries that pull full table snapshots or large result sets, which can cause significant performance degradation on the source MySQL database and cannot achieve near real-time latency. Option C is wrong because Amazon RDS for MySQL with a read replica is a database migration or read scaling solution, not a data ingestion service to S3; it does not natively stream data to S3 without additional tooling. Option D is wrong because AWS Schema Conversion Tool (SCT) is designed for converting database schemas between different database engines (e.g., Oracle to Aurora), not for ingesting data into S3.

772
MCQhard

A financial services company stores transaction data in Amazon RDS for PostgreSQL. The company requires that all changes to the database be logged for audit purposes, including before and after images of updated rows. Which feature should the data engineer enable?

A.Enable automated backups and export logs to Amazon S3
B.Enable Enhanced Monitoring and publish logs to CloudWatch Logs
C.Set up logical replication using pglogical or native publication/subscription
D.Enable Multi-AZ deployment and read replicas
AnswerC

Logical replication provides row-level changes with before and after images.

Why this answer

Logical replication, using either pglogical or native PostgreSQL publication/subscription, captures row-level changes (INSERT, UPDATE, DELETE) and can include both the old and new values of updated rows. This meets the audit requirement for before-and-after images, as logical replication decodes the write-ahead log (WAL) to produce a change stream that includes full row snapshots.

Exam trap

The trap here is that candidates confuse database-level logging features (like Enhanced Monitoring or automated backups) with row-level change data capture, assuming any logging mechanism will capture before-and-after images, when only logical replication (or triggers with audit tables) provides that granularity.

How to eliminate wrong answers

Option A is wrong because automated backups capture point-in-time snapshots of the entire database, not a continuous, row-level change stream with before-and-after images; exporting logs to S3 provides error logs or slow query logs, not row-level audit trails. Option B is wrong because Enhanced Monitoring collects OS-level metrics (CPU, memory, disk I/O) and publishes them to CloudWatch Logs, not database row changes. Option D is wrong because Multi-AZ deployment provides high availability via synchronous standby replication, and read replicas serve read traffic; neither logs individual row modifications or provides before-and-after images.

773
MCQeasy

A data engineer needs to audit all AWS KMS key usage in the account. Which AWS service should be used to record KMS API calls?

A.AWS CloudTrail
B.AWS Config
C.Amazon CloudWatch Logs
D.Amazon GuardDuty
AnswerA

AWS CloudTrail records every KMS API call as a management event, capturing the caller identity, key ARN, timestamp and source IP. This directly satisfies the audit requirement, since KMS operations such as Encrypt, Decrypt and CreateKey are logged automatically to the account's event history and any configured trail.

Why this answer

AWS CloudTrail records API calls for KMS. Option B (AWS Config) is wrong because it records resource changes, not API calls. Option C (Amazon CloudWatch Logs) is wrong because it stores logs but does not record API calls.

Option D (Amazon GuardDuty) is wrong because it is for threat detection.

774
MCQeasy

A data engineer needs to store archival data that is rarely accessed but must be retained for 7 years. The data should be retrievable within 12 hours. Which Amazon S3 storage class is MOST cost-effective?

A.S3 Intelligent-Tiering
B.S3 Glacier Flexible Retrieval
C.S3 Standard
D.S3 Glacier Deep Archive
AnswerD

S3 Glacier Deep Archive offers the lowest storage cost of all S3 classes and supports a standard retrieval time within 12 hours, matching the stated access and retrieval constraints. Its 7-year retention suitability makes it the most cost-effective choice for rarely accessed archival data.

Why this answer

S3 Glacier Deep Archive is the most cost-effective storage class for archival data that is rarely accessed and requires a 7-year retention period, with retrieval times up to 12 hours. It offers the lowest storage cost among S3 classes, making it ideal for long-term retention of data that does not need immediate access.

Exam trap

The trap here is that candidates often confuse S3 Glacier Flexible Retrieval (which offers faster retrieval but higher cost) with S3 Glacier Deep Archive, failing to recognize that the 12-hour retrieval requirement is easily met by Deep Archive's standard retrieval, making it the most cost-effective choice for long-term archival.

How to eliminate wrong answers

Option A is wrong because S3 Intelligent-Tiering is designed for data with unknown or changing access patterns, automatically moving data between tiers based on usage, which incurs monitoring and automation costs that are unnecessary for rarely accessed archival data. Option B is wrong because S3 Glacier Flexible Retrieval offers retrieval times from minutes to hours (typically 1-5 minutes for expedited, 3-5 hours for standard), but its storage cost is higher than Glacier Deep Archive, making it less cost-effective for data that only needs retrieval within 12 hours. Option C is wrong because S3 Standard is designed for frequently accessed data with millisecond retrieval times, and its storage cost is significantly higher than archival classes, making it prohibitively expensive for data that is rarely accessed and retained for 7 years.

775
MCQeasy

A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data should be stored cost-effectively. Which storage solution meets these requirements?

A.Amazon S3 Glacier Flexible Retrieval storage class.
B.Amazon S3 Glacier Deep Archive storage class.
C.Amazon S3 Standard-Infrequent Access (S3 Standard-IA) storage class.
D.Amazon S3 Standard storage class.
AnswerC

S3 Standard-IA matches the stated access pattern: infrequent retrieval with millisecond availability, satisfying the "within minutes" constraint. It costs less than S3 Standard for storage while retaining the same low-latency access, unlike Glacier tiers, which impose retrieval delays measured in minutes to hours.

Why this answer

Amazon S3 Standard-Infrequent Access (S3 Standard-IA) is the correct choice because it is designed for data accessed infrequently but requires rapid retrieval (within milliseconds). It offers lower storage costs than S3 Standard while maintaining low-latency access, meeting the requirement of retrievability within minutes cost-effectively.

Exam trap

The trap here is that candidates often confuse retrieval time with retrieval cost, assuming that 'infrequent access' implies slower retrieval, but S3 Standard-IA provides the same low-latency access as S3 Standard, unlike Glacier classes which have significantly longer retrieval times.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Glacier Flexible Retrieval is optimized for archival data where retrieval times range from minutes to hours, but it incurs higher retrieval costs and is not designed for frequent or rapid access within minutes. Option B is wrong because Amazon S3 Glacier Deep Archive is the lowest-cost storage class for long-term archival, but retrieval times are typically 12 hours or more, far exceeding the 'within minutes' requirement. Option D is wrong because Amazon S3 Standard is designed for frequently accessed data with high durability and low latency, but it is more expensive than S3 Standard-IA for infrequently accessed data, making it less cost-effective for this use case.

776
Multi-Selectmedium

A data engineer is monitoring an Amazon Kinesis Data Analytics for Apache Flink application that processes streaming data. The application is falling behind (increasing 'MillisBehindLatest') and the CPU utilization of the Flink task managers is consistently above 80%. Which THREE actions should the engineer take to improve performance? (Choose THREE.)

Select 3 answers
A.Increase the number of shards in the Kinesis data stream.
B.Decrease the checkpoint interval to reduce state size.
C.Enable auto-scaling for the Flink application.
D.Decrease the number of task managers to reduce CPU contention.
E.Increase the Flink application's parallelism.
AnswersA, C, E

Adding shards raises the stream's ingest and read parallelism, letting more Flink subtasks consume concurrently and reducing the backlog driving MillisBehindLatest. However, shards only relieve source-side throughput; the sustained CPU above 80% on task managers still requires scaling parallelism or KPUs, so this alone is insufficient.

Why this answer

Option A is correct because increasing the number of shards in the Kinesis data stream raises the stream's read throughput capacity (each shard supports up to 1 MB/s or 1,000 records/s for reads), which directly addresses the growing MillisBehindLatest by allowing the Flink source to consume data faster. Option C is correct because enabling auto-scaling for the Flink application (via Kinesis Data Analytics' automatic scaling of KPUs) dynamically adds parallel task slots and CPU resources when utilization is high, relieving the sustained >80% CPU pressure on task managers. Option E is correct because increasing the Flink application's parallelism distributes the workload across more parallel subtasks and task slots, lowering per-task CPU load and improving overall processing throughput.

Option B is not appropriate because shortening the checkpoint interval increases checkpointing overhead and I/O rather than reducing state size, potentially worsening performance. Option D is wrong because decreasing the number of task managers reduces available CPU and parallelism, which would increase contention and further degrade the application.

Exam trap

The trap here is that candidates often confuse decreasing checkpoint intervals with improving performance, not realizing that more frequent checkpoints increase CPU and I/O overhead, making the lag worse.

777
Multi-Selecteasy

A data engineer is setting up a new Amazon Redshift cluster for a data warehouse. The engineer wants to ensure data durability and high availability. Which THREE features should the engineer consider? (Choose three.)

Select 3 answers
A.S3 Cross-Region Replication for Redshift data.
B.Cross-Region snapshot copy.
C.Multi-node cluster with data replication.
D.Multi-AZ deployment for automatic failover.
E.Automated snapshots to Amazon S3.
AnswersB, C, E

Cross-Region snapshot copy replicates automated and manual snapshots to a secondary AWS Region, so cluster data survives a full regional outage. This directly satisfies the durability and high availability requirement, since Redshift snapshots are stored in Amazon S3 and copying them across Regions provides disaster recovery beyond a single Region's failure.

Why this answer

Option B (Cross-Region snapshot copy) is correct because Redshift can automatically copy cluster snapshots to an S3 bucket in another AWS Region, providing disaster recovery and durability if an entire Region becomes unavailable. Option C (Multi-node cluster with data replication) is correct because a multi-node Redshift cluster distributes data across compute nodes and maintains replicas of each block on other nodes, so the cluster can continue serving queries if a node or disk fails. Option E (Automated snapshots to Amazon S3) is correct because Redshift automatically takes snapshots of the cluster and stores them in Amazon S3, enabling point-in-time recovery and restoring the cluster if data is lost or corrupted.

Option A is not correct because S3 Cross-Region Replication is an S3 feature for replicating S3 objects and is not a Redshift durability mechanism. Option D is not correct because Redshift does not offer a Multi-AZ deployment mode with automatic failover; high availability within a Region is achieved through multi-node replication and automated snapshots.

Exam trap

The trap is assuming Redshift offers RDS-style Multi-AZ automatic failover, when its availability model is multi-node replication plus snapshots and cross-region copy.

778
MCQmedium

A company stores sensitive data in Amazon S3. They need to ensure that all objects are encrypted at rest. Which approach meets this requirement with minimal effort?

A.Use client-side encryption before uploading
B.Enable default encryption on the S3 bucket with SSE-S3
C.Enable S3 Versioning and MFA Delete
D.Use a bucket policy to deny PutObject without encryption
AnswerB

Enabling SSE-S3 default encryption on the bucket applies AES-256 encryption to every object at rest automatically, including future uploads, with no application changes. This meets the requirement with minimal operational effort compared to client-side or KMS-based approaches.

Why this answer

Enabling default encryption on an S3 bucket with SSE-S3 (Server-Side Encryption with S3-Managed Keys) automatically encrypts all objects at rest using AES-256, with no additional effort from the user. This ensures that any object uploaded without explicit encryption headers is encrypted by default, meeting the requirement with minimal configuration overhead.

Exam trap

The trap here is that candidates often confuse enforcing encryption at upload (via bucket policy) with automatically encrypting data at rest, leading them to choose option D, which requires additional policy management and does not guarantee encryption of all objects without explicit headers.

How to eliminate wrong answers

Option A is wrong because client-side encryption requires the application to manage encryption keys and perform encryption before upload, adding significant operational effort and complexity, which contradicts the 'minimal effort' requirement. Option C is wrong because S3 Versioning and MFA Delete provide data protection against accidental deletion and overwrites, but they do not encrypt objects at rest. Option D is wrong because a bucket policy to deny PutObject without encryption only enforces encryption on upload but does not encrypt existing objects or objects uploaded without the required headers; it also requires additional policy management and does not automatically encrypt data at rest.

779
MCQmedium

A data engineer needs to create a table in Amazon Athena that reads JSON data stored in Amazon S3. The JSON records are stored in a single file, one JSON object per line. The engineer wants Athena to automatically discover the schema and create the table without manually defining columns. Which AWS service or feature should the engineer use?

A.Amazon Athena CREATE TABLE AS SELECT (CTAS) statement
B.AWS Glue crawler
C.AWS Glue DataBrew
D.Amazon S3 Inventory
AnswerB

AWS Glue crawler scans data in S3, infers the schema, and populates the AWS Glue Data Catalog. Athena can then query the table using the catalog metadata. This meets the requirement of automatic schema discovery without manual column definition.

Why this answer

An AWS Glue crawler automatically scans data in Amazon S3, infers the schema, and creates table definitions in the AWS Glue Data Catalog. Athena uses this catalog to query the data without manual schema definition. This is the standard method for automatic schema discovery in a data lake.

Exam trap

The trap here is assuming that Athena can automatically infer schemas from raw data without a crawler or manual DDL.

780
MCQeasy

A data engineer is designing a data lake on Amazon S3. The data includes customer PII that must be encrypted at rest. The company also requires that the encryption keys be rotated automatically every year. Which encryption solution should the engineer use?

A.SSE-KMS with automatic key rotation enabled
B.SSE-S3
C.SSE-C
D.Client-side encryption with AWS KMS
AnswerA

SSE-KMS encrypts objects at rest using AWS KMS keys and supports automatic annual key rotation, meeting both the PII encryption and rotation requirements. SSE-S3 rotation is managed by AWS without configurable schedules, so it cannot satisfy the stated yearly rotation constraint.

Why this answer

SSE-KMS with automatic key rotation enabled meets both requirements: it encrypts data at rest in S3 and allows the company to automatically rotate the customer master key (CMK) every year. AWS KMS supports automatic annual rotation for symmetric CMKs, which satisfies the compliance need without manual intervention.

Exam trap

The trap here is that candidates often assume SSE-S3 provides customer-controlled key rotation, but SSE-S3 uses keys fully managed by AWS with no customer control over rotation frequency, whereas SSE-KMS with automatic key rotation lets the company meet the explicit annual rotation requirement.

How to eliminate wrong answers

Option B (SSE-S3) is wrong because while it encrypts data at rest, it does not support automatic key rotation; the encryption keys are managed and rotated by S3 but not on a customer-defined schedule. Option C (SSE-C) is wrong because it requires the customer to provide and manage their own encryption keys, and AWS does not handle key rotation, making it unsuitable for automated annual rotation. Option D (Client-side encryption with AWS KMS) is wrong because it encrypts data before sending it to S3, but the key rotation applies only to the KMS key used for client-side encryption, not to the S3-side encryption; moreover, client-side encryption adds complexity and does not directly address the requirement for encryption at rest within S3.

781
Multi-Selecteasy

A data engineer is designing a data pipeline that processes streaming data. The pipeline must be able to handle duplicate records and ensure exactly-once processing semantics. Which THREE AWS services or features should the engineer consider? (Choose three.)

Select 3 answers
A.Amazon EMR with Apache Flink for exactly-once semantics.
B.Amazon Kinesis Data Firehose with automatic retries.
C.Amazon Kinesis Data Streams with sequence numbers for deduplication.
D.Amazon DynamoDB Streams for change data capture.
E.Amazon Kinesis Data Analytics for Apache Flink with idempotent sinks.
AnswersA, C, E

Apache Flink on Amazon EMR provides checkpointing and two-phase commit sinks, giving exactly-once state consistency across operator failures. This satisfies the stem's requirement to tolerate duplicate records while guaranteeing each record affects downstream state only once.

Why this answer

Option A is correct because Amazon EMR running Apache Flink supports exactly-once state consistency via its checkpointing mechanism (distributed snapshots aligned with barriers), so a streaming pipeline built on Flink can guarantee exactly-once processing even with duplicate or replayed records. Option C is correct because Kinesis Data Streams assigns each record a unique sequence number within a shard, and consumers can use those sequence numbers (or the KCL's checkpointing) to detect and skip already-processed records, enabling deduplication and effectively-once consumption. Option E is correct because Kinesis Data Analytics for Apache Flink inherits Flink's exactly-once checkpointing and, when combined with idempotent sinks (e.g., writing to DynamoDB with conditional writes or to S3 with deterministic keys), prevents duplicate records from producing duplicate downstream effects.

Option B is not appropriate because Kinesis Data Firehose delivers at-least-once and its automatic retries can actually introduce duplicates rather than deduplicate them. Option D is not appropriate because DynamoDB Streams provides change data capture with at-least-once delivery and no built-in deduplication or exactly-once guarantee for a streaming pipeline.

Exam trap

DEA-C01 often tests the misconception that Firehose retries provide exactly-once, but retries can cause duplicates; candidates must recognize that exactly-once requires specific features like Flink checkpointing and idempotent sinks.

782
MCQmedium

A data engineer is troubleshooting an AWS Glue ETL job that fails with a 'java.lang.OutOfMemoryError: Java heap space' error. The job processes a 50 GB Parquet file from an S3 bucket. The job uses a G.1X DPU (16 GB memory) and default parameters. Which action should the engineer take to resolve the issue?

A.Change the worker type to G.2X (32 GB memory).
B.Increase the number of workers from 2 to 4.
C.Increase the 'batch size' parameter in the DynamicFrame reader.
D.Convert the input data from Parquet to JSON format.
AnswerA

The OutOfMemoryError arises because each G.1X executor has only 16 GB of heap for the 50 GB Parquet dataset. Switching to G.2X doubles memory to 32 GB per worker, giving the JVM sufficient heap to process the partitions without exhausting memory during the shuffle.

Why this answer

Changing the worker type to G.2X (32 GB memory) doubles the memory per worker, directly addressing the Java heap space error. Option B is incorrect because increasing the number of workers does not increase memory per worker; each G.1X DPU still has only 16 GB. Option C is incorrect because increasing the batch size would increase the amount of data loaded into memory per worker, potentially worsening the memory issue.

Option D is incorrect because converting from Parquet to JSON typically increases file size and memory usage due to lack of compression and columnar storage.

783
MCQhard

A company uses Amazon DynamoDB with provisioned capacity. During a sales event, write traffic spikes and some requests receive ProvisionedThroughputExceeded exceptions. The reads are within limits. The data engineer needs to minimize latency for the spike without manual intervention. Which solution is MOST cost-effective?

A.Use Amazon SQS to buffer write requests and process them in batches.
B.Disable auto scaling and set write capacity to the peak observed value.
C.Enable DynamoDB auto scaling for write capacity with a target utilization of 70%.
D.Enable DynamoDB Accelerator (DAX) to cache write operations.
AnswerC

DynamoDB auto scaling adjusts provisioned write capacity automatically in response to CloudWatch utilisation metrics, directly resolving the ProvisionedThroughputExceeded exceptions during the spike without manual intervention. A 70% target keeps headroom while avoiding over-provisioning, satisfying the cost-effectiveness constraint. Reads already sit within limits, so scaling writes alone is sufficient.

Why this answer

DynamoDB auto scaling for write capacity automatically adjusts the provisioned write capacity units (WCUs) based on the actual traffic pattern, using a target utilization of 70% to balance cost and performance. This eliminates manual intervention and handles spikes efficiently by scaling up before throttling occurs, while remaining cost-effective since capacity scales down when traffic subsides.

Exam trap

The trap here is that candidates may confuse DAX as a solution for write performance, but DAX only accelerates reads (via caching) and does not mitigate write throttling, leading to an incorrect choice of Option D.

How to eliminate wrong answers

Option A is wrong because using Amazon SQS to buffer write requests introduces additional latency for processing batches, which contradicts the requirement to minimize latency during the spike, and it adds complexity and cost for queue management. Option B is wrong because disabling auto scaling and setting write capacity to the peak observed value is wasteful and costly, as it permanently allocates high capacity that is only needed during spikes, and it requires manual intervention to adjust. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for read operations, not writes; it does not reduce write throttling or ProvisionedThroughputExceeded exceptions, and it adds cost without addressing the write capacity issue.

784
MCQhard

A data engineer is designing a streaming pipeline that ingests IoT sensor data from 10,000 devices. Each device sends a 1 KB message every second. The data must be processed in near real-time and stored in S3 for analytics. Which combination of services provides the most cost-effective solution?

A.AWS Data Pipeline with periodic S3 copy.
B.Amazon Kinesis Data Streams with Kinesis Data Firehose delivery to S3.
C.Amazon MSK (Managed Streaming for Kafka) with Kafka Connect S3 sink.
D.Amazon SQS FIFO queue with Lambda consumers writing to S3.
AnswerB

Handles high throughput, Firehose batches to S3.

Why this answer

B is correct because Kinesis Data Streams ingests high-throughput IoT data (10,000 messages/sec at 1 KB each) with low latency, and Kinesis Data Firehose automatically batches and compresses data before delivering it to S3, eliminating the need for custom code or manual scaling. This combination provides the most cost-effective near-real-time solution by leveraging Firehose's built-in buffering and compression to minimize S3 storage costs and reduce the number of PUT requests.

Exam trap

The trap here is that candidates often choose MSK (Option C) thinking it is more scalable or flexible, but they overlook the higher operational cost and complexity for a simple S3 sink use case, where Kinesis Data Firehose's fully managed batching, compression, and direct S3 integration is more cost-effective and simpler to maintain.

How to eliminate wrong answers

Option A is wrong because AWS Data Pipeline is a batch-oriented orchestration service, not designed for near-real-time streaming ingestion; it would introduce significant latency and require periodic polling, failing the near-real-time requirement. Option C is wrong because Amazon MSK (Managed Streaming for Apache Kafka) introduces unnecessary operational overhead and cost for this use case, as it requires managing Kafka clusters, brokers, and the Kafka Connect S3 sink, while Kinesis Data Firehose provides a simpler, fully managed, and more cost-effective direct S3 delivery. Option D is wrong because SQS FIFO queues are designed for exactly-once processing and low throughput (300 transactions per second by default), making them unsuitable for 10,000 messages per second; additionally, Lambda consumers would incur high costs due to the large number of invocations and lack built-in batching and compression for S3 writes.

785
MCQhard

A company wants to audit all changes to IAM policies in their AWS account. Which combination of services should be used to achieve this?

A.AWS Config and Amazon SNS
B.AWS CloudTrail and Amazon CloudWatch Logs
C.Amazon CloudWatch Logs and Amazon SNS
D.AWS CloudTrail and Amazon DynamoDB
AnswerB

AWS CloudTrail records every IAM policy change as a management event, capturing the API caller, timestamp and request parameters. Streaming those trails into Amazon CloudWatch Logs satisfies the audit requirement, enabling metric filters and alarms on specific `PutPolicy` or `CreatePolicy` calls for continuous monitoring.

Why this answer

AWS CloudTrail records all API calls, including IAM policy changes, and can deliver logs to Amazon CloudWatch Logs for monitoring and alerting. This combination provides a comprehensive audit trail and real-time analysis capability.

Exam trap

The trap is confusing AWS Config with CloudTrail. Config is for configuration history, not API auditing. Candidates might also think CloudWatch Logs alone can capture API calls without CloudTrail.

How to eliminate wrong answers

Option A is wrong because AWS Config records resource configurations but not API calls; it can track IAM policy changes but lacks the detailed event history of CloudTrail. Option C is wrong because CloudWatch Logs alone does not capture API activity; it needs a source like CloudTrail. Option D is wrong because DynamoDB is a database, not an audit or logging service.

786
MCQeasy

A data engineer creates an Amazon DynamoDB table using the CloudFormation snippet in the exhibit. The application writes 200 items per second to the table. The engineer notices that many write requests are being throttled. What is the MOST likely reason?

A.The table does not have a sort key, causing hot partitions.
B.The attribute type for OrderID should be numeric for better performance.
C.The table name 'Orders' conflicts with an existing table.
D.The provisioned write capacity is too low for the application's write rate.
AnswerD

Provisioned write capacity is set below the 200 items per second the application writes, so DynamoDB throttles excess requests once consumed capacity units are exhausted. DynamoDB rejects writes exceeding the table's provisioned WCU rather than queuing them, making insufficient capacity the direct cause of the throttling observed.

Why this answer

The DynamoDB table is provisioned with a fixed write capacity, and the application writes 200 items per second. If the provisioned write capacity units (WCUs) are less than required, throttling occurs. Each write of an item up to 1 KB consumes 1 WCU, so 200 writes per second require at least 200 WCUs.

If the table's provisioned write capacity is below this, requests will be throttled.

Exam trap

DEA-C01 often tests the misconception that missing sort keys or attribute types cause throttling, when the primary cause is insufficient provisioned capacity or hot partitions due to poor partition key design.

How to eliminate wrong answers

Option A is wrong because the absence of a sort key does not inherently cause hot partitions; hot partitions occur when the partition key has low cardinality or uneven access patterns. Option B is wrong because the attribute type for OrderID (numeric vs. string) does not affect performance; it only affects storage and query capabilities. Option C is wrong because a table name conflict would prevent creation, not cause throttling; CloudFormation would fail if the table already exists.

787
MCQeasy

A company stores critical financial data in Amazon DynamoDB. To meet compliance requirements, the data must be encrypted at rest with a customer-managed key. Which solution should the data engineer implement?

A.Configure the DynamoDB table to use a customer managed key from AWS KMS.
B.Use AWS CloudHSM to generate a key and import it into DynamoDB.
C.Enable default encryption on the DynamoDB table using S3-managed keys.
D.Use AWS Certificate Manager to issue a certificate and configure TLS.
AnswerA

DynamoDB supports encryption at rest using AWS KMS customer managed keys, configured per table. Selecting a customer managed key rather than the AWS owned default key satisfies the compliance requirement for a key the company controls and can audit.

Why this answer

AWS DynamoDB integrates with AWS KMS to support encryption at rest using customer-managed keys (CMKs). By selecting a customer-managed key from KMS during table creation or via an update, the company can meet compliance requirements for controlling the encryption key lifecycle, including rotation and access policies. This approach ensures that the data is encrypted using AES-256 encryption, with the key material managed by the customer rather than AWS.

Exam trap

The trap here is that candidates may confuse encryption at rest with encryption in transit (TLS) or assume that CloudHSM can be directly used with DynamoDB, when in fact DynamoDB only supports KMS for encryption at rest, and CloudHSM requires a custom integration layer.

How to eliminate wrong answers

Option B is wrong because AWS CloudHSM provides hardware-based key storage but does not directly integrate with DynamoDB for encryption at rest; DynamoDB only supports KMS keys, not CloudHSM-generated keys imported via the HSM. Option C is wrong because DynamoDB does not use S3-managed keys; default encryption in DynamoDB uses AWS-owned keys or AWS-managed keys, not S3-managed keys, and S3-managed keys are specific to Amazon S3. Option D is wrong because AWS Certificate Manager and TLS are used for encryption in transit, not encryption at rest, and do not address the requirement for encrypting stored data with a customer-managed key.

788
Multi-Selecthard

A data engineer is designing a data lake on Amazon S3 using AWS Lake Formation. The engineer needs to grant fine-grained access to specific columns and rows of a table to different analysts. Which two actions should the engineer take to meet these requirements? (Choose two.)

Select 2 answers
A.Define a Lake Formation data filter that specifies column and row expressions.
B.Use AWS Glue crawlers to update the Data Catalog with new partitions.
C.Register the S3 data location with Lake Formation.
D.Enable S3 Block Public Access on the bucket.
E.Create an IAM policy that allows s3:GetObject on the entire bucket.
AnswersA, C

Data filters in Lake Formation allow you to define column-level and row-level access controls. You can grant permissions on a filtered view of the table, restricting which columns and rows an analyst can see. This directly meets the requirement for fine-grained access.

Why this answer

To enforce column- and row-level security with Lake Formation, the engineer must first register the S3 location so Lake Formation can manage the data. Then, data filters can be created to define which columns and rows are accessible. These two steps together enable fine-grained permissions, while other options do not provide the required access controls.

Exam trap

The trap here is assuming that broad IAM policies or S3 Block Public Access can provide fine-grained access, when Lake Formation requires explicit registration and data filters.

789
MCQmedium

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes only new data added since the last run. The engineer notices that the job is reprocessing old data, increasing runtime and cost. Which AWS Glue feature should be enabled to ensure only new data is processed?

A.AWS Glue DataBrew
B.Job bookmarks
C.Amazon S3 event notifications
D.AWS Glue Crawler incremental crawl
AnswerB

AWS Glue job bookmarks track the data that has already been processed in previous runs. When enabled, the job reads only new data based on the last modified time of objects in S3. This prevents reprocessing of old data, reducing runtime and cost. It is the standard feature for incremental data processing in AWS Glue.

Why this answer

Job bookmarks are the AWS Glue feature designed to track processed data and enable incremental processing. When enabled, the job reads only new data based on S3 object metadata. This reduces runtime and cost by avoiding reprocessing.

The other options do not provide this capability: DataBrew is for data preparation, crawler incremental crawl updates the catalog, and S3 event notifications are for triggering actions.

Exam trap

The trap here is confusing incremental crawls or event-driven triggers with job bookmarks, which are specifically needed for incremental ETL processing within AWS Glue.

790
MCQmedium

A data engineer is using Amazon Redshift and needs to improve the performance of complex queries that join large tables. The engineer has already set the distribution style to KEY on the join columns. What additional step should the engineer take to optimize the join performance?

A.Change the distribution style to ALL for all large tables.
B.Enable concurrency scaling to handle concurrent queries.
C.Set the sort key on the columns used in the join and in the WHERE clause.
D.Increase the number of nodes in the Redshift cluster to add more compute resources.
AnswerC

Sort keys enable efficient range scans and merge joins by co-locating data on disk in sorted order. When joining large tables, having the join columns as sort keys can allow Redshift to use merge joins instead of hash joins, reducing data movement and improving performance. This is a best practice for join optimization.

Why this answer

Sort keys on join and filter columns allow Redshift to perform merge joins and skip unnecessary data blocks, reducing I/O and data movement. This optimizes join performance without adding resources. The other options either scale resources or change distribution in ways that are not optimal for large tables.

Exam trap

The trap here is assuming that adding more nodes or enabling concurrency scaling will automatically improve join performance, when the real gain comes from sort keys and proper data layout.

791
MCQhard

A company is using Amazon DynamoDB with on-demand capacity for a gaming application. During a new game launch, write traffic spikes to 50,000 writes per second, but the application experiences throttling. The DynamoDB table has a partition key of 'game_id' and a sort key of 'timestamp'. What is the MOST likely cause of throttling?

A.The table has not enabled auto-scaling for writes.
B.The table's on-demand capacity is insufficient for the write spike.
C.Hot partitions due to a skewed access pattern on the partition key 'game_id'.
D.The sort key is not optimal for write-heavy workloads.
AnswerC

Skewed writes concentrate on a few popular game_id values, so individual partitions exceed their per-partition throughput ceiling even though on-demand capacity scales table-wide. DynamoDB distributes capacity per partition, and a single partition key value cannot exceed roughly 1,000 write capacity units per second, causing throttling despite aggregate headroom.

Why this answer

DynamoDB on-demand capacity automatically scales to handle traffic spikes, but it still has per-partition throughput limits. With 'game_id' as the partition key, a single popular game can create a hot partition where all writes target the same partition, exceeding the partition's maximum write capacity (1,000 write capacity units per partition) and causing throttling, even though the overall table capacity is sufficient.

Exam trap

The trap here is that candidates assume on-demand capacity eliminates all throttling, but they overlook DynamoDB's per-partition throughput limits, which can cause throttling on hot partitions even with on-demand mode.

How to eliminate wrong answers

Option A is wrong because on-demand capacity does not use auto-scaling; it automatically adjusts capacity without needing auto-scaling enabled. Option B is wrong because on-demand capacity is designed to handle sudden spikes without manual provisioning, so insufficient capacity is not the issue—the problem is partition-level limits. Option D is wrong because the sort key does not affect write throughput distribution; partition key selection determines write distribution, and a sort key is irrelevant to throttling caused by hot partitions.

792
MCQmedium

A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for ad-hoc querying using SQL. Which AWS service should be used?

A.Amazon Athena with data in S3
B.Amazon DynamoDB
C.Amazon RDS for MySQL
D.Amazon Kinesis Data Analytics
AnswerA

Amazon Athena queries JSON directly from S3 using schema-on-read, so no transformation or loading is needed before ad-hoc SQL analysis. S3 provides the cheapest durable storage for semi-structured logs, satisfying the cost-effectiveness constraint, while Athena's pay-per-query model avoids provisioning clusters for intermittent querying.

Why this answer

Amazon Athena is the correct choice because it allows you to query semi-structured JSON logs stored in S3 directly using standard SQL, without needing to load or transform the data. Athena's schema-on-read approach and pay-per-query pricing make it highly cost-effective for ad-hoc analysis of large volumes of log data, as you only pay for the data scanned during queries.

Exam trap

The trap here is that candidates often confuse Amazon Athena with Amazon Kinesis Data Analytics, mistakenly thinking that Kinesis is the go-to service for SQL-based log analysis, when in fact Kinesis is for real-time streaming and Athena is the correct serverless query service for stored data in S3.

How to eliminate wrong answers

Option B (Amazon DynamoDB) is wrong because it is a NoSQL key-value and document database optimized for low-latency, high-throughput transactional workloads, not for ad-hoc SQL querying of semi-structured logs; it lacks native SQL support and would require expensive scanning of large datasets. Option C (Amazon RDS for MySQL) is wrong because it requires you to predefine a schema, load the JSON logs into relational tables, and pay for provisioned compute and storage even when idle, making it less cost-effective for sporadic ad-hoc queries compared to Athena's serverless model. Option D (Amazon Kinesis Data Analytics) is wrong because it is designed for real-time stream processing and analytics on streaming data using SQL, not for querying stored JSON logs in S3; it would require continuous ingestion and incurs ongoing costs regardless of query frequency.

793
MCQmedium

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job must run incrementally and process only new data since the last run. The source data is partitioned by date in S3, and new partitions are added daily. Which AWS Glue feature should the engineer enable to track previously processed data?

A.Data Catalog partition index
B.Glue DataBrew profile
C.Glue FindMatches
D.Job bookmarks
AnswerD

Job bookmarks in AWS Glue track the state of data that has already been processed. When enabled, Glue stores bookmark state for each job run, allowing the job to process only new data since the last successful run. For S3 sources, bookmarks track objects and partitions based on their last modified time and other metadata. This directly addresses the requirement to process only new data incrementally without reprocessing old partitions.

Why this answer

Job bookmarks are the AWS Glue feature designed to track previously processed data. When enabled, Glue maintains state for each job, so subsequent runs process only new data, such as new S3 partitions or files. This is the standard way to implement incremental ETL in Glue without reprocessing historical data.

Exam trap

The trap here is confusing Data Catalog partition indexes, which optimize metadata queries, with job bookmarks, which track processed data for incremental runs.

794
MCQeasy

A company is using Amazon S3 as a data lake. The data engineer needs to ensure that all objects uploaded to a specific bucket are automatically replicated to a bucket in another AWS Region for disaster recovery. Which configuration should the engineer implement?

A.Enable S3 Same-Region Replication (SRR) on the source bucket.
B.Enable S3 Cross-Region Replication (CRR) on the source bucket.
C.Use S3 Transfer Acceleration to copy objects to the destination.
D.Use S3 Batch Operations to copy existing objects.
AnswerB

S3 Cross-Region Replication automatically copies objects to a bucket in a different AWS Region, satisfying the disaster recovery requirement. Configuring CRR on the source bucket replicates new uploads asynchronously, providing cross-Region durability without application changes. Versioning must be enabled on both source and destination buckets for replication to function.

Why this answer

S3 Cross-Region Replication (CRR) is the correct choice because it automatically replicates objects from a source bucket in one AWS Region to a destination bucket in a different AWS Region, meeting the disaster recovery requirement for geographic separation. CRR requires versioning to be enabled on both buckets and replicates new objects asynchronously after upload.

Exam trap

The trap here is that candidates confuse S3 Transfer Acceleration (which speeds up uploads) with replication, or assume S3 Batch Operations can be used for ongoing replication, when only CRR provides automatic, cross-region object replication for disaster recovery.

How to eliminate wrong answers

Option A is wrong because S3 Same-Region Replication (SRR) replicates objects within the same AWS Region, not across regions, so it does not provide disaster recovery across geographic boundaries. Option C is wrong because S3 Transfer Acceleration speeds up uploads over long distances using AWS edge locations but does not replicate objects to another bucket; it only improves transfer performance for clients. Option D is wrong because S3 Batch Operations is used for bulk actions like copying existing objects or tagging, but it is a one-time operation, not an automatic, ongoing replication configuration for new objects.

795
MCQmedium

A company needs to enforce encryption in transit for all data moving between its Amazon S3 bucket and a fleet of Amazon EC2 instances. The data is accessed via S3 API calls over the internet. Which configuration ensures encryption in transit?

A.Enable SSE-S3 on the bucket.
B.Enable S3 Transfer Acceleration.
C.Use a VPC endpoint for S3.
D.Configure the bucket policy to deny requests that do not use HTTPS.
AnswerD

A bucket policy with a `aws:SecureTransport` condition set to false denies any request made over plain HTTP, forcing all S3 API calls from the EC2 fleet to use HTTPS/TLS. This directly enforces encryption in transit for internet-based access, satisfying the stem's requirement.

Why this answer

Configuring the bucket policy to deny requests that do not use HTTPS ensures encryption in transit for S3 API calls. Option A is wrong because SSE-S3 encrypts data at rest, not in transit. Option B is wrong because S3 Transfer Acceleration uses a global network but does not enforce encryption; HTTPS must still be used.

Option C is wrong because a VPC endpoint for S3 uses AWS network but does not enforce encryption; the bucket policy must explicitly require HTTPS.

796
MCQhard

A company is using Amazon DynamoDB with auto scaling enabled. During a marketing campaign, write traffic spikes, and some write requests fail with ProvisionedThroughputExceededException. The auto scaling policy has a target utilization of 70% and a maximum capacity that is high enough. What is the most likely cause of the throttling?

A.The table has a global secondary index that is throttling.
B.Auto scaling cannot react quickly enough to sudden traffic spikes.
C.The table does not have enough maximum capacity.
D.The auto scaling policy is not configured correctly.
AnswerB

DynamoDB auto scaling adjusts capacity through CloudWatch alarms over minutes, so it lags behind abrupt spikes. The maximum capacity being sufficient is irrelevant; the reaction delay itself causes ProvisionedThroughputExceededException when demand outpaces the scaling response.

Why this answer

Auto scaling in DynamoDB adjusts capacity based on the average utilization over a period (typically 5-10 minutes). When a sudden traffic spike occurs, the write requests can exceed the current provisioned capacity before the auto scaling policy has time to react and increase the capacity. This delay causes ProvisionedThroughputExceededException errors, even though the maximum capacity is set high enough.

Exam trap

The trap here is that candidates assume a correctly configured auto scaling policy with sufficient maximum capacity will always prevent throttling, ignoring the inherent latency in auto scaling's reaction to sudden, short-lived traffic spikes.

How to eliminate wrong answers

Option A is wrong because a throttling global secondary index (GSI) would cause its own ProvisionedThroughputExceededException, but the question states the write requests fail directly on the table, and a GSI throttling would typically manifest as errors on writes that affect the index, not necessarily all table writes. Option C is wrong because the question explicitly states that the maximum capacity is high enough, so insufficient maximum capacity is not the cause. Option D is wrong because the auto scaling policy is configured with a target utilization of 70% and a high enough maximum capacity, which is a standard and correct configuration; the issue is the inherent lag in auto scaling's response to sudden spikes, not a misconfiguration.

797
Multi-Selectmedium

A company is building a data pipeline that ingests sensitive customer data from an on-premises database into Amazon S3 using AWS DMS. The data must be encrypted at rest in S3 and in transit. The security team requires that the encryption keys be managed by the company (not AWS). Which TWO actions should the data engineer take to meet these requirements? (Choose TWO.)

Select 2 answers
A.Enable encryption at rest using the default DMS encryption settings.
B.Configure the S3 bucket to use server-side encryption with AWS KMS (SSE-KMS) using a customer managed key.
C.Configure the S3 bucket to use server-side encryption with S3 managed keys (SSE-S3).
D.Enable SSL/TLS encryption on the DMS source and target endpoints.
E.Create an AWS KMS key and use it in the DMS endpoint to encrypt data in transit.
AnswersB, D

SSE-KMS with a customer managed key encrypts objects at rest while the company retains control over key rotation and access policy, satisfying the requirement that keys not be AWS-managed. It also complements TLS for the in-transit requirement.

Why this answer

Option B is correct because SSE-KMS with a customer managed key lets the company own and control the KMS key that protects the data at rest in Amazon S3, satisfying the requirement that keys be managed by the company rather than AWS. Option D is correct because enabling SSL/TLS on the DMS source and target endpoints encrypts data in transit between the on-premises database and Amazon S3, which is the required in-transit protection. Option A is incorrect because the default DMS encryption settings do not provide company-managed keys for S3 data at rest.

Option C is incorrect because SSE-S3 uses S3-managed keys, which AWS controls, not the company. Option E is incorrect because KMS keys are used for encryption at rest, not for encrypting DMS data in transit; in-transit encryption is handled by SSL/TLS on the endpoints.

Exam trap

The trap here is that candidates often confuse encryption at rest with encryption in transit, and mistakenly think that KMS keys can be used for both, or that default DMS encryption or SSE-S3 satisfies the customer-managed key requirement.

798
MCQhard

A company is designing a data lake on Amazon S3. The data includes personal identifiable information (PII). The data engineer must ensure that only authorized users can access the data, and that access is logged for auditing. Which combination of services should the data engineer use?

A.S3 bucket policies with IAM policies and AWS CloudTrail with data events
B.Amazon S3 access points and VPC endpoints
C.Amazon Macie to discover PII and S3 Object Lock to prevent deletion
D.AWS KMS to encrypt data and AWS CloudTrail to log access
AnswerA

S3 bucket policies and IAM policies together enforce least-privilege authorisation on the PII objects, while CloudTrail data events capture object-level API activity such as GetObject, satisfying the auditing requirement that management events alone would not record.

Why this answer

S3 bucket policies combined with IAM policies provide fine-grained access control to restrict who can access the data, while AWS CloudTrail with data events logs every S3 object-level operation (e.g., GetObject, PutObject) for auditing. This combination directly meets the requirements of authorized access and logging for PII data.

Exam trap

The trap here is that candidates often assume CloudTrail automatically logs all S3 operations, but it only logs management events by default; data events must be explicitly enabled, and many overlook this distinction when designing for auditing.

How to eliminate wrong answers

Option B is wrong because Amazon S3 access points and VPC endpoints control network-level access and simplify bucket management, but they do not provide the required logging of data access for auditing. Option C is wrong because Amazon Macie discovers and classifies PII, and S3 Object Lock prevents deletion or overwriting, but neither service enforces access control or logs data access events. Option D is wrong because AWS KMS encrypts data at rest, which protects confidentiality but does not control who can access the data, and while AWS CloudTrail logs API calls, it does not log data events by default; without enabling data event logging, object-level access (e.g., reading a file) is not recorded.

799
MCQhard

A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains a column with inconsistent date formats (for example, '2023-01-15', '01/15/2023', and '15-Jan-2023'). The engineer needs to standardize all values to ISO 8601 format and then write the cleaned data back to S3. Which approach should the engineer use?

A.Use the 'Standardize date' transform in DataBrew, select the column, and choose the target format as ISO 8601.
B.Use the 'Format' transform in DataBrew to change the column data type to date.
C.Use the 'Replace value' transform to manually replace each non-ISO date string with its ISO equivalent.
D.Write a custom Python function in a DataBrew recipe step using the 'Custom transform' option.
AnswerA

AWS Glue DataBrew provides a built-in 'Standardize date' transform that can parse multiple date formats and convert them to a consistent target format such as ISO 8601. The engineer selects the column, configures the transform, and DataBrew applies it across the dataset. The cleaned data can then be written back to S3 as part of a DataBrew recipe or job.

Why this answer

AWS Glue DataBrew includes a 'Standardize date' transform that parses multiple date formats and converts them to a specified target format such as ISO 8601. This transform is designed for exactly this scenario, where a column contains mixed date representations. Using it avoids custom code and ensures consistent, repeatable cleaning across the dataset.

Exam trap

The trap here is assuming that a custom transform is needed for mixed date formats, when DataBrew already provides a built-in date standardization transform.

800
MCQhard

A data engineer is troubleshooting an access denied error when an AWS Lambda function tries to decrypt an object encrypted with the KMS key 'abc123'. The Lambda function's execution role has the above policy attached. What is the likely cause of the error?

A.The Deny statement blocks all decrypt requests
B.The Lambda function does not have permission to call kms:GenerateDataKey
C.The KMS key policy does not grant the Lambda role decrypt permission
D.The IAM policy does not include kms:Decrypt permission
AnswerC

The execution role's identity policy only permits the Lambda to request decryption; the KMS key policy must also allow that principal. If the key policy omits the Lambda role, KMS denies the Decrypt call, producing the access denied error despite the attached policy.

Why this answer

The error occurs because KMS requires both the IAM policy and the key policy to grant the necessary permissions. While the IAM policy attached to the Lambda execution role includes kms:Decrypt, the KMS key policy for 'abc123' does not explicitly grant the Lambda role permission to call kms:Decrypt. Since KMS key policies act as a resource-based policy, they must allow the principal (the Lambda role) to perform the action; otherwise, the request is denied even if the IAM policy allows it.

Exam trap

The trap here is that candidates assume IAM permissions alone are sufficient for KMS operations, overlooking that KMS key policies must explicitly grant access to the IAM role, which is a common source of access denied errors in cross-account or cross-service scenarios.

How to eliminate wrong answers

Option A is wrong because the Deny statement in the policy only blocks decrypt requests that do not include the encryption context 'Project=Alpha', not all decrypt requests; the error is likely due to missing key policy permissions, not a blanket Deny. Option B is wrong because the error is about decrypting an object, not generating a data key; kms:GenerateDataKey is used for encryption operations, not decryption, and the Lambda function is trying to decrypt, not encrypt. Option D is wrong because the IAM policy shown in the question includes kms:Decrypt permission (the policy lists kms:Decrypt as an allowed action), so the issue is not a missing IAM permission but rather the KMS key policy not granting the Lambda role decrypt permission.

801
MCQeasy

A company uses Amazon DynamoDB as the primary data store for a gaming application. The application stores user profiles and game state. During peak hours, the application experiences throttling on writes to the UserProfiles table. The table's read capacity is underutilized. Which solution should resolve the write throttling?

A.Increase the provisioned write capacity units for the table.
B.Enable DynamoDB Accelerator (DAX) on the table.
C.Add a global secondary index (GSI) to the table.
D.Configure auto scaling for read capacity units.
AnswerA

Increasing provisioned write capacity units directly raises the table's write throughput ceiling, eliminating the throttling caused by write requests exceeding the current WCU allocation during peak hours. Since read capacity is underutilised, only write capacity needs adjustment, making this the precise fix for the stated write bottleneck.

Why this answer

Write throttling occurs when the number of write requests exceeds the provisioned write capacity units (WCUs) for the DynamoDB table. Since the read capacity is underutilized, the correct solution is to increase the provisioned WCUs to accommodate the peak write traffic. This directly addresses the capacity deficit without affecting read operations.

Exam trap

The trap here is that candidates may confuse read and write capacity solutions, such as selecting DAX (which only helps reads) or auto scaling for reads, when the issue is specifically write throttling.

How to eliminate wrong answers

Option B is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that improves read performance, not write throughput; it does not increase write capacity or reduce write throttling. Option C is wrong because adding a global secondary index (GSI) consumes additional write capacity from the base table and can actually increase write throttling, not resolve it. Option D is wrong because auto scaling for read capacity units does not affect write throttling; write throttling requires adjusting write capacity, not read capacity.

802
MCQmedium

A data engineer is designing a data lake on Amazon S3. The data consists of sensitive personally identifiable information (PII) that must be encrypted at rest. The company requires that encryption keys be rotated every 90 days and that access to the keys be logged. Which encryption solution meets these requirements?

A.Use server-side encryption with customer-provided keys (SSE-C).
B.Use client-side encryption with a master key stored in AWS Secrets Manager.
C.Use server-side encryption with S3 managed keys (SSE-S3).
D.Use server-side encryption with AWS KMS (SSE-KMS) and enable automatic key rotation.
AnswerD

SSE-KMS provides customer-managed keys with rotation and logging via CloudTrail.

Why this answer

SSE-KMS with automatic key rotation (Option D) meets the requirements because it encrypts data at rest in S3, allows key rotation every 90 days via AWS KMS automatic rotation, and logs all key usage in AWS CloudTrail for auditing. This provides the necessary encryption, rotation, and access logging without managing keys externally.

Exam trap

The trap here is that candidates often confuse SSE-S3's automatic annual rotation with the required 90-day rotation, or assume SSE-C or client-side encryption can meet logging and rotation requirements without realizing they lack native AWS rotation and auditing capabilities.

How to eliminate wrong answers

Option A is wrong because SSE-C requires the customer to provide and manage their own encryption keys, and AWS does not support automatic key rotation or logging of key access for customer-provided keys. Option B is wrong because client-side encryption encrypts data before it reaches S3, but storing the master key in AWS Secrets Manager does not provide automatic key rotation every 90 days (Secrets Manager rotation is configurable but not native to KMS key rotation) and does not log key usage in CloudTrail as KMS does. Option C is wrong because SSE-S3 uses S3-managed keys that are rotated annually (not every 90 days) and do not provide granular access logging for key usage.

803
MCQhard

A financial services company uses a multi-account AWS Organization with hundreds of accounts. The data engineering team needs to enable cross-account access to an encrypted S3 bucket in the data lake account (account ID 111111111111) for a Glue ETL job running in the analytics account (account ID 222222222222). The S3 bucket uses AWS KMS customer managed key (CMK) for server-side encryption (SSE-KMS). The Glue job fails with an AccessDenied error when trying to read data from the bucket. The IAM roles in both accounts have the necessary S3 permissions and the bucket policy allows access from the analytics account. What is the most likely cause of the failure?

A.The KMS key policy does not grant the analytics account's IAM role permission to use the key for decryption.
B.The S3 bucket is in a different region than the Glue job.
C.The Glue job does not have an IAM role assigned.
D.The S3 bucket policy does not allow the s3:GetObject action for the analytics account's IAM role.
AnswerA

With SSE-KMS, the caller needs kms:Decrypt on the CMK in addition to S3 permissions. The bucket policy grants S3 access, but the key policy must separately authorise the analytics account's role, so its absence causes the AccessDenied failure.

Why this answer

When an S3 bucket is encrypted with SSE-KMS using a customer managed key (CMK), cross-account access requires that the KMS key policy explicitly grants the external account's IAM role permission to use the key for decryption. Even if the S3 bucket policy allows access, KMS enforces its own key policy, and without a grant for kms:Decrypt, the Glue job in the analytics account cannot decrypt the objects. This is the most likely cause because the scenario states that S3 permissions and bucket policy are already correctly configured.

Exam trap

DEA-C01 often tests the misconception that S3 bucket policies alone are sufficient for cross-account access to encrypted data, ignoring the separate KMS key policy requirement for decryption.

How to eliminate wrong answers

Option B is wrong because a region mismatch would typically produce a different error (e.g., bucket not found or endpoint mismatch) and the scenario does not indicate any regional configuration issue. Option C is wrong because the Glue job must have an IAM role to run at all; if no role were assigned, the job would fail immediately upon creation or with a different error, not an AccessDenied on S3 read. Option D is wrong because the scenario explicitly states that the bucket policy allows access from the analytics account, so the S3 bucket policy is not the missing piece.

804
MCQhard

A data engineer is troubleshooting a Lambda function that reads from a Kinesis Data Stream, processes records, and writes to a Kinesis Data Firehose delivery stream. The Firehose delivery stream is configured to deliver data to an S3 bucket. The Lambda function is failing with an access denied error. The IAM policy attached to the Lambda execution role is shown in the exhibit. Which permission is missing?

A.firehose:PutRecord on the Firehose delivery stream
B.firehose:DescribeDeliveryStream on the Firehose delivery stream
C.logs:CreateLogGroup and logs:CreateLogStream on the CloudWatch log group
D.s3:PutObjectAcl on the S3 bucket
AnswerA

Firehose ingestion uses the PutRecord (and PutRecordBatch) API, distinct from kinesis:PutRecord used for Kinesis Data Streams. Without firehose:PutRecord on the delivery stream, the Lambda execution role's write attempt is denied, producing the access denied error.

Why this answer

The Lambda function is failing due to a missing permission to write to the Kinesis Data Firehose delivery stream. The required permission is firehose:PutRecord (or firehose:PutRecordBatch). The IAM policy likely lacks this permission, causing an access denied error when the Lambda attempts to write records.

Option A correctly identifies this missing permission.

Exam trap

Candidates often confuse permissions for Kinesis Data Streams and Kinesis Data Firehose. Writing to a Firehose delivery stream requires firehose:PutRecord, not kinesis:PutRecord, which is for Kinesis Data Streams.

How to eliminate wrong answers

Option A is wrong because the Lambda function reads from the Kinesis Data Stream (requiring `kinesis:GetRecords`, `kinesis:DescribeStream`, etc.), not writes to it, so `kinesis:PutRecord` is irrelevant. Option C is wrong because CloudWatch Logs permissions (`logs:CreateLogGroup`, `logs:CreateLogStream`) are needed for logging but would cause a different error (e.g., 'Unable to create log stream'), not an access denied on Firehose. Option D is wrong because `s3:PutObjectAcl` is not required for Firehose to deliver to S3; Firehose uses `s3:PutObject` with bucket owner full control by default, and ACLs are not involved in this scenario.

805
MCQmedium

A data engineer is building a data lake on Amazon S3 and must choose the optimal file format for a dataset that is queried by Amazon Athena. The queries typically select a few columns from wide tables containing hundreds of columns, and the data volume is in terabytes. The engineer wants to minimize query scan costs and improve performance. Which file format should the engineer use?

A.Apache Avro
B.CSV
C.Apache Parquet
D.JSON
AnswerC

Parquet is a columnar format that stores data by column, enabling Athena to read only the columns referenced in a query. This drastically reduces the amount of data scanned, lowering costs and improving performance for wide tables where only a few columns are accessed. It also supports efficient compression and encoding, further reducing storage and scan volume, making it ideal for this analytical workload.

Why this answer

Apache Parquet is a columnar storage format that allows Amazon Athena to read only the columns needed for a query, significantly reducing the amount of data scanned. For wide tables where queries access a small subset of columns, this column pruning leads to lower costs and faster performance. Parquet also offers efficient compression, further reducing storage and scan volume, making it the best choice for this scenario.

Exam trap

The trap here is assuming that any columnar format is equally optimal, but Parquet specifically excels for Athena due to its widespread support and predicate pushdown capabilities.

806
MCQeasy

A company wants to ingest streaming data from thousands of IoT devices into AWS for real-time processing. Each device sends JSON payloads of about 2 KB at a rate of 1 message per second. The data must be processed with a durable, ordered stream per device. Which service should the company use as the ingestion layer?

A.Amazon Simple Queue Service (Amazon SQS) with a FIFO queue.
B.Amazon Kinesis Data Streams.
C.Amazon Simple Notification Service (Amazon SNS) with a Lambda subscriber.
D.Amazon Kinesis Data Firehose with Direct Put.
AnswerB

Amazon Kinesis Data Streams partitions records by partition key, giving each device a durable, strictly ordered shard sequence, satisfying the per-device ordering constraint. It also ingests thousands of concurrent 2 KB/second producers without provisioning brokers, unlike Amazon MSK, which requires cluster management, or SQS, which offers no ordering guarantee.

Why this answer

Amazon Kinesis Data Streams is the correct choice because it provides durable, ordered stream processing per shard, which can be partitioned by device ID to maintain message order for each device. It supports real-time ingestion from thousands of IoT devices at 2 KB per message per second, with a retention period of up to 365 days and the ability to reprocess data via multiple consumers.

Exam trap

The trap here is that candidates often confuse Kinesis Data Streams with Kinesis Data Firehose, assuming Firehose can handle real-time ordered streams, but Firehose does not provide per-record ordering or real-time consumer access; it is a delivery stream, not a stream processing layer.

How to eliminate wrong answers

Option A is wrong because Amazon SQS FIFO queues guarantee exactly-once processing and strict ordering within a message group, but they are designed for decoupling microservices, not for high-throughput streaming ingestion from thousands of devices; FIFO throughput is limited to 300 transactions per second (with batching) and does not support multiple consumers reading the same stream in real-time. Option C is wrong because Amazon SNS is a pub/sub messaging service that does not provide ordered delivery or durable stream storage; it pushes messages to subscribers like Lambda, but ordering is not guaranteed and messages are not persisted for replay. Option D is wrong because Amazon Kinesis Data Firehose is a near-real-time delivery service that buffers data and writes it to destinations like S3 or Redshift, but it does not support ordered stream processing per device and cannot be consumed by multiple real-time applications directly; it is designed for batch loading, not for real-time ordered stream consumption.

807
MCQhard

Refer to the exhibit. A data engineer is configuring an AWS Lambda function to process records from a Kinesis stream. The function is set up with an event source mapping, but no records are being processed. The Lambda function's IAM role has the policy shown. What is the most likely reason for the issue?

A.The policy does not grant permission to describe the Kinesis stream.
B.The IAM policy does not include all the necessary Kinesis actions for the event source mapping to work.
C.The policy includes too many actions, which causes a conflict.
D.The resource ARN for the Lambda function in the policy is incorrect.
AnswerB

The event source mapping requires kinesis:GetRecords, GetShardIterator, DescribeStream, ListStreams and ListShards. If the role's policy omits any of these, Lambda cannot poll the stream, so no records are read or invoked despite the mapping existing.

Why this answer

The most likely reason is that the IAM policy does not include all necessary Kinesis actions for the event source mapping to work. Lambda's event source mapping requires permissions to perform actions like `kinesis:GetRecords`, `kinesis:GetShardIterator`, `kinesis:DescribeStream`, and `kinesis:ListStreams`. If any of these are missing, the mapping cannot poll the stream, resulting in no records processed.

Exam trap

DEA-C01 often tests the specific IAM permissions required for Lambda event source mappings; candidates may assume that only `GetRecords` is needed, but multiple actions are required.

How to eliminate wrong answers

Option A is wrong because while `kinesis:DescribeStream` is necessary, it is not the only missing action; the policy likely lacks multiple actions. Option C is wrong because including too many actions does not cause conflicts; IAM policies are additive. Option D is wrong because the resource ARN for the Lambda function is not relevant in a policy attached to the Lambda execution role; the policy should reference the Kinesis stream ARN, not the Lambda function ARN.

808
MCQhard

A CloudFormation template includes this IAM policy for a cross-account S3 upload use case. What is the purpose of the condition?

A.To enforce server-side encryption with KMS.
B.To limit the size of objects that can be uploaded.
C.To restrict uploads to only a specific AWS account.
D.To ensure that uploaded objects grant full control to the bucket owner.
AnswerD

The condition checks the canned ACL on uploaded objects, ensuring the bucket owner receives full control. This satisfies the cross-account upload scenario, where objects written by another account would otherwise remain owned solely by the uploading account.

Why this answer

The condition in the IAM policy uses the `s3:x-amz-acl` key with a value of `bucket-owner-full-control`. This ensures that any object uploaded to the S3 bucket explicitly grants the bucket owner full control over the object, overriding the default behavior where the uploading account retains ownership. This is critical in cross-account uploads to prevent the uploading account from retaining exclusive access to the objects.

Exam trap

The trap here is that candidates confuse the `s3:x-amz-acl` condition key with account-level restrictions or encryption settings, when in fact it specifically controls the Access Control List (ACL) applied to the uploaded object.

How to eliminate wrong answers

Option A is wrong because server-side encryption with KMS is enforced using the `s3:x-amz-server-side-encryption-aws:kms` condition key, not the `s3:x-amz-acl` key. Option B is wrong because object size limits are enforced using the `s3:content-length-range` condition key, not the ACL-related condition shown. Option C is wrong because restricting uploads to a specific AWS account is done using the `aws:SourceAccount` or `aws:SourceArn` condition keys, not the `s3:x-amz-acl` key which controls object ACL permissions.

809
MCQmedium

A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data is then processed by an AWS Lambda function that transforms the records and writes them to an Amazon S3 bucket. Recently, the Lambda function has been timing out and the S3 bucket is not receiving all expected data. The Kinesis stream is not throttling and has sufficient shards. Which step should the company take to resolve this issue?

A.Increase the Lambda function's reserved concurrency.
B.Increase the Lambda function's timeout and memory allocation.
C.Increase the number of shards in the Kinesis stream.
D.Enable enhanced fan-out on the Kinesis stream to reduce latency.
AnswerB

Lambda timeouts arise when per-invocation processing exceeds the configured limit, and memory allocation proportionally increases CPU, accelerating transformation. Raising both values lets each batch complete before the timeout, so all Kinesis records reach S3 without the stream itself being throttled.

Why this answer

The Lambda function is timing out, indicating that it cannot process the records within the allotted time. Increasing the Lambda function's timeout and memory allocation (Option B) provides more execution time and CPU resources, which can help process larger or more frequent records before timing out. Option A is incorrect because reserved concurrency limits the maximum number of concurrent instances, not the execution duration of a single instance.

Option C is incorrect because increasing shards is unnecessary; the stream is not throttling, and more shards do not resolve Lambda timeouts. Option D is incorrect because enhanced fan-out reduces latency for multiple consumers but does not prevent a single Lambda function from timing out due to processing time.

810
Multi-Selecthard

A data engineer is troubleshooting a failed AWS Glue ETL job that reads from an S3 bucket. The job logs show the following error: 'java.lang.RuntimeException: java.lang.ClassNotFoundException: Class org.apache.hadoop.fs.s3a.S3AFileSystem not found'. Which TWO actions will resolve this issue?

Select 2 answers
A.Enable VPC S3 endpoint for the Glue job.
B.Include the hadoop-aws jar as an extra jar in the Glue job configuration.
C.Update the IAM role to allow access to S3.
D.Use a Glue version that includes the S3A filesystem library (e.g., Glue 3.0 or later).
E.Change the S3 access mode from S3A to EMRFS.
AnswersB, D

The ClassNotFoundException means the S3A filesystem implementation is absent from the job's classpath. Adding the hadoop-aws jar as an extra jar supplies the missing org.apache.hadoop.fs.s3a.S3AFileSystem class, letting the Glue job resolve and read from S3.

Why this answer

The error 'ClassNotFoundException: org.apache.hadoop.fs.s3a.S3AFileSystem' means the S3A filesystem implementation class is not on the Glue job's classpath, so the fix must supply that library. Option B is correct because adding the hadoop-aws jar (which contains org.apache.hadoop.fs.s3a.S3AFileSystem) as an extra jar via the Glue job's --extra-jars configuration puts the missing S3A class on the classpath. Option D is correct because newer AWS Glue versions (Glue 3.0 and later, built on Spark 3.x/Hadoop 3.x) bundle the S3A filesystem library, so running the job on such a version provides the class natively.

Option A is wrong because a VPC S3 endpoint only fixes network routing to S3, not a missing Java class. Option C is wrong because IAM permissions govern authorization, and a ClassNotFoundException is a classpath problem, not an access-denied error. Option E is wrong because EMRFS is an EMR-specific filesystem and is not a valid S3 access mode to switch to within AWS Glue.

Exam trap

DEA-C01 often tests the confusion between classpath errors (missing JAR) and permission errors (IAM), leading candidates to choose IAM or VPC options when the error is clearly a ClassNotFoundException.

811
MCQeasy

Refer to the exhibit. An IAM policy includes the above statement to allow decryption of a KMS key under specific conditions. What does this policy allow?

A.Decrypt any data encrypted with any KMS key
B.Decrypt data that was encrypted with the encryption context {"aws:pi":"db-123"}
C.Encrypt data with the KMS key using the specified encryption context
D.Decrypt data encrypted with the KMS key without any encryption context
AnswerB

The policy's condition matches the encryption context key-value pair aws:pi equal to db-123, so kms:Decrypt succeeds only for ciphertext whose encryption context includes that exact pair, binding decryption to data encrypted under the same context.

Why this answer

Option B is correct because the policy allows the kms:Decrypt action only when the encryption context matches the specified key-value pair. The condition likely uses the StringEquals condition operator on the kms:EncryptionContext:aws:pi key with value db-123. This means decryption is permitted only for ciphertext that was encrypted with that exact encryption context.

Exam trap

DEA-C01 often tests the misunderstanding of encryption context conditions in IAM policies. Candidates may think the policy allows decryption of any data with the key, but the condition restricts it to a specific context.

How to eliminate wrong answers

Option A is wrong because the policy is scoped to a specific KMS key and condition, not any key. Option C is wrong because the action is kms:Decrypt, not kms:Encrypt. Option D is wrong because the condition requires the encryption context to be present and match; decryption without any encryption context would fail the condition.

812
MCQhard

A CloudFormation template defines an AWS Glue job. The job fails during execution with the error 'Unable to locate script: s3://scripts-bucket/etl-script.py'. The S3 bucket 'scripts-bucket' exists and the script file is present. What is the most likely cause?

A.The script location path is incorrect; it should include the bucket's region.
B.The IAM role for the Glue job does not have s3:GetObject permission on the scripts bucket.
C.The Glue job requires Python version 2, but the script uses Python 3 syntax.
D.The S3 bucket is in a different AWS region than the Glue job.
AnswerB

Glue retrieves the script from S3 using the job's IAM role, so a missing s3:GetObject permission on scripts-bucket produces exactly this 'Unable to locate script' error even though the object exists. The bucket policy and object presence are irrelevant if the role is denied.

Why this answer

The Glue job fails to locate the script because the IAM role assigned to the job lacks the s3:GetObject permission on the scripts-bucket. Even though the bucket and object exist, AWS Glue requires the execution role to have explicit read access to the S3 object to download and execute the script. Without this permission, the job cannot retrieve the file, resulting in the 'Unable to locate script' error.

Exam trap

The trap here is that candidates assume the error is about the script path or region mismatch, but the real cause is almost always missing IAM permissions for the Glue execution role to read the script from S3.

How to eliminate wrong answers

Option A is wrong because S3 object paths do not include the bucket's region; the path format is s3://bucket-name/key, and region is irrelevant to the path. Option C is wrong because Python version compatibility would cause a syntax error during execution, not a 'Unable to locate script' error, which is a file access issue. Option D is wrong because S3 buckets and Glue jobs can operate across regions as long as the Glue job's IAM role has appropriate cross-region permissions; the error message specifically indicates a missing object, not a region mismatch.

813
MCQmedium

A data engineer runs the above command and gets the output. What does the 'MFADelete' setting imply?

A.Any modification to an object requires MFA.
B.To permanently delete a version of an object, the user must provide MFA.
C.MFA is required for all read operations as well.
D.All operations on the bucket require MFA authentication.
AnswerB

With MFADelete enabled on a versioned bucket, deleting a specific object version requires an MFA token in the request; without it, the delete is denied. This protects against accidental or malicious permanent version removal, satisfying the stem's requirement to understand what the setting enforces.

Why this answer

The 'MFADelete' setting on an S3 bucket versioning configuration requires multi-factor authentication to permanently delete an object version. This means that when a user issues a DELETE request with a version ID (a permanent delete), they must include a valid MFA token in the request headers. It does not apply to creating new versions, reading objects, or other operations.

Exam trap

The DEA-C01 exam often tests the distinction between 'MFA Delete' (which only applies to permanent version deletion) and general MFA enforcement on all bucket operations, leading candidates to overgeneralize the scope of the setting.

How to eliminate wrong answers

Option A is wrong because 'MFADelete' does not require MFA for any modification (e.g., PUT, POST, or COPY operations); it only applies to permanent deletes of specific versions. Option C is wrong because read operations (GET, HEAD) are never subject to MFA requirements under this setting. Option D is wrong because 'MFADelete' is a versioning-specific sub-setting and does not enforce MFA on all bucket operations, only on permanent version deletion.

814
MCQeasy

A data engineer needs to ensure that all data in an S3 bucket is encrypted at rest. The bucket currently contains unencrypted objects from past uploads. Which action will encrypt these existing objects without re-uploading them?

A.Attach a bucket policy requiring SSE-S3
B.Enable default encryption on the bucket
C.Use the S3 console to select all objects and apply encryption
D.Use S3 Batch Operations with an encryption job
AnswerD

S3 Batch Operations applies a batch job that re-encrypts existing objects in place using SSE-KMS or SSE-S3, satisfying the stem's constraint of encrypting past uploads without re-uploading. It reads each object and writes it back encrypted, preserving keys and metadata.

Why this answer

S3 Batch Operations can run an encryption job across existing objects, applying SSE-S3, SSE-KMS, or SSE-C encryption without re-uploading the data. This is the only option that retroactively encrypts already-stored objects in place. Default encryption and bucket policies only affect new uploads, not existing objects.

Exam trap

The trap is assuming that enabling default encryption or a bucket policy retroactively encrypts existing objects — the exam tests whether you know that only a copy/rewrite operation (via Batch Operations or CLI copy) changes encryption on already-stored objects.

How to eliminate wrong answers

Option A is wrong because a bucket policy requiring SSE-S3 only rejects unencrypted PUT requests; it does not encrypt objects already in the bucket. Option B is wrong because enabling default encryption applies only to new objects uploaded after the setting is enabled, leaving existing objects unencrypted. Option C is wrong because the S3 console does not provide a bulk 'apply encryption' action on existing objects; you would have to copy each object, which is effectively re-uploading.

815
MCQhard

A company has an AWS Glue ETL job that reads data from an S3 bucket, transforms it, and writes to another S3 bucket. The security team requires that data in transit between the Glue job and S3 be encrypted using TLS. The Glue job runs in a VPC with a VPC endpoint for S3. Which configuration ensures TLS encryption for all data transfer?

A.Use an S3 Gateway Endpoint and ensure the Glue job uses HTTP instead of HTTPS.
B.Use an S3 Interface Endpoint and disable TLS.
C.Use an S3 Gateway Endpoint and ensure the Glue job uses HTTPS.
D.Enable SSE-KMS encryption on both source and destination S3 buckets.
AnswerC

An S3 Gateway Endpoint keeps traffic to S3 on the AWS private network, so TLS is not applied automatically. Forcing the Glue connection to use HTTPS (s3a:// with SSL enabled) encrypts data in transit, satisfying the security team's TLS requirement.

Why this answer

An S3 Gateway Endpoint keeps traffic between the VPC and S3 on the AWS private network, and when the Glue job is configured to use HTTPS (the default for the AWS SDK and Glue S3 connections), the data in transit is encrypted with TLS. The gateway endpoint itself does not encrypt traffic, but it does not prevent TLS; the key is that the client must use the HTTPS scheme. Combining the gateway endpoint with HTTPS satisfies the requirement for TLS encryption in transit.

Exam trap

The trap is confusing encryption at rest (SSE-KMS) with encryption in transit (TLS); candidates see 'encryption' in the option and pick it without checking whether it applies to data in motion.

How to eliminate wrong answers

Option A is wrong because using HTTP explicitly disables TLS, which directly violates the requirement for encryption in transit. Option B is wrong because disabling TLS on an interface endpoint removes encryption, and interface endpoints are for services like KMS or Secrets Manager, not the standard S3 data path. Option D is wrong because SSE-KMS is encryption at rest, not in transit; it protects objects on disk but does nothing for the TLS requirement between Glue and S3.

816
MCQmedium

A company uses AWS Data Pipeline to copy data from DynamoDB to S3 daily. Recently, the pipeline started failing with 'ThrottlingException' errors. The DynamoDB table has on-demand capacity. Which action should be taken to resolve the issue?

A.Increase the write capacity units of the DynamoDB table.
B.Replace Data Pipeline with AWS Glue using a DynamoDB connector.
C.Configure the pipeline to use a retry strategy with exponential backoff.
D.Disable the pipeline's retry logic and increase the timeout.
AnswerC

On-demand DynamoDB tables still throttle when request rates spike above what the table can absorb, and Data Pipeline's default retries are too aggressive. Adding a retry strategy with exponential backoff spaces out retries, letting the table recover and the copy complete without manual intervention.

Why this answer

ThrottlingException errors in AWS Data Pipeline when reading from DynamoDB indicate that the pipeline's read requests are exceeding the table's available throughput. Since the table uses on-demand capacity, which can handle spikes but has a per-second throughput limit, implementing exponential backoff in the pipeline's retry strategy allows it to reduce request rate upon throttling, aligning with AWS SDK best practices for handling DynamoDB throttling.

Exam trap

The trap here is that candidates assume on-demand capacity eliminates all throttling, but it only handles traffic spikes within a per-second limit, so throttling can still occur with sustained high read rates, and the correct fix is to implement exponential backoff in the pipeline's retry strategy rather than modifying capacity or switching tools.

How to eliminate wrong answers

Option A is wrong because DynamoDB on-demand capacity does not use provisioned write capacity units; increasing write capacity units is irrelevant and would require switching to provisioned mode, which is unnecessary. Option B is wrong because replacing Data Pipeline with AWS Glue using a DynamoDB connector does not inherently resolve throttling; Glue also uses the same DynamoDB read APIs and would face the same throttling issue without proper retry handling. Option D is wrong because disabling retry logic and increasing the timeout would cause the pipeline to fail permanently on the first throttling error, as it would not retry the request, and a longer timeout does not prevent throttling.

817
MCQhard

A company runs an Amazon RDS for PostgreSQL instance that stores financial data. The company requires point-in-time recovery (PITR) with a retention period of 35 days. Additionally, the company needs to create a new database from a specific snapshot every night for testing. Which combination of actions should the data engineer take to meet these requirements?

A.Enable automated backups with a 35-day retention period and create a manual snapshot each night for testing.
B.Create a read replica and promote it to a new instance for testing each night.
C.Enable Multi-AZ and use the standby instance for testing.
D.Disable automated backups to reduce storage costs and take manual snapshots with 35-day retention.
AnswerA

Automated backups provide continuous PITR up to 35 days, meeting the retention requirement. Manual snapshots are independent of the backup retention window and can be created nightly, giving a stable source for the test database without affecting PITR.

Why this answer

Automated backups in Amazon RDS for PostgreSQL support a maximum retention period of 35 days, which satisfies the PITR requirement. Additionally, creating a manual snapshot each night provides a stable, independent copy for testing without interfering with the automated backup schedule or the source database's performance.

Exam trap

The trap here is that candidates often confuse the purpose of Multi-AZ standby instances (which are not directly usable for testing) or assume that manual snapshots alone can provide PITR, but automated backups are strictly required for point-in-time recovery in RDS.

How to eliminate wrong answers

Option B is wrong because a read replica is designed for read scaling and high availability, not for creating a nightly test database; promoting a read replica each night would disrupt replication and require re-creating the replica, which is inefficient and does not meet the PITR retention requirement. Option C is wrong because Multi-AZ provides high availability and automatic failover, but the standby instance is not directly accessible for testing; it cannot be used to create a new database without promoting it, which would break the Multi-AZ configuration. Option D is wrong because disabling automated backups eliminates the ability to perform point-in-time recovery (PITR), and manual snapshots alone do not support PITR; automated backups are required for transaction log retention and restore to any point within the retention window.

818
MCQeasy

A data engineer needs to store log files from multiple applications in a centralized location. The logs are generated in JSON format and each log entry is about 1 KB. The engineer needs to query the logs occasionally using SQL-like queries. Which AWS service is most appropriate?

A.Amazon DynamoDB
B.Amazon Redshift
C.Amazon Athena with data stored in S3
D.Amazon RDS for MySQL
AnswerC

Athena queries JSON in S3 directly using standard SQL, charging only per query scanned, which suits occasional log analysis. S3 provides the centralised, durable store for the 1 KB JSON entries, so no database loading or cluster is required.

Why this answer

Amazon Athena is the most appropriate service because it allows you to query log files stored in S3 directly using standard SQL, without needing to load or transform the data. Since the logs are in JSON format and each entry is about 1 KB, Athena's schema-on-read approach works perfectly for occasional SQL-like queries, and you only pay for the data scanned per query, making it cost-effective for infrequent access.

Exam trap

The trap here is that candidates often choose Amazon Redshift or RDS because they think 'SQL-like queries' require a traditional database, overlooking Athena's ability to query data directly in S3 without loading it, which is a key serverless pattern for log analytics.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for low-latency, high-throughput access patterns, not for ad-hoc SQL-like queries on large volumes of log data, and it would require schema design and provisioning. Option B is wrong because Amazon Redshift is a petabyte-scale data warehouse designed for complex analytical queries on structured, transformed data, which is overkill and costly for occasional log queries on JSON files stored in S3. Option D is wrong because Amazon RDS for MySQL is a relational database that requires schema definition, data loading, and ongoing management, making it unsuitable for storing raw JSON log files directly without ETL, and it lacks the serverless, pay-per-query model for infrequent access.

819
MCQeasy

A company wants to centrally manage access to multiple AWS accounts for its data engineers. The company already uses AWS Organizations. Which AWS service should be used to define fine-grained permissions across accounts?

A.AWS IAM
B.AWS IAM Identity Center (AWS Single Sign-On)
C.AWS Resource Access Manager (AWS RAM)
D.AWS Key Management Service (AWS KMS)
AnswerB

IAM Identity Center centralises workforce access across all accounts in AWS Organizations, assigning permission sets to users and groups from one place. It provides the cross-account fine-grained permissions the company needs without per-account IAM users.

Why this answer

AWS IAM Identity Center (formerly AWS Single Sign-On) is designed to centrally manage access to multiple AWS accounts within AWS Organizations. It allows you to define fine-grained permissions using permission sets and assign them to users or groups across accounts.

Exam trap

The trap is confusing AWS RAM (resource sharing) with identity management; candidates may also think IAM alone can centrally manage multi-account access, but IAM is per-account and lacks the centralized SSO and permission set features of IAM Identity Center.

How to eliminate wrong answers

Option A is wrong because AWS IAM is account-specific; while it can be used with roles for cross-account access, it does not provide a central place to manage access across multiple accounts in an organization. Option C is wrong because AWS RAM shares resources (e.g., subnets, transit gateways) across accounts, not user permissions. Option D is wrong because AWS KMS manages encryption keys, not user access permissions.

820
MCQeasy

A company uses Amazon Redshift for data warehousing. The security team requires that all data in transit between the Redshift cluster and clients be encrypted. Which feature should be enabled?

A.Client-side VPN
B.SSL/TLS encryption
C.AWS KMS key
D.VPC peering
AnswerB

SSL/TLS encryption secures data in transit between clients and the Redshift cluster, directly satisfying the requirement that all traffic be encrypted. Redshift supports SSL connections, and enabling the `require_ssl` parameter forces every client connection to use TLS, preventing unencrypted access.

Why this answer

Amazon Redshift supports SSL/TLS encryption for client connections to ensure data in transit is encrypted. Option A (Client-side VPN) is not a Redshift feature for encrypting client connections. Option C (AWS KMS key) is used for encrypting data at rest, not in transit.

Option D (VPC peering) does not provide encryption of data in transit between the cluster and clients.

821
MCQeasy

A company stores sensitive data in Amazon S3 and uses AWS Lake Formation to manage fine-grained access control. A data engineer notices that users are able to access data in S3 directly via the AWS Management Console, bypassing Lake Formation permissions. What should the engineer do to enforce Lake Formation access controls for all access methods?

A.Add a bucket policy that denies all access except from Lake Formation.
B.Disable AWS CloudTrail logging for S3 access.
C.Register the S3 location in Lake Formation and disable IAM access control for the registered location.
D.Enable S3 Block Public Access on the bucket.
AnswerC

Registering the S3 location in Lake Formation makes it a governed table, and disabling IAM access control removes the S3 bucket-policy and IAM identity-based paths that let users read objects directly through the console. This satisfies the stem's requirement to enforce Lake Formation permissions across all access methods, not just integrated engines.

Why this answer

Registering the S3 location in Lake Formation and disabling IAM access control for that registered location enforces Lake Formation permissions for all access methods, including direct S3 access via the console. This ensures that Lake Formation acts as the single source of truth for fine-grained access control.

Exam trap

DEA-C01 often tests the misconception that S3 bucket policies alone can enforce Lake Formation permissions, when in fact Lake Formation must be configured to override IAM access for registered locations.

How to eliminate wrong answers

Option A is wrong because a bucket policy denying all access except from Lake Formation is not sufficient; IAM policies can still grant direct access unless Lake Formation is configured to override them. Option B is wrong because disabling CloudTrail logging reduces auditability and does not enforce access controls. Option D is wrong because S3 Block Public Access only prevents public access, not fine-grained access control bypass by authenticated users.

822
MCQeasy

A company is using Amazon EMR to process large datasets stored in Amazon S3. The data engineer wants to reduce the time it takes to read data from S3 by optimizing the data format. Which file format should the engineer recommend?

A.CSV
B.Parquet
C.ORC
D.JSON
AnswerB

Parquet is columnar and compressed, so EMR reads only the columns referenced by the query rather than scanning entire rows. This column pruning plus predicate pushdown and efficient encoding cuts the volume of data transferred from S3, directly reducing read time.

Why this answer

Parquet is the correct choice because it is a columnar storage format that significantly reduces the amount of data read from Amazon S3 during analytical queries. By storing data column-wise, Parquet enables predicate pushdown and compression, which minimizes I/O and speeds up data processing in Amazon EMR, especially for large datasets.

Exam trap

The trap here is that candidates often assume ORC is the default or preferred format for all big data engines. However, for Amazon EMR, Parquet is generally recommended because of its superior performance with Spark and its ability to handle complex nested data structures efficiently.

How to eliminate wrong answers

Option A is wrong because CSV is a row-oriented text format that requires full file scans and offers no compression or predicate pushdown, leading to slower reads. Option C is wrong because ORC is also a columnar format optimized for Hive workloads, but it is not natively as performant with Spark and EMR as Parquet, and the question asks for the best recommendation for EMR. Option D is wrong because JSON is a row-oriented, self-describing format that is verbose and lacks efficient compression or columnar access patterns, resulting in high I/O and slower processing.

823
MCQhard

A data engineer is designing a data ingestion pipeline for real-time financial transactions. The pipeline must ensure exactly-once processing semantics and must handle duplicate records that may occur due to retries. Which combination of AWS services can achieve exactly-once processing?

A.Amazon Kinesis Data Streams with Amazon Kinesis Data Analytics for Apache Flink
B.Amazon SQS with AWS Lambda
C.Amazon MSK with AWS Lambda
D.Amazon Kinesis Data Firehose with AWS Lambda
AnswerA

Flink supports exactly-once processing with KDS.

Why this answer

Amazon Kinesis Data Streams with Kinesis Data Analytics for Apache Flink enables exactly-once processing by leveraging Flink's built-in checkpointing and two-phase commit protocols. Kinesis Data Streams provides a durable, ordered record store, while Flink uses its internal state and idempotent sinks to deduplicate records that arise from retries, ensuring each record is processed exactly once.

Exam trap

The trap here is that candidates often assume SQS FIFO or Lambda's idempotency can achieve exactly-once, but they overlook that without a stream processing framework with checkpointing and transactional sinks, retries can still introduce duplicates in distributed systems.

How to eliminate wrong answers

Option B is wrong because Amazon SQS with AWS Lambda provides at-least-once delivery by default; SQS does not guarantee deduplication without a FIFO queue and idempotent Lambda logic, and even then, exactly-once is not natively supported across retries. Option C is wrong because Amazon MSK (Kafka) with AWS Lambda requires custom checkpointing and idempotency logic in the Lambda function; MSK does not natively provide exactly-once semantics without a stream processing framework like Flink or Kafka Streams. Option D is wrong because Amazon Kinesis Data Firehose with AWS Lambda delivers records at least once and cannot guarantee exactly-once processing; Firehose buffers and batches data but does not support transactional commits or deduplication across retries.

824
Multi-Selecthard

Which TWO are benefits of using Amazon S3 Object Lock? (Choose TWO.)

Select 2 answers
A.Helps meet regulatory requirements for write-once-read-many (WORM) storage.
B.Encrypts objects at rest using AWS KMS.
C.Prevents objects from being deleted or overwritten for a fixed time.
D.Automatically transitions objects to lower-cost storage classes.
E.Enables automatic versioning of objects.
AnswersA, C

Object Lock provides WORM semantics, meaning objects written once cannot be altered or removed during retention. This directly satisfies the stem's regulatory compliance benefit, since auditors accept S3 Object Lock as evidence of tamper-proof, immutable record retention.

Why this answer

Amazon S3 Object Lock helps meet regulatory requirements for write-once-read-many (WORM) storage by allowing you to set retention periods and legal holds on objects. This ensures that data cannot be deleted or overwritten for a specified duration, which is a common requirement for compliance frameworks such as SEC Rule 17a-4 or FINRA.

Exam trap

The trap here is that candidates confuse S3 Object Lock with S3 Versioning or S3 Lifecycle policies, mistakenly thinking Object Lock handles encryption or storage tier transitions, when in reality it is solely focused on preventing object deletion or overwrite for compliance-driven WORM scenarios.

825
MCQhard

A company uses AWS Lake Formation to manage data lakes on Amazon S3. The data engineer needs to grant a data analyst access to query specific columns in a table using Amazon Athena, but deny access to columns containing personally identifiable information (PII). Which Lake Formation feature should be used?

A.Row-level security filters.
B.Column-level permissions in Lake Formation.
C.Tag-based access control with Lake Formation tags.
D.Cell-level security with AWS Glue.
AnswerB

Column-level permissions in Lake Formation restrict access at the individual column granularity, so the analyst can query non-PII columns while PII columns remain denied. This satisfies the stem's requirement to grant selective column access through Athena without exposing sensitive data, since Lake Formation enforces these controls centrally across integrated analytics services.

Why this answer

Lake Formation column-level permissions allow a data engineer to grant SELECT on specific columns of a table while excluding others, which is exactly what is needed to hide PII columns from the analyst. When the analyst queries via Athena, Lake Formation enforces the column filter at query time, returning only the permitted columns. This is the native, fine-grained access control mechanism for column-level restrictions in Lake Formation.

Exam trap

DEA-C01 often tests the confusion between row-level filters (horizontal) and column-level permissions (vertical), and candidates may incorrectly choose tag-based access control when the question asks for the specific column-restriction feature.

How to eliminate wrong answers

Option A is wrong because row-level security filters restrict which rows are visible based on a filter expression, not which columns — they address horizontal, not vertical, data restriction. Option C is wrong because tag-based access control (LF-TBAC) uses tags to scale permission management across many resources, but the question asks for the specific feature that grants column-level access; tags are a management mechanism, not the column-permission primitive itself. Option D is wrong because 'cell-level security with AWS Glue' is not a real Lake Formation feature — Glue does not provide cell-level access control, and Lake Formation is the correct service for fine-grained column/row control.

Page 10

Page 11 of 18

Page 12