Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 676–750

1321 questions total · 18pages · All types, answers revealed

Page 9

Page 10 of 18

Page 11
676
MCQmedium

A company uses Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format. The delivery stream is configured with a buffer size of 5 MB and a buffer interval of 60 seconds. However, the data engineer notices that S3 objects are being created with sizes much smaller than 5 MB. What is a likely cause?

A.The data is being compressed before delivery, reducing object size.
B.The incoming data rate is too low, causing the buffer interval to trigger before reaching the buffer size.
C.The data transformation lambda is splitting records into smaller ones.
D.The S3 bucket is configured with a lifecycle policy that splits objects.
AnswerB

Firehose flushes a buffer when either the 5 MB size or the 60-second interval elapses, whichever occurs first. At a low incoming data rate, the interval triggers repeatedly before the buffer fills, producing S3 objects far smaller than 5 MB.

Why this answer

Kinesis Data Firehose delivers data to S3 when either the buffer size (5 MB) or the buffer interval (60 seconds) is reached, whichever occurs first. If the incoming data rate is low, the buffer interval will expire before accumulating 5 MB of data, resulting in smaller S3 objects.

Exam trap

The trap here is that candidates may assume the buffer size is a hard minimum that must be reached before delivery, but Firehose uses an 'or' condition between buffer size and buffer interval, so low data rate causes interval-based delivery of small objects.

How to eliminate wrong answers

Option A is wrong because compression reduces the size of data after buffering, but the buffer size limit is based on the uncompressed data; compression does not cause smaller objects to be created before the buffer interval triggers. Option C is wrong because a data transformation Lambda can modify records but does not inherently split records into smaller ones; it processes records as a batch and returns them, and any splitting would be a custom logic not default behavior. Option D is wrong because S3 lifecycle policies manage object transitions or deletions after objects are created; they do not split objects during delivery.

677
MCQmedium

Refer to the exhibit. An IAM policy for an AWS Lambda function. The Lambda function is triggered by an S3 event (object created) and needs to read from a Kinesis stream. However, the function fails with access denied when trying to read from Kinesis. What is the most likely cause?

A.The Lambda function is not in the same region as the Kinesis stream
B.The Lambda function does not have permission to list S3 buckets
C.The Kinesis stream is encrypted with a customer managed KMS key, and the Lambda function lacks kms:Decrypt permission
D.The S3 bucket policy denies access to the Lambda function
AnswerC

Reading from a Kinesis stream encrypted with a customer managed KMS key requires kms:Decrypt in the Lambda execution role. The S3 event trigger and stream read permissions may be correct, but the missing KMS grant causes the access denied error.

Why this answer

When a Kinesis stream is encrypted with a customer managed KMS key, the Lambda function must have the `kms:Decrypt` permission on that key to read data from the stream. Without this permission, the Lambda function will receive an access denied error even if it has the necessary Kinesis actions (e.g., `kinesis:GetRecords`) allowed in its IAM policy. The S3 event trigger only invokes the function; it does not grant Kinesis access.

Exam trap

The DEA-C01 exam often tests the interaction between Kinesis SSE-KMS and Lambda IAM permissions, trapping candidates who assume that Kinesis read permissions alone are sufficient without considering the KMS key policy.

How to eliminate wrong answers

Option A is wrong because Lambda functions can access Kinesis streams across regions as long as the IAM permissions and network connectivity (e.g., VPC endpoints) are correctly configured; region mismatch does not inherently cause access denied. Option B is wrong because the Lambda function is triggered by an S3 event and only needs permission to read from Kinesis; listing S3 buckets is irrelevant to the Kinesis read failure. Option D is wrong because the S3 bucket policy controls access to the S3 bucket itself, not to Kinesis; the error occurs when reading from Kinesis, not when the S3 event triggers the function.

678
MCQeasy

A data engineer needs to audit data access events in Amazon S3. Which AWS service should be used to record and monitor API calls for S3 buckets?

A.AWS CloudTrail
B.AWS Config
C.Amazon Macie
D.Amazon GuardDuty
AnswerA

AWS CloudTrail records S3 data events such as GetObject and PutObject, capturing the identity, source IP and timestamp of each API call. Enabling data-event logging on the buckets satisfies the audit requirement, which S3 server access logs alone cannot match for API-level detail.

Why this answer

AWS CloudTrail records API activity across AWS services, including S3 data-plane and management-plane events, capturing who made each request, from where, and when. For auditing data access to S3 buckets, CloudTrail (with S3 data events enabled) is the authoritative service. It integrates with CloudWatch Logs and S3 for long-term retention and analysis.

Exam trap

DEA-C01 often tests the confusion between CloudTrail (API audit logging) and AWS Config (configuration compliance tracking), since both provide visibility into account activity but serve fundamentally different purposes.

How to eliminate wrong answers

Option B is wrong because AWS Config evaluates resource configuration compliance and tracks configuration changes over time, but it does not record individual API calls or data-access events. Option C is wrong because Amazon Macie discovers and classifies sensitive data in S3 (e.g., PII) using machine learning; it is a data-security posture tool, not an API audit log. Option D is wrong because Amazon GuardDuty is a threat-detection service that analyzes CloudTrail, VPC Flow Logs, and DNS logs for malicious behavior — it consumes audit data rather than being the source of it.

679
Matchingmedium

Match each AWS storage class to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Frequent access, low latency

Auto-moves data between tiers

Archive retrieval in minutes to hours

Lowest cost, 12-hour retrieval

Infrequent access, single AZ

Why these pairings

Correct matches: S3 Standard for frequently accessed data, S3 Intelligent-Tiering for automatic cost optimization, and S3 Glacier for archival. Common confusions involve mixing up Standard-IA and Glacier Deep Archive definitions.

680
Multi-Selectmedium

A data engineer is designing a data lake on Amazon S3 that will be accessed by multiple AWS Glue ETL jobs. The engineer needs to ensure that the data is organized efficiently for querying and that sensitive columns are masked for certain users. Which TWO actions should the engineer take? (Choose TWO.)

Select 2 answers
A.Use AWS Lake Formation to define column-level permissions for sensitive data.
B.Configure AWS Glue Data Catalog to automatically mask sensitive columns in table definitions.
C.Organize data in S3 using a partition structure like 'year=YYYY/month=MM/day=DD/region=XX/'.
D.Use S3 object tags to label sensitive data and apply bucket policies to restrict access.
E.Implement S3 lifecycle policies to transition sensitive data to S3 Glacier after 30 days.
AnswersA, C

Lake Formation provides column-level security to mask sensitive columns.

Why this answer

AWS Lake Formation provides fine-grained access control at the column level, allowing you to mask or restrict sensitive columns (e.g., PII) for specific IAM roles or users without altering the underlying data in S3. This is achieved through Lake Formation’s column-level permissions and data filtering, which integrate directly with the AWS Glue Data Catalog and query engines like Athena and Redshift Spectrum.

Exam trap

The trap here is that candidates often confuse S3 object tags or bucket policies with fine-grained column-level access control, or assume the Glue Data Catalog can natively mask columns, when in fact only Lake Formation provides that capability.

681
MCQeasy

A data engineer is setting up an Amazon Kinesis Data Firehose delivery stream to load data into Amazon Redshift. The data is coming from an application that produces JSON records. The engineer needs to transform the data to match the Redshift table schema. Which approach is the MOST cost-effective and requires the least operational overhead?

A.Use AWS Glue as a transformation step between Firehose and Redshift, with a trigger on S3.
B.Use Kinesis Data Firehose with direct PUT to Redshift and rely on Redshift's COPY command to transform.
C.Configure a Lambda function in the Firehose delivery stream to transform records before delivery.
D.Use the Kinesis Client Library (KCL) to consume the stream, transform in an EC2 instance, and then load to Redshift.
AnswerC

Firehose invokes a Lambda function inline to transform each JSON record before loading, so no intermediate storage or separate processing cluster is needed. This satisfies the least-operational-overhead and cost-effectiveness constraints, since you pay only per invocation.

Why this answer

Kinesis Data Firehose natively supports invoking a Lambda function as a transformation step within the delivery stream. This allows the engineer to write a simple Lambda function that parses the incoming JSON records and transforms them to match the Redshift table schema, all without provisioning or managing any additional infrastructure. This approach is the most cost-effective (pay per invocation) and requires the least operational overhead since Firehose handles the orchestration, retries, and delivery to Redshift automatically.

Exam trap

The trap here is that candidates often overestimate the transformation capabilities of Redshift's COPY command, mistakenly believing it can perform complex record-level transformations, when in fact it only supports basic data mapping and format parsing, not arbitrary JSON restructuring.

How to eliminate wrong answers

Option A is wrong because inserting AWS Glue as an intermediate step between Firehose and Redshift introduces unnecessary complexity, cost (Glue jobs run on a per-DPU-hour basis), and latency, as Glue is designed for batch ETL, not real-time streaming transformations. Option B is wrong because Redshift's COPY command does not perform record-level transformations; it only maps source fields to target columns and can apply basic data format conversions (e.g., JSON parsing via 'jsonpaths'), but it cannot restructure or compute new fields from the JSON payload. Option D is wrong because using the Kinesis Client Library (KCL) on an EC2 instance requires manual provisioning, scaling, and management of the EC2 fleet, which incurs significant operational overhead and cost compared to the serverless Lambda integration within Firehose.

682
MCQmedium

A company is building a data lake on Amazon S3 and wants to ingest data from multiple AWS services (CloudTrail, VPC Flow Logs, and ALB logs). The data should be stored in a central S3 bucket with a common partitioning scheme. Which service can be used to collect and centralize this data with minimal configuration?

A.Use AWS Data Pipeline to copy logs from each source S3 bucket to the central bucket.
B.Use AWS Glue to crawl the logs from each source and write to a central S3 bucket.
C.Set up Amazon Kinesis Data Firehose to ingest logs from each service and write to S3.
D.Configure each source service to deliver logs directly to the central S3 bucket.
AnswerD

CloudTrail, VPC Flow Logs and ALB each natively publish logs to S3, so pointing them at one bucket centralises delivery with no pipeline service to manage. This satisfies the minimal-configuration constraint, though each source writes its own prefix scheme.

Why this answer

CloudTrail, VPC Flow Logs, and ALB logs can each be configured to deliver logs directly to a specified S3 bucket, including a central bucket, with no intermediary service required. This approach minimizes configuration overhead and avoids data movement costs, as each service writes natively to S3 using its own built-in delivery mechanism.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing a data pipeline or ETL service (like Data Pipeline or Glue) when the simplest and most efficient method is to configure each source service to write directly to the central S3 bucket, leveraging native AWS integrations.

How to eliminate wrong answers

Option A is wrong because AWS Data Pipeline is designed for scheduled data movement and transformation between data stores, not for real-time log ingestion from multiple AWS services; it would require custom pipeline definitions and adds unnecessary complexity. Option B is wrong because AWS Glue is an ETL service for crawling, cataloging, and transforming data, not a log collection or delivery service; it cannot natively ingest logs from CloudTrail, VPC Flow Logs, or ALB logs without first having the data in S3. Option C is wrong because Amazon Kinesis Data Firehose can ingest streaming data but does not natively subscribe to CloudTrail, VPC Flow Logs, or ALB logs; these services do not send data to Firehose directly, requiring additional setup like CloudWatch Logs subscriptions or custom agents.

683
MCQmedium

A data engineer manages an Amazon S3 data lake with millions of small JSON files. To improve query performance with Amazon Athena, the engineer wants to compact these files into larger Parquet files. The engineer must also minimize ongoing storage costs. Which solution should the engineer implement?

A.Create an AWS Lambda function that triggers on each S3 PUT event to merge JSON files into larger files.
B.Use AWS Glue ETL jobs to read the JSON files, transform them to Parquet, and write the output to a new S3 prefix. Then, configure an S3 Lifecycle rule to expire the original JSON objects after a retention period.
C.Use Amazon Kinesis Data Firehose to stream the JSON files into Parquet format in real time.
D.Enable S3 Transfer Acceleration on the bucket to speed up read operations for Athena queries.
AnswerB

AWS Glue ETL can efficiently convert JSON to Parquet, reducing file count and improving Athena query performance. Writing to a new prefix preserves the original data until verified. An S3 Lifecycle rule to expire the original JSON objects after a retention period reduces storage costs without immediate data loss. This approach is scalable and aligns with best practices for data lake optimization.

Why this answer

Converting JSON to Parquet with AWS Glue ETL compacts files and enables columnar storage, which improves Athena query performance. Writing to a new prefix and using an S3 Lifecycle rule to expire original JSON files after a retention period reduces storage costs while preserving data integrity. This combination addresses both performance and cost requirements.

Exam trap

The trap here is assuming that S3 Transfer Acceleration or streaming services like Kinesis Data Firehose can optimize existing batch data in S3, when they are designed for different use cases.

684
MCQmedium

A data engineer needs to store semi-structured JSON data that is accessed infrequently but must be retrievable within minutes. The data is generated by IoT devices and each object is about 500 KB. The engineer wants the most cost-effective storage solution. Which AWS service should be used?

A.Amazon S3 Glacier Deep Archive
B.Amazon S3 Standard
C.Amazon S3 Standard-Infrequent Access (S3 Standard-IA)
D.Amazon Elastic Block Store (EBS)
AnswerC

S3 Standard-IA charges lower storage rates than S3 Standard while retaining millisecond retrieval, meeting the minutes-level access requirement. At 500 KB per object, each exceeds the 128 KB minimum billable size, so per-object overhead stays negligible for infrequent IoT data.

Why this answer

Amazon S3 Standard-Infrequent Access (S3 Standard-IA) is the correct choice because it is designed for data that is accessed infrequently but requires rapid retrieval (within minutes). The 500 KB JSON objects from IoT devices fit the use case, and S3 Standard-IA offers lower storage costs than S3 Standard while maintaining the same low-latency retrieval performance, making it the most cost-effective option for this scenario.

Exam trap

The trap here is that candidates often confuse 'infrequent access' with 'archival' and choose Glacier Deep Archive, overlooking the retrieval time requirement of 'within minutes' which S3 Standard-IA satisfies but Glacier does not.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Glacier Deep Archive is intended for long-term archival data with retrieval times of 12 hours or more, not within minutes, and its retrieval costs are higher for urgent access. Option B is wrong because Amazon S3 Standard is optimized for frequently accessed data with higher storage costs, making it less cost-effective for infrequently accessed IoT data. Option D is wrong because Amazon Elastic Block Store (EBS) is a block-level storage service designed for EC2 instances, not for storing semi-structured JSON objects as a standalone data store, and it incurs costs even when not in use.

685
MCQmedium

A data engineer is configuring an Amazon Redshift cluster to encrypt data at rest. The company policy requires that encryption keys be stored in AWS CloudHSM. Which integration should the engineer use to meet this requirement?

A.Use AWS KMS with a customer managed key.
B.Configure Redshift to use an HSM for encryption.
C.Enable encryption using the AWS Redshift SSL/TLS feature.
D.Use Redshift automatic key rotation.
AnswerB

Configuring Redshift with a hardware security module directly satisfies the CloudHSM key-storage mandate, since Redshift supports HSM connections for at-rest encryption via CloudHSM. This keeps keys inside the dedicated HSM rather than AWS-managed KMS, meeting the policy's explicit key custody constraint.

Why this answer

Amazon Redshift supports encryption at rest using either AWS KMS or a hardware security module (HSM) via CloudHSM. When company policy mandates that encryption keys be stored in AWS CloudHSM, the engineer must configure Redshift to use an HSM connection for encryption, which is the only option that satisfies the CloudHSM key storage requirement.

Exam trap

The trap here is confusing encryption in transit (SSL/TLS) with encryption at rest, and assuming AWS KMS keys are stored in CloudHSM when they are actually stored in AWS-managed HSMs.

How to eliminate wrong answers

Option A is wrong because AWS KMS with a customer managed key stores keys in KMS-managed HSMs, not in the customer's dedicated AWS CloudHSM cluster, so it fails the policy requirement. Option C is wrong because SSL/TLS encrypts data in transit, not data at rest, and does not address key storage. Option D is wrong because automatic key rotation is a key lifecycle feature, not a key storage mechanism, and does not involve CloudHSM.

686
Multi-Selectmedium

A data engineer is managing an Amazon S3 data lake that contains raw JSON data. The engineer needs to optimize the data lake for query performance and cost when using Amazon Athena. The data is currently stored in a single S3 prefix without partitioning, and queries often filter on `event_type` and `event_date`. The engineer wants to implement best practices for Athena. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Enable S3 Transfer Acceleration on the bucket.
B.Convert the JSON data to Apache Parquet format.
C.Compress the JSON files using gzip.
D.Partition the data by `event_type` and `event_date` in S3.
E.Use Amazon S3 Select to filter data before querying with Athena.
AnswersB, D

Converting JSON to Parquet reduces the amount of data scanned by Athena because Parquet is columnar and compressed. Athena can read only the columns needed for a query, significantly lowering cost and improving performance. JSON is row-based and not splittable, so queries scan more data. Parquet also supports efficient compression and encoding. This is a fundamental optimization for Athena.

Why this answer

The two most effective actions are converting JSON to Parquet and partitioning by `event_type` and `event_date`. Parquet's columnar format reduces data scanned, and partitioning enables partition pruning. Together, they minimize query cost and improve performance.

Other options either do not affect Athena queries or are less effective. These are core best practices for optimizing Athena on S3 data lakes.

Exam trap

The trap here is considering gzip compression as sufficient, but it does not provide the columnar benefits of Parquet and may not be splittable.

687
MCQeasy

A data engineer needs to transfer 50 TB of historical data from an on-premises HDFS cluster to Amazon S3. The network bandwidth is limited to 100 Mbps. The transfer must be completed within one week. Which service should be used?

A.AWS DataSync
B.AWS Snowball
C.Amazon CloudFront
D.AWS Database Migration Service (DMS)
AnswerB

At 100 Mbps, transferring 50 TB over the network takes roughly 46 days, far exceeding the one-week deadline. AWS Snowball ships data physically on a rugged appliance, bypassing the bandwidth constraint and completing the migration within the required timeframe.

Why this answer

AWS Snowball is the correct choice because transferring 50 TB over a 100 Mbps network would take approximately 46 days (50 TB * 8 bits/byte / 100 Mbps / 86400 seconds/day), far exceeding the one-week deadline. Snowball provides a physical storage device that can be shipped to the on-premises location, allowing data to be loaded locally and shipped to AWS, bypassing network bandwidth constraints entirely.

Exam trap

The trap here is that candidates may underestimate the time required for online transfer and choose AWS DataSync, failing to calculate that 50 TB at 100 Mbps takes over 46 days, not one week, and overlooking Snowball's physical shipping approach for offline data transfer.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is designed for online data transfer over the network, and at 100 Mbps, it would take over 46 days to transfer 50 TB, which does not meet the one-week requirement. Option C is wrong because Amazon CloudFront is a content delivery network (CDN) for caching and distributing content to edge locations, not a data transfer service for ingesting large volumes of historical data into S3. Option D is wrong because AWS Database Migration Service (DMS) is specialized for migrating databases (e.g., relational, NoSQL) and does not support transferring HDFS files or large-scale file-based data to S3.

688
MCQmedium

A company stores sensitive data in Amazon S3 and requires that all data be encrypted at rest. The data is accessed by multiple AWS services. Which solution meets the encryption requirement with the LEAST operational overhead?

A.Use server-side encryption with AWS KMS (SSE-KMS)
B.Use client-side encryption with AWS KMS
C.Use server-side encryption with customer-provided keys (SSE-C)
D.Enable S3 default encryption with SSE-S3
AnswerD

SSE-S3 lets Amazon S3 manage the encryption keys entirely, so no key policies, grants or per-service configuration are needed. It satisfies the at-rest encryption requirement for data accessed by multiple AWS services while adding the least operational overhead, unlike SSE-KMS or client-side encryption.

Why this answer

Enabling S3 default encryption with SSE-S3 applies AES-256 encryption automatically to all objects at rest with zero key management overhead. AWS manages the keys, and no per-request KMS calls are needed, making it the lowest operational overhead solution that meets the encryption-at-rest requirement.

Exam trap

DEA-C01 often tests the trade-off between security control and operational overhead, causing candidates to choose SSE-KMS for its stronger key management when the question emphasizes 'least operational overhead,' where SSE-S3 is the correct answer.

How to eliminate wrong answers

Option A is wrong because SSE-KMS, while secure, adds operational overhead: KMS key policies must be managed, and each request incurs a KMS API call (with potential throttling and cost), which is more overhead than SSE-S3. Option B is wrong because client-side encryption requires the application to encrypt data before uploading and manage keys, adding significant operational complexity. Option C is wrong because SSE-C requires the customer to provide and manage encryption keys on every request, which is the highest operational overhead among the options.

689
MCQmedium

A company uses AWS Glue to catalog data in Amazon S3. The data arrives in Parquet format, but the crawler fails to update the schema when new columns are added. What is the most likely cause?

A.The Glue Data Catalog is not configured to accept schema changes.
B.Parquet files do not support schema evolution.
C.The crawler configuration has 'Update the table definition' set to 'Ignore the change'.
D.The S3 bucket has versioning disabled.
AnswerC

If set to ignore changes, the crawler will not add new columns.

Why this answer

AWS Glue crawlers have a table-update behavior controlled by the 'Update the table definition in the Data Catalog' setting (SchemaChangePolicy.UpdateBehavior). If it is set to 'Ignore the change' (LOG), the crawler detects new columns but does not modify the table definition, so the schema appears stale. Setting it to 'Update the table definition' (UPDATE_IN_DATABASE) allows the crawler to add new columns to the catalog table.

Exam trap

DEA-C01 often tests the assumption that Parquet cannot evolve schemas or that the Data Catalog has a global setting — candidates miss that the crawler's per-run schema change policy is the actual control.

How to eliminate wrong answers

Option A is wrong because the Glue Data Catalog does accept schema changes — there is no global 'accept schema changes' toggle; the behavior is controlled per-crawler via the schema change policy. Option B is wrong because Parquet absolutely supports schema evolution (adding columns is a core feature of the format, and readers can handle missing columns with nulls). Option D is wrong because S3 bucket versioning is unrelated to Glue crawler schema detection — versioning affects object retention, not catalog updates.

690
MCQeasy

A company stores application logs in Amazon S3 and uses AWS Glue crawlers to populate the AWS Glue Data Catalog. A data engineer needs to query the logs with Amazon Athena. The logs are partitioned by year/month/day in S3, but Athena queries are scanning all partitions and returning errors about missing partitions. What should the engineer do to enable partition pruning?

A.Run the AWS Glue crawler with the option to update the table's partition metadata and ensure the crawler has permission to list the partition folders.
B.Convert the logs to Parquet format and re-crawl the data.
C.Increase the Athena query timeout and retry the queries.
D.Move the logs to a single prefix without date partitions and update the crawler.
AnswerA

The crawler must be able to list the partition folders and add partition metadata to the Data Catalog. If partitions are missing, Athena cannot prune and may error. Configuring the crawler to detect partitions and granting S3 list permissions ensures the catalog is updated, enabling Athena to use partition pruning and avoid scanning all data.

Why this answer

For Athena to use partition pruning, the AWS Glue Data Catalog must contain accurate partition metadata. The crawler needs permission to list the S3 partition folders and must be configured to update partitions. Once partitions are registered, Athena can prune and avoid scanning irrelevant data, resolving the errors and improving performance.

Exam trap

The trap here is focusing on file format or query settings instead of ensuring the crawler can discover and register partitions in the Data Catalog.

691
MCQhard

A media company stores millions of small JSON files in an Amazon S3 bucket and queries them with Amazon Athena. Analysts report that queries scan far more data than expected, and costs are rising. The data engineer confirms that the files are uncompressed, use no partitioning, and are stored as newline-delimited JSON. Which change will MOST reduce the data scanned per query?

A.Convert the files to Apache Parquet, compress with Snappy, and partition the S3 prefix by event date.
B.Add more small files to increase parallelism across Athena workers.
C.Enable S3 Transfer Acceleration on the bucket.
D.Increase the Athena workgroup's data usage control limit.
AnswerA

Parquet is columnar, so Athena reads only the columns referenced in the query, and Snappy compression reduces bytes scanned further. Partitioning by event date lets the query engine prune irrelevant prefixes entirely. Together these three changes attack data scanned from three angles: column pruning, compression, and partition pruning, delivering the largest reduction.

Why this answer

Converting to a columnar format, compressing with a splittable codec, and partitioning by a common filter column collectively minimize bytes read. Parquet enables column pruning, Snappy reduces stored size, and date partitioning enables partition pruning so only relevant prefixes are read. This combination produces the largest reduction in data scanned and therefore in Athena query cost.

Exam trap

The trap here is treating operational knobs such as Transfer Acceleration or workgroup limits as performance optimizations for Athena scan volume.

692
MCQhard

A data engineer runs an AWS Glue job that reads Parquet files from Amazon S3 partitioned by year/month/day and writes to another S3 prefix. The job currently processes all historical partitions on every run, causing long runtimes and high cost. The engineer wants subsequent runs to process only new data. Which configuration should the engineer apply?

A.Configure an S3 Lifecycle rule to transition old partitions to S3 Glacier Instant Retrieval so the Glue job reads them faster.
B.Increase the number of AWS Glue DPUs allocated to the job so it can process all partitions faster.
C.Enable AWS Glue job bookmarks and ensure the job's source and target paths are stable across runs.
D.Set the Glue job's MaxConcurrentRuns parameter to 1 to prevent overlapping executions from reprocessing data.
AnswerC

Job bookmarks persist state about which partitions and files have already been processed, so reruns skip previously handled data. This directly reduces runtime and cost by limiting each run to new partitions. Bookmarks require stable source and target paths and a consistent transformation script to remain effective.

Why this answer

Job bookmarks maintain persistent state across runs, allowing Glue to skip files and partitions that were already processed in prior executions. This is the native mechanism for incremental processing in Glue ETL, directly reducing scanned data, runtime, and cost without changing the transformation logic or adding external state tracking.

Exam trap

The trap here is confusing concurrency or compute scaling controls with incremental processing, when only job bookmarks track previously processed partitions.

693
Multi-Selecthard

A data engineer is designing a data lake on Amazon S3 for a retail company. The company ingests point-of-sale data as small JSON files every few minutes, totaling about 5 GB per day. Analysts query the data with Amazon Athena, and costs are rising due to many small files and full scans. The engineer wants to reduce Athena query costs and improve performance while keeping the data in S3. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Increase the Athena workgroup data usage control limit to allow larger scans.
B.Run an AWS Glue ETL job or AWS Glue compaction to merge small files into larger files of 128 MB or more.
C.Enable Amazon S3 Intelligent-Tiering on the bucket to automatically move data to lower-cost storage classes.
D.Convert the JSON files to Apache Parquet and store them partitioned by date and store ID.
E.Enable S3 Transfer Acceleration on the bucket to speed up Athena queries.
AnswersB, D

Compacting many small files into larger ones reduces the number of S3 GET requests and metadata overhead, which lowers Athena query latency and cost. Glue jobs or Glue compaction can perform this merge while keeping the data in S3, complementing the Parquet and partition changes.

Why this answer

Athena cost and performance depend on bytes scanned and the number of files read. Converting JSON to partitioned Parquet enables column pruning and partition pruning, while compacting small files into larger ones reduces request overhead and metadata processing. These two actions together cut scanned bytes and improve query speed without leaving S3.

Exam trap

The trap here is confusing storage-cost optimizations like Intelligent-Tiering with query-cost optimizations, which depend on scanned bytes and file layout.

694
MCQmedium

A company ingests application logs into Amazon S3 through Amazon Kinesis Data Firehose. The logs arrive as newline-delimited JSON, and analysts query them with Amazon Athena. Query performance is poor because the JSON files are small and uncompressed. The engineer must improve Athena query performance while keeping the raw JSON available for a downstream legacy system. Which change should the engineer make?

A.Enable S3 Transfer Acceleration on the target bucket to speed up Athena scans.
B.Configure the Athena table with a SerDe for JSON and add a WHERE clause on the timestamp column in every query.
C.Increase the Firehose buffer size and interval so larger objects are delivered to S3.
D.Add a Firehose record format conversion to Apache Parquet and keep the raw JSON in a separate S3 prefix using a second delivery stream.
AnswerD

Converting to Parquet with a Firehose record format conversion gives Athena a columnar, compressed format that scans far less data, directly improving query performance. Delivering the raw JSON to a separate prefix through a second stream preserves the legacy feed, satisfying both requirements without duplicating logic in the consumers.

Why this answer

Athena performance improves when it scans less data, and converting the stream to Parquet with a record format conversion makes the stored objects columnar and compressed. A second delivery stream that leaves the raw JSON in place satisfies the legacy consumer, so the engineer gets both performance and compatibility.

Exam trap

The trap here is assuming buffer tuning or query-side filters can substitute for changing the storage format, when only a columnar, compressed layout reduces the bytes Athena actually reads.

695
MCQhard

A data engineer is using AWS Glue to run an ETL job that reads from an Amazon RDS for PostgreSQL database and writes to Amazon S3. The job is configured with a JDBC connection to RDS. The engineer notices that the job fails intermittently with a 'Connection timed out' error. The RDS instance is in a private subnet, and the Glue job has been configured with a VPC connection. Which action should the engineer take to resolve the timeout issue?

A.Enable the Glue job's VPC connection to use a NAT gateway for outbound traffic.
B.Increase the size of the Glue job's DPU capacity to handle more concurrent connections.
C.Ensure the Glue job's VPC connection has a security group that allows outbound traffic to the RDS instance's security group on the database port.
D.Configure the Glue job to use a public internet gateway to connect to RDS.
AnswerC

When AWS Glue runs in a VPC, it uses an elastic network interface (ENI) with a security group. To connect to RDS, the RDS security group must allow inbound traffic from the Glue ENI's security group on the database port. Additionally, the Glue security group must allow outbound traffic to RDS. Misconfigured security groups are a common cause of connection timeouts.

Why this answer

The intermittent connection timeout is likely caused by security group rules blocking traffic between the AWS Glue VPC connection and the RDS instance. The Glue job's ENI security group must allow outbound traffic to RDS, and the RDS security group must allow inbound traffic from the Glue security group on the database port. Ensuring these rules are correctly configured resolves the issue.

Exam trap

The trap here is assuming that scaling resources or using NAT gateways will fix network connectivity, when the issue is typically security group or subnet routing misconfiguration.

696
MCQeasy

A company wants to enforce that all data written to an S3 bucket is encrypted with a customer-managed AWS KMS key. The data engineer has created the KMS key and attached an S3 bucket policy. However, users are still able to upload objects without specifying the KMS key. What is the most likely cause?

A.The S3 bucket policy does not include a condition that denies s3:PutObject without the correct encryption
B.The S3 bucket has default encryption enabled with SSE-S3
C.The KMS key policy does not grant the users kms:Encrypt permission
D.The IAM role for the users does not have s3:PutObject permission
AnswerA

The bucket policy lacks a Deny statement with the `s3:PutObject` action and a `StringNotEquals` condition on `s3:x-amz-server-side-encryption` (and the KMS key ARN), so uploads without the customer-managed key are never refused. Without an explicit deny, S3 permits unencrypted or default-encrypted writes despite the policy.

Why this answer

The most likely cause is that the S3 bucket policy does not include a condition that denies s3:PutObject requests unless the correct encryption (e.g., aws:kms with the specific key) is specified. Without an explicit deny condition, users can still upload objects using the default encryption or no encryption if their IAM permissions allow it. The bucket policy must enforce the encryption requirement by denying non-compliant uploads.

Exam trap

DEA-C01 often tests the difference between default encryption and policy-enforced encryption. Candidates may think that enabling default encryption is sufficient, but it only applies when no encryption is specified, so users can still upload without the KMS key if they don't specify encryption.

How to eliminate wrong answers

Option B is wrong because default encryption with SSE-S3 would not prevent users from uploading without specifying a KMS key; it would just encrypt with SSE-S3 if no encryption is specified. Option C is wrong because the KMS key policy granting kms:Encrypt is necessary for users to use the key, but if they are not specifying the key, the issue is not the key policy. Option D is wrong because if the IAM role lacked s3:PutObject, users would not be able to upload at all, not just without the KMS key.

697
MCQmedium

A company uses AWS Glue to process streaming data from Amazon Kinesis Data Streams. The job fails intermittently with a 'MemoryError'. What is the MOST likely cause?

A.The Glue job worker type is too small for the data volume
B.The Glue job uses too many DynamicFrames
C.The S3 output bucket is in a different region
D.The Kinesis stream has insufficient shards
AnswerA

MemoryError indicates the worker's heap is exhausted during processing. A too-small worker type provides insufficient memory for the Kinesis stream's volume, so the job fails intermittently as load spikes. Scaling the worker type directly addresses the memory constraint.

Why this answer

The 'MemoryError' in AWS Glue indicates that the worker type allocated to the job does not have sufficient memory to process the data volume. Glue workers (Standard, G.1X, G.2X) have fixed memory allocations (e.g., 16 GB for Standard), and if the streaming data from Kinesis exceeds this, the job fails. Increasing the worker type or the number of workers resolves this.

Exam trap

The trap here is that candidates confuse memory errors with throttling or connectivity issues, leading them to pick insufficient shards (Option D) or cross-region problems (Option C), when the root cause is almost always an undersized worker type for the data volume.

How to eliminate wrong answers

Option B is wrong because using too many DynamicFrames does not directly cause a MemoryError; DynamicFrames are lazy transformations and memory issues arise from data volume or worker size, not the number of frames. Option C is wrong because an S3 output bucket in a different region would cause a cross-region access error (e.g., AccessDenied or timeout), not a MemoryError. Option D is wrong because insufficient Kinesis shards cause throttling (ProvisionedThroughputExceededException) or data latency, not a memory exhaustion in the Glue job.

698
MCQhard

A data engineer is using AWS Glue to process a large dataset stored in Amazon S3 in Parquet format. The Glue job performs a join between two tables and writes the result back to S3. The engineer notices that the job is running slowly and consuming excessive DPU hours. The job has 10 workers of type G.1X. Which action should the engineer take to improve performance and reduce cost?

A.Use the 'ApplyMapping' transformation to rename columns before the join.
B.Enable AWS Glue job bookmarks to avoid reprocessing old data.
C.Increase the number of workers to 20 and keep the worker type as G.1X.
D.Partition the Parquet data by the join key and use broadcast join if one table is small.
AnswerD

Partitioning the data by the join key can reduce data shuffling during the join, and using a broadcast join for a small table avoids a full shuffle altogether. These optimizations can significantly improve performance and reduce DPU hours. This approach directly addresses the join bottleneck and is a best practice in AWS Glue.

Why this answer

The performance issue is likely due to inefficient join operations causing large data shuffles. Partitioning the Parquet data by the join key aligns data with the join operation, reducing shuffling. If one table is small enough to fit in memory, a broadcast join eliminates the shuffle entirely.

These optimizations reduce runtime and DPU consumption. Other options either do not address the join bottleneck or could increase cost without solving the problem.

Exam trap

The trap here is assuming that adding more workers always improves performance, when the real issue may be data skew or lack of partitioning.

699
MCQmedium

A data engineer needs to migrate an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB and has a continuous stream of write operations. The migration should minimize downtime. Which AWS service should be used?

A.AWS DataSync
B.AWS Database Migration Service (DMS)
C.AWS Snowball Edge
D.AWS Glue
AnswerB

AWS Database Migration Service performs continuous replication from the on-premises PostgreSQL source to Amazon RDS for PostgreSQL while the source stays live, then cuts over during a brief window. This satisfies the requirement to minimise downtime for a 2 TB database with ongoing writes.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it supports ongoing replication (change data capture) from an on-premises PostgreSQL source to Amazon RDS for PostgreSQL, enabling a near-zero downtime migration. DMS can handle the 2 TB dataset and continuous write stream by performing a full load followed by continuous replication of changes until the cutover. Other services lack the ability to perform live, transactional replication with minimal interruption.

Exam trap

The trap here is that candidates often choose AWS DataSync (Option A) because they confuse it with a database migration tool, but DataSync cannot replicate live transactional changes and is meant for file or object storage, not relational databases with ongoing writes.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is designed for one-time or periodic bulk data transfers between on-premises storage and AWS, not for continuous database replication or minimizing downtime during a live database migration. Option C is wrong because AWS Snowball Edge is a physical device for offline data transfer, which would require stopping writes to the database to export the data, causing significant downtime and not supporting ongoing replication. Option D is wrong because AWS Glue is a serverless data integration service for ETL (extract, transform, load) jobs, not a database migration tool; it cannot perform live replication or handle continuous write streams from a source database.

700
MCQeasy

A company stores sensitive financial data in Amazon S3 and requires that all data be encrypted at rest using customer-managed keys. A data engineer configures the S3 bucket to use SSE-KMS with a customer-managed KMS key. The security team now wants to audit all API calls that use the KMS key to decrypt data. Which AWS service should the engineer use to capture and review these KMS API calls?

A.AWS Config
B.AWS Trusted Advisor
C.Amazon CloudWatch Logs
D.AWS CloudTrail
AnswerD

AWS CloudTrail logs all API activity, including KMS operations such as Decrypt, Encrypt, and GenerateDataKey. By enabling CloudTrail, the engineer can capture and review these calls for auditing. CloudTrail is the standard service for auditing API activity across AWS services.

Why this answer

AWS CloudTrail is the service that records API activity in your AWS account, including KMS operations. By enabling CloudTrail, you can audit who used the KMS key and when. AWS Config, CloudWatch Logs, and Trusted Advisor do not provide the same level of API call logging for KMS.

Exam trap

The trap here is assuming that AWS Config or CloudWatch Logs automatically capture API calls, when actually CloudTrail is the dedicated service for API auditing.

701
MCQeasy

A data engineer needs to restrict access to an S3 bucket so that only users from a specific AWS account can read objects. Which S3 bucket policy element should be used?

A.Action
B.Principal
C.Resource
D.Condition
AnswerB

The Principal element names the AWS account permitted to read objects, directly satisfying the cross-account restriction. Bucket policies evaluate Principal against the requesting identity, so specifying the trusted account's ARN grants access solely to its users, excluding all other accounts.

Why this answer

The Principal element in an S3 bucket policy specifies the AWS account, user, or role that the policy applies to. By setting the Principal to a specific AWS account ID, only users from that account can read objects. Option A is wrong because the Action element specifies the allowed or denied operations (e.g., s3:GetObject), not the account.

Option C is wrong because the Resource element identifies the bucket or objects, not the requester. Option D is wrong because the Condition element adds additional constraints (e.g., IP address), but it is not the primary element for specifying the allowed account; Principal is the correct element for that purpose.

702
MCQmedium

A data engineer needs to ingest data from an Amazon DynamoDB table into an Amazon S3 data lake. The table is updated frequently and the engineer must capture all item-level changes in near real-time without impacting table performance. The ingested data must be stored in a format that preserves the change type (INSERT, MODIFY, REMOVE). Which solution meets these requirements with the LEAST operational overhead?

A.Enable DynamoDB Accelerator (DAX) and use it to stream changes to Amazon S3.
B.Schedule an AWS Glue job to periodically scan the entire DynamoDB table and write the full dataset to Amazon S3.
C.Use AWS Database Migration Service (AWS DMS) with ongoing replication to read from DynamoDB and write to Amazon S3.
D.Enable DynamoDB Streams on the table and use AWS Lambda to read the stream and write the records to Amazon S3 in JSON format.
AnswerD

DynamoDB Streams captures item-level changes in near real-time and Lambda can process these records and write them to S3. This preserves the change type and requires no server management, meeting the low operational overhead requirement. It is the standard serverless pattern for DynamoDB change data capture to S3.

Why this answer

DynamoDB Streams provides a time-ordered sequence of item-level modifications in near real-time, and AWS Lambda can process these records and write them to S3. This serverless approach requires minimal operational overhead and preserves the change type. Other options either do not capture change type, are batch-oriented, or add unnecessary management overhead.

Exam trap

The trap here is assuming that AWS DMS captures change type for DynamoDB, but it does not; it replicates the current state without indicating the operation.

703
Multi-Selecthard

A company runs an Amazon RDS for PostgreSQL instance for an OLTP application. The database size is 500 GB. The company wants to minimize downtime during backups and ensure point-in-time recovery (PITR) for the last 7 days. Which TWO features should the company use? (Choose TWO.)

Select 2 answers
A.Enable Multi-AZ deployment for high availability.
B.Create a read replica in a different Availability Zone.
C.Enable automated backups with a retention period of 7 days.
D.Create daily manual snapshots and copy them to another region.
E.Enable Enhanced Monitoring to track backup progress.
AnswersA, C

Multi-AZ reduces downtime during automated backups by taking backups from the standby.

Why this answer

Multi-AZ deployment for Amazon RDS provides high availability by automatically provisioning and maintaining a synchronous standby replica in a different Availability Zone. This minimizes downtime during backups by allowing automated backups to be taken from the standby instance, eliminating I/O suspension on the primary. Option C is correct because enabling automated backups with a retention period of 7 days enables point-in-time recovery (PITR) within that window, restoring the database to any second within the retention period using transaction logs.

Exam trap

The trap here is that candidates often confuse read replicas or manual snapshots with backup and recovery features, failing to recognize that only automated backups with a retention period enable point-in-time recovery, and that Multi-AZ is required to minimize downtime during backups by offloading them to the standby instance.

704
MCQhard

A company uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data is transformed using an AWS Lambda function. Some records fail transformation and are lost because the Lambda function throws an exception. The data engineer needs to capture the failed records for analysis without affecting the pipeline. What should the engineer do?

A.Configure the Firehose delivery stream to send failed records to a backup S3 bucket
B.Increase the buffer size of the Firehose stream
C.Disable the Lambda transformation and process all records in batch later
D.Modify the Lambda function to write failed records to Amazon DynamoDB
AnswerA

Firehose can be configured to send failed records to a backup S3 bucket.

Why this answer

Amazon Kinesis Data Firehose natively supports a 'backup' or 'error output' configuration: when a Lambda transformation fails (e.g., throws an exception), Firehose can automatically route the failed records to a designated Amazon S3 bucket for failed data. This allows the engineer to capture and analyze the failed records without blocking or altering the main delivery pipeline, as the stream continues processing successfully transformed records.

Exam trap

The trap here is that candidates often confuse Firehose's 'backup' mode (which copies all records) with the 'error output' configuration (which only captures failed records), or they assume that modifying the Lambda function to handle failures is the only option, missing the native Firehose feature that requires no code changes.

How to eliminate wrong answers

Option B is wrong because increasing the buffer size (e.g., MBs or interval) only affects how long Firehose waits before delivering a batch; it does not capture or handle Lambda transformation failures. Option C is wrong because disabling the Lambda transformation would stop all data transformation, which is a core requirement of the pipeline, and processing records in batch later would not recover records already lost during the stream. Option D is wrong because while writing failed records to DynamoDB from within the Lambda function is possible, it requires modifying the Lambda code and does not leverage Firehose's built-in error handling; moreover, DynamoDB has a 400 KB item size limit and is not designed for high-volume streaming failure capture, making it less suitable than the native S3 backup bucket.

705
MCQeasy

A company uses Amazon Redshift for data warehousing. They notice that query performance has degraded over time. Which maintenance operation should be performed to improve performance?

A.Run the VACUUM command
B.Drop and recreate the table
C.Run the REINDEX command
D.Run the ANALYZE command
AnswerA

VACUUM reclaims space from deleted rows and re-sorts unsorted data, restoring the physical ordering that Redshift's zone maps rely on for block skipping. Degraded performance from accumulated deletes and unsorted loads is exactly the maintenance gap it addresses, so it satisfies the stem's requirement directly.

Why this answer

The VACUUM command in Amazon Redshift reclaims space from deleted rows and re-sorts rows to restore the sort order, which improves query performance by reducing I/O and enabling more efficient scans. Over time, as data is updated or deleted, the physical storage becomes fragmented and sort order degrades, leading to slower queries. Running VACUUM (or VACUUM SORT ONLY, VACUUM DELETE ONLY, etc.) is the recommended maintenance operation to address this degradation.

Exam trap

DEA-C01 often tests the distinction between VACUUM (space reclamation and sort) and ANALYZE (statistics update), causing candidates to choose ANALYZE when the issue is performance degradation due to fragmentation.

How to eliminate wrong answers

Option B is wrong because dropping and recreating a table is a drastic, disruptive action that would lose data and require reloading, and it is not a standard maintenance operation for performance improvement. Option C is wrong because REINDEX is not a Redshift command; Redshift does not have a REINDEX operation like some other databases. Option D is wrong because ANALYZE updates table statistics used by the query planner, which can improve query plans, but it does not reclaim space or restore sort order; it is complementary to VACUUM but not the primary operation for performance degradation due to fragmentation.

706
MCQeasy

A company needs to transfer 20 TB of historical data from an on-premises Hadoop cluster to Amazon S3. The network bandwidth is limited and the transfer must complete within one week. Which service should the company use?

A.Amazon S3 Transfer Acceleration
B.AWS Snowball Edge
C.AWS Direct Connect with DataSync
D.AWS DataSync over a VPN connection
AnswerB

AWS Snowball Edge ships a physical appliance to the on-premises site, bypassing the limited network bandwidth constraint. It transfers 20 TB within a week, whereas online transfer methods would be too slow over the constrained link.

Why this answer

AWS Snowball Edge is a petabyte-scale data transport solution that uses physical devices to transfer large amounts of data into and out of AWS, bypassing network bandwidth limitations. For 20 TB over a limited network within one week, Snowball Edge is the most practical and cost-effective choice. It supports up to 80 TB usable capacity per device and provides encryption and tamper-resistant security.

Exam trap

DEA-C01 often tests the threshold where network transfer becomes impractical, so candidates may overlook Snowball for 20 TB and incorrectly choose online transfer methods like DataSync or Transfer Acceleration.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration speeds up transfers over the internet but still relies on available bandwidth, which is limited. Option C is wrong because AWS Direct Connect with DataSync requires a dedicated network connection that may not be feasible within the timeframe and does not solve bandwidth constraints. Option D is wrong because DataSync over VPN is also limited by network bandwidth and would likely not complete 20 TB within a week.

707
MCQmedium

A company stores sensitive data in an Amazon S3 bucket. A compliance requirement mandates that all data must be encrypted at rest with a key that is automatically rotated every year. The company also needs to maintain an audit trail of who used the key. Which solution meets these requirements?

A.Use AWS KMS customer managed keys (SSE-KMS) with automatic key rotation enabled.
B.Use customer-provided encryption keys (SSE-C) and rotate keys manually.
C.Use S3 managed keys (SSE-S3) and enable S3 server access logs.
D.Configure a bucket policy to enforce encryption using the 'aws:SecureTransport' condition.
AnswerA

AWS KMS customer managed keys with SSE-KMS satisfy both constraints: automatic annual rotation is configurable, and every cryptographic operation is logged to CloudTrail, providing the required audit trail of key usage. AWS managed keys rotate automatically but cannot be audited per-customer or configured to a yearly schedule.

Why this answer

AWS KMS customer managed keys (SSE-KMS) with automatic key rotation enabled satisfies both requirements: it encrypts data at rest in S3 and automatically rotates the KMS key every year. Additionally, KMS integrates with AWS CloudTrail to log every API call (e.g., Decrypt, GenerateDataKey) that uses the key, providing an audit trail of who used the key and when.

Exam trap

The trap here is that candidates confuse SSE-S3's automatic rotation (which is invisible and lacks audit trails) with SSE-KMS's automatic rotation (which provides both rotation and CloudTrail logging), or they mistakenly think SSE-C or bucket policies can satisfy the audit trail requirement.

How to eliminate wrong answers

Option B is wrong because SSE-C requires the customer to provide and manage their own encryption keys, and AWS does not support automatic rotation for customer-provided keys — rotation must be done manually, which violates the compliance requirement for automatic yearly rotation. Option C is wrong because SSE-S3 uses S3-managed keys that are automatically rotated by AWS, but it does not provide a per-key audit trail of who used the key; S3 server access logs only record requests to the bucket, not granular key usage events. Option D is wrong because the 'aws:SecureTransport' condition enforces encryption in transit (HTTPS), not encryption at rest, and it does not involve key rotation or audit trails for key usage.

708
MCQmedium

A company uses AWS Glue to transform data in S3. The transformation job reads Parquet files, filters rows, and writes to another S3 bucket. The job takes longer than expected. Which change would MOST likely reduce the job execution time?

A.Use a single large file instead of multiple small files.
B.Reduce the number of partitions in the output data.
C.Convert the input files from Parquet to CSV format.
D.Increase the number of DPUs allocated to the Glue job.
AnswerD

Glue distributes work across DPUs, so adding DPUs increases parallel task capacity and shortens the filter-and-write job's runtime. This directly addresses the execution-time constraint, provided the job is not bottlenecked by a single small input partition.

Why this answer

Increasing the number of DPUs (Data Processing Units) allocated to the Glue job directly increases the parallelism of the Apache Spark-based execution environment. Since the job reads Parquet files, filters rows, and writes output, a bottleneck in compute capacity is the most likely cause of prolonged execution time. More DPUs allow Spark to distribute the workload across more executors, reducing overall runtime.

Exam trap

The trap here is that candidates often assume optimizing file format or output partitioning will always improve performance, but for a compute-bound transformation job, increasing parallelism via DPUs is the most direct solution.

How to eliminate wrong answers

Option A is wrong because using a single large file instead of multiple small files would reduce parallelism in Spark's file scanning, potentially increasing execution time due to less efficient task distribution. Option B is wrong because reducing the number of partitions in the output data affects the write phase but does not address the root cause of slow processing during the read and filter stages. Option C is wrong because converting from Parquet to CSV would increase I/O and CPU overhead due to CSV's lack of compression, schema enforcement, and columnar storage, making the job slower, not faster.

709
MCQhard

A company uses AWS Glue to process data from Amazon RDS MySQL into Amazon S3. The Glue job uses a JDBC connection and runs on a schedule. Recently, the job has been failing with a 'Communications link failure' error. The RDS instance is in a private subnet. Which troubleshooting step should the data engineer take FIRST?

A.Check the Glue job's DPU allocation; increase if too low.
B.Review the Glue job script for data type mismatches.
C.Verify that the Glue job's VPC subnet and security group allow outbound traffic to RDS.
D.Increase the RDS instance's max_connections parameter.
AnswerC

Glue jobs in a VPC need a subnet route and security group egress permitting TCP 3306 to the RDS instance; without it, the JDBC handshake fails with 'Communications link failure'. Verifying outbound connectivity from the Glue ENI directly addresses the private-subnet constraint before deeper checks.

Why this answer

The 'Communications link failure' error typically indicates a network connectivity issue between the Glue job and the RDS instance. Since the RDS instance is in a private subnet, the Glue job must be configured with a VPC subnet and security group that allows outbound traffic to the RDS instance's security group on port 3306 (MySQL). Without this network path, the JDBC connection cannot be established, making verifying the VPC and security group configuration the first logical troubleshooting step.

Exam trap

The trap here is that candidates often assume 'Communications link failure' is a database-side issue (like connection limits or timeouts) and jump to tuning RDS parameters, when in fact it is most commonly a network connectivity problem in a VPC environment.

How to eliminate wrong answers

Option A is wrong because DPU allocation affects job parallelism and memory, not network connectivity; increasing DPUs would not resolve a 'Communications link failure' caused by a missing network path. Option B is wrong because data type mismatches typically cause runtime errors during data conversion, not connection-level failures like 'Communications link failure'. Option D is wrong because increasing max_connections addresses connection limits on the RDS side, but the error indicates the Glue job cannot even reach the database, not that connections are being refused due to exhaustion.

710
MCQmedium

A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream and writes results to an Amazon S3 bucket. Recently, the application has been failing with 'ResourceNotFoundException' for the S3 bucket. What is the MOST likely cause?

A.The IAM role used by the application does not have s3:PutObject permission.
B.The S3 bucket ARN is incorrectly specified in the application configuration.
C.The Flink application code specifies the wrong AWS Region for the S3 bucket.
D.The S3 bucket has versioning disabled.
AnswerB

A malformed or incorrect S3 bucket ARN in the application configuration causes the Flink application's S3 sink to fail resolving the target bucket, producing ResourceNotFoundException. Correcting the ARN restores access without changing IAM or network settings.

Why this answer

The 'ResourceNotFoundException' for the S3 bucket indicates that the application cannot find the bucket. This is most likely due to an incorrectly specified bucket name or ARN in the application configuration. If the bucket name is misspelled or the ARN is malformed, the Kinesis Data Analytics application cannot resolve the bucket resource.

Option A is incorrect because missing 's3:PutObject' permission would cause an 'AccessDenied' error, not 'ResourceNotFoundException'. Option C is incorrect because the AWS Region for S3 is determined by the bucket's location, and specifying a wrong region in Flink code would typically result in a different error (e.g., 'IllegalArgumentException' or cross-region access issues) but not 'ResourceNotFoundException'. Option D is incorrect because versioning is unrelated to bucket existence; disabling versioning does not cause a resource not found error.

Exam trap

Candidates often confuse permission errors with resource-not-found errors. A missing IAM permission typically results in 'AccessDenied', not 'ResourceNotFoundException'.

711
MCQmedium

A data engineer is designing a solution to ingest streaming data from Amazon Kinesis Data Streams into an Amazon Redshift cluster for near-real-time analytics. The engineer needs to ensure that data is loaded efficiently and that the Redshift cluster can handle the ingestion load without impacting query performance. Which approach should the engineer use?

A.Use the Amazon Redshift Streaming Ingestion feature to directly ingest from Kinesis Data Streams into Redshift materialized views.
B.Use Amazon Kinesis Client Library (KCL) on an Amazon EC2 instance to read from Kinesis and insert data into Redshift using INSERT statements.
C.Use Amazon Kinesis Data Firehose to deliver the stream to an Amazon S3 bucket, then use a COPY command to load data into Redshift.
D.Use AWS Glue streaming ETL to read from Kinesis and write to Redshift using JDBC connections.
AnswerA

Redshift Streaming Ingestion allows you to create materialized views that directly consume from Kinesis Data Streams, providing low-latency ingestion without intermediate storage. This reduces load on the cluster and enables near-real-time analytics by querying the materialized views, which can be refreshed automatically.

Why this answer

Amazon Redshift Streaming Ingestion is designed to ingest data directly from Kinesis Data Streams into Redshift materialized views with low latency. It eliminates the need for intermediate storage and complex ETL, and it offloads ingestion from the cluster's query processing. This meets the near-real-time analytics requirement while minimizing impact on query performance.

Exam trap

The trap here is assuming that Kinesis Data Firehose with S3 and COPY is the only way to load streaming data into Redshift, overlooking the native streaming ingestion feature.

712
MCQmedium

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job needs to run daily and process only new data since the last run. The data in S3 is partitioned by date in the format `year=YYYY/month=MM/day=DD/`. Which feature should the data engineer use to track processed partitions and avoid reprocessing old data?

A.AWS Glue job bookmarks
B.AWS Glue crawler update schedules
C.AWS Glue triggers with a schedule
D.AWS Glue Data Catalog partition indexes
AnswerA

AWS Glue job bookmarks track the data that has already been processed in previous job runs. When reading from S3 with partitions, bookmarks keep track of the last processed partition and only process new partitions in subsequent runs. This is the recommended way to handle incremental processing without manual tracking.

Why this answer

AWS Glue job bookmarks are designed to track the state of data processed by a job. When reading from partitioned S3 data, bookmarks record the last processed partition and only process new partitions in subsequent runs. This enables incremental processing without manual intervention, making it the correct choice for avoiding reprocessing of old data.

Exam trap

The trap here is confusing scheduling or catalog maintenance features with state management, assuming that a trigger or crawler schedule would automatically handle incremental processing.

713
MCQeasy

Refer to the exhibit. A data engineer runs this CLI command on an S3 bucket. The data is ingested from multiple sources. Which AWS service would be best to process these files in a single batch transformation?

A.AWS Lambda
B.Amazon Kinesis Data Analytics
C.Amazon Athena
D.AWS Glue
AnswerD

AWS Glue is a serverless, managed extract-transform-load service that handles schema discovery via crawlers and runs distributed Spark jobs, so it can transform all files from multiple sources in one batch job without provisioning or managing servers.

Why this answer

AWS Glue is designed for batch ETL processing, capable of handling multiple files of varying sizes in a single transformation job. Option A is wrong because Lambda has limitations on execution time (15 minutes) and memory (10 GB), making it unsuitable for large-scale batch transforms. Option B is wrong because Kinesis Data Analytics is for real-time stream processing, not batch.

Option C is wrong because Athena is an interactive query service for ad-hoc analysis of data in S3, not for batch transformations.

714
MCQmedium

A data engineer must ingest data from an Amazon DynamoDB table into an Amazon S3 bucket for analytics. The table receives continuous writes, and the ingestion must capture changes with minimal latency. The engineer wants a fully managed solution that requires no server provisioning. Which approach should the engineer use?

A.Use AWS Database Migration Service (AWS DMS) with ongoing replication to replicate the DynamoDB table to Amazon S3.
B.Enable DynamoDB Streams on the table, and use AWS Lambda to read the stream and write the data to Amazon S3.
C.Use Amazon Kinesis Data Firehose to read directly from DynamoDB and deliver to Amazon S3.
D.Schedule an AWS Glue job to run every minute, reading the entire DynamoDB table and writing to Amazon S3.
AnswerB

DynamoDB Streams captures item-level changes in near real time, and Lambda can process these events without managing servers. This serverless pattern provides low-latency ingestion into S3. The engineer does not need to provision or scale any infrastructure, and the solution automatically scales with the stream volume, making it ideal for continuous writes.

Why this answer

DynamoDB Streams provides a time-ordered sequence of item-level changes, and Lambda can process these events in near real time. Writing to S3 from Lambda is a common serverless pattern that requires no server management and scales automatically. This combination meets the latency and managed-service requirements without the overhead of provisioning or full-table scans.

Exam trap

The trap here is assuming that AWS DMS or Kinesis Data Firehose can directly capture DynamoDB changes with low latency, when actually DynamoDB Streams is the native change capture mechanism.

715
Multi-Selecteasy

A company wants to audit API calls made to its Amazon S3 buckets. Which AWS services can be used to achieve this? (Choose TWO.)

Select 2 answers
A.IAM Access Analyzer
B.AWS Config
C.VPC Flow Logs
D.AWS CloudTrail
E.Amazon S3 server access logs
AnswersD, E

AWS CloudTrail records API activity in the account, including S3 management and, with data events, object-level calls. It logs the identity, time and source of each request, providing the API audit trail the company requires.

Why this answer

AWS CloudTrail (D) is correct because it records S3 data events such as GetObject, PutObject, and DeleteObject, providing an audit trail of API calls made to S3 buckets when data events are enabled. Amazon S3 server access logs (E) are also correct because they capture detailed records of requests made to a bucket, including the requester, bucket name, request time, request action, response status, and error code, which directly supports auditing API calls. IAM Access Analyzer (A) is not correct here because it identifies resource policies that grant external access, not a log of API calls.

AWS Config (B) tracks resource configuration changes and compliance, not individual S3 API request activity. VPC Flow Logs (C) capture IP traffic metadata for network interfaces, not S3 API-level operations.

716
Multi-Selecthard

A data engineering team is designing a batch processing workflow using AWS Glue. The job reads from an S3 bucket, transforms data, and writes to another S3 bucket. The job runs daily and processes new data incrementally. Which THREE features should they use to optimize performance and cost?

Select 3 answers
A.Convert all input data to Parquet format before processing.
B.Enable Glue job autoscaling.
C.Manually increase the number of DPUs for each run.
D.Use predicate pushdown and column pruning in the script.
E.Enable job bookmarks to process only new data.
AnswersB, D, E

Glue job autoscaling dynamically adds and removes workers based on the stage of the shuffle and partition workload, so the daily incremental job does not over-provision capacity for its peak. This directly cuts DPU-hour cost while meeting the performance requirement for variable-volume runs.

Why this answer

Option B is correct because Glue job autoscaling dynamically adjusts the number of workers (DPUs) based on the workload, which reduces cost during lighter stages and improves throughput during heavier ones for a daily incremental job. Option D is correct because predicate pushdown and column pruning let Glue read only the needed partitions, rows, and columns from Parquet/ORC on S3, cutting I/O, shuffle, and DPU-hours. Option E is correct because Glue job bookmarks persist state so the job processes only new or changed data since the last run, which is exactly the incremental daily pattern described and avoids reprocessing old S3 objects.

Option A is not marked correct because converting all input data to Parquet is a broad storage-format change, not a Glue job performance/cost feature, and it may add conversion overhead. Option C is not marked correct because manually increasing DPUs for every run is static over-provisioning that raises cost rather than optimizing it, unlike autoscaling.

717
Multi-Selecteasy

A company is designing a data ingestion pipeline for real-time IoT sensor data. The data volume peaks at 10,000 messages per second. The pipeline must process messages in order per sensor and persist raw data to Amazon S3 for archival. Which TWO services should be used together to meet these requirements? (Choose TWO.)

Select 2 answers
A.Amazon Kinesis Data Firehose
B.Amazon Kinesis Data Streams
C.AWS Lambda
D.Amazon AppFlow
E.Amazon Simple Queue Service (Amazon SQS)
AnswersA, B

Delivers streaming data to S3.

Why this answer

Amazon Kinesis Data Streams (KDS) provides the ordered, real-time ingestion layer required for per-sensor message ordering, as it partitions data by a partition key (e.g., sensor ID) and guarantees order within a shard. Amazon Kinesis Data Firehose then reliably reads from the KDS stream and delivers the raw data to Amazon S3 for archival, handling buffering, compression, and automatic retries without requiring custom code.

Exam trap

The trap here is that candidates often choose AWS Lambda as a processing step without realizing it is not required for the core ingestion and archival pipeline, and they may overlook that Kinesis Data Streams is needed for ordering while Firehose handles the S3 delivery.

718
MCQmedium

A company stores sensitive data in Amazon S3 and uses AWS KMS customer-managed keys for encryption. The security team wants to monitor and audit all KMS API calls that involve the key, including who used the key and when. They also want to receive alerts if the key is used by an unauthorized principal. Which AWS service should the data engineer use to meet these requirements?

A.Amazon GuardDuty with KMS protection, and AWS Lambda functions to remediate unauthorized access.
B.AWS CloudTrail with KMS key usage logs, and AWS Security Hub for automated alerts on unauthorized access.
C.AWS CloudTrail with KMS key usage logs, and Amazon CloudWatch alarms based on CloudTrail metrics.
D.AWS Config with KMS key configuration rules, and Amazon SNS notifications for noncompliance.
AnswerC

AWS CloudTrail captures all KMS API calls as events, including the identity of the caller, time, and key ID. These events can be delivered to an S3 bucket and also to CloudWatch Logs. By creating metric filters and CloudWatch alarms, you can alert on unauthorized usage patterns. This provides both auditing and alerting, meeting the requirements.

Why this answer

AWS CloudTrail logs all KMS API calls, providing a detailed audit trail. To receive alerts on unauthorized usage, you can create CloudWatch metric filters on CloudTrail logs and set up CloudWatch alarms. This combination meets both the auditing and alerting requirements.

Other services like AWS Config, GuardDuty, or Security Hub do not provide the same level of detailed API call logging and direct alerting for KMS usage.

Exam trap

The trap here is confusing configuration compliance (AWS Config) or threat detection (GuardDuty) with API-level auditing and alerting, which CloudTrail and CloudWatch provide.

719
MCQmedium

A data engineer is designing a pipeline to ingest change data capture (CDC) events from an Amazon RDS for PostgreSQL database into Amazon S3. The CDC events are captured using AWS DMS. The data must be available for querying within 5 minutes of the change. Which approach meets these requirements?

A.Export the database to S3 using pg_dump and then use AWS Glue to load into S3 in Parquet format.
B.Use AWS DMS to replicate data to Amazon Redshift, then unload to S3.
C.Use AWS DMS to replicate data directly to S3 in near real-time.
D.Use AWS DMS to replicate data to an SQS queue, then process with Lambda to write to S3.
AnswerC

AWS DMS continuous replication writes CDC changes straight to Amazon S3, avoiding intermediate staging that would add latency. This satisfies the five-minute availability constraint because DMS applies ongoing changes as they occur, rather than relying on scheduled batch exports or hourly snapshots. Target endpoint configuration handles the S3 delivery directly.

Why this answer

AWS DMS supports continuous replication (change data capture) directly to Amazon S3 as a target endpoint. DMS can write CDC events to S3 in near real-time (typically seconds to minutes), meeting the 5-minute latency requirement without intermediate services. The data is stored in comma-separated value (CSV) or Parquet format, ready for querying via Athena or Glue.

Exam trap

The trap here is that candidates may overcomplicate the solution by introducing intermediate services (like Redshift or SQS) when DMS’s native S3 target endpoint already provides near-real-time CDC replication, and they may confuse pg_dump (a batch export tool) with a CDC mechanism.

How to eliminate wrong answers

Option A is wrong because pg_dump is a one-time logical backup tool, not a CDC solution; it cannot capture ongoing changes and would require full exports, exceeding the 5-minute latency window. Option B is wrong because replicating to Redshift then unloading to S3 adds unnecessary complexity and latency; Redshift is optimized for analytics, not as a CDC staging area, and the unload operation introduces additional delay. Option D is wrong because DMS does not natively support SQS as a target endpoint; DMS can write to S3, Kinesis, or other targets, but not directly to SQS, and the proposed architecture would require custom integration, breaking the near-real-time requirement.

720
MCQhard

A company uses AWS Lake Formation to manage access to data in a data lake. A new data engineer has been granted SELECT permission on a table but receives an 'AccessDeniedException' when querying via Amazon Athena. The table is registered in Lake Formation and the data is encrypted with SSE-KMS. Which of the following is the MOST likely cause?

A.The table's resource-based policy does not include the engineer's IAM role.
B.The S3 bucket policy denies access to the engineer's IAM role.
C.The AWS Glue Data Catalog has not been granted permission to the engineer's role.
D.The IAM role used by Athena does not have kms:Decrypt permission on the KMS key.
AnswerD

Lake Formation grants table-level SELECT, but Athena still requires underlying AWS Lake Formation and KMS permissions for encrypted data. Without kms:Decrypt on the SSE-KMS key, Athena cannot read the encrypted objects, producing AccessDeniedException. The stem specifies SSE-KMS encryption, so the missing KMS decrypt permission is the most likely cause.

Why this answer

The IAM role used by Athena does not have kms:Decrypt permission on the KMS key. When data is encrypted with SSE-KMS, Athena's IAM role requires kms:Decrypt to read the data from S3. Even if Lake Formation grants SELECT permission, the query fails without KMS access.

Option A is incorrect because Lake Formation does not use resource-based policies on tables; it uses LF-Tags or resource links to grant permissions. Option B is incorrect because while an S3 bucket policy could block access, the most likely issue with encrypted data is missing KMS permissions. Option C is incorrect because the Glue Data Catalog does not enforce data access; Lake Formation is the service that manages fine-grained access.

721
MCQmedium

A company uses AWS Glue DataBrew to clean and normalize data from an Amazon S3 bucket before loading it into Amazon Redshift. The data contains PII such as social security numbers. A compliance policy requires that PII be masked in the DataBrew output. Which DataBrew transformation should the engineer use to replace the last four digits of each SSN with asterisks?

A.Use the 'Encrypt' transformation with AWS KMS to encrypt the SSN column.
B.Use the 'Remove columns' transformation to drop the SSN column entirely.
C.Use the 'Mask value' transformation with a custom pattern to mask the last four characters.
D.Use the 'Replace value or pattern' transformation to replace the entire SSN with a fixed string.
AnswerC

DataBrew's 'Mask value' transformation supports custom patterns and can mask specific characters. By defining a pattern that targets the last four digits, the engineer can replace them with asterisks while preserving the rest of the SSN. This meets the compliance requirement precisely and is a built-in DataBrew feature.

Why this answer

AWS Glue DataBrew provides a 'Mask value' transformation that can apply custom masking patterns to specific characters. By configuring a pattern to mask the last four digits of the SSN with asterisks, the engineer preserves the first five digits for analytics while complying with the PII masking policy. Other transformations either over-mask, encrypt, or remove the data, failing the specific requirement.

Exam trap

The trap here is assuming that encryption is a form of masking; masking replaces characters with symbols, while encryption transforms data into ciphertext.

722
MCQeasy

An IAM policy includes the above resource ARN for CloudWatch Logs. A data engineer needs to allow a Lambda function to create log streams and put logs to the log group 'my-log-group'. However, the Lambda function is failing with access denied. What is the issue?

A.The region in the ARN does not match the Lambda function's region.
B.The Lambda function does not have an execution role.
C.The ARN does not include the log-stream portion.
D.The ARN is incorrectly formatted because of the wildcard.
AnswerC

CloudWatch Logs actions such as CreateLogStream and PutLogEvents require the ARN to include the log-stream wildcard suffix (log-group:my-log-group:log-stream:*). The policy's ARN stops at the log-group level, so the Lambda function lacks permission for stream-level operations.

Why this answer

The ARN provided in the policy is for the log group itself (arn:aws:logs:region:account-id:log-group:my-log-group), but CloudWatch Logs requires separate permissions for the log-stream subresource. To allow a Lambda function to create log streams and put logs, the ARN must include the log-stream portion, typically with a wildcard like arn:aws:logs:region:account-id:log-group:my-log-group:log-stream:*. Without this, the IAM policy does not grant access to the log-stream actions (CreateLogStream, PutLogEvents), causing an access denied error even if the log group ARN is correct.

Exam trap

The trap here is that candidates assume granting access to the log group ARN implicitly covers log streams, but AWS IAM requires explicit resource-level permissions for each subresource in the ARN hierarchy.

How to eliminate wrong answers

Option A is wrong because the region in the ARN must match the Lambda function's region for the policy to apply, but the question does not indicate a region mismatch; the error is specifically about log-stream permissions. Option B is wrong because a Lambda function must have an execution role to access AWS resources, and the question implies a role exists (the policy is attached), but the role lacks the necessary log-stream permissions. Option D is wrong because the wildcard in the ARN is correctly placed for the log group name (my-log-group) and is not the cause of the failure; the issue is the missing log-stream component, not the wildcard format.

723
MCQeasy

A data engineer needs to store large volumes of semi-structured JSON data in Amazon S3 and query it with Amazon Athena. The engineer wants to minimize query costs and improve performance. Which action should be taken?

A.Increase the number of partitions to one per day.
B.Store the data in a single large JSON file.
C.Use gzip compression on each JSON file.
D.Convert the data to columnar format like Apache Parquet and partition it.
AnswerD

Converting to a columnar format such as Parquet allows Athena to read only the columns referenced in the query, reducing data scanned and thus cost. Partitioning further limits the amount of data scanned by filtering on partition keys. This combination significantly improves performance and reduces query costs.

Why this answer

Using a columnar format like Parquet with partitioning allows Athena to scan only the necessary columns and partitions, drastically reducing the amount of data scanned and thus the cost. This is the recommended best practice for optimizing Athena queries on S3 data.

Exam trap

The trap here is focusing only on compression or partitioning while overlooking the benefit of columnar storage for column pruning.

724
Multi-Selecthard

A data engineer is troubleshooting slow query performance on an Amazon Redshift cluster. The cluster has 10 nodes and is using automatic distribution style. The engineer suspects that data distribution is causing excessive data movement. Which steps should the engineer take to diagnose and resolve the issue? (Choose THREE.)

Select 3 answers
A.Choose appropriate distribution keys for large tables
B.Use the EXPLAIN command to analyze query plans
C.Run the VACUUM command to reclaim space
D.Query the STL_DIST and STL_BCAST system tables
E.Increase the number of nodes in the cluster
AnswersA, B, D

Proper distribution keys minimize data movement.

Why this answer

Choosing appropriate distribution keys for large tables ensures that data is evenly distributed across the cluster slices, minimizing the need for data redistribution during joins and aggregations. Automatic distribution style may not always select the optimal key, leading to excessive data movement and slow query performance.

Exam trap

The trap here is that candidates often confuse VACUUM (which only reorganizes data within slices) with distribution optimization, or assume scaling out nodes automatically resolves distribution-related performance issues without addressing the underlying key choice.

725
MCQeasy

A company uses Amazon S3 to store log files from multiple applications. The logs are written in JSON format. A data engineer wants to use Amazon Athena to query these logs. The logs are stored in a bucket with the following structure: 's3://logs/app1/date=2021-01-01/'. The engineer creates an Athena table with partitions. However, when querying, Athena returns zero results for partitions that exist. The engineer has run MSCK REPAIR TABLE to add partitions. What is the most likely cause of the issue?

A.The MSCK REPAIR TABLE command failed silently.
B.The partition key name in the table definition does not match the S3 folder naming convention.
C.The log files are in JSON format and Athena does not support JSON.
D.The log files need to be copied to a different bucket in the same region.
AnswerB

This is the correct answer. The S3 folder structure uses 'date=' as the partition key prefix, so the Athena table must define a partition key named 'date' exactly. If it is named differently, MSCK REPAIR will not register those folders as partitions.

Why this answer

The most likely cause is that the partition key name in the Athena table definition does not match the S3 folder naming convention. When using MSCK REPAIR TABLE, Athena relies on the partition folder structure (e.g., 'date=2021-01-01') to automatically add partitions. If the table's partition key is named differently (e.g., 'dt' instead of 'date'), MSCK REPAIR will not recognize the folders and will not register the partitions, resulting in zero results.

Option A is incorrect because MSCK REPAIR does not fail silently; it either adds partitions or reports none if the structure doesn't match. Option C is incorrect because Athena fully supports JSON format. Option D is incorrect because the bucket location does not affect partition registration; data can be queried in any bucket as long as the table points to it.

726
Multi-Selecthard

A data engineer is building a governed data lake in AWS Lake Formation. The security team wants to detect sensitive data such as credit card numbers in newly registered S3 tables and automatically apply column-level access restrictions to those columns. Which TWO actions should the engineer take to meet these requirements? (Choose two.)

Select 2 answers
A.Create an EventBridge rule that invokes an AWS Lambda function to apply Lake Formation column-level grants or LF-Tags based on the Macie findings.
B.Enable AWS CloudTrail data events on the S3 bucket and use them to identify objects containing credit card numbers.
C.Use Amazon Macie to run sensitive data discovery jobs against the S3 bucket and publish findings to Amazon EventBridge.
D.Configure an AWS Glue crawler with a custom classifier to detect credit card number patterns and tag the columns in the Data Catalog.
E.Configure S3 Block Public Access on the bucket and require SSE-KMS encryption for all objects.
AnswersA, C

An EventBridge rule matching Macie findings can trigger a Lambda function that calls Lake Formation APIs to restrict columns or attach LF-Tags. This automates the response so newly detected sensitive columns receive restrictive permissions without manual review, satisfying the automatic restriction requirement.

Why this answer

Amazon Macie performs sensitive data discovery on S3 objects and emits findings that can flow through EventBridge, providing the detection layer. A Lambda function triggered by those findings can then apply Lake Formation column-level grants or LF-Tags, automating enforcement so sensitive columns are restricted as data is registered and classified.

Exam trap

The trap here is assuming CloudTrail or Glue classifiers detect sensitive values in object contents, when content inspection belongs to Macie.

727
MCQhard

An application uses the 'orders' DynamoDB table with the schema and provisioned throughput shown in the exhibit. The application frequently queries by customer_id (range key) without specifying the order_id (partition key). What is the most likely impact on performance?

A.Queries will require a full table scan, consuming significant read capacity.
B.Queries will be throttled because the table does not have a global secondary index.
C.Queries will be fast because the sort key is indexed.
D.Queries will cause hot partitions on the table.
AnswerA

DynamoDB queries must target a single partition key value; omitting order_id forces a Scan across every partition, reading all items and consuming large amounts of provisioned read capacity. The range key alone cannot route the request, so performance degrades proportionally to table size.

Why this answer

The application queries by customer_id (the sort key) without specifying order_id (the partition key). In DynamoDB, a Query operation requires the partition key to be specified; without it, the only way to retrieve items is a full table Scan, which reads every item in the table. This consumes read capacity proportional to the entire table size, leading to high latency and cost.

Exam trap

The trap here is that candidates assume the sort key alone can be used for efficient queries, forgetting that DynamoDB's indexing requires the partition key to be specified for a Query operation.

How to eliminate wrong answers

Option B is wrong because throttling is not caused by the absence of a GSI; throttling occurs when consumed capacity exceeds provisioned throughput, and a Scan can cause throttling indirectly by consuming high capacity, but the lack of a GSI itself does not throttle queries. Option C is wrong because the sort key (customer_id) is only indexed within the context of a specific partition key (order_id); without the partition key, the sort key index cannot be used for efficient lookup. Option D is wrong because hot partitions are caused by uneven access patterns on a single partition key, not by queries that omit the partition key; a Scan reads all partitions evenly, so it does not create hot spots.

728
MCQmedium

A company is using Amazon RDS for MySQL with Multi-AZ deployment. The database size is 2 TB and the workload is read-heavy. To improve read performance, which option should be used?

A.Use Amazon ElastiCache to cache database queries
B.Increase the instance size to 16xlarge
C.Create Read Replicas in the same or different regions
D.Enable Multi-AZ on additional instances
AnswerC

Read Replicas asynchronously copy the primary via MySQL's binlog, and applications can direct read-only queries to them, offloading the read-heavy workload from the primary. Multi-AZ standby exists solely for failover and cannot serve reads, so it adds no read throughput.

Why this answer

Amazon RDS for MySQL Read Replicas offload read traffic from the primary DB instance, directly improving read performance for a read-heavy workload. With a 2 TB database, Read Replicas can be created in the same or different regions, providing horizontal read scaling without impacting the primary instance's write capacity. Multi-AZ deployment already provides high availability but does not improve read performance, as the standby instance is not used for reads.

Exam trap

The trap here is that candidates often confuse Multi-AZ with read scaling, assuming the standby instance can serve reads, but Multi-AZ is strictly for high availability and disaster recovery, not for read traffic.

How to eliminate wrong answers

Option A is wrong because Amazon ElastiCache caches query results, reducing database load, but it does not directly improve read performance for all queries, especially those that cannot be cached (e.g., write-heavy or dynamic queries), and it introduces cache invalidation complexity. Option B is wrong because increasing the instance size to 16xlarge vertically scales the database, which can improve performance, but it is less cost-effective and does not provide the same read scalability as horizontal scaling with Read Replicas; it also does not leverage the Multi-AZ deployment's standby for reads. Option D is wrong because enabling Multi-AZ on additional instances is not a valid configuration; Multi-AZ is a feature of the primary instance that provisions a standby in a different Availability Zone for failover, not for read scaling, and additional instances cannot have Multi-AZ enabled independently.

729
MCQeasy

A data engineer needs to grant an IAM user read-only access to an S3 bucket named 'data-lake'. Which IAM policy statement should be used?

A.{"Effect":"Allow","Action":["s3:PutObject","s3:DeleteObject"],"Resource":"arn:aws:s3:::data-lake/*"}
B.{"Effect":"Allow","Action":"s3:*","Resource":"*"}
C.{"Effect":"Allow","Action":["s3:ListBucket","s3:GetObject"],"Resource":["arn:aws:s3:::data-lake","arn:aws:s3:::data-lake/*"]}
D.{"Effect":"Allow","Action":"s3:ListBucket","Resource":"arn:aws:s3:::data-lake"}
AnswerC

The policy pairs bucket-level `s3:ListBucket` with object-level `s3:GetObject`, mapping each action to the correct ARN: `arn:aws:s3:::data-lake` for the bucket itself and `arn:aws:s3:::data-lake/*` for its objects. This satisfies the read-only requirement, since listing needs bucket scope while retrieval needs object scope.

Why this answer

It allows ListBucket on the bucket and GetObject on objects, enabling read-only access. Option A is wrong because it grants write actions (PutObject, DeleteObject). Option B is wrong because it allows all S3 actions (s3:*).

Option D is wrong because it only allows ListBucket, not GetObject.

730
MCQhard

A data engineer is using AWS Glue to transform data from an Amazon S3 bucket. The Glue job reads JSON files, applies complex transformations, and writes the output to another S3 bucket in Parquet format. The job runs daily and must complete within a 2-hour window. The engineer notices that the job is taking longer than expected and wants to optimize performance. The source data is partitioned by date, and the job uses a dynamic frame. Which optimization should the engineer implement to improve performance?

A.Increase the number of DPUs allocated to the Glue job.
B.Enable job bookmarks to track processed data and avoid reprocessing.
C.Convert the dynamic frame to a Spark DataFrame and use Spark SQL for transformations.
D.Use pushdown predicates to filter data at the source based on partition columns.
AnswerD

Pushdown predicates allow Glue to filter data at the source, reducing the amount of data read and processed. Since the data is partitioned by date, the engineer can specify a predicate to read only the relevant partitions. This minimizes I/O and speeds up the job, often more effectively than adding resources. It directly addresses the performance bottleneck by reducing data volume.

Why this answer

Pushdown predicates enable AWS Glue to filter data at the source based on partition columns, significantly reducing the amount of data read from S3. Since the data is partitioned by date, specifying a predicate for the desired date range ensures only relevant partitions are processed, leading to faster job completion and lower cost. This is a targeted optimization that addresses the performance bottleneck without unnecessary resource scaling.

Exam trap

The trap here is assuming that adding more DPUs is always the best way to speed up a Glue job, when in fact reducing the amount of data read via pushdown predicates is often more effective and cost-efficient.

731
MCQmedium

A data pipeline uses AWS Glue to process data from Amazon S3 and write results to Amazon Redshift. The pipeline fails intermittently with the error 'S3ServiceException: Access Denied'. The IAM role used by Glue has permissions to read from the S3 bucket. What is the most likely cause of this error?

A.The S3 bucket is in a different AWS Region than the Glue job
B.S3 Server Access Logging is enabled and blocking requests
C.The S3 bucket policy denies access to the Glue job's IAM role
D.The S3 bucket has S3 Transfer Acceleration enabled
AnswerC

An explicit Deny in the S3 bucket policy overrides the IAM role's Allow permissions. Even though the Glue role can read S3, the bucket policy denying that role causes the intermittent Access Denied error described in the stem.

Why this answer

S3 bucket policies can explicitly deny access even if the IAM role allows it. Option A is wrong because S3 bucket region does not cause access denied errors. Option B is wrong because S3 Server Access Logging does not affect access permissions.

Option D is wrong because S3 Transfer Acceleration is not related to access denied errors.

732
MCQmedium

Refer to the exhibit. A data engineer is troubleshooting an AWS Glue ETL job that fails with an access denied error when writing to S3. The IAM role attached to the Glue job has the policy shown. What is the most likely cause of the error?

A.The Glue job is missing the s3:ListBucket permission on the bucket
B.The Glue job is writing to an S3 bucket that is not included in the Resource ARN
C.The S3 bucket is encrypted with AWS KMS and the policy does not include kms:Decrypt permissions
D.The Glue job does not have permission to call glue:StartJobRun
AnswerB

S3 access is granted per bucket ARN in the IAM policy. If the Glue job's target bucket falls outside the Resource ARN, the write is denied even though the role itself is valid. The error is authorisation scope, not credentials or networking.

Why this answer

The IAM policy shown in the exhibit explicitly restricts the `s3:PutObject` action to a specific ARN pattern (e.g., `arn:aws:s3:::example-bucket/*`). If the Glue job attempts to write to a different S3 bucket—one not matching that ARN—the request will be denied with an access denied error. AWS Glue ETL jobs use the attached IAM role's permissions, and a mismatch between the target bucket and the resource ARN in the policy is the most common cause of such failures.

Exam trap

The trap here is that candidates often overlook the explicit resource ARN in the policy and assume the error is due to missing permissions like `s3:ListBucket` or KMS, when the real issue is a bucket name mismatch in the resource specification.

How to eliminate wrong answers

Option A is wrong because the `s3:ListBucket` permission is required for listing objects (e.g., `s3:ListBucket` on the bucket resource), not for writing objects; the error occurs during a write operation, and the policy already includes `s3:PutObject` on the bucket objects. Option C is wrong because the error message is specifically 'access denied when writing to S3', not a KMS-related error; if KMS encryption were the issue, the error would typically mention 'KMS' or 'decrypt' and the policy would need `kms:GenerateDataKey` or `kms:Decrypt`, not just `kms:Decrypt`. Option D is wrong because `glue:StartJobRun` is a permission to start a Glue job run, not a permission for writing to S3; the error occurs during the job execution (writing to S3), not during job initiation.

733
MCQhard

A data engineer is using AWS Glue to catalog data stored in Amazon S3. The engineer needs to run an AWS Glue ETL job that reads from a large dataset in Parquet format, performs transformations, and writes the output to Amazon Redshift. The job must handle data skew and optimize performance. Which AWS Glue feature should the engineer use to address data skew during the join operation?

A.Configure the job to use the `spark.sql.adaptive.enabled` and `spark.sql.adaptive.skewJoin.enabled` properties.
B.Use the AWS Glue DynamicFrame and apply the `Relationalize` transform to flatten nested data.
C.Enable job bookmarks to track processed data and avoid reprocessing.
D.Increase the number of DPUs allocated to the job to provide more resources for handling skewed data.
AnswerA

Enabling adaptive query execution (AQE) with skew join optimization allows Spark to dynamically handle data skew during joins. It detects skewed partitions and splits them into smaller sub-partitions, balancing the workload. This directly addresses the skew issue and improves job performance. These properties are supported in AWS Glue jobs running Spark 3.0 and later.

Why this answer

Data skew in joins can be mitigated by enabling Spark's adaptive query execution (AQE) with skew join optimization. This feature dynamically detects and splits skewed partitions, balancing the workload across executors. Other options do not directly address skew: job bookmarks track processed data, Relationalize flattens nested data, and adding DPUs increases capacity without redistributing data.

Exam trap

The trap here is thinking that adding more resources (DPUs) will automatically resolve data skew, when in fact skew requires specific algorithmic handling.

734
MCQmedium

A data engineer is troubleshooting a slow-running query on Amazon Redshift. The query scans a large table but returns few rows. Which diagnostic step should be taken first?

A.Use EXPLAIN to review the query plan.
B.Run ANALYZE on the table.
C.Check the concurrency scaling status.
D.Run VACUUM on the table.
AnswerA

EXPLAIN reveals the query plan, exposing sequential scans, missing sort keys, or distribution distortions causing the large scan. Reviewing it first identifies the actual bottleneck before altering the query or cluster, satisfying the stem's requirement to diagnose the slow query.

Why this answer

When a query scans a large table but returns few rows, the most likely cause is an inefficient query plan—such as a full table scan instead of using indexes or zone maps. Using EXPLAIN first reveals the execution plan, allowing the engineer to identify whether the query is performing unnecessary sequential scans, missing filter pushdown, or using suboptimal join strategies. This diagnostic step should always precede tuning actions like ANALYZE or VACUUM, which address data distribution or storage bloat rather than query planning.

Exam trap

The trap here is that candidates often jump to performance-tuning commands like ANALYZE or VACUUM without first diagnosing the query plan, but the DEA-C01 exam emphasizes that EXPLAIN is the foundational step for identifying inefficient scan patterns before applying any corrective actions.

How to eliminate wrong answers

Option B is wrong because ANALYZE updates table statistics for the query optimizer, but if the query plan is already suboptimal (e.g., missing a WHERE clause filter), fresh statistics won't fix the root cause—EXPLAIN must be checked first. Option C is wrong because concurrency scaling handles increased query load by adding cluster capacity, but it does not improve the efficiency of a single slow query that scans many rows unnecessarily. Option D is wrong because VACUUM reclaims disk space and sorts rows for better compression, but it does not change the query execution path—a full table scan will remain a full table scan even after a vacuum.

735
Multi-Selecthard

A company is migrating a large Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime and preserve data consistency. Which THREE AWS services or features should be used?

Select 3 answers
A.Amazon RDS for Oracle as the target
B.AWS DataSync for initial load
C.AWS Schema Conversion Tool (SCT) for schema conversion
D.Amazon Aurora PostgreSQL as the target database
E.AWS Database Migration Service (DMS) for continuous replication
AnswersC, D, E

AWS Schema Conversion Tool converts Oracle PL/SQL, sequences and data types into Aurora PostgreSQL-compatible DDL, satisfying the heterogeneous-engine schema transformation the migration demands. Without it, stored procedures and proprietary constructs would fail on PostgreSQL, blocking the minimal-downtime cutover that AWS DMS alone cannot address.

Why this answer

Option C (AWS Schema Conversion Tool) is correct because migrating from Oracle to Aurora PostgreSQL requires converting the Oracle schema, stored procedures, and PL/SQL objects into PostgreSQL-compatible DDL, which SCT automates. Option D (Amazon Aurora PostgreSQL as the target database) is correct because the scenario explicitly states the migration target is Aurora PostgreSQL, so the target engine must be Aurora PostgreSQL. Option E (AWS DMS for continuous replication) is correct because DMS supports ongoing change data capture (CDC) from Oracle to Aurora PostgreSQL, enabling minimal downtime by replicating changes until cutover.

Option A is incorrect because RDS for Oracle is the source engine, not the target, and would not accomplish the migration to PostgreSQL. Option B is incorrect because AWS DataSync is a file/object transfer service for data like NFS/SMB/S3, not a database migration tool, and cannot preserve relational data consistency for an Oracle-to-PostgreSQL migration.

Exam trap

The trap here is that candidates often confuse AWS DataSync (a file-transfer service) with database migration tools, or mistakenly think RDS for Oracle can serve as a migration target when the question explicitly specifies Aurora PostgreSQL.

736
MCQmedium

A data engineer manages an AWS Glue ETL job that reads JSON files from Amazon S3, transforms the data, and writes to an Amazon Redshift table. The job recently started failing with the error 'Communication link failure: connection reset'. The Redshift cluster is healthy and the IAM role used by the Glue job has the necessary permissions. The engineer notices that the job runs longer than before and the Redshift cluster's WLM queue is often full. Which action should the engineer take to resolve the failure?

A.Modify the Redshift WLM configuration to increase concurrency or add a new queue for the Glue job.
B.Configure the Glue job to use the Redshift JDBC driver with a longer connection timeout and retry logic.
C.Increase the number of AWS Glue DPUs allocated to the job to speed up processing.
D.Switch the Glue job to write to Amazon S3 first, then use a Redshift COPY command to load the data.
AnswerA

The error 'Communication link failure: connection reset' often occurs when Redshift terminates connections due to WLM queue timeouts. Increasing concurrency or adding a dedicated queue for the Glue job reduces wait times and prevents connection resets. This directly addresses the root cause of the failure, allowing the job to complete successfully without changing Glue resources.

Why this answer

The failure is caused by Redshift's workload management (WLM) queue being full, which leads to connection timeouts and resets when the AWS Glue job attempts to write data. Adjusting the WLM configuration to increase concurrency or create a dedicated queue for the Glue job reduces wait times and prevents the connection resets. This directly addresses the root cause without unnecessary changes to the Glue job or architecture.

Exam trap

The trap here is assuming that connection reset errors are always network-related and can be fixed by increasing timeouts or retries, when in fact they often stem from Redshift workload management limits.

737
MCQmedium

A data engineer must transform data in an AWS Glue ETL job. The transform requires calling an external REST API for each record to enrich the data. The Glue job runs on AWS Glue 4.0 with Python. The engineer wants to minimize the number of API calls and improve performance. Which approach should the engineer take?

A.Convert the DynamicFrame to a Spark DataFrame, use foreachPartition to batch records and call the API once per batch.
B.Use Glue DynamicFrame's map method to call the API for each record.
C.Use AWS Glue's built-in transform 'ApplyMapping' to call the API.
D.Configure the Glue job to use a larger number of DPUs to parallelize API calls.
AnswerA

Converting to a Spark DataFrame and using foreachPartition allows processing records in batches per partition. The engineer can accumulate records into batches and make a single API call per batch, reducing the number of calls. This leverages Spark's distributed processing and is efficient for external API enrichment. It minimizes API calls and improves performance by batching.

Why this answer

To minimize API calls in an AWS Glue job, the engineer should batch records and call the API once per batch. Converting the DynamicFrame to a Spark DataFrame and using foreachPartition allows processing each partition, accumulating records into batches, and making a single API call per batch. This reduces the number of calls and improves performance.

Other options either call the API per record or do not support API calls.

Exam trap

The trap here is assuming that increasing DPUs or using map will automatically optimize API calls, but neither reduces the number of calls per record.

738
MCQmedium

A company stores sensitive data in Amazon Redshift. The security team requires that all data in the cluster be encrypted at rest using a customer managed key in AWS KMS, and that the key be rotated annually. The data engineer needs to configure the Redshift cluster accordingly. Which action should the engineer take?

A.Use Redshift Spectrum to query data in S3 that is encrypted with SSE-KMS, and enable key rotation on the S3 bucket's default key.
B.Enable Redshift cluster encryption using AWS KMS and set the cluster parameter require_ssl to true.
C.Enable Redshift cluster encryption using AWS KMS and create a scheduled AWS Lambda function that calls kms:RotateKey on the key annually.
D.Enable Redshift cluster encryption using AWS KMS and select the customer managed key. Configure automatic key rotation for the KMS key with a 365-day rotation period.
AnswerD

Amazon Redshift supports encryption at rest with AWS KMS customer managed keys. When creating or modifying a cluster, you can choose a customer managed key. Separately, in AWS KMS, you can enable automatic rotation for that key with a custom period, such as 365 days. This meets both the encryption and rotation requirements. The key rotation is a property of the KMS key, not the Redshift cluster.

Why this answer

To encrypt a Redshift cluster at rest with a customer managed KMS key, you enable encryption on the cluster and select the key. To rotate the key annually, you enable automatic rotation on that KMS key with a 365-day period. These are separate configurations but together satisfy the requirements.

Exam trap

The trap here is confusing encryption in transit (SSL) with encryption at rest, or assuming that key rotation is configured on the Redshift cluster rather than in KMS.

739
MCQeasy

A data engineer is using Amazon Redshift to store sales data. The engineer needs to ensure that the data is encrypted at rest and that encryption keys are managed by AWS. The engineer also wants to minimize administrative overhead. Which Redshift encryption option should the engineer use?

A.AWS Key Management Service (AWS KMS) customer managed key
B.AWS KMS AWS managed key (aws/redshift)
C.Client-side encryption with a custom key
D.Hardware security module (HSM) encryption
AnswerB

The AWS managed key for Redshift is automatically created and managed by AWS, providing encryption at rest with minimal administrative effort. You do not need to manage key rotation or policies; AWS handles these tasks. This meets the requirement for encryption at rest while minimizing overhead, making it the ideal choice for this scenario.

Why this answer

Amazon Redshift supports encryption at rest using AWS KMS. The AWS managed key (aws/redshift) is automatically created and managed by AWS, providing encryption without the need for manual key management. This minimizes administrative overhead while ensuring data is encrypted at rest.

Customer managed keys offer more control but require more management, and HSM or client-side encryption add complexity.

Exam trap

The trap here is equating 'managed by AWS' with 'customer managed key', which actually requires more administrative effort.

740
MCQeasy

A company stores sensitive data in Amazon S3 and needs to ensure that data is encrypted at rest. Which AWS service can be used to manage the encryption keys?

A.AWS Key Management Service (KMS)
B.AWS Secrets Manager
C.AWS Identity and Access Management (IAM)
D.AWS Certificate Manager (ACM)
AnswerA

AWS Key Management Service creates, rotates and controls the customer managed keys used for server-side encryption of S3 objects, including SSE-KMS. It satisfies the requirement to manage encryption keys centrally while keeping data encrypted at rest in Amazon S3.

Why this answer

AWS Key Management Service (KMS) is the managed service for creating and controlling encryption keys used to encrypt data at rest in Amazon S3. Option A is correct. Option B (Secrets Manager) is for managing secrets like database passwords, not encryption keys.

Option C (IAM) manages access permissions, not encryption keys. Option D (Certificate Manager) handles SSL/TLS certificates, not encryption keys for data at rest.

741
MCQmedium

A company is using Amazon RDS for PostgreSQL with Multi-AZ deployment. The primary instance fails and a failover occurs. After the failover, the application cannot connect to the database. What is the MOST likely cause?

A.The database instance is in a 'stopped' state after failover.
B.The Multi-AZ failover requires manual intervention to complete.
C.The security group for the RDS instance was not updated during failover.
D.The application is using the old primary instance endpoint instead of the RDS CNAME.
AnswerD

RDS failover repoints the cluster's CNAME to the standby, which becomes the new primary. Applications hard-coded to the former primary's instance endpoint keep hitting a host that is no longer writable, so connections fail until they use the DNS name that RDS updates automatically.

Why this answer

After a Multi-AZ failover in Amazon RDS for PostgreSQL, the DNS CNAME record automatically updates to point to the new primary instance in the standby Availability Zone. If the application hardcodes the old primary instance's endpoint (specific IP or DNS name) instead of using the RDS CNAME (which remains constant), it will attempt to connect to the failed instance, causing connectivity loss. The CNAME is the stable connection point that always resolves to the current primary instance.

Exam trap

The trap here is that candidates may assume security groups or instance state are the issue, but AWS explicitly tests the concept that the RDS CNAME is the correct connection target and that hardcoding endpoints leads to failover failures.

How to eliminate wrong answers

Option A is wrong because RDS Multi-AZ failover does not stop the database instance; the new primary is promoted and remains in an 'available' state. Option B is wrong because Multi-AZ failover is fully automated and requires no manual intervention to complete. Option C is wrong because security groups are associated with the RDS instance itself, not with a specific AZ or IP, and they remain unchanged during failover; the new primary inherits the same security group configuration.

742
Multi-Selectmedium

A data engineer is designing a data ingestion pipeline to load clickstream data from an Amazon S3 bucket into an Amazon Redshift cluster. The data arrives in 5-minute batches. Which TWO actions should the engineer take to ensure data consistency and avoid duplicates? (Select TWO.)

Select 2 answers
A.Define a SORTKEY on the target table to improve deduplication.
B.Disable workload management (WLM) to maximize resources.
C.Use the STL_LOAD_ERRORS system table to monitor and resolve load errors.
D.Load data into a single slice to maintain order.
E.Use a staging table and perform a MERGE operation to avoid duplicates.
AnswersC, E

STL_LOAD_ERRORS records every row Redshift rejected during COPY, including the raw error reason and the S3 source file. Reviewing it lets the engineer correct malformed clickstream records and reload them, preventing silent data loss that would otherwise break consistency across the 5-minute batches.

Why this answer

Option E is correct because loading each 5-minute batch into a staging table and then running a MERGE (or DELETE+INSERT) against the target table lets you match on a business key and update or insert rows idempotently, which prevents duplicate records when batches are retried or overlap. Option C is correct because monitoring STL_LOAD_ERRORS (and STL_LOAD_ERRORS is populated per load with err_code, err_reason, and raw_field_value) lets the engineer detect and resolve failed or partially loaded rows so the pipeline maintains consistent data. Option A is not correct because a SORTKEY only affects physical ordering and query performance; it does not deduplicate rows.

Option B is not correct because disabling WLM removes query queue management and does not address duplicates or consistency. Option D is not correct because Redshift distributes data across slices automatically, and forcing data into a single slice would hurt performance without providing deduplication.

Exam trap

DEA-C01 often tests the confusion between performance tuning (SORTKEY, distribution keys) and data consistency mechanisms (staging tables, MERGE). Candidates may think SORTKEY deduplicates, but it does not. Also, disabling WLM is never a correct answer for consistency.

743
MCQmedium

A company uses AWS Glue to process sensitive customer data stored in S3. The security team requires that all data be encrypted at rest using a customer-managed KMS key and that access to the key be auditable. Which solution meets these requirements?

A.Encrypt the data client-side before uploading to S3.
B.Configure the S3 bucket to use SSE-KMS with a customer-managed KMS key and enable CloudTrail for KMS events.
C.Enable default SSE-S3 encryption on the S3 bucket.
D.Use SSE-C with a customer-provided key.
AnswerB

SSE-KMS with a customer-managed key encrypts S3 objects at rest under your own key policy, satisfying the customer-managed requirement. CloudTrail captures every KMS API call, including Decrypt and GenerateDataKey, delivering the auditable key-access trail the security team demands. SSE-S3 cannot provide either control.

Why this answer

SSE-KMS with a customer-managed KMS key provides encryption at rest and, when combined with CloudTrail logging for KMS events, offers full auditability of key usage. Option A (client-side encryption) does not encrypt data at rest within S3; it encrypts before upload. Option C (SSE-S3) uses AWS-managed keys, which do not allow customer audit of key access.

Option D (SSE-C) relies on customer-provided keys that are not managed by KMS and cannot be audited via CloudTrail.

744
MCQhard

A company has a data lake in Amazon S3 with millions of objects. The security team wants to enforce that all objects are encrypted with a specific customer-managed KMS key. The data engineer configures an S3 bucket policy to deny PutObject if the encryption is not set to that key. However, some existing objects are not encrypted with that key. What is the most efficient way to remediate the existing objects?

A.Use S3 Cross-Region Replication to replicate objects to a new bucket with the correct encryption.
B.Write a script using the AWS SDK to iterate over all objects and re-upload them with the correct encryption.
C.Use S3 Batch Operations to copy objects in the same bucket with the new encryption settings.
D.Use S3 Object Lambda to dynamically encrypt objects on read.
AnswerC

S3 Batch Operations applies a single job across millions of objects, performing a copy-in-place that rewrites each object under the specified KMS key. This remediates existing objects at scale without custom scripts or per-object API calls, meeting the requirement that all objects use the customer-managed key.

Why this answer

S3 Batch Operations is designed to perform bulk actions (like copying objects) across millions of objects efficiently, and it supports specifying new encryption settings during the copy. By copying objects in the same bucket with the desired KMS key, the existing objects are re-encrypted with the customer-managed key, satisfying the bucket policy requirement. This is the most efficient, server-side, managed approach for large-scale remediation.

Exam trap

The trap is choosing a custom SDK script or replication — but S3 Batch Operations is the managed, scalable solution for bulk object remediation, and replication does not fix encryption on existing objects in place.

How to eliminate wrong answers

Option A is wrong because Cross-Region Replication replicates objects to a different bucket/region and does not remediate encryption on existing objects in the current bucket; it also incurs cross-region data transfer costs and does not change the original objects. Option B is wrong because writing a custom SDK script to iterate and re-upload millions of objects is inefficient, error-prone, and requires managing concurrency and throttling — S3 Batch Operations is purpose-built for this. Option D is wrong because S3 Object Lambda transforms data on read for specific use cases (e.g., redaction) and does not change the stored encryption of existing objects.

745
MCQmedium

A company is storing sensitive user data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using a customer-managed key stored in AWS KMS. The bucket policy must deny any PUT request that does not include the appropriate encryption header. Which bucket policy condition key should be used?

A.s3:x-amz-server-side-encryption-aws-kms-key-id
B.s3:x-amz-server-side-encryption
C.s3:x-amz-acl
D.aws:SourceArn
AnswerA

The condition key `s3:x-amz-server-side-encryption-aws-kms-key-id` validates the exact KMS key ID supplied in the `x-amz-server-side-encryption-aws-kms-key-id` header, satisfying the requirement that objects use a customer-managed key. Denying PUT requests lacking this header enforces encryption at rest with the mandated key, not merely any SSE-KMS key.

Why this answer

The `s3:x-amz-server-side-encryption-aws-kms-key-id` condition key allows the bucket policy to enforce that PUT requests include a specific customer-managed KMS key ID in the `x-amz-server-side-encryption-aws-kms-key-id` header, ensuring encryption at rest with the required key. This directly meets the security team's requirement to deny PUT requests that lack the appropriate encryption header tied to a customer-managed KMS key.

Exam trap

The trap here is that candidates often confuse `s3:x-amz-server-side-encryption` (which only checks the encryption algorithm, not the key) with `s3:x-amz-server-side-encryption-aws-kms-key-id` (which checks the specific KMS key ID), leading them to pick option B when the requirement explicitly demands a customer-managed key.

How to eliminate wrong answers

Option B is wrong because `s3:x-amz-server-side-encryption` only checks whether the `x-amz-server-side-encryption` header is present (e.g., `AES256` or `aws:kms`), but it cannot enforce that a specific customer-managed KMS key ID is used; it would allow any KMS key, including AWS-managed keys. Option C is wrong because `s3:x-amz-acl` is used to control access control list (ACL) headers in requests, not encryption headers, so it is irrelevant to encryption enforcement. Option D is wrong because `aws:SourceArn` is a global condition key used to restrict requests based on the ARN of the source resource (e.g., an SNS topic or Lambda function), not to enforce encryption headers in S3 PUT requests.

746
Multi-Selectmedium

A data engineer is managing an Amazon Kinesis Data Firehose delivery stream that writes to an Amazon S3 bucket. The engineer notices that some records are being delivered to the S3 bucket with a prefix of 'errors/' instead of the intended 'data/' prefix. The Firehose stream is configured with an AWS Lambda function for data transformation, and the S3 bucket has a lifecycle policy that transitions objects to Glacier after 30 days. Which TWO actions should the engineer take to ensure that only successfully transformed records are delivered to the 'data/' prefix and that failed records are handled appropriately? (Choose two.)

Select 2 answers
A.Specify an error prefix in the Firehose stream's S3 destination configuration so that failed records are delivered to a separate prefix.
B.Enable the 'S3 backup' setting in the Firehose stream to back up all incoming records to a separate S3 bucket before transformation.
C.Increase the Firehose stream's buffer size and buffer interval to reduce the number of failed records.
D.Configure the Lambda function to return the transformed record with a 'result' status of 'Ok' for successful records and 'Dropped' or 'ProcessingFailed' for failed records.
E.Configure the Firehose stream's S3 destination to use a dynamic prefix based on the record's result status, such as 'data/' for success and 'errors/' for failure.
AnswersA, D

The S3 destination configuration for Kinesis Data Firehose includes an optional error prefix. When a record fails transformation (Lambda returns 'ProcessingFailed') or delivery, Firehose writes that record to the specified error prefix instead of the main prefix. By setting the error prefix to 'errors/', failed records are segregated, and only successfully transformed records (with 'Ok' status) are delivered to the 'data/' prefix. This is the correct way to handle failed records.

Why this answer

To ensure only successful records go to the 'data/' prefix, the Lambda function must return a 'result' of 'Ok' for successful transformations and 'Dropped' or 'ProcessingFailed' for failures. Additionally, configuring an error prefix in the Firehose S3 destination directs failed records to a separate 'errors/' prefix. Together, these actions segregate successful and failed records correctly.

Exam trap

The trap here is assuming that Firehose can dynamically route records based on status, when in fact the error prefix is a static configuration and the Lambda result status determines delivery.

747
MCQmedium

A data engineer is troubleshooting a data pipeline that uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The engineer notices that the S3 bucket contains many small files (less than 1 MB). This is causing performance issues in downstream processing. What is the BEST way to reduce the number of small files?

A.Increase the buffer size to at least 128 MB in the Firehose delivery stream configuration.
B.Use an AWS Lambda function to transform the data before delivery.
C.Change the compression format from GZIP to Snappy.
D.Decrease the buffer interval in the Firehose delivery stream configuration.
AnswerA

Firehose buffers incoming records and only delivers once the buffer size or interval is reached. Raising the buffer size to 128 MB or more forces larger, less frequent S3 objects, directly reducing the sub-1 MB file count causing downstream performance issues.

Why this answer

Increasing the Firehose buffer size to at least 128 MB allows more data to accumulate before delivery, resulting in fewer, larger files in S3. This directly addresses the small file problem by batching more records per delivery.

Exam trap

DEA-C01 often tests the confusion between buffer size and buffer interval, where candidates might think decreasing interval reduces file count, when it actually increases it.

How to eliminate wrong answers

Option B is wrong because using Lambda to transform data does not inherently reduce file count; it may even increase processing overhead without addressing buffering. Option C is wrong because changing compression format affects file size but not the number of files; small files would still be created. Option D is wrong because decreasing the buffer interval causes more frequent deliveries, resulting in even more small files.

748
Multi-Selectmedium

A company is designing a data lake on Amazon S3 for analytics. The data includes sensitive personally identifiable information (PII). Which TWO actions should the company take to protect the data? (Choose TWO.)

Select 2 answers
A.Enable S3 Block Public Access.
B.Enable Requester Pays.
C.Enable S3 Transfer Acceleration.
D.Enable cross-region replication.
E.Enable default encryption with SSE-KMS.
AnswersA, E

S3 Block Public Access overrides bucket policies and ACLs that would otherwise expose objects publicly, preventing accidental PII leakage. This satisfies the protection requirement by enforcing account- and bucket-level guardrails against public exposure of the data lake.

Why this answer

S3 Block Public Access (Option A) prevents any public access to S3 buckets and objects, which is critical for protecting PII from unintended exposure. Default encryption with SSE-KMS (Option E) ensures that all data written to S3 is encrypted at rest using AWS KMS-managed keys, providing both encryption and centralized key management for sensitive data.

Exam trap

The trap here is that candidates often confuse operational features like Requester Pays or Transfer Acceleration with security controls, or they think replication alone provides data protection, when in fact encryption and access blocking are the direct mechanisms for safeguarding PII.

749
MCQmedium

A data engineer runs the above CLI command to describe the DynamoDB table 'Orders'. The table has a partition key 'OrderID' and sort key 'CustomerID'. Which query operation is most efficient for retrieving all orders for a specific customer?

A.Query the table using CustomerID as the partition key
B.Scan the table and filter by CustomerID
C.Use GetItem with CustomerID as the key
D.Create a Global Secondary Index on CustomerID and query the index
AnswerD

Queries must target the partition key, so retrieving orders by CustomerID needs a Global Secondary Index with CustomerID as its partition key. A Scan reads the whole table; a Query on OrderID cannot filter efficiently by customer.

Why this answer

A Global Secondary Index (GSI) on CustomerID allows you to query efficiently using CustomerID as the partition key, avoiding a full table scan. Since the base table's primary key is (OrderID, CustomerID), you cannot directly query by CustomerID alone; a GSI provides an alternative access pattern optimized for this query.

Exam trap

The trap here is that candidates assume the sort key can be used as a query filter without an index, but DynamoDB requires the partition key for Query operations, and a Scan is often mistakenly chosen as a simpler alternative despite its performance cost.

How to eliminate wrong answers

Option A is wrong because CustomerID is the sort key, not the partition key, so a Query operation requires the partition key (OrderID) to be specified; you cannot query using only the sort key. Option B is wrong because a Scan reads every item in the table, which is inefficient and costly for large datasets, especially when a targeted query is possible. Option C is wrong because GetItem requires both the partition key and sort key to retrieve a single item; it cannot return multiple orders for a customer.

750
MCQhard

A company uses AWS Lake Formation to manage data lake permissions. The data engineer notices that a user with SELECT permission on a table can also query the underlying data in Amazon S3 directly. How can the engineer enforce that access to the S3 data is only through Lake Formation?

A.Use S3 Access Points with a policy that restricts access to only Lake Formation
B.Grant the user permissions only through Lake Formation and remove any IAM policies that allow direct S3 access to the data location
C.Enable S3 Block Public Access on the bucket
D.Change the S3 bucket policy to deny all access except from Lake Formation
AnswerB

Lake Formation permissions govern only requests routed through Lake Formation; direct S3 GETs are authorised by IAM. Removing IAM policies that permit s3:GetObject on the data location closes that bypass, so the user's only viable path to the data is via Lake Formation's credential vending.

Why this answer

Lake Formation permissions are enforced only when the caller accesses data through a Lake Formation-integrated engine (Athena, Redshift Spectrum, EMR, Glue). If the user also has IAM permissions on the underlying S3 path, they can bypass Lake Formation entirely by reading S3 directly. The fix is to remove those direct S3 IAM permissions so the only path to the data is through Lake Formation-governed services.

Exam trap

The trap is assuming that granting Lake Formation permissions automatically overrides or supersedes IAM — in reality IAM and Lake Formation are additive, and any direct S3 IAM allow defeats LF governance.

How to eliminate wrong answers

Option A is wrong because S3 Access Points do not by themselves prevent a user with existing IAM s3:GetObject permissions on the bucket from reading the data; they add an access path, not a restriction. Option C is wrong because S3 Block Public Access only blocks public/anonymous access and does nothing about an authenticated IAM principal with explicit allow. Option D is wrong because a bucket policy denying all except Lake Formation would also break other legitimate access and is not how Lake Formation enforcement is designed — Lake Formation uses its own permission model layered on top of IAM, not a blanket bucket-policy deny.

Page 9

Page 10 of 18

Page 11