Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 1–75

1321 questions total · 18pages · All types, answers revealed

Page 1 of 18

Page 2
1
MCQeasy

A data engineer needs to capture change data capture (CDC) events from an Amazon RDS for PostgreSQL database and stream them to Amazon S3 in near real-time. Which AWS service should be used?

A.Amazon S3 Transfer Acceleration
B.Amazon Athena
C.AWS Database Migration Service (AWS DMS)
D.Amazon Kinesis Data Streams
AnswerC

DMS supports ongoing replication (CDC) from databases to S3.

Why this answer

AWS DMS supports continuous replication from PostgreSQL source databases using logical replication slots to capture CDC events in near real-time. It can directly stream these changes to Amazon S3 as a target, making it the correct choice for this use case.

Exam trap

The trap here is that candidates may confuse Kinesis Data Streams as a direct CDC solution, but it requires a separate CDC tool or custom application to capture PostgreSQL changes, whereas AWS DMS natively supports this integration.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Transfer Acceleration is a feature that speeds up uploads to S3 over long distances using edge locations, but it does not capture or stream CDC events from a database. Option B is wrong because Amazon Athena is an interactive query service for analyzing data in S3 using SQL, not a tool for capturing or streaming database changes. Option D is wrong because Amazon Kinesis Data Streams is a real-time streaming service that can ingest data, but it cannot directly capture CDC events from an RDS for PostgreSQL database without additional configuration or a separate CDC connector.

2
MCQeasy

A company stores application logs in Amazon S3 and wants to analyze them using Amazon Athena. The logs are in JSON format and are compressed with gzip. The data engineer needs to create an Athena table that can query these logs efficiently. The logs are stored in an S3 bucket with the prefix logs/year=2023/month=10/day=15/. The engineer wants to minimize query costs and improve performance. Which action should the engineer take?

A.Create an external table with the JSON SerDe, specifying the S3 location as s3://bucket/logs/ and using the OpenCSVSerDe for better performance.
B.Create an external table with the CSV SerDe, specifying the S3 location as s3://bucket/logs/ and defining partitions for year, month, and day.
C.Create an external table with the JSON SerDe, specifying the S3 location as s3://bucket/logs/year=2023/month=10/day=15/ and no partitions.
D.Create an external table with the JSON SerDe, specifying the S3 location as s3://bucket/logs/ and defining partitions for year, month, and day.
AnswerD

Creating an external table with the JSON SerDe and defining partitions allows Athena to read the JSON logs and use partition pruning to scan only relevant data. Specifying the root S3 location and adding partitions (either manually or via MSCK REPAIR) enables efficient queries and reduces cost by limiting data scanned.

Why this answer

To query JSON logs efficiently in Athena, the table must use the JSON SerDe to parse the data correctly. Defining partitions on year, month, and day allows Athena to prune partitions based on query filters, reducing the amount of data scanned and lowering costs. The table location should point to the root prefix so that all partitions are included.

This combination provides accurate parsing and cost-effective queries.

Exam trap

The trap here is selecting a SerDe based on compression or performance assumptions rather than matching it to the actual data format, which is JSON.

3
MCQhard

A company uses AWS Database Migration Service (DMS) to continuously replicate data from an on-premises Oracle database to Amazon S3 in Parquet format. The replication is used for near-real-time analytics. Recently, the DMS task started failing with an error indicating insufficient memory. The source database is large (2 TB). What should a data engineer do to resolve this issue while minimizing changes to the existing architecture?

A.Change the target format to JSON to reduce memory usage.
B.Split the DMS task into multiple smaller tasks.
C.Use Change Data Capture (CDC) only, without full load.
D.Increase the DMS replication instance size.
AnswerD

DMS memory exhaustion during large-table replication is resolved by scaling the replication instance, adding RAM and CPU. This satisfies the 2 TB source constraint while preserving the existing task, endpoints and Parquet target architecture unchanged.

Why this answer

The error indicates the DMS replication instance is running out of memory during continuous replication of a 2 TB Oracle database to S3 in Parquet format. Increasing the replication instance size (Option D) directly addresses the memory constraint by providing more RAM and processing capacity, which is necessary for handling large volumes of Change Data Capture (CDC) data and Parquet conversion overhead. This solution requires minimal architectural changes, as it only involves modifying the instance class in the DMS task settings.

Exam trap

The trap here is that candidates may think splitting tasks or changing formats reduces memory usage, but the root cause is insufficient instance resources, and AWS DMS tasks require adequate instance sizing for large-scale CDC workloads.

How to eliminate wrong answers

Option A is wrong because changing the target format to JSON would not reduce memory usage; JSON is typically larger than Parquet and would increase memory consumption during serialization, not decrease it. Option B is wrong because splitting the DMS task into multiple smaller tasks would increase complexity and overhead, potentially causing additional memory pressure from multiple connections and task management, and does not directly resolve the insufficient memory error. Option C is wrong because using CDC only without full load ignores the fact that the task is already failing during continuous replication (CDC phase), and the full load may have already completed; disabling full load does not address the memory issue in CDC processing.

4
MCQhard

A data engineer is building an AWS Glue job that reads semi-structured JSON from Amazon S3 and must flatten nested arrays into relational columns before writing to Amazon Redshift. The transformation logic is complex and the engineer wants to unit test it locally without provisioning a cluster. Which Glue capability should the engineer use to develop and test this transformation logic?

A.AWS Glue DataBrew recipe steps executed against a sample dataset.
B.AWS Glue interactive sessions with an AWS Glue Studio notebook.
C.The AWS Glue ETL library (awsglue) run locally with a Python environment and sample datasets.
D.AWS Glue Studio visual job editor with a custom transform node.
AnswerC

The AWS Glue ETL library can be installed and run in a local Python environment, allowing the engineer to execute transformation functions against sample DataFrames without any Glue cluster. This supports true local unit testing of complex flattening logic. It matches the requirement to develop and test without provisioning managed infrastructure, and the same code can later run on Glue with minimal changes.

Why this answer

The AWS Glue ETL library can be installed locally so developers run the same transformation code against sample data without any managed cluster. This enables genuine unit tests of complex nested-array flattening. The visual editor, interactive sessions, and DataBrew all depend on AWS-hosted execution and therefore cannot satisfy the requirement for local testing before deployment.

Exam trap

The trap here is equating AWS Glue Studio notebooks or interactive sessions with local development, when only the installable Glue ETL library actually runs on the developer's machine.

5
Multi-Selectmedium

A data engineer is optimizing an Amazon S3 data lake for cost and performance. The data lake contains large volumes of CSV files that are queried by Amazon Athena. The engineer wants to reduce query costs and improve query performance. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Partition the data by a frequently filtered column, such as date.
B.Enable S3 server access logging on the bucket.
C.Enable S3 Transfer Acceleration on the bucket.
D.Convert the CSV files to Apache Parquet format.
E.Increase the number of Athena workgroups.
AnswersA, D

Partitioning by a commonly filtered column allows Athena to scan only the relevant partitions, reducing the amount of data read. This lowers query cost and improves performance by skipping irrelevant data. It is a best practice for large datasets in S3 when queries filter on that column. Combined with Parquet, it maximizes savings.

Why this answer

Converting CSV to Parquet and partitioning the data by a frequently filtered column are the two most effective actions to reduce Athena query costs and improve performance. Parquet reduces scanned data through columnar compression and predicate pushdown, while partitioning limits the data scanned to relevant partitions. The other actions do not affect query execution or data scanned.

Exam trap

The trap here is confusing S3 bucket-level features like Transfer Acceleration or access logging with query optimization, when the real gains come from data format and partitioning.

6
MCQhard

A company stores sensitive financial data in an Amazon S3 bucket. The data engineering team must ensure that all data is encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys are rotated annually. The team also needs to audit key usage. Which solution meets these requirements?

A.Enable default encryption on the S3 bucket using SSE-S3, and configure S3 server access logging to track key usage.
B.Use client-side encryption with a customer provided key stored in AWS Secrets Manager, and enable AWS CloudTrail to log S3 API calls.
C.Enable default encryption on the S3 bucket using SSE-KMS with an AWS managed key, and configure S3 Inventory to report on encryption status.
D.Enable default encryption on the S3 bucket using SSE-KMS with a customer managed key, enable automatic key rotation for the KMS key, and enable AWS CloudTrail to log KMS API calls.
AnswerD

SSE-KMS with a customer managed key satisfies the encryption requirement. Enabling automatic key rotation on the KMS key rotates the key annually. AWS CloudTrail logs all KMS API calls, providing an audit trail of key usage. This combination meets all specified requirements: encryption with customer managed keys, annual rotation, and auditing.

Why this answer

The correct solution uses SSE-KMS with a customer managed key to encrypt data at rest, enables automatic key rotation for that key to meet the annual rotation requirement, and enables AWS CloudTrail to log KMS API calls for auditing key usage. Other options either use incorrect key types (SSE-S3 or AWS managed keys) or do not provide the necessary auditing capabilities.

Exam trap

The trap here is confusing AWS managed keys with customer managed keys and assuming that S3 server access logging or S3 Inventory can audit KMS key usage.

7
MCQmedium

Refer to the exhibit. This log snippet is from a failed AWS Glue job. The job processes a large dataset in memory. What is the MOST likely cause of the OutOfMemoryError?

A.The Glue job is running with insufficient DPUs or worker type.
B.The input data is in an unsupported file format.
C.The job is attempting to join two tables with mismatched keys.
D.The job has too many partitions.
AnswerA

Glue allocates memory per executor based on DPU count and worker type; a large in-memory dataset exceeding that allocation triggers OutOfMemoryError. Insufficient DPUs or an undersized worker type directly caps heap available to the job, so scaling them addresses the constraint stated in the stem.

Why this answer

An OutOfMemoryError in AWS Glue typically occurs when the allocated DPUs or worker type are insufficient for the in-memory processing of a large dataset. Option B is incorrect because unsupported file formats cause parsing errors, not memory errors. Option C is incorrect because mismatched keys in a join cause data skew or incorrect results, but not directly an OutOfMemoryError.

Option D is incorrect because too many partitions usually lead to small file overhead, not heap space exhaustion.

8
MCQmedium

A company has a large volume of CSV files in S3 that need to be transformed into Parquet using AWS Glue. The files are partitioned by date. The engineer wants to minimize costs by processing only new files each day. Which approach should be used?

A.Use S3 partition discovery to automatically read new partitions.
B.Schedule the Glue job to run daily and process all files.
C.Enable job bookmarks in the Glue job.
D.Use S3 Event Notifications to trigger the Glue job on each new file.
AnswerC

Job bookmarks persist state across runs, tracking which S3 objects and partitions were already processed. This satisfies the requirement to process only new files daily, avoiding full re-scans of historical data and cutting Glue DPU costs.

Why this answer

AWS Glue job bookmarks track the state of data processed by a job, so subsequent runs only process new or changed data. When enabled, bookmarks persist the last processed path or partition, allowing the daily job to skip already-transformed files and reduce runtime and cost. This directly satisfies the requirement to process only new files each day.

Exam trap

DEA-C01 often tests the misconception that S3 partition discovery or event notifications provide incremental processing, when only Glue job bookmarks track processed state to avoid reprocessing.

How to eliminate wrong answers

Option A is wrong because S3 partition discovery reads partitions but does not track which files were already processed, so the job would reprocess old data. Option B is wrong because processing all files daily increases cost and runtime, contradicting the goal of minimizing costs. Option D is wrong because S3 Event Notifications trigger the job per file, which can cause excessive job runs and does not inherently track processed state; it also adds complexity and potential duplicate processing.

9
MCQmedium

A data engineer needs to ingest data from an on-premises Oracle database to Amazon S3 daily. The data volume is 500 GB per day, and the network bandwidth is 200 Mbps. The requirement is to minimize the impact on the source database and ensure data integrity. Which combination of AWS services should be used?

A.AWS Database Migration Service (DMS) with S3 as target
B.AWS Glue ETL jobs with JDBC connection
C.Amazon Kinesis Data Firehose with Oracle as source
D.AWS Data Pipeline with SQLActivity
AnswerA

AWS DMS minimizes source impact by using change data capture and supports S3 as a target.

Why this answer

AWS DMS with S3 as target is correct because it supports continuous change data capture (CDC) from Oracle, minimizing impact on the source database by reading redo logs instead of querying tables directly. It can handle 500 GB/day over 200 Mbps (which yields ~2.16 TB/day theoretical max) and ensures data integrity via transactional consistency and validation checksums. DMS also automatically partitions large datasets and can resume from failures, making it ideal for daily bulk loads.

Exam trap

The trap here is that candidates assume AWS Glue or Data Pipeline are suitable for database ingestion, but they lack the CDC and low-impact features of DMS, which is specifically designed for minimal source database load during large-scale migrations or replication.

How to eliminate wrong answers

Option B is wrong because AWS Glue ETL jobs with JDBC connection would pull data via full table scans, placing significant load on the Oracle database and lacking native CDC capabilities, which violates the requirement to minimize source impact. Option C is wrong because Amazon Kinesis Data Firehose cannot use Oracle as a direct source; it ingests from streaming sources like Kinesis Data Streams, not relational databases via JDBC. Option D is wrong because AWS Data Pipeline with SQLActivity uses a polling-based approach that repeatedly queries the source, causing unnecessary overhead, and does not support CDC or optimized large-volume transfers like DMS.

10
MCQmedium

A company uses Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The data is JSON and each record is about 2 KB. The delivery stream is configured to buffer incoming data to 5 MB or 60 seconds, whichever comes first. The data engineering team notices that the S3 bucket contains many small files (average 2 MB), which makes subsequent processing inefficient. They need to reduce the number of small files without increasing the latency beyond 5 minutes. Which solution should they implement?

A.Enable compression (GZIP) on the delivery stream.
B.Increase the buffer size to 50 MB and the buffer interval to 300 seconds.
C.Use a Lambda function to merge small files after delivery.
D.Decrease the buffer size to 1 MB and the buffer interval to 60 seconds.
AnswerB

Raising the buffer to 50 MB lets Firehose accumulate far more records before each S3 PUT, directly cutting the small-file count. The 300-second interval still satisfies the stem's five-minute latency ceiling, so larger objects arrive without breaching the stated constraint.

Why this answer

Increasing the buffer size to 50 MB and the buffer interval to 300 seconds directly addresses the root cause: the current 5 MB buffer size triggers a flush too frequently, producing many 2 MB files. By raising the buffer size to 50 MB, each flush will contain more data, resulting in larger S3 objects (up to ~50 MB uncompressed), while the 300-second interval ensures latency stays within the 5-minute requirement. This reduces the number of small files without requiring additional services or post-processing.

Exam trap

The trap here is that candidates often assume compression (Option A) reduces file count, but compression only reduces file size, not the number of files; the real issue is the flush frequency controlled by buffer size and interval.

How to eliminate wrong answers

Option A is wrong because enabling GZIP compression reduces the storage size of each file but does not change the buffer size or flush behavior; the delivery stream will still flush at 5 MB or 60 seconds, producing the same number of small files (just compressed). Option C is wrong because using a Lambda function to merge small files after delivery adds complexity, cost, and latency (Lambda invocation delays, S3 event processing), and does not address the root cause of premature flushes; it also violates the goal of not increasing latency beyond 5 minutes due to the merging overhead. Option D is wrong because decreasing the buffer size to 1 MB and keeping the buffer interval at 60 seconds would make the problem worse, producing even smaller files (average ~1 MB) and more frequent flushes, increasing the number of small files.

11
MCQmedium

A data engineer is building an Amazon Redshift data warehouse. The cluster will ingest data from Amazon S3 using the COPY command. The engineer needs to ensure that the data is loaded in a way that maximizes query performance for future complex analytical queries. The data is currently stored as uncompressed CSV files in S3. Which action should the engineer take to optimize the load and subsequent query performance?

A.Use the COPY command with the GZIP option to load compressed files, which automatically distributes the data evenly across all nodes.
B.Load the data into a table with no compression and then run a VACUUM command to reorganize and compress the data.
C.Load the data into a single large staging table and then use CREATE TABLE AS (CTAS) to redistribute it into multiple tables.
D.Use the COPY command with the COMPUPDATE ON option to automatically apply compression, and define an appropriate distribution key and sort key on the target table.
AnswerD

Using COPY with COMPUPDATE ON allows Redshift to analyze the data and apply optimal compression encodings automatically, which reduces storage and improves I/O performance. Defining a distribution key ensures even data distribution across nodes, and a sort key improves query performance by enabling efficient range scans and minimizing data movement during joins. This combination directly addresses the need for optimized load and query performance.

Why this answer

The correct approach is to use COPY with COMPUPDATE ON to automatically apply compression, and to define an appropriate distribution key and sort key. This optimizes storage, reduces I/O, and improves query performance by ensuring even data distribution and efficient data access patterns. Other options either do not address compression, distribution, or sort keys adequately, or rely on commands that do not achieve the desired optimization.

Exam trap

The trap here is assuming that loading compressed files or running VACUUM will automatically optimize the table's physical design for query performance.

12
Multi-Selecteasy

Which TWO features of Amazon S3 help protect data from accidental deletion or modification? (Choose two.)

Select 2 answers
A.Lifecycle policies
B.Default encryption
C.S3 MFA Delete
D.S3 Object Versioning
E.Cross-Region Replication
AnswersC, D

S3 MFA Delete enforces a second authentication factor on delete operations, requiring the bucket owner's MFA token to permanently remove object versions or change versioning state. This directly satisfies the accidental-deletion constraint by blocking unauthorised or mistaken deletions.

Why this answer

C is correct because S3 MFA Delete requires multi-factor authentication for permanent deletion of objects or suspension of versioning, adding a critical layer of protection against accidental or malicious deletions. D is correct because S3 Object Versioning preserves every version of an object, allowing recovery from accidental overwrites or deletions by restoring a previous version.

Exam trap

The trap here is that candidates often confuse data protection features (like encryption or replication) with deletion prevention, mistakenly selecting Lifecycle policies or Cross-Region Replication because they think 'protection' includes backup or security, when the question specifically asks about preventing accidental deletion or modification.

13
MCQmedium

A data engineer manages an Amazon S3 data lake that ingests millions of small JSON files daily from an IoT fleet. Query performance in Amazon Athena has degraded significantly, and each query scans far more data than expected. The engineer wants to reduce per-query cost and improve performance without changing the raw data. Which solution should the engineer implement?

A.Convert the JSON files to Apache Parquet, partition the data, and compact small files with AWS Glue ETL jobs.
B.Move the data into an Amazon Redshift cluster and query it with Redshift Spectrum.
C.Increase the Athena workgroup data usage control limit and rerun the queries.
D.Enable S3 Transfer Acceleration on the ingestion bucket.
AnswerA

Athena charges by bytes scanned, so columnar Parquet with compression and partition pruning dramatically reduces data read. Compacting millions of tiny files into larger row groups removes per-object overhead and improves parallel scan throughput. Partitioning by a common filter column such as device date or region lets Athena skip irrelevant prefixes entirely, cutting both latency and cost.

Why this answer

Athena performance and cost are driven by how many bytes a query scans, so converting row-oriented JSON to columnar Parquet, partitioning on common filter columns, and compacting small files directly reduce scanned data. These three changes work together: columnar storage reads only referenced columns, partitions prune entire prefixes, and compaction removes per-object overhead. The result is faster queries at lower cost without altering the source data in S3.

Exam trap

The trap here is assuming that a networking or workgroup setting such as S3 Transfer Acceleration or a higher data usage limit can fix slow Athena queries, when the real driver is bytes scanned and file layout.

14
MCQeasy

A data engineer is monitoring an Amazon Redshift cluster and notices that the disk space usage is increasing rapidly. The engineer wants to reclaim space from deleted rows. Which command should the engineer run?

A.VACUUM
B.ANALYZE
C.UNLOAD
D.COPY
AnswerA

VACUUM reclaims storage occupied by deleted rows and re-sorts unsorted data in Amazon Redshift, directly addressing rapidly growing disk usage. It is the specific maintenance command that returns freed blocks to the system for reuse.

Why this answer

VACUUM in Amazon Redshift reclaims disk space occupied by rows marked for deletion and re-sorts rows to restore sort order. When rows are deleted, they are only logically marked until VACUUM runs, so disk usage keeps growing. Running VACUUM (or VACUUM DELETE ONLY) releases that space back to the system.

Exam trap

DEA-C01 often tests the distinction between VACUUM (space reclamation and re-sort) and ANALYZE (statistics refresh) — candidates who conflate maintenance commands pick ANALYZE, which does nothing for deleted-row disk usage.

How to eliminate wrong answers

Option B is wrong because ANALYZE only updates table statistics used by the query planner — it does not reclaim any disk space from deleted rows. Option C is wrong because UNLOAD exports query results from Redshift to Amazon S3; it is an outbound data operation, not a space-reclamation command. Option D is wrong because COPY loads data into Redshift from S3, DynamoDB, or other sources — it adds data rather than reclaiming deleted space.

15
Multi-Selectmedium

A data engineer is building an AWS Glue ETL job that reads data from Amazon S3 and must write the output partitioned by year, month, and day for efficient downstream querying in Amazon Athena. The engineer wants the job to create the partition folders and register them in the Glue Data Catalog automatically. (Choose two.)

Select 2 answers
A.Call the Glue catalog's create_partition API for each partition after the job writes files.
B.Set the DynamicFrame write option partitionKeys to the year, month, and day columns.
C.Enable the job's 'Update the Data Catalog' option so new partitions are added during the write.
D.Run an AWS Glue crawler on the output prefix after every job run.
E.Store the output as a single Parquet file per run to preserve partition metadata.
AnswersB, C

The partitionKeys write option tells Glue how to split output files into directory hierarchies based on the specified columns. Setting it to year, month, and day produces the year=/month=/day= folder structure that Athena and other engines expect. This is the primary mechanism for generating partitioned output from a Glue job.

Why this answer

Producing partitioned output and registering it in the Glue Data Catalog from a Glue job is achieved by specifying partitionKeys on the write and enabling the job to update the Data Catalog. Together these create the year/month/day folder structure and add the corresponding catalog partitions during the run. Manual API calls or a separate crawler can work but are not the automatic in-job mechanism the scenario requires.

Exam trap

The trap here is assuming a post-run crawler is the only way to register partitions, when the Glue write itself can update the catalog if configured to do so.

16
Multi-Selecthard

A company is ingesting Apache logs from multiple web servers into AWS. The logs are sent via Amazon CloudWatch Logs to a subscription filter that delivers to a Lambda function. The Lambda function parses the logs and writes to Amazon S3. However, there is a significant backlog. Which THREE actions can reduce the backlog?

Select 3 answers
A.Route the CloudWatch Logs subscription to an Amazon SQS queue first
B.Increase the Lambda function memory allocation
C.Increase the Lambda function reserved concurrency
D.Change the Lambda function runtime from Python to Node.js
E.Increase the Lambda function maximum concurrency (unreserved account concurrency)
AnswersB, C, E

More memory also increases CPU, speeding up processing.

Why this answer

Increasing the Lambda function's memory allocation also increases its CPU allocation, allowing the function to process each log event faster. This reduces the per-invocation processing time, enabling the function to handle more log data per unit time and thus reduce the backlog.

Exam trap

The trap here is that candidates may confuse 'reserved concurrency' with 'maximum concurrency' or think that adding an SQS queue always improves throughput, when in fact it can add latency and does not address the root cause of slow per-invocation processing.

17
MCQmedium

A data engineer is loading data from Amazon S3 into an Amazon Redshift cluster using the COPY command. The S3 bucket contains 500 Parquet files, each about 200 MB, in a single prefix. The COPY job is running slowly and consuming excessive cluster resources. The engineer wants to improve performance without changing the data format or the cluster size. Which action should the engineer take?

A.Use the COPY command with the PARALLEL OFF option to reduce the number of slices used.
B.Split the data into multiple prefixes and run multiple COPY commands in parallel, or use a manifest file to distribute the load across slices.
C.Convert the Parquet files to CSV and load them with the CSV option to enable faster parsing.
D.Add the COMPUPDATE OFF and STATUPDATE OFF options to the COPY command to skip compression and statistics updates.
AnswerB

Redshift COPY parallelizes across slices by reading multiple files concurrently. When many files are in one prefix, the leader node can become a bottleneck and slice distribution may be uneven. Splitting into multiple prefixes or using a manifest file lets COPY distribute files more evenly across slices, improving throughput and reducing resource contention without changing the data format or cluster size.

Why this answer

Redshift COPY achieves high throughput by having multiple slices read files in parallel. When hundreds of files sit in one prefix, the leader node may serialize listing and assignment, and some slices may read more files than others, causing skew and resource contention. Distributing files across prefixes or using a manifest gives COPY finer-grained control over parallel reads, balancing work across slices and improving load speed without altering the format or cluster size.

Exam trap

The trap here is assuming that COPY always parallelizes perfectly regardless of file layout, when in fact a single prefix with many files can create a leader-node and slice-distribution bottleneck.

18
MCQmedium

A data engineer is configuring an AWS Glue job that reads from an Amazon RDS for MySQL database and writes to Amazon S3. The security team requires that the data be encrypted in transit between AWS Glue and Amazon RDS. Which action should the engineer take to meet this requirement?

A.Attach an IAM policy to the Glue job role that allows rds:DescribeDBInstances and rds:Connect, which enforces SSL.
B.Configure the AWS Glue connection to use SSL by setting the JDBC URL with the sslMode=REQUIRED parameter and providing the RDS root certificate.
C.Use AWS Glue's built-in encryption context to encrypt the data before writing to RDS.
D.Enable encryption at rest on the RDS instance, which automatically encrypts data in transit as well.
AnswerB

For JDBC connections to RDS for MySQL, encryption in transit is enabled by setting sslMode=REQUIRED in the JDBC URL and providing the RDS CA certificate for validation. This ensures that the connection between AWS Glue and RDS is encrypted using SSL/TLS. This is the standard method to enforce encryption in transit for Glue JDBC connections.

Why this answer

To encrypt data in transit between AWS Glue and Amazon RDS for MySQL, the JDBC connection must be configured to use SSL. Setting sslMode=REQUIRED in the JDBC URL and providing the RDS root certificate ensures that the connection is encrypted and the server certificate is validated. This is the correct approach for meeting the encryption in transit requirement.

Exam trap

The trap here is confusing encryption at rest with encryption in transit, or assuming that IAM policies can enforce SSL for database connections.

19
MCQmedium

A data engineer manages an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. Both buckets are encrypted with SSE-KMS using customer managed keys. The Glue job execution role has permissions to read from the source bucket and write to the target bucket, and has kms:Decrypt permission on the source key, but the job fails with an error indicating it cannot write to the target bucket due to encryption. What is the MOST likely cause?

A.The Glue job execution role lacks kms:GenerateDataKey permission on the target bucket's KMS key.
B.The Glue job execution role lacks s3:PutObjectAcl permission on the target bucket.
C.The target S3 bucket's default encryption uses SSE-S3 instead of SSE-KMS, causing a conflict.
D.The Glue job execution role lacks kms:Decrypt permission on the target bucket's KMS key.
AnswerA

When writing to an SSE-KMS encrypted S3 bucket, the writer must call kms:GenerateDataKey to obtain a data key for encryption. Without this permission, the write fails. The role already has read access and decrypt permission on the source key, but the target key requires GenerateDataKey to encrypt the data.

Why this answer

Writing to an SSE-KMS encrypted S3 bucket requires the kms:GenerateDataKey permission on the KMS key used for the target bucket. The Glue job role already has read and decrypt permissions on the source, but lacks the necessary permission to encrypt data for the target. Without GenerateDataKey, the S3 PutObject operation fails.

Exam trap

The trap here is assuming that kms:Decrypt is sufficient for both reading and writing encrypted data, when in fact writing requires kms:GenerateDataKey.

20
MCQhard

A data engineer notices that an S3 bucket policy allows access to a user from another AWS account, but the access is being denied. What could be the reason?

A.The bucket policy does not include KMS permissions
B.The other account's IAM user does not have permissions to access the bucket
C.S3 does not support cross-account access
D.The bucket is in a different region
AnswerB

Cross-account S3 access requires both the bucket policy to allow the principal and the principal's own IAM identity policy to permit the S3 action. The bucket policy alone is insufficient, so the missing identity-based permission causes the denial.

Why this answer

For cross-account S3 access to succeed, both the bucket policy (resource-based policy) and the IAM user policy (identity-based policy) in the other account must grant the necessary permissions. Option B is correct because even if the bucket policy allows access from the other account, the IAM user in that account must have an explicit IAM policy that permits the S3 action (e.g., s3:GetObject) on the bucket. Without this, the request is denied by the other account's own IAM evaluation.

Exam trap

The trap here is that candidates assume a bucket policy alone is sufficient for cross-account access, forgetting that the requesting account's IAM user must also have explicit permissions, which is a classic AWS cross-account authorization nuance.

How to eliminate wrong answers

Option A is wrong because KMS permissions are only required if the bucket uses SSE-KMS encryption; the question does not mention encryption, and a missing KMS permission would cause a different error (e.g., AccessDenied with KMS key context). Option C is wrong because S3 fully supports cross-account access via bucket policies and ACLs, as documented in the AWS S3 User Guide. Option D is wrong because S3 is a global service and cross-region access is allowed; region does not inherently block cross-account access.

21
MCQeasy

A data engineer is using Amazon Athena to query data stored in Amazon S3. The engineer wants to reduce query costs and improve performance for a table that is frequently queried with filters on a date column. The data is stored as uncompressed CSV files partitioned by year/month/day. Which action should the engineer take?

A.Convert the data to Parquet format and use partitioning by date, then run MSCK REPAIR TABLE.
B.Use Amazon Redshift Spectrum to query the S3 data instead of Athena.
C.Increase the number of S3 buckets storing the data to parallelize reads.
D.Enable Athena query result reuse and set a short cache TTL.
AnswerA

Converting to Parquet reduces the amount of data scanned because Athena reads only the necessary columns, and partitioning by date allows Athena to skip irrelevant partitions. Running MSCK REPAIR TABLE updates the partition metadata so queries can leverage the partitions, significantly lowering cost and improving performance.

Why this answer

For Athena, the most effective way to reduce cost and improve performance is to use a columnar format like Parquet and partition the data. Parquet allows column pruning, and partitioning enables partition pruning, so queries scan less data. Updating partitions with MSCK REPAIR TABLE ensures the metadata is current, allowing these optimizations to take effect.

Exam trap

The trap here is focusing on caching or infrastructure changes instead of the fundamental storage format and partitioning strategy that directly impact data scanned.

22
MCQhard

A company has a multi-account AWS environment with a centralized data lake in the Security account. Data producers in other accounts use AWS Glue to write data to S3 buckets in the Security account. The Security account uses AWS Lake Formation to manage permissions. The data engineer is setting up cross-account access so that users in the Producer account can query the data using Athena in their own account. The engineer has registered the S3 buckets and Data Catalog tables in Lake Formation. The IAM roles in the Producer account have the necessary permissions. However, when a user in the Producer account tries to query the table, they get an AccessDenied error. The error message indicates that the principal is not authorized to perform lakeformation:GetTable on the resource. What is the most likely cause?

A.The Glue Data Catalog resource policy is missing a statement to allow cross-account access.
B.The S3 bucket policy does not allow the Producer account's IAM role to read the data.
C.The KMS key policy does not allow the Producer account's IAM role to decrypt objects.
D.The Lake Formation permissions in the Security account do not include a grant to the Producer account's IAM role.
AnswerD

Cross-account Lake Formation access requires two grants: the resource owner grants permissions to the external principal, and the recipient account grants its role access. The Security account never granted the Producer role, so lakeformation:GetTable fails despite correct IAM permissions.

Why this answer

The error explicitly states the principal is not authorized to perform lakeformation:GetTable, which is a Lake Formation permission check, not an S3, KMS, or Glue resource policy check. In Lake Formation cross-account access, the data owner (Security account) must grant table and data location permissions directly to the external IAM principal (the Producer account's role) using the Lake Formation GrantPermissions API or console. Without this explicit grant, Lake Formation denies the GetTable call before any S3 or KMS access is even attempted.

Since the IAM roles already have the necessary permissions, the missing piece is the Lake Formation grant in the Security account.

Exam trap

DEA-C01 often tests the misconception that IAM policies alone are sufficient for cross-account Lake Formation access, when in fact Lake Formation requires an explicit grant to the external principal, and the error message's mention of lakeformation:GetTable is the key clue that distinguishes it from S3 or KMS issues.

How to eliminate wrong answers

Option A is wrong because the Glue Data Catalog resource policy is not the authorization mechanism when Lake Formation manages the catalog; Lake Formation intercepts and authorizes Glue API calls (including GetTable) via its own permission model, so a Glue resource policy would not resolve a lakeformation:GetTable denial. Option B is wrong because the error occurs at the Lake Formation authorization stage, before any S3 data access is attempted; an S3 bucket policy issue would produce an S3 AccessDenied on GetObject, not a lakeformation:GetTable error. Option C is wrong because KMS decryption happens only after Lake Formation authorizes the table access and returns temporary credentials; a KMS key policy problem would surface as a KMS AccessDenied or decryption failure during data retrieval, not during the GetTable call.

23
MCQeasy

A company wants to audit all data access events in their S3 buckets, including who accessed objects and from which IP address. Which AWS service should be used to capture these events?

A.AWS CloudTrail with data events enabled
B.Amazon CloudWatch Logs
C.AWS Config
D.Amazon S3 Server Access Logs
AnswerA

CloudTrail data events record object-level S3 operations, capturing the requester's identity and source IP address for every GetObject and PutObject call. Management events alone omit these object reads, so enabling data events satisfies the audit requirement for who accessed which object and from where.

Why this answer

AWS CloudTrail with data events enabled captures object-level S3 operations (GetObject, PutObject, DeleteObject) along with the identity of the caller and the source IP address, which is exactly what is needed for a full data access audit. Management events alone do not capture object reads/writes, so data events must be explicitly turned on.

Exam trap

DEA-C01 often tests the difference between CloudTrail management events and data events — candidates who assume default CloudTrail captures object reads miss that data events must be explicitly enabled for S3 object-level auditing.

How to eliminate wrong answers

Option B is wrong because CloudWatch Logs is a log aggregation and monitoring service — it does not natively capture S3 data access events unless CloudTrail or S3 access logs are routed to it. Option C is wrong because AWS Config tracks resource configuration changes and compliance, not individual object access events. Option D is wrong because S3 Server Access Logs capture request details but are delivered on a best-effort basis, do not include IAM identity context the way CloudTrail does, and lack the same fidelity for security auditing.

24
MCQmedium

A data engineering team stores clickstream events in an Amazon S3 bucket under the prefix s3://analytics/raw/. New objects arrive continuously, and the team wants Amazon Athena queries to scan only the events for the current day without scanning the entire prefix. The events are written as JSON files partitioned by year/month/day, but queries still scan all partitions because the partition metadata is not registered. Which action should the data engineer take to enable partition pruning in Athena?

A.Run MSCK REPAIR TABLE on the Athena table to add the year/month/day partitions to the AWS Glue Data Catalog.
B.Create an Amazon CloudFront distribution in front of the S3 bucket and point Athena at the distribution.
C.Enable S3 Transfer Acceleration on the bucket so Athena can retrieve the daily events faster.
D.Convert the JSON files to Apache Parquet and rely on columnar storage alone to limit the scan to one day.
AnswerA

MSCK REPAIR TABLE scans the S3 prefix, discovers Hive-style partition folders such as year=2024/month=03/day=15, and registers each partition in the Data Catalog so Athena can prune non-matching partitions. This directly addresses the missing partition metadata that causes full-prefix scans in this scenario.

Why this answer

Partition pruning requires the query engine to know which partition values exist and where their data lives. Registering the year/month/day folders as partitions in the Data Catalog lets Athena build a plan that reads only the matching day, cutting scanned bytes and cost. Physical optimizations such as acceleration or columnar formats do not supply that partition metadata.

Exam trap

The trap here is assuming that a columnar file format or faster data transfer alone limits how much data Athena scans, when partition pruning depends on registered partition metadata.

25
MCQeasy

A data engineer needs to transform JSON data from Amazon S3 into Parquet format using AWS Glue. The data contains nested fields. Which Glue feature should the engineer use to define the schema and handle the nested structure?

A.Use the 'FindMatches' transform to identify duplicates.
B.Use the 'DropFields' transform to remove nested fields.
C.Use the 'Relationalize' transform in a Glue ETL script.
D.Use the 'Spigot' transform to write sample data.
AnswerC

Relationalize flattens nested JSON structures into separate relational tables linked by keys, letting Glue's DynamicFrame handle arrays and structs that a flat schema cannot represent. This satisfies the requirement to define schema and process nested fields before writing Parquet.

Why this answer

The Relationalize transform in AWS Glue is specifically designed to convert nested JSON data into a relational format. It flattens nested structures by creating separate tables for arrays and nested objects, which can then be written to Parquet. This is the appropriate feature to handle nested fields when transforming JSON to Parquet.

Exam trap

DEA-C01 often tests the misconception that other transforms like DropFields or FindMatches can handle nested structures, when Relationalize is the specific tool for that purpose.

How to eliminate wrong answers

Option A is wrong because FindMatches is used for deduplication and record matching, not for schema definition or handling nested structures. Option B is wrong because DropFields simply removes fields, which does not help in defining a schema for nested data. Option D is wrong because Spigot is used for debugging by writing sample data to S3, not for transforming nested structures.

26
MCQeasy

A data engineer needs to run a transformation on a small dataset of 500 MB stored in Amazon S3 and load the result into Amazon Redshift. The transformation logic is simple column renaming and filtering. The engineer wants to minimize operational overhead and avoid managing servers. Which approach is most appropriate?

A.Launch an Amazon EMR cluster, run a Spark job, and terminate the cluster after the job completes.
B.Create an AWS Lambda function to read the S3 object and write directly to Redshift.
C.Use AWS Glue with a Python shell job to perform the transformation and load the data.
D.Provision an Amazon EC2 instance, install Apache Spark, and run the transformation as a scheduled cron job.
AnswerC

AWS Glue Python shell jobs are serverless and designed for small to medium workloads that do not require the distributed processing of Apache Spark. For a 500 MB dataset with simple column renaming and filtering, a Python shell job minimizes operational overhead, starts quickly, and integrates with Redshift for loading. This matches the requirement to avoid managing servers.

Why this answer

AWS Glue Python shell jobs are serverless and optimized for small to medium ETL workloads that do not need Spark's distributed processing. For a 500 MB dataset with simple transformations, they provide the lowest operational overhead while integrating with S3 and Redshift. EMR, EC2, and Lambda each introduce either management burden or technical limits that make them less suitable for this scenario.

Exam trap

The trap here is assuming that any serverless option works equally well, when Lambda's 15-minute timeout and limited resources make it a poor fit for ETL and loading into Redshift.

27
Multi-Selectmedium

Which TWO of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose TWO.)

Select 2 answers
A.Improves write throughput
B.Provides microsecond read latency
C.Offloads read traffic from the DynamoDB table
D.Provides data durability across Availability Zones
E.Reduces storage costs
AnswersB, C

DAX is an in-memory write-through cache sitting in front of DynamoDB, serving eventually consistent reads from RAM in microseconds rather than the single-digit millisecond latency of direct DynamoDB access. This satisfies the microsecond read latency benefit.

Why this answer

Option B is correct because DAX is an in-memory write-through cache placed in front of DynamoDB that serves eventually consistent reads from RAM, delivering microsecond read latency for cached items instead of the single-digit millisecond latency of direct DynamoDB reads. Option C is correct because read requests that hit the DAX cache are answered by DAX itself, so they never reach the DynamoDB table, which offloads read traffic and reduces the read capacity units consumed on the underlying table. Option A is not correct because DAX does not accelerate writes: write operations are still written through to DynamoDB synchronously, so write throughput and write latency remain governed by the table.

Option D is not correct because durability across Availability Zones is provided by DynamoDB's own multi-AZ replication of table data, not by DAX, which is a cache and can lose cached data. Option E is not correct because DAX does not reduce the storage consumed by the DynamoDB table; it is an additional caching layer that itself incurs cost.

Exam trap

The trap here is that candidates often confuse DAX's read acceleration with write performance improvements, or assume that a cache provides durability guarantees similar to the underlying database.

28
MCQhard

A financial services company stores trade records in Amazon DynamoDB. An application performs many reads per second for individual trades by tradeId and also runs a nightly analytics job that must read every trade for a given trading day. The table uses tradeId as the partition key only. The nightly job currently performs a full table scan and is slow and expensive. The company wants to optimize the nightly access pattern with minimal application changes. Which change should the data engineer make?

A.Add a local secondary index on the base table with tradeId as the sort key and query it for each day.
B.Create a global secondary index with a partition key of tradingDay and a sort key of tradeId, then query that index for each day.
C.Increase the table's provisioned read capacity units and schedule the full table scan during off-peak hours.
D.Enable DynamoDB Streams on the table and have the nightly job consume the stream to reconstruct the day's trades.
AnswerB

A global secondary index with tradingDay as partition key groups all trades for a day into one logical partition, so the nightly job can issue efficient Query operations instead of scanning the whole table. Keeping tradeId as the sort key allows ordering and range access. The base table partition key and application writes remain unchanged, meeting the minimal-change requirement.

Why this answer

DynamoDB access is most efficient when queries target a specific partition key value. Because the base table partitions by tradeId, retrieving all trades for a day requires a Scan. A global secondary index partitioned by tradingDay gives each day its own logical partition, so the nightly job can use Query to read only that day's items.

The base table schema and write path stay the same, so application changes are minimal.

Exam trap

The trap here is thinking that a local secondary index can change the partition grouping, when it must reuse the base table partition key and therefore cannot group by trading day.

29
MCQmedium

A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The engineer notices that some data records are missing from S3, and the Firehose delivery stream metrics show increased DeliveryToS3.DataFreshness. The engineer needs to ensure all records are delivered. Which action should the engineer take?

A.Increase the number of shards in the source Kinesis data stream to improve throughput.
B.Enable Amazon S3 server access logging on the destination bucket to capture failed deliveries.
C.Enable error logging on the Firehose delivery stream and review the error output in Amazon CloudWatch Logs.
D.Configure the Firehose delivery stream to use a longer buffering interval to reduce the number of S3 PUT requests.
AnswerC

Enabling error logging for Firehose sends delivery errors to CloudWatch Logs, providing visibility into why records are not delivered. Common causes include insufficient S3 permissions, KMS key issues, or malformed records. Reviewing these logs allows the engineer to identify and fix the root cause, ensuring all records are delivered and reducing data freshness.

Why this answer

Enabling error logging on the Firehose delivery stream is crucial for diagnosing why records are not delivered to S3. The logs in CloudWatch will show specific errors such as access denied or invalid records, allowing the engineer to take corrective action. This directly addresses the missing data and high data freshness.

Exam trap

The trap here is assuming that increasing buffering or adding shards will fix missing data, but the real issue is delivery errors that require logging to diagnose.

30
MCQeasy

A company needs to audit all API calls made in their AWS account, including actions performed by the root user. Which AWS service should be used?

A.VPC Flow Logs
B.AWS CloudTrail
C.Amazon CloudWatch Logs
D.AWS Config
AnswerB

CloudTrail records API activity as management and data events, capturing the identity, source IP, timestamp and request details for every call, including root user actions. This satisfies the requirement to audit all API calls across the account.

Why this answer

AWS CloudTrail records all API calls made in an AWS account, including root user actions, for auditing and compliance. Option A is incorrect because VPC Flow Logs capture network traffic metadata, not API calls. Option C is incorrect because Amazon CloudWatch Logs stores log data but does not natively capture API calls; it can ingest logs from other sources.

Option D is incorrect because AWS Config tracks configuration changes to AWS resources, not API calls.

31
MCQmedium

A data engineer needs to ingest data from an on-premises Apache Kafka cluster into Amazon S3. The data volume is about 10 TB per day. The engineer wants to set up a managed Kafka connector. Which AWS service should they use?

A.AWS Database Migration Service
B.AWS Lambda with Kafka trigger
C.Amazon MSK Connect
D.Amazon Kinesis Data Streams
AnswerC

Amazon MSK Connect is the managed Kafka Connect service, so it runs connectors that stream from the on-premises Kafka cluster into Amazon S3 without you operating Connect workers, meeting the managed-connector and 10 TB per day requirement.

Why this answer

Amazon MSK Connect is a managed Kafka connector service that integrates with Amazon MSK or self-managed Apache Kafka clusters to stream data into Amazon S3 using Kafka Connect. It handles the 10 TB/day volume efficiently with auto-scaling and checkpointing, making it the correct choice for a managed connector setup.

Exam trap

The trap here is that candidates confuse Amazon MSK Connect with Amazon Kinesis Data Streams, thinking both are streaming services, but MSK Connect is specifically a managed Kafka connector service for Kafka-to-S3 ingestion, while Kinesis is a separate streaming platform.

How to eliminate wrong answers

Option A is wrong because AWS Database Migration Service (DMS) is designed for database migrations and continuous replication, not for ingesting data from Kafka into S3. Option B is wrong because AWS Lambda with Kafka trigger is event-driven and not suitable for high-volume, continuous streaming of 10 TB/day due to concurrency limits and lack of managed checkpointing for large-scale Kafka ingestion. Option D is wrong because Amazon Kinesis Data Streams is a separate streaming service, not a managed Kafka connector; it cannot directly connect to an on-premises Kafka cluster as a connector.

32
MCQhard

A company is building a data lake on S3 and needs to ingest data from on-premises Oracle database. The data is 5 TB and changes incrementally. The ingestion must capture changes in near real-time (less than 1 minute latency) and be cost-effective. Which approach should be used?

A.Use AWS Database Migration Service (DMS) with ongoing replication to S3
B.Use Amazon Kinesis Data Firehose with an Oracle JDBC connector
C.Use AWS Glue to perform a full table export daily
D.Use AWS DataSync to sync the Oracle data files to S3
AnswerA

DMS supports CDC and can replicate changes to S3 with low latency.

Why this answer

AWS DMS with ongoing replication captures incremental changes from Oracle using its native change data capture (CDC) mechanism, such as Oracle LogMiner or binary logs, and streams them to S3 in near real-time with latency under 1 minute. This approach is cost-effective because DMS charges only for the compute resources used during replication, and S3 storage is inexpensive, making it ideal for a 5 TB dataset with continuous changes.

Exam trap

The trap here is that candidates often confuse Kinesis Data Firehose's ability to accept data from custom sources with native JDBC support, leading them to choose Option B, but Firehose lacks built-in CDC connectors for relational databases like Oracle.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Firehose does not natively support a JDBC connector for Oracle; it ingests data from sources like Kinesis Data Streams, AWS IoT, or custom HTTP endpoints, and using a JDBC connector would require custom code and add complexity without guaranteeing sub-minute latency for CDC. Option C is wrong because AWS Glue performing a full table export daily cannot meet the near real-time requirement of less than 1 minute latency; it is a batch-oriented service designed for periodic ETL jobs, not continuous change capture. Option D is wrong because AWS DataSync is designed for one-time or scheduled bulk data transfers of files or objects, not for capturing incremental database changes from Oracle; it syncs data at the file level, not the row-level CDC needed for a database.

33
MCQmedium

A data engineer notices that an Amazon Redshift cluster is experiencing slow query performance. The engineer suspects that tables are not properly sorted. Which diagnostic query should the engineer run to identify unsorted rows?

A.SELECT * FROM SVV_TABLE_INFO ORDER BY unsorted DESC;
B.SELECT * FROM PG_CATALOG;
C.SELECT * FROM STV_TBL_PERM;
D.SELECT * FROM STL_LOAD_ERRORS;
AnswerA

SVV_TABLE_INFO exposes per-table statistics including the unsorted percentage, so ordering by unsorted DESC surfaces the tables with the most unsorted rows. That directly identifies where a VACUUM SORT is needed to address the slow query performance.

Why this answer

The `SVV_TABLE_INFO` system view in Amazon Redshift provides metadata about each table, including the `unsorted` column which shows the percentage of unsorted rows. By ordering by `unsorted DESC`, the engineer can quickly identify tables with the highest proportion of unsorted data, which directly impacts query performance due to inefficient zone maps and scan pruning.

Exam trap

The trap here is that candidates may confuse `SVV_TABLE_INFO` with `STV_TBL_PERM` (which shows block counts) or `STL_LOAD_ERRORS` (which is for load debugging), missing that only `SVV_TABLE_INFO` exposes the `unsorted` column specifically designed for sort health analysis.

How to eliminate wrong answers

Option B is wrong because `PG_CATALOG` is a system schema containing PostgreSQL catalog tables (e.g., `pg_class`, `pg_attribute`), not a diagnostic view for unsorted rows; it lacks the `unsorted` metric. Option C is wrong because `STV_TBL_PERM` provides block-level storage information (e.g., number of blocks per slice) but does not include a column for unsorted row percentage. Option D is wrong because `STL_LOAD_ERRORS` logs errors from COPY and INSERT operations, such as data type mismatches or malformed CSV rows, and has no relevance to sort key efficiency.

34
MCQeasy

A company uses Amazon DynamoDB for a gaming application. They need to store player session data that expires after 24 hours. Which DynamoDB feature should they use to automatically delete expired items?

A.Time to Live (TTL)
B.DynamoDB auto scaling
C.DynamoDB Streams
D.Point-in-time recovery
AnswerA

Time to Live (TTL) lets you define an attribute holding an expiry timestamp; DynamoDB then deletes each item automatically once that timestamp passes, with no write capacity consumed. This directly satisfies the 24-hour automatic expiry requirement for player session data, eliminating manual cleanup jobs.

Why this answer

DynamoDB Time to Live (TTL) is the correct feature because it allows you to define a per-item timestamp attribute (e.g., `expireAt`) that DynamoDB automatically deletes once that timestamp is reached. This is ideal for expiring session data after 24 hours without requiring custom code or scheduled jobs to scan and delete items, reducing cost and operational overhead.

Exam trap

The trap here is that candidates may confuse TTL with DynamoDB Streams, thinking streams can automatically delete items, but streams only notify of changes and require separate logic to perform deletions.

How to eliminate wrong answers

Option B (DynamoDB auto scaling) is wrong because it manages throughput capacity (read/write units) based on traffic, not item expiration or deletion. Option C (DynamoDB Streams) is wrong because it captures item-level changes (inserts, updates, deletes) in near real-time for downstream processing, but does not automatically delete items. Option D (Point-in-time recovery) is wrong because it provides continuous backups to restore a table to any point within the last 35 days, but does not handle automatic deletion of expired data.

35
Multi-Selecteasy

A data engineer needs to ingest data from a SaaS application that sends webhooks in JSON format. The data must be stored in S3 for batch analysis. Which AWS services can receive the webhooks and store the data in S3 with minimal custom code? (Choose TWO.)

Select 2 answers
A.AWS Lambda with S3 SDK
B.AWS Glue with a Python shell job
C.Amazon Kinesis Data Streams
D.Amazon API Gateway with S3 integration
E.Amazon API Gateway with Kinesis Data Firehose integration
AnswersD, E

API Gateway can directly write to S3.

Why this answer

Amazon API Gateway with S3 integration (Option D) allows you to create a REST API that directly writes incoming webhook payloads to an S3 bucket without any custom code, using a proxy integration with an AWS service. This minimizes custom code because the integration handles the mapping and storage automatically.

Exam trap

The trap here is that candidates often assume Lambda is the only serverless option for webhook ingestion, overlooking API Gateway's direct S3 integration, or they mistakenly think Kinesis Data Streams can directly receive HTTP requests without a custom producer.

36
MCQhard

A company is using Amazon Redshift for data warehousing. The data engineer notices that the STL_ALERT_EVENT_LOG table shows many 'missing statistics' alerts. What is the best course of action to address this issue?

A.Increase the WLM concurrency slots.
B.Run VACUUM on the tables.
C.Enable compression on the tables.
D.Run ANALYZE on the tables.
AnswerD

ANALYZE refreshes table statistics that the query planner uses to build optimal execution plans. Missing statistics force suboptimal plans, degrading performance and generating STL_ALERT_EVENT_LOG entries. Running ANALYZE directly resolves the alerts by restoring accurate cardinality estimates.

Why this answer

The STL_ALERT_EVENT_LOG table records alerts about query performance issues, including 'missing statistics' alerts. This indicates that the query optimizer lacks up-to-date table statistics, leading to suboptimal query plans. Running the ANALYZE command updates table statistics, enabling the optimizer to generate efficient execution plans.

Therefore, option D is the correct course of action.

Exam trap

The trap here is that candidates often confuse VACUUM (which reorganizes data) with ANALYZE (which updates statistics), assuming both are needed for query performance, but only ANALYZE directly resolves 'missing statistics' alerts.

How to eliminate wrong answers

Option A is wrong because increasing WLM concurrency slots does not address missing statistics; it only allows more queries to run simultaneously, which could worsen performance if statistics are outdated. Option B is wrong because VACUUM reclaims disk space and sorts rows but does not update table statistics; it is used for managing data storage, not query optimization. Option C is wrong because enabling compression reduces storage and I/O but does not provide the optimizer with the statistical metadata needed for efficient query planning.

37
MCQmedium

A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 data source, applies a filter transformation, and writes to Amazon S3 in Parquet. The engineer notices that the job is reading all files in the prefix, including files that do not match the expected schema, causing job failures. Which action should the engineer take to ensure only valid files are processed?

A.Enable job bookmarks and set the transformation context to filter invalid records
B.Increase the number of DPUs and enable auto-scaling to handle the invalid files
C.Use a Glue crawler with a custom classifier to catalog only valid files and update the job to read from the catalog table
D.Configure the S3 data source with a path that includes only valid partitions or use a glob pattern to exclude invalid files
AnswerD

Narrowing the S3 source path with a glob pattern or restricting to specific partitions ensures the Glue job only reads files that match the expected schema. This directly prevents invalid files from being processed and avoids job failures. It is the simplest and most targeted fix for the described problem.

Why this answer

The root cause is that the job reads all objects under the prefix, including files that do not match the schema. Restricting the S3 source path with a glob pattern or limiting to valid partitions ensures only conforming files are read. This is a configuration change that directly prevents the failures without adding compute or custom classification logic.

Exam trap

The trap here is reaching for job bookmarks or crawlers to solve a schema mismatch; those features handle incremental processing and cataloging, not selective file exclusion at read time.

38
Multi-Selecthard

A data engineer is designing a data transformation pipeline using AWS Glue. The source data is in Amazon S3 in Parquet format, and the transformed output must be written to another S3 bucket in Parquet format partitioned by year, month, day. The pipeline should handle incremental updates efficiently. Which three features should the engineer use? (Choose THREE.)

Select 3 answers
A.AWS Glue job bookmarks to track processed data
B.Use AWS Glue JobWatch for monitoring job progress
C.Use DynamicFrames instead of Spark DataFrames for schema handling
D.Enable partition pruning in the Glue job
E.Use Spark SQL for transformations
AnswersA, C, D

Enables incremental processing.

Why this answer

AWS Glue job bookmarks track processed data by recording the state of previously processed files and partitions, enabling incremental processing of new or changed data in subsequent runs. This is essential for efficiently handling incremental updates without reprocessing the entire dataset.

Exam trap

The trap here is that candidates may confuse general-purpose tools like Spark SQL or monitoring concepts with the specific AWS Glue features designed for incremental processing and partitioning, leading them to select options that are technically possible but not the three required features.

39
MCQmedium

A company is using Amazon RDS for MySQL and needs to reduce read latency for a global user base. Which AWS feature should be implemented?

A.Multi-AZ deployment
B.Aurora Auto Scaling
C.Read Replicas
D.Cross-Region Replication
AnswerC

Read Replicas allow offloading read queries to reduce latency.

Why this answer

Amazon RDS Read Replicas allow you to offload read traffic from the primary DB instance to one or more read-only copies, which can be placed in different AWS Regions to reduce read latency for a global user base. Unlike Multi-AZ, which is designed for high availability, Read Replicas directly address read performance and latency by distributing read queries closer to users. Cross-Region Replication for RDS MySQL is achieved through Read Replicas, making option C the correct choice for reducing read latency globally.

Exam trap

The trap here is that candidates often confuse Multi-AZ (high availability) with Read Replicas (read scaling), or they assume Cross-Region Replication is a separate feature when it is actually implemented via Read Replicas in RDS MySQL.

How to eliminate wrong answers

Option A is wrong because Multi-AZ deployment provides high availability and automatic failover by maintaining a standby replica in a different Availability Zone, but it does not offload read traffic or reduce read latency for a global user base. Option B is wrong because Aurora Auto Scaling automatically adjusts the number of Aurora Replicas based on workload, but it is a feature of Amazon Aurora, not Amazon RDS for MySQL, and the question specifies RDS for MySQL. Option D is wrong because Cross-Region Replication is not a standalone feature for RDS MySQL; it is implemented using Read Replicas in a different region, so it is a subset of the correct answer, not a separate feature.

40
MCQmedium

A company uses AWS Glue crawlers to populate the Data Catalog from data in Amazon S3. The crawler fails to update the schema when new columns are added to the CSV files. What is the most likely cause?

A.The S3 bucket has versioning enabled.
B.The crawler is configured to only crawl new partitions.
C.The IAM role for the crawler lacks permissions to read the new columns.
D.The crawler uses a custom classifier that defines a fixed schema.
AnswerD

A custom classifier overrides the crawler's built-in schema inference, so it applies the fixed column definitions you supplied rather than re-reading the CSV headers. New columns in the S3 files are therefore ignored, which is exactly the failure described. Removing the classifier or updating its schema restores automatic detection.

Why this answer

AWS Glue crawlers use classifiers to infer the schema of data in Amazon S3. If a custom classifier is used and it defines a fixed schema, the crawler will not detect new columns because the classifier's schema is static. The crawler relies on the classifier to parse the data and determine the schema; if the classifier does not include the new columns, they will be ignored.

Therefore, the most likely cause is that the crawler uses a custom classifier with a fixed schema.

Exam trap

DEA-C01 often tests the misconception that IAM permissions or S3 versioning affect schema detection, when the real issue is the use of a custom classifier with a fixed schema that prevents dynamic schema inference.

How to eliminate wrong answers

Option A is wrong because S3 bucket versioning does not affect the crawler's ability to detect schema changes; versioning is about object versions, not schema inference. Option B is wrong because configuring the crawler to only crawl new partitions would affect which partitions are crawled, but if new columns are added to existing partitions, they might still be detected if the classifier is dynamic. However, the question states the crawler fails to update the schema when new columns are added, which is more directly related to the classifier.

Option C is wrong because if the IAM role lacks permissions to read the new columns, the crawler would fail to read the data entirely, not just fail to update the schema; also, permissions are typically at the bucket/object level, not column level.

41
MCQmedium

A company runs a multi-AZ Amazon RDS for PostgreSQL instance. They need to run a one-time analytical query that will take several hours and consume significant I/O. The query should not impact the primary workload. What should the data engineer do?

A.Create a read replica of the RDS instance and run the query on the replica.
B.Run the query directly on the primary instance during off-peak hours.
C.Increase the instance size to handle the load.
D.Enable Multi-AZ and run the query on the standby instance.
AnswerA

A read replica offloads the analytical query onto a separate database instance, isolating its heavy I/O from the primary. Because replication is asynchronous, the primary workload keeps running unaffected, satisfying the requirement that the query must not impact production. Note that a multi-AZ standby cannot serve read traffic, so it would not work here.

Why this answer

Creating a read replica of the RDS for PostgreSQL instance allows the analytical query to run on a separate database engine without affecting the primary workload. Read replicas in Amazon RDS use asynchronous replication from the source instance, so the replica can handle heavy I/O and long-running queries independently. This ensures the primary instance remains available for the production workload without performance degradation.

Exam trap

The trap here is that candidates often confuse the Multi-AZ standby instance with a read replica, assuming the standby can be used for queries, but in Amazon RDS, the standby is only for high availability and is not accessible for read operations.

How to eliminate wrong answers

Option B is wrong because running the query directly on the primary instance, even during off-peak hours, still consumes significant I/O and CPU resources on that instance, which can impact the primary workload and potentially cause performance issues or increased latency. Option C is wrong because increasing the instance size only adds more resources to the same single instance; the analytical query would still compete with the primary workload for I/O and memory, and scaling up does not isolate the workload. Option D is wrong because the standby instance in a Multi-AZ deployment is not directly accessible for read or write operations; it is a synchronous replica used only for automatic failover, and Amazon RDS does not allow connecting to the standby for queries.

42
MCQhard

A company stores regulated records in an Amazon S3 bucket and must prove that individual objects cannot be deleted or overwritten for 365 days after creation, even by the account root user. The compliance team also needs to retain the ability to delete the bucket itself after the retention window expires. Which configuration meets these requirements?

A.Enable S3 Versioning and apply an S3 Object Lock default retention rule in compliance mode with a 365-day period.
B.Enable S3 Object Lock in governance mode with a 365-day retention period and grant s3:BypassGovernanceRetention to the security team.
C.Apply a bucket policy that denies s3:DeleteObject and s3:PutObject to all principals for 365 days using a date condition.
D.Enable S3 Versioning and add a lifecycle rule that transitions objects to S3 Glacier Deep Archive after 365 days.
AnswerA

Object Lock in compliance mode prevents any principal, including the root user, from deleting or overwriting a protected object version until the retain-until date passes. A default retention rule applies the 365-day period automatically to new objects. Because the lock protects object versions rather than the bucket, the bucket can still be removed once all versions age out and are deleted.

Why this answer

S3 Object Lock in compliance mode provides write-once-read-many protection that no principal, including the root user, can override before the retention date. A default retention rule applies the period automatically to newly created object versions. Since the lock applies to object versions, the bucket itself can still be deleted after retention lapses and versions are removed.

Exam trap

The trap here is confusing governance mode with compliance mode, when only compliance mode is immune to privileged deletion and retention changes.

43
MCQmedium

A data engineer is building an AWS Lake Formation governed data lake. The security team wants to grant a group of analysts access to only the non-sensitive columns of a table in the Data Catalog, while denying access to columns containing Social Security numbers. The analysts use Amazon Athena to query the data. Which Lake Formation permission model should the data engineer use?

A.Grant table-level SELECT permission to the analysts and rely on Athena to hide the sensitive columns.
B.Use an IAM policy that denies athena:GetQueryResults for queries that reference the sensitive column.
C.Create a separate Data Catalog table that contains only non-sensitive columns and grant access to that table.
D.Grant column-level SELECT permission on the non-sensitive columns and exclude the Social Security number column.
AnswerD

Lake Formation supports column-level permissions, allowing you to grant SELECT on specific columns of a table. By granting only the non-sensitive columns, the analysts can query those columns in Athena while the Social Security number column remains inaccessible. This directly implements the least-privilege requirement without duplicating data.

Why this answer

Lake Formation column-level permissions let you grant SELECT on a subset of columns in a Data Catalog table. Analysts can then query the permitted columns through Athena, and any attempt to select the Social Security number column is denied. Duplicating tables or using broad IAM denies does not achieve the same precise, least-privilege control.

Exam trap

The trap here is thinking that table-level SELECT plus Athena behavior hides sensitive columns, when table-level access exposes every column.

44
MCQhard

A data engineer is designing a multi-region disaster recovery plan for an Amazon DynamoDB table. The table stores critical user profile data and must have a Recovery Point Objective (RPO) of less than 1 minute and a Recovery Time Objective (RTO) of less than 5 minutes. Which solution meets these requirements?

A.Configure DynamoDB Streams and a Lambda function to replicate data to another region.
B.Use DynamoDB on-demand backup and restore to another region.
C.Use DynamoDB global tables to replicate data to another region.
D.Enable point-in-time recovery (PITR) and restore to another region.
AnswerC

DynamoDB global tables use multi-region, active-active replication that propagates writes across regions, typically within one second. This satisfies the sub-one-minute RPO, while the replica's continuous availability enables failover well inside the five-minute RTO without restore operations.

Why this answer

DynamoDB global tables provide active-active multi-region replication with sub-second latency, ensuring that data written in one region is automatically replicated to other regions within seconds. This meets the RPO of less than 1 minute and RTO of less than 5 minutes because the table is already available in the secondary region for immediate reads and writes, with no manual restore or failover steps required.

Exam trap

The trap here is that candidates confuse point-in-time recovery (PITR) or on-demand backups with multi-region replication, not realizing that restore operations are manual and time-consuming, whereas global tables provide automatic, near-real-time replication.

How to eliminate wrong answers

Option A is wrong because DynamoDB Streams with a Lambda function introduces asynchronous replication that can have variable latency and potential data loss if the Lambda fails or is throttled, making it unreliable for a sub-1-minute RPO. Option B is wrong because on-demand backup and restore is a manual process that takes minutes to hours to complete, far exceeding the 5-minute RTO and 1-minute RPO. Option D is wrong because point-in-time recovery (PITR) only allows restoring to a point within the last 35 days in the same region, and restoring to another region requires exporting and re-importing data, which cannot achieve sub-5-minute RTO.

45
MCQeasy

A data engineer needs to transform JSON data from an S3 bucket using AWS Glue. The JSON contains nested arrays and objects. Which Glue transform is best suited for flattening nested structures?

A.Unnest
B.ResolveChoice
C.Relationalize
D.Map
AnswerC

Relationalize converts nested JSON arrays and objects into separate, related DynamoDB-style tables, unnesting the hierarchy. This flattening produces the tabular structure Glue and Athena require, which ApplyMapping or Filter alone cannot achieve for deeply nested data.

Why this answer

The Relationalize transform is specifically designed to flatten nested JSON structures (arrays and objects) into a set of related tables, making it ideal for this use case. It automatically handles complex nesting by creating separate DataFrames for each nested level and linking them via foreign keys, which is exactly what is needed when ingesting JSON with nested arrays and objects into a relational format.

Exam trap

The trap here is that candidates often confuse the generic Spark SQL function `explode` (or the concept of 'unnesting') with a named AWS Glue transform, leading them to select 'Unnest' even though it does not exist as a Glue transform and would require manual handling of multiple nesting levels.

How to eliminate wrong answers

Option A (Unnest) is wrong because AWS Glue does not have a built-in transform named 'Unnest'; this is a Spark SQL function (e.g., `explode`) but not a named Glue transform, and it would require manual handling of multiple nesting levels. Option B (ResolveChoice) is wrong because it is used to resolve schema ambiguities (e.g., when a column has mixed types like string and int) and does not flatten nested structures. Option D (Map) is wrong because it applies a function to each record in a DynamicFrame for row-wise transformations, but it does not inherently flatten nested arrays or objects—you would need to write custom logic to handle the nesting.

46
MCQmedium

A data engineer is troubleshooting an issue where an AWS Glue ETL job fails when trying to read data from an S3 bucket encrypted with SSE-KMS. The job has an IAM role that includes `kms:Decrypt` permission. What is the most likely reason for the failure?

A.The IAM role does not have s3:GetObject permission
B.The KMS key policy does not allow the Glue job to use the key
C.The S3 bucket is in a different AWS region than the Glue job
D.The Glue job is not configured to use the KMS key for decryption
AnswerB

Correct. The most likely cause is that the KMS key policy does not grant the Glue job's IAM role permission to use the key. Even with kms:Decrypt in the role, the key policy must allow it.

Why this answer

Even if the IAM role has `kms:Decrypt` permission, the KMS key policy must also allow the Glue job's role to use the key. KMS key policies are resource-based policies that control access to the key, and if they do not explicitly grant permission to the Glue job's role, the decrypt operation will fail. This is a common oversight when using SSE-KMS with services like AWS Glue.

Exam trap

DEA-C01 often tests the dual requirement of IAM policies and KMS key policies for accessing encrypted data, and candidates may overlook the key policy.

How to eliminate wrong answers

Option A is wrong because the question states the job has an IAM role that includes `kms:Decrypt` permission, but it does not mention s3:GetObject; however, the failure is specifically about reading data encrypted with SSE-KMS, and the most likely reason is the KMS key policy, not missing S3 permissions. Option C is wrong because S3 and Glue can be in different regions, but cross-region access is possible if configured; it is not the most likely reason for failure. Option D is wrong because the Glue job does not need to be configured to use the KMS key for decryption; the S3 service handles decryption on behalf of the job if permissions are correct.

47
MCQmedium

A company has an Amazon RDS for PostgreSQL DB instance with a large table that is frequently updated. The data engineer needs to reduce storage costs by archiving old records that are no longer accessed. The archived records must be retained for 7 years due to compliance requirements. Which solution is MOST cost-effective?

A.Use RDS native backup and restore to keep a separate backup.
B.Export old records using pg_dump and store in S3 Glacier Deep Archive.
C.Enable storage autoscaling on the RDS instance.
D.Move old records to a separate table in the same RDS instance.
AnswerB

pg_dump extracts the old records, and S3 Glacier Deep Archive provides the lowest-cost long-term retention tier suitable for seven-year compliance archiving. This removes the data from expensive RDS storage while preserving it durably and cheaply.

Why this answer

Exporting old records via pg_dump and storing them in S3 Glacier Deep Archive provides the lowest-cost storage for data that must be retained for 7 years but is never accessed. S3 Glacier Deep Archive offers retrieval times of 12–48 hours at a storage cost of approximately $0.00099/GB/month, far cheaper than any RDS storage tier. This approach removes the archived data from the RDS instance, reducing provisioned storage costs while meeting compliance requirements.

Exam trap

The trap here is that candidates confuse 'archiving' with 'backup' or 'storage autoscaling,' failing to recognize that only moving data out of the RDS instance to a low-cost storage class like S3 Glacier Deep Archive actually reduces ongoing storage costs.

How to eliminate wrong answers

Option A is wrong because RDS native backup and restore creates full or snapshot backups of the entire DB instance, not just the old records, and storing these backups for 7 years would incur high costs for redundant data and storage fees. Option C is wrong because enabling storage autoscaling only increases storage capacity when thresholds are reached, which does not reduce costs or archive old records—it merely prevents out-of-space errors. Option D is wrong because moving old records to a separate table in the same RDS instance does not reduce storage costs; the data still occupies the same provisioned storage, and the table remains part of the instance's billed capacity.

48
MCQeasy

Refer to the exhibit. A data engineer runs this CLI command to check an object's metadata. The engineer wants to verify if the object is eligible for lifecycle transition to S3 Glacier based on its age. What additional information is needed?

A.The current date
B.The ETag value
C.The metadata archive flag
D.The ContentLength value
AnswerA

Lifecycle transition eligibility depends on the object's age, calculated as the current date minus LastModified. The CLI output supplies LastModified, so the current date is the missing value needed to determine whether the transition threshold has elapsed.

Why this answer

To determine whether an S3 object is eligible for lifecycle transition to S3 Glacier based on age, you need the object's creation date (LastModified) and the current date to compute its age. The exhibit presumably shows the object's LastModified timestamp, so the missing piece is the current date to calculate how many days old the object is relative to the lifecycle rule's transition threshold.

Exam trap

DEA-C01 often tests whether candidates focus on object metadata like ETag or size when the actual lifecycle eligibility depends on object age (LastModified vs. current date) and rule thresholds.

How to eliminate wrong answers

Option B is wrong because the ETag is a hash of the object's content (or multipart upload identifier) and has no bearing on lifecycle transition eligibility, which depends on age. Option C is wrong because there is no 'metadata archive flag' in S3 object metadata that governs lifecycle transitions; lifecycle rules are evaluated by the S3 service based on object age and rule configuration. Option D is wrong because ContentLength (object size) does not determine lifecycle transition eligibility unless the lifecycle rule explicitly filters by size, which is not the scenario described.

49
Multi-Selecthard

A data engineer is optimizing an AWS Glue ETL job that reads a large dataset from Amazon S3 and writes to Amazon Redshift. The job currently runs slowly and consumes many DPUs. The engineer wants to improve performance and reduce cost. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Increase the number of DPUs to the maximum allowed for the job
B.Partition the source data in Amazon S3 and use predicate pushdown in the Glue job
C.Enable job bookmarks to process only new data on subsequent runs
D.Convert the output to CSV instead of Parquet to reduce write time
E.Disable auto-scaling to keep the job at a fixed capacity
AnswersB, C

Partitioning the S3 data and using predicate pushdown allows Glue to read only the partitions needed for the query, reducing I/O and the amount of data processed. This improves job performance and lowers DPU usage. It is a targeted optimization that addresses the root cause of slow reads from large datasets.

Why this answer

Job bookmarks reduce the data read on subsequent runs by tracking processed files, and partitioning with predicate pushdown reduces the data scanned per run. Together they lower runtime and DPU consumption, improving performance while reducing cost. Increasing DPUs or converting to CSV would increase cost or hurt performance, and disabling auto-scaling removes a cost-saving feature.

Exam trap

The trap here is thinking that adding more DPUs is always the answer to a slow Glue job; the question asks to improve performance and reduce cost, which requires reducing data processed rather than scaling compute.

50
Multi-Selecthard

A data engineering team is building a data lake on Amazon S3. They need to catalog data and make it queryable by Amazon Athena and Amazon Redshift Spectrum. The data arrives in multiple formats and the schema evolves frequently. Which TWO actions should the team take to support schema evolution and efficient querying? (Choose two.)

Select 2 answers
A.Define tables in the AWS Glue Data Catalog and use AWS Glue crawlers to infer and update schemas as new data arrives.
B.Store all data as uncompressed CSV to maximize compatibility with all query engines.
C.Create a separate S3 bucket for each schema version and require analysts to query the correct bucket manually.
D.Use open columnar formats such as Apache Parquet or ORC with partition prefixes, and register partitions in the Data Catalog.
E.Enable S3 Object Lock on the data lake bucket to preserve schema versions.
AnswersA, D

The AWS Glue Data Catalog is the central metastore used by Athena and Redshift Spectrum. Crawlers inspect data, infer schemas, and update table definitions, so new columns and partitions are reflected automatically. This supports frequent schema evolution without manual DDL, and both query engines can immediately use the updated metadata.

Why this answer

The AWS Glue Data Catalog with crawlers keeps table and partition metadata current as schemas evolve, and both Athena and Redshift Spectrum read from it. Using open columnar formats such as Parquet or ORC with partitioned prefixes enables column projection and partition pruning, which cut scanned bytes and cost. Together, these two actions support evolving schemas and efficient querying across both engines.

Exam trap

The trap here is assuming that a single format or manual bucket separation solves schema evolution, when the Data Catalog plus columnar, partitioned storage is what enables both evolution and efficient querying.

51
MCQhard

A data pipeline uses AWS Glue to read from an Amazon S3 bucket containing millions of small CSV files (each < 1 MB). The ETL job is slow. Which optimization would most improve performance?

A.Write the ETL script using PySpark instead of Scala
B.Increase the number of Glue workers
C.Use the G.1X worker type for more memory
D.Use S3 file grouping to combine small files
AnswerD

Glue's S3 file grouping (groupFiles) coalesces many small objects into larger partitions per task, cutting per-file listing, opening and request overhead. Millions of sub-1 MB CSVs otherwise dominate runtime, so grouping directly addresses the small-file constraint.

Why this answer

AWS Glue (and Spark generally) performs poorly with millions of tiny files because each file requires a separate S3 GET request, metadata operation, and task scheduling overhead. S3 file grouping (via 'groupFiles' and 'groupSize' job parameters) coalesces multiple small files into a single read per Spark partition, drastically reducing the number of S3 operations and task overhead. This directly addresses the root cause of the slowness.

Exam trap

DEA-C01 often tests the instinct to 'add more workers' for any slow Glue job — candidates miss that small-file overhead is an I/O/metadata problem that horizontal scaling cannot fix.

How to eliminate wrong answers

Option A is wrong because PySpark vs. Scala is a language choice — both compile to the same Spark execution engine, so performance is essentially identical for this workload. Option B is wrong because adding workers increases parallelism but each worker still opens millions of tiny files, so the per-file overhead remains and may even worsen due to more concurrent S3 requests.

Option C is wrong because G.1X provides more memory per worker, but the bottleneck is I/O and metadata overhead, not memory — more memory does not reduce the number of file opens.

52
Multi-Selecthard

A data engineer is implementing a CDC (Change Data Capture) pipeline from a relational database to Amazon S3 using AWS Database Migration Service (DMS). Which TWO configurations are required for continuous replication?

Select 2 answers
A.Define transformation rules in the DMS task.
B.Enable binary logging on the source database.
C.Configure a VPC endpoint for DMS.
D.Enable 'Full load' and 'Ongoing replication' in the task.
E.Pre-create the target table in S3.
AnswersB, D

Binary logging records every row-level change on the source, which DMS reads to capture ongoing inserts, updates and deletes. Without it, the source retains no change history, so continuous replication cannot proceed beyond the initial load.

Why this answer

Option B is correct because AWS DMS continuous replication (CDC) requires the source relational database to expose its change stream — for MySQL/MariaDB that means binary logging (log_bin) enabled with binlog_format=ROW, and for other engines the equivalent (e.g., Oracle supplemental logging, SQL Server MS-CDC). Option D is correct because the DMS task's migration type must be set to 'Full load and ongoing replication' (or 'Ongoing replication' if data is already loaded) so DMS performs the initial load and then continuously applies CDC changes to Amazon S3. Option A is not required — transformation rules are optional and only used to rename, filter, or modify schema/data during migration.

Option C is not required — a VPC endpoint is only needed for private connectivity to services like S3 when using a VPC-attached replication instance, not as a general CDC requirement. Option E is not required — DMS writes to S3 as CSV or Parquet files and does not need a pre-created target table.

Exam trap

DEA-C01 often tests that 'continuous replication' requires both source-side logging and task-level CDC configuration — candidates may select transformation or networking options that sound relevant but don't enable change capture.

53
MCQeasy

A data engineer needs to ensure that an Amazon Redshift cluster encrypts all data at rest. Which setting must be enabled when creating the cluster?

A.Enable automated snapshots
B.Enable encryption
C.Enable SSL/TLS
D.Enable VPC
AnswerB

Enabling encryption on the cluster applies AWS KMS-backed encryption to all data at rest, including the cluster's storage volumes and snapshots. This directly satisfies the stem's requirement that the Redshift cluster encrypt all data at rest, and it is the setting exposed during cluster creation.

Why this answer

Amazon Redshift encryption at rest is enabled during cluster creation by selecting the encryption option. Option A is incorrect because enabling automated snapshots is for backup and recovery, not encryption. Option C is incorrect because SSL/TLS ensures encryption in transit, not at rest.

Option D is incorrect because VPC is for network isolation, not encryption.

54
MCQhard

A data engineer is troubleshooting an Amazon Redshift cluster that is not allowing connections from a specific IP range. The engineer verified that the cluster's security group allows inbound traffic from the IP range. What is the next step to resolve the issue?

A.Modify the Redshift cluster parameter group to enable public accessibility.
B.Verify that the cluster's security group is attached to the Redshift cluster.
C.Check the IAM role associated with the Redshift cluster.
D.Check the network ACL (NACL) associated with the Redshift cluster's subnet.
AnswerD

Security groups are stateful and evaluated after subnet-level filtering; a restrictive NACL on the cluster's subnet silently drops inbound traffic before it reaches the security group. Checking the NACL is therefore the correct next diagnostic step.

Why this answer

Even if the security group allows inbound traffic from a specific IP range, the network ACL (NACL) associated with the Redshift cluster's subnet can block traffic at the subnet level. NACLs are stateless and can override security group rules. Option A is incorrect because modifying the cluster parameter group does not control network-level access; public accessibility is a separate setting.

Option B is incorrect because the engineer already verified the security group, but even if it is correctly attached, the NACL could still block traffic. Option C is incorrect because IAM roles control authentication and authorization, not network connectivity.

55
Multi-Selectmedium

A company uses Amazon Redshift to store customer data. The security team requires that all queries are logged for auditing purposes. Which step should be taken to meet this requirement? (Select ONE.)

Select 1 answer
A.Enable AWS CloudTrail database audit logging.
B.Use AWS CloudTrail to log Redshift API calls.
C.Enable logging on the Redshift security group.
D.Enable VPC Flow Logs for the Redshift cluster.
E.Enable Amazon Redshift audit logging to an S3 bucket.
AnswersE

Audit logging captures connection, user, and query activity, then delivers it to Amazon S3 for durable retention. This directly satisfies the security team's requirement that all queries are logged for auditing, since S3 provides the persistent, reviewable store auditors need.

Why this answer

The requirement is to log all queries for auditing. Amazon Redshift's native audit logging captures connection logs, user activity logs, and query logs, and can be exported to an S3 bucket. This is the only step that directly logs SQL queries.

AWS CloudTrail does not log SQL queries; it logs management API calls (e.g., CreateCluster, ModifyCluster). Therefore, only Option E meets the requirement.

Exam trap

The trap is that many candidates assume AWS CloudTrail can log SQL queries, but it only logs API calls. The correct answer is solely Amazon Redshift's native audit logging.

56
MCQhard

A company uses Amazon DynamoDB as the primary data store for a gaming application. The application experiences sudden spikes in traffic. The data engineer notices that write requests are throttled during peak times. The partition keys are well-distributed. What should the data engineer do to reduce throttling?

A.Use DynamoDB global tables to distribute writes across regions.
B.Configure DynamoDB auto scaling to adjust write capacity automatically.
C.Increase the number of partition keys to improve write distribution.
D.Enable DynamoDB Accelerator (DAX) to cache write operations.
AnswerB

Auto scaling adjusts provisioned write capacity in response to consumed capacity, directly addressing the sudden traffic spikes causing throttling. Since partition keys are already well-distributed, the bottleneck is provisioned throughput rather than hot partitions, so scaling write capacity automatically absorbs peak demand without manual intervention.

Why this answer

DynamoDB auto scaling allows the table to automatically adjust its provisioned write capacity based on actual traffic patterns, preventing throttling during sudden spikes without manual intervention. Since the partition keys are already well-distributed, throttling is likely due to insufficient write capacity units, which auto scaling can dynamically increase.

Exam trap

The trap here is that candidates may confuse throttling due to hot partitions (uneven key distribution) with throttling due to insufficient overall capacity, leading them to incorrectly choose option C even when the question explicitly states partition keys are well-distributed.

How to eliminate wrong answers

Option A is wrong because DynamoDB global tables replicate data across regions for disaster recovery and low-latency reads, but they do not increase write capacity within a single region; writes are still subject to the same per-table capacity limits. Option C is wrong because the partition keys are already well-distributed, so adding more partition keys would not resolve throttling caused by insufficient provisioned write capacity; throttling occurs when write requests exceed the table's write capacity units, not due to partition key distribution. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for read operations only, not writes; it cannot reduce write throttling.

57
MCQeasy

A data engineer stores sensitive records in an Amazon S3 bucket. The security team wants to guarantee that every object is encrypted before it is written to disk and that the bucket automatically rejects any unencrypted PUT request, regardless of which IAM principal sends it. Which configuration should the data engineer apply?

A.Enable S3 Block Public Access at the account level so that only encrypted objects can be uploaded.
B.Enable default bucket encryption with SSE-S3 and rely on the S3 console warning for unencrypted uploads.
C.Configure an S3 Lifecycle rule to transition objects to S3 Glacier Instant Retrieval, which encrypts data at rest.
D.Attach a bucket policy that denies s3:PutObject when the request lacks the s3:x-amz-server-side-encryption condition key.
AnswerD

A bucket policy with a Deny effect and the s3:x-amz-server-side-encryption condition key rejects any PUT that does not carry a server-side encryption header. This enforces encryption at the bucket level for every principal, including the root account, and works with SSE-S3, SSE-KMS, or DSSE-KMS. It directly satisfies the requirement to block unencrypted writes.

Why this answer

The only mechanism that blocks an unencrypted PUT before the object is stored is a bucket policy using a Deny effect with the s3:x-amz-server-side-encryption condition key. Default encryption and console warnings are passive safeguards that cannot stop an explicit unencrypted request from an SDK or CLI. Block Public Access and lifecycle rules address different concerns entirely.

Exam trap

The trap here is assuming that enabling default bucket encryption prevents unencrypted uploads, when it only supplies encryption for requests that omit the header.

58
MCQmedium

A company uses AWS Lake Formation to manage data lake permissions. A data analyst cannot query a table in Athena, although the table appears in the catalog. The analyst has IAM permissions to run Athena. What is the MOST likely cause?

A.The Glue Data Catalog does not have the table registered.
B.The S3 bucket policy denies access to the analyst's IAM role.
C.The analyst lacks Lake Formation permissions on the table.
D.The Athena workgroup is not configured with the correct output location.
AnswerC

Lake Formation enforces its own table and column-level grants on top of IAM. The analyst's IAM policy permits Athena API calls, but without a Lake Formation SELECT grant on the table, the query is denied — explaining why the table appears in the catalog yet remains unqueryable.

Why this answer

The most likely cause is that the analyst lacks Lake Formation permissions on the table. Even with IAM permissions to run Athena, Lake Formation enforces fine-grained access control at the table and column level. If the analyst has not been granted SELECT permission on the table via Lake Formation, Athena queries will fail with an access denied error, even though the table appears in the Glue Data Catalog.

Exam trap

The trap is assuming that IAM permissions alone are sufficient to query data in a Lake Formation-governed data lake. Candidates might overlook Lake Formation's role and pick S3 bucket policy or Glue catalog issues. The exam tests the layered security model where Lake Formation permissions are required in addition to IAM.

How to eliminate wrong answers

Option A is wrong because if the table appears in the catalog, it is registered; the issue is not registration. Option B is wrong because while S3 bucket policies can deny access, Lake Formation manages access to underlying data, and if the analyst lacks Lake Formation permissions, that is the more direct cause. Option D is wrong because an incorrect Athena output location would cause query failures due to write permission issues, but the error would typically mention the output location, not access to the table.

59
MCQmedium

A company runs a data pipeline that uses AWS Glue to process data from an Amazon DynamoDB table and write results to Amazon S3. The Glue job runs on a schedule every hour. Recently, the job started failing intermittently with 'ProvisionedThroughputExceededException' errors from DynamoDB. What is the BEST solution?

A.Use DynamoDB Accelerator (DAX) to reduce read latency.
B.Change the Glue job schedule to run every 2 hours.
C.Implement exponential backoff and retries in the Glue job for DynamoDB operations.
D.Increase the read capacity units of the DynamoDB table.
AnswerC

Exponential backoff with jitter retries throttled DynamoDB requests, letting the table's provisioned capacity recover between attempts. This directly addresses the intermittent ProvisionedThroughputExceededException, which signals transient throttling rather than a persistent capacity shortfall. Retries absorb those spikes without over-provisioning, satisfying the requirement to stop hourly Glue job failures.

Why this answer

ProvisionedThroughputExceededException is a throttling error that occurs when request rate exceeds the table's provisioned read/write capacity, and AWS's documented best practice is to implement exponential backoff and retry logic so the client retries with increasing delays instead of hammering the table. This is the correct fix for intermittent throttling because it lets the Glue job ride out short bursts of capacity contention without failing. It also works whether the table is in provisioned or on-demand mode, and it is the least disruptive, most cost-effective change.

Exam trap

DEA-C01 often tests whether candidates jump to 'add more capacity' when the AWS-recommended answer for transient throttling is retry with exponential backoff — the exam wants the resilience pattern, not the brute-force capacity increase.

How to eliminate wrong answers

Option A is wrong because DAX reduces read latency for cached items but does not eliminate throttling on the underlying table for uncached reads or for writes, and Glue's DynamoDB connector does not automatically route through DAX without custom configuration. Option B is wrong because changing the schedule to every 2 hours only reduces frequency; it does not address the bursty request pattern that causes the throttle, and it delays data freshness. Option D is wrong because increasing RCUs may help but is a guess at the root cause, costs more money, and does not address write throttling or burst behavior — and the question asks for the BEST solution, which is the AWS-recommended retry pattern.

60
MCQmedium

Refer to the exhibit. A data engineer is troubleshooting an AWS Lambda function that processes data from Amazon S3. The function is triggered by S3 events, but no logs appear in CloudWatch Logs. The engineer runs the AWS CLI command shown. What is the MOST likely reason for the missing logs?

A.The Lambda execution role does not have permissions to create log groups and write logs.
B.The Lambda function is configured to log to a different log group.
C.The Lambda function is not being invoked by S3 events.
D.The log retention policy is set to 7 days, causing logs to expire immediately.
AnswerA

Without `logs:CreateLogGroup`, `logs:CreateLogStream` and `logs:PutLogEvents` in the execution role, Lambda cannot create the function's log group or stream, so invocations produce no CloudWatch output at all. This directly satisfies the stem's missing-logs constraint, since the trigger itself is firing normally.

Why this answer

The most likely reason for missing CloudWatch Logs is that the Lambda execution role lacks the necessary permissions to create log groups and write log streams. When a Lambda function is invoked, it attempts to create a log group named /aws/lambda/<function-name> and a log stream, then write logs. If the execution role does not include actions like logs:CreateLogGroup, logs:CreateLogStream, and logs:PutLogEvents, the function cannot write any logs, resulting in no log entries.

This is a common misconfiguration, especially when roles are custom-created without the default AWSLambdaBasicExecutionRole policy.

Exam trap

DEA-C01 often tests the misconception that Lambda automatically has permissions to write logs, but the execution role must explicitly allow it; candidates may overlook this and choose other options like log retention or invocation issues.

How to eliminate wrong answers

Option B is wrong because Lambda does not support configuring a different log group; logs are always sent to the log group /aws/lambda/<function-name> unless you use a custom logging solution, which is not indicated here. Option C is wrong because if the function were not invoked, there would be no logs, but the question states the function is triggered by S3 events; moreover, the exhibit likely shows a CLI command that confirms invocations, making this less likely. Option D is wrong because a 7-day retention policy would not cause logs to expire immediately; logs would still be visible for 7 days, and the issue is missing logs entirely, not expired ones.

61
Multi-Selecthard

A data engineer is designing a data lake on Amazon S3 that will store sensitive financial data. The security team requires that access to the data be audited, that data be encrypted at rest with customer-managed keys, and that the engineer be able to identify which IAM principals accessed specific objects. Which TWO AWS services or features should the engineer use to meet these requirements? (Choose two.)

Select 2 answers
A.AWS CloudTrail data events for S3
B.Amazon S3 server access logging
C.Amazon Macie for sensitive data discovery
D.AWS KMS customer managed keys with SSE-KMS
E.AWS Secrets Manager for storing encryption keys
AnswersA, D

AWS CloudTrail data events capture object-level API activity, such as GetObject and PutObject, including the identity of the caller and the object accessed. This provides the audit trail required to identify which IAM principals accessed specific objects. Management events alone do not log object-level access, so data events are necessary for this granular auditing.

Why this answer

CloudTrail data events provide the necessary audit trail for object-level access, identifying IAM principals. SSE-KMS with customer managed keys satisfies the encryption at rest requirement with customer-managed keys. Together, they meet the auditing and encryption needs.

Exam trap

The trap here is confusing server access logging with CloudTrail data events; while both log access, CloudTrail data events are integrated with IAM and CloudTrail Lake for easier analysis of principal activity.

62
MCQmedium

A data engineer manages an AWS Glue ETL job that processes JSON files from Amazon S3 and writes to Amazon Redshift. The job fails with the error 'Unable to find a suitable JDBC driver in the classpath'. The engineer has included the Redshift JDBC driver as a job parameter in the Glue job configuration. Which step should the engineer take to resolve the error?

A.Add the JDBC driver as an additional jar in the Glue job's 'Dependent jars path' and ensure the job's IAM role has s3:GetObject permission on the jar location.
B.Modify the job script to include the JDBC driver URL directly in the 'connection_options' when calling the write_dynamic_frame method.
C.Increase the number of workers in the Glue job to ensure the driver is loaded in parallel across all executors.
D.Convert the Glue job to use a development endpoint and manually install the JDBC driver on the underlying EC2 instance.
AnswerA

AWS Glue requires that JDBC drivers be provided as additional jars via the 'Dependent jars path' and accessible from S3. Simply setting a connection parameter is insufficient. The job's IAM role must also have permission to read the jar from S3. This is the documented approach for using custom JDBC drivers with Glue.

Why this answer

To use a custom JDBC driver with AWS Glue, you must upload the driver jar to Amazon S3 and specify its path in the job's 'Dependent jars path' parameter. Additionally, the IAM role associated with the job must have s3:GetObject permission for that jar. This ensures the driver is available in the classpath at runtime, resolving the 'Unable to find a suitable JDBC driver' error.

Exam trap

The trap here is assuming that adding the JDBC driver as a job parameter or connection option automatically makes it available to the job's classpath.

63
MCQeasy

A data engineer needs to load data from an Amazon S3 bucket into an Amazon Redshift cluster as part of an ETL pipeline. The source files are already in Parquet format and the engineer wants the fastest load with minimal transformation. Which Redshift load method should the engineer use?

A.Use the COPY command with the PARQUET format option.
B.Use an AWS Glue ETL job to read the Parquet files and write to Redshift using the Redshift connector.
C.Use the Redshift UNLOAD command to move the Parquet files into the cluster.
D.Use an Amazon Kinesis Data Firehose delivery stream with Redshift as the destination.
AnswerA

The Redshift COPY command is the native, high-throughput bulk load utility and it reads Parquet directly from Amazon S3. Using the PARQUET format option lets Redshift parse the columnar file natively, which is faster than loading text formats and requires no intermediate transformation. It is the recommended approach for loading Parquet files into Redshift.

Why this answer

The Redshift COPY command is the native bulk loader and supports Parquet directly via the PARQUET format option, which reads columnar data efficiently from S3. For Parquet files with no required transformation, COPY is faster and simpler than a Glue job, and the other options either move data the wrong direction or add streaming overhead.

Exam trap

The trap here is reaching for a managed ETL service or streaming delivery when the native COPY command already loads Parquet from S3 directly and is the fastest option for untransformed data.

64
MCQmedium

A data pipeline uses AWS Glue to process data from Amazon S3. The job fails with an 'OutOfMemoryError' during the transformation phase. Which action should the data engineer take to resolve this issue?

A.Enable S3 server-side encryption.
B.Increase the number of partitions in the input data.
C.Change the data format from CSV to Parquet.
D.Increase the number of DPUs (Data Processing Units) for the Glue job.
AnswerD

Adding DPUs increases the total executor memory and parallelism available to the Glue job, letting Spark distribute transformation partitions across more workers. This directly addresses the OutOfMemoryError caused by insufficient memory per executor during the transformation phase.

Why this answer

The OutOfMemoryError occurs because the Glue job does not have enough memory allocated. Increasing the number of DPUs (Data Processing Units) increases both memory and processing capacity, directly resolving the issue. Option A (S3 server-side encryption) affects data security, not memory.

Option B (increasing data partitions) may help parallelism but does not directly increase memory per executor. Option C (changing to Parquet) can reduce data volume but does not guarantee sufficient memory for transformation.

65
Multi-Selecteasy

A data engineer is setting up an AWS Glue job to process data from an Amazon S3 bucket. The job fails with an 'Access Denied' error. Which TWO IAM permissions are MOST likely missing from the Glue job's IAM role?

Select 2 answers
A.s3:PutObject
B.kms:Decrypt
C.dynamodb:GetItem
D.glue:StartJobRun
E.s3:GetObject
AnswersA, E

Writing transformed output back to Amazon S3 requires s3:PutObject on the target prefix. Without it, the Glue job's write stage returns Access Denied even when reading succeeds, so this permission must be added to the job's IAM role.

Why this answer

The Glue job needs to read the source data from the S3 bucket, which requires the s3:GetObject permission on the bucket's objects, so option E is correct. The job also needs to write its output (or intermediate results) back to S3, which requires the s3:PutObject permission, making option A correct. Option B (kms:Decrypt) is not necessarily missing unless the S3 objects are encrypted with a customer-managed KMS key and the role lacks decrypt rights, but the scenario does not state that.

Option C (dynamodb:GetItem) is unrelated because the job processes data from S3, not DynamoDB. Option D (glue:StartJobRun) is a permission for triggering jobs, not for the job's execution role to access data, so it is not the cause of the Access Denied error.

66
MCQmedium

A data engineer needs to transform data in Amazon S3 using SQL statements without managing any infrastructure. The transformations are simple projections and filters, and the engineer wants the results written back to S3 in Parquet. Which AWS service should be used?

A.Amazon Athena with a CREATE TABLE AS SELECT statement
B.AWS Glue ETL with a PySpark script
C.Amazon EMR with Hive on a transient cluster
D.Amazon Redshift Spectrum with an external schema
AnswerA

Athena is serverless and runs standard SQL against data in S3, and CREATE TABLE AS SELECT writes the transformed result back to S3 in a specified format such as Parquet. This matches the requirement for simple SQL transformations with no infrastructure to manage.

Why this answer

Amazon Athena is fully serverless and executes SQL directly against S3 data. A CREATE TABLE AS SELECT statement reads the source, applies projections and filters, and writes the result to S3 in a chosen format like Parquet. This satisfies the SQL and no-infrastructure requirements without cluster or job management.

Exam trap

The trap here is equating SQL-on-S3 with Glue or EMR, when Athena is the serverless option purpose-built for running SQL directly against S3 data.

67
MCQmedium

A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The migration must have minimal downtime. Which service should be used to perform the migration?

A.AWS Glue
B.AWS DataSync
C.AWS Database Migration Service (DMS)
D.Amazon S3 Transfer Acceleration
AnswerC

AWS Database Migration Service performs continuous change data capture replication from MongoDB to Amazon DocumentDB, keeping the target synchronised until cutover. This satisfies the minimal-downtime constraint, since the source stays live during the initial full load and ongoing replication, unlike offline dump-and-restore approaches that require stopping writes.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it supports continuous replication from MongoDB to Amazon DocumentDB using change data capture (CDC), enabling near-zero downtime migrations. DMS can perform a full load of existing data and then apply ongoing changes from the source MongoDB oplog, keeping the target DocumentDB synchronized until the cutover.

Exam trap

The trap here is that candidates may confuse AWS DMS with AWS DataSync or AWS Glue, assuming any 'migration' or 'data transfer' service can handle live database replication, but only DMS provides the necessary CDC engine for heterogeneous database migrations with minimal downtime.

How to eliminate wrong answers

Option A is wrong because AWS Glue is a serverless data integration service for ETL (extract, transform, load) jobs, not designed for live database migration with minimal downtime; it lacks native CDC support for MongoDB to DocumentDB replication. Option B is wrong because AWS DataSync is optimized for moving large volumes of file data (e.g., NFS, SMB) to AWS storage services like S3 or EFS, not for heterogeneous database migrations or ongoing replication. Option D is wrong because Amazon S3 Transfer Acceleration is a feature that speeds up uploads to S3 buckets over long distances using edge locations; it has no capability to migrate or replicate a MongoDB database to DocumentDB.

68
MCQeasy

A company stores raw customer records in an Amazon S3 bucket and processes them with AWS Glue. A governance requirement states that a specific tag named DataClass must exist on every catalog table, and any table missing that tag must not be queryable. Where should the data engineer enforce this requirement with the least operational effort?

A.Enable AWS CloudTrail data events on the S3 bucket and alert when untagged tables are queried.
B.Write an AWS Lambda function that scans the Data Catalog hourly and deletes any table missing the DataClass tag.
C.Attach an IAM policy to all analyst roles that denies glue:GetTable unless the table has the DataClass tag.
D.Use AWS Lake Formation tag-based access control by defining an LF-Tag and granting table permissions only when the DataClass tag value matches.
AnswerD

Lake Formation tag-based access control lets an LF-Tag named DataClass be attached to catalog resources, and permissions are granted based on tag values rather than per-table grants. Tables without the required tag value receive no matching grant, so they are not queryable, which enforces the rule centrally with minimal ongoing effort.

Why this answer

Lake Formation tag-based access control centralizes permission decisions on tags instead of individual tables. Defining an LF-Tag named DataClass and granting access only for matching tag values means any table lacking the required tag value has no applicable grant and is therefore not queryable, satisfying the governance rule with minimal administrative effort.

Exam trap

The trap here is assuming IAM can evaluate Glue Data Catalog tags in a policy condition, when tag-based enforcement belongs to Lake Formation.

69
MCQmedium

A company needs to transform JSON data from an Amazon S3 bucket into Parquet format and load it into an Amazon Redshift cluster. The transformation includes joining with a reference table stored in Amazon RDS. Which AWS service is BEST suited for this task?

A.AWS Data Pipeline
B.AWS Glue ETL job
C.Amazon Athena
D.Amazon EMR with Spark
AnswerB

AWS Glue ETL jobs run Apache Spark, so they can read JSON from S3, join it against the RDS reference table via a JDBC connection, and write Parquet into Redshift. This satisfies the transformation and cross-source join requirement in one managed job.

Why this answer

(AWS Glue ETL job) is the best choice because it natively integrates with S3, RDS, and Redshift. Glue can read JSON from S3, connect to RDS via JDBC to join with the reference table, transform the data to Parquet using its built-in converter, and write directly to Redshift. Option A (AWS Data Pipeline) is older and less integrated for this purpose.

Option C (Amazon Athena) can query S3 and convert to Parquet but cannot natively join with RDS without additional services. Option D (Amazon EMR with Spark) is possible but requires more setup and maintenance.

70
MCQmedium

A company uses Amazon Kinesis Data Firehose to ingest application logs into an Amazon S3 bucket. The logs are in JSON format. The data engineering team wants to convert the logs from JSON to Parquet format before landing in S3. What is the most cost-effective way to achieve this?

A.Use Amazon Athena to query the JSON data and write results in Parquet format.
B.Configure the Firehose delivery stream to convert the data to Parquet using a schema from AWS Glue.
C.Use an AWS Lambda function to transform each record to Parquet and send to Firehose.
D.Use an AWS Glue ETL job to run on a schedule and convert JSON to Parquet in S3.
AnswerB

Firehose performs in-flight format conversion to Parquet using an AWS Glue Data Catalog schema, so no separate ETL job or compute cluster is needed. This satisfies the cost-effectiveness constraint by eliminating additional processing infrastructure while landing data directly in S3.

Why this answer

Amazon Kinesis Data Firehose supports record format conversion natively, allowing a delivery stream to convert incoming JSON records to Parquet or ORC using a schema stored in the AWS Glue Data Catalog. This is the most cost-effective approach because the conversion happens within Firehose without provisioning or paying for separate compute such as Lambda or Glue ETL jobs.

Exam trap

The trap is overlooking Firehose's built-in record format conversion and instead choosing Lambda or Glue ETL, which are more expensive and complex; candidates must know that Firehose can convert JSON to Parquet natively using a Glue schema.

How to eliminate wrong answers

Option A is wrong because Athena is a query service, not a transformation pipeline; using it to read JSON and write Parquet would require additional orchestration and cost, and it is not a Firehose-integrated conversion method. Option C is wrong because a Lambda function would need to parse and convert each record to Parquet, which is inefficient, adds latency, and incurs Lambda costs; Firehose already provides this capability. Option D is wrong because a scheduled Glue ETL job would run periodically and incur Glue DPU costs, and it would not convert data before landing in S3 as required; it would process data after it is already in S3.

71
MCQeasy

A company is storing large amounts of log data in Amazon S3. The data is accessed frequently for the first 30 days, then rarely after that. The company wants to automatically transition the data to a lower-cost storage class after 30 days. Which S3 feature should the data engineer use?

A.S3 Intelligent-Tiering
B.S3 Lifecycle policies
C.S3 Cross-Region Replication
D.S3 Batch Operations
AnswerB

Lifecycle policies can transition objects after a specified number of days.

Why this answer

S3 Lifecycle policies allow you to define rules that automatically transition objects between storage classes based on age or other criteria. In this scenario, a lifecycle rule can be configured to transition objects from S3 Standard to a lower-cost class like S3 Glacier Deep Archive after 30 days, directly meeting the requirement for automated cost optimization.

Exam trap

The trap here is that candidates confuse S3 Intelligent-Tiering's automatic cost optimization with the ability to enforce a fixed time-based transition, when in fact Intelligent-Tiering monitors access patterns and may not align with a strict 30-day policy.

How to eliminate wrong answers

Option A is wrong because S3 Intelligent-Tiering automatically moves data between access tiers based on changing access patterns, but it does not allow you to set a fixed 30-day transition rule; it monitors usage and may not transition data that is rarely accessed after exactly 30 days. Option C is wrong because S3 Cross-Region Replication is used to copy objects to a different AWS region for disaster recovery or compliance, not to transition objects to a lower-cost storage class within the same region. Option D is wrong because S3 Batch Operations is designed for bulk actions like copying, tagging, or restoring objects, not for automating storage class transitions based on time.

72
MCQmedium

A data engineer is designing a data lake on S3 and needs to ensure that data is encrypted at rest using customer-managed KMS keys. The engineer also needs to audit all access to the KMS keys. Which combination of services should be used?

A.SSE-KMS with AWS CloudTrail
B.SSE-C with CloudWatch Logs
C.SSE-KMS with S3 Inventory
D.SSE-S3 with S3 server access logs
AnswerA

SSE-KMS encrypts S3 objects with customer-managed KMS keys, satisfying the encryption-at-rest requirement. CloudTrail records every KMS API call, including Encrypt, Decrypt and GenerateDataKey, delivering the key-access audit trail. Together they meet both constraints: customer-managed keys plus auditable key usage.

Why this answer

SSE-KMS allows customer-managed KMS keys, and AWS CloudTrail logs all KMS API calls (e.g., Decrypt, GenerateDataKey), enabling auditing. Option B is incorrect because SSE-C uses customer-provided encryption keys, not KMS, and CloudWatch Logs is for application logs, not KMS access. Option C is incorrect because S3 Inventory provides object metadata but does not audit KMS access.

Option D is incorrect because SSE-S3 uses AWS-managed keys, not customer-managed, and S3 server access logs do not capture KMS API calls.

73
MCQhard

Refer to the exhibit. A data engineer is using a Kinesis Data Stream with 2 shards. The producer uses a partition key that is the user ID (a UUID). The consumer is falling behind. Which change would improve throughput?

A.Switch to Kinesis Data Firehose
B.Increase the number of shards
C.Increase the retention period
D.Change the partition key to a constant value
AnswerB

A stream's throughput ceiling is set by its shard count, and two shards cap parallel consumption. Adding shards increases that ceiling and spreads the UUID-keyed records across more consumers, letting the lagging consumer catch up.

Why this answer

The consumer is falling behind because the total throughput of the stream (1 MB/s or 1,000 records/s per shard for writes, and 2 MB/s per shard for reads) is insufficient for the incoming data volume. Increasing the number of shards scales both the write and read capacity linearly, allowing the consumer to process records faster and catch up. Changing the partition key or retention period does not increase throughput, and switching to Firehose changes the delivery model but does not inherently solve the consumer lag.

Exam trap

The trap here is that candidates may think changing the partition key to a constant value would simplify processing, but it actually destroys parallelism and reduces throughput to a single shard, making the lag worse.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Firehose is a fully managed delivery service that buffers and loads data into destinations like S3 or Redshift; it does not increase the read throughput for a consumer that is falling behind, and it removes the ability for custom consumers to process records in real time. Option C is wrong because increasing the retention period (default 24 hours, max 365 days) only keeps records longer in the stream; it does not increase the ingestion or consumption rate, so the consumer will still lag. Option D is wrong because changing the partition key to a constant value would cause all records to go to a single shard, drastically reducing throughput and making the lag worse, as the other shard would be idle.

74
Multi-Selecthard

A company needs to ingest data from a MySQL database into Amazon S3 using AWS DMS. The data changes frequently and the requirement is to capture changes in near real-time. Which THREE configurations are necessary?

Select 3 answers
A.Create a VPC endpoint for S3.
B.Create an S3 target endpoint in DMS.
C.Enable binary logging (binlog) on the MySQL source database.
D.Create an AWS DMS replication instance.
E.Configure an S3 event notification to trigger DMS.
AnswersB, C, D

DMS requires a target endpoint defining Amazon S3 as the destination, including bucket, folder and IAM role, before any task can write replicated data. Without it, the replication instance has nowhere to deliver the ingested MySQL changes.

Why this answer

Option B is correct because AWS DMS requires an explicitly defined target endpoint that describes the S3 bucket, folder, and IAM role used to write the migrated data. Option C is correct because ongoing change data capture (CDC) from MySQL relies on the source's binary log (binlog), which must be enabled with settings such as log_bin and binlog_format=ROW so DMS can read committed changes in near real-time. Option D is correct because a replication instance is the managed compute resource that runs the DMS migration and CDC tasks between the source and target endpoints.

Option A is not required, since a VPC endpoint for S3 is only an optional networking choice and not a mandatory DMS configuration. Option E is not required, because DMS tasks are started and monitored by DMS itself, not by S3 event notifications.

75
MCQhard

A company uses Amazon Kinesis Data Streams with a Lambda consumer. The Lambda function is failing with 'ProvisionedThroughputExceededException' when writing to a DynamoDB table. Which action should the data engineer take to resolve this without losing data?

A.Reduce the number of Kinesis shards to lower the ingestion rate.
B.Increase the DynamoDB table's read capacity.
C.Configure a dead-letter queue (DLQ) on the Lambda function and increase the DynamoDB write capacity.
D.Disable retries on the Lambda function to avoid throttling.
AnswerC

The DLQ captures failed records so Kinesis retries do not discard them, while raising DynamoDB write capacity removes the throttling that triggers ProvisionedThroughputExceededException. Together they satisfy the no-data-loss constraint, though the DLQ alone would not fix the underlying capacity shortfall.

Why this answer

The Lambda function is throttled by DynamoDB because the write capacity is insufficient for the incoming Kinesis stream rate. Increasing the DynamoDB write capacity resolves the throttling, and configuring a DLQ on the Lambda function ensures that any records that fail after retries are captured for later processing, preventing data loss. This combination directly addresses the root cause while providing a safety net.

Exam trap

DEA-C01 often tests the misconception that increasing read capacity resolves write throttling, or that disabling retries prevents throttling, when in fact it causes data loss.

How to eliminate wrong answers

Option A is wrong because reducing Kinesis shards would lower the ingestion rate but does not address the DynamoDB write capacity issue and could cause data loss if the stream is scaled down. Option B is wrong because the error is ProvisionedThroughputExceededException on writes, so increasing read capacity has no effect. Option D is wrong because disabling retries would cause immediate data loss on throttling, worsening the problem.

Page 1 of 18

Page 2