Courseiva

CCNA Data Operations and Support Questions

75 of 270 questions · Page 2/4 · Data Operations and Support · Answers revealed

76
MCQmedium

A data engineer is designing a data pipeline that processes sensitive personal data. The data is ingested via Amazon Kinesis Data Firehose and stored in Amazon S3. The pipeline must ensure that the data is encrypted at rest and in transit. The engineer also needs to audit access to the data. Which combination of services meets these requirements?

A.AWS KMS for encryption at rest, Kinesis Data Analytics for in-transit encryption, and AWS CloudTrail for auditing.
B.AWS KMS for encryption at rest, Amazon CloudWatch Logs for auditing, and TLS for in-transit encryption.
C.S3 server-side encryption (SSE-S3) for at-rest encryption, HTTPS for in-transit encryption, and AWS CloudTrail for auditing.
D.S3 client-side encryption, AWS Config for auditing, and TLS for in-transit encryption.
AnswerC

SSE-S3 encrypts objects at rest, HTTPS secures data in transit from Firehose to S3, and CloudTrail records API activity for auditing. Together these three satisfy the encryption and access-audit requirements for the sensitive personal data pipeline.

Why this answer

Option C correctly combines S3 server-side encryption (SSE-S3) for data at rest, HTTPS (TLS) for data in transit, and AWS CloudTrail for auditing access. SSE-S3 provides AES-256 encryption managed by S3, HTTPS ensures secure ingestion and retrieval, and CloudTrail logs all API calls to S3, enabling audit trails. This meets all three requirements: encryption at rest, encryption in transit, and auditability.

Exam trap

DEA-C01 often tests the confusion between services that provide auditing (CloudTrail) versus monitoring (CloudWatch) or compliance (Config), and between encryption mechanisms for data at rest (SSE-S3, SSE-KMS) versus in transit (TLS/HTTPS).

How to eliminate wrong answers

Option A is wrong because Kinesis Data Analytics is an analytics service, not an encryption mechanism for in-transit data; in-transit encryption for Firehose is handled by HTTPS/TLS, not Kinesis Data Analytics. Option B is wrong because Amazon CloudWatch Logs is for monitoring and logging application/system metrics, not for auditing access to S3 data; auditing requires CloudTrail. Option D is wrong because AWS Config is a configuration compliance service, not an audit trail for data access; CloudTrail is the correct service for auditing access.

77
MCQmedium

A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format, and the engineer needs to convert it to Parquet before storage to optimize Athena queries. The Firehose delivery stream is configured with an AWS Lambda function for record transformation. However, the transformed data is still in JSON format in S3. What is the likely cause?

A.The Firehose delivery stream is not configured with the 'Convert record format' option to convert JSON to Parquet using an AWS Glue table.
B.The Lambda function is missing the necessary IAM permissions to write Parquet files to S3.
C.The Lambda function is not correctly converting the JSON to Parquet; it must output Parquet-formatted records.
D.The S3 bucket is configured with an S3 Lifecycle policy that converts objects to Parquet after a certain period.
AnswerA

Kinesis Data Firehose supports record format conversion from JSON to Parquet or ORC using an AWS Glue Data Catalog table. This feature must be explicitly enabled in the delivery stream settings. If it is not configured, the data remains in its original JSON format regardless of any Lambda transformation. The engineer needs to enable this option and specify the Glue table schema.

Why this answer

Kinesis Data Firehose can convert JSON to Parquet using the record format conversion feature, which relies on an AWS Glue table for schema. This option must be enabled in the delivery stream; otherwise, data remains in JSON even if a Lambda function transforms records. The Lambda function cannot perform the conversion itself, so enabling the Firehose setting is the correct fix.

Exam trap

The trap here is believing that a Lambda transformation function can output Parquet directly, when actually Firehose's built-in format conversion must be enabled.

78
MCQmedium

A data engineer manages an AWS Glue job that processes JSON files from Amazon S3 and writes Parquet to another S3 location. The job intermittently fails with 'Unable to find catalog table' errors, even though the table exists in the Glue Data Catalog. The job's IAM role has full S3 access but only limited Glue permissions. Which action will resolve the failure with the LEAST privilege?

A.Enable AWS Glue Data Catalog encryption and update the job's security configuration.
B.Add glue:GetTable and glue:GetDatabase permissions for the specific database and table to the job's IAM role.
C.Attach the AWSGlueServiceRole managed policy to the job's IAM role.
D.Modify the job to use a different IAM role that has s3:GetObject permissions on the catalog bucket.
AnswerB

The error indicates the job cannot retrieve table metadata from the Data Catalog. Granting glue:GetTable and glue:GetDatabase scoped to the exact database and table provides the necessary read access without excessive permissions. This aligns with least privilege and directly addresses the missing permission that causes the catalog lookup to fail during job execution.

Why this answer

The job fails because its IAM role cannot call Glue Data Catalog APIs to retrieve table metadata. The minimal fix is to grant glue:GetTable and glue:GetDatabase for the specific database and table. This resolves the error while adhering to least privilege, unlike broad managed policies or irrelevant S3 permissions.

Exam trap

The trap here is assuming that S3 permissions alone are sufficient for AWS Glue jobs, when in fact the Glue Data Catalog requires its own IAM permissions.

79
MCQmedium

An Amazon Kinesis Data Streams application is lagging behind. The data records are small (1 KB) and the shard count is 10. The consumer uses the KCL with default configuration. Which action will MOST effectively reduce the consumer lag?

A.Increase the number of KCL workers per shard (e.g., 2 workers per shard).
B.Use Enhanced Fan-Out to provide dedicated throughput.
C.Increase the number of shards to 20.
D.Reduce the record size by compressing the data.
AnswerB

Enhanced Fan-Out provides each consumer with dedicated throughput (2 MB/s per shard) and push-based delivery, which reduces latency and lag directly. This is the most effective action to reduce consumer lag given the scenario.

Why this answer

Enhanced Fan-Out provides each consumer with dedicated throughput (2 MB/s per shard) and eliminates the need for polling, which reduces latency and lag. Given that the consumer is lagging with default configuration and small records, Enhanced Fan-Out directly addresses the consumer-side bottleneck by providing dedicated read throughput, making it the most effective option. Option A is incorrect because the KCL does not support multiple workers per shard; each shard is processed by a single worker in a single-threaded manner.

Option C may help if the shard is saturated with incoming data, but the problem is consumer lag, not data ingestion. Option D reduces data size but does not address the consumer processing speed.

Exam trap

Candidates often assume that adding more shards (Option C) will always reduce lag, but the bottleneck is the consumer's processing capacity. Enhanced Fan-Out (Option B) is specifically designed to reduce consumer lag by providing dedicated throughput per consumer.

How to eliminate wrong answers

Option B is wrong because Enhanced Fan-Out provides dedicated throughput per consumer (up to 2 MB/s per shard per consumer), but the issue here is processing lag, not throttling or throughput limits—the default KCL already handles the 1 KB records easily, so dedicated throughput does not address the processing bottleneck. Option C is wrong because increasing shards to 20 would increase the number of parallel processing units, but each shard still has only one KCL worker by default, so the per-shard processing capacity remains unchanged; this would only help if the shard were overloaded with data, which is not the case with small records. Option D is wrong because compressing data reduces the size of records, but the records are already only 1 KB, and the bottleneck is processing time per record, not network or storage throughput; compression adds CPU overhead and does not reduce lag.

80
MCQmedium

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. One of the Glue jobs occasionally fails due to transient issues, such as a temporary network glitch. The engineer wants the Step Functions state machine to automatically retry that specific Glue job up to three times before moving to a failure state. Which Step Functions feature should be used to implement this?

A.Configure a Catch field to transition to a retry state on failure.
B.Use a Parallel state to run the Glue job multiple times concurrently.
C.Increase the Glue job's timeout and memory allocation.
D.Add a Retry field to the state that invokes the Glue job.
AnswerD

Step Functions states support a Retry field where you can specify the error types to retry, the interval between retries, the maximum number of attempts, and a backoff rate. This is exactly designed for transient failures. By configuring MaxAttempts to 3, the state will automatically retry the Glue job invocation three times before failing.

Why this answer

The Retry field in Step Functions is specifically designed to handle transient errors by automatically retrying a state. It allows you to specify the error types to retry (e.g., States.TaskFailed), the maximum number of attempts, the interval between retries, and a backoff rate. This is the standard way to implement retry logic for a Glue job invocation within a state machine.

Exam trap

The trap here is confusing error handling with retry logic; Catch handles errors by branching, while Retry actually re-executes the state.

81
MCQmedium

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is ongoing, and the engineer needs to ensure that changes made to the source database during the migration are replicated to the target. The engineer has set up a full load plus change data capture (CDC) task. However, after the full load completes, the CDC task fails with an error indicating that it cannot find the archive log files. What should the engineer do to resolve this issue?

A.Increase the number of AWS DMS replication instances to handle the CDC load.
B.Restart the AWS DMS task with the 'Stop task after full load completes' option enabled.
C.Configure the source Oracle database to retain archive log files for a longer period and ensure they are accessible to AWS DMS.
D.Change the target endpoint to Amazon RDS for Oracle to match the source database engine.
AnswerC

AWS DMS CDC requires access to Oracle archive log files to capture changes. If the logs are purged too quickly or not available, the CDC task fails. Increasing the retention period and ensuring accessibility allows DMS to read the logs and continue replication without interruption.

Why this answer

AWS DMS CDC for Oracle relies on reading archive log files. If these logs are not retained or accessible, the CDC task fails. The engineer must ensure the source database retains archive logs for a sufficient period and that DMS can access them, typically by configuring Oracle to keep logs and granting necessary permissions.

Exam trap

The trap here is thinking that scaling DMS resources or changing the target engine would fix a CDC error, when the error clearly points to missing archive logs on the source.

82
MCQhard

A data engineer is troubleshooting a failed AWS Glue ETL job that reads from a JDBC source. The error log shows 'java.sql.SQLException: Connection timed out'. The job previously ran successfully. Which of the following is the MOST likely cause?

A.The JDBC connection string has incorrect credentials.
B.The source database schema has changed.
C.The Glue job's timeout setting is too low.
D.The security group for the source database no longer allows traffic from the Glue job's IP range.
AnswerD

A security group rule change blocking the Glue job's ENI traffic causes the JDBC connection to time out, matching the sudden failure after prior success. Network ACLs or credentials would produce different errors, so the revoked inbound rule is the likely cause.

Why this answer

The error 'Connection timed out' indicates a network-level failure, not an authentication or schema issue. Since the job previously ran successfully, the most likely cause is that the security group for the source database no longer allows inbound traffic from the Glue job's IP range. AWS Glue ETL jobs run in a VPC with elastic network interfaces, and the security group rules must permit traffic on the JDBC port (e.g., 5432 for PostgreSQL, 3306 for MySQL).

Exam trap

AWS often tests the distinction between authentication errors (wrong credentials) and network connectivity errors (timeout), and candidates may confuse the Glue job timeout setting with a network timeout.

How to eliminate wrong answers

Option A is wrong because incorrect credentials would produce an authentication error (e.g., 'Access denied for user'), not a timeout. Option B is wrong because a schema change would cause a data type mismatch or column-not-found error, not a connection timeout. Option C is wrong because the Glue job's timeout setting controls how long the job can run before being terminated, not the network connection timeout to the JDBC source.

83
MCQmedium

A data engineer is tasked with designing a disaster recovery solution for a data lake stored in Amazon S3. The data lake contains sensitive customer data that must be replicated to a different AWS Region. The engineer needs to ensure that all objects, including those with encryption using SSE-KMS, are replicated. Which solution meets the requirements?

A.Use S3 Batch Operations to copy objects to the destination bucket.
B.Enable S3 Cross-Region Replication (CRR) with the appropriate KMS key and IAM role.
C.Use S3 Transfer Acceleration to copy objects across regions.
D.Use the AWS CLI s3 sync command scheduled in a cron job.
AnswerB

CRR replicates objects across Regions, but SSE-KMS-encrypted objects require the destination bucket's KMS key permissions and an IAM role allowing decrypt and encrypt. Configuring both satisfies the requirement that all objects, including KMS-encrypted ones, replicate successfully.

Why this answer

S3 Cross-Region Replication (CRR) with the appropriate KMS key and IAM role is the correct solution because it automatically replicates objects, including those encrypted with SSE-KMS, to a destination bucket in another region. To replicate SSE-KMS encrypted objects, you must specify a KMS key in the destination region and grant the necessary IAM permissions to the replication role. This meets the requirement for disaster recovery of sensitive data.

Exam trap

DEA-C01 often tests the misconception that S3 Transfer Acceleration or Batch Operations can serve as replication solutions, when in fact CRR is the only automated, continuous replication feature.

How to eliminate wrong answers

Option A is wrong because S3 Batch Operations can copy objects but does not provide continuous replication and requires manual initiation; it is not a disaster recovery solution with automatic replication. Option C is wrong because S3 Transfer Acceleration speeds up uploads to S3 over long distances but does not replicate data across regions; it is for accelerating transfers, not for disaster recovery. Option D is wrong because using AWS CLI s3 sync in a cron job is a manual, scripted approach that lacks the automation, reliability, and metadata preservation of CRR; it also does not handle SSE-KMS encryption seamlessly.

84
MCQmedium

A company uses Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The data is then processed by a scheduled AWS Glue ETL job that loads it into an Amazon Redshift table. Recently, the Glue job has been failing with the error: 'S3ServiceException: Access Denied'. The Firehose delivery stream is configured with a prefix and error logging to the same S3 bucket. The Glue job uses the same IAM role that has s3:GetObject and s3:ListBucket permissions on the bucket. What is the most likely cause?

A.The Glue job expects a different data format than what Firehose writes.
B.The Glue job's IAM role does not have s3:GetObjectVersion permission.
C.The Glue job is using the wrong IAM role that does not have permissions to the S3 bucket.
D.The S3 bucket has default encryption enabled with AWS KMS (SSE-KMS), and the Glue job's IAM role lacks kms:Decrypt permission.
AnswerD

With SSE-KMS default encryption, reading objects requires kms:Decrypt in addition to s3:GetObject. The Glue role holds only S3 permissions, so decryption fails and surfaces as Access Denied, despite the bucket and prefix configuration being correct.

Why this answer

The Glue job's IAM role has s3:GetObject and s3:ListBucket, which are sufficient for reading objects from S3 in the absence of encryption. However, when the bucket uses SSE-KMS default encryption, reading an object requires the kms:Decrypt permission on the KMS key in addition to s3:GetObject. The 'Access Denied' error is the classic symptom of a missing kms:Decrypt grant, making option D the most likely cause.

Exam trap

DEA-C01 often tests the misconception that s3:GetObject alone is sufficient to read encrypted objects, ignoring the separate KMS permission requirement for SSE-KMS.

How to eliminate wrong answers

Option A is wrong because a data format mismatch would produce parsing or schema errors (e.g., Glue DataFormatException), not an S3ServiceException: Access Denied. Option B is wrong because s3:GetObjectVersion is only needed when accessing specific object versions; the Glue job reads the current version, so its absence would not cause Access Denied. Option C is wrong because the scenario explicitly states the Glue job uses the same IAM role that has s3:GetObject and s3:ListBucket on the bucket, so the role is not the issue.

85
Multi-Selectmedium

A data engineer is troubleshooting an AWS Glue job that fails with 'java.lang.OutOfMemoryError: Java heap space'. The job processes a large dataset. Which TWO configuration changes should the engineer consider to resolve this issue? (Choose TWO.)

Select 2 answers
A.Change the output format from Parquet to CSV.
B.Increase the Spark shuffle partitions configuration (spark.sql.shuffle.partitions).
C.Reduce the number of partitions in the source data.
D.Increase the number of DPUs allocated to the Glue job.
E.Disable job bookmarks to avoid incremental processing.
AnswersB, D

Raising spark.sql.shuffle.partitions splits shuffle data into more, smaller partitions, so each task's working set fits within executor heap. This directly addresses the Java heap space exhaustion caused by oversized partitions during wide transformations on the large dataset.

Why this answer

Option B is correct because increasing spark.sql.shuffle.partitions creates more, smaller partitions during shuffles, which reduces the amount of data each executor task must hold in memory and helps prevent Java heap space exhaustion on large datasets. Option D is correct because allocating more DPUs adds more executors and memory to the Glue job, giving Spark more heap capacity to process the large dataset without running out of memory. Option A is incorrect because switching from Parquet to CSV increases data size and I/O, worsening memory pressure rather than relieving it.

Option C is incorrect because reducing source partitions concentrates more data per partition, increasing per-task memory usage and the risk of OOM. Option E is incorrect because disabling job bookmarks only affects incremental processing state and does not address heap memory consumption.

Exam trap

The trap is that candidates might think reducing partitions or changing output format would help, but these can worsen memory issues; the key is to increase resources and parallelism.

86
Multi-Selectmedium

A data engineer is designing a pipeline that ingests streaming data into Amazon S3 using Amazon Kinesis Data Firehose. The data must be delivered to S3 in Parquet format and partitioned by date. The engineer needs to configure the Firehose delivery stream. Which two actions are required to meet these requirements? (Choose two.)

Select 2 answers
A.Enable Amazon CloudWatch Logs for the Firehose delivery stream.
B.Set the Firehose buffer size to 128 MB and buffer interval to 900 seconds.
C.Enable record format conversion in the Firehose delivery stream and select Apache Parquet as the output format.
D.Configure dynamic partitioning with a JQ expression that extracts the date from each record.
E.Attach an IAM role to Firehose that allows s3:PutObject and glue:GetTable.
AnswersC, D

Firehose supports record format conversion from JSON to Parquet using an AWS Glue Data Catalog table for the schema. Enabling this feature and selecting Parquet as the output format is necessary to deliver data in columnar format, which is the requirement. Without it, Firehose delivers raw JSON.

Why this answer

To deliver Parquet, Firehose must use record format conversion, which requires selecting Parquet and referencing a Glue Data Catalog table for the schema. To partition by date, Firehose must use dynamic partitioning with a JQ expression that extracts the date from each record and includes it in the S3 prefix. Together, these two configurations satisfy the format and partitioning requirements.

Exam trap

The trap here is confusing supporting IAM or monitoring configurations with the specific features that enable Parquet conversion and dynamic partitioning.

87
MCQhard

A data engineer is using AWS Lake Formation to manage permissions on a data lake in Amazon S3. The engineer grants SELECT permission on a table to an IAM role used by an Amazon Athena user. However, the user reports that queries against the table return an 'Access Denied' error. The engineer verifies that the IAM role has the necessary S3 permissions and that Lake Formation permissions are correctly set. What is the most likely cause of the error?

A.The Athena workgroup has a query result location that the IAM role cannot access.
B.The AWS Glue Data Catalog is not integrated with Lake Formation.
C.The S3 bucket policy does not allow access from the Athena service.
D.The IAM role lacks the lakeformation:GetDataAccess permission.
AnswerD

AWS Lake Formation requires the IAM principal to have the lakeformation:GetDataAccess permission to obtain temporary credentials for accessing data. Even with SELECT permission on the table, without this permission, Athena cannot retrieve data from S3. This is a common oversight when configuring Lake Formation permissions.

Why this answer

When Lake Formation manages access, IAM principals need the lakeformation:GetDataAccess permission to obtain temporary credentials for data access. Without it, Athena cannot read data even if table permissions are granted. This permission is often overlooked because it is not part of standard S3 or Glue permissions.

Exam trap

The trap here is focusing on S3 bucket policies or Data Catalog integration, while the missing lakeformation:GetDataAccess permission is the subtle requirement for data access with Lake Formation.

88
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs and Amazon EMR steps as part of a nightly ETL pipeline. The pipeline occasionally fails due to transient issues such as Amazon S3 throttling or temporary network errors. The engineer wants to make the workflow more resilient without duplicating the entire state machine. Which Step Functions feature should be used to automatically retry failed states?

A.Configure Retry and Catch fields on individual states.
B.Enable X-Ray tracing on the state machine.
C.Set the state machine's TimeoutSeconds to a large value.
D.Use a Map state to iterate over each Glue job.
AnswerA

The Retry field on a state allows you to define automatic retries with backoff and max attempts for specific error types, while Catch can route to a fallback state after retries are exhausted. This directly addresses transient failures without modifying the state machine structure. It is the intended mechanism for handling intermittent errors such as S3 throttling in Step Functions workflows.

Why this answer

Step Functions provides Retry and Catch fields at the state level to handle transient errors. By configuring Retry with a backoff strategy and maximum attempts, the workflow can automatically re-execute a failed state, such as a Glue job start or EMR step, without manual intervention. Catch can then define a fallback if retries are exhausted.

This is the standard approach for making orchestration resilient to intermittent service issues.

Exam trap

The trap here is confusing observability features like X-Ray tracing or execution timeouts with actual error-handling and retry capabilities.

89
MCQhard

A data engineer is monitoring an Amazon Redshift cluster and notices that the 'WLM query wait time' metric is consistently high during peak hours. The cluster uses automatic WLM. The engineer wants to reduce query wait times without changing the cluster size. Which action is MOST effective?

A.Enable concurrency scaling.
B.Change WLM to manual mode and increase the number of queues.
C.Increase the maximum number of queries per queue.
D.Enable short query acceleration (SQA).
AnswerA

Concurrency scaling adds transient Redshift compute capacity automatically when queries queue, directly cutting WLM query wait time during peak hours. It absorbs burst concurrency without resizing the cluster, satisfying the constraint of reducing wait times at unchanged cluster size.

Why this answer

Enabling concurrency scaling (Option A) is the most effective action because it automatically adds transient cluster capacity during peak loads, allowing more queries to run concurrently without increasing wait times. This is specifically designed to reduce WLM query wait time. Option B (manual WLM) requires tuning and does not add capacity.

Option C (increasing max queries per queue) could increase concurrency but may lead to resource contention and longer wait times if the cluster is already saturated. Option D (short query acceleration) prioritizes short queries, which does not address overall wait times for all queries. Therefore, A is correct.

90
MCQhard

A data engineer is troubleshooting a DMS task that is replicating data from an on-premises Oracle database to an RDS for MySQL instance. The task is failing with 'ORA-1555: snapshot too old' error. What is the best course of action?

A.Disable full supplemental logging on the source tables.
B.Increase the size of the redo logs on the source database.
C.Enable batch optimized apply on the DMS task.
D.Increase the UNDO tablespace size and set UNDO_RETENTION to a higher value.
AnswerD

ORA-1555 occurs when Oracle cannot reconstruct a consistent read image because undo data was overwritten before the DMS read completed. Enlarging the UNDO tablespace and raising UNDO_RETENTION preserves that undo long enough for the long-running extraction.

Why this answer

ORA-1555 'snapshot too old' occurs when Oracle cannot reconstruct a consistent read image because the required undo data has been overwritten. Increasing the UNDO tablespace size and raising UNDO_RETENTION gives Oracle more undo history to retain, allowing the long-running DMS full-load read to complete without losing its snapshot. This directly addresses the root cause of insufficient undo retention on the source.

Exam trap

The trap is confusing redo logs (crash recovery) with undo tablespace (read consistency), leading candidates to increase redo log size when the actual fix is undo retention.

How to eliminate wrong answers

Option A is wrong because disabling supplemental logging would break change data capture for ongoing replication and does nothing to fix undo retention; supplemental logging is required for DMS CDC. Option B is wrong because redo logs protect against instance failure and support recovery, not consistent-read reconstruction — the snapshot-too-old error is an undo problem, not a redo problem. Option C is wrong because batch optimized apply is a target-side apply optimization for MySQL and does not affect the source Oracle undo retention that causes ORA-1555.

91
MCQeasy

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs that process data in Amazon S3. The engineer needs to ensure that if a Glue job fails, the entire workflow stops and an Amazon SNS notification is sent. Which Step Functions state type should be used to handle the error and send the notification?

A.A Choice state that evaluates the output of the Glue job and branches based on success or failure.
B.A Task state with a Catch field that transitions to a state that publishes to an Amazon SNS topic.
C.A Parallel state that runs the Glue job and an SNS notification task simultaneously.
D.A Task state with a Retry field that retries the Glue job up to three times.
AnswerB

The Catch field in a Task state is used to catch errors and transition to a fallback state. By configuring a Catch that leads to a state which publishes to an SNS topic, the engineer can stop the workflow and send a notification. This is the standard way to handle errors in Step Functions and meet the requirement.

Why this answer

In AWS Step Functions, the Catch field on a Task state allows you to define a fallback state when an error occurs. By transitioning to a state that publishes to an Amazon SNS topic, the workflow can stop and send a notification. This is the correct pattern for error handling and alerting in Step Functions.

Exam trap

The trap here is confusing error handling with retry logic or conditional branching; only the Catch field provides the mechanism to transition to a notification state on failure.

92
MCQmedium

A data engineer manages an AWS Glue ETL job that reads CSV files from Amazon S3 and writes Parquet to another S3 prefix. The job recently started failing with the error 'AnalysisException: Unable to infer schema for CSV.' The engineer confirms the S3 path contains files and the IAM role has s3:GetObject permissions. The job's script uses glueContext.create_dynamic_frame.from_catalog with a database and table name. What is the MOST likely cause?

A.The AWS Glue job's bookmarks are enabled, causing it to skip files that were already processed and leaving no data to infer schema from.
B.The IAM role lacks s3:ListBucket permission on the source bucket, so the job cannot list the CSV files.
C.The CSV files are compressed with gzip, and AWS Glue cannot infer schema from compressed files.
D.The AWS Glue Data Catalog table's SerDe parameters are misconfigured, so the crawler did not correctly identify the CSV delimiter or header.
AnswerD

The error 'Unable to infer schema for CSV' occurs when the Glue job reads from the Data Catalog and the table metadata lacks a valid SerDe or has incorrect parameters such as separatorType or header. The job relies on the catalog table definition, not direct file inspection, so misconfigured metadata prevents schema inference.

Why this answer

When an AWS Glue job uses create_dynamic_frame.from_catalog, it depends entirely on the AWS Glue Data Catalog table definition to determine the schema. If the table's SerDe parameters are incorrect—such as a wrong delimiter, missing header, or incorrect classification—the job cannot infer the schema and throws the AnalysisException. Verifying and repairing the catalog table resolves the failure.

Exam trap

The trap here is assuming that direct S3 file access or compression causes the schema inference error, when the job is actually reading metadata from the Data Catalog.

93
MCQeasy

A data engineer needs to transform a large dataset stored in Amazon S3 using Apache Spark. The engineer wants to minimize startup time and use a serverless approach. Which AWS service should the engineer use?

A.Amazon Redshift
B.Amazon EMR
C.AWS Glue
D.Amazon Athena
AnswerC

AWS Glue provides a serverless Apache Spark environment, so the engineer submits Spark jobs without provisioning or waiting for cluster startup. This directly satisfies the requirement to minimise startup time while avoiding infrastructure management, unlike Amazon EMR, which requires cluster provisioning.

Why this answer

AWS Glue provides a serverless Spark environment with fast startup. Option A is wrong because Amazon Redshift is a data warehouse, not a Spark environment. Option B is wrong because Amazon EMR requires cluster provisioning, which increases startup time.

Option D is wrong because Amazon Athena is for querying data, not for transforming with Spark.

94
Multi-Selecteasy

Which TWO actions are effective ways to monitor the health of an Amazon DynamoDB table? (Choose two.)

Select 2 answers
A.Use AWS S3 inventory to track table size.
B.Use EC2 instance status checks.
C.Enable DynamoDB Streams and process with Lambda to detect failures.
D.Set up Amazon CloudWatch alarms on ConsumedReadCapacityUnits.
E.Monitor the 'TableHealth' metric in CloudWatch.
AnswersC, D

DynamoDB Streams capture item-level changes in near real time; processing them with Lambda lets you detect write failures, anomalies or unexpected patterns. This satisfies the monitoring requirement by providing change-level visibility that capacity metrics alone cannot reveal.

Why this answer

Option C is correct because DynamoDB Streams captures item-level changes in near real time, and a Lambda consumer can inspect those records to detect anomalies or failed writes, providing an event-driven health signal for the table. Option D is correct because CloudWatch publishes DynamoDB metrics such as ConsumedReadCapacityUnits (and ConsumedWriteCapacityUnits), and alarms on these metrics reveal throttling or unexpected capacity consumption that indicates table health issues. Option A is wrong because S3 Inventory reports on S3 bucket objects, not DynamoDB table size or health.

Option B is wrong because EC2 instance status checks only reflect the health of EC2 instances, not a managed DynamoDB table. Option E is wrong because DynamoDB does not expose a 'TableHealth' metric in CloudWatch; health must be inferred from real metrics like throttled requests, consumed capacity, and latency.

Exam trap

DEA-C01 often tests the difference between monitoring tools for DynamoDB; candidates might incorrectly assume that EC2 status checks or S3 inventory apply, or invent non-existent metrics like 'TableHealth'.

95
Multi-Selectmedium

A company runs a data processing pipeline on Amazon EMR. The pipeline reads data from S3, processes it with Spark, and writes results back to S3. The engineer notices that the cluster is underutilized and wants to reduce costs. Which TWO actions should the engineer take? (Choose TWO.)

Select 2 answers
A.Use Spot instances for task nodes.
B.Configure the cluster to terminate after the job completes.
C.Change the master node to a larger instance type.
D.Enable EMRFS consistent view.
E.Increase the number of core nodes to improve parallelism.
AnswersA, B

Spot instances suit task nodes because EMR task nodes hold no HDFS data, so interruption only loses in-flight Spark tasks, which rerun. This directly addresses the underutilised cluster's cost problem by cutting compute spend on the elastic portion of the cluster.

Why this answer

Option A is correct because using Spot instances for task nodes is a standard EMR cost-optimization technique: task nodes perform only HDFS-non-persistent work (Spark executors), so they can be interrupted without data loss, and Spot capacity typically costs significantly less than On-Demand. Option B is correct because configuring the cluster to terminate after the job completes (a transient cluster, e.g., via --auto-terminate or a termination-protected=false setting) stops paying for idle EC2 instances once the batch pipeline finishes, directly addressing the underutilization. Option C is wrong because enlarging the master node does not improve processing throughput and increases cost, since the master only manages the cluster.

Option D is wrong because EMRFS consistent view is a data-consistency feature for S3 reads/writes, not a cost-reduction mechanism. Option E is wrong because adding core nodes increases cost and capacity rather than reducing spend, and the cluster is already underutilized.

Exam trap

The trap here is that candidates may confuse cost optimization features like Spot instances and auto-termination with performance improvements or data consistency settings, leading them to select options that increase resources or enable features unrelated to cost reduction.

96
MCQeasy

A data engineer uses AWS CloudTrail to investigate a security incident. The engineer runs the command shown in the exhibit. What does the output indicate?

A.A file was downloaded from the S3 bucket.
B.A file was deleted from the S3 bucket.
C.A batch of files was listed from the S3 bucket.
D.A file was uploaded to the S3 bucket.
AnswerD

The CloudTrail event records an S3 PutObject API call, indicating that a file was uploaded to the bucket. The event source and action name confirm the upload operation rather than a download, deletion or bucket-level configuration change.

Why this answer

The CloudTrail event shown is an S3 PutObject API call, which is the operation used to upload an object to an S3 bucket. The event name 'PutObject' and the presence of request parameters like 'bucketName' and 'key' confirm that a file was successfully written to the bucket. Therefore, the output indicates that a file was uploaded to the S3 bucket.

Exam trap

DEA-C01 often tests the ability to differentiate between S3 API operations based on CloudTrail event names, so candidates must memorize that PutObject corresponds to upload, GetObject to download, and DeleteObject to deletion.

How to eliminate wrong answers

Option A is wrong because downloading a file from S3 would generate a GetObject event, not PutObject. Option B is wrong because deleting a file from S3 would generate a DeleteObject event, not PutObject. Option C is wrong because listing files in an S3 bucket would generate a ListObjects or ListObjectsV2 event, not PutObject.

97
Multi-Selecthard

A data engineer is using AWS Step Functions to orchestrate an ETL workflow that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The workflow sometimes fails due to transient errors, and the engineer wants to implement a retry strategy that avoids duplicate data processing. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Implement idempotency in the ETL tasks so that retries do not produce duplicate data.
B.Configure the Step Functions execution to run in Express mode for faster retries.
C.Use a Catch block to redirect to a cleanup state that deletes partially written data before retrying the entire workflow.
D.Add a Retry policy on the individual Task states with a backoff rate and max attempts to handle transient failures.
E.Increase the timeout of the entire state machine to allow more time for tasks to complete.
AnswersA, D

Idempotency ensures that re-executing a task produces the same result without side effects, such as duplicate records. By designing Glue jobs, EMR steps, and Redshift procedures to be idempotent, retries from Step Functions will not cause data duplication. This is essential when combined with retry policies to safely handle transient failures without compromising data integrity.

Why this answer

To handle transient errors without duplicating data, the engineer should add Retry policies to Task states and ensure ETL tasks are idempotent. Retry policies automatically re-execute failed tasks, while idempotency guarantees that repeated executions do not create duplicate records. Together, they provide robust error handling and data integrity.

Exam trap

The trap here is thinking that simply increasing timeouts or switching to Express workflows will solve transient errors, when the real solution requires retries and idempotency.

98
MCQhard

Refer to the exhibit. An IAM policy is attached to an IAM role used by an application. The application needs to read objects from 'my-bucket' that have the tag 'classification=public'. The application account is 123456789012. However, the application is getting 'Access Denied' errors. What is the most likely reason?

A.The Deny statement uses StringNotEquals, which incorrectly denies the application account.
B.The policy does not grant s3:ListBucket permission, so the application cannot list objects.
C.The object being accessed does not have the tag 'classification=public'.
D.The Deny statement blocks all access from accounts other than 123456789012, but the application is in that account.
AnswerC

Tag-based access control evaluates the tag attached to the object itself, not the bucket. If the requested object lacks `classification=public`, the condition in the IAM policy fails and S3 returns Access Denied, even though the role's permissions and bucket policy are otherwise valid.

Why this answer

The most likely reason for Access Denied is that the object being accessed does not have the tag 'classification=public'. The policy's Allow statement is conditioned on s3:ExistingObjectTag/classification equaling 'public', so if the tag is missing or has a different value, the Allow does not match and the request is denied. The Deny statement in the policy is a separate guardrail and is not the cause here.

Exam trap

DEA-C01 often tests whether candidates read the policy conditions carefully — the trap is blaming the Deny statement or missing permissions when the real issue is that the object lacks the required tag, making the Allow condition false.

How to eliminate wrong answers

Option A is wrong because StringNotEquals in a Deny statement is a common pattern to block access from accounts other than the trusted one — it does not incorrectly deny the application account if the condition is structured correctly (e.g., denying when aws:PrincipalAccount != 123456789012). Option B is wrong because s3:ListBucket is only needed for listing operations; GetObject on a specific key does not require ListBucket, so its absence would not cause Access Denied for a read. Option D is wrong because the Deny statement is designed to block other accounts, not the application account — if the application is in 123456789012, the Deny does not apply to it.

99
MCQmedium

A data engineer manages an AWS Glue ETL job that processes JSON files from an S3 bucket and writes Parquet to another bucket. The job uses a Glue DynamicFrame with a specified schema. During execution, the job fails with the error: 'AnalysisException: cannot resolve column 'transaction_id' given input columns: [txn_id, amount, timestamp]'. The source data has a column named 'txn_id', but the Glue job's script references 'transaction_id'. The job's catalog table for the source points to the correct S3 location and has the correct schema. What is the most likely cause of this error?

A.The Glue job's IAM role lacks permission to read the S3 bucket, resulting in an empty DynamicFrame.
B.The Glue Data Catalog table for the source has an incorrect schema definition.
C.The Glue job's script uses a hardcoded schema that does not match the actual data columns.
D.The S3 bucket contains files with inconsistent schemas, causing Glue to infer a different column name.
AnswerC

The error indicates that the column 'transaction_id' is not present in the input data, while 'txn_id' exists. If the script applies a hardcoded schema or a mapping that expects 'transaction_id', the Spark engine cannot resolve it. This is a common issue when the script is written with a static schema that diverges from the actual source structure, even if the catalog is correct.

Why this answer

The error occurs because the Glue script references a column that does not exist in the source data. Even though the Glue Data Catalog is correct, the script may have been written with a hardcoded schema or mapping that expects 'transaction_id' instead of 'txn_id'. To resolve, the engineer should update the script to use the correct column name or apply a mapping to rename 'txn_id' to 'transaction_id'.

Exam trap

The trap here is assuming that the Glue Data Catalog schema is always used by the job, but the script may define its own schema or mappings that override the catalog.

100
MCQhard

A data engineer manages an Amazon Kinesis Data Stream with multiple shards. The stream is experiencing high throughput, and the engineer notices that some shards are throttling while others are underutilized. The engineer needs to redistribute the data evenly across shards to avoid throttling. Which action should the engineer take?

A.Use the MergeShards API to combine underutilized shards and reduce the total number of shards.
B.Modify the partition key strategy in the producer to ensure even distribution across shards.
C.Increase the number of shards using the UpdateShardCount API to add more capacity.
D.Enable enhanced fan-out for the stream to provide dedicated throughput for consumers.
AnswerB

Throttling on specific shards while others are underutilized indicates a skewed partition key. By changing the partition key strategy—such as adding a random suffix or using a more uniform key—the producer can distribute records more evenly across all shards, resolving the throttling without changing shard count.

Why this answer

Uneven shard utilization with throttling on some shards is typically caused by a partition key that does not distribute records uniformly. Changing the partition key strategy—for example, by incorporating a random value or using a composite key—ensures records are spread evenly across all shards. This addresses the root cause without altering shard count or read capacity.

Exam trap

The trap here is thinking that adding more shards or enabling enhanced fan-out will fix throttling, when the real issue is data distribution due to partition key skew.

101
MCQeasy

A data engineer needs to grant an AWS Lambda function permission to read objects from a specific Amazon S3 bucket. The Lambda function assumes an IAM role. Which policy should the engineer attach to the IAM role to allow the Lambda function to read objects from the bucket?

A.An IAM policy with s3:PutObject and s3:GetObject actions, with resource set to the bucket ARN and objects ARN.
B.An IAM policy with s3:ListBucket action and resource set to the objects ARN only.
C.An IAM policy with s3:GetObject and s3:ListBucket actions, with the resource set to the bucket ARN and objects ARN.
D.An IAM policy with s3:GetObject action and resource set to the bucket ARN only.
AnswerC

To read objects from an S3 bucket, the Lambda function needs s3:GetObject on the objects and s3:ListBucket on the bucket. The resource for s3:GetObject should be the object ARN (e.g., arn:aws:s3:::bucket/*) and for s3:ListBucket the bucket ARN (arn:aws:s3:::bucket). This policy grants the necessary permissions correctly.

Why this answer

To read objects from S3, the Lambda function requires s3:GetObject on the object ARN and s3:ListBucket on the bucket ARN. This combination allows listing the bucket and retrieving objects. Policies that omit the object ARN for GetObject or include unnecessary write permissions are incorrect.

The correct policy follows the principle of least privilege.

Exam trap

The trap here is confusing the resource ARN requirements for bucket-level versus object-level S3 actions, leading to policies that either do not work or grant excessive permissions.

102
Multi-Selectmedium

Which THREE are best practices for managing data in Amazon S3 for a data lake? (Choose three.)

Select 3 answers
A.Enable S3 Versioning to protect against accidental deletions.
B.Configure lifecycle policies to transition data to colder storage tiers.
C.Enable S3 Snapshot for point-in-time recovery.
D.Disable S3 server access logging to reduce costs.
E.Use bucket policies to restrict access based on IAM roles.
AnswersA, B, E

Versioning provides data protection.

Why this answer

Enabling S3 Versioning is a best practice for data lakes because it protects against accidental deletions or overwrites by preserving all versions of an object, including deletions (which are recorded as delete markers). This allows you to recover previous object states and is essential for data governance and auditability in a data lake environment.

Exam trap

The trap here is that candidates may confuse S3 Versioning with a non-existent 'S3 Snapshot' feature, or mistakenly think disabling server access logging is a cost-saving best practice, when in fact it undermines security auditing.

103
Drag & Dropmedium

Arrange the steps to set up a streaming ETL pipeline using Amazon Kinesis Data Firehose to Amazon S3.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First, create the Firehose stream, configure source, set S3 destination, enable optional Lambda transformation, and test.

104
MCQeasy

A data engineer needs to monitor the number of records processed by an AWS Glue ETL job. Which CloudWatch metric should the engineer use?

A.glue.driver.aggregate.elapsedTime
B.glue.driver.aggregate.numRecords
C.glue.driver.aggregate.bytesRead
D.glue.driver.aggregate.recordsRead
AnswerB

The glue.driver.aggregate.numRecords metric reports the total record count processed by an AWS Glue ETL job, aggregated at the driver. It directly answers the requirement to monitor how many records the job handles, unlike per-executor or duration metrics.

Why this answer

Glue emits a 'glue.driver.aggregate.numRecords' metric for the number of records processed. Option A is wrong because 'glue.driver.aggregate.elapsedTime' is for time. Option C is wrong because 'glue.driver.aggregate.bytesRead' is for bytes.

Option D is wrong because 'glue.driver.aggregate.recordsRead' is not a standard metric.

105
MCQhard

A data engineer is investigating intermittent failures in an AWS Step Functions state machine that orchestrates a nightly ETL workflow. The state machine invokes an AWS Glue job, then an Amazon EMR step, then an AWS Lambda function. Occasionally a task fails transiently and the entire workflow stops instead of retrying. The engineer needs the workflow to automatically retry failed tasks with exponential backoff before alerting. What should the engineer do?

A.Set the state machine's execution role to include the 'states:Retry' IAM action and re-run the workflow.
B.Configure the Glue job, EMR step, and Lambda function each with their own internal retry logic and remove error handling from the state machine.
C.Add a Retry field to the relevant state with ErrorEquals, IntervalSeconds, MaxAttempts, and BackoffRate values.
D.Enable the state machine's 'Retry on failure' setting in the Amazon CloudWatch console.
AnswerC

Step Functions supports a Retry field on any state, where you specify which errors to match and how many times to retry with an interval and a backoff multiplier. Adding Retry with ErrorEquals, IntervalSeconds, MaxAttempts, and BackoffRate implements automatic exponential-backoff retries within the state machine, satisfying the requirement without external tooling.

Why this answer

Step Functions state machines are defined in Amazon States Language, and each state can include a Retry field that matches specific error names and retries with an interval that grows by a BackoffRate multiplier. This declarative approach centralizes retry logic in the orchestrator, so transient failures in Glue, EMR, or Lambda are retried automatically before any alert is raised.

Exam trap

The trap here is assuming the retry policy lives in IAM or CloudWatch rather than being declared inside the state machine definition itself.

106
MCQeasy

A data engineer is using Amazon Athena to query data stored in Amazon S3. The engineer notices that queries are returning incorrect results, specifically missing some rows that are known to exist in the underlying data. The data is stored in Parquet format and is partitioned by date. The engineer runs a query with a WHERE clause on the date partition and finds that some dates are missing from the results. The S3 bucket contains folders for each date, but some folders are empty. What is the MOST likely cause of the missing rows?

A.Athena is using a stale metadata cache and needs to be refreshed.
B.The Parquet files are corrupted, causing Athena to skip them.
C.The Athena table is not configured with the correct partition projection settings.
D.The empty folders in S3 indicate that the data for those dates was never written or was deleted, so Athena correctly returns no rows for those dates.
AnswerD

If the S3 folders for certain dates are empty, there is no data for Athena to query. Athena reads files from S3; empty folders contain no files, so queries for those dates return no rows. This is expected behavior. The missing rows are due to missing data files, not an Athena configuration issue. The engineer should investigate why those folders are empty, possibly due to upstream ETL failures.

Why this answer

The missing rows correspond to dates for which the S3 folders are empty. Athena queries data directly from S3, so if there are no files in a partition folder, no rows will be returned for that partition. This is not an Athena misconfiguration; it is a data availability issue.

The engineer should check the upstream processes that write data to those partitions.

Exam trap

The trap here is assuming that Athena is malfunctioning when the underlying data is simply absent, leading to unnecessary troubleshooting of Athena settings.

107
Multi-Selecthard

A company runs a data processing pipeline using Amazon EMR with Spark. The pipeline reads from S3, processes data, and writes to S3. Recently, the job started failing with 'S3AccessDeniedException' even though the EMR role has appropriate S3 permissions. Which TWO actions should the data engineer take to resolve this issue? (Choose TWO.)

Select 2 answers
A.Enable S3 versioning on the bucket to allow multiple access methods.
B.Verify that the EMR service role has the necessary S3 permissions in IAM.
C.Disable S3 Block Public Access settings on the bucket.
D.Check the S3 bucket policy for explicit deny statements that may override the IAM role.
E.Ensure the EMR cluster is launched in a VPC with an S3 VPC endpoint.
AnswersB, D

The EMR service role is the identity Spark uses to call S3, so its attached IAM policy must explicitly allow the required s3:GetObject and s3:PutObject actions on the bucket and prefix. Confirming these permissions rules out an identity-based policy gap before investigating bucket-level controls.

Why this answer

Option B is correct because even when the EMR role appears to have S3 permissions, the actual attached IAM service role (EMR_DefaultRole or a custom EC2 instance profile for EMR) must explicitly allow the required s3:GetObject, s3:PutObject, s3:ListBucket, and related actions on the specific bucket and object ARNs; a missing or mis-scoped policy on the role produces exactly the S3AccessDeniedException seen here. Option D is correct because an explicit Deny in the S3 bucket policy always overrides any Allow granted through the IAM role, so a deny statement (for example, one enforcing TLS, VPC endpoint, or encryption conditions) would cause access failures despite the role having appropriate permissions. Option A is incorrect because S3 versioning only affects object version retention and does not grant or alter access permissions.

Option C is incorrect because disabling Block Public Access would only matter for public/anonymous access and would not fix a role-based access denial; it would also weaken security. Option E is incorrect because a missing S3 VPC endpoint causes connectivity/timeout issues, not an S3AccessDeniedException, and the error indicates the request reached S3 and was explicitly denied by permissions.

108
MCQmedium

A data engineer is using AWS Step Functions to orchestrate a complex ETL workflow that includes multiple AWS Glue jobs, Amazon EMR steps, and AWS Lambda functions. The engineer notices that on rare occasions, the entire workflow fails due to a transient error in one of the Lambda functions. The engineer wants to make the workflow more resilient without changing the overall architecture. Which approach is the MOST effective?

A.Enable AWS X-Ray tracing for the Lambda function to identify the root cause of the errors.
B.Modify the Lambda function code to catch exceptions and return a success status to avoid workflow failure.
C.Add a Retry state in the Step Functions state machine for the Lambda task with exponential backoff and a maximum number of attempts.
D.Configure the Lambda function to have a longer timeout and increase its memory size.
AnswerC

Adding a Retry state in Step Functions allows the workflow to automatically retry the Lambda function when it fails due to transient errors. Configuring exponential backoff and a maximum attempt count ensures that temporary issues are handled gracefully without manual intervention. This directly improves the workflow's resilience and is a best practice for handling transient failures.

Why this answer

Adding a Retry state in Step Functions with exponential backoff and a maximum attempts limit is the most effective way to handle transient errors in Lambda functions. It automatically retries the failed task, increasing the likelihood of success without manual intervention, and is a standard resilience pattern in Step Functions.

Exam trap

The trap here is thinking that increasing Lambda resources or enabling tracing will resolve transient errors, when the real solution is to implement automated retries.

109
Multi-Selectmedium

Which TWO actions should a data engineer take to optimize Amazon S3 query performance for Amazon Athena when dealing with large Parquet files? (Choose 2.)

Select 2 answers
A.Store data in a single large file without partitioning
B.Use GZIP compression on the Parquet files
C.Split large files into many small files
D.Optimize file sizes to be around 64 MB to 256 MB
E.Partition the data by frequently filtered columns
AnswersD, E

Athena splits Parquet scans across files; many small files add per-file overhead, while very large files limit parallelism. Targeting 64–256 MB balances scan parallelism against request overhead, directly improving query performance on large Parquet datasets in S3.

Why this answer

Option D is correct because Athena's performance degrades with very large Parquet files since a single file can only be read by one split; targeting roughly 64 MB to 256 MB per file allows Athena's split-based parallel reads to distribute work across more nodes. Option E is correct because partitioning data by frequently filtered columns (e.g., date, region) enables partition pruning, so Athena scans only the relevant S3 prefixes instead of the entire dataset, drastically reducing bytes scanned and query time. Option A is wrong because a single large unpartitioned file forces full-table scans and prevents parallelism.

Option B is wrong because Parquet already uses columnar compression internally, and GZIP is not splittable, which can actually hurt parallelism for large files. Option C is wrong because splitting into many tiny files creates excessive per-file overhead (metadata, open/close costs) and degrades Athena performance rather than improving it.

110
MCQhard

Refer to the exhibit. A data engineer is reviewing the configuration of an Amazon Redshift cluster. The engineer wants to ensure that the cluster can be restored to a point in time up to 35 days in the past. Based on the exhibit, what change is needed?

A.Increase the automated snapshot retention period to 35 days.
B.Change the cluster subnet group to a custom one.
C.Enable encryption on the cluster.
D.Increase the number of nodes to 6.
AnswerA

Automated snapshots are the mechanism that supports point-in-time restoration in Amazon Redshift. The retention period defaults to one day and can be extended to a maximum of 35 days, so raising it to 35 days directly satisfies the stated requirement to restore to any point within that window.

Why this answer

Amazon Redshift automated snapshots have a default retention period of 1 day, but can be configured up to 35 days. To restore to a point in time up to 35 days, the retention period must be increased to 35 days.

Exam trap

DEA-C01 often tests the default and maximum snapshot retention periods, leading candidates to assume the default is already 35 days or that other cluster settings affect retention.

How to eliminate wrong answers

Option B is wrong because the cluster subnet group affects network placement, not snapshot retention. Option C is wrong because encryption is for data security, not backup retention. Option D is wrong because the number of nodes affects performance and storage, not snapshot retention.

111
Multi-Selecthard

Which THREE are valid considerations when troubleshooting data loss in an AWS Glue ETL job? (Choose three.)

Select 3 answers
A.Job bookmarks may be skipping new data if not configured properly.
B.Server-side encryption is disabled on the S3 bucket.
C.The job timeout is set too low.
D.Dynamic frame transformations may drop rows with errors.
E.The mapping of source columns to target columns may be incorrect.
AnswersA, D, E

Job bookmarks track previously processed data using a state store. If misconfigured — for example, wrong transformation context or a changed source path — the bookmark may mark new files or rows as already processed, silently skipping them and producing apparent data loss in the Glue ETL output.

Why this answer

Option A is correct because AWS Glue job bookmarks track previously processed data; if a bookmark is misconfigured, stale, or reset, the job can skip new or changed records, causing apparent data loss. Option D is correct because Glue DynamicFrame transformations and write operations can drop or quarantine rows that fail schema or type validation, so records with errors may silently disappear unless error handling is configured. Option E is correct because an incorrect ApplyMapping or column mapping can write source columns into the wrong target fields or omit them, making data appear lost in the target.

Option B is not a valid consideration because S3 server-side encryption protects data at rest and does not cause records to be dropped during ETL processing. Option C is not a valid consideration because a low job timeout causes the job to fail or stop, which is a job failure rather than silent data loss.

Exam trap

DEA-C01 often tests whether candidates can distinguish data-loss causes from job-failure or security issues — options like low timeout or disabled encryption sound plausible but cause failures or posture problems, not silent data loss.

112
Multi-Selecteasy

A data engineer is setting up a data pipeline using Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data must be transformed using an AWS Lambda function before delivery. Which THREE steps are required to configure this?

Select 3 answers
A.Create a Lambda@Edge function in the same Region.
B.Create an AWS Lambda function that transforms the data.
C.Attach an IAM role to the Firehose delivery stream that grants permission to invoke the Lambda function.
D.Configure an S3 event notification to trigger the Lambda function when new data arrives.
E.Configure the Kinesis Data Firehose delivery stream to use the Lambda function as a data transformation source.
AnswersB, C, E

Firehose requires a Lambda function to perform record transformation; the function must exist and be selected in the delivery stream's processing configuration. Creating it is therefore a mandatory step before Firehose can invoke it on incoming records.

Why this answer

Option B is correct because Firehose data transformation requires an AWS Lambda function (a standard regional Lambda function, not Lambda@Edge) that receives batches of records and returns transformed records in the expected Firehose format. Option C is correct because the Firehose delivery stream assumes an IAM role, and that role must include lambda:InvokeFunction permission on the transformation function so Firehose can call it. Option E is correct because you must enable and select the Lambda function in the Firehose delivery stream's processing configuration (the Lambda transformation processor) so records are transformed before being written to Amazon S3.

Option A is wrong because Lambda@Edge runs at CloudFront edge locations for viewer/origin request and response events, not for Firehose record transformation. Option D is wrong because S3 event notifications trigger actions after objects land in S3; Firehose invokes the transformation Lambda itself, so no S3 event notification is needed.

Exam trap

DEA-C01 often tests the misconception that S3 event notifications are needed to trigger Lambda for Firehose transformation, but Firehose directly invokes Lambda; candidates must remember the direct integration.

113
MCQmedium

A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Aurora MySQL. After the migration, the data in Aurora is inconsistent with the source. The engineer needs to ensure ongoing replication with minimal downtime. Which solution should the engineer implement?

A.Use AWS Schema Conversion Tool (SCT) to convert the schema
B.Export the data from Oracle and import into Aurora using mysqldump
C.Configure a DMS task with change data capture (CDC)
D.Perform a full load migration again
AnswerC

Change data capture reads the source Oracle redo logs continuously, replicating only committed inserts, updates and deletes to Aurora MySQL after the initial full load. This satisfies the ongoing replication requirement with minimal downtime, since no bulk re-copy or application outage is needed to keep the target consistent.

Why this answer

Configuring a DMS task with Change Data Capture (CDC) enables ongoing replication of changes from the source Oracle database to the target Aurora MySQL database with minimal downtime, ensuring consistency. Option A is incorrect because AWS Schema Conversion Tool (SCT) only converts schema and does not handle data replication. Option B is incorrect because mysqldump provides a one-time export/import, not ongoing replication.

Option D is incorrect because performing a full load migration again would disrupt operations and would not capture ongoing changes.

114
MCQhard

A data engineer is monitoring an AWS Glue job that reads from an Amazon S3 bucket and writes to Amazon Redshift. The job has been running for 2 hours, which is longer than usual. The engineer checks the Glue job's metrics and sees that the number of active executors is high, but the job is not making progress. The engineer suspects a data skew issue. Which action should the engineer take to diagnose and mitigate the skew?

A.Use the Spark UI to inspect the stage details and identify tasks that are taking significantly longer than others.
B.Increase the number of DPUs for the Glue job to add more executors and reduce the impact of skew.
C.Enable AWS Glue job metrics and examine the 'glue.driver.aggregate.bytesRead' metric to identify skewed partitions.
D.Modify the Glue job to use the 'groupFiles' option to combine small files and improve parallelism.
AnswerA

The Spark UI provides detailed information about stages and tasks, including task duration and input size. By examining the stage details, you can identify tasks that are processing much more data than others, indicating data skew. This is the most direct way to diagnose skew in a Glue job.

Why this answer

Data skew in Spark jobs manifests as a few tasks taking much longer than others. The Spark UI is the primary tool to diagnose skew by showing task-level metrics. Once identified, mitigation strategies include salting keys, using broadcast joins, or repartitioning.

The correct action here is to use the Spark UI to inspect stage details.

Exam trap

The trap here is assuming that adding more resources (DPUs) will solve skew, when in fact skew requires redistributing data, not just adding compute.

115
MCQmedium

A data engineer is using Amazon Kinesis Data Streams to ingest real-time data. The stream has 4 shards and is receiving 2 MB/s of data. The engineer notices that the WriteProvisionedThroughputExceeded metric is increasing. The engineer wants to resolve this issue with minimal changes. What should the engineer do?

A.Increase the retention period of the stream to 168 hours.
B.Enable enhanced fan-out for the stream.
C.Increase the number of shards in the stream to 8.
D.Use a more random partition key to distribute data evenly across shards.
AnswerD

Correct. With 4 shards, the stream can handle up to 4 MB/s (1 MB/s per shard). Since the incoming rate is 2 MB/s, the throughput exceedance suggests that data is not evenly distributed, causing some shards to hit their 1 MB/s limit. Using a more random partition key, such as a UUID, distributes records evenly, resolving the issue without adding shards.

Why this answer

The WriteProvisionedThroughputExceeded metric indicates that some shards are exceeding their 1 MB/s write capacity. With 4 shards, total capacity is 4 MB/s, which is sufficient for 2 MB/s if evenly distributed. The issue is likely hot shards due to a non-uniform partition key.

Using a more random partition key, such as a UUID, ensures even distribution and resolves the throttling without adding shards.

Exam trap

The trap here is assuming that adding shards is always the first step, when the actual problem might be uneven data distribution across existing shards.

116
Multi-Selectmedium

A company uses Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format, and each record is approximately 5 KB. The company has set the buffer interval to 60 seconds and the buffer size to 5 MB. However, the data engineer observes that the delivery to S3 is delayed by up to 5 minutes during peak traffic. The engineer wants to reduce the delivery latency to under 1 minute. Which TWO actions should the engineer take? (Choose TWO.)

Select 2 answers
A.Enable GZIP compression for the delivery stream.
B.Reduce the buffer size to 1 MB.
C.Increase the buffer size to 50 MB.
D.Convert the data format to Apache Parquet before delivery.
E.Reduce the buffer interval to 10 seconds.
AnswersB, E

Firehose flushes when either the buffer size or buffer interval is reached, whichever comes first. Lowering the size to 1 MB means the 5 MB threshold is hit sooner, triggering delivery earlier during peak traffic and cutting latency below one minute.

Why this answer

Option B is correct because Kinesis Data Firehose flushes data to S3 when either the buffer size OR the buffer interval is reached; lowering the buffer size from 5 MB to 1 MB makes the size threshold trigger much sooner, so smaller batches are delivered more frequently and latency drops. Option E is correct because reducing the buffer interval from 60 seconds to 10 seconds forces a flush at least every 10 seconds even if the buffer size is not reached, directly capping delivery latency well under 1 minute. Together, B and E ensure the delivery stream flushes on the shorter of the two thresholds, which is exactly what is needed to cut the observed 5-minute delays.

Option A does not help because GZIP compression only reduces payload size and can actually add CPU overhead, not reduce flush latency. Option C is wrong because increasing the buffer size to 50 MB would delay flushes further, worsening latency. Option D is wrong because converting to Parquet changes the storage format and query efficiency, not the Firehose buffering/flush timing that governs delivery latency.

Exam trap

The trap is assuming compression or format conversion (GZIP, Parquet) reduces delivery latency, when only the buffer size and buffer interval thresholds control when Firehose actually flushes data to S3.

117
MCQhard

A company runs a data warehouse on Amazon Redshift. The data engineer notices that some queries are running slowly. Upon reviewing the system tables, the engineer finds that the 'svv_table_info' shows high 'unsorted' percentage for several large tables. What is the MOST effective action to improve query performance?

A.Run the ANALYZE command on the tables.
B.Run the VACUUM command on the tables.
C.Change the distribution style of the tables to ALL.
D.Increase the number of nodes in the Redshift cluster.
AnswerB

Running VACUUM re-sorts rows into the table's defined sort key order, directly reducing the high unsorted percentage that svv_table_info reports. Because those large tables were loaded without sorted data, range-restricted scans read far more blocks than necessary; restoring sort order lets Redshift apply zone-map block pruning, cutting I/O for the slow queries.

Why this answer

A high 'unsorted' percentage in svv_table_info means many rows were added after the last sort and are not in sort-key order, forcing Redshift to scan more blocks and slowing range-restricted queries. VACUUM re-sorts the table and reclaims space, restoring the benefits of the sort key. This is the direct, most effective fix for the reported symptom.

Exam trap

DEA-C01 often tests the distinction between VACUUM (re-sorts and reclaims space) and ANALYZE (updates statistics), tricking candidates into choosing ANALYZE when the symptom is a high unsorted percentage.

How to eliminate wrong answers

Option A is wrong because ANALYZE only refreshes table statistics for the query planner; it does not re-sort rows or reduce the unsorted percentage. Option C is wrong because changing distribution style to ALL replicates the entire table to every node, which increases storage and load time and does not address unsorted data. Option D is wrong because adding nodes increases compute capacity but does not fix the underlying unsorted data that causes inefficient block scans.

118
Multi-Selectmedium

A data engineer is troubleshooting a slow-running Amazon Athena query. The query scans a large amount of data. Which TWO actions can improve query performance? (Choose TWO.)

Select 2 answers
A.Convert the data to Parquet or ORC format.
B.Enable encryption at rest.
C.Increase the Athena query timeout.
D.Partition the table on frequently filtered columns.
E.Use SELECT * to retrieve all columns.
AnswersA, D

Parquet and ORC are columnar formats, so Athena reads only the columns referenced by the query rather than every field in each row. This reduces the bytes scanned from S3, which is the dominant cost for the large scans described in the stem.

Why this answer

Option A is correct because converting data to columnar formats like Parquet or ORC lets Athena read only the columns referenced in the query and benefits from compression and predicate pushdown, drastically reducing the bytes scanned and thus improving performance and lowering cost. Option D is correct because partitioning the table on frequently filtered columns (e.g., date or region) enables partition pruning, so Athena skips scanning irrelevant partitions instead of reading the entire dataset. Option B is incorrect because encryption at rest protects stored data but does not reduce the amount of data scanned or speed up query execution.

Option C is incorrect because increasing the query timeout only allows a slow query to run longer; it does not make the query faster. Option E is incorrect because using SELECT * retrieves all columns, which increases the data scanned and worsens performance, especially with columnar formats.

119
MCQhard

A data engineer is troubleshooting a slow Amazon Redshift query. The query scans a large table with interleaved sort keys. The engineer notices that the query plan shows a sequential scan instead of a range-restricted scan. What is the MOST likely reason?

A.The table has not been vacuumed and reindexed after large data loads.
B.The table has a poor distribution key (DISTKEY) causing data skew.
C.The table uses compression encodings that prevent range-restricted scans.
D.The workload management (WLM) queue is configured with too few query slots.
AnswerA

Interleaved sort keys rely on accurate zone maps; after large loads, unsorted or deleted rows degrade them, so Redshift cannot skip blocks and falls back to a sequential scan. Vacuum and reindex restore the range-restricted scan.

Why this answer

Interleaved sort keys in Amazon Redshift rely on zone maps that must be refreshed via VACUUM REINDEX after significant data changes. Without reindexing, the zone maps become stale and the query planner cannot skip blocks, forcing a sequential scan instead of a range-restricted scan. Running VACUUM REINDEX rebuilds the interleaved sort metadata so range predicates can prune blocks effectively.

Exam trap

The trap is blaming distribution or compression for scan-type problems; the exam tests whether you know that interleaved sort keys specifically require VACUUM REINDEX to keep zone maps usable.

How to eliminate wrong answers

Option B is wrong because a poor DISTKEY causes data skew and uneven node utilization, but it does not prevent range-restricted scans — that is a sort-key/zone-map concern. Option C is wrong because compression encodings (e.g., AZ64, ZSTD) do not block range-restricted scans; Redshift can filter on compressed blocks using zone maps. Option D is wrong because WLM queue slot configuration affects query concurrency and queuing, not the planner's ability to perform range-restricted scans.

120
MCQhard

A data engineer is using AWS Glue job bookmarks to process incremental data from Amazon S3. The job reads from a partitioned S3 path and writes to Amazon Redshift. After a recent run, the engineer notices that some new partitions were not processed. The job bookmark state shows that the job has already processed up to a certain timestamp. What is the most likely reason for the missing partitions?

A.The new partitions were created with a timestamp earlier than the last processed timestamp.
B.The Glue job bookmark was not enabled for the S3 data source.
C.The Redshift cluster's copy command does not support partitioned data.
D.The Glue job's IAM role lacks permissions to read the new partitions.
AnswerA

Glue job bookmarks track the last processed timestamp or partition. If new data arrives with an older timestamp (e.g., late-arriving data), it may be considered already processed and skipped. This is a common issue with bookmarks when data is not strictly ordered by time. The engineer should adjust the bookmark logic or handle late data.

Why this answer

Glue job bookmarks use timestamps or partition values to determine what data has already been processed. If new partitions have timestamps older than the last processed timestamp, they are considered already processed and skipped. This often happens with late-arriving data.

The engineer should either adjust the bookmark to use partition-based tracking or handle late data separately.

Exam trap

The trap here is assuming that all new data will be processed regardless of timestamp, but bookmarks rely on ordering and can miss late-arriving partitions.

121
MCQmedium

A data engineer is using AWS Glue to run a PySpark ETL job that processes millions of small JSON files stored in an Amazon S3 bucket. The job is experiencing high memory usage and failing with 'Container killed by YARN for exceeding memory limits'. The engineer wants to optimize the job to handle the data more efficiently without changing the source data format. Which solution will MOST effectively reduce memory usage and improve performance?

A.Convert the JSON files to Parquet format using an AWS Glue crawler before running the ETL job.
B.Use AWS Glue's groupFiles and groupSize options to combine small files into larger chunks during the read operation.
C.Increase the number of DPUs allocated to the AWS Glue job to provide more memory per executor.
D.Enable AWS Glue job bookmarks to track processed files and avoid reprocessing.
AnswerB

The groupFiles and groupSize options in AWS Glue allow the job to coalesce multiple small files into larger groups, reducing the number of tasks and memory overhead per file. This directly addresses the inefficiency of processing many small files by minimizing the number of Spark partitions and the associated memory pressure, leading to better performance and lower resource consumption.

Why this answer

The groupFiles and groupSize options in AWS Glue are specifically designed to handle large numbers of small files by grouping them into larger chunks during read, which reduces the number of Spark partitions and memory overhead. This optimization directly targets the root cause of the memory failures and improves job efficiency without altering the source data format.

Exam trap

The trap here is assuming that increasing DPUs or enabling bookmarks will solve the memory issue, when the real problem is the inefficiency of processing many small files.

122
MCQhard

A data engineer is troubleshooting an AWS Glue job that writes data to an Amazon S3 bucket in Parquet format. The job runs successfully but the output files are smaller than the configured 'groupFiles' size. The engineer has set 'groupFiles' to 'inPartition' and 'groupSize' to 1 GB. The input data is 10 GB in a single partition. What is the most likely reason for the small files?

A.The 'groupFiles' parameter is deprecated in the current Glue version.
B.The 'groupFiles' parameter only affects the input read phase, not the output write phase.
C.The 'groupFiles' parameter is misspelled or set incorrectly.
D.The engineer must also set 'repartition' to 1 to merge output files.
AnswerB

Grouping coalesces small input files during reading but does not control output file size.

Why this answer

'groupFiles' only works when the input data is already small and needs to be coalesced. However, if the input is large and the job writes output, the output file size is determined by the number of Spark partitions, not grouping. The grouping feature only applies to reading input files.

Option A is wrong because the setting is correct. Option C is wrong because grouping is a read-time feature, not write-time. Option D is wrong because grouping does not require repartitioning.

123
MCQmedium

A data engineer sees the CloudWatch log entry in the exhibit for a Lambda function that processes data from an Amazon SQS queue. What is the MOST likely cause of the timeout?

A.The Lambda function's reserved concurrency is set too low.
B.The Lambda function is running out of memory.
C.The Lambda function's timeout is too short for the processing required.
D.The SQS queue's visibility timeout is set too low.
AnswerC

Lambda terminates execution once the configured timeout elapses, so a function whose processing exceeds that limit fails mid-batch. The SQS-triggered invocation needs enough duration for the workload; extending the timeout directly addresses the premature termination shown in the log.

Why this answer

The CloudWatch log entry shows a Lambda invocation timing out, which directly indicates the function's configured timeout is shorter than the time needed to process the SQS message. Lambda timeouts are a hard limit (max 15 minutes), and when exceeded, the invocation is terminated and the message returns to the queue. Increasing the timeout (or optimizing the code) resolves the issue.

Exam trap

DEA-C01 often tests the confusion between Lambda timeout, SQS visibility timeout, and concurrency throttling — candidates must read the log message carefully to distinguish a timeout from a throttling or memory error.

How to eliminate wrong answers

Option A is wrong because low reserved concurrency causes throttling (429 errors) and messages to remain in the queue, not invocation timeouts — the log would show throttling, not a timeout. Option B is wrong because running out of memory produces an 'OutOfMemoryError' or 'Runtime exited with error: signal: killed' message, not a timeout. Option D is wrong because a low SQS visibility timeout causes duplicate processing (messages reappear before processing completes), not a Lambda timeout — the Lambda would still complete its invocation.

124
MCQhard

Refer to the exhibit. A data engineer runs the command on an object in S3. The engineer expected the object to have a tag 'type=raw' but sees no metadata. What is the likely cause?

A.Object tags are not returned by head-object; use get-object-tagging instead
B.The S3 bucket is in a different AWS Region
C.The bucket policy blocks reading tags
D.The object was created without tags because of lifecycle rules
AnswerA

HeadObject returns object metadata and system headers, not the object's tag set, so the engineer sees no 'type=raw' tag despite it existing. GetObjectTagging is the API that retrieves tags, satisfying the expectation of confirming the object's tagging.

Why this answer

The head-object command does not return object tags; you must use the get-object-tagging command to retrieve tags. Option B is incorrect because the head-object command succeeds regardless of region, and region does not affect tag visibility. Option C is incorrect because bucket policies can deny access but do not prevent tags from being returned by head-object; they would affect get-object-tagging instead.

Option D is incorrect because lifecycle rules do not remove tags from objects; they may transition or expire objects but do not strip metadata.

125
MCQhard

A data engineer is troubleshooting an AWS Glue ETL job that suddenly started failing with 'An error occurred while calling o103.pyWriteDynamicFrame. Unknown error'. The job writes data to an Amazon Redshift table. Which step should the engineer take FIRST?

A.Recreate the Redshift table with a different distribution style.
B.Test the job with a small sample dataset to isolate the issue.
C.Update the Redshift JDBC driver version in the Glue job.
D.Review the job's CloudWatch Logs for detailed error messages.
AnswerD

The error message 'An error occurred while calling o103.pyWriteDynamicFrame. Unknown error' is generic and does not specify the root cause. The first troubleshooting step should be to review the job's CloudWatch Logs, which provide detailed error messages, stack traces, and other diagnostic information.

Why this answer

The error message 'An error occurred while calling o103.pyWriteDynamicFrame. Unknown error' is a generic wrapper thrown by the Glue DynamicFrame writer when the underlying cause is not surfaced. The first and most reliable step is to inspect the job's CloudWatch Logs, where the full stack trace, Redshift error code, and driver-level messages are recorded.

This aligns with standard AWS troubleshooting guidance: always check logs before making changes. Reviewing logs is non-destructive, fast, and directly reveals the root cause (e.g., permission issue, connection timeout, schema mismatch).

Exam trap

DEA-C01 often tests the principle of 'logs first' in troubleshooting scenarios, where candidates are tempted to jump to speculative fixes like changing configurations or drivers instead of gathering diagnostic data.

How to eliminate wrong answers

Option A is wrong because changing the distribution style is a speculative fix that does not address the immediate need to diagnose the error; it could also introduce performance regressions. Option B is wrong because testing with a small dataset might reproduce the error but does not provide the detailed diagnostic information that logs offer, and it consumes additional time and resources. Option C is wrong because updating the JDBC driver is a potential solution only after confirming a driver incompatibility, which the logs would reveal; doing it first is premature and may not resolve the issue.

126
MCQhard

A data engineer is using AWS Lake Formation to manage access to a data lake stored in Amazon S3. The engineer grants SELECT permission on a table in the AWS Glue Data Catalog to an IAM role used by an Amazon Athena user. However, the user still cannot query the table and receives an 'Access Denied' error. The engineer verifies that the IAM role has the necessary AWS Lake Formation permissions and that the S3 bucket policy allows access. What is the most likely cause of the issue?

A.The IAM role lacks permissions to the AWS Glue Data Catalog.
B.The S3 bucket policy does not grant access to the IAM role.
C.The IAM role requires an explicit DENY in the S3 bucket policy.
D.The AWS Lake Formation data lake location is not registered.
AnswerD

For Lake Formation to enforce and grant access to data in S3, the S3 location must be registered with Lake Formation. If the location is not registered, Lake Formation cannot manage permissions on the underlying data, and queries will fail with 'Access Denied' even if table-level permissions are granted. Registering the location allows Lake Formation to assume control of S3 access for that path.

Why this answer

The root cause is that the S3 data location is not registered with AWS Lake Formation. Even with Lake Formation table permissions, if the underlying S3 path is not registered, Lake Formation cannot manage access, and Athena queries fail with 'Access Denied'. Registering the location enables Lake Formation to control S3 permissions and grant access based on its policies.

Exam trap

The trap here is assuming that granting Lake Formation table permissions alone is sufficient, overlooking the need to register the S3 location with Lake Formation for it to manage data access.

127
MCQmedium

A data engineer maintains an Amazon Kinesis Data Firehose delivery stream that writes JSON records to Amazon S3 and then invokes an AWS Lambda function for transformation. The Lambda function occasionally times out, causing records to be delivered untransformed. The engineer must ensure failed records are captured for later reprocessing without blocking delivery. What should the engineer do?

A.Enable the Kinesis Data Firehose data transformation failure option to send failed records to a separate S3 bucket for later reprocessing.
B.Increase the Lambda function's timeout to 15 minutes and memory to 10 GB so transformations always complete.
C.Configure the delivery stream to use Amazon Redshift as the destination and enable S3 backup for all records.
D.Attach a dead-letter queue to the Lambda function and configure Kinesis Data Firehose to read from that queue.
AnswerA

Kinesis Data Firehose supports a processing configuration with a Lambda function and a failure destination for records that fail transformation. Enabling that option writes failed records to a designated S3 bucket so they can be reprocessed later, while successful records continue to the main destination. This meets the requirement without blocking delivery.

Why this answer

Kinesis Data Firehose transformation supports a failure destination for records that the Lambda function cannot process. Enabling it sends failed records to a separate S3 bucket for later reprocessing while successful records continue to the primary destination. The other options either mask timeouts, change the destination without capturing failures, or use a dead-letter queue mechanism that does not apply to synchronous Firehose invocations.

Exam trap

The trap here is assuming a Lambda dead-letter queue captures Kinesis Data Firehose transformation failures, when Firehose invokes the function synchronously and cannot consume from an SQS queue.

128
MCQhard

A data engineer is using Amazon Managed Workflows for Apache Airflow (MWAA) to orchestrate a data pipeline. The pipeline includes a task that runs an AWS Glue job. The engineer notices that the Glue job occasionally fails due to transient issues, and the Airflow task fails immediately without retrying. The engineer wants to configure the Airflow task to retry the Glue job up to 2 times with a 5-minute delay between retries. Which configuration in the Airflow DAG should the engineer use?

A.Set retries=2 and retry_delay=timedelta(minutes=5) in the default_args dictionary.
B.Set max_retries=2 and retry_interval=300 in the Glue job's default arguments.
C.Use the retry_exponential_backoff=True parameter in the DAG definition.
D.Configure the Airflow task to use the GlueJobOperator with retry_limit=2 and retry_delay=300.
AnswerA

In Apache Airflow, the retries and retry_delay parameters in default_args control the number of retries and the delay between them for all tasks in the DAG. Setting retries to 2 and retry_delay to 5 minutes will automatically retry the failed Glue job task up to two times with the specified delay. This is the standard way to configure retry behavior in Airflow.

Why this answer

In Amazon MWAA, retry behavior for tasks is controlled by Airflow's built-in retry parameters. Setting retries and retry_delay in default_args applies to all tasks, including those running Glue jobs. This ensures that transient failures are retried with the specified delay, improving pipeline resilience.

Exam trap

The trap here is confusing retry configuration at the orchestration layer (Airflow) with retry settings on the Glue job itself, leading to attempts to set non-existent Glue parameters.

129
MCQmedium

A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an Amazon RDS for MySQL database to an Amazon S3 bucket in Parquet format. The replication task is configured with full load plus change data capture (CDC). After several hours, the engineer notices that the S3 bucket contains only the full load data and no incremental changes. Which action should the engineer take to ensure CDC changes are captured?

A.Increase the DMS replication instance size to handle the change volume.
B.Enable Multi-AZ on the source RDS for MySQL instance.
C.Modify the DMS task to use a different target endpoint for CDC data.
D.Enable binary logging on the source RDS for MySQL instance and set the binlog_format parameter to ROW.
AnswerD

AWS DMS requires binary logging enabled on the source MySQL instance for CDC, and the binlog format must be ROW to capture row-level changes. Without this, DMS cannot read the transaction log and will only perform the initial full load. Setting binlog_format to ROW is a prerequisite for ongoing replication, so this action directly enables CDC.

Why this answer

For AWS DMS to perform ongoing replication from MySQL, the source must have binary logging enabled and binlog_format set to ROW. Without these settings, DMS cannot read the transaction log to capture changes. The engineer must enable binary logging and set the format to ROW, then restart the task.

Other actions like scaling the instance or enabling Multi-AZ do not enable CDC.

Exam trap

The trap here is assuming that CDC will work automatically once the DMS task is set to full load plus CDC, without verifying source database prerequisites like binary logging.

130
MCQhard

A data engineer is responsible for a data pipeline that uses Amazon S3 as a data lake, AWS Glue for ETL, and Amazon Athena for ad-hoc queries. The pipeline ingests CSV files from an external partner via SFTP into an S3 bucket. The files are then processed by a Glue job that converts them to Parquet and writes to a separate S3 bucket partitioned by date. The Glue job runs daily and is triggered by a scheduled CloudWatch Events rule. Recently, the data engineer noticed that some days the Glue job fails because of memory errors, and on those days the Athena queries that rely on the data return incomplete results. The engineer needs to ensure that the pipeline is resilient and that Athena queries always see a complete view of the data, even if the Glue job fails mid-run. The engineer also needs to minimize re-processing of data. Which course of action should the engineer take?

A.Increase the number of workers and the worker type to G.2X to handle the memory errors, and enable job retries.
B.Replace the Glue job with an AWS Lambda function that processes the CSV files and writes Parquet to S3, and use S3 Event Notifications to trigger the function.
C.Modify the Glue job to use job bookmarks for incremental processing and write the Parquet output to a temporary location, then use an S3 copy operation to move the data into the final partitioned location only after the job completes successfully.
D.Use Athena partition projection to automatically discover partitions and set up a retry mechanism using AWS Step Functions.
AnswerC

Writing Parquet to a temporary prefix and copying only after successful completion keeps the final partitioned location free of partial output, so Athena never reads incomplete data. Glue job bookmarks track previously processed S3 objects, satisfying the minimise re-processing constraint by skipping already-ingested files on retry.

Why this answer

Writing Parquet output to a temporary location and only copying it into the final partitioned path after the job succeeds ensures Athena never sees partial data from a failed run. Job bookmarks enable incremental processing so already-processed files aren't reprocessed, minimizing re-work. This combination directly addresses both resilience (atomic visibility) and efficiency (no duplicate processing).

Exam trap

DEA-C01 often tests whether candidates conflate 'fix the failure' (more workers, retries) with 'ensure atomic visibility' — the trap is choosing a scaling fix that doesn't address partial-data exposure to Athena.

How to eliminate wrong answers

Option A is wrong because increasing workers and retries only mitigates memory errors — it doesn't prevent Athena from seeing partial output if a run fails mid-write. Option B is wrong because Lambda has a 15-minute timeout and limited memory, making it unsuitable for large CSV-to-Parquet ETL, and it doesn't solve atomic visibility. Option D is wrong because partition projection only speeds up Athena partition discovery — it does nothing to prevent partial data visibility or reduce Glue reprocessing.

131
Multi-Selecthard

A data engineer is optimizing an AWS Glue ETL job that processes large Parquet files in Amazon S3. The job currently takes several hours to complete. The engineer wants to improve performance by tuning the job's execution parameters. Which TWO actions will MOST effectively reduce the job's runtime? (Choose two.)

Select 2 answers
A.Partition the input data in Amazon S3 and use predicate pushdown in the Glue job.
B.Use the Glue DynamicFrame instead of Spark DataFrame for all transformations.
C.Increase the number of DPUs allocated to the Glue job.
D.Increase the Glue job's timeout value to allow more time for completion.
E.Enable job bookmarks to track processed data.
AnswersA, C

Partitioning the input data (e.g., by date or category) allows Glue to read only relevant partitions, reducing I/O. Predicate pushdown pushes filter conditions to the data source, so only matching data is read. Together, they minimize the amount of data scanned and processed, significantly improving runtime for large datasets. This is a best practice for optimizing Glue ETL jobs.

Why this answer

To reduce the runtime of a Glue ETL job processing large Parquet files, increasing DPUs provides more compute resources for parallel processing, and partitioning the input data with predicate pushdown reduces the amount of data read. These two actions directly address performance bottlenecks. Job bookmarks, DynamicFrames, and timeout adjustments do not effectively reduce runtime for a single large job run.

Exam trap

The trap here is confusing features that improve incremental processing (like job bookmarks) with those that improve single-run performance, and assuming that DynamicFrames are always faster.

132
MCQmedium

A data engineer is troubleshooting an AWS Glue ETL job that fails intermittently with the error 'Rate exceeded.' The job reads from an Amazon RDS for MySQL source and writes to Amazon S3. What is the MOST likely cause of this error?

A.The Glue job is using Amazon Kinesis Data Streams as a source, which has a shard throughput limit.
B.The number of Glue job workers or parallel queries is exceeding the maximum connections or IOPS of the RDS instance.
C.The Amazon S3 bucket has a bucket policy that limits the number of objects written per second.
D.The IAM role attached to the Glue job does not have sufficient permissions to read from RDS.
AnswerB

Glue workers open concurrent JDBC connections to the RDS for MySQL source. When worker count or parallel read queries exceed the instance's maximum connections or provisioned IOPS, the source rejects requests, surfacing as 'Rate exceeded' — an RDS-side throttling limit, not a Glue service quota.

Why this answer

The 'Rate exceeded' error in AWS Glue typically occurs when the job opens too many concurrent connections to the source database. With RDS for MySQL, each Glue worker or parallel read task establishes a connection, and exceeding the instance's max_connections or IOPS limits triggers throttling. This is the most likely cause when reading from RDS and writing to S3.

Exam trap

The trap is assuming the error comes from the target (S3) or from a misconfigured IAM role; candidates must recognize that 'Rate exceeded' often points to source database connection limits when using JDBC sources like RDS.

How to eliminate wrong answers

Option A is wrong because the source is explicitly Amazon RDS for MySQL, not Kinesis Data Streams; Kinesis shard limits would produce a different error and are not relevant here. Option C is wrong because S3 does not have a bucket policy that limits objects per second; S3 scales automatically and throttling would be reported as 503 Slow Down, not 'Rate exceeded' from RDS. Option D is wrong because insufficient IAM permissions would cause an AccessDenied error, not a rate limit error.

133
MCQeasy

A data engineer is configuring an S3 bucket for a data lake. The engineer runs the command shown in the exhibit. What does the output indicate about the bucket?

A.Versioning is enabled on the bucket.
B.The bucket retains only the latest version of each object.
C.Versioning is suspended on the bucket.
D.MFA Delete is enabled for the bucket.
AnswerA

The command's output reports a Versioning configuration, and its Status field reads Enabled, confirming that object versions are retained rather than overwritten. This distinguishes versioning from bucket logging, encryption or lifecycle settings, which appear under different configuration keys.

Why this answer

The output of the command (likely 'aws s3api get-bucket-versioning') shows a Status of 'Enabled', indicating that versioning is enabled on the bucket. This means multiple versions of objects are retained.

Exam trap

DEA-C01 often tests the difference between versioning enabled and suspended; the output 'Enabled' means versioning is active, not suspended.

How to eliminate wrong answers

Option B is wrong because if versioning is enabled, all versions are retained, not just the latest. Option C is wrong because 'Suspended' would be the status if versioning were suspended. Option D is wrong because MFA Delete is a separate setting that requires MFA for deletion and is not indicated by the versioning status alone.

134
Multi-Selectmedium

A data engineer is designing a data pipeline that ingests streaming data from an IoT device fleet. The data must be processed in near real-time and stored in Amazon S3 for long-term analytics. Which TWO AWS services should the engineer use together to achieve this?

Select 2 answers
A.Amazon Athena
B.AWS Glue
C.Amazon Kinesis Data Firehose
D.Amazon Kinesis Data Streams
E.Amazon Simple Queue Service (SQS)
AnswersC, D

Kinesis Data Firehose provides a fully managed delivery stream that ingests streaming records and automatically batches, transforms and writes them into Amazon S3, satisfying the near real-time processing and long-term S3 storage requirements without custom consumer code.

Why this answer

Amazon Kinesis Data Streams (D) is correct because it provides a highly scalable, low-latency ingestion service for real-time streaming data from thousands of IoT devices, allowing custom consumers to process records in near real-time. Amazon Kinesis Data Firehose (C) is correct because it can consume that stream (or receive data directly) and reliably deliver it in near real-time to Amazon S3 for long-term analytics, handling batching, compression, and format conversion. Together they form the canonical AWS pattern for real-time IoT ingestion plus durable S3 storage.

Amazon Athena (A) is only a query service over S3 and does not ingest or process streaming data. AWS Glue (B) is a serverless ETL/catalog service, not a real-time streaming ingestion or delivery mechanism. Amazon SQS (E) is a message queue for decoupling applications, not designed for high-throughput real-time streaming ingestion into S3.

Exam trap

DEA-C01 often tests whether candidates can distinguish ingestion services (Kinesis) from storage/query services (Athena, Glue) and decoupling services (SQS), so picking Athena or Glue for the ingestion leg is the common mistake.

135
MCQhard

A data engineer is troubleshooting an AWS Glue ETL job that fails with an 'Access Denied' error when trying to write to an S3 bucket. The IAM role used by the job has the policy shown in the exhibit. The bucket 'my-bucket' uses S3 default encryption with AWS KMS. What is the most likely missing permission?

A.s3:GetObjectVersion
B.glue:GetObject
C.s3:ListBucketMultipartUploads
D.s3:PutObjectAcl
E.kms:GenerateDataKey and kms:Decrypt
AnswerE

Writing to an SSE-KMS encrypted bucket requires the role to call KMS GenerateDataKey to obtain a data key for encrypting each object, plus Decrypt for reads and multipart uploads. The IAM policy grants only S3 actions, so the missing kms:GenerateDataKey and kms:Decrypt permissions cause the Access Denied.

Why this answer

When an S3 bucket uses AWS KMS for default encryption, any write operation requires the IAM role to have kms:GenerateDataKey and kms:Decrypt permissions on the KMS key. The policy in the exhibit grants s3:PutObject but does not include any KMS actions, resulting in an 'Access Denied' error. Option A (s3:GetObjectVersion) is not needed for writing.

Option B (glue:GetObject) is not a valid AWS action. Option C (s3:ListBucketMultipartUploads) is not required for a write operation. Option D (s3:PutObjectAcl) is unnecessary unless the job explicitly sets ACLs.

136
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. The engineer notices that when a Glue job fails, the Step Functions execution also fails, but the engineer wants to retry the failed job up to three times with exponential backoff before failing the entire workflow. The engineer needs to implement this with minimal changes to the state machine. What should the engineer do?

A.Configure the Glue job to automatically retry on failure by setting the MaxRetries parameter in the job's default arguments.
B.Create an Amazon CloudWatch alarm that triggers an AWS Lambda function to restart the Glue job on failure.
C.Modify the state machine to use a Map state that iterates over the Glue job three times.
D.Add a Retry field with MaxAttempts: 3 and BackoffRate: 2.0 to the state that invokes the Glue job.
AnswerD

The Retry field in Amazon States Language allows specifying retry behavior for a state, including maximum attempts and exponential backoff. Setting MaxAttempts to 3 and BackoffRate to 2.0 will retry the Glue job up to three times with exponential backoff. This is the standard and minimal change to achieve the requirement without altering the job itself.

Why this answer

AWS Step Functions provides native retry capabilities via the Retry field in the state definition. Specifying MaxAttempts and BackoffRate on the state that calls the Glue job enables automatic retries with exponential backoff. This is the simplest and most direct method to meet the requirement without altering the Glue job or adding external components.

Exam trap

The trap here is confusing retry mechanisms at the job level with those at the orchestration level, leading to attempts to configure retries in Glue instead of Step Functions.

137
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. During a recent run, a Glue job failed due to a transient network issue, and the Step Functions execution stopped. The engineer needs to ensure that the workflow can automatically retry the failed Glue job up to three times before considering the step failed. Which Step Functions state configuration should the engineer implement?

A.Set the TimeoutSeconds and HeartbeatSeconds fields in the Task state to trigger a retry on failure.
B.Add a Catch field to the Task state with ErrorEquals set to States.ALL and Next set to a retry state.
C.Add a Retry field to the Task state with ErrorEquals set to States.ALL and MaxAttempts set to 3.
D.Configure the Glue job itself to have a maximum retry count of 3 in its job definition.
AnswerC

The Retry field in a Step Functions Task state allows automatic retries on specified errors. Setting ErrorEquals to States.ALL catches all errors, and MaxAttempts to 3 retries the task up to three times. This is the correct way to handle transient failures in Glue jobs orchestrated by Step Functions. It ensures the workflow does not fail immediately and provides resilience.

Why this answer

The Retry field in Step Functions is designed to automatically retry a task when specified errors occur. By setting ErrorEquals to States.ALL and MaxAttempts to 3, the Glue job will be retried up to three times on any error, including transient network issues. The Catch field is for fallback, not retry.

Glue job definitions do not have a retry count, and timeout/heartbeat fields are for monitoring, not retry logic.

Exam trap

The trap here is confusing the Catch field with the Retry field, or assuming Glue jobs have built-in retry settings.

138
MCQeasy

A data engineer notices that an Amazon Kinesis Data Firehose delivery stream is failing to deliver data to an Amazon S3 bucket. The CloudWatch metrics show 'DeliveryToS3.Success' is 0 and 'S3.BucketExists' is 1. What is the MOST likely cause?

A.The S3 bucket has an ACL that denies access to Firehose.
B.The Firehose delivery stream Lambda transformation function is failing.
C.The IAM role for Firehose lacks s3:PutObject permission.
D.The S3 bucket does not exist.
AnswerC

S3.BucketExists being 1 confirms the bucket is reachable, so the failure lies in authorisation. Firehose requires s3:PutObject in its IAM role to write objects; without it, every delivery attempt is denied and DeliveryToS3.Success stays at 0.

Why this answer

The metric 'S3.BucketExists' is 1, confirming the S3 bucket exists, so the issue is not bucket existence. With 'DeliveryToS3.Success' at 0, the failure is in the write operation. The IAM role assumed by Firehose must have the s3:PutObject permission to deliver data; lacking it would cause all delivery attempts to fail silently, matching the observed metrics.

Exam trap

The trap here is that candidates may confuse 'S3.BucketExists' with successful delivery, or assume a missing bucket is the issue when the metric clearly shows the bucket exists, leading them to overlook the IAM permission gap.

How to eliminate wrong answers

Option A is wrong because S3 bucket ACLs are not evaluated when the IAM role grants the s3:PutObject permission via a bucket policy or identity-based policy; ACLs are legacy and Firehose uses IAM for authorization. Option B is wrong because a failing Lambda transformation function would cause 'DeliveryToS3.Success' to be 0 only if the transformation is mandatory, but the metric 'S3.BucketExists' would still be 1, and the failure would be logged as 'Lambda.ExecutionErrors' or similar, not directly as a delivery failure. Option D is wrong because 'S3.BucketExists' is 1, which explicitly indicates the bucket exists, so the bucket not existing cannot be the cause.

139
MCQhard

A data engineer is using Amazon Athena to query data stored in Amazon S3 in Parquet format. The engineer notices that a specific query is scanning much more data than expected, resulting in high costs and slow performance. The query filters on a column named 'event_date' which is a string in 'YYYY-MM-DD' format. The table is partitioned by 'year', 'month', and 'day' as separate string columns. The engineer wants to reduce the amount of data scanned. Which action should the engineer take?

A.Compress the Parquet files using Snappy compression.
B.Enable Amazon Athena query result reuse to cache the results.
C.Convert the 'event_date' column to a date type and use it in the WHERE clause.
D.Use the partition columns 'year', 'month', and 'day' in the WHERE clause instead of 'event_date'.
AnswerD

Athena uses partition pruning to limit the data scanned based on partition columns in the WHERE clause. Since the table is partitioned by year, month, and day, filtering on these columns allows Athena to scan only the relevant partitions, drastically reducing data scanned. This is the most effective way to optimize the query and reduce costs.

Why this answer

The table is partitioned by year, month, and day, so filtering on those partition columns in the WHERE clause enables Athena's partition pruning. This limits the data scanned to only the relevant partitions, significantly reducing cost and improving performance. Filtering on a non-partition column like event_date does not provide the same benefit.

Exam trap

The trap here is assuming that converting a column to date type or enabling caching will reduce data scanned, when the key optimization is to use the existing partition columns in the filter.

140
MCQeasy

A team uses Amazon Kinesis Data Analytics to process streaming data. They notice that the application's output is delayed. Which AWS service can be used to monitor the application's performance and identify bottlenecks?

A.AWS CloudTrail
B.Amazon CloudWatch
C.Amazon Athena
D.AWS X-Ray
AnswerB

Amazon CloudWatch collects Kinesis Data Analytics metrics such as millisBehindLatest, input/output records and processing time, exposing exactly where the pipeline stalls. Alarms and dashboards surface the bottleneck causing the output delay, satisfying the stem's requirement to monitor application performance.

Why this answer

Amazon CloudWatch is the native monitoring service for AWS resources and applications, and Kinesis Data Analytics (now Amazon Managed Service for Apache Flink) publishes metrics such as millisBehindLatest, KPU utilization, and input/output throughput directly to CloudWatch. These metrics let the team pinpoint whether the delay stems from input backlog, insufficient KPUs, or downstream sink throttling. CloudWatch alarms and dashboards then provide the operational visibility needed to remediate the bottleneck.

Exam trap

The trap here is confusing observability services: candidates see 'monitor performance' and pick X-Ray or CloudTrail, but DEA-C01 expects you to know CloudWatch is the metrics/alarms backbone for Kinesis Data Analytics, while X-Ray is for request tracing and CloudTrail is for API auditing.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API activity and audit events (who called what, when), not runtime performance metrics, so it cannot reveal processing latency or throughput bottlenecks. Option C is wrong because Amazon Athena is a serverless query service for data in S3 — it analyzes historical data at rest, not live streaming application performance. Option D is wrong because AWS X-Ray traces requests through distributed applications (e.g., Lambda, API Gateway, ECS) and, while useful for latency analysis in some architectures, it is not the primary monitoring surface for Kinesis Data Analytics application metrics like millisBehindLatest or KPU usage.

141
MCQmedium

A company stores sensitive data in Amazon S3 and requires that all data be encrypted at rest. The data is accessed by multiple AWS services. Which solution meets the encryption requirement with the LEAST operational overhead?

A.Use server-side encryption with AWS KMS (SSE-KMS)
B.Use client-side encryption with AWS KMS
C.Use server-side encryption with customer-provided keys (SSE-C)
D.Enable S3 default encryption with SSE-S3
AnswerD

SSE-S3 lets Amazon S3 manage the encryption keys entirely, so no key policies, grants or per-service configuration are needed. It satisfies the at-rest encryption requirement for data accessed by multiple AWS services while adding the least operational overhead, unlike SSE-KMS or client-side encryption.

Why this answer

Enabling S3 default encryption with SSE-S3 applies AES-256 encryption automatically to all objects at rest with zero key management overhead. AWS manages the keys, and no per-request KMS calls are needed, making it the lowest operational overhead solution that meets the encryption-at-rest requirement.

Exam trap

DEA-C01 often tests the trade-off between security control and operational overhead, causing candidates to choose SSE-KMS for its stronger key management when the question emphasizes 'least operational overhead,' where SSE-S3 is the correct answer.

How to eliminate wrong answers

Option A is wrong because SSE-KMS, while secure, adds operational overhead: KMS key policies must be managed, and each request incurs a KMS API call (with potential throttling and cost), which is more overhead than SSE-S3. Option B is wrong because client-side encryption requires the application to encrypt data before uploading and manage keys, adding significant operational complexity. Option C is wrong because SSE-C requires the customer to provide and manage encryption keys on every request, which is the highest operational overhead among the options.

142
MCQhard

A data engineer is using AWS Glue to run an ETL job that reads from an Amazon RDS for PostgreSQL database and writes to Amazon S3. The job is configured with a JDBC connection to RDS. The engineer notices that the job fails intermittently with a 'Connection timed out' error. The RDS instance is in a private subnet, and the Glue job has been configured with a VPC connection. Which action should the engineer take to resolve the timeout issue?

A.Enable the Glue job's VPC connection to use a NAT gateway for outbound traffic.
B.Increase the size of the Glue job's DPU capacity to handle more concurrent connections.
C.Ensure the Glue job's VPC connection has a security group that allows outbound traffic to the RDS instance's security group on the database port.
D.Configure the Glue job to use a public internet gateway to connect to RDS.
AnswerC

When AWS Glue runs in a VPC, it uses an elastic network interface (ENI) with a security group. To connect to RDS, the RDS security group must allow inbound traffic from the Glue ENI's security group on the database port. Additionally, the Glue security group must allow outbound traffic to RDS. Misconfigured security groups are a common cause of connection timeouts.

Why this answer

The intermittent connection timeout is likely caused by security group rules blocking traffic between the AWS Glue VPC connection and the RDS instance. The Glue job's ENI security group must allow outbound traffic to RDS, and the RDS security group must allow inbound traffic from the Glue security group on the database port. Ensuring these rules are correctly configured resolves the issue.

Exam trap

The trap here is assuming that scaling resources or using NAT gateways will fix network connectivity, when the issue is typically security group or subnet routing misconfiguration.

143
MCQeasy

A company uses Amazon Redshift for data warehousing. They notice that query performance has degraded over time. Which maintenance operation should be performed to improve performance?

A.Run the VACUUM command
B.Drop and recreate the table
C.Run the REINDEX command
D.Run the ANALYZE command
AnswerA

VACUUM reclaims space from deleted rows and re-sorts unsorted data, restoring the physical ordering that Redshift's zone maps rely on for block skipping. Degraded performance from accumulated deletes and unsorted loads is exactly the maintenance gap it addresses, so it satisfies the stem's requirement directly.

Why this answer

The VACUUM command in Amazon Redshift reclaims space from deleted rows and re-sorts rows to restore the sort order, which improves query performance by reducing I/O and enabling more efficient scans. Over time, as data is updated or deleted, the physical storage becomes fragmented and sort order degrades, leading to slower queries. Running VACUUM (or VACUUM SORT ONLY, VACUUM DELETE ONLY, etc.) is the recommended maintenance operation to address this degradation.

Exam trap

DEA-C01 often tests the distinction between VACUUM (space reclamation and sort) and ANALYZE (statistics update), causing candidates to choose ANALYZE when the issue is performance degradation due to fragmentation.

How to eliminate wrong answers

Option B is wrong because dropping and recreating a table is a drastic, disruptive action that would lose data and require reloading, and it is not a standard maintenance operation for performance improvement. Option C is wrong because REINDEX is not a Redshift command; Redshift does not have a REINDEX operation like some other databases. Option D is wrong because ANALYZE updates table statistics used by the query planner, which can improve query plans, but it does not reclaim space or restore sort order; it is complementary to VACUUM but not the primary operation for performance degradation due to fragmentation.

144
MCQeasy

A company uses Amazon S3 to store log files from multiple applications. The logs are written in JSON format. A data engineer wants to use Amazon Athena to query these logs. The logs are stored in a bucket with the following structure: 's3://logs/app1/date=2021-01-01/'. The engineer creates an Athena table with partitions. However, when querying, Athena returns zero results for partitions that exist. The engineer has run MSCK REPAIR TABLE to add partitions. What is the most likely cause of the issue?

A.The MSCK REPAIR TABLE command failed silently.
B.The partition key name in the table definition does not match the S3 folder naming convention.
C.The log files are in JSON format and Athena does not support JSON.
D.The log files need to be copied to a different bucket in the same region.
AnswerB

This is the correct answer. The S3 folder structure uses 'date=' as the partition key prefix, so the Athena table must define a partition key named 'date' exactly. If it is named differently, MSCK REPAIR will not register those folders as partitions.

Why this answer

The most likely cause is that the partition key name in the Athena table definition does not match the S3 folder naming convention. When using MSCK REPAIR TABLE, Athena relies on the partition folder structure (e.g., 'date=2021-01-01') to automatically add partitions. If the table's partition key is named differently (e.g., 'dt' instead of 'date'), MSCK REPAIR will not recognize the folders and will not register the partitions, resulting in zero results.

Option A is incorrect because MSCK REPAIR does not fail silently; it either adds partitions or reports none if the structure doesn't match. Option C is incorrect because Athena fully supports JSON format. Option D is incorrect because the bucket location does not affect partition registration; data can be queried in any bucket as long as the table points to it.

145
MCQmedium

A data pipeline uses AWS Glue to process data from Amazon S3 and write results to Amazon Redshift. The pipeline fails intermittently with the error 'S3ServiceException: Access Denied'. The IAM role used by Glue has permissions to read from the S3 bucket. What is the most likely cause of this error?

A.The S3 bucket is in a different AWS Region than the Glue job
B.S3 Server Access Logging is enabled and blocking requests
C.The S3 bucket policy denies access to the Glue job's IAM role
D.The S3 bucket has S3 Transfer Acceleration enabled
AnswerC

An explicit Deny in the S3 bucket policy overrides the IAM role's Allow permissions. Even though the Glue role can read S3, the bucket policy denying that role causes the intermittent Access Denied error described in the stem.

Why this answer

S3 bucket policies can explicitly deny access even if the IAM role allows it. Option A is wrong because S3 bucket region does not cause access denied errors. Option B is wrong because S3 Server Access Logging does not affect access permissions.

Option D is wrong because S3 Transfer Acceleration is not related to access denied errors.

146
MCQmedium

A data engineer manages an AWS Glue ETL job that reads JSON files from Amazon S3, transforms the data, and writes to an Amazon Redshift table. The job recently started failing with the error 'Communication link failure: connection reset'. The Redshift cluster is healthy and the IAM role used by the Glue job has the necessary permissions. The engineer notices that the job runs longer than before and the Redshift cluster's WLM queue is often full. Which action should the engineer take to resolve the failure?

A.Modify the Redshift WLM configuration to increase concurrency or add a new queue for the Glue job.
B.Configure the Glue job to use the Redshift JDBC driver with a longer connection timeout and retry logic.
C.Increase the number of AWS Glue DPUs allocated to the job to speed up processing.
D.Switch the Glue job to write to Amazon S3 first, then use a Redshift COPY command to load the data.
AnswerA

The error 'Communication link failure: connection reset' often occurs when Redshift terminates connections due to WLM queue timeouts. Increasing concurrency or adding a dedicated queue for the Glue job reduces wait times and prevents connection resets. This directly addresses the root cause of the failure, allowing the job to complete successfully without changing Glue resources.

Why this answer

The failure is caused by Redshift's workload management (WLM) queue being full, which leads to connection timeouts and resets when the AWS Glue job attempts to write data. Adjusting the WLM configuration to increase concurrency or create a dedicated queue for the Glue job reduces wait times and prevents the connection resets. This directly addresses the root cause without unnecessary changes to the Glue job or architecture.

Exam trap

The trap here is assuming that connection reset errors are always network-related and can be fixed by increasing timeouts or retries, when in fact they often stem from Redshift workload management limits.

147
MCQeasy

A company stores sensitive data in Amazon S3 and needs to ensure that data is encrypted at rest. Which AWS service can be used to manage the encryption keys?

A.AWS Key Management Service (KMS)
B.AWS Secrets Manager
C.AWS Identity and Access Management (IAM)
D.AWS Certificate Manager (ACM)
AnswerA

AWS Key Management Service creates, rotates and controls the customer managed keys used for server-side encryption of S3 objects, including SSE-KMS. It satisfies the requirement to manage encryption keys centrally while keeping data encrypted at rest in Amazon S3.

Why this answer

AWS Key Management Service (KMS) is the managed service for creating and controlling encryption keys used to encrypt data at rest in Amazon S3. Option A is correct. Option B (Secrets Manager) is for managing secrets like database passwords, not encryption keys.

Option C (IAM) manages access permissions, not encryption keys. Option D (Certificate Manager) handles SSL/TLS certificates, not encryption keys for data at rest.

148
Multi-Selectmedium

A data engineer is managing an Amazon Kinesis Data Firehose delivery stream that writes to an Amazon S3 bucket. The engineer notices that some records are being delivered to the S3 bucket with a prefix of 'errors/' instead of the intended 'data/' prefix. The Firehose stream is configured with an AWS Lambda function for data transformation, and the S3 bucket has a lifecycle policy that transitions objects to Glacier after 30 days. Which TWO actions should the engineer take to ensure that only successfully transformed records are delivered to the 'data/' prefix and that failed records are handled appropriately? (Choose two.)

Select 2 answers
A.Specify an error prefix in the Firehose stream's S3 destination configuration so that failed records are delivered to a separate prefix.
B.Enable the 'S3 backup' setting in the Firehose stream to back up all incoming records to a separate S3 bucket before transformation.
C.Increase the Firehose stream's buffer size and buffer interval to reduce the number of failed records.
D.Configure the Lambda function to return the transformed record with a 'result' status of 'Ok' for successful records and 'Dropped' or 'ProcessingFailed' for failed records.
E.Configure the Firehose stream's S3 destination to use a dynamic prefix based on the record's result status, such as 'data/' for success and 'errors/' for failure.
AnswersA, D

The S3 destination configuration for Kinesis Data Firehose includes an optional error prefix. When a record fails transformation (Lambda returns 'ProcessingFailed') or delivery, Firehose writes that record to the specified error prefix instead of the main prefix. By setting the error prefix to 'errors/', failed records are segregated, and only successfully transformed records (with 'Ok' status) are delivered to the 'data/' prefix. This is the correct way to handle failed records.

Why this answer

To ensure only successful records go to the 'data/' prefix, the Lambda function must return a 'result' of 'Ok' for successful transformations and 'Dropped' or 'ProcessingFailed' for failures. Additionally, configuring an error prefix in the Firehose S3 destination directs failed records to a separate 'errors/' prefix. Together, these actions segregate successful and failed records correctly.

Exam trap

The trap here is assuming that Firehose can dynamically route records based on status, when in fact the error prefix is a static configuration and the Lambda result status determines delivery.

149
MCQmedium

A data engineer is troubleshooting a data pipeline that uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The engineer notices that the S3 bucket contains many small files (less than 1 MB). This is causing performance issues in downstream processing. What is the BEST way to reduce the number of small files?

A.Increase the buffer size to at least 128 MB in the Firehose delivery stream configuration.
B.Use an AWS Lambda function to transform the data before delivery.
C.Change the compression format from GZIP to Snappy.
D.Decrease the buffer interval in the Firehose delivery stream configuration.
AnswerA

Firehose buffers incoming records and only delivers once the buffer size or interval is reached. Raising the buffer size to 128 MB or more forces larger, less frequent S3 objects, directly reducing the sub-1 MB file count causing downstream performance issues.

Why this answer

Increasing the Firehose buffer size to at least 128 MB allows more data to accumulate before delivery, resulting in fewer, larger files in S3. This directly addresses the small file problem by batching more records per delivery.

Exam trap

DEA-C01 often tests the confusion between buffer size and buffer interval, where candidates might think decreasing interval reduces file count, when it actually increases it.

How to eliminate wrong answers

Option B is wrong because using Lambda to transform data does not inherently reduce file count; it may even increase processing overhead without addressing buffering. Option C is wrong because changing compression format affects file size but not the number of files; small files would still be created. Option D is wrong because decreasing the buffer interval causes more frequent deliveries, resulting in even more small files.

150
MCQmedium

A data engineering team notices that an AWS Glue ETL job, which processes hourly data from an S3 bucket, is taking progressively longer to run. The job reads Parquet files partitioned by date and hour. Which action is MOST likely to improve the job's performance?

A.Enable pushdown predicate filtering on the job's data source.
B.Convert Parquet files to CSV to improve read performance.
C.Increase the number of DPUs for the job.
D.Switch from Spark to Python shell for simpler processing.
AnswerA

Pushdown predicates filter partitions and row groups at the S3/Parquet source, so Glue reads only matching date and hour data instead of scanning the full dataset. This cuts I/O and shuffle volume, directly addressing the growing runtime.

Why this answer

Pushdown predicate filtering allows AWS Glue to push filter conditions down to the data source, so only the relevant partitions and rows are read from S3. Since the data is partitioned by date and hour, predicate pushdown minimizes the amount of data scanned, reducing I/O and improving job performance. This is especially effective for Parquet, which supports columnar pruning and predicate pushdown.

Exam trap

DEA-C01 often tests the misconception that adding more DPUs is always the answer to performance issues, but the exam expects candidates to identify data pruning and predicate pushdown as the primary optimization for partitioned data.

How to eliminate wrong answers

Option B is wrong because converting Parquet to CSV would remove the columnar storage benefits and increase file size, making reads slower and more expensive. Option C is wrong because increasing DPUs adds more compute resources but does not address the root cause of reading unnecessary data; it may temporarily speed up the job but is not the most efficient or cost-effective solution. Option D is wrong because Python shell jobs are limited in scalability and are not suitable for complex ETL transformations; switching from Spark to Python shell would likely degrade performance for large datasets.

← PreviousPage 2 of 4 · 270 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Operations and Support questions.