Courseiva

CCNA Data Operations and Support Questions

75 of 270 questions · Page 1/4 · Data Operations and Support · Answers revealed

1
MCQmedium

Refer to the exhibit. This log snippet is from a failed AWS Glue job. The job processes a large dataset in memory. What is the MOST likely cause of the OutOfMemoryError?

A.The Glue job is running with insufficient DPUs or worker type.
B.The input data is in an unsupported file format.
C.The job is attempting to join two tables with mismatched keys.
D.The job has too many partitions.
AnswerA

Glue allocates memory per executor based on DPU count and worker type; a large in-memory dataset exceeding that allocation triggers OutOfMemoryError. Insufficient DPUs or an undersized worker type directly caps heap available to the job, so scaling them addresses the constraint stated in the stem.

Why this answer

An OutOfMemoryError in AWS Glue typically occurs when the allocated DPUs or worker type are insufficient for the in-memory processing of a large dataset. Option B is incorrect because unsupported file formats cause parsing errors, not memory errors. Option C is incorrect because mismatched keys in a join cause data skew or incorrect results, but not directly an OutOfMemoryError.

Option D is incorrect because too many partitions usually lead to small file overhead, not heap space exhaustion.

2
MCQeasy

A data engineer is monitoring an Amazon Redshift cluster and notices that the disk space usage is increasing rapidly. The engineer wants to reclaim space from deleted rows. Which command should the engineer run?

A.VACUUM
B.ANALYZE
C.UNLOAD
D.COPY
AnswerA

VACUUM reclaims storage occupied by deleted rows and re-sorts unsorted data in Amazon Redshift, directly addressing rapidly growing disk usage. It is the specific maintenance command that returns freed blocks to the system for reuse.

Why this answer

VACUUM in Amazon Redshift reclaims disk space occupied by rows marked for deletion and re-sorts rows to restore sort order. When rows are deleted, they are only logically marked until VACUUM runs, so disk usage keeps growing. Running VACUUM (or VACUUM DELETE ONLY) releases that space back to the system.

Exam trap

DEA-C01 often tests the distinction between VACUUM (space reclamation and re-sort) and ANALYZE (statistics refresh) — candidates who conflate maintenance commands pick ANALYZE, which does nothing for deleted-row disk usage.

How to eliminate wrong answers

Option B is wrong because ANALYZE only updates table statistics used by the query planner — it does not reclaim any disk space from deleted rows. Option C is wrong because UNLOAD exports query results from Redshift to Amazon S3; it is an outbound data operation, not a space-reclamation command. Option D is wrong because COPY loads data into Redshift from S3, DynamoDB, or other sources — it adds data rather than reclaiming deleted space.

3
MCQeasy

A data engineer is using Amazon Athena to query data stored in Amazon S3. The engineer wants to reduce query costs and improve performance for a table that is frequently queried with filters on a date column. The data is stored as uncompressed CSV files partitioned by year/month/day. Which action should the engineer take?

A.Convert the data to Parquet format and use partitioning by date, then run MSCK REPAIR TABLE.
B.Use Amazon Redshift Spectrum to query the S3 data instead of Athena.
C.Increase the number of S3 buckets storing the data to parallelize reads.
D.Enable Athena query result reuse and set a short cache TTL.
AnswerA

Converting to Parquet reduces the amount of data scanned because Athena reads only the necessary columns, and partitioning by date allows Athena to skip irrelevant partitions. Running MSCK REPAIR TABLE updates the partition metadata so queries can leverage the partitions, significantly lowering cost and improving performance.

Why this answer

For Athena, the most effective way to reduce cost and improve performance is to use a columnar format like Parquet and partition the data. Parquet allows column pruning, and partitioning enables partition pruning, so queries scan less data. Updating partitions with MSCK REPAIR TABLE ensures the metadata is current, allowing these optimizations to take effect.

Exam trap

The trap here is focusing on caching or infrastructure changes instead of the fundamental storage format and partitioning strategy that directly impact data scanned.

4
MCQmedium

A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The engineer notices that some data records are missing from S3, and the Firehose delivery stream metrics show increased DeliveryToS3.DataFreshness. The engineer needs to ensure all records are delivered. Which action should the engineer take?

A.Increase the number of shards in the source Kinesis data stream to improve throughput.
B.Enable Amazon S3 server access logging on the destination bucket to capture failed deliveries.
C.Enable error logging on the Firehose delivery stream and review the error output in Amazon CloudWatch Logs.
D.Configure the Firehose delivery stream to use a longer buffering interval to reduce the number of S3 PUT requests.
AnswerC

Enabling error logging for Firehose sends delivery errors to CloudWatch Logs, providing visibility into why records are not delivered. Common causes include insufficient S3 permissions, KMS key issues, or malformed records. Reviewing these logs allows the engineer to identify and fix the root cause, ensuring all records are delivered and reducing data freshness.

Why this answer

Enabling error logging on the Firehose delivery stream is crucial for diagnosing why records are not delivered to S3. The logs in CloudWatch will show specific errors such as access denied or invalid records, allowing the engineer to take corrective action. This directly addresses the missing data and high data freshness.

Exam trap

The trap here is assuming that increasing buffering or adding shards will fix missing data, but the real issue is delivery errors that require logging to diagnose.

5
MCQeasy

Refer to the exhibit. A data engineer runs this CLI command to check an object's metadata. The engineer wants to verify if the object is eligible for lifecycle transition to S3 Glacier based on its age. What additional information is needed?

A.The current date
B.The ETag value
C.The metadata archive flag
D.The ContentLength value
AnswerA

Lifecycle transition eligibility depends on the object's age, calculated as the current date minus LastModified. The CLI output supplies LastModified, so the current date is the missing value needed to determine whether the transition threshold has elapsed.

Why this answer

To determine whether an S3 object is eligible for lifecycle transition to S3 Glacier based on age, you need the object's creation date (LastModified) and the current date to compute its age. The exhibit presumably shows the object's LastModified timestamp, so the missing piece is the current date to calculate how many days old the object is relative to the lifecycle rule's transition threshold.

Exam trap

DEA-C01 often tests whether candidates focus on object metadata like ETag or size when the actual lifecycle eligibility depends on object age (LastModified vs. current date) and rule thresholds.

How to eliminate wrong answers

Option B is wrong because the ETag is a hash of the object's content (or multipart upload identifier) and has no bearing on lifecycle transition eligibility, which depends on age. Option C is wrong because there is no 'metadata archive flag' in S3 object metadata that governs lifecycle transitions; lifecycle rules are evaluated by the S3 service based on object age and rule configuration. Option D is wrong because ContentLength (object size) does not determine lifecycle transition eligibility unless the lifecycle rule explicitly filters by size, which is not the scenario described.

6
MCQmedium

A company uses AWS Lake Formation to manage data lake permissions. A data analyst cannot query a table in Athena, although the table appears in the catalog. The analyst has IAM permissions to run Athena. What is the MOST likely cause?

A.The Glue Data Catalog does not have the table registered.
B.The S3 bucket policy denies access to the analyst's IAM role.
C.The analyst lacks Lake Formation permissions on the table.
D.The Athena workgroup is not configured with the correct output location.
AnswerC

Lake Formation enforces its own table and column-level grants on top of IAM. The analyst's IAM policy permits Athena API calls, but without a Lake Formation SELECT grant on the table, the query is denied — explaining why the table appears in the catalog yet remains unqueryable.

Why this answer

The most likely cause is that the analyst lacks Lake Formation permissions on the table. Even with IAM permissions to run Athena, Lake Formation enforces fine-grained access control at the table and column level. If the analyst has not been granted SELECT permission on the table via Lake Formation, Athena queries will fail with an access denied error, even though the table appears in the Glue Data Catalog.

Exam trap

The trap is assuming that IAM permissions alone are sufficient to query data in a Lake Formation-governed data lake. Candidates might overlook Lake Formation's role and pick S3 bucket policy or Glue catalog issues. The exam tests the layered security model where Lake Formation permissions are required in addition to IAM.

How to eliminate wrong answers

Option A is wrong because if the table appears in the catalog, it is registered; the issue is not registration. Option B is wrong because while S3 bucket policies can deny access, Lake Formation manages access to underlying data, and if the analyst lacks Lake Formation permissions, that is the more direct cause. Option D is wrong because an incorrect Athena output location would cause query failures due to write permission issues, but the error would typically mention the output location, not access to the table.

7
MCQmedium

A company runs a data pipeline that uses AWS Glue to process data from an Amazon DynamoDB table and write results to Amazon S3. The Glue job runs on a schedule every hour. Recently, the job started failing intermittently with 'ProvisionedThroughputExceededException' errors from DynamoDB. What is the BEST solution?

A.Use DynamoDB Accelerator (DAX) to reduce read latency.
B.Change the Glue job schedule to run every 2 hours.
C.Implement exponential backoff and retries in the Glue job for DynamoDB operations.
D.Increase the read capacity units of the DynamoDB table.
AnswerC

Exponential backoff with jitter retries throttled DynamoDB requests, letting the table's provisioned capacity recover between attempts. This directly addresses the intermittent ProvisionedThroughputExceededException, which signals transient throttling rather than a persistent capacity shortfall. Retries absorb those spikes without over-provisioning, satisfying the requirement to stop hourly Glue job failures.

Why this answer

ProvisionedThroughputExceededException is a throttling error that occurs when request rate exceeds the table's provisioned read/write capacity, and AWS's documented best practice is to implement exponential backoff and retry logic so the client retries with increasing delays instead of hammering the table. This is the correct fix for intermittent throttling because it lets the Glue job ride out short bursts of capacity contention without failing. It also works whether the table is in provisioned or on-demand mode, and it is the least disruptive, most cost-effective change.

Exam trap

DEA-C01 often tests whether candidates jump to 'add more capacity' when the AWS-recommended answer for transient throttling is retry with exponential backoff — the exam wants the resilience pattern, not the brute-force capacity increase.

How to eliminate wrong answers

Option A is wrong because DAX reduces read latency for cached items but does not eliminate throttling on the underlying table for uncached reads or for writes, and Glue's DynamoDB connector does not automatically route through DAX without custom configuration. Option B is wrong because changing the schedule to every 2 hours only reduces frequency; it does not address the bursty request pattern that causes the throttle, and it delays data freshness. Option D is wrong because increasing RCUs may help but is a guess at the root cause, costs more money, and does not address write throttling or burst behavior — and the question asks for the BEST solution, which is the AWS-recommended retry pattern.

8
MCQmedium

Refer to the exhibit. A data engineer is troubleshooting an AWS Lambda function that processes data from Amazon S3. The function is triggered by S3 events, but no logs appear in CloudWatch Logs. The engineer runs the AWS CLI command shown. What is the MOST likely reason for the missing logs?

A.The Lambda execution role does not have permissions to create log groups and write logs.
B.The Lambda function is configured to log to a different log group.
C.The Lambda function is not being invoked by S3 events.
D.The log retention policy is set to 7 days, causing logs to expire immediately.
AnswerA

Without `logs:CreateLogGroup`, `logs:CreateLogStream` and `logs:PutLogEvents` in the execution role, Lambda cannot create the function's log group or stream, so invocations produce no CloudWatch output at all. This directly satisfies the stem's missing-logs constraint, since the trigger itself is firing normally.

Why this answer

The most likely reason for missing CloudWatch Logs is that the Lambda execution role lacks the necessary permissions to create log groups and write log streams. When a Lambda function is invoked, it attempts to create a log group named /aws/lambda/<function-name> and a log stream, then write logs. If the execution role does not include actions like logs:CreateLogGroup, logs:CreateLogStream, and logs:PutLogEvents, the function cannot write any logs, resulting in no log entries.

This is a common misconfiguration, especially when roles are custom-created without the default AWSLambdaBasicExecutionRole policy.

Exam trap

DEA-C01 often tests the misconception that Lambda automatically has permissions to write logs, but the execution role must explicitly allow it; candidates may overlook this and choose other options like log retention or invocation issues.

How to eliminate wrong answers

Option B is wrong because Lambda does not support configuring a different log group; logs are always sent to the log group /aws/lambda/<function-name> unless you use a custom logging solution, which is not indicated here. Option C is wrong because if the function were not invoked, there would be no logs, but the question states the function is triggered by S3 events; moreover, the exhibit likely shows a CLI command that confirms invocations, making this less likely. Option D is wrong because a 7-day retention policy would not cause logs to expire immediately; logs would still be visible for 7 days, and the issue is missing logs entirely, not expired ones.

9
MCQmedium

A data engineer manages an AWS Glue ETL job that processes JSON files from Amazon S3 and writes to Amazon Redshift. The job fails with the error 'Unable to find a suitable JDBC driver in the classpath'. The engineer has included the Redshift JDBC driver as a job parameter in the Glue job configuration. Which step should the engineer take to resolve the error?

A.Add the JDBC driver as an additional jar in the Glue job's 'Dependent jars path' and ensure the job's IAM role has s3:GetObject permission on the jar location.
B.Modify the job script to include the JDBC driver URL directly in the 'connection_options' when calling the write_dynamic_frame method.
C.Increase the number of workers in the Glue job to ensure the driver is loaded in parallel across all executors.
D.Convert the Glue job to use a development endpoint and manually install the JDBC driver on the underlying EC2 instance.
AnswerA

AWS Glue requires that JDBC drivers be provided as additional jars via the 'Dependent jars path' and accessible from S3. Simply setting a connection parameter is insufficient. The job's IAM role must also have permission to read the jar from S3. This is the documented approach for using custom JDBC drivers with Glue.

Why this answer

To use a custom JDBC driver with AWS Glue, you must upload the driver jar to Amazon S3 and specify its path in the job's 'Dependent jars path' parameter. Additionally, the IAM role associated with the job must have s3:GetObject permission for that jar. This ensures the driver is available in the classpath at runtime, resolving the 'Unable to find a suitable JDBC driver' error.

Exam trap

The trap here is assuming that adding the JDBC driver as a job parameter or connection option automatically makes it available to the job's classpath.

10
MCQmedium

A data pipeline uses AWS Glue to process data from Amazon S3. The job fails with an 'OutOfMemoryError' during the transformation phase. Which action should the data engineer take to resolve this issue?

A.Enable S3 server-side encryption.
B.Increase the number of partitions in the input data.
C.Change the data format from CSV to Parquet.
D.Increase the number of DPUs (Data Processing Units) for the Glue job.
AnswerD

Adding DPUs increases the total executor memory and parallelism available to the Glue job, letting Spark distribute transformation partitions across more workers. This directly addresses the OutOfMemoryError caused by insufficient memory per executor during the transformation phase.

Why this answer

The OutOfMemoryError occurs because the Glue job does not have enough memory allocated. Increasing the number of DPUs (Data Processing Units) increases both memory and processing capacity, directly resolving the issue. Option A (S3 server-side encryption) affects data security, not memory.

Option B (increasing data partitions) may help parallelism but does not directly increase memory per executor. Option C (changing to Parquet) can reduce data volume but does not guarantee sufficient memory for transformation.

11
Multi-Selecteasy

A data engineer is setting up an AWS Glue job to process data from an Amazon S3 bucket. The job fails with an 'Access Denied' error. Which TWO IAM permissions are MOST likely missing from the Glue job's IAM role?

Select 2 answers
A.s3:PutObject
B.kms:Decrypt
C.dynamodb:GetItem
D.glue:StartJobRun
E.s3:GetObject
AnswersA, E

Writing transformed output back to Amazon S3 requires s3:PutObject on the target prefix. Without it, the Glue job's write stage returns Access Denied even when reading succeeds, so this permission must be added to the job's IAM role.

Why this answer

The Glue job needs to read the source data from the S3 bucket, which requires the s3:GetObject permission on the bucket's objects, so option E is correct. The job also needs to write its output (or intermediate results) back to S3, which requires the s3:PutObject permission, making option A correct. Option B (kms:Decrypt) is not necessarily missing unless the S3 objects are encrypted with a customer-managed KMS key and the role lacks decrypt rights, but the scenario does not state that.

Option C (dynamodb:GetItem) is unrelated because the job processes data from S3, not DynamoDB. Option D (glue:StartJobRun) is a permission for triggering jobs, not for the job's execution role to access data, so it is not the cause of the Access Denied error.

12
MCQhard

A company uses Amazon Kinesis Data Streams with a Lambda consumer. The Lambda function is failing with 'ProvisionedThroughputExceededException' when writing to a DynamoDB table. Which action should the data engineer take to resolve this without losing data?

A.Reduce the number of Kinesis shards to lower the ingestion rate.
B.Increase the DynamoDB table's read capacity.
C.Configure a dead-letter queue (DLQ) on the Lambda function and increase the DynamoDB write capacity.
D.Disable retries on the Lambda function to avoid throttling.
AnswerC

The DLQ captures failed records so Kinesis retries do not discard them, while raising DynamoDB write capacity removes the throttling that triggers ProvisionedThroughputExceededException. Together they satisfy the no-data-loss constraint, though the DLQ alone would not fix the underlying capacity shortfall.

Why this answer

The Lambda function is throttled by DynamoDB because the write capacity is insufficient for the incoming Kinesis stream rate. Increasing the DynamoDB write capacity resolves the throttling, and configuring a DLQ on the Lambda function ensures that any records that fail after retries are captured for later processing, preventing data loss. This combination directly addresses the root cause while providing a safety net.

Exam trap

DEA-C01 often tests the misconception that increasing read capacity resolves write throttling, or that disabling retries prevents throttling, when in fact it causes data loss.

How to eliminate wrong answers

Option A is wrong because reducing Kinesis shards would lower the ingestion rate but does not address the DynamoDB write capacity issue and could cause data loss if the stream is scaled down. Option B is wrong because the error is ProvisionedThroughputExceededException on writes, so increasing read capacity has no effect. Option D is wrong because disabling retries would cause immediate data loss on throttling, worsening the problem.

13
MCQmedium

A data engineer is using Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to orchestrate a data pipeline. The engineer needs to ensure that the Airflow environment can access an Amazon S3 bucket to read and write data. The S3 bucket is in the same AWS account and Region as the MWAA environment. Which configuration is required to allow MWAA to access the S3 bucket?

A.Attach an IAM policy to the MWAA execution role that grants the necessary S3 permissions.
B.Configure the S3 bucket policy to allow the MWAA service principal.
C.Enable public access on the S3 bucket and use the S3 REST API with presigned URLs.
D.Store S3 credentials in AWS Secrets Manager and reference them in the Airflow connection.
AnswerA

Amazon MWAA uses an execution role that grants the Airflow environment permissions to access AWS services. To allow MWAA to read and write to S3, you must attach an IAM policy to this execution role with the appropriate s3:GetObject, s3:PutObject, and s3:ListBucket permissions. This is the standard method for granting MWAA access to S3.

Why this answer

MWAA environments use an execution role that is assumed by the Airflow components. To grant access to S3, you attach an IAM policy to this role with the necessary S3 permissions. This is the standard and recommended approach.

The execution role must also have a trust policy that allows MWAA to assume it, but that is created automatically when you create the environment.

Exam trap

The trap here is thinking that a bucket policy or Secrets Manager is needed, when the MWAA execution role's IAM policy is the primary mechanism.

14
MCQmedium

A data engineer maintains an AWS Glue ETL job that processes millions of small JSON files stored in Amazon S3. The job's runtime has increased significantly, and CloudWatch logs show many small executor tasks and frequent garbage collection. The engineer wants to improve job performance by reducing the number of small files processed per task. Which action should the engineer take?

A.Convert the source files to Parquet format using an AWS Glue crawler before running the ETL job.
B.Use the Glue ETL job's 'groupFiles' option to group multiple small files into a single partition.
C.Enable job bookmarks to track previously processed files and skip them on subsequent runs.
D.Increase the number of DPUs allocated to the job to provide more memory per executor.
AnswerB

The 'groupFiles' option in AWS Glue ETL allows grouping multiple small files into a single partition, reducing the number of tasks and improving read efficiency. This directly addresses the issue of many small files causing overhead. By setting groupFiles to 'inPartition' or 'inPartitionAcrossBuckets', the job can process larger chunks of data per task, reducing garbage collection and improving performance.

Why this answer

Grouping small files into larger partitions reduces the number of tasks and the overhead of task scheduling and garbage collection. This is a specific optimization for AWS Glue ETL jobs that read many small files. Enabling job bookmarks or increasing DPUs does not address the root cause.

Converting formats via a crawler is not a valid approach because crawlers do not transform data.

Exam trap

The trap here is assuming that adding more resources or enabling bookmarks will fix small-file inefficiency, when the real solution is to group files.

15
Multi-Selectmedium

A data engineer is troubleshooting a slow-running Amazon Athena query on a large dataset stored in S3. The query scans many small files. Which TWO actions can improve query performance?

Select 2 answers
A.Increase the number of files to increase parallelism
B.Disable S3 server-side encryption
C.Concatenate small files into larger files
D.Partition the data by a frequently filtered column
E.Convert files from CSV to JSON
AnswersC, D

Merging many small files reduces per-file overhead: Athena must open, list and read metadata for each object, so fewer larger files cut S3 request costs and task scheduling overhead. This directly addresses the small-file scanning inefficiency described in the stem.

Why this answer

Option C is correct because Athena performance is heavily degraded by many small files, since each file incurs overhead for opening, listing, and reading metadata; concatenating small files into larger files (typically 128 MB or more) reduces this per-file overhead and lets Athena scan data more efficiently. Option D is correct because partitioning the data by a frequently filtered column allows Athena to use partition pruning, so it reads only the relevant S3 prefixes instead of scanning the entire dataset, dramatically reducing the amount of data scanned and query time. Option A is incorrect because adding more small files increases metadata and open/close overhead rather than improving performance, even though parallelism is a factor.

Option B is incorrect because S3 server-side encryption is transparent to Athena and disabling it does not affect query performance. Option E is incorrect because converting CSV to JSON does not inherently reduce scanned data or file count and JSON is typically more verbose, so it would not improve performance.

16
MCQmedium

A data engineer needs to set up a data catalog for a new data lake in AWS Glue. The data resides in S3 in Parquet format. The engineer wants to ensure that the schema is automatically detected and updated when new columns are added to the data. Which configuration should the engineer use?

A.Add a partition index to the Glue Data Catalog table.
B.Configure the crawler's 'Schema updates' option to 'Update the table schema'.
C.Set the crawler's 'Database' output to a new database.
D.Enable partition indexing on the table.
AnswerB

The Glue crawler's schema-update behaviour controls whether re-crawls revise the Data Catalog table definition. Setting it to update the table schema lets the crawler add newly detected columns automatically, satisfying the requirement that schema changes in the Parquet data be picked up without manual intervention.

Why this answer

The Glue crawler's 'Schema updates' option controls what happens when the crawler detects schema changes in the data. Setting it to 'Update the table schema' allows the crawler to add new columns and modify the existing table definition in the Data Catalog when new columns appear in the Parquet files. This is the correct configuration for automatic schema evolution.

Exam trap

DEA-C01 often tests the confusion between crawler schema update behaviour and partition indexing, tricking candidates into picking an optimization feature when the question is about schema evolution.

How to eliminate wrong answers

Option A is wrong because a partition index improves query performance on partitioned tables but does not handle schema detection or updates. Option C is wrong because setting the crawler's database output only determines where the table metadata is stored, not how schema changes are handled. Option D is wrong because partition indexing, like option A, is a query optimization feature and has no role in schema evolution.

17
MCQmedium

A data engineer manages an AWS Glue ETL job that processes millions of small JSON files in Amazon S3. The job is slow and often fails with an OutOfMemory error on the driver. The engineer wants to improve performance without changing the output format. Which solution should the engineer implement?

A.Enable job bookmarks to track processed files and skip already-processed data.
B.Use AWS Glue's groupFiles and groupSize options to combine small files into larger groups.
C.Convert the JSON files to Parquet format using an AWS Glue crawler before running the ETL job.
D.Increase the number of DPUs allocated to the job to provide more memory per executor.
AnswerB

The groupFiles and groupSize options in AWS Glue allow the job to coalesce many small files into larger groups before processing. This reduces the number of input partitions and the memory overhead on the driver, directly addressing the OutOfMemory error and improving performance. It is the recommended approach for large numbers of small files without changing the output format.

Why this answer

Grouping small files with groupFiles and groupSize is the most effective way to handle millions of small files in AWS Glue. It reduces the number of input splits, lowers driver memory pressure, and improves overall job performance. Other options either do not address the driver memory issue or require changing the output format.

Exam trap

The trap here is assuming that adding more DPUs will always resolve OutOfMemory errors, but driver memory is not proportional to the number of DPUs.

18
MCQhard

A data team runs a daily AWS Glue ETL job that processes data from an Amazon Redshift cluster and writes results to Amazon S3. The job completes successfully but takes 2 hours longer than expected. The job uses the JDBC connection to Redshift. The Redshift cluster is 4 dc2.large nodes. The Glue job has 10 workers of type G.1X. Which change would MOST likely reduce the job duration?

A.Use Redshift Spectrum to query data directly from S3
B.Use the S3 staging option in the Glue connection to unload data from Redshift to S3 first
C.Increase the Redshift cluster size to 8 nodes
D.Increase the number of Glue workers to 20
AnswerB

The JDBC connector reads Redshift row-by-row through the driver, which is slow at this scale. Using the S3 staging option runs a Redshift UNLOAD to S3 in parallel, then Glue reads S3 directly, removing the JDBC bottleneck and cutting job duration.

Why this answer

The JDBC connection in AWS Glue reads data row-by-row from Redshift, which is slow for large datasets. By enabling the S3 staging option in the Glue connection, the job uses Redshift's UNLOAD command to export data to S3 in parallel, then Glue reads from S3. This bypasses the JDBC bottleneck and leverages Redshift's massively parallel processing (MPP) to export data much faster.

Exam trap

The trap here is that candidates assume the bottleneck is either Redshift compute (C) or Glue parallelism (D), when in fact the JDBC driver's single-threaded row-by-row fetch is the primary performance limiter.

How to eliminate wrong answers

Option A is wrong because Redshift Spectrum queries data directly from S3, but the source data is in Redshift, not S3; Spectrum does not help extract data from Redshift. Option C is wrong because the bottleneck is the JDBC connection, not Redshift compute capacity; adding more Redshift nodes would not speed up a single-threaded JDBC read. Option D is wrong because increasing Glue workers only helps if the job is CPU-bound or parallelizable; the JDBC read is I/O-bound and limited by the single connection, so more workers would not reduce the 2-hour delay.

19
MCQeasy

A data engineer is responsible for monitoring an AWS Glue ETL job that runs daily. The job reads data from an Amazon S3 bucket and writes to an Amazon Redshift table. The engineer wants to receive an alert if the job fails or if it takes longer than expected to complete. Which AWS service should the engineer use to set up these alerts with the LEAST operational overhead?

A.Amazon CloudWatch Events (EventBridge) to trigger an AWS Lambda function that sends an email via Amazon SES.
B.AWS Glue job bookmarks to track job progress and send notifications on failure.
C.AWS CloudTrail to log AWS Glue API calls and analyze logs for failures.
D.Amazon CloudWatch Alarms based on AWS Glue job metrics, with Amazon SNS notifications.
AnswerD

CloudWatch Alarms can monitor AWS Glue job metrics such as glue.driver.aggregate.elapsedTime and glue.driver.aggregate.numFailedTasks. When a threshold is breached, the alarm can publish to an SNS topic, sending notifications. This requires minimal setup and no custom code, making it the least operational overhead solution for alerting on job failures and durations.

Why this answer

Amazon CloudWatch Alarms can monitor AWS Glue job metrics and trigger Amazon SNS notifications when thresholds are exceeded. This setup requires no custom code and provides timely alerts on job failures or prolonged execution, making it the most efficient solution with minimal operational overhead.

Exam trap

The trap here is overcomplicating the solution with Lambda and SES, when CloudWatch Alarms with SNS directly provide the needed alerting.

20
MCQeasy

A data analyst needs to query a large Amazon S3 bucket containing CSV files using Amazon Athena. The bucket has millions of small files (less than 1 MB each). The analyst reports that queries are very slow and often time out. The data is partitioned by date and the partition columns are defined in the table. What is the most effective way to improve query performance?

A.Convert the files to Apache Parquet format using an AWS Glue ETL job.
B.Run a compaction job to consolidate small files into fewer larger files (e.g., 128 MB each).
C.Add more partitions by including hour and minute as partition keys.
D.Use S3 Select to push down filtering to S3 before Athena processes the data.
AnswerB

Millions of sub-1 MB files force Athena to open each object individually, so per-file overhead dominates and queries time out. Compacting into roughly 128 MB objects cuts the number of GET requests and partition listings, letting Athena scan far fewer, larger files.

Why this answer

Many small files (under 1 MB) cause high overhead because each file requires a separate read operation and metadata call. Consolidating them into fewer larger files (e.g., 128 MB each) reduces the number of read operations and improves I/O efficiency, directly addressing the root cause of slowdowns and timeouts. Option A (converting to Parquet) improves storage efficiency and query performance but does not reduce the file count; it is a beneficial addition but not the most effective standalone fix for the small file problem.

Option C (adding more partitions) would increase overhead by creating even more directories/files to scan. Option D (S3 Select) applies within individual files and does not mitigate overhead from file quantity.

21
Matchingmedium

Match each AWS data analytics service to its primary function.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Serverless SQL query on S3

Business intelligence and dashboards

Data lake setup and access control

Real-time SQL on streaming data

Query data in S3 from Redshift

Why these pairings

The correct matches are: Amazon Athena for serverless SQL querying on S3, Amazon Redshift for data warehousing, Amazon EMR for big data processing, and Amazon Kinesis for real-time streaming. Common confusions include swapping the roles of Athena and Redshift.

22
Multi-Selecteasy

A data engineer is setting up a data pipeline using AWS Glue. The engineer wants to monitor job failures and receive notifications. Which TWO services can be used together for this purpose?

Select 2 answers
A.AWS Step Functions
B.Amazon CloudWatch
C.Amazon SNS
D.Amazon Kinesis Data Streams
E.Amazon SQS
AnswersB, C

Amazon CloudWatch captures Glue job metrics and emits events on state changes such as FAILED, satisfying the requirement to detect job failures. Combined with Amazon SNS for notification delivery, it forms the monitoring half of the pair. Glue natively publishes logs and metrics to CloudWatch, so no custom instrumentation is needed.

Why this answer

Amazon CloudWatch (B) is correct because AWS Glue automatically publishes job metrics and logs to CloudWatch, and you can create CloudWatch alarms that trigger when a Glue job fails (for example, on the glue.driver.aggregate.numFailedTasks metric or a FAILED job run state). Amazon SNS (C) is correct because it is the notification service that CloudWatch alarms target via an SNS topic, delivering email, SMS, or HTTP notifications to subscribers when a Glue job failure alarm fires. Together, CloudWatch detects the failure and SNS delivers the alert, which is the standard AWS pattern for Glue job failure notifications.

AWS Step Functions (A) can orchestrate Glue jobs but is not a monitoring/notification service, Amazon Kinesis Data Streams (D) is for real-time streaming ingestion, and Amazon SQS (E) is a message queue for decoupling applications, so none of these provide the alarm-and-notify capability required here.

Exam trap

The trap here is that candidates may confuse AWS Step Functions (A) as a monitoring tool because it can orchestrate retries, but it does not natively send notifications and is not the primary service for monitoring Glue job failures.

23
MCQeasy

A data engineer is monitoring an Amazon EMR cluster and notices that the cluster is running out of disk space on the core nodes. Which action can be taken to resolve this issue?

A.Reduce the retention period of data stored on HDFS
B.Change the core node instance type to a compute-optimized type
C.Increase the EBS volume size attached to core nodes
D.Use Spot Instances for core nodes
AnswerC

Core nodes hold HDFS data and shuffle output, so exhausted local disk blocks task execution. Expanding each core node's attached EBS volume increases available block storage, directly resolving the disk-space constraint without altering cluster topology or losing HDFS data.

Why this answer

Increasing the EBS volume size attached to core nodes directly adds storage capacity, resolving the disk space issue. Option A is wrong because reducing HDFS data retention may free space but does not increase available disk space; it could result in data loss. Option B is wrong because changing to a compute-optimized instance type affects CPU and memory, not storage.

Option D is wrong because Spot Instances are a pricing model and do not add disk space.

24
Multi-Selecthard

A data engineer is using AWS Step Functions to orchestrate a daily ETL workflow that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The workflow occasionally fails and the engineer needs to troubleshoot and recover. Which TWO actions should the engineer take to identify the failure and resume from the failed step? (Choose two.)

Select 2 answers
A.Delete the state machine and recreate it with a retry policy on every state.
B.Use the Step Functions 'Redrive' capability to restart the execution from the failed state after fixing the underlying issue.
C.Enable AWS Step Functions logging to CloudWatch Logs and inspect the execution history for the failed state.
D.Configure the state machine to use an Amazon SQS dead-letter queue for failed states and reprocess messages manually.
E.Manually trigger the AWS Glue job and Amazon EMR step from the console, then restart the Step Functions execution from the beginning.
AnswersB, C

Redrive allows resuming a failed execution from the point of failure, preserving previous successful steps and their outputs. This avoids re-running the entire workflow and is the recommended way to recover after correcting the root cause. It requires the state machine to have the necessary IAM permissions and the execution to be in a failed state.

Why this answer

To troubleshoot a failed Step Functions execution, enabling logging and reviewing execution history provides the necessary error details. Once the root cause is fixed, the Redrive feature allows resuming from the failed state without re-running successful steps. These two actions together enable efficient diagnosis and recovery.

Exam trap

The trap here is assuming that restarting the entire workflow or manually rerunning individual jobs is the only recovery path, while overlooking Step Functions' native logging and Redrive capabilities.

25
MCQhard

A data engineer is using AWS Step Functions to orchestrate a daily ETL pipeline that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The pipeline occasionally fails with the error 'States.TaskFailed' from the Glue job, but the Glue job's own logs show that it completed successfully. The Step Functions execution history shows that the Glue job task timed out after 15 minutes, while the Glue job actually ran for 18 minutes. The Step Functions state machine uses the optimized Glue service integration with a TaskTimeout of 900 seconds. Which change will allow the pipeline to complete successfully without reducing the Glue job's runtime?

A.Change the Step Functions integration from the optimized Glue service integration to the AWS SDK integration with a longer timeout.
B.Add a Retry policy with a maximum of 3 attempts and an interval of 60 seconds to the Glue job task.
C.Reduce the Glue job's runtime by increasing the number of DPUs so it finishes within 15 minutes.
D.Increase the Step Functions TaskTimeout to a value greater than the Glue job's maximum expected runtime, such as 3600 seconds.
AnswerD

The Step Functions task timed out because its TaskTimeout was set to 900 seconds, but the Glue job took 18 minutes (1080 seconds). The optimized Glue service integration waits for the job to finish, and if the task exceeds the TaskTimeout, Step Functions fails the task even if the Glue job continues running. Increasing the TaskTimeout to a value larger than the job's maximum runtime, such as 3600 seconds, allows Step Functions to wait for completion without timing out.

Why this answer

The Step Functions task failed because its TaskTimeout was shorter than the Glue job's actual runtime. The optimized Glue service integration waits for the job to complete, but if the task exceeds the TaskTimeout, Step Functions marks it as failed even though the Glue job continues. Setting the TaskTimeout to a value greater than the job's maximum expected runtime ensures the state machine waits for the job to finish successfully.

Exam trap

The trap here is believing that a Retry policy or more DPUs will solve a timeout, when the real issue is that the Step Functions TaskTimeout is shorter than the job's runtime.

26
MCQeasy

A data engineer is running an AWS Glue ETL job that reads from an Amazon RDS MySQL database and writes to Amazon S3. The job fails with a 'Communications link failure' error. The security group for the RDS instance allows inbound traffic from the Glue job's security group. What is the most likely cause of the failure?

A.The JDBC connection string in the Glue job does not include the database name.
B.The Glue job is using the wrong JDBC driver.
C.The Glue job's security group does not allow outbound traffic to the RDS security group on port 3306.
D.The IAM role used by the Glue job does not have rds:Connect permission.
AnswerC

Security groups are stateful, but the Glue job's own outbound rules must still permit traffic to the RDS security group on MySQL port 3306. Inbound access alone is insufficient, so the missing egress rule explains the communications link failure.

Why this answer

AWS Glue ETL jobs run in a VPC that requires outbound security group rules to initiate connections to RDS. Even if the RDS security group allows inbound traffic from the Glue security group, the Glue security group must also have an outbound rule allowing traffic to the RDS security group on port 3306 (MySQL default port). Without this outbound rule, the TCP handshake from Glue to RDS fails, causing a 'Communications link failure'.

Exam trap

The trap here is that candidates assume only inbound rules matter for security groups, but outbound rules are equally critical for initiating connections from the client (Glue) to the server (RDS).

How to eliminate wrong answers

Option A is wrong because omitting the database name from the JDBC connection string would cause a different error (e.g., 'Unknown database' or connection rejection), not a 'Communications link failure', which indicates a network-level issue. Option B is wrong because AWS Glue automatically includes the correct JDBC driver for MySQL (compatible with Amazon RDS MySQL) when using the Glue connection type 'MySQL'; using the wrong driver would typically produce a class-not-found or driver-incompatibility error, not a communications link failure. Option D is wrong because IAM permissions for Glue jobs use actions like 'glue:GetConnection' and 'rds:DescribeDBInstances' to retrieve connection metadata, but there is no 'rds:Connect' IAM action; database authentication is handled via username/password in the Glue connection, not IAM.

27
Multi-Selecthard

A data engineer is using AWS Glue to process data from an Amazon Kinesis Data Stream. The Glue job is configured to run every 15 minutes and uses job bookmarks to track processed data. Recently, the job started reprocessing old data, leading to duplicate records in the target Amazon S3 bucket. The engineer verifies that the job bookmark is enabled and the job is not being run manually. Which TWO actions should the engineer take to resolve the duplicate processing issue? (Choose two.)

Select 2 answers
A.Check that the Glue job's bookmark key is not being changed between runs, as a changed bookmark key resets the bookmark state.
B.Increase the Kinesis stream's retention period to 168 hours to allow the Glue job more time to process data.
C.Configure the Glue job to use the '--job-bookmark-option' parameter set to 'job-bookmark-enable'.
D.Ensure that the Glue job is not being run concurrently with the same job bookmark, as concurrent runs can cause duplicate processing.
E.Verify that the Kinesis stream's shard iterator type is set to LATEST in the Glue job's connection options.
AnswersA, D

The job bookmark key is used to identify the bookmark state. If the key changes, the job treats it as a new job and starts processing from the beginning, causing duplicates. Ensuring the bookmark key remains consistent is crucial for proper bookmark functionality.

Why this answer

Duplicate processing with AWS Glue job bookmarks often occurs when the bookmark key changes or when concurrent runs interfere with bookmark updates. Verifying the bookmark key remains constant and preventing concurrent runs with the same bookmark are key steps. These actions ensure the bookmark state is preserved and correctly updated, preventing reprocessing of already-processed data.

Exam trap

The trap here is assuming that enabling job bookmarks is sufficient to prevent duplicates, when in fact bookmark key changes or concurrent runs can still cause reprocessing.

28
MCQmedium

A company uses AWS Glue to run ETL jobs on a schedule. Recently, a job failed with the error: 'AnalysisException: cannot resolve '`column_name`' given input columns: ...'. The job reads from an Amazon S3 source that has a schema defined in the AWS Glue Data Catalog. What is the MOST likely cause?

A.The schema of the source data has changed and is not reflected in the Data Catalog.
B.The source data file is corrupted and cannot be parsed.
C.The IAM role associated with the Glue job does not have permissions to read the S3 bucket.
D.The data type of the column in the source does not match the Data Catalog definition.
AnswerA

The AnalysisException arises because Spark resolves column names against the schema registered in the Data Catalog, not the actual S3 objects. If the source's schema evolved—columns renamed, added or dropped—without a crawler or manual update refreshing the catalog table, the job's query references columns the catalog no longer exposes.

Why this answer

The AnalysisException 'cannot resolve column_name given input columns' means the Spark/Glue job's code references a column that does not exist in the DataFrame schema derived from the Data Catalog. The most likely cause is that the source data's schema has evolved (new/renamed/dropped columns) but the Glue Data Catalog table definition was not updated, so the job's schema and the actual data are out of sync.

Exam trap

DEA-C01 often tests whether candidates can distinguish schema-resolution errors from permission or corruption errors, so the trap is picking IAM or file-corruption options for a clearly schema-related exception.

How to eliminate wrong answers

Option B is wrong because a corrupted file typically produces a parse/format error (e.g., 'MalformedInputException'), not a column-resolution error — Spark would fail reading the file, not resolving a column name. Option C is wrong because missing IAM permissions produce an 'Access Denied' error at the S3 read stage, not an AnalysisException about column resolution. Option D is wrong because a data-type mismatch would produce a cast/conversion error, not a 'cannot resolve column' error — the column would be found but with the wrong type.

29
MCQhard

A data engineer is troubleshooting an AWS Glue ETL job that fails with the error 'java.lang.OutOfMemoryError: Java heap space'. The job processes a large number of small files in Amazon S3. Which action would MOST effectively resolve the issue?

A.Enable S3 groupFiles option in the Glue job
B.Change the worker type to G.1X
C.Increase the number of workers in the Glue job
D.Use a G.2X worker type with more memory
AnswerA

Grouping coalesces many small S3 objects into larger input partitions, so Glue reads far fewer files and holds less per-file metadata and buffer overhead in the driver and executors. This directly relieves the Java heap exhaustion caused by the large number of small files described in the stem.

Why this answer

The 'java.lang.OutOfMemoryError: Java heap space' in AWS Glue when processing many small files is typically caused by the driver or executor accumulating too many file metadata objects. Enabling the S3 groupFiles option (with groupSize and groupFiles parameters) consolidates small files into larger groups, reducing the number of objects processed and alleviating heap pressure. This directly addresses the root cause of the memory issue.

Exam trap

DEA-C01 often tests the misconception that scaling up worker type or count solves OutOfMemory errors, when the root cause (many small files) requires the groupFiles optimization instead.

How to eliminate wrong answers

Option B is wrong because changing to G.1X workers increases memory per worker but does not reduce the number of small files being processed, so the heap pressure from file metadata remains. Option C is wrong because increasing the number of workers adds parallelism but does not solve the per-executor memory issue caused by many small files; it may even worsen driver memory pressure. Option D is wrong because G.2X workers provide more memory per worker, but again, the fundamental problem of too many small files persists, and the driver may still run out of heap.

30
MCQmedium

A data engineer is troubleshooting a failed AWS Glue job that reads from an Amazon RDS for MySQL table. The error message indicates 'java.sql.SQLException: No suitable driver'. What is the most likely cause?

A.The MySQL JDBC driver JAR is not included in the Glue job's dependencies.
B.The Glue job is using the wrong JDBC driver class name.
C.The Glue job's VPC subnet does not have a route to the RDS instance.
D.The RDS instance is not publicly accessible.
AnswerA

Glue's JDBC connection requires the MySQL driver JAR to be supplied via the job's dependent JARs path; the bundled environment lacks it, so the driver class cannot load and the connection throws 'No suitable driver'. Adding the MySQL Connector/J JAR resolves the missing driver.

Why this answer

The error 'java.sql.SQLException: No suitable driver' indicates that the Java application (the Glue job) cannot find a suitable JDBC driver for the database URL. In AWS Glue, you must include the JDBC driver JAR as a dependency for the job. If the MySQL JDBC driver JAR is not included, the job cannot load the driver class, resulting in this error.

Exam trap

DEA-C01 often tests the confusion between connectivity errors (timeouts, network) and driver errors, where 'No suitable driver' specifically points to a missing or misconfigured JDBC driver JAR.

How to eliminate wrong answers

Option B is wrong because if the wrong JDBC driver class name were used, the error would typically be 'ClassNotFoundException' or 'No suitable driver' only if the class is not found, but the more common cause is the missing JAR. Option C is wrong because a network routing issue would result in a connection timeout or 'Communications link failure', not 'No suitable driver'. Option D is wrong because if the RDS instance is not publicly accessible, the error would be a connection timeout or 'Unknown host', not a driver error.

31
MCQmedium

An AWS Glue job that performs data transformation on large Parquet files in Amazon S3 is taking a long time to complete. The job uses the default number of DPUs. Which change would most likely improve the job's performance?

A.Increase 'Max capacity' (number of DPUs) for the job.
B.Use 'coalesce' to reduce the number of output files.
C.Reduce the number of partitions in the source data.
D.Change the input format from Parquet to CSV.
AnswerA

The job is bottlenecked by compute, not partitioning or file size. Raising Max capacity adds more DPUs, distributing the Parquet transformation across additional executors and reducing wall-clock time, directly addressing the stem's constraint that the job runs on the default DPU allocation.

Why this answer

AWS Glue allocates DPUs (Data Processing Units) to run the job, and increasing Max capacity adds more compute and memory resources that can parallelize the transformation across partitions. For large Parquet files with the default DPU count, the job is likely CPU- or memory-bound, so scaling DPUs is the most direct performance improvement. Parquet is already a columnar, compressed format, so the bottleneck is compute, not I/O format.

Exam trap

DEA-C01 often tests the misconception that reducing output files or changing file formats improves Glue performance, when the real lever for compute-bound transformations is scaling DPUs and ensuring adequate partitioning.

How to eliminate wrong answers

Option B is wrong because coalesce reduces the number of output files, which can actually hurt parallelism and does not address the compute bottleneck during transformation. Option C is wrong because reducing source partitions decreases parallelism, making the job slower, not faster. Option D is wrong because switching from Parquet to CSV increases I/O volume and loses columnar pushdown benefits, degrading performance rather than improving it.

32
MCQmedium

A company is using Kinesis Data Firehose to deliver data to an S3 bucket. The delivery stream is failing with 'S3 bucket access denied' errors. The bucket policy allows the Firehose service principal. What could be the issue?

A.The S3 bucket is in a different VPC
B.The S3 bucket uses SSE-KMS and Firehose does not have KMS permissions
C.The S3 bucket name contains invalid characters
D.The IAM role assigned to Firehose lacks s3:PutObject permission
AnswerD

Correct. The IAM role assigned to the Firehose delivery stream must have the s3:PutObject permission to write objects to the S3 bucket. The bucket policy allowing the service principal is not sufficient; the role also needs the appropriate S3 action.

Why this answer

Even though the S3 bucket policy allows the Firehose service principal, Kinesis Data Firehose uses an IAM role to write data. This role must have the s3:PutObject permission. Without it, Firehose will receive an access denied error.

Option A is incorrect because VPC differences affect network connectivity, not IAM permissions. Option B is incorrect because SSE-KMS requires KMS permissions, but the error here is specifically about S3 access. Option C is incorrect because bucket name validation occurs during stream creation, not during data delivery.

Exam trap

Candidates often confuse the bucket policy and the IAM role permissions. The bucket policy allowing the service principal is necessary but not sufficient; the delivery role must also have s3:PutObject.

33
Multi-Selecthard

A data engineer is designing an Amazon Redshift data warehouse for a high-traffic analytics workload. The engineer needs to ensure fast query performance and minimize data movement. Which THREE design decisions should be made? (Choose THREE.)

Select 3 answers
A.Choose DISTSTYLE KEY for tables that are frequently joined.
B.Use the default distribution style for all tables.
C.Use DISTSTYLE ALL for all large fact tables.
D.Apply appropriate compression encodings to columns.
E.Define SORT KEYs on columns used in WHERE clauses.
AnswersA, D, E

DISTSTYLE KEY distributes rows according to a chosen column's values, so matching join keys co-locate on the same slice. For frequently joined tables, this eliminates cross-node data movement during joins, satisfying the requirement to minimise data movement and speed queries.

Why this answer

Option A is correct because choosing DISTSTYLE KEY on the columns used in frequent joins colocates matching rows on the same slice, so join processing avoids broadcasting or redistributing data across nodes. Option D is correct because applying appropriate compression encodings (for example AZ64, ZSTD, or LZO) reduces the amount of data read from disk and moved across the network, directly improving query performance. Option E is correct because defining SORT KEYs on columns used in WHERE clauses enables zone-map block skipping, so the query reads only the relevant blocks instead of scanning the whole table.

Option B is not appropriate because the default distribution style (AUTO) may not optimally colocate join keys, and relying on it for all tables can cause unnecessary data movement. Option C is wrong because DISTSTYLE ALL replicates the entire table to every node, which is suitable only for small dimension tables, not large fact tables, since it would multiply storage and load/redistribution overhead.

34
MCQmedium

A data engineer is monitoring an AWS Glue ETL job that processes data from an S3 bucket and writes to a Redshift table. The job completes successfully but takes longer than expected. The engineer notices that the job uses 10 DPUs and the data size is 500 GB. The job runs in standard mode. Which change would MOST reduce job duration?

A.Increase the number of DPUs to 20.
B.Use a smaller worker type like G.1X.
C.Change the output format from Parquet to CSV.
D.Reduce the number of partitions in the data.
AnswerA

Scaling DPUs adds parallel executors, so the 500 GB shuffle and Redshift write are processed concurrently rather than serially. With 10 DPUs the job is compute-bound in standard mode; doubling to 20 directly halves wall-clock duration, satisfying the requirement to most reduce job duration.

Why this answer

Increasing the number of DPUs from 10 to 20 allows the job to process data in parallel, reducing execution time. AWS Glue standard mode scales linearly with DPUs for ETL jobs that are not I/O bound. In this case, with 500 GB of data and 10 DPUs, the job is likely CPU-bound and can benefit from additional parallelism.

Option B is incorrect because using a smaller worker type (G.1X) reduces available memory and CPU, worsening performance. Option C is incorrect because changing the output from Parquet (columnar, compressed) to CSV (row-based, uncompressed) increases data size and I/O, slowing the job. Option D is incorrect because reducing partitions can cause data skew and reduce parallelism, increasing runtime.

Exam trap

Candidates may assume that increasing DPUs always helps, but for small datasets or I/O-bound jobs, diminishing returns occur. However, for large datasets like 500 GB in standard mode, increasing DPUs typically reduces duration linearly up to a point.

35
Multi-Selecthard

A data engineer is designing a data pipeline that ingests data from multiple sources into Amazon S3, then processes it with AWS Glue and loads it into Amazon Redshift. Which THREE practices should be implemented to ensure data quality?

Select 3 answers
A.Implement data validation checks at the ingestion stage
B.Use AWS Glue DataBrew for data profiling and schema enforcement
C.Compress data files to reduce storage costs
D.Use manual sampling to check data quality periodically
E.Set up Amazon CloudWatch alarms for pipeline failures and data anomalies
AnswersA, B, E

Validating records as they land in Amazon S3 catches malformed, incomplete or out-of-range data before Glue transforms it, preventing corrupt rows from propagating into Redshift. This satisfies the requirement to ensure data quality across the multi-source ingestion pipeline at its earliest stage.

Why this answer

Option A is correct because validating data at ingestion (for example, checking types, nulls, ranges, and referential integrity before writing to Amazon S3) prevents corrupt or malformed records from propagating downstream into Glue ETL jobs and Redshift tables. Option B is correct because AWS Glue DataBrew provides visual data profiling, column statistics, and built-in transformations/rules that can enforce schema consistency and detect anomalies before loading into Redshift. Option E is correct because Amazon CloudWatch alarms on Glue job metrics, Redshift load events, and custom anomaly metrics enable proactive detection of pipeline failures and unexpected data patterns, which is essential for ongoing data quality assurance.

Option C is not correct because compressing files reduces storage and I/O costs but does not validate or improve data accuracy, completeness, or consistency. Option D is not correct because manual sampling is ad hoc, non-scalable, and not a reliable or automated data quality control compared with systematic validation, profiling, and monitoring.

Exam trap

DEA-C01 often tests the confusion between cost optimisation (compression) and data quality practices, so candidates pick compression or manual sampling instead of automated validation, profiling, and monitoring.

36
MCQmedium

A team uses Amazon Redshift for analytics. They notice that some queries are slow and the system shows high disk usage. The team wants to improve query performance without adding more nodes. Which action should they take first?

A.Run the VACUUM and ANALYZE commands on the tables.
B.Enable compression on all columns.
C.Redistribute the tables by changing the distribution key to a column with high cardinality.
D.Modify the workload management (WLM) queue to increase concurrency.
AnswerA

VACUUM reclaims space from deleted rows and re-sorts data, while ANALYZE updates table statistics used by the query planner. Both address the high disk usage and slow queries without adding nodes, making them the correct first action.

Why this answer

VACUUM and ANALYZE are the first actions to take because they reclaim disk space from deleted rows and update table statistics, which can significantly improve query performance and reduce disk usage. Option B is incorrect because enabling compression on all columns is not a first step; compression is typically set during table creation and may not address current disk usage issues. Option C is incorrect: redistributing tables by changing the distribution key requires recreating the table, which is an invasive operation and not the first troubleshooting step.

Option D is incorrect: modifying the WLM queue affects concurrency management, not disk usage or query performance directly related to high disk usage.

37
MCQeasy

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is running but the engineer observes that some tables are not being replicated. The DMS task logs show no errors, and the task status is 'Running'. Which action should the engineer take to identify the missing tables?

A.Check the DMS task's table mappings to ensure all required tables are included in the selection rules.
B.Enable CloudWatch Logs for the DMS task and look for warnings about table selection.
C.Increase the DMS replication instance size to handle more tables concurrently.
D.Restart the DMS task with the 'Reload target' option to force a full load of all tables.
AnswerA

In AWS DMS, table mappings define which tables are included or excluded from the migration. If some tables are missing, it is likely that they are not covered by the selection rules. Reviewing and updating the table mappings to include the missing tables will ensure they are replicated. This is a common oversight when setting up DMS tasks.

Why this answer

Table mappings in AWS DMS determine which tables are migrated. If tables are missing and no errors are present, the most likely cause is that the tables are not included in the selection rules. Reviewing and correcting the table mappings is the direct way to ensure all required tables are replicated.

Exam trap

The trap here is assuming that a running task with no errors is replicating all tables, but table mappings may exclude some tables by design.

38
MCQeasy

A company is using AWS Glue to catalog data stored in Amazon S3. The data is partitioned by year, month, and day. A data analyst reports that new partitions are not automatically discovered by the Glue crawler. The crawler runs on a schedule every hour. What is the MOST likely reason for the missing partitions?

A.The IAM role used by the crawler does not have permission to list the S3 bucket.
B.The Glue Data Catalog is not configured to use a Hive metastore.
C.The number of partitions exceeds the Glue catalog limit of 100,000.
D.The crawler schedule is set to run too frequently.
AnswerA

Without s3:ListBucket permission, the crawler cannot enumerate prefixes under the bucket, so newly written year/month/day partitions remain invisible even though the hourly schedule runs. The missing IAM action is the specific cause of undiscovered partitions.

Why this answer

For a Glue crawler to discover new partitions in S3, its IAM role must have s3:ListBucket (and GetObject) permissions on the bucket and prefix. If the role lacks ListBucket, the crawler cannot enumerate the partition folders and will silently miss new partitions even though it runs on schedule. This is the most common cause of 'crawler runs but doesn't find new data' issues.

Exam trap

DEA-C01 often tests the misconception that crawler scheduling or catalog limits cause missing partitions, when the most common root cause is insufficient S3 ListBucket permission on the crawler's IAM role.

How to eliminate wrong answers

Option B is wrong because the Glue Data Catalog is a managed Hive-compatible metastore by default; it does not require an external Hive metastore to discover partitions. Option C is wrong because while Glue has soft limits on partitions per table (and 100,000 is not the relevant hard limit for this symptom), exceeding a limit would typically produce errors, not silent omission of new partitions. Option D is wrong because running the crawler more frequently does not prevent discovery — if anything, it would discover partitions sooner; schedule frequency is not the cause of missing partitions.

39
MCQeasy

A data engineer is using AWS Step Functions to orchestrate a daily pipeline that runs several AWS Glue jobs in sequence. One Glue job intermittently fails due to a transient Amazon S3 503 error. The engineer wants the state machine to automatically retry only that Glue job up to three times with exponential backoff, without retrying the other jobs. What should the engineer do?

A.Set the Glue job's MaximumRetries property to 3 in the job definition, then add a Catch field in the state machine to handle final failures.
B.Add a Retry field to the Glue job task state with ErrorEquals set to States.TaskFailed, MaxAttempts set to 3, and an appropriate IntervalSeconds and BackoffRate.
C.Configure an Amazon CloudWatch alarm on Glue job failures and use an Application Auto Scaling policy to restart the state machine execution.
D.Wrap the entire state machine in a Map state and set MaxConcurrency to 1 so that failures are retried automatically by Step Functions.
AnswerB

Step Functions Retry is configured per state, so adding it to the specific Glue job task retries only that task up to three times. Exponential backoff is achieved with IntervalSeconds and BackoffRate, and ErrorEquals can include States.TaskFailed to catch transient failures. This precisely satisfies the requirement without affecting other jobs in the sequence.

Why this answer

Step Functions retry logic is declared inside the state definition, so attaching a Retry field to the Glue job task scopes retries to that task alone. Specifying MaxAttempts of 3 with IntervalSeconds and BackoffRate produces exponential backoff for transient failures. The other choices either apply retries at the wrong layer, change pipeline structure, or use services that cannot restart executions.

Exam trap

The trap here is conflating AWS Glue job-level MaximumRetries with Step Functions state-level Retry, which are separate mechanisms with different scope and backoff behavior.

40
MCQhard

A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Redshift. The migration is successful, but after a few days, data in Redshift becomes inconsistent with the source due to ongoing changes. The company needs to keep Redshift synchronized with minimal latency. Which approach should the data engineer use?

A.Configure DMS with ongoing replication using change data capture (CDC).
B.Use Amazon Redshift COPY with S3 staging and AWS Lambda triggers.
C.Schedule a full DMS load every night.
D.Set up Amazon Redshift Spectrum to query the Oracle database directly.
AnswerA

Ongoing replication with CDC lets AWS DMS continuously apply source Oracle changes to Redshift, satisfying the requirement to keep Redshift synchronised with minimal latency after the initial full load. Without CDC, only a one-time migration occurs, so later source changes never propagate.

Why this answer

AWS DMS supports ongoing replication using change data capture (CDC), which captures incremental changes from the Oracle source (via Oracle LogMiner or binary logs) and applies them to Amazon Redshift in near real-time. This approach ensures that Redshift remains synchronized with the source database with minimal latency, meeting the requirement for ongoing consistency after the initial full load.

Exam trap

The trap here is that candidates may confuse Amazon Redshift Spectrum's federated querying capability with actual data replication, or assume that nightly batch loads (Option C) are sufficient for 'minimal latency' requirements, when DMS CDC is the only option that provides continuous, low-latency synchronization.

How to eliminate wrong answers

Option B is wrong because Amazon Redshift COPY with S3 staging and AWS Lambda triggers requires manual or event-driven extraction of data from Oracle, which introduces latency and complexity, and does not provide native CDC-based continuous replication. Option C is wrong because scheduling a full DMS load every night would result in significant data loss between loads (up to 24 hours of inconsistency) and does not achieve minimal latency. Option D is wrong because Amazon Redshift Spectrum queries external data directly from Oracle via federated querying, but it does not replicate or synchronize data into Redshift; it only provides a query-time view, which incurs high latency and does not maintain a consistent local copy.

41
MCQhard

A company runs a time-series forecasting model that writes results to an S3 bucket every 5 minutes. A downstream ETL job reads this data, but sometimes fails because it encounters incomplete files (zero bytes). What is the MOST reliable way to ensure the ETL job only processes complete files?

A.Set an S3 Lifecycle policy to delete files smaller than 1 MB.
B.Use S3 Copy to move files to a 'processed' folder after the ETL job reads them.
C.Configure S3 Select to query the files and only return rows if the file is complete.
D.Use S3 Event Notifications to trigger a Lambda function that checks file size and then moves the file to a 'ready' prefix.
AnswerD

S3 Event Notifications fire on object creation, letting a Lambda verify the object is non-zero before relocating it to a ready prefix. The ETL job then reads only that prefix, so it never encounters the zero-byte partial writes.

Why this answer

S3 Event Notifications can trigger a Lambda function upon object creation, which can check the file size (e.g., > zero bytes) and then copy the file to a 'ready' prefix, ensuring the ETL job only processes complete files. Option A is wrong because an S3 Lifecycle policy can delete small files but does not prevent the ETL from reading incomplete files. Option B is wrong because S3 Copy does not verify completeness.

Option C is wrong because S3 Select still reads the file even if it is incomplete; it doesn't guarantee completeness.

42
MCQhard

A company uses Amazon DynamoDB as the primary data store for a high-traffic application. Recently, read latency has increased significantly. The DynamoDB table has on-demand capacity mode. Which action is MOST effective to reduce read latency?

A.Add a DynamoDB Accelerator (DAX) cluster in front of the table
B.Switch the table to provisioned capacity mode with higher read capacity
C.Increase the read capacity units in the table's auto scaling settings
D.Enable DynamoDB Global Tables to distribute reads across regions
AnswerA

DynamoDB Accelerator is an in-memory cache that serves eventually consistent reads in microseconds, absorbing the repeated read traffic that drives latency on the on-demand table. Placing DAX in front reduces read latency by orders of magnitude without changing the table's capacity mode.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache purpose-built for DynamoDB that reduces read latency from single-digit milliseconds to microseconds for eventually consistent reads. For a high-traffic, read-heavy application on on-demand mode, adding a DAX cluster in front of the table is the most direct and effective way to cut read latency without changing capacity mode or table architecture.

Exam trap

DEA-C01 often tests the distinction between throughput scaling (RCUs, on-demand) and latency reduction (DAX, caching), so candidates who equate 'more capacity' with 'lower latency' pick the wrong option.

How to eliminate wrong answers

Option B is wrong because switching to provisioned mode with higher RCUs does not reduce per-request latency — it only changes capacity management and can even cause throttling if misconfigured; on-demand already scales automatically. Option C is wrong because auto scaling settings apply to provisioned mode, not on-demand, and increasing RCUs addresses throughput, not latency. Option D is wrong because Global Tables replicate data across regions for disaster recovery and local read performance, but they do not reduce latency for reads in the primary region and add replication cost and complexity.

43
Multi-Selectmedium

A data engineer is monitoring an AWS Glue ETL job that intermittently fails with 'Container killed by YARN for exceeding memory limits' during a large shuffle stage. The job reads from Amazon S3, performs a groupByKey aggregation, and writes to Amazon S3. The engineer wants to reduce the chance of executor memory exhaustion without changing the source data. (Choose two.)

Select 2 answers
A.Enable the Glue job bookmark to skip previously processed files.
B.Replace the groupByKey operation with reduceByKey or aggregateByKey to combine values before the shuffle.
C.Change the output write format from Parquet to uncompressed JSON.
D.Set the job's max concurrent runs to 1 to avoid overlapping executions.
E.Increase the number of workers allocated to the Glue job.
AnswersB, E

groupByKey shuffles all values for a key before aggregation, which can produce enormous intermediate data. reduceByKey and aggregateByKey perform map-side combiners that pre-aggregate values locally before the shuffle, dramatically shrinking the data moved across the network. Less shuffled data means smaller per-executor memory footprints during the aggregation, addressing the root cause of the container kills.

Why this answer

The container kill originates from a single executor holding too much data during the shuffle. Adding workers spreads the shuffle across more containers, reducing per-executor memory demand. Replacing groupByKey with reduceByKey or aggregateByKey introduces map-side combining so far less data crosses the network.

Both changes reduce peak memory in the shuffle stage without touching the source data.

Exam trap

The trap here is attributing an executor memory kill to total job volume rather than to how much data a single shuffle partition must hold.

44
MCQhard

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is ongoing, and the engineer notices that some transactions are not being replicated to the target. The DMS task is configured for full load plus change data capture (CDC). Which action should the engineer take to troubleshoot the missing transactions?

A.Check the DMS task's CloudWatch logs for errors related to unsupported data types or DDL changes.
B.Enable AWS DMS validation to compare source and target rows and identify discrepancies.
C.Modify the DMS task to use 'Drop and create' target table preparation mode.
D.Increase the DMS replication instance size to handle the transaction volume.
AnswerA

AWS DMS logs errors for issues like unsupported data types, DDL changes, or transformation failures in CloudWatch Logs. These errors can cause specific transactions to be skipped during CDC. Reviewing the logs is the first step to identify why transactions are missing. The logs provide detailed messages that help pinpoint the root cause, such as a data type not supported by the target.

Why this answer

When AWS DMS CDC skips transactions, it is often due to errors such as unsupported data types or DDL changes. These errors are logged in CloudWatch Logs. Reviewing the logs allows the engineer to identify the specific cause and take corrective action, such as excluding the problematic table or adjusting the task settings.

This is the most direct troubleshooting step.

Exam trap

The trap here is assuming that performance tuning or validation will fix missing transactions, when the real issue is often logged errors that need to be addressed.

45
Multi-Selectmedium

A data engineer needs to ensure that sensitive data stored in Amazon S3 is encrypted at rest. Which TWO options meet this requirement? (Choose TWO.)

Select 2 answers
A.Server-Side Encryption with AWS KMS-Managed Keys (SSE-KMS)
B.Server-Side Encryption with S3-Managed Keys (SSE-S3)
C.Using a VPC to restrict network access
D.Enabling MFA Delete on the S3 bucket
E.Client-Side Encryption with SSL/TLS
AnswersA, B

SSE-KMS encrypts objects at rest using keys managed in AWS KMS, satisfying the encryption-at-rest requirement. It provides envelope encryption with an independent data key per object, plus audit trails via CloudTrail and granular access control through key policies — unlike SSE-S3, which offers no separate key permissions or key-usage auditing.

Why this answer

Options A (SSE-KMS) and B (SSE-S3) are correct because both are server-side encryption mechanisms that encrypt S3 objects at rest: SSE-KMS uses AWS KMS customer master keys (CMKs) to generate and manage data keys, while SSE-S3 uses AES-256 keys fully managed by Amazon S3. Both satisfy the requirement that sensitive data stored in S3 be encrypted at rest, and each is applied per-object when the object is written to the bucket. Option C is incorrect because a VPC only controls network-level access to S3 (via endpoints and policies) and does not encrypt data at rest.

Option D is incorrect because MFA Delete only adds an authentication requirement for deleting objects or changing versioning state; it provides no encryption. Option E is incorrect because SSL/TLS encrypts data in transit, not at rest, and client-side encryption is a separate approach not represented by that option.

Exam trap

The trap here is that candidates often confuse encryption in transit (SSL/TLS) with encryption at rest, or they mistakenly think network controls like VPCs or access controls like MFA Delete provide data encryption, when they only address different security domains.

46
Multi-Selecteasy

A data engineer is monitoring Amazon CloudWatch metrics for an Amazon Redshift cluster and notices high CPU utilization. The engineer wants to reduce CPU usage. Which TWO actions should the engineer take?

Select 2 answers
A.Enable concurrency scaling to offload read queries to additional clusters.
B.Increase the number of nodes in the cluster.
C.Optimize the table design by using sort keys and compression.
D.Run the VACUUM command on all tables.
E.Enable audit logging to monitor queries.
AnswersA, C

Concurrency scaling adds transient Redshift clusters that serve read queries, so eligible read workloads move off the main cluster and its CPU utilisation drops. This directly addresses the high CPU constraint by offloading read query processing rather than resizing or tuning the existing cluster.

Why this answer

Option A is correct because concurrency scaling automatically adds transient cluster capacity to handle bursts of read (SELECT) queries, offloading that work from the main cluster and thereby reducing its CPU utilization. Option C is correct because choosing appropriate sort keys and compression encodings reduces the amount of data scanned and I/O performed, which directly lowers CPU work during query execution. Option B is not the best fit because adding nodes increases capacity but does not address inefficient queries or design, and it is a costlier, less targeted remedy for CPU pressure.

Option D is not correct because VACUUM reclaims space and re-sorts rows after deletes/updates; it is maintenance rather than a primary CPU-reduction action and can itself consume significant resources. Option E is not correct because audit logging only records connection and user activity for monitoring/compliance and does nothing to reduce CPU usage.

Exam trap

DEA-C01 often tests whether candidates confuse scaling out (adding nodes) with optimizing query and table design, so 'add more nodes' looks appealing but does not reduce CPU utilization per node.

47
MCQhard

A data pipeline uses AWS Glue ETL jobs to process data from Amazon RDS for MySQL to Amazon S3. Recently, the jobs have been failing with the error 'Communications link failure' during the connection phase. The RDS instance is in a private subnet, and the Glue job uses a VPC endpoint for S3. What is the most likely cause?

A.The RDS database has reached the maximum number of connections.
B.The Glue job does not have IAM permissions to decrypt the RDS database using AWS KMS.
C.The JDBC driver used by Glue is incompatible with the MySQL version.
D.The Glue job does not have a network path to the RDS instance because it is not attached to the same VPC subnet.
AnswerD

AWS Glue jobs run inside a VPC only when attached to a subnet; without that attachment they cannot route to a private RDS instance. The S3 VPC endpoint does not provide a path to RDS, so the connection fails during the handshake.

Why this answer

The most likely cause is that the AWS Glue job does not have a network path to the RDS instance because it is not attached to the same VPC subnet. AWS Glue jobs run in a serverless environment that, by default, is outside the customer's VPC. To connect to a private RDS instance, the Glue job must be configured with a VPC connection that specifies the VPC, subnet, and security group, allowing it to reach the RDS instance.

Without this, the JDBC connection fails with 'Communications link failure' during the connection phase.

Exam trap

DEA-C01 often tests the misconception that enabling a VPC endpoint for S3 automatically provides network access to other VPC resources like RDS, causing candidates to overlook the need for a separate VPC connection for Glue.

How to eliminate wrong answers

Option A is wrong because reaching the maximum number of connections would typically produce a 'Too many connections' error, not 'Communications link failure'; also, the error occurs during the connection phase, which suggests a network issue rather than connection exhaustion. Option B is wrong because IAM permissions to decrypt the RDS database using KMS are not required for a JDBC connection; KMS decryption is for encrypted data at rest, and the error would be an access denied or encryption-related error, not a communications link failure. Option C is wrong because JDBC driver incompatibility would typically manifest as a version mismatch error or a specific SQL exception, not a generic 'Communications link failure'; moreover, AWS Glue includes compatible JDBC drivers for MySQL.

48
MCQhard

A data engineer is troubleshooting a failed AWS Glue job that reads from an Apache Hive metastore in an Amazon EMR cluster. The error message indicates 'ClassNotFoundException: org.apache.hadoop.hive.ql.metadata.HiveException'. The Glue job uses a custom Python shell script. What is the most likely cause of this error?

A.Check the network connectivity between Glue and the EMR cluster.
B.Include the Hive JAR files in the 'Python library path' or use a Glue version with Hive support.
C.Modify the Python script to import the Hive libraries manually.
D.Update the IAM role to allow 'hive:Describe*' actions.
AnswerB

The ClassNotFoundException for the Hive metastore class means the Hive client JARs are absent from the Python shell job's classpath. Adding the Hive JARs to the Python library path, or using a Glue version bundling Hive support, supplies the missing classes.

Why this answer

The 'ClassNotFoundException' for a Hive class indicates the Hive JARs are not on the classpath at runtime. AWS Glue's Python shell jobs run in an environment that does not include Hive libraries by default, so the engineer must either add the Hive JARs to the Python library path or use a Glue version/configuration that bundles Hive support. This is a classpath/dependency issue, not a network or IAM problem.

Exam trap

DEA-C01 often tests the confusion between dependency/classpath errors and network or IAM errors — candidates see 'Hive' and reach for IAM or connectivity fixes when the error is a missing JAR on the classpath.

How to eliminate wrong answers

Option A is wrong because a network connectivity failure would produce a connection timeout or 'UnknownHostException', not a 'ClassNotFoundException' — the JVM cannot even find the class locally. Option C is wrong because manually importing Hive libraries in Python does not resolve a missing JAR; the class must be present on the JVM classpath, and Python imports cannot load Java classes that are absent. Option D is wrong because 'hive:Describe*' is not a valid IAM action for Hive metastore access, and IAM permission errors manifest as 'AccessDeniedException', not 'ClassNotFoundException'.

49
MCQmedium

A data engineer maintains an AWS Glue job that reads JSON files from Amazon S3, applies a transform, and writes Parquet to a second bucket. The job's bookmark was enabled at creation, but each nightly run reprocesses all previously handled files, and downstream tables now contain duplicate rows. The job script has not been modified and the S3 prefix is unchanged. Which action will MOST directly resolve the duplicate processing?

A.Reprocess the prefix with a job that has job bookmarks disabled and rely on the S3 object LastModified timestamp to filter files.
B.Confirm the transformation_ctx parameter is passed to each source and sink call, then reset the job bookmark and rerun once to rebuild state.
C.Increase the number of AWS Glue DPUs allocated to the job so the run completes before the next scheduled trigger.
D.Change the job's output write mode to append and add a deduplication step that drops rows whose keys already exist.
AnswerB

Job bookmarks rely on the transformation_ctx value to namespace state per source and sink; when it is absent or has changed between runs, Glue cannot correlate prior state and falls back to reading everything. Supplying a stable transformation_ctx on the S3 source and sink, then resetting the bookmark to clear stale state, restores correct incremental processing.

Why this answer

AWS Glue job bookmarks persist per-source and per-sink state keyed by the transformation_ctx argument. If that context is missing or inconsistent, the job cannot determine which objects were already processed and re-reads the entire prefix, producing duplicates. Passing a stable transformation_ctx and resetting the bookmark to rebuild the state store fixes the incremental behavior without changing the transform logic.

Exam trap

The trap here is assuming duplicated output always means the transformation is non-deterministic, when the usual cause is bookmark state that cannot be matched to a source or sink.

50
MCQmedium

A data engineer notices that an AWS Glue ETL job processing data from Amazon S3 to Amazon Redshift has been failing intermittently with the error 'S3ServiceException: SlowDown'. Which action is MOST likely to resolve this issue?

A.Increase the number of partitions in the Glue job to parallelize reads.
B.Switch from a Standard to a G.2X large Glue worker type.
C.Implement exponential backoff and retry logic in the Glue job.
D.Enable S3 Transfer Acceleration on the source bucket.
AnswerC

S3 SlowDown is a throttling response signalling too many concurrent requests to a partition. Exponential backoff with retries spaces requests progressively, letting the Glue job ride out the throttle rather than failing, which resolves the intermittent S3ServiceException errors.

Why this answer

The 'S3ServiceException: SlowDown' error indicates that the AWS Glue job is making requests to Amazon S3 at a rate that exceeds the bucket's request rate limits. Implementing exponential backoff and retry logic (option C) is the most effective solution because it reduces the effective request rate by introducing delays between retries, allowing S3 to recover from throttling. Option A is incorrect because increasing partitions would likely increase the number of concurrent requests, exacerbating throttling.

Option B is incorrect because switching to a larger worker type does not affect the rate of S3 requests. Option D is incorrect because S3 Transfer Acceleration improves network transfer speed but does not reduce request throttling.

51
MCQhard

A company runs an Amazon Redshift cluster for analytics. During peak hours, query performance degrades significantly. The data engineer notices that disk space usage is above 80% on many nodes. Which of the following is the MOST effective long-term solution to improve query performance?

A.Increase the workload management (WLM) queue slots.
B.Resize the cluster to include additional nodes.
C.Apply compression encoding to all columns.
D.Run the VACUUM command to reclaim space.
AnswerB

Adding nodes increases both storage capacity and compute parallelism across the cluster, relieving the 80% disk pressure while distributing query workload. This addresses the root cause of degradation rather than temporarily masking it, providing a durable performance improvement for peak-hour analytics.

Why this answer

Resizing the cluster to include additional nodes increases both storage and compute capacity, directly addressing the high disk usage and improving query performance. Increasing WLM queue slots (Option A) only manages concurrency but does not add capacity. Compression encoding (Option C) reduces storage but may not alleviate immediate performance degradation, and is not a long-term solution for capacity.

Running VACUUM (Option D) reclaims space from deleted rows but does not add new capacity.

52
MCQeasy

A company runs an Amazon RDS for PostgreSQL database and wants to capture change data (inserts, updates, deletes) to stream into Amazon Kinesis Data Streams for real-time processing. Which AWS service should be used to capture the changes directly from the database?

A.Amazon RDS automated snapshots
B.AWS Glue ETL job scheduled to run every minute
C.Amazon Kinesis Agent
D.AWS Database Migration Service (DMS) with ongoing replication
AnswerD

DMS ongoing replication reads the PostgreSQL write-ahead log via logical replication slots, capturing inserts, updates and deletes as they occur and streaming them to Kinesis. This satisfies the requirement to capture change data directly from the database without application changes.

Why this answer

AWS DMS with ongoing replication (change data capture) is the correct service because it can continuously capture insert, update, and delete operations from the PostgreSQL transaction logs (WAL) and stream them to a Kinesis Data Streams endpoint. This allows real-time processing without modifying the source database or requiring application-level triggers.

Exam trap

The trap here is that candidates confuse scheduled polling (Glue) or file-based agents (Kinesis Agent) with true CDC, failing to recognize that only DMS ongoing replication can stream row-level changes directly from the database transaction log in real time.

How to eliminate wrong answers

Option A is wrong because Amazon RDS automated snapshots are point-in-time backups of the entire database, not a mechanism to capture individual row-level changes in real time. Option B is wrong because an AWS Glue ETL job scheduled every minute introduces at least 60 seconds of latency and cannot capture every single change as it happens, making it unsuitable for true real-time streaming. Option C is wrong because Amazon Kinesis Agent is designed to stream log files (e.g., from EC2 instances) to Kinesis, not to connect directly to a database and read transactional changes from its WAL.

53
MCQeasy

A data engineer runs a Spark job on Amazon EMR that reads data from Amazon S3 and writes results back to S3. The job fails with an 'S3AccessDenied' error. The engineer verifies that the IAM role attached to the EMR cluster has s3:GetObject and s3:PutObject permissions on the relevant buckets. What is the MOST likely cause of the error?

A.S3 Transfer Acceleration is not enabled on the bucket.
B.EMRFS consistent view is not configured.
C.The S3 bucket is in a different AWS Region than the EMR cluster.
D.The IAM role does not have s3:ListBucket permission on the bucket.
AnswerD

Spark's S3A filesystem lists the bucket or prefix before reading and writing objects, and that listing call requires s3:ListBucket on the bucket resource. GetObject and PutObject alone are insufficient, so the missing ListBucket permission causes the S3AccessDenied failure.

Why this answer

The IAM role attached to the EMR cluster must have the s3:ListBucket permission on the bucket to allow the Spark job to enumerate objects when reading from S3. Without this permission, even with s3:GetObject and s3:PutObject, the job fails with an 'S3AccessDenied' error because the S3 list operation is required for directory listing and file discovery.

Exam trap

The trap here is that candidates often assume GetObject and PutObject are sufficient for S3 read/write operations, overlooking that the ListBucket permission is required for directory listing and file discovery in Spark jobs.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration is a feature for faster uploads over long distances and is not required for basic read/write operations; its absence does not cause an access denied error. Option B is wrong because EMRFS consistent view is a consistency mechanism for eventually consistent S3 buckets, not a permission or access control feature; its absence would not produce an S3AccessDenied error. Option C is wrong because while cross-region access can cause latency or additional costs, it does not inherently cause an access denied error as long as the IAM role has the correct permissions and the bucket policy allows cross-region access.

54
MCQeasy

A data engineer receives an alert that a Kinesis Data Stream has a 'WriteProvisionedThroughputExceeded' error. The stream has 5 shards with 1 MB/s write capacity per shard. The producer application is sending data at 8 MB/s sustained. What should the engineer do to resolve the issue?

A.Reduce the record size to below 1 MB per record.
B.Enable enhanced fan-out on the stream.
C.Increase the number of shards from 5 to 10.
D.Use Kinesis Firehose as an intermediary to buffer data.
AnswerC

The producer writes 8 MB/s while five shards provide only 5 MB/s aggregate, causing WriteProvisionedThroughputExceeded. Doubling to ten shards raises capacity to 10 MB/s, comfortably absorbing the sustained 8 MB/s load and clearing the error.

Why this answer

The 'WriteProvisionedThroughputExceeded' error indicates that the total write throughput to the Kinesis Data Stream exceeds the provisioned capacity. With 5 shards, each offering 1 MB/s write capacity, the total write capacity is 5 MB/s. The producer is sending 8 MB/s, which is above this limit.

Increasing the number of shards to 10 raises the total write capacity to 10 MB/s, accommodating the sustained 8 MB/s throughput and resolving the throttling.

Exam trap

The trap here is that candidates confuse write-side throttling with read-side limitations, leading them to choose enhanced fan-out (a read-side optimization) instead of scaling shards to increase write capacity.

How to eliminate wrong answers

Option A is wrong because reducing record size below 1 MB does not address the throughput limit; the error is about aggregate write throughput exceeding shard capacity, not individual record size limits. Option B is wrong because enhanced fan-out is a feature for increasing read throughput (up to 2 MB/s per shard per consumer) and does not affect write capacity or resolve write-side throttling. Option D is wrong because Kinesis Firehose is a delivery service that reads from a Kinesis stream; it cannot buffer data before it is written to the stream, so it does not solve the write throughput exceedance at the producer side.

55
MCQhard

A data engineer is designing a solution to move data from an on-premises Oracle database to Amazon S3 using AWS DMS. The engineer needs to ensure that data changes are replicated continuously with minimal latency. Which DMS configuration is most appropriate?

A.Use AWS SCT to convert the schema and then use DMS for full load
B.Use a full-load task with ongoing replication (CDC)
C.Use a full-load task that runs daily
D.Use Kinesis Data Streams to capture changes and write to S3
AnswerB

Full load plus ongoing replication (CDC) migrates existing rows and then continuously applies change data capture records from the Oracle redo logs, keeping S3 synchronised with minimal latency. A full-load-only task would leave subsequent changes unreplicated.

Why this answer

DMS with continuous replication (CDC) captures ongoing changes with low latency. Option A is wrong because AWS SCT is used for schema conversion, not data movement; DMS with full load only does an initial copy. Option C is wrong because a daily full-load task does not provide continuous replication or low latency.

Option D is wrong because it describes Kinesis Data Streams, which is not a DMS configuration—DMS itself supports CDC to S3.

56
MCQhard

A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream and writes results to an S3 bucket. The application is consistently running out of memory and failing. The operator has already increased the Parallelism and TaskManager memory. What is the next BEST step to troubleshoot?

A.Change the processing mode from exactly-once to at-least-once
B.Reduce the number of shards in the source stream
C.Enable Apache Flink metrics in Amazon CloudWatch to monitor heap and checkpoint details
D.Increase the buffer timeout for the S3 sink
AnswerC

Enabling Flink metrics in CloudWatch exposes heap usage, garbage-collection pressure and checkpoint behaviour, revealing whether state growth or backpressure—not raw parallelism—causes the out-of-memory failures. Since TaskManager memory and parallelism were already raised without success, these per-operator metrics identify the actual bottleneck before further resource changes.

Why this answer

When a Flink application on Kinesis Data Analytics runs out of memory despite increased parallelism and TaskManager memory, the next best step is to enable Apache Flink metrics in CloudWatch to observe heap usage, garbage collection, and checkpoint behavior. These metrics reveal whether the issue is memory leaks, backpressure, or state growth. Without observability, further tuning is guesswork.

Exam trap

DEA-C01 often tests the impulse to keep scaling resources (parallelism, memory) without first enabling metrics, when the correct troubleshooting step is to gain visibility into heap and checkpoint behavior.

How to eliminate wrong answers

Option A is wrong because changing from exactly-once to at-least-once affects checkpointing semantics and may reduce overhead slightly, but it does not diagnose or resolve the root cause of memory exhaustion. Option B is wrong because reducing source shards lowers parallelism and throughput, potentially worsening the problem and not addressing memory pressure. Option D is wrong because increasing the S3 sink buffer timeout only affects write batching, not the application's memory footprint or the underlying cause of OOM failures.

57
MCQeasy

A data engineer notices that an Amazon RDS for PostgreSQL instance's CPU utilization is consistently above 90% during business hours. The database is used for reporting queries. Which action should be taken FIRST to improve performance?

A.Enable Multi-AZ deployment for automatic failover.
B.Enable Performance Insights and review slow queries.
C.Create a read replica to offload reporting queries.
D.Increase the instance size to a larger instance class.
AnswerB

Performance Insights identifies the specific SQL statements and wait events driving high CPU, so slow reporting queries can be targeted rather than guessed at. This diagnostic step precedes changes such as indexing or instance scaling, satisfying the requirement to act first on evidence.

Why this answer

The first step in diagnosing high CPU utilization on an RDS for PostgreSQL instance used for reporting queries is to identify the root cause. Enabling Performance Insights provides a detailed view of database load, wait events, and SQL query performance, allowing the data engineer to pinpoint slow or inefficient queries that are consuming CPU resources. Without this diagnostic data, any other action would be premature and could lead to unnecessary cost or complexity.

Exam trap

The trap here is that candidates often jump to scaling solutions (like increasing instance size or adding a read replica) without first diagnosing the root cause, but AWS emphasizes observability and optimization before capacity changes.

How to eliminate wrong answers

Option A is wrong because enabling Multi-AZ deployment improves availability and failover, not performance; it does not reduce CPU utilization or address query performance issues. Option C is wrong because creating a read replica offloads read traffic but does not fix the underlying inefficient queries that are causing high CPU on the source instance; the replica would also suffer from the same workload if queries are poorly optimized. Option D is wrong because increasing the instance size may temporarily mask the problem by providing more CPU capacity, but it does not resolve the root cause of inefficient queries and incurs higher costs without guaranteeing sustained performance improvement.

58
MCQeasy

A data engineer has an AWS Glue job that processes data from an Amazon S3 bucket and writes to an Amazon Redshift cluster. The job is scheduled to run daily. Recently, the job started failing with the error: 'java.sql.SQLException: [Amazon](500310) Invalid operation: Spectrum Scan Error: S3 Access Denied'. The engineer verifies that the IAM role associated with the Glue job has full access to the S3 bucket. What is the most likely cause of this error?

A.The Glue job's script is using an incorrect S3 path for the Redshift COPY command.
B.The Amazon Redshift cluster's IAM role lacks permission to access the S3 bucket.
C.The Amazon Redshift cluster is in a different AWS Region than the S3 bucket, causing access latency.
D.The S3 bucket policy explicitly denies access to the Glue job's IAM role.
AnswerB

When AWS Glue writes to Amazon Redshift using the COPY command, Redshift itself must access the S3 bucket. If the Redshift cluster's IAM role does not have the necessary S3 permissions, Redshift cannot read the data, resulting in a Spectrum Scan Error with S3 Access Denied. The Glue job's role permissions are separate from Redshift's role.

Why this answer

The error occurs because Amazon Redshift, not AWS Glue, is attempting to access S3 during the COPY operation. The Redshift cluster's IAM role must have permissions to read the S3 bucket. Even if the Glue job's role has access, Redshift uses its own role.

The engineer should attach a policy to the Redshift cluster's IAM role granting S3 read access.

Exam trap

The trap here is assuming that the Glue job's IAM role is the only one involved, but Redshift uses its own IAM role for COPY commands.

59
Multi-Selectmedium

A data engineer is troubleshooting an AWS Glue ETL job that fails with the error: 'An error occurred while calling o123.pyWriteDynamicFrame. Access Denied when writing to S3 bucket: my-bucket'. The job uses a Glue service role named 'GlueServiceRole'. Which TWO actions should the engineer take to resolve the issue? (Choose TWO.)

Select 2 answers
A.Disable S3 Block Public Access on the bucket.
B.Grant the GlueServiceRole permission to write to the AWS Glue Data Catalog.
C.Check if the S3 bucket policy denies access from the GlueServiceRole.
D.Verify that the IAM policy attached to GlueServiceRole includes s3:PutObject on the bucket.
E.Ensure the Glue job is in the same VPC as the S3 bucket.
AnswersC, D

An explicit Deny in the bucket policy overrides any Allow in the identity policy, producing Access Denied on s3:PutObject. Inspecting the bucket policy for a deny targeting GlueServiceRole identifies that blocking statement, satisfying the need to find the actual cause.

Why this answer

The error 'Access Denied when writing to S3 bucket' during pyWriteDynamicFrame indicates an S3 authorization failure for the Glue job's execution role, so option D is correct: the engineer must verify that the IAM policy attached to GlueServiceRole includes s3:PutObject (and typically s3:PutObjectAcl) on the target bucket/prefix, since Glue writes output objects to S3 using that role. Option C is also correct because an explicit Deny in the S3 bucket policy overrides any Allow in the IAM policy, so the engineer must check whether the bucket policy denies access from GlueServiceRole. Option A is wrong because disabling S3 Block Public Access is unrelated to a role-based write failure and would weaken security without fixing the permission issue.

Option B is wrong because the failure is writing to S3, not to the Glue Data Catalog, so Data Catalog permissions would not resolve the S3 Access Denied error. Option E is wrong because S3 is accessed via AWS public service endpoints and does not require the Glue job to be in the same VPC as the bucket; VPC placement only matters for private endpoint configurations, not for this authorization error.

Exam trap

The trap here is that candidates may confuse S3 access errors with network or VPC issues, but S3 is a global service and access is governed by IAM and bucket policies, not VPC placement.

60
MCQeasy

A data engineer needs to monitor an AWS Glue ETL job that runs daily. The job sometimes fails due to missing partitions in the Data Catalog. The engineer wants to receive an alert when the job fails. What is the MOST operationally efficient way to achieve this?

A.Enable AWS Glue job bookmarks and configure an Amazon CloudWatch Logs subscription filter to email the engineer on error patterns.
B.Use AWS Glue job event notifications via Amazon EventBridge to trigger an AWS Lambda function that sends an Amazon SNS notification.
C.Create an Amazon CloudWatch alarm on the Glue job's 'glue.driver.aggregate.numFailedTasks' metric and notify an Amazon SNS topic.
D.Schedule a daily AWS Lambda function that calls the Glue GetJobRuns API and sends an email if the last run failed.
AnswerB

AWS Glue emits job state change events to Amazon EventBridge. You can create an EventBridge rule that matches Glue job state changes (e.g., FAILED, TIMEOUT) and targets a Lambda function or SNS topic directly. This is a serverless, operationally efficient approach that requires no polling and provides near-real-time alerts on job failures, including those caused by missing partitions.

Why this answer

Amazon EventBridge can capture AWS Glue job state change events and route them to targets like Lambda or SNS. This provides immediate, event-driven notifications on job failures without polling or custom log parsing. It is the most operationally efficient method for alerting on Glue job failures, including those due to missing partitions.

Exam trap

The trap here is assuming that CloudWatch metrics or log filters are the primary way to alert on Glue job failures, when EventBridge events are the most direct and efficient mechanism.

61
MCQeasy

A company stores sensitive data in Amazon S3. To meet compliance requirements, they need to ensure that any data older than 1 year is automatically moved to a lower-cost storage class. Which S3 feature should they use?

A.S3 Replication
B.S3 Lifecycle policies
C.S3 Glacier
D.S3 Intelligent-Tiering
AnswerB

S3 Lifecycle policies define transition rules that automatically move objects to a lower-cost storage class once they reach a specified age, such as 365 days. This satisfies the stem's compliance requirement that data older than one year be moved automatically without manual intervention.

Why this answer

S3 Lifecycle policies let you define rules that automatically transition objects to a lower-cost storage class (e.g., S3 Standard-IA, S3 Glacier Instant Retrieval, S3 Glacier Flexible Retrieval) after a specified age, such as 365 days. This is the native, declarative mechanism for age-based storage class transitions and satisfies the compliance requirement without custom code.

Exam trap

DEA-C01 often tests the difference between a storage class (Glacier) and the feature that moves data into it (Lifecycle policy) — the trap is choosing 'S3 Glacier' as if it were an automated tiering mechanism.

How to eliminate wrong answers

Option A is wrong because S3 Replication copies objects to another bucket or Region for durability/DR or latency — it does not change storage class based on age and does not reduce cost by itself. Option B is correct. Option C is wrong because S3 Glacier is a storage class, not a feature — you cannot 'use S3 Glacier' to move data automatically; you must use a lifecycle policy (or Intelligent-Tiering) to transition objects into Glacier.

Option D is wrong because S3 Intelligent-Tiering automatically moves objects between access tiers based on changing access patterns, not based on a fixed age threshold like 1 year, and it charges a monitoring fee — it does not meet a deterministic 'older than 1 year' compliance rule.

62
Multi-Selecteasy

A data engineer is setting up a data pipeline to ingest streaming data from an IoT fleet. The data must be processed in near real-time and stored in Amazon S3 for analytics. Which THREE AWS services should the engineer consider using?

Select 3 answers
A.Amazon EMR
B.Amazon Kinesis Data Firehose
C.AWS Lambda
D.AWS Glue
E.Amazon Kinesis Data Streams
AnswersB, C, E

Amazon Kinesis Data Firehose satisfies the near real-time ingestion and S3 delivery constraints by buffering streaming records and writing them directly to Amazon S3 without custom consumer code. It handles scaling, batching and format conversion automatically, so IoT telemetry lands in S3 for analytics with minimal operational overhead.

Why this answer

Amazon Kinesis Data Firehose (B) is correct because it is a fully managed service designed to reliably load streaming data directly into Amazon S3 (and other destinations) with near real-time delivery, requiring no server management. Amazon Kinesis Data Streams (E) is correct because it ingests and buffers high-throughput streaming data from IoT fleets in real time, allowing custom consumers to process the data before it is stored in S3. AWS Lambda (C) is correct because it can be invoked by Kinesis to process streaming records in near real-time, enabling serverless transformation or enrichment of the IoT data before it lands in S3.

Amazon EMR (A) is not the best fit here because it is a batch-oriented big data processing platform (Hadoop/Spark) rather than a streaming ingestion service, and AWS Glue (D) is primarily a serverless ETL and data catalog service for batch and some streaming jobs, not a dedicated real-time ingestion pipeline component for this scenario.

Exam trap

DEA-C01 often tests the distinction between real-time streaming ingestion (Kinesis Data Streams) and near-real-time delivery to S3 (Kinesis Data Firehose), causing candidates to overlook Lambda as the processing layer or incorrectly select batch-oriented services like EMR or Glue.

63
MCQeasy

A data engineer needs to schedule a recurring AWS Glue ETL job that must run every night at 02:00 UTC and must not start a new run while a previous run is still executing. The engineer wants the simplest managed scheduling option that integrates natively with Glue job run state. Which approach should the engineer use?

A.Create an AWS Step Functions state machine with a Wait state that loops every 24 hours and starts the Glue job.
B.Use an AWS Lambda function invoked by a CloudWatch Events rule to call StartJobRun on a fixed schedule.
C.Configure a Glue trigger of type SCHEDULED with a cron expression and set the job's maximum concurrency to 1.
D.Create an Amazon EventBridge scheduled rule with a cron expression that invokes the Glue job via a target.
AnswerC

A SCHEDULED Glue trigger natively starts the job on a cron schedule, and setting the job's maximum concurrency to 1 prevents a new run from starting while a previous run is still active. Both settings live within Glue, so no external orchestration service is required. This is the simplest managed approach that meets the time-based schedule and the no-overlap requirement.

Why this answer

Glue scheduled triggers use cron expressions to start jobs at fixed times, and the job's maximum concurrency setting controls whether overlapping runs are allowed. Setting maximum concurrency to 1 ensures that if a run is still active when the next trigger fires, the new run does not start. This keeps scheduling and concurrency control entirely within Glue, which is the simplest managed solution.

Exam trap

The trap here is reaching for an external scheduler like EventBridge or Lambda when Glue already provides native cron triggers and concurrency limits.

64
MCQmedium

A data engineer uses Amazon EMR to run a Spark job that reads from S3 and writes to HDFS on the cluster. The job fails with an 'OutOfMemoryError: Java heap space' error in the executors. Which parameter adjustment should be made to resolve this?

A.Increase spark.default.parallelism
B.Increase spark.sql.shuffle.partitions
C.Increase spark.executor.memory
D.Increase spark.driver.memory
AnswerC

Raising spark.executor.memory enlarges each executor's JVM heap, directly addressing the 'OutOfMemoryError: Java heap space' thrown during Spark execution. Since the failure occurs in executors rather than the driver, this parameter targets the constrained component, giving shuffle and aggregation buffers sufficient headroom to complete the S3-to-HDFS job.

Why this answer

The OutOfMemoryError: Java heap space in Spark executors indicates that the executor JVM heap is insufficient to hold the data being processed. Increasing spark.executor.memory allocates more heap space to each executor, allowing it to handle larger partitions or aggregations without running out of memory. This directly addresses the root cause of the error.

Exam trap

DEA-C01 often tests the confusion between driver and executor memory, and the misconception that increasing parallelism or shuffle partitions directly solves heap memory errors.

How to eliminate wrong answers

Option A is wrong because increasing spark.default.parallelism changes the number of partitions for RDD operations, which can increase parallelism but does not increase the memory available to each executor; it may even create more tasks that compete for the same heap. Option B is wrong because increasing spark.sql.shuffle.partitions controls the number of partitions after a shuffle, which can reduce the size of each partition but does not increase executor heap; it might help if partitions are too large, but the error is specifically about heap space, so increasing memory is more direct. Option D is wrong because increasing spark.driver.memory affects the driver JVM, not the executors; the error is in the executors, so this would not resolve it.

65
MCQhard

A company uses AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration completes successfully, but data validation shows some tables have missing rows. The task is configured for ongoing replication using change data capture (CDC). What is the MOST likely cause of the missing rows?

A.Source database archive log retention period too short
B.Large objects (LOBs) not supported by the target
C.Source tables missing primary keys
D.Insufficient storage on the DMS replication instance
AnswerC

DMS applies CDC changes using the source table's primary key to identify and update the correct target rows. Tables lacking a primary key cannot be matched reliably, so updates and deletes are dropped or misapplied, producing missing rows in the PostgreSQL target after replication.

Why this answer

AWS DMS requires a primary key or unique index on source tables to reliably identify and apply row changes during CDC. Without a primary key, DMS cannot uniquely match rows for updates and deletes, and it may also fail to capture all inserts during the initial load plus CDC handoff, resulting in missing rows.

Exam trap

DEA-C01 often tests the misconception that DMS CDC works on any table, when missing primary keys silently break change capture and cause missing rows.

How to eliminate wrong answers

Option A is wrong because insufficient archive log retention would cause the CDC task to fail with a specific error about missing log files, not silently drop rows; DMS would stop rather than skip. Option B is wrong because unsupported LOBs cause LOB truncation or task failure with explicit errors, not missing rows in non-LOB tables. Option D is wrong because insufficient replication instance storage causes the task to fail with a storage-full error, not selective row loss.

66
MCQmedium

A company uses Amazon EMR to run Spark jobs on data stored in S3. After upgrading the EMR cluster to a new release, one of the Spark jobs fails with 'OutOfMemoryError' in the executor. Which configuration change is MOST likely to resolve this issue?

A.Increase the number of core nodes in the EMR cluster.
B.Decrease spark.sql.shuffle.partitions to reduce overhead.
C.Increase spark.driver.memory in the Spark configuration.
D.Increase spark.executor.memory to allocate more memory per executor.
AnswerD

Spark executor memory is set by spark.executor.memory; the OutOfMemoryError arises because the new EMR release changed default executor sizing, leaving too little heap. Raising this value directly expands executor heap, satisfying the job's memory constraint without altering cluster hardware.

Why this answer

Increasing spark.executor.memory directly allocates more memory per executor, addressing the OutOfMemoryError in the executor. Option A (increasing the number of core nodes) adds more cluster capacity but does not increase the memory available to individual executors, so it may not resolve the OOM if the existing executors are already memory-constrained. Option B (decreasing spark.sql.shuffle.partitions) reduces the number of shuffle partitions, which can increase the size of each partition and potentially cause more memory pressure, not less.

Option C (increasing spark.driver.memory) only helps the driver process, not the executor, so it does not fix executor OOM errors.

67
MCQmedium

A data engineer runs an AWS Glue ETL job that reads CSV files from an Amazon S3 bucket, applies transformations, and writes Parquet output to another S3 bucket. The job fails with the error 'AnalysisException: Unable to infer schema for CSV. It must be specified manually.' The CSV files are stored with a header row, and the job's script uses the default Glue DynamicFrame reader without specifying format options. What is the MOST likely cause of the failure?

A.The CSV files are compressed with gzip, and Glue cannot infer schema from compressed CSV files.
B.The CSV files have inconsistent or malformed data that prevents Glue from inferring a schema, or the header option is not set correctly.
C.The S3 bucket containing the CSV files does not have the correct bucket policy allowing AWS Glue to read objects.
D.The Glue job's IAM role lacks permissions to read the CSV files from the S3 bucket.
AnswerB

Glue's schema inference can fail if CSV files have inconsistent columns, missing headers, or if the 'withHeader' option is not set to true. In this scenario, the header row exists but the default reader may not treat it as a header, leading to type conflicts. Setting 'withHeader' to true and ensuring consistent data resolves the error.

Why this answer

The error 'Unable to infer schema for CSV' occurs when AWS Glue cannot determine column names and types from the source data. This often happens when the header option is not enabled or when the CSV data has inconsistencies such as varying column counts or mixed data types. Ensuring the reader is configured with 'withHeader' set to true and that data is well-formed allows Glue to infer the schema correctly.

Exam trap

The trap here is assuming that a schema inference error is caused by permissions or compression, when it actually stems from data formatting or missing header configuration.

68
MCQeasy

A data engineer has set up an AWS Lambda function that processes files uploaded to an S3 bucket. The function is triggered by S3 event notifications. However, the function is not being invoked when a file is uploaded. The engineer checks the Lambda function's CloudWatch Logs and finds no execution logs. What should the engineer check FIRST?

A.Check the Lambda function's code for errors.
B.Verify that the Lambda function's IAM role has permissions to read from S3.
C.Verify that the S3 bucket has an event notification configured for the Lambda function.
D.Check if the Lambda function is attached to a VPC.
AnswerC

Without an S3 event notification targeting the Lambda function, uploads never generate an invocation, so no CloudWatch execution logs appear. Confirming the notification configuration first isolates whether the trigger exists before investigating permissions or code.

Why this answer

The symptom is that the Lambda function is not invoked at all and there are no CloudWatch execution logs — meaning the function never ran. The first thing to verify is whether the S3 bucket actually has an event notification configured to trigger the Lambda function, since without that configuration no invocation occurs. This is the most direct cause of 'no invocation, no logs.'

Exam trap

DEA-C01 often tests the misconception that a Lambda failure is always a code or IAM issue — candidates must remember that 'no invocation, no logs' points to the trigger configuration (S3 event notification) being missing or misconfigured.

How to eliminate wrong answers

Option A is wrong because code errors would only manifest after the function is invoked — and the absence of execution logs proves the function never ran. Option B is wrong because IAM permissions on the Lambda role matter only after invocation; they cannot prevent the S3 event from triggering the function. Option D is wrong because VPC attachment affects network access from the function, not whether the function is invoked — and again, no logs means no invocation.

69
MCQeasy

A data engineer notices that a nightly AWS Glue ETL job has been failing for the past three days with the error 'Unable to locate credentials'. The job uses an IAM role for execution. What is the most likely cause of this error?

A.The IAM role does not have an access key attached.
B.The S3 bucket name in the job parameters is misspelled.
C.The IAM role's trust policy does not include glue.amazonaws.com as a trusted entity.
D.The JDBC connection string contains an incorrect password.
AnswerC

Glue assumes its execution role via AWS Security Token Service, which validates the role's trust policy. If glue.amazonaws.com is absent as a trusted principal, AssumeRole is denied, producing the 'Unable to locate credentials' error rather than a permissions failure.

Why this answer

The error 'Unable to locate credentials' indicates that the AWS Glue job cannot obtain AWS credentials to authenticate API calls. Since the job uses an IAM role for execution, the most likely cause is that the trust policy of that IAM role does not include 'glue.amazonaws.com' as a trusted entity. Without this trust relationship, AWS Glue cannot assume the role and thus has no credentials to sign requests.

Exam trap

AWS often tests the distinction between IAM role trust policies (who can assume the role) and IAM role permission policies (what actions the role can perform), and candidates mistakenly focus on permission policies when the error is about credential acquisition.

How to eliminate wrong answers

Option A is wrong because IAM roles do not use access keys; they use temporary security credentials obtained via the AWS Security Token Service (STS). Option B is wrong because a misspelled S3 bucket name would cause a 'NoSuchBucket' or 'Access Denied' error, not a credentials-related error. Option D is wrong because an incorrect JDBC password would result in a connection failure or authentication error from the database, not an 'Unable to locate credentials' error from AWS.

70
Multi-Selecthard

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query hours. The engineer notices that some queries are waiting in the queue for a long time, and the WLM (Workload Management) configuration is set to auto. The engineer wants to implement manual WLM to improve query throughput and ensure that short-running queries are not blocked by long-running ones. Which TWO actions should the engineer take to achieve this? (Choose two.)

Select 2 answers
A.Set the WLM timeout for the long-running query queue to automatically terminate queries that exceed a specified time.
B.Create separate WLM queues for short-running and long-running queries, and assign appropriate memory percentages to each queue.
C.Use query monitoring rules (QMR) to log queries that exceed resource thresholds and send alerts.
D.Increase the number of nodes in the Redshift cluster to provide more memory and CPU resources.
E.Enable concurrency scaling to automatically add cluster capacity for bursts of queries.
AnswersA, B

Setting a WLM timeout on the long-running queue prevents those queries from monopolizing resources indefinitely. When a query exceeds the timeout, it is terminated, freeing up resources for other queries. This helps maintain throughput and ensures that short queries are not starved. It is a key configuration in manual WLM to manage runaway queries.

Why this answer

Manual WLM with separate queues for short and long queries, along with WLM timeouts for long queries, directly addresses the issue of short queries being blocked. Separate queues provide resource isolation, while timeouts prevent long queries from consuming resources indefinitely. These two actions together improve throughput and ensure responsiveness during peak hours.

Exam trap

The trap here is confusing concurrency scaling or QMR as solutions for queue wait times, when they do not provide the necessary workload isolation and resource management.

71
MCQeasy

A data engineer needs to schedule a daily AWS Glue job that extracts data from Amazon S3 and loads it into Amazon Redshift. The engineer wants to ensure the job runs at 2:00 AM UTC every day and can be monitored for failures. What is the simplest way to achieve this?

A.Use AWS Glue triggers to schedule the job with a cron expression.
B.Create an AWS Lambda function that triggers the Glue job using a cron expression in Amazon EventBridge.
C.Use Amazon EventBridge Scheduler to invoke the Glue job directly.
D.Create an AWS Step Functions state machine with a Wait state and a Glue job task.
AnswerA

Correct. AWS Glue triggers allow you to schedule jobs using cron expressions. You can create a trigger of type SCHEDULED, set the cron expression to '0 2 * * ? *' for 2:00 AM UTC daily, and attach it to the Glue job. This is the simplest and most native way to schedule Glue jobs, with built-in monitoring through Glue job run metrics and CloudWatch.

Why this answer

AWS Glue triggers are the native scheduling mechanism for Glue jobs. Creating a scheduled trigger with a cron expression allows the job to run automatically at specified times. This approach is simple, requires no additional services, and provides monitoring through Glue job run statuses and CloudWatch metrics.

The cron expression '0 2 * * ? *' schedules the job for 2:00 AM UTC daily.

Exam trap

The trap here is overcomplicating the solution by involving additional services like Lambda or Step Functions when AWS Glue provides a built-in scheduling feature.

72
MCQmedium

A data engineer is using AWS Step Functions to orchestrate an ETL workflow that includes an AWS Glue job. The Glue job occasionally fails due to transient issues, such as network timeouts. The engineer wants the Step Functions state machine to automatically retry the Glue job up to three times with exponential backoff before failing the workflow. Which Step Functions state configuration should the engineer use?

A.Configure the Glue job to have a retry policy in its job definition, such as MaxRetries: 3.
B.Use a Retry field with MaxAttempts: 3 and BackoffRate: 2.0 in the task state that invokes the Glue job.
C.Use a Catch field with ErrorEquals: ["States.ALL"] and Next: "RetryState" to redirect to a retry state.
D.Set the TimeoutSeconds and HeartbeatSeconds fields in the task state to trigger a retry.
AnswerB

The Retry field in a Step Functions task state allows you to specify retry behavior for failed tasks. Setting MaxAttempts to 3 and BackoffRate to 2.0 will retry the Glue job up to three times with exponential backoff (doubling the wait time between retries). This matches the requirement.

Why this answer

The Retry field in AWS Step Functions is specifically designed to handle retries on task failures, including automatic backoff. By specifying MaxAttempts and BackoffRate, the state machine will retry the Glue job invocation with exponential backoff, meeting the requirement without custom logic.

Exam trap

The trap here is confusing error handling with retry logic; Catch is for fallback, not for automatic retries with backoff.

73
MCQeasy

A data engineer is running an Amazon EMR cluster with Spark to process log files. The cluster uses instance fleets with m5.xlarge core nodes. The engineer observes that the Spark job is running slower than expected. CloudWatch metrics show that the cluster's CPU utilization is below 20% but memory utilization is near 90%. Which configuration change would most likely improve performance?

A.Use memory-optimized instances (r5.xlarge) for core nodes.
B.Increase the number of core nodes from 5 to 10.
C.Increase the number of Spark shuffle partitions.
D.Decrease the number of core nodes to reduce overhead.
AnswerA

Memory-optimised r5.xlarge instances provide a higher memory-to-vCPU ratio than m5.xlarge, directly relieving the near-90% memory utilisation that is throttling Spark executors. Since CPU sits below 20%, the bottleneck is memory capacity, not compute, so swapping core nodes to r5.xlarge lets executors hold larger partitions without spilling to disk.

Why this answer

The CloudWatch metrics show CPU below 20% while memory is near 90%, which is the classic signature of a memory-bound Spark workload. Spark executors on m5.xlarge (16 GiB RAM) are spilling to disk or GC-thrashing because the working set exceeds available heap. Switching core nodes to r5.xlarge (memory-optimized, 32 GiB RAM) doubles the memory per node, reducing spills and GC pressure, which directly addresses the bottleneck.

Exam trap

The trap here is assuming that 'slower than expected' always means insufficient compute, so candidates add nodes or partitions instead of reading the CloudWatch signal that memory — not CPU — is the saturated resource.

How to eliminate wrong answers

Option B is wrong because adding more core nodes increases aggregate memory but does not change the per-executor memory-to-core ratio; if each executor is already memory-starved, more nodes with the same instance type will still spill and the shuffle/IO overhead may even grow. Option C is wrong because increasing shuffle partitions addresses data skew or large partition sizes, not a memory bottleneck — it can actually increase overhead and small-file pressure. Option D is wrong because reducing core nodes lowers total cluster memory and parallelism, worsening the memory pressure and slowing the job further.

74
Multi-Selecthard

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query hours. The engineer needs to identify and resolve issues related to workload management (WLM). Which TWO actions should the engineer take to improve query performance? (Choose two.)

Select 2 answers
A.Increase the number of nodes in the cluster to add more compute resources.
B.Set the WLM query slots to the maximum value for all queues to increase concurrency.
C.Enable short query acceleration (SQA) to prioritize short-running queries.
D.Disable concurrency scaling to prevent additional clusters from being added.
E.Configure a manual WLM queue with a higher memory allocation for the queue handling complex queries.
AnswersC, E

Short query acceleration (SQA) uses machine learning to predict query execution time and runs short queries in a dedicated space, preventing them from waiting behind long-running queries. Enabling SQA can improve overall throughput and reduce latency for short queries during peak periods, making it an effective action to resolve WLM-related performance issues.

Why this answer

To improve query performance in Amazon Redshift during peak hours from a workload management perspective, the engineer should configure manual WLM queues with appropriate memory allocation and enable short query acceleration. Manual WLM allows fine-grained control over memory and concurrency, ensuring complex queries get necessary resources. SQA prioritizes short queries, reducing wait times.

These actions directly address WLM-related bottlenecks.

Exam trap

The trap here is confusing general scaling actions with WLM-specific tuning, such as adding nodes or maximizing concurrency, which can be counterproductive.

75
Multi-Selecthard

A company is migrating its on-premises data warehouse to Amazon Redshift. The data includes tables with up to 100 columns and 500 million rows. The migration involves a full load followed by incremental updates. The company needs to minimize downtime during the final cutover. Which THREE strategies should the data engineer use to facilitate the migration? (Choose THREE.)

Select 3 answers
A.Increase the number of WLM queues to allow more concurrent loads.
B.Use the COPY command to load data from Amazon S3.
C.Use columnar format (e.g., Parquet) for the data files in S3.
D.Run VACUUM and ANALYZE commands after loading the data.
E.Disable distribution keys on the target tables to simplify loading.
AnswersB, C, D

The COPY command loads data in parallel from Amazon S3 into Redshift, directly satisfying the requirement to minimise cutover downtime during the full load. Its massively parallel architecture ingests the 500-million-row tables far faster than row-by-row INSERT statements, compressing the migration window before incremental updates begin.

Why this answer

Option B is correct because the COPY command is the most efficient, parallelized way to bulk-load large datasets from Amazon S3 into Redshift, and it supports loading from multiple files in parallel to speed up the full load and minimize cutover downtime. Option C is correct because storing the source files in a columnar format such as Parquet lets COPY read only needed columns, compresses data heavily, and reduces I/O and load time compared with row-based CSV or JSON. Option D is correct because after a large load, running VACUUM re-sorts rows and reclaims space (restoring sort-key performance) and ANALYZE refreshes table statistics so the query planner produces efficient plans for the subsequent incremental workload.

Option A is not appropriate because adding WLM queues does not by itself accelerate a single large load and can even reduce per-query resources; queue design is about workload isolation, not migration throughput. Option E is not appropriate because disabling distribution keys removes the ability to co-locate joins and causes data redistribution at query time, hurting performance; distribution keys should be chosen deliberately, not disabled.

Exam trap

DEA-C01 often tests whether candidates know that WLM queues and distribution keys are performance/concurrency constructs, not migration accelerators — the trap is picking 'more queues' or 'disable distribution keys' as shortcuts, when the real levers are COPY from S3, columnar formats, and post-load VACUUM/ANALYZE.

Page 1 of 4 · 270 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Operations and Support questions.