Courseiva

CCNA Data Ingestion Transformation Questions

75 of 447 questions · Page 2/6 · Data Ingestion Transformation topic · Answers revealed

76
MCQmedium

An e-commerce company ingests clickstream data from their website into Amazon S3. The data is in JSON format, and each file is about 10 MB. They need to transform the data into a columnar format for analytics and load it into Amazon Redshift nightly. The transformation should be cost-effective and require minimal operational overhead. Which approach meets these requirements?

A.Use AWS Glue ETL job to convert to Parquet and load into Redshift.
B.Use Amazon Redshift COPY command to load JSON directly.
C.Use Amazon EMR with Spark to transform and load data.
D.Use AWS Lambda to transform each file and write to Redshift.
AnswerA

AWS Glue ETL jobs convert JSON to Parquet on serverless Spark infrastructure, eliminating cluster provisioning and satisfying the minimal-operational-overhead constraint. Parquet's columnar layout and compression reduce Redshift storage and scan costs, meeting the cost-effectiveness requirement for the nightly 10 MB file transformations.

Why this answer

AWS Glue ETL is the correct choice because it is a serverless, managed service that can efficiently convert JSON to Parquet (a columnar format optimized for Redshift) and load the data into Redshift with minimal operational overhead. The nightly batch processing of 10 MB files is well-suited for Glue's pay-per-use pricing, making it cost-effective without requiring infrastructure management.

Exam trap

The trap here is that candidates may choose Amazon EMR or Lambda because they are familiar with Spark or serverless functions, but they overlook the operational overhead of EMR and the execution limits of Lambda for batch workloads, while Glue provides a balanced, managed solution for this specific use case.

How to eliminate wrong answers

Option B is wrong because the Redshift COPY command can load JSON directly, but it does not transform the data into a columnar format like Parquet; it loads JSON as-is, which is less efficient for analytics and may require additional schema handling. Option C is wrong because Amazon EMR with Spark introduces significant operational overhead for managing clusters, tuning, and monitoring, which is unnecessary for a simple nightly transformation of small 10 MB files. Option D is wrong because AWS Lambda has a maximum execution timeout of 15 minutes and limited memory (up to 10 GB), making it unsuitable for batch processing multiple files or handling large datasets; it is designed for event-driven, short-lived tasks, not nightly ETL workloads.

77
MCQeasy

A data engineer is tasked with transforming JSON data from an S3 bucket into Parquet format for efficient querying. The transformation should run on a schedule every hour. Which AWS service is best suited for this task?

A.AWS Lambda
B.Amazon Athena
C.AWS Glue
D.Amazon EMR
AnswerC

AWS Glue satisfies the hourly schedule constraint through time-based triggers, and its ETL engine converts JSON to Parquet using built-in transforms that infer schemas automatically. Crawlers catalogue the S3 source, while the Parquet output is written directly to S3 for efficient downstream querying by Athena or Redshift Spectrum.

Why this answer

AWS Glue is the best choice because it is a fully managed ETL service designed specifically for transforming and cataloging data at scale. It can natively read JSON from S3, convert it to Parquet, and run on a scheduled hourly basis using a Glue job with a trigger, without requiring server management or custom infrastructure.

Exam trap

The trap here is that candidates often confuse Athena's ability to query Parquet with the ability to transform data into Parquet, but Athena is a query engine, not an ETL service, and cannot perform scheduled data format conversions.

How to eliminate wrong answers

Option A is wrong because AWS Lambda has a maximum execution timeout of 15 minutes and a 10 GB memory limit, making it unsuitable for processing large JSON datasets or running long-running hourly transformations. Option B is wrong because Amazon Athena is an interactive query service for analyzing data directly in S3, not a transformation engine; it cannot convert JSON to Parquet and write the output back to S3 in a scheduled, automated manner. Option D is wrong because Amazon EMR requires provisioning and managing a cluster of EC2 instances, which adds operational overhead and cost, whereas the task calls for a serverless, scheduled transformation with minimal management.

78
MCQmedium

A company uses AWS Glue ETL jobs to process data from an S3 data lake. The job reads data in CSV format, transforms it, and writes to Parquet. The job runs daily and takes 2 hours to complete. The data volume is increasing by 20% each month. The engineer wants to reduce the job runtime. Which action is most effective?

A.Increase the number of DPUs for the Glue job
B.Enable compression on the input CSV files
C.Switch from Python Shell to Spark ETL
D.Partition the input data in S3 by date and use partition pruning in the job
AnswerD

Partitioning S3 input by date lets Glue read only the relevant partitions, cutting the bytes scanned each run. This addresses the growing data volume constraint far more effectively than scaling workers or changing formats alone.

Why this answer

Most effective because partitioning the input data by date and using partition pruning allows the Glue ETL job to read only the relevant partitions instead of scanning the entire S3 data lake. This drastically reduces the amount of data processed, which directly addresses the growing data volume and shortens job runtime. Partition pruning is a core optimization for Spark-based Glue jobs, as it leverages Hive-style partitioning to skip unnecessary files.

Exam trap

The trap here is that candidates often assume increasing DPUs or enabling compression is the universal fix, but they fail to recognize that reducing the data scanned via partition pruning is the most impactful optimization for growing datasets in S3-based Glue jobs.

How to eliminate wrong answers

Option A is wrong because increasing DPUs (Data Processing Units) adds more parallelism but does not reduce the volume of data read; it may help only if the job is CPU-bound, but the primary bottleneck here is the increasing data volume, not compute capacity. Option B is wrong because enabling compression on input CSV files reduces storage size and I/O overhead, but CSV is not splittable when compressed (e.g., Gzip), which can actually harm parallelism and increase runtime; moreover, the job still reads all data. Option C is wrong because the question states the job already uses AWS Glue ETL, which is Spark-based by default; switching from Python Shell to Spark ETL would be a regression, as Python Shell is single-node and slower for large datasets, but the current job is already using Spark (implied by Glue ETL), so this change is irrelevant or counterproductive.

79
MCQmedium

A streaming application sends data to Amazon Kinesis Data Streams. The data must be enriched with reference data from an Amazon DynamoDB table in real-time. Which AWS service can be used to perform this enrichment with minimal latency?

A.Amazon Kinesis Data Analytics for Apache Flink
B.Amazon Kinesis Data Firehose with Lambda transformation
C.AWS Lambda function triggered by Kinesis Data Streams
D.AWS Glue streaming ETL
AnswerA

Amazon Kinesis Data Analytics for Apache Flink runs continuous SQL or Flink applications directly over the stream, joining each record against the DynamoDB reference table in-flight. This satisfies the real-time enrichment requirement with minimal latency, avoiding the batching delay of Lambda-based or Firehose-based approaches.

Why this answer

Amazon Kinesis Data Analytics for Apache Flink is correct because it allows you to run Apache Flink applications that can read from a Kinesis data stream, perform stateful stream processing, and enrich records in real-time by joining with reference data stored in DynamoDB. Flink's asynchronous I/O and managed state enable sub-second enrichment latency without the cold-start delays or concurrency limits of Lambda-based approaches.

Exam trap

The DEA-C01 exam often tests the distinction between real-time stream processing (Kinesis Data Analytics for Flink) and near-real-time or batch-oriented services (Firehose, Glue ETL), leading candidates to choose Lambda because they assume serverless functions are always the lowest-latency option, ignoring concurrency and cold-start limitations in streaming contexts.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Firehose is a near-real-time delivery service with a minimum buffer interval of 60 seconds, making it unsuitable for real-time enrichment with minimal latency. Option C is wrong because AWS Lambda triggered by Kinesis Data Streams has a maximum concurrency limit per shard (e.g., 10 concurrent invocations per shard) and incurs cold-start latency, which can cause backpressure and increased processing delays for high-throughput streaming workloads. Option D is wrong because AWS Glue streaming ETL is based on Apache Spark Structured Streaming, which introduces higher startup overhead and micro-batch latency (typically seconds), making it less optimal for sub-second real-time enrichment compared to Flink's event-at-a-time processing.

80
MCQhard

A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer notices that the job fails with an error indicating that the Redshift table does not exist. The engineer has already created the Redshift cluster and database. What should the engineer do to resolve the error?

A.Create the target table in Amazon Redshift with the appropriate schema before running the job.
B.Add a 'ResolveChoice' transform before the Redshift writer to handle nested types.
C.Use the 'ApplyMapping' transform to map columns to Redshift data types.
D.Configure the Redshift writer to use the 'CREATE TABLE' option and specify the table name.
AnswerA

AWS Glue jobs writing to Amazon Redshift require the target table to exist beforehand. The error indicates the table is missing. The engineer must create the table in Redshift with a schema that matches the output of the Glue job, including column names and data types. This is a common prerequisite when using the Redshift connector in Glue Studio.

Why this answer

When writing to Amazon Redshift from AWS Glue, the target table must already exist. Glue does not automatically create Redshift tables. The engineer should create the table with the correct schema in Redshift, then rerun the job.

The other transforms mentioned are for data manipulation within the job and do not address the missing table.

Exam trap

The trap here is thinking that AWS Glue can create the Redshift table automatically, similar to how it might create a Data Catalog table, when in fact Redshift tables must be pre-created.

81
MCQeasy

A data engineer needs to ingest data from an external partner's FTP server to Amazon S3. The data arrives once daily as a CSV file. Which AWS service should be used for this ingestion?

A.AWS DataSync
B.Amazon Kinesis Data Firehose
C.Amazon AppFlow
D.AWS Transfer Family
AnswerD

AWS Transfer Family provides a fully managed FTP endpoint that writes directly into Amazon S3, satisfying the daily CSV ingestion requirement without custom code or servers. Unlike AWS DataSync, which synchronises existing data stores, Transfer Family natively speaks FTP, letting the external partner push files straight to S3.

Why this answer

AWS Transfer Family provides fully managed support for file transfers over SFTP, FTPS, and FTP protocols, making it the correct choice for ingesting CSV files from an external partner's FTP server. It integrates directly with Amazon S3 as a destination, enabling automated, secure, and scheduled transfers without custom infrastructure.

Exam trap

The trap here is that candidates often confuse AWS DataSync (which is for NFS/SMB, not FTP) with a general-purpose file transfer service, or they incorrectly assume Kinesis Data Firehose can handle batch file ingestion from external sources.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is designed for high-speed, large-scale data transfers between on-premises storage and AWS, but it does not support the FTP protocol; it uses its own agent-based architecture over NFS/SMB. Option B is wrong because Amazon Kinesis Data Firehose is a streaming ingestion service for real-time data (e.g., logs, events) and cannot connect to an FTP server or handle scheduled batch file transfers. Option C is wrong because Amazon AppFlow supports SaaS application integrations (e.g., Salesforce, Slack) and does not support FTP as a source or destination.

82
MCQeasy

A company is using Amazon Kinesis Data Firehose to ingest data into Amazon S3. The data must be transformed from JSON to Parquet format before delivery. Which feature should be enabled on the Firehose delivery stream?

A.Amazon Kinesis Data Analytics
B.Amazon S3 event notifications
C.Format conversion (Parquet/ORC)
D.AWS Lambda transformation
AnswerC

Format conversion performs schema-aware JSON-to-Parquet transcoding inside the Firehose delivery stream, using a referenced AWS Glue Data Catalog table to map source fields to Parquet columns. This satisfies the stem's requirement that records be transformed before delivery, removing the need for downstream ETL or Lambda-based conversion.

Why this answer

Amazon Kinesis Data Firehose has a built-in format conversion feature that can automatically convert input data from JSON to Parquet or ORC format before delivery to Amazon S3. Option A (Amazon Kinesis Data Analytics) is for real-time stream processing, not format conversion within Firehose. Option B (Amazon S3 event notifications) triggers notifications on S3 events, not data transformation.

Option D (AWS Lambda transformation) allows custom code for data transformation but is not specifically for converting JSON to Parquet; the built-in format conversion is the appropriate feature for this task.

83
Multi-Selectmedium

A data engineer needs to design a data ingestion pipeline that captures streaming data from mobile app events into Amazon S3 for analytics. The pipeline must support real-time processing of events and allow for schema evolution over time. Which AWS services should the engineer use? (Choose THREE.)

Select 3 answers
A.Amazon Kinesis Data Analytics
B.Amazon Kinesis Data Firehose
C.AWS Glue ETL jobs
D.Amazon Kinesis Data Streams
E.AWS AppFlow
AnswersA, B, D

Enables real-time processing and schema evolution.

Why this answer

Amazon Kinesis Data Analytics is correct because it enables real-time processing of streaming data using SQL or Apache Flink, allowing the engineer to analyze mobile app events as they arrive. This supports the requirement for real-time processing before the data is stored in Amazon S3 for analytics.

Exam trap

The trap here is that candidates often confuse AWS Glue ETL jobs as a streaming solution, but Glue is fundamentally batch-oriented and cannot meet real-time processing requirements, while AppFlow is mistakenly chosen for its integration capabilities despite lacking streaming ingestion support.

84
MCQhard

A data engineer uses AWS Glue DataBrew to clean a dataset stored in Amazon S3. The recipe includes a 'Filter' step that removes rows where the 'status' column equals 'INVALID'. After running the recipe, the engineer notices that the output still contains rows with status 'INVALID'. The recipe was published and the job ran successfully. What is the most likely cause?

A.The filter step was configured as a 'Remove' transformation instead of a 'Filter' transformation, so it only flagged rows without deleting them.
B.The recipe was applied to a sample of the data during job execution, so only a subset was filtered.
C.The DataBrew job was run in profile mode instead of recipe mode, so transformations were not applied.
D.The filter condition used a case-sensitive match and the actual values are 'invalid' in lowercase, so the condition did not match any rows.
AnswerD

DataBrew filter conditions are case-sensitive by default. If the source data contains 'invalid' rather than 'INVALID', a condition checking for 'INVALID' will not match, so those rows are retained. The engineer should either normalize case or adjust the condition. This is a common cause of filters appearing to have no effect when the job succeeds but output is unchanged.

Why this answer

DataBrew filter conditions are case-sensitive, so a condition matching 'INVALID' will not remove rows containing 'invalid'. The engineer should either standardize the case of the status column or adjust the filter condition to match the actual values. This explains why the job succeeded but the unwanted rows remained.

Exam trap

The trap here is assuming that DataBrew filter conditions are case-insensitive or that a successful job run guarantees the filter matched the intended rows.

85
Multi-Selecteasy

A data engineer is designing a serverless data ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data must be transformed using AWS Lambda before being written to S3. Which two steps are required to enable this transformation? (Select TWO.)

Select 2 answers
A.Set up an S3 event notification to trigger the Lambda function on object creation.
B.Configure a Lambda function as a data transformation source in the Firehose delivery stream.
C.Ensure the Lambda function returns the transformed data in the format required by Firehose.
D.Subscribe the Lambda function to the CloudWatch Logs log group for the Firehose stream.
E.Have the Lambda function write the transformed data directly to the S3 bucket.
AnswersB, C

Firehose invokes a Lambda function only when it is attached as the delivery stream's data transformation source, which enables the buffered records to be processed before delivery. Without this configuration, Firehose writes raw records straight to S3, so transformation never occurs.

Why this answer

Option B is correct because Firehose supports Lambda-based data transformation by letting you specify a Lambda function as the processor in the delivery stream's transform configuration (via the console or the ProcessingConfiguration/TransformParameters in the API), which Firehose then invokes synchronously for each buffered batch. Option C is correct because the Lambda function must return the transformed records in the exact structure Firehose expects — a JSON object containing records with recordId, result (Ok, Dropped, or ProcessingFailed), and base64-encoded data — otherwise Firehose cannot continue delivery. Option A is wrong because S3 event notifications trigger actions on object creation and play no role in Firehose's inline transformation; Firehose itself invokes the Lambda function.

Option D is wrong because subscribing Lambda to the Firehose CloudWatch Logs log group is not a configuration step for transformation and would not enable it. Option E is wrong because the Lambda function must return transformed data to Firehose, not write directly to S3; Firehose remains responsible for delivering the records to the destination bucket.

Exam trap

The trap here is that candidates often confuse post-delivery transformations (using S3 event notifications) with in-stream transformations (using Firehose's built-in Lambda integration), leading them to select Option A instead of the correct Firehose-specific configuration.

86
MCQmedium

A data engineer is using AWS Glue to run an ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job processes data in CSV format and the engineer wants to ensure the output is partitioned by year, month, and day based on a timestamp column in the data. The engineer needs to optimize the job for performance and cost. Which approach should the engineer take?

A.Write the output to a single CSV file and then use Amazon Athena to create a partitioned table.
B.Use the AWS Glue job's bookmark feature to automatically partition the output by timestamp.
C.Use the AWS Glue ResolveChoice transformation to split the timestamp column into separate year, month, and day columns, then write to S3 with partitioning.
D.Use the AWS Glue DynamicFrame partitionColumns parameter to specify year, month, and day as partition keys.
AnswerD

The partitionColumns parameter in AWS Glue's write_dynamic_frame method allows you to specify columns to partition the output by. When writing to S3, Glue will create a directory structure based on these columns, such as year=2023/month=01/day=01. This is the standard and efficient way to partition output data in Glue, improving query performance and reducing costs for downstream analytics.

Why this answer

The partitionColumns parameter in AWS Glue's write operation is the correct way to partition output data by specified columns. It creates a hierarchical directory structure in S3, enabling efficient querying with services like Athena and Redshift Spectrum. Other options either misuse transformations or misunderstand the purpose of bookmarks and single-file output.

Exam trap

The trap here is confusing job bookmarks with partitioning, or thinking that ResolveChoice can derive new columns, when actually partitionColumns is the direct method for output partitioning.

87
MCQmedium

A data engineer is using AWS Glue Studio to create a visual ETL job that reads from an Amazon S3 bucket containing JSON files, applies a filter transformation, and writes the output to Amazon Redshift. The job must run daily. The engineer notices that the job is taking a long time to complete and wants to improve performance. Which action should the engineer take to optimize the job?

A.Convert the source JSON files to Parquet format before running the ETL job.
B.Use a larger number of smaller files instead of fewer large files.
C.Increase the number of AWS Glue DPUs allocated to the job.
D.Enable job bookmarks to track processed files and avoid reprocessing.
AnswerA

Parquet is a columnar format that offers better compression and faster read performance compared to JSON. Converting the source data to Parquet reduces I/O and CPU overhead during the ETL job, significantly improving performance. This is a common best practice for AWS Glue jobs reading from S3, especially when the data is large. It directly addresses the slow read and transform steps.

Why this answer

Converting JSON to Parquet improves ETL performance because Parquet is columnar, compressed, and requires less I/O and CPU to parse. AWS Glue jobs benefit significantly from columnar formats when reading from S3. While increasing DPUs or enabling bookmarks can help in some cases, the most direct and effective optimization for slow JSON processing is to use a more efficient file format.

Exam trap

The trap here is assuming that adding more DPUs is always the first step to improve AWS Glue job performance, when data format optimization often yields greater benefits at lower cost.

88
Multi-Selecthard

A company uses Amazon Kinesis Data Firehose to deliver data to an S3 bucket. The data contains personally identifiable information (PII) that must be redacted before storage. Which TWO actions can achieve this requirement? (Choose TWO.)

Select 2 answers
A.Use AWS Glue ETL to read from the S3 bucket and write redacted data to another S3 bucket.
B.Use Amazon Athena to query the data and redact PII on the fly.
C.Use Amazon Macie to discover and automatically redact PII before storage.
D.Use AWS Database Migration Service (AWS DMS) to replicate data and apply transformations.
E.Use an AWS Lambda function as a transformation in the Firehose delivery stream.
AnswersA, E

Correct. AWS Glue ETL can perform schema-aware transformations, including redacting PII fields, and write to a target S3 bucket.

Why this answer

Option A is correct because AWS Glue ETL jobs can read the raw data from the S3 bucket, apply transformation logic (such as masking or removing PII fields), and write the redacted output to another S3 bucket, satisfying the requirement that PII be redacted before storage in the destination. Option E is correct because Kinesis Data Firehose supports Lambda-based record transformation, where a Lambda function is invoked on each incoming record to redact PII before Firehose delivers the data to S3, which is the most direct in-stream solution. Option B is not correct because Amazon Athena is a query service that reads data already stored in S3; it cannot redact PII before storage.

Option C is not correct because Amazon Macie discovers and classifies sensitive data but does not automatically redact or modify it. Option D is not correct because AWS DMS is designed for database migration and replication, not for transforming and redacting PII in a Firehose-to-S3 data pipeline.

Exam trap

The trap is that candidates may assume Macie can automatically redact PII, but it only discovers and classifies; it cannot modify data without additional services like Lambda. Thus, Option C is not a correct answer.

89
MCQeasy

A data engineer is using AWS Glue to run an ETL job that reads data from Amazon DynamoDB and writes to Amazon Redshift. The job fails with a 'ThroughputExceededException' error. What is the most likely cause?

A.The Glue job has a timeout setting that is too low
B.The Redshift cluster's concurrency scaling is insufficient
C.The DynamoDB table's read capacity is insufficient for the Glue job's read rate
D.The S3 bucket where Glue writes temporary data does not have proper permissions
AnswerC

DynamoDB returns ThroughputExceededException when read requests exceed the table's provisioned or on-demand read capacity. Glue's parallel executors consume read capacity units rapidly, exhausting the table's limits. This directly satisfies the stem's constraint: the read rate from DynamoDB surpasses available capacity, causing throttling rather than a Redshift or Glue configuration fault.

Why this answer

A ThroughputExceededException from DynamoDB indicates that the read or write request rate exceeded the provisioned throughput (RCUs/WCUs) or the burst capacity of the table or index. When AWS Glue reads from DynamoDB, it consumes read capacity units; if the table's read capacity is too low for the parallel read rate of the Glue job, DynamoDB throttles the requests and the job fails.

Exam trap

DEA-C01 often tests whether candidates correctly attribute throttling errors to the source database (DynamoDB) rather than the target (Redshift) or the Glue job configuration; the exception name itself points to DynamoDB throughput.

How to eliminate wrong answers

Option A is wrong because a Glue job timeout would produce a timeout error, not a ThroughputExceededException from DynamoDB. Option B is wrong because Redshift concurrency scaling affects query performance on the write side, not DynamoDB read throttling; the error originates from DynamoDB. Option D is wrong because S3 permissions issues would manifest as AccessDenied or similar errors when writing temporary data, not as a DynamoDB throughput exception.

90
MCQhard

A data engineer is troubleshooting a Kinesis Data Firehose delivery stream that is experiencing high error rates when writing to an S3 bucket. The error logs indicate 'AccessDenied' errors. The S3 bucket policy allows access from the Firehose service, but the errors persist. What is the most likely cause?

A.The S3 bucket has a lifecycle policy that is deleting objects too quickly
B.The IAM role assumed by Firehose does not have the s3:PutObject permission
C.The S3 bucket has default encryption enabled
D.The S3 bucket uses an AWS KMS key for encryption and Firehose does not have kms:Decrypt permission
AnswerB

Firehose requires both the bucket policy and its assumed IAM role to grant access. The bucket policy alone is insufficient; if the role lacks s3:PutObject, every delivery attempt returns AccessDenied, which matches the persistent errors despite the policy appearing correct.

Why this answer

The most likely cause is that the IAM role assumed by Kinesis Data Firehose lacks the `s3:PutObject` permission. Even if the S3 bucket policy allows access from the Firehose service, the IAM role must explicitly grant the necessary S3 write permissions for Firehose to deliver data. Without this permission, Firehose receives 'AccessDenied' errors when attempting to write objects to the bucket.

Exam trap

The trap here is that candidates assume a bucket policy allowing Firehose access is sufficient, but the IAM role assumed by Firehose must also explicitly grant the write permissions, as AWS evaluates both identity-based and resource-based policies.

How to eliminate wrong answers

Option A is wrong because a lifecycle policy that deletes objects too quickly would not cause 'AccessDenied' errors; it would cause data to be deleted after delivery, not prevent writes. Option C is wrong because default encryption on the S3 bucket does not block write access; Firehose can write encrypted objects as long as it has the necessary permissions. Option D is wrong because the error is 'AccessDenied', not a KMS-related error; if the issue were KMS permissions, the error would typically be 'KMS.AccessDeniedException' or similar, and Firehose would need `kms:GenerateDataKey` (not `kms:Decrypt`) to encrypt objects with SSE-KMS.

91
MCQeasy

A data engineer is using AWS Glue Studio to build a visual ETL job that reads JSON files from Amazon S3, applies a mapping transform, and writes Parquet to another S3 location. The engineer notices that the job is writing many small files, which hurts downstream query performance. Which action should the engineer take to reduce the number of output files without changing the source data?

A.Use the repartition or coalesce transform before the S3 target to reduce the number of output partitions.
B.Increase the number of DPUs for the Glue job so more workers write files in parallel.
C.Enable the Glue job bookmark so the job only processes new files on each run.
D.Change the output format from Parquet to CSV so the files are smaller and easier to manage.
AnswerA

Repartition or coalesce changes the number of partitions in the DataFrame, which directly controls how many files are written. Coalesce is especially useful when reducing partitions without a full shuffle. Applying it before the target write consolidates output into fewer, larger files and improves downstream query performance.

Why this answer

The number of files written by a Glue job corresponds to the number of partitions in the DataFrame at write time. Repartition or coalesce reduces that partition count, so the job writes fewer, larger files. This directly addresses the small-file issue without altering the source data or the output format.

Exam trap

The trap here is assuming that more compute resources will consolidate output, when adding workers usually increases the number of output files.

92
Multi-Selecteasy

Which TWO AWS services can be used to transform data in transit during ingestion? (Choose 2.)

Select 2 answers
A.Amazon S3 Transfer Acceleration
B.Amazon Kinesis Data Firehose with Lambda transformation
C.AWS Glue ETL
D.Amazon Athena
E.AWS Data Pipeline
AnswersB, C

Firehose supports invoking a Lambda function to transform each record in flight before delivery, satisfying the stem's in-transit transformation requirement. The transformation occurs during ingestion rather than after landing, so downstream consumers receive already-processed data.

Why this answer

Amazon Kinesis Data Firehose can invoke an AWS Lambda function to transform streaming data in real time before delivering it to a destination, making it suitable for transforming data in transit during ingestion. AWS Glue ETL can be used for both batch and streaming transformations; specifically, AWS Glue supports streaming ETL jobs that can transform data in transit as it is ingested from sources like Amazon MSK or Kinesis Data Streams. Therefore, both services can transform data during ingestion, albeit with different use cases and latency characteristics.

Exam trap

Candidates often mistakenly think that AWS Glue ETL is only for batch processing and cannot transform data in transit, but AWS Glue supports streaming ETL jobs that can transform data during ingestion, making it a valid choice for in-transit transformation.

93
Multi-Selecthard

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to improve the performance of the Glue job, which currently takes several hours to complete. The job reads large Parquet files, performs joins, and writes to Redshift. Which two actions should the engineer take to improve performance? (Choose two.)

Select 2 answers
A.Partition the source Parquet data in Amazon S3 by commonly filtered columns and use partition pruning in the Glue job.
B.Enable job bookmarks to avoid reprocessing previously processed data.
C.Increase the number of AWS Glue DPUs allocated to the job to add more Spark executors.
D.Convert the Parquet files to CSV to reduce storage size and speed up reading.
E.Use the 'Relationalize' transform to flatten the Parquet data before joining.
AnswersA, C

Partitioning the source data by columns used in filters allows Glue to read only relevant partitions, reducing I/O and the amount of data shuffled during joins. This can dramatically improve performance for large datasets. The engineer should ensure the Glue job uses predicate pushdown and that the partition columns are used in the query. This is a best practice for optimizing Glue ETL on S3.

Why this answer

Increasing DPUs adds compute resources to parallelize the job, and partitioning the source data enables partition pruning to reduce the amount of data read. Both actions directly address performance bottlenecks in a large Glue job. Other options either do not affect single-run performance or would degrade it.

Exam trap

The trap here is assuming that job bookmarks improve performance for a single large run, when they only help avoid reprocessing data across runs.

94
MCQmedium

A data engineer needs to transform CSV files arriving in an S3 bucket into Parquet format and store them in another S3 bucket. The transformation is simple and on-demand, triggered by data arrival. Which solution is the MOST cost-effective and requires the least operational overhead?

A.Use Amazon EMR with Spark streaming
B.Use Amazon Athena to create a new table with Parquet format
C.Use AWS Glue ETL jobs scheduled to run every hour
D.Use S3 Events to trigger an AWS Lambda function that transforms the data
AnswerD

S3 event notifications invoke Lambda directly on object arrival, so no cluster runs between executions and no polling infrastructure is needed. Lambda's per-invocation billing suits sporadic, on-demand CSV-to-Parquet conversion, and the service manages scaling and patching, satisfying the minimal operational overhead constraint.

Why this answer

Using S3 Events to trigger an AWS Lambda function is the most cost-effective and operationally lightweight solution for simple, on-demand CSV-to-Parquet transformations. Lambda scales automatically with each S3 PUT event, incurs no idle cost, and requires no cluster management, making it ideal for event-driven, low-volume transformations.

Exam trap

The DEA-C01 exam often tests the misconception that AWS Glue is always the best choice for ETL, but for simple, event-driven transformations with minimal overhead, Lambda is more cost-effective and operationally simpler than Glue's managed Spark environment.

How to eliminate wrong answers

Option A is wrong because Amazon EMR with Spark streaming introduces significant operational overhead (cluster provisioning, scaling, and management) and cost (even with auto-scaling, you pay for running instances) for a simple, on-demand transformation that does not require real-time streaming. Option B is wrong because Amazon Athena cannot transform data into Parquet format; it is a query engine that can read from and write to Parquet tables via CTAS statements, but it does not provide a direct, event-driven transformation trigger and incurs per-query costs that can be higher than Lambda for frequent small files. Option C is wrong because AWS Glue ETL jobs scheduled every hour introduce unnecessary latency (up to 1 hour delay) and cost (minimum billing per DPU hour) for an on-demand workload triggered by data arrival, and the scheduled polling approach is less efficient than event-driven invocation.

95
Drag & Dropmedium

Order the steps to troubleshoot a failed AWS Glue job that reads from JDBC and writes to S3.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Start with logs to identify errors, then check connectivity, IAM permissions, test connection, and review script.

96
Multi-Selectmedium

A data engineering team uses AWS Glue to extract, transform, and load (ETL) data from Amazon RDS for MySQL to Amazon S3. The job runs daily and processes incremental data. The team notices that the job is taking longer than expected. Which TWO actions can improve the job performance? (Choose two.)

Select 2 answers
A.Change the worker type to Standard (single node).
B.Use pushdown predicates to filter data at the source.
C.Add more transformations to the ETL script to clean data.
D.Increase the number of DPUs for the Glue job.
E.Disable compression on the output data to reduce CPU usage.
AnswersB, D

Pushdown predicates translate filter conditions into SQL WHERE clauses executed by RDS for MySQL, so only matching incremental rows are read over JDBC. Less data is transferred and processed in Glue, directly reducing the job's runtime.

Why this answer

Option B is correct because pushdown predicates let AWS Glue push filtering logic down to the source RDS for MySQL database, so only the required incremental rows are read over JDBC instead of the entire table, reducing I/O and shuffle work in the job. Option D is correct because increasing the number of DPUs adds more Apache Spark executors and parallel task slots, which improves throughput for a large, daily incremental ETL workload that is currently resource-bound. Option A is not appropriate because switching to a Standard single-node worker removes distributed processing and would generally slow the job rather than improve performance.

Option C is not appropriate because adding more transformations increases CPU and memory work in the ETL script, which would make the job slower, not faster. Option E is not appropriate because disabling output compression increases the volume of data written to Amazon S3 and read downstream, raising I/O and cost rather than improving job performance.

Exam trap

DEA-C01 often tests the misconception that more transformations or disabling compression improve ETL speed, when in fact pushdown predicates and additional DPUs are the canonical performance levers.

97
MCQmedium

A data engineer is designing a pipeline that ingests JSON logs from an application into Amazon S3. The logs contain a timestamp field. The pipeline must partition the data by date in S3 (e.g., year=2024/month=10/day=01). Which approach minimizes transformation effort?

A.Use Amazon Kinesis Data Firehose with dynamic partitioning
B.Use AWS Glue crawlers to infer schema and create partitions
C.Use AWS Lambda to process each object and copy to the appropriate prefix
D.Use Amazon Athena to create partitions on the existing data
AnswerA

Kinesis Data Firehose dynamic partitioning extracts the timestamp field via a jq expression and writes records into year=/month=/day=/ prefixes automatically, so no downstream ETL job is needed to reorganise objects. This directly satisfies the stem's requirement to minimise transformation effort while delivering date-partitioned JSON into Amazon S3.

Why this answer

Amazon Kinesis Data Firehose with dynamic partitioning can automatically partition incoming JSON data based on the timestamp field without requiring custom transformation code. It evaluates the timestamp using a JQ expression or inline parsing, then writes records directly to S3 prefixes like year=2024/month=10/day=01. This minimizes transformation effort because the partitioning logic is configured declaratively in the Firehose delivery stream, eliminating the need for Lambda functions or post-ingestion processing.

Exam trap

The trap here is that candidates confuse metadata partitioning (e.g., using Glue crawlers or Athena) with physical partitioning in S3, assuming that catalog operations alone reorganize the data, when in fact only ingestion-time partitioning (like Firehose dynamic partitioning) creates the folder structure without extra transformation effort.

How to eliminate wrong answers

Option B is wrong because AWS Glue crawlers infer schema and create partition metadata in the Glue Data Catalog, but they do not physically reorganize data into partitioned S3 prefixes; they only add partition keys to the catalog after data is already stored. Option C is wrong because using AWS Lambda to process each object and copy it to the appropriate prefix introduces significant transformation effort, including writing custom code for parsing, partitioning logic, and handling retries, which contradicts the goal of minimizing effort. Option D is wrong because Amazon Athena can create partitions on existing data using ALTER TABLE ADD PARTITION or MSCK REPAIR TABLE, but this only updates the catalog metadata and does not physically partition the data in S3; the data remains in a flat structure, and Athena queries still scan all files unless partitions are manually created.

98
MCQhard

A data pipeline ingests streaming data from Kinesis Data Streams into S3 via Kinesis Data Firehose. Occasionally, small files are written to S3, increasing downstream processing costs. What is the most efficient way to reduce the number of small files?

A.Use a Lambda function to aggregate records before sending to Firehose.
B.Use the Kinesis Client Library (KCL) to write larger batches to S3 directly.
C.Run a daily AWS Glue job to concatenate small files.
D.Increase the Firehose buffering interval to 300 seconds and buffering size to 64 MB.
AnswerD

Firehose buffers incoming records before delivering them, so raising the interval to 300 seconds and the size to 64 MB lets more records accumulate per delivery, producing fewer, larger S3 objects and cutting downstream processing overhead.

Why this answer

Kinesis Data Firehose allows you to configure buffering hints (size and interval) to control when data is delivered to S3. By increasing the buffering interval to 300 seconds and the buffering size to 64 MB, Firehose accumulates more records before writing, which reduces the number of small files. This is the most efficient approach as it requires no additional infrastructure or post-processing.

Exam trap

The trap here is that candidates may think a Lambda pre-aggregation (Option A) or a Glue job (Option C) is necessary, when in fact Firehose's built-in buffering configuration is the simplest and most cost-effective solution to control file sizes.

How to eliminate wrong answers

Option A is wrong because using a Lambda function to aggregate records before sending to Firehose adds latency and complexity, and Firehose already has built-in buffering capabilities that can be tuned without extra services. Option B is wrong because the Kinesis Client Library (KCL) is designed for consuming and processing records from a stream, not for writing directly to S3; it would require custom code to batch and write to S3, which is less efficient and not a managed solution. Option C is wrong because running a daily AWS Glue job to concatenate small files is a reactive, post-processing approach that does not prevent small files from being created in the first place, and it incurs additional compute costs and delays.

99
MCQeasy

A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is stored as CSV files and is updated daily. The engineer wants to load only new data each day without duplicating existing records. Which AWS service or feature should the engineer use to automate this process?

A.Amazon Redshift Spectrum to query S3 data directly and insert new records.
B.Amazon Kinesis Data Firehose to stream S3 data to Redshift.
C.AWS Glue with job bookmarks enabled to track processed files.
D.AWS Database Migration Service (DMS) with ongoing replication from S3 to Redshift.
AnswerC

AWS Glue job bookmarks track the state of data processed in previous runs, allowing the job to process only new or changed files in subsequent runs. This is ideal for daily incremental loads from S3 to Redshift, as it prevents reprocessing and duplication. The engineer can create a Glue ETL job that reads from S3, applies transformations, and writes to Redshift using the COPY command. This automates the incremental load process efficiently.

Why this answer

AWS Glue job bookmarks are designed to track processed data across job runs, enabling incremental processing. For daily loads from S3 to Redshift, a Glue ETL job with bookmarks ensures only new files are processed each day, preventing duplication. This automates the process and integrates with Redshift via the COPY command.

Other options do not provide automatic incremental load tracking for S3 files.

Exam trap

The trap here is assuming that Redshift Spectrum or DMS can automatically handle incremental loads from S3 without additional logic.

100
MCQmedium

A data engineer must ingest a 4 TB Oracle database into Amazon S3 nightly. The database is on-premises and the network link supports only 200 Mbps. The engineer wants to minimize the total transfer time and avoid impacting production. Which approach should the engineer use?

A.Use Amazon S3 Transfer Acceleration with multipart uploads from an on-premises script.
B.Use AWS DataSync with an on-premises agent to transfer the data over the internet.
C.Use AWS Snowball Edge devices to transfer the data offline and import into Amazon S3.
D.Use AWS Database Migration Service (AWS DMS) with a replication instance and an S3 target endpoint.
AnswerC

AWS Snowball Edge is designed for large-scale offline data transfer. With 4 TB and a 200 Mbps link, online transfer would take roughly 44 hours or more. Snowball Edge can copy the data locally and ship it, avoiding the network bottleneck entirely. This minimizes transfer time and avoids saturating the production link.

Why this answer

For multi-terabyte datasets and limited bandwidth, offline transfer with AWS Snowball Edge is the recommended approach. It bypasses the network bottleneck, reduces transfer time, and avoids impacting production systems. Online methods like DMS, DataSync, or S3 Transfer Acceleration are constrained by the available bandwidth and are not optimal for this volume.

Exam trap

The trap here is assuming that an online acceleration service like S3 Transfer Acceleration or DataSync can overcome a slow origin uplink, when the bottleneck is the on-premises network itself.

101
MCQmedium

Refer to the exhibit. A data engineer deploys this CloudFormation template to create an AWS Glue job. The job fails on the first run with an error: 'AccessDeniedException: User: arn:aws:sts::123456789012:assumed-role/GlueServiceRole/... is not authorized to perform: s3:GetObject on resource: s3://my-bucket/scripts/etl.py'. What is the most likely cause?

A.The ExecutionProperty MaxConcurrentRuns is set to 1, preventing the job from running.
B.The IAM role associated with the Glue job does not have an S3 GetObject permission for the script location.
C.The MaxRetries is set to 0, so the job does not retry on failure.
D.The script location is incorrectly specified; it should be an S3 URI with bucket and key.
AnswerB

AWS Glue assumes the job's IAM role to fetch the ETL script from S3 before execution begins. The AccessDeniedException naming s3:GetObject on the script path shows the role's policy lacks that permission, so the identity-based policy must grant GetObject on the script location.

Why this answer

The error message indicates that the IAM role 'GlueServiceRole' assumed by the AWS Glue job does not have the s3:GetObject permission for the script object at s3://my-bucket/scripts/etl.py. AWS Glue requires the execution role to have read access to the script location specified in the 'ScriptLocation' parameter. Without this permission, the job fails immediately on startup because it cannot download and execute the ETL script.

Exam trap

The DEA-C01 exam often tests the distinction between permissions errors and configuration errors, where candidates might incorrectly focus on script location format or job parameters instead of recognizing that an AccessDeniedException is a clear IAM permissions issue.

How to eliminate wrong answers

Option A is wrong because ExecutionProperty MaxConcurrentRuns controls how many concurrent runs of the job are allowed, not whether the job can start; it would not cause an AccessDeniedException. Option C is wrong because MaxRetries determines how many times the job retries after a failure, but the job fails on the first run with an access denied error, not a retry-related issue. Option D is wrong because the script location is already specified as an S3 URI (s3://my-bucket/scripts/etl.py), which is the correct format; the error is about permissions, not format.

102
MCQmedium

A data engineer is using AWS Glue Studio to build a visual ETL job that joins a large Amazon S3 dataset with a small reference dataset of country codes. The join is currently implemented as a standard join, and the job runs slowly and shuffles large amounts of data. The engineer wants to optimize performance without changing the output. Which change should the engineer make?

A.Increase the job's timeout setting so the existing standard join has more time to complete.
B.Move the small reference dataset into the Glue Data Catalog as a view and reference it in the join.
C.Replace the standard join with a broadcast join so the small reference dataset is replicated to each executor.
D.Convert the large S3 dataset to a single file before the join to eliminate partitioning overhead.
AnswerC

When one side of a join is small, broadcasting it to every executor avoids shuffling the large dataset across the cluster. Glue Studio exposes a join type option that maps to Spark's broadcast join behavior. This reduces network and shuffle cost while producing the same joined output, directly addressing the slowness described.

Why this answer

A broadcast join replicates the small reference dataset to each executor, eliminating the shuffle of the large dataset that a standard join incurs. Because the output rows are unchanged, this is a pure performance optimization. Timeout increases, file consolidation, and catalog views do not alter the join execution strategy and therefore do not address the shuffle bottleneck.

Exam trap

The trap here is thinking a longer timeout fixes a slow join, when the real cost is shuffle volume that only a broadcast join can remove.

103
MCQhard

A company uses Kinesis Data Analytics for SQL-based real-time analytics on streaming data. They notice that the application is processing data slower than the incoming rate, causing increased latency. Which action is MOST likely to improve the throughput?

A.Increase the number of Kinesis Processing Units (KPUs) for the application
B.Increase the number of shards in the Kinesis data stream
C.Enable auto-scaling on the Kinesis data stream
D.Decrease the retention period of the Kinesis data stream
AnswerA

Kinesis Data Analytics parallelism scales with KPU count; each KPU supplies a fixed slice of CPU and memory, so raising KPUs lifts the application's processing ceiling above the incoming stream rate, directly addressing the throughput shortfall causing latency.

Why this answer

Kinesis Data Analytics for SQL applications processes data using Kinesis Processing Units (KPUs), which define the compute and memory resources available. When the incoming data rate exceeds the processing capacity, increasing the number of KPUs directly scales the application's parallelism and throughput, allowing it to keep up with the stream. This is the most direct way to reduce latency caused by insufficient processing power.

Exam trap

The trap here is that candidates often confuse scaling the source stream (shards) with scaling the analytics application (KPUs), assuming that more shards automatically improve processing throughput, when in fact the application's compute resources are the limiting factor.

How to eliminate wrong answers

Option B is wrong because increasing the number of shards in the Kinesis data stream increases the ingestion capacity and parallelism of the source stream, but it does not directly increase the processing capacity of the Kinesis Data Analytics application; the application must also be scaled (e.g., via KPUs) to consume the additional shards. Option C is wrong because enabling auto-scaling on the Kinesis data stream only adjusts the number of shards based on throughput, which again does not address the application's processing bottleneck. Option D is wrong because decreasing the retention period of the Kinesis data stream only reduces how long data is stored in the stream; it does not affect the processing rate or throughput of the analytics application.

104
MCQhard

A company uses AWS Glue to run ETL jobs that process data from Amazon RDS for MySQL and load it into Amazon S3. The job runs daily and processes incremental changes using the JDBC connection. Recently, the job has been failing with a 'Communications link failure' error. The RDS instance is in a private subnet. Which step should the engineer take first to diagnose the issue?

A.Verify that the IAM role used by Glue has the correct permissions to access RDS.
B.Change the Glue job type from Spark to Python shell.
C.Check the security group and network ACL rules for the RDS instance and the Glue connection.
D.Check that the JDBC driver is compatible with the Glue version.
AnswerC

A 'Communications link failure' means the JDBC connection cannot reach the private RDS endpoint. Verifying security group and network ACL rules for both the RDS instance and the Glue connection confirms whether the required port and subnet traffic are actually permitted.

Why this answer

The 'Communications link failure' error typically indicates a network connectivity issue between AWS Glue and the RDS instance. Since the RDS instance is in a private subnet, the Glue job must be able to reach it via a VPC endpoint or a Glue connection that uses network configuration. Checking the security group (inbound rules for the RDS instance allowing traffic from Glue's elastic network interfaces) and network ACLs (ensuring ephemeral ports are open) is the first logical step to diagnose connectivity.

Exam trap

The trap here is that candidates often jump to IAM permissions or JDBC driver issues first, but the 'Communications link failure' error is a classic network connectivity symptom that requires checking security groups and network ACLs before anything else.

How to eliminate wrong answers

Option A is wrong because IAM permissions control authentication and authorization to AWS services, not network-level connectivity; a 'Communications link failure' is a network error, not an access denied error. Option B is wrong because changing the job type from Spark to Python shell does not resolve network connectivity issues; it only changes the execution environment and may even introduce new limitations for JDBC connections. Option D is wrong because JDBC driver compatibility would cause a different error (e.g., 'No suitable driver' or class not found), not a 'Communications link failure', which is a network timeout or connection reset.

105
Multi-Selectmedium

A company is building a data lake on Amazon S3. Data arrives from multiple sources in JSON, CSV, and Avro formats. The data must be transformed to Parquet and partitioned by date and source. Which TWO services can perform this transformation with minimal custom code? (Choose TWO.)

Select 2 answers
A.Amazon EMR with Spark
B.AWS Lake Formation
C.Amazon Athena CTAS queries
D.AWS Glue ETL jobs
E.Amazon Kinesis Data Firehose
AnswersA, D

Amazon EMR with Spark provides native readers for JSON, CSV and Avro plus a Parquet writer, and supports partitionBy on date and source columns. Transformations are expressed declaratively in Spark SQL or DataFrames, requiring little bespoke code.

Why this answer

Amazon EMR with Spark (A) is correct because Spark natively reads JSON, CSV, and Avro, and can write Parquet with partitionBy('date','source') using built-in DataFrame APIs, requiring only a small script rather than a custom transformation engine. AWS Glue ETL jobs (D) are correct because Glue provides managed, serverless Spark with built-in DynamicFrame readers/writers for JSON, CSV, and Avro, plus automatic schema inference and native Parquet output with partition keys, minimizing custom code. AWS Lake Formation (B) is a permission and metadata/catalog governance layer, not a data transformation engine.

Amazon Athena CTAS (C) can convert formats and partition results, but it is query-oriented and less suited to general multi-format ETL pipelines. Amazon Kinesis Data Firehose (E) is a streaming delivery service that can convert to Parquet via Glue schema, but it does not perform the required multi-source batch transformation and partitioning logic.

Exam trap

The trap here is that candidates often confuse AWS Lake Formation's data catalog and permission features with actual data transformation capabilities, or they assume Kinesis Data Firehose can transform existing S3 objects when it only processes streaming data in transit.

106
MCQmedium

A social media company ingests user activity data from multiple sources using Amazon Kinesis Data Firehose. The data is delivered to Amazon S3 in near-real-time. The company wants to transform the data by adding a timestamp and masking email addresses before storing it in S3. The transformation should be applied to all records. What is the most cost-effective way to implement this transformation?

A.Use Amazon Athena to run a CTAS query that transforms the data and writes to a new location.
B.Use AWS Glue to schedule a batch job every 5 minutes to transform the data.
C.Use Amazon S3 Events to trigger a Lambda function whenever a new object is created.
D.Configure the Firehose delivery stream to invoke a Lambda function for data transformation.
AnswerD

Firehose supports inline Lambda transformation, invoking the function on each batch before delivery to Amazon S3. This applies the timestamp addition and email masking to all records without managing servers, satisfying the stem's cost-effectiveness and universal-transformation constraints more cheaply than separate processing infrastructure.

Why this answer

Amazon Kinesis Data Firehose supports invoking an AWS Lambda function for inline data transformation before delivering data to the destination. This allows adding timestamps and masking email addresses in near-real-time as data flows through the delivery stream. It is the most cost-effective and operationally efficient way because it avoids additional storage, batch processing, or separate compute resources.

Exam trap

The trap is choosing batch or post-storage transformation methods (Glue, S3 Events) instead of inline transformation, which is more efficient and cost-effective for near-real-time streaming data.

How to eliminate wrong answers

Option A is wrong because Amazon Athena is a query service for analyzing data in S3, and running CTAS queries would require additional steps and is not designed for continuous near-real-time transformation; it would also incur costs for scanning data. Option B is wrong because AWS Glue batch jobs every 5 minutes introduce latency and require additional infrastructure and cost, and they are not integrated directly with Firehose. Option C is wrong because S3 Events triggering Lambda would process data after it is stored, adding latency and requiring additional Lambda invocations and possibly reprocessing, and it is not as seamless as Firehose transformation.

107
MCQhard

A data engineer reviews the Glue job configuration. The job fails when processing large datasets. The error message indicates out-of-memory in the executors. Which change to the job configuration will most directly address this issue?

A.Change the worker type from Standard to G.2X.
B.Increase the timeout from 30 to 60 minutes.
C.Increase the number of workers from 5 to 10.
D.Set MaxRetries to 3.
AnswerA

G.2X workers have more memory (8 GB vs 4 GB), directly addressing OOM.

Why this answer

The job fails due to out-of-memory errors in the executors. The current configuration uses 5 Standard workers, each with 16 GB of memory. Changing the worker type to G.2X provides 32 GB per worker, doubling the memory per executor and directly addressing the OOM issue.

Increasing the number of workers (option C) adds more executors but does not increase memory per executor, which may not resolve the issue if each executor still runs out of memory. Options B and D do not affect memory.

108
Multi-Selecthard

A data engineer is configuring an AWS Glue crawler to catalog data in an Amazon S3 bucket. The bucket contains CSV files organized in folders by year and month, and new files are added daily. The engineer wants the crawler to detect schema changes automatically and avoid reprocessing unchanged files on subsequent runs. (Choose two.)

Select 2 answers
A.Enable the crawler's incremental crawling feature to identify and process only new folders and files since the last crawl.
B.Set the crawler's S3 target to exclude the folders that have already been cataloged using an exclude pattern.
C.Enable the crawler's schema change policy to update the table definition in the Data Catalog when new columns are detected.
D.Configure the crawler to use S3 event notifications so it runs only when new objects are created.
E.Change the crawler's output to create a separate table for each partition folder.
AnswersA, C

Incremental crawling lets the crawler track previously crawled S3 paths and process only new folders and files on subsequent runs. This directly satisfies the requirement to avoid reprocessing unchanged files, reducing crawl time and cost. Combined with a schema change policy, the crawler keeps the catalog current while minimizing redundant work.

Why this answer

Incremental crawling lets the crawler process only new folders and files since the last run, avoiding redundant scanning of unchanged data. The schema change policy, when set to update the table in the Data Catalog, automatically incorporates new columns detected during a crawl. Together these settings satisfy both requirements: detecting schema evolution and skipping already-processed files.

Exam trap

The trap here is assuming that S3 event notifications or exclude patterns will make the crawler incremental, when only the built-in incremental crawling feature tracks previously crawled paths and files.

109
Multi-Selectmedium

A data engineer is building an AWS Glue ETL job that reads large Parquet datasets from Amazon S3 and must optimize performance and cost. The engineer wants to reduce the number of small files written to the target S3 prefix and improve read efficiency. (Choose two.)

Select 2 answers
A.Set the Glue job's worker type to G.1X and increase the number of workers to the maximum allowed.
B.Convert the output format to CSV so that multiple small Parquet files can be merged by the S3 service automatically.
C.Use the AWS Glue groupFiles and groupSize options to coalesce small input files before processing.
D.Call coalesce or repartition on the DynamicFrame before writing to reduce the number of output files.
E.Enable the Glue job bookmark to skip previously processed files and reduce the number of files read.
AnswersC, D

groupFiles and groupSize allow the Glue reader to combine many small input files into larger groups, reducing the number of tasks and improving read efficiency. This directly addresses the small-file problem on the input side, which lowers overhead and improves job performance when reading Parquet datasets.

Why this answer

Input-side coalescing with groupFiles and groupSize reduces the number of read tasks, while output-side coalesce or repartition controls how many files are written. Together they attack the small-file problem from both ends, improving performance and lowering cost for Parquet datasets in S3.

Exam trap

The trap here is focusing on worker scaling or bookmarks for small-file issues, when the fix is input grouping and output partition control.

110
MCQeasy

A data engineering team needs to transform CSV files to Parquet format after they land in an S3 bucket. The transformation should be triggered automatically as soon as a new file arrives. Which AWS service is best suited for this task?

A.AWS Batch job submitted by S3 event
B.Amazon EMR cluster running continuously
C.AWS Lambda function triggered by S3 event
D.AWS Glue ETL job scheduled every 5 minutes
AnswerC

S3 event notifications invoke a Lambda function within milliseconds of each object landing, satisfying the immediate-trigger requirement. Lambda natively transforms CSV to Parquet using libraries such as pandas with pyarrow, writing output back to S3, with no cluster provisioning or polling needed for this short, event-driven task.

Why this answer

AWS Lambda functions can be directly triggered by S3 events (e.g., `s3:ObjectCreated:*`) to process newly uploaded CSV files. This serverless approach provides near-instantaneous, event-driven transformation to Parquet without managing any infrastructure, making it the most cost-effective and simplest solution for this specific use case.

Exam trap

The trap here is that candidates often choose AWS Glue (Option D) because it is a dedicated ETL service, but they overlook the requirement for immediate, event-driven processing, which Glue's scheduled jobs cannot provide without additional event-bridge triggers.

How to eliminate wrong answers

Option A is wrong because AWS Batch requires provisioning compute resources and a job queue, adding latency and complexity for a simple file transformation that can be handled by a lightweight Lambda function. Option B is wrong because an Amazon EMR cluster running continuously incurs ongoing costs and management overhead, and is overkill for a simple CSV-to-Parquet conversion triggered by file arrival. Option D is wrong because a scheduled AWS Glue ETL job every 5 minutes introduces unnecessary polling and potential latency (up to 5 minutes), whereas the requirement is for immediate, event-driven processing.

111
Multi-Selectmedium

Which TWO actions can improve the performance of an AWS Glue ETL job that processes large datasets in Amazon S3? (Choose two.)

Select 2 answers
A.Increase the frequency of the Glue crawler.
B.Use a single Availability Zone for the S3 bucket.
C.Increase the number of DPUs allocated to the job.
D.Use columnar file formats like Parquet or ORC.
E.Use a single large file instead of many small files.
AnswersC, D

Increasing DPUs adds parallel executors, so more partitions are processed concurrently during the shuffle and write stages. This directly addresses the stem's large-dataset constraint, where a single worker count throttles throughput. More DPUs raise aggregate memory and CPU, reducing spills to disk and shortening each stage's runtime.

Why this answer

Increasing the number of DPUs allocates more processing power to the Glue job, which can speed up data processing for large datasets. Option D is correct because columnar file formats like Parquet or ORC are more efficient for analytical queries, reduce I/O, and allow better compression compared to row-based formats. Option A is incorrect: increasing crawler frequency only affects the metadata catalog update frequency, not the ETL job performance.

Option B is incorrect: using a single Availability Zone for the S3 bucket does not improve performance and may reduce availability. Option E is incorrect: using a single large file can reduce parallelism, as distributed processing benefits from splitting data into multiple files to be processed in parallel by different executors.

112
Multi-Selectmedium

A data engineer is configuring an AWS Glue crawler to catalog CSV files stored in Amazon S3. The files are organized in prefixes by year and month, and the engineer wants the crawler to detect new partitions automatically and avoid re-crawling unchanged partitions. (Choose two.)

Select 2 answers
A.Set the crawler's schema change policy to update the table definition in the Data Catalog.
B.Configure the crawler to use incremental crawling so it processes only folders added since the last crawl.
C.Ensure the S3 prefixes follow a Hive-style partition naming convention with key=value pairs.
D.Add a path in the crawler configuration that points to each year and month prefix individually.
E.Create a partition index on the Data Catalog table to speed up partition filtering.
AnswersB, C

Incremental crawling lets a Glue crawler examine only new partitions or folders that appeared since the previous run, rather than rescanning the entire S3 prefix. This directly reduces crawl time and cost for a partitioned layout organized by year and month, satisfying the requirement to detect new partitions automatically while avoiding re-crawling unchanged partitions.

Why this answer

Incremental crawling limits each crawler run to newly added folders, so previously cataloged partitions are not rescanned, reducing time and cost. Hive-style key=value prefixes make the year and month directories recognizable as partitions so the crawler can populate them correctly in the Data Catalog. Together they deliver automatic partition detection without reprocessing unchanged data.

Exam trap

The trap here is believing that schema change policies or partition indexes control crawler partition discovery, when incremental crawling and Hive-style layout do.

113
MCQmedium

A data engineer needs to run an AWS Glue extract, transform, and load (ETL) job that joins an Amazon S3-based Parquet dataset with a slowly changing dimension table in Amazon Redshift. The Redshift cluster is in a private subnet and cannot be reached over the public internet. The engineer wants the Glue job to read from Redshift without exposing credentials in the job script. Which combination of actions should the engineer take to meet these requirements?

A.Copy the Redshift data to Amazon S3 using an UNLOAD command, then have the Glue job read the S3 copy, and delete the S3 copy after the job completes.
B.Attach the Glue job to a VPC connection with a NAT gateway, store the Redshift credentials in AWS Secrets Manager, and reference the secret in the Glue job's connection options.
C.Use the Redshift Data API from the Glue job, attaching an IAM role to the job that has redshift-data:ExecuteStatement permissions, and run all transformations in Redshift.
D.Create an AWS Glue connection of type JDBC to the Redshift cluster, attach the job to a VPC connection in the same VPC, and store the Redshift credentials in AWS Secrets Manager for the connection to use.
AnswerD

A JDBC Glue connection combined with a VPC connection places the Glue job's elastic network interfaces in the same VPC as Redshift, enabling private connectivity. Storing credentials in Secrets Manager lets the connection retrieve them at runtime without embedding them in the script, satisfying both the network and credential requirements.

Why this answer

A JDBC connection plus a VPC connection gives the Glue job a private network path to the Redshift cluster, which is required because the cluster is in a private subnet. Storing credentials in Secrets Manager and referencing them through the connection avoids hardcoding secrets in the job script. Together these satisfy both the private connectivity and credential management requirements.

Exam trap

The trap here is assuming that adding a NAT gateway to a VPC connection is enough to reach a private Redshift cluster, when the job actually needs network interfaces in the same VPC.

114
MCQhard

An e-commerce company uses AWS Glue to run ETL jobs that transform clickstream data from Amazon S3. The job reads Parquet files, performs aggregations, and writes the results to Amazon Redshift. The job runs successfully but takes longer than expected. The data volume is increasing. Which design change would MOST improve the job's performance?

A.Write the aggregated results to a single large file instead of multiple partitions.
B.Convert the Parquet files to CSV to simplify the schema.
C.Replace the Redshift target with Amazon Redshift Spectrum.
D.Increase the number of Glue worker nodes (DPUs) for the job.
AnswerD

Adding DPUs increases the number of executors available to process partitions in parallel, directly addressing the growing data volume that is stretching job duration. Glue scales horizontally, so more workers reduce per-node workload and shorten the aggregation and write phases.

Why this answer

Increasing the number of Glue worker nodes (DPUs) directly scales the distributed processing capacity of the ETL job, allowing it to process larger volumes of Parquet data in parallel. This is the most straightforward way to reduce execution time when data volume is growing, as AWS Glue automatically partitions the workload across the additional workers.

Exam trap

The trap here is that candidates assume increasing DPUs always increases cost without considering that the job's runtime reduction often lowers total cost, and they mistakenly choose a data format or target change that does not address the core parallelism issue.

How to eliminate wrong answers

Option A is wrong because writing to a single large file eliminates parallelism in downstream reads and can cause bottlenecks in Redshift's COPY operation, which benefits from multiple files for concurrent loading. Option B is wrong because converting Parquet to CSV increases file size and I/O overhead due to lack of columnar compression and predicate pushdown, degrading performance. Option C is wrong because replacing Redshift with Redshift Spectrum would offload query processing to S3 but does not address the ETL job's performance bottleneck; the job still writes to Redshift, and Spectrum is a query engine, not a write target.

115
MCQhard

A data engineer is building an AWS Glue job that reads from a large Parquet dataset in Amazon S3 partitioned by year/month/day and writes aggregated results to Amazon Redshift. The job currently reads all partitions and takes several hours. The engineer wants the job to process only partitions from the last seven days and reduce runtime. Which change should the engineer make?

A.Pass a pushdown predicate in the from_catalog call that filters on the year, month, and day partition columns for the last seven days.
B.Convert the source data to ORC format to reduce the bytes read per partition.
C.Increase the number of DPUs allocated to the Glue job and enable auto scaling.
D.Enable job bookmarks so the job processes only new files since the last successful run.
AnswerA

A pushdown predicate on partition columns is evaluated by the Glue Data Catalog and S3 reader before data is loaded, so only the matching partitions are listed and read. Restricting to the last seven days dramatically reduces the number of files scanned, cutting runtime and cost while satisfying the requirement.

Why this answer

Partition pruning is the key optimization for partitioned datasets on S3. Supplying a pushdown predicate that references the year, month, and day partition columns lets the Glue reader skip non-matching partitions entirely, so only seven days of data are listed and read. More workers, format conversion, and bookmarks do not restrict the read window and therefore do not solve the runtime problem.

Exam trap

The trap here is reaching for more DPUs or bookmarks, when the actual bottleneck is scanning partitions that fall outside the desired date range.

116
Multi-Selecthard

A data pipeline uses AWS Glue to process large CSV files. The team notices that some jobs fail with out-of-memory errors. Which TWO configuration changes can help mitigate this issue?

Select 2 answers
A.Reduce the number of DPUs to limit concurrency.
B.Increase the number of DPUs for the Glue job.
C.Enable Glue job autoscaling.
D.Convert input files from CSV to Parquet.
E.Enable job bookmarks.
AnswersB, C

Adding DPUs allocates more executors and memory per worker, directly relieving the heap pressure causing out-of-memory failures when processing large CSV files. Horizontal scaling suits Glue's distributed Spark engine, satisfying the stem's memory constraint without altering the transformation logic itself.

Why this answer

Options B and C are correct: increasing the number of DPUs provides more memory, and enabling autoscaling allows the job to automatically scale resources as needed. Option A (reducing DPUs) would worsen the problem by limiting resources. Option D (converting to Parquet) can improve performance but is not a direct configuration change for the Glue job itself.

Option E (job bookmarks) is for incremental processing and does not affect memory.

117
MCQmedium

A data engineer is designing a data ingestion pipeline to load millions of small JSON files from an on-premises FTP server into Amazon S3. The pipeline should minimize cost and operational overhead. Which approach is most suitable?

A.Use S3 Transfer Acceleration to upload files directly from the FTP server
B.Deploy AWS DataSync to transfer files from the FTP server to S3
C.Use AWS Snowball Edge to ship the data to AWS
D.Set up an AWS Direct Connect connection and use AWS CLI to copy files
AnswerB

AWS DataSync is designed for efficient data transfer from on-premises to AWS, handling small files well with minimal operational overhead.

Why this answer

AWS DataSync is the most suitable option because it is designed to efficiently transfer large volumes of data from on-premises storage (including FTP servers) to AWS, handling millions of small files with minimal operational overhead. It automates data transfer, retries, and validation, and it is cost-effective as you pay only for the data transferred, with no need for additional infrastructure or complex scripting.

Exam trap

The trap here is that candidates often assume S3 Transfer Acceleration is a general-purpose acceleration tool for any source, but it only accelerates the upload leg from the client to AWS and does not address the FTP-to-S3 protocol conversion or the orchestration of millions of small files.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration is a feature that speeds up uploads over the internet by using AWS edge locations, but it does not handle the protocol mismatch between FTP and S3; you would still need a client to read from FTP and write to S3, and it adds cost per GB transferred without solving the file ingestion logic. Option C is wrong because AWS Snowball Edge is designed for large-scale, offline data transfers (typically terabytes to petabytes) and is overkill and costly for millions of small files; it also introduces significant latency for shipping and manual handling. Option D is wrong because AWS Direct Connect provides a dedicated network connection but does not automate the transfer of files from an FTP server; you would still need to write custom scripts using AWS CLI to copy files, increasing operational overhead and complexity.

118
MCQhard

A data engineer is troubleshooting a daily batch ingestion pipeline that uses AWS Glue to read CSV files from Amazon S3 and write Parquet files to another S3 bucket. The job runs successfully but takes significantly longer than expected. The engineer notices that the input data is highly skewed with many small files. Which is the most effective optimization to reduce job duration?

A.Change the output format to JSON
B.Enable the 'groupFiles' option in the S3 source configuration
C.Increase the number of DPUs allocated to the job
D.Enable the 'use_glue_schema_registry' option
AnswerB

The 'groupFiles' option coalesces multiple small CSV files into larger partitions during the S3 read, reducing per-file overhead and task count. This directly addresses the many-small-files skew identified in the stem, cutting the job's overall duration.

Why this answer

The 'groupFiles' option in AWS Glue's S3 source configuration allows Glue to combine small files into larger splits, reducing the number of tasks and the overhead associated with processing many small files. This directly addresses the high skew and small file issue, significantly reducing job duration. Option A is incorrect because changing the output format from Parquet to JSON would likely increase file size and processing time, not reduce it.

Option C is incorrect because increasing DPUs may improve parallelism but does not solve the file-level overhead; it could even be wasteful if the bottleneck is task scheduling. Option D is incorrect because enabling the schema registry is unrelated to file grouping; it is used for schema management and validation.

119
MCQmedium

A data engineer must transform data in Amazon S3 using Apache Spark. The transformation logic needs to be reused across multiple AWS Glue jobs, and the engineer wants to version-control the code and run it in a serverless environment without managing clusters. Which approach should the engineer take?

A.Create an AWS Lambda function containing the transformation logic and invoke it from each AWS Glue job using boto3.
B.Package the transformation logic as a Python wheel file stored in Amazon S3, reference it in AWS Glue job parameters, and import it as a module in the job script.
C.Provision an Amazon EMR cluster, store the transformation code in a Git repository, and run Spark jobs on the cluster.
D.Create an AWS Glue job with a script that contains all transformation logic, and copy-paste the code into each job that needs it.
AnswerB

AWS Glue supports referencing additional Python modules via the --extra-py-files job parameter, which can point to a wheel file in Amazon S3. This enables code reuse across jobs, allows version control of the wheel artifact, and keeps the execution serverless. The engineer writes the transformation once, packages it, and imports it in any job that needs it.

Why this answer

AWS Glue supports referencing external Python libraries through the --extra-py-files job parameter. Packaging transformation logic as a wheel file in Amazon S3 allows the code to be versioned, reused across multiple Glue jobs, and executed in a serverless Spark environment. This satisfies both the reusability and serverless requirements without duplicating code or introducing cluster management overhead.

Exam trap

The trap here is assuming that AWS Glue cannot import external code and that all transformation logic must reside in a single job script.

120
MCQmedium

A retail company uses Amazon Kinesis Data Firehose to ingest clickstream data from its website into an Amazon S3 bucket. The data includes fields: user_id, event_type, timestamp, page_url. Recently, the data engineering team noticed that some records have malformed JSON (missing commas, extra brackets) causing delivery failures to S3. The Firehose delivery stream is configured to retry failed records for 300 seconds, after which the records are sent to an S3 bucket for failed records. The team wants to transform the data to correct malformed JSON before delivery to the main S3 bucket. They need a solution that does not require managing servers and can handle high throughput. What should the team do?

A.Configure an AWS Lambda function as a data transformation in Kinesis Data Firehose to correct malformed JSON.
B.Set up an Amazon EMR cluster with Apache Spark to process the data in micro-batches and fix JSON errors.
C.Use an AWS Glue streaming ETL job to read from Firehose and write corrected data to S3.
D.Use Amazon Kinesis Data Analytics with a SQL application to parse and fix JSON.
AnswerA

Lambda transformation runs inline within Firehose, letting you repair malformed JSON before delivery to the main S3 bucket. It is fully serverless and scales with throughput, satisfying the no-server-management and high-volume constraints while preventing records reaching the failure bucket.

Why this answer

Kinesis Data Firehose supports Lambda-based data transformation natively: you attach a Lambda function to the delivery stream, and Firehose invokes it on each batch before delivery. This is serverless, scales with throughput, and lets you parse and repair malformed JSON before it reaches the main S3 bucket.

Exam trap

DEA-C01 often tests whether candidates know Firehose's built-in Lambda transformation versus external processing services — candidates pick Glue or EMR for a problem that Firehose solves natively.

How to eliminate wrong answers

Option B is wrong because EMR requires managing clusters (even with managed scaling) and is overkill for simple per-record JSON repair; it also adds latency and cost. Option C is wrong because Glue streaming ETL reads from Kinesis Data Streams, not directly from Firehose, and would require re-architecting the ingestion path. Option D is wrong because Kinesis Data Analytics (now Managed Service for Apache Flink) processes streams with SQL/Flink but is not the Firehose-native transformation mechanism and adds operational complexity.

121
Drag & Dropmedium

Arrange the steps to create an AWS Glue job that transforms data from Amazon S3 to Amazon Redshift in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First, catalog the source data with a crawler. Then, prepare the ETL script. Configure the job with connections, run it, and finally verify the results in Redshift.

122
Multi-Selectmedium

A company uses AWS Glue to transform data in S3. The Glue job fails with memory errors. Which THREE actions can help resolve this?

Select 3 answers
A.Optimize the transformation to use pushdown predicates.
B.Use a larger worker type (e.g., G.2X).
C.Increase the number of DPUs.
D.Increase the job timeout.
E.Decrease the number of DPUs.
AnswersA, B, C

Pushdown predicates filter rows in the source system before loading, so fewer records reach the Glue workers, directly easing the memory pressure causing the failures. This reduces the data volume shuffled and held in memory during transformation, satisfying the stem's constraint of resolving out-of-memory errors without simply adding capacity.

Why this answer

Options A, B, and C are correct. Using pushdown predicates (A) reduces the amount of data read by filtering at the data source, which can alleviate memory pressure. Using a larger worker type (B), such as G.2X, increases the memory available per worker, directly addressing out-of-memory errors.

Increasing the number of DPUs (C) adds more workers, distributing the memory load. Option D (increasing job timeout) does not solve memory issues, and Option E (decreasing DPUs) would reduce available memory, making the problem worse.

123
MCQhard

A company needs to ingest real-time clickstream data from a web application into Amazon Redshift with minimal latency. The data volume is high and requires processing before loading. Which architecture is MOST appropriate?

A.AWS Glue ETL jobs scheduled every 5 minutes -> Redshift
B.S3 -> Lambda -> Redshift
C.DynamoDB Streams -> Lambda -> Redshift
D.Kinesis Data Streams -> Kinesis Data Firehose -> Redshift
AnswerD

Provides real-time ingestion with transformation capability.

Why this answer

D is correct because Kinesis Data Streams captures high-volume clickstream data in real time, and Kinesis Data Firehose can buffer, transform (e.g., with Lambda), and load the data directly into Amazon Redshift with near-zero latency. This architecture is purpose-built for streaming ingestion with minimal overhead, unlike batch or intermediary storage approaches.

Exam trap

The trap here is that candidates often confuse 'real-time' with 'near-real-time' and choose a batch option like Glue (A) or an indirect streaming path like S3 -> Lambda (B), failing to recognize that Kinesis Data Firehose is the only AWS service that natively integrates streaming ingestion with Redshift without additional latency or complexity.

How to eliminate wrong answers

Option A is wrong because AWS Glue ETL jobs scheduled every 5 minutes introduce batch latency, which violates the 'minimal latency' requirement for real-time clickstream data. Option B is wrong because S3 -> Lambda -> Redshift requires Lambda to write to Redshift, which is inefficient for high-volume streaming data due to Lambda's invocation limits and lack of native streaming buffering, plus S3 adds an unnecessary intermediate storage hop. Option C is wrong because DynamoDB Streams are designed for change data capture from DynamoDB tables, not for ingesting raw clickstream data from a web application; it would require an additional service to capture the data into DynamoDB first, adding complexity and latency.

124
MCQmedium

A data engineer is configuring an AWS Glue crawler to catalog data stored in an Amazon S3 bucket. The data is partitioned by year, month, and day in a Hive-style structure (for example, s3://bucket/data/year=2023/month=01/day=15/). The engineer wants the crawler to recognize the partitions and add them to the AWS Glue Data Catalog. What should the engineer do?

A.Manually create a partitioned table in the AWS Glue Data Catalog using the AWS CLI before running the crawler.
B.Configure the crawler with the 'Add new columns only' option and enable partition detection.
C.Ensure the crawler's include path points to the bucket root and that partition detection is enabled in the crawler configuration.
D.Create a separate crawler for each partition prefix and run them individually.
AnswerC

AWS Glue crawlers automatically detect Hive-style partitions when the include path points to the root of the partitioned data and partition detection is enabled. The crawler parses the key=value structure in the S3 prefixes and adds partition metadata to the Data Catalog. This allows queries in Athena and Redshift Spectrum to use partition pruning for efficiency.

Why this answer

AWS Glue crawlers can automatically detect Hive-style partitions when the include path covers the partitioned data and partition detection is enabled. The crawler reads the key=value structure in the S3 prefixes, creates partition metadata, and updates the Data Catalog. This enables partition pruning in query engines like Amazon Athena and Redshift Spectrum.

Exam trap

The trap here is assuming that partition detection is automatic regardless of crawler configuration or include path scope.

125
Matchingmedium

Match each AWS database service to its primary use case.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Relational database with managed operations

NoSQL key-value and document database

In-memory caching for low latency

Graph database for connected data

Time-series data for IoT and analytics

Why these pairings

Correct matches: RDS -> OLTP relational, DynamoDB -> NoSQL low-latency, Redshift -> data warehousing, ElastiCache -> in-memory caching. Common confusions include swapping RDS and DynamoDB use cases.

126
MCQhard

A company is using AWS Database Migration Service (DMS) to migrate a 2 TB MySQL database to Amazon Aurora MySQL. The migration is taking longer than expected. The source database is in a different AWS region. Which change would MOST likely improve the migration speed?

A.Use a smaller DMS replication instance to reduce costs.
B.Use a Multi-AZ deployment for the DMS replication instance in the target region.
C.Increase the number of parallel tables being migrated.
D.Disable binary logging on the source MySQL database.
AnswerC

DMS parallel load splits large tables into concurrent threads, so raising the number of tables migrated simultaneously uses more of the available bandwidth and replication capacity. This directly addresses the slow 2 TB transfer across regions, where serial table processing underuses throughput.

Why this answer

DMS migrates tables in parallel, and the number of tables loaded concurrently is controlled by the replication task's parallel load settings. Increasing the number of parallel tables (or enabling parallel load with a higher thread count) directly increases throughput when the bottleneck is per-table serialization, which is the most likely cause of a slow 2 TB MySQL-to-Aurora migration. This is the change most likely to improve speed without altering the source or target architecture.

Exam trap

DEA-C01 often tests whether candidates blame network or instance size for slow DMS migrations when the real bottleneck is DMS's default serialized table loading, leading them to pick instance or Multi-AZ changes instead of parallel load tuning.

How to eliminate wrong answers

Option A is wrong because a smaller replication instance reduces CPU, memory, and network capacity, which would slow the migration further rather than speed it up. Option B is wrong because Multi-AZ on the replication instance improves availability, not throughput; the standby does not participate in the migration workload. Option D is wrong because disabling binary logging on the source MySQL database would break ongoing replication (CDC) and is not a supported or safe way to accelerate a DMS migration; binary logs are required for change data capture.

127
MCQeasy

A company needs to ingest data from a relational database into Amazon S3 for analytics. The database is an Amazon RDS MySQL instance. Which AWS service should be used for a one-time historical data load?

A.AWS Database Migration Service (DMS)
B.AWS Glue ETL
C.Amazon Athena
D.Amazon Kinesis Data Firehose
AnswerA

AWS DMS performs bulk migrations and can replicate existing table data from Amazon RDS MySQL into Amazon S3, making it suited to a one-time historical load. It reads source data directly without custom extraction code.

Why this answer

AWS Database Migration Service (DMS) is designed for migrating data from relational databases like Amazon RDS MySQL into targets such as Amazon S3, supporting one-time full loads as well as ongoing replication. For a one-time historical data load, DMS can extract the existing data and land it in S3 efficiently. This makes DMS the correct service for the described scenario.

Exam trap

DEA-C01 often tests the distinction between DMS (database migration/ingestion) and Glue (ETL transformation), catching candidates who pick Glue for a straightforward one-time database-to-S3 load.

How to eliminate wrong answers

Option B is wrong because AWS Glue ETL is a serverless ETL service for transforming data, but it is not the primary tool for bulk database migration; it would require custom connectors and more setup for a one-time RDS-to-S3 load. Option C is wrong because Amazon Athena is a query service that reads data in S3, not a data ingestion or migration tool. Option D is wrong because Kinesis Data Firehose is for streaming data delivery to destinations like S3, not for extracting historical data from a relational database.

128
MCQmedium

A data engineer is using AWS Glue ETL to transform data from an S3 data lake. The job fails with a memory error. Which approach should be used to resolve this issue without major code changes?

A.Rewrite the ETL script in PySpark instead of Scala
B.Change the input file format from CSV to Parquet
C.Increase the number of DPUs allocated to the Glue job
D.Use Amazon EMR instead of AWS Glue
AnswerC

Increasing the number of DPUs allocated to the Glue job directly increases memory and parallelism, which helps resolve memory errors without major code changes.

Why this answer

Increasing the number of DPUs (Data Processing Units) allocated to the Glue job provides more memory and parallelism. Option A is wrong because rewriting in PySpark is a major code change. Option B is wrong because using a smaller file format may not address memory issues.

Option D is wrong because using a different service is unnecessary.

129
MCQeasy

A company uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data delivery is delayed by up to 5 minutes. The engineer wants to reduce the delay to under 1 minute. Which parameter should be adjusted?

A.Enable error logging to CloudWatch.
B.Increase the buffer size in Kinesis Data Firehose.
C.Enable data compression.
D.Decrease the buffer interval in Kinesis Data Firehose.
AnswerD

Firehose buffers incoming records and delivers them when the buffer interval elapses or the buffer size fills. Lowering the buffer interval shortens the maximum wait, reducing delivery latency below one minute, whereas the size parameter alone cannot guarantee that timing.

Why this answer

Decreasing the buffer interval reduces the time Kinesis Data Firehose waits before delivering a batch, thus lowering latency to under 1 minute. Option A is incorrect because error logging to CloudWatch does not affect delivery timing. Option B is incorrect because increasing the buffer size would actually increase the delay as Firehose waits for more data to accumulate.

Option C is incorrect because enabling data compression reduces storage size but has no impact on delivery frequency.

130
MCQeasy

A data pipeline uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The delivery occasionally fails with 'Firehose is throttled'. What should be done to reduce throttling?

A.Enable compression on the Firehose delivery stream
B.Increase the buffer size and buffer interval
C.Decrease the buffer size to flush more frequently
D.Increase the number of shards in the Kinesis stream
AnswerB

Larger buffer reduces the number of write requests.

Why this answer

Increasing the buffer size and buffer interval gives Kinesis Data Firehose more time and data volume to accumulate before delivering to S3, reducing the frequency of PutRecord.Batch calls to the underlying Kinesis stream. This directly mitigates throttling by lowering the request rate, as Firehose throttling typically occurs when the per-shard write throughput limit (1,000 records/second or 1 MB/second) is exceeded.

Exam trap

The DEA-C01 exam often tests the misconception that Firehose throttling is resolved by scaling shards (like in Kinesis Data Streams), but Firehose manages its own internal shards and the correct fix is to adjust buffer settings to reduce API call frequency.

How to eliminate wrong answers

Option A is wrong because enabling compression reduces the data size sent to S3 but does not reduce the number of API calls or the request rate to the Kinesis stream, so it does not address throttling at the stream level. Option C is wrong because decreasing the buffer size causes more frequent flushes, which increases the request rate and exacerbates throttling rather than reducing it. Option D is wrong because Kinesis Data Firehose does not use a Kinesis data stream as its source by default; it uses its own internal stream with a fixed number of shards (default 1), and increasing shards is not a configurable option for Firehose—this option confuses Firehose with Kinesis Data Streams.

131
Multi-Selecthard

A data engineer is designing an ingestion pipeline where AWS Lambda processes records from an Amazon Kinesis Data Stream. During peak traffic, records are being reprocessed and some are lost. The engineer needs to make the consumer resilient to failures and avoid duplicate processing. (Choose two.)

Select 2 answers
A.Enable a dead-letter queue or on-failure destination for the event source mapping.
B.Configure the event source mapping with a bisect-on-error function response and a maximum retry count.
C.Increase the Lambda function's reserved concurrency to match the number of shards.
D.Reduce the batch size to one record to eliminate any possibility of duplicates.
E.Set the starting position to LATEST so the consumer only reads new records.
AnswersA, B

Configuring an on-failure destination, such as an SQS queue or SNS topic, captures records that exceed the maximum retry count so they are not silently lost. This gives operators visibility and a path to reprocess failed records after fixing the root cause, which supports resilience.

Why this answer

Bisect-on-error with a bounded retry count isolates and stops poison records from blocking a shard, while an on-failure destination captures records that exhaust retries so they are not lost. Together they make the Kinesis consumer resilient and give operators a recovery path. Concurrency, starting position, and batch size changes do not address the underlying failure-handling gaps.

Exam trap

The trap here is assuming higher concurrency or smaller batches fix Kinesis consumer reliability, when the real issue is bounded retries and a failure destination for undeliverable records.

132
MCQmedium

A company uses AWS Glue ETL jobs to transform data from Amazon RDS to Amazon S3 daily. The job recently started failing with memory errors. The data volume has grown 3x in the past month. Which change should the data engineer make to resolve the issue?

A.Increase the size of the Amazon RDS instance
B.Switch the Glue job type from Python Shell to Spark
C.Partition the output data in Amazon S3 by date
D.Increase the number of DPUs allocated to the Glue job
AnswerD

Glue allocates memory per DPU, so tripling data volume exhausts the current worker capacity. Increasing DPUs scales the compute and memory available to each task, resolving the out-of-memory failures. This directly addresses the resource constraint created by the threefold data growth.

Why this answer

The Glue job is failing with memory errors due to a 3x increase in data volume. Increasing the number of DPUs (Data Processing Units) allocated to the job provides more memory and compute resources, directly addressing the out-of-memory condition without changing the job logic or architecture.

Exam trap

The trap here is that candidates may confuse scaling the source database (RDS) with scaling the ETL compute (Glue), or assume that output partitioning (S3) will fix an in-memory processing error, when the actual solution is to increase the compute resources allocated to the Glue job.

How to eliminate wrong answers

Option A is wrong because increasing the RDS instance size does not affect the memory available to the Glue ETL job; the bottleneck is in the Glue execution environment, not the source database. Option B is wrong because switching from Python Shell to Spark would change the execution model but does not inherently resolve memory errors; Python Shell jobs are limited to a single executor with fixed memory, while Spark jobs distribute work but still require sufficient DPUs to handle the data volume. Option C is wrong because partitioning output data in S3 by date improves query performance and cost but does not reduce the memory footprint of the Glue job during the transformation phase; the memory error occurs during processing, not during writing.

133
Multi-Selecthard

A data engineer needs to transform data in Amazon S3 using AWS Glue. The job must handle schema evolution and partition pruning. Which THREE features should be used?

Select 3 answers
A.AWS Glue Data Catalog
B.AWS Glue job bookmarks
C.AWS Glue FindMatches transform
D.AWS Glue crawlers
E.Partition indexes
AnswersA, D, E

The AWS Glue Data Catalog stores table definitions, schemas and partition metadata centrally, so Glue jobs read current schema versions and prune partitions using that metadata. It is the component that lets the job handle evolving schemas and avoid scanning irrelevant partitions.

Why this answer

AWS Glue Data Catalog (A) is correct because it stores the table definitions, schemas, and partition metadata that Glue ETL jobs reference, and it is the component that must be updated when schema evolution occurs so the job can read new columns. AWS Glue crawlers (D) are correct because they automatically scan S3 data, detect new or changed columns, and update the Data Catalog with revised schemas and newly discovered partitions, which is exactly how schema evolution is handled. Partition indexes (E) are correct because they accelerate partition pruning by allowing Glue and Athena to look up partitions without listing the entire catalog, dramatically reducing planning time for highly partitioned tables.

AWS Glue job bookmarks (B) only track previously processed data to support incremental loads and do not address schema evolution or partition pruning. AWS Glue FindMatches transform (C) is a machine-learning transform for deduplicating and matching records, which is unrelated to schema evolution or partition pruning.

Exam trap

The DEA-C01 exam often tests the distinction between 'incremental processing' (job bookmarks) and 'schema evolution' (Data Catalog + crawlers), leading candidates to incorrectly select job bookmarks for schema changes.

134
MCQhard

A data engineer is designing a data ingestion pipeline for a social media analytics platform. The pipeline must ingest tweets in real-time, perform sentiment analysis, and store results in Amazon S3. The sentiment analysis is compute-intensive and must be done as the data arrives. The estimated throughput is 10,000 tweets per second. Which architecture is most suitable?

A.Amazon SQS with AWS Lambda pollers to process tweets and store in S3.
B.Amazon EMR with Spark Streaming to process tweets and write to S3.
C.Amazon Kinesis Data Streams with Amazon Kinesis Data Analytics for sentiment analysis, then Kinesis Data Firehose to S3.
D.Amazon API Gateway with AWS Lambda to process each tweet and store in S3.
AnswerC

Scalable real-time stream processing.

Why this answer

The most suitable because Amazon Kinesis Data Streams can ingest up to 10,000 records per second per shard (with shard-level scaling), and Kinesis Data Analytics provides built-in, low-latency stream processing for compute-intensive sentiment analysis using SQL or Apache Flink. Kinesis Data Firehose then reliably buffers and writes the processed results to Amazon S3 without custom code, ensuring near-real-time delivery.

Exam trap

The trap here is that candidates often choose SQS+Lambda (Option A) for simplicity, underestimating the throughput ceiling and polling overhead, while overlooking Kinesis Data Analytics as the only AWS-managed service that natively supports real-time, compute-intensive stream processing without custom infrastructure.

How to eliminate wrong answers

Option A is wrong because Amazon SQS with Lambda pollers introduces polling latency and cannot efficiently handle 10,000 tweets per second; Lambda has a maximum concurrency limit and SQS batch sizes are capped at 10 messages, leading to throttling and backpressure. Option B is wrong because Amazon EMR with Spark Streaming is designed for large-scale batch and micro-batch processing, not for true real-time, per-record sentiment analysis at 10,000 TPS; it incurs startup overhead and is better suited for historical analysis. Option D is wrong because Amazon API Gateway with Lambda processes each tweet synchronously, which cannot sustain 10,000 requests per second without aggressive throttling and cold starts; it also lacks built-in stream buffering and ordering for real-time ingestion.

135
MCQmedium

A company uses AWS Glue to process streaming data from Amazon Kinesis Data Streams. The data is JSON formatted and includes a timestamp field. The company wants to partition the output in Amazon S3 by date and hour, and ensure exactly-once processing semantics. Which combination of configurations should be used?

A.Disable checkpointing and use the 'exactly_once' delivery option in Kinesis Data Streams.
B.Enable checkpointing in the AWS Glue streaming job and specify an S3 location for checkpoint data.
C.Use Amazon DynamoDB as a checkpoint store by configuring the Glue job with a DynamoDB connection.
D.Use Kinesis Client Library (KCL) checkpointing with a DynamoDB table.
AnswerB

Glue streaming jobs support checkpointing to S3 for exactly-once processing.

Why this answer

AWS Glue streaming jobs require checkpointing to track the progress of data consumption from Kinesis Data Streams and to ensure exactly-once processing semantics. By enabling checkpointing and specifying an S3 location, Glue periodically saves the state of processed records, allowing it to resume from the last committed offset in case of failures, thus preventing duplicates or data loss.

Exam trap

The trap here is that candidates confuse the checkpointing mechanism of AWS Glue (which uses S3) with the Kinesis Client Library (KCL) pattern (which uses DynamoDB), leading them to select option D or C, even though Glue streaming jobs do not support DynamoDB for checkpointing.

How to eliminate wrong answers

Option A is wrong because disabling checkpointing removes the mechanism for tracking processed records, making exactly-once semantics impossible; the 'exactly_once' delivery option in Kinesis Data Streams refers to producer-side delivery guarantees, not consumer-side processing semantics. Option C is wrong because AWS Glue streaming jobs do not support DynamoDB as a checkpoint store; they only support S3 for checkpoint data. Option D is wrong because Kinesis Client Library (KCL) checkpointing with DynamoDB is a pattern for custom applications, not for AWS Glue streaming jobs, which manage checkpointing internally via S3.

136
Multi-Selecthard

A company is building a data lake on S3. They have a large volume of CSV files (hundreds of GB) in a source bucket. They need to convert them to Parquet, partition by date, and ensure the data is encrypted at rest with SSE-KMS. The pipeline must be triggered automatically when new files arrive. Which THREE steps should be part of the solution? (Choose THREE.)

Select 3 answers
A.Configure S3 Event Notification to send events to an SQS queue
B.Use Amazon Kinesis Data Firehose to ingest new files
C.Create an AWS Glue ETL job that converts to Parquet and partitions by date
D.Use Amazon Athena CTAS query to convert files in batch
E.Configure the Glue job to use a KMS key for server-side encryption in S3
AnswersA, C, E

SQS can buffer events and trigger a Lambda or Step Functions workflow.

Why this answer

S3 Event Notifications can be configured to send events to an SQS queue when new CSV files arrive. This decouples the ingestion pipeline, allowing the Glue job to poll the queue for new file notifications and trigger processing without tight coupling or polling the S3 bucket directly. SQS provides reliable, scalable message delivery that can trigger downstream ETL workflows.

Exam trap

The trap here is that candidates often confuse batch conversion tools like Athena CTAS with event-driven pipelines, or assume Kinesis Firehose can process existing S3 files, when in fact Firehose only ingests streaming data and cannot read from S3 buckets.

137
MCQeasy

A data engineering team needs to ingest streaming data from an application into Amazon S3 for analytics. The data volume is moderate and the team wants the lowest operational overhead. Which AWS service should they use?

A.Amazon SQS
B.AWS Glue
C.Amazon Kinesis Data Streams
D.Amazon Kinesis Data Firehose
AnswerD

Kinesis Data Firehose is fully managed, automatically scaling and delivering streaming data to Amazon S3 without managing clusters or consumers, meeting the lowest operational overhead constraint. Kinesis Data Streams would require provisioning shards and custom consumers, adding operational burden for moderate volume.

Why this answer

Amazon Kinesis Data Firehose is a fully managed service for loading streaming data into S3 with no code required and minimal operational overhead. Option A is incorrect because Amazon SQS is a message queue service, not designed for streaming data ingestion into S3. Option B is incorrect because AWS Glue is primarily a batch ETL service, not suitable for real-time streaming.

Option C is incorrect because Amazon Kinesis Data Streams requires custom consumers and more management, increasing operational overhead.

138
MCQeasy

A marketing analytics team needs to ingest customer transaction data from an on-premises PostgreSQL database into Amazon S3 for analysis. The data volume is about 10 GB daily, and the team wants to perform full refresh daily (truncate and load) into S3 as Parquet files. The company has a Direct Connect connection to AWS. The team needs a simple, managed solution that minimizes operational overhead. What should the team use?

A.Set up AWS Database Migration Service (DMS) to continuously replicate data to S3 in Parquet format.
B.Use Amazon EMR with a Spark job that reads from PostgreSQL and writes to S3.
C.Use an AWS Glue ETL job with a JDBC connection to the PostgreSQL database, extract data, and write to S3 in Parquet format.
D.Use AWS Data Pipeline with a SQLActivity to extract data and copy to S3.
AnswerC

AWS Glue is serverless and managed, using a JDBC connection to read PostgreSQL and writing Parquet to S3, which minimises operational overhead for the daily 10 GB full refresh. It satisfies the stem's simplicity and managed-solution constraints without managing servers.

Why this answer

AWS Glue ETL jobs provide a serverless, managed environment to extract data from JDBC sources like PostgreSQL, transform it, and write to S3 in Parquet format. It minimizes operational overhead as it handles provisioning, scaling, and job execution. For a daily full refresh, a Glue job can be scheduled to truncate and load data.

Exam trap

DEA-C01 often tests the choice between DMS and Glue for batch ingestion; DMS is for continuous replication, while Glue is for batch ETL with transformations.

How to eliminate wrong answers

Option A is wrong because DMS is designed for continuous replication, not full refresh truncate-and-load; it would require additional steps to truncate and may not produce Parquet directly without transformation. Option B is wrong because Amazon EMR requires managing clusters, which increases operational overhead. Option D is wrong because AWS Data Pipeline is a legacy service and less managed than Glue; it also requires more configuration.

139
Multi-Selectmedium

A data engineer is configuring an AWS Glue job bookmark on a job that reads partitioned Parquet data from Amazon S3 and writes to another S3 location. The engineer notices that reprocessing keeps occurring and wants the bookmark to correctly skip already-processed data. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Ensure the job reads from a source that supports bookmarks, such as S3 or the Glue Data Catalog, rather than an unsupported source
B.Increase the number of worker nodes so the bookmark commits faster between runs
C.Enable the Glue Data Catalog schema evolution setting on the job
D.Set the job's bookmark option to enable and confirm no manual state reset was performed between runs
E.Convert the output to a single unpartitioned file so the bookmark can track it
AnswersA, D

Job bookmarks only track state for supported sources, primarily Amazon S3 and JDBC/Glue Data Catalog tables. If the transform reads from an unsupported origin, Glue cannot persist progress and the job reprocesses everything each run. Confirming the source is a bookmark-capable type is a prerequisite for the feature to function at all in this pipeline.

Why this answer

Reprocessing under a bookmark usually means the feature is not actually tracking the source. The bookmark must be enabled on the job and the source must be a supported type such as S3 or the Glue Data Catalog, with no manual reset in between. Confirming both restores incremental processing so already-consumed partitions are skipped.

Exam trap

The trap here is treating job bookmarks as automatic, when they must be explicitly enabled and only work against supported sources with intact state.

140
Multi-Selectmedium

A data engineer is designing a data ingestion pipeline for IoT sensor data. The data is generated at a high velocity and must be processed in near real-time. The pipeline must also handle bursty traffic. Which TWO AWS services should be combined to achieve this? (Choose TWO.)

Select 2 answers
A.Amazon S3
B.Amazon Kinesis Data Analytics
C.Amazon Simple Queue Service (SQS)
D.AWS Glue
E.Amazon Kinesis Data Streams
AnswersB, E

Kinesis Data Analytics runs SQL or Apache Flink over streaming data, delivering the near real-time transformation the pipeline demands. Paired with a stream, it consumes bursty IoT input continuously, satisfying the high-velocity processing constraint without batch delays.

Why this answer

Amazon Kinesis Data Streams is designed for real-time, high-velocity data ingestion, providing durable, ordered data streams that can handle bursty traffic by scaling shard capacity. Amazon Kinesis Data Analytics can process these streams in near real-time using SQL or Apache Flink, enabling immediate transformations and analytics without needing to store data first.

Exam trap

The DEA-C01 exam often tests the distinction between streaming services (Kinesis Data Streams) and batch/queue services (SQS, S3), so the trap here is assuming SQS can handle real-time streaming or that S3 can serve as a primary ingestion point for high-velocity data.

141
Multi-Selecthard

A data engineer is building an AWS Glue ETL job that must read from an Amazon S3 bucket in the same account and write to an Amazon Redshift cluster in a private VPC. The job must not traverse the public internet and must use least-privilege credentials. (Choose two.)

Select 2 answers
A.Store the Redshift user password in the job script as a plaintext variable so the connection can authenticate.
B.Set the Glue job's connection type to JDBC and rely on the default Glue service role for all S3 and Redshift permissions.
C.Create an internet gateway and a NAT gateway in the VPC so the Glue job can reach Redshift endpoints.
D.Grant the Glue job an IAM role with an S3 bucket policy that allows access only to the required prefixes.
E.Attach the Glue job to a VPC connection that includes a subnet with a route to the Redshift cluster, and configure the required Glue security group rules.
AnswersD, E

Least privilege is enforced through the IAM role attached to the job plus a bucket policy scoped to the needed prefixes. This limits what the job can read and write even if the code is changed later, and it is the credential control the scenario asks for rather than broad account-wide S3 permissions.

Why this answer

Private connectivity to Redshift requires the Glue job to run inside the VPC through a VPC connection with correct subnet routing and security group rules. Least privilege is then enforced by a scoped IAM role and bucket policy, so the two together satisfy both the no-public-internet and least-privilege requirements.

Exam trap

The trap here is assuming a JDBC connection by itself keeps traffic private, when the Glue job must actually be attached to a VPC connection with proper routing and security groups.

142
MCQmedium

A company is using AWS Database Migration Service (DMS) to migrate a 2 TB Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime. The source database is highly active with continuous writes. Which DMS migration type and additional configuration should the engineer use?

A.Use a CDC-only migration task to capture changes from the source.
B.Use a full load migration task and stop the source database before starting.
C.Use a full load migration task with task restart enabled.
D.Use a full load migration task followed by ongoing replication (CDC).
AnswerD

A full load plus change data capture (CDC) copies existing rows then continuously applies ongoing source changes, keeping Aurora PostgreSQL synchronised until cutover. This satisfies the minimal-downtime constraint for a highly active 2 TB Oracle source, unlike full load alone.

Why this answer

A full load migration followed by ongoing replication (CDC) allows the initial data copy to complete while continuously capturing and applying incremental changes from the highly active source. This minimizes downtime by keeping the target database nearly synchronized, requiring only a brief cutover window to stop writes and finalize replication.

Exam trap

The trap here is that candidates often assume a CDC-only task can handle both initial load and ongoing changes, but DMS requires a full load phase to populate the target before CDC can start, and they overlook that a full load alone cannot capture writes occurring during the migration.

How to eliminate wrong answers

Option A is wrong because a CDC-only task cannot migrate the existing 2 TB of data; it only captures ongoing changes, so the target would lack the initial dataset. Option B is wrong because stopping the source database before starting the migration would cause significant downtime, which contradicts the requirement for minimal downtime. Option C is wrong because a full load migration task with task restart enabled only retries the full load on failure; it does not capture ongoing writes after the initial load, so changes made during the migration would be lost, leading to data inconsistency.

143
MCQmedium

A data engineer is building a pipeline that ingests records from an Amazon Kinesis data stream and writes them to Amazon S3 in Parquet format. The engineer wants to use AWS Glue to perform the transformation and needs the pipeline to handle records that arrive out of order and to deduplicate based on a record ID. Which combination of features should the engineer use?

A.Use Amazon Kinesis Data Firehose with a Lambda function for record transformation, and rely on Firehose's built-in deduplication to handle duplicates.
B.Use AWS Glue batch ETL with an S3 source, schedule the job every minute, and use the DropDuplicates transform on the record ID.
C.Use AWS Glue streaming ETL with a Kinesis source, set the job to process records in order, and use the ResolveChoice transform to merge duplicate records.
D.Use AWS Glue streaming ETL with a Kinesis source, apply a windowed deduplication using the record ID and event timestamp, and write the results to S3 in Parquet.
AnswerD

Glue streaming ETL can read from Kinesis and process micro-batches. Applying a windowed deduplication on the record ID and event timestamp handles out-of-order arrival and removes duplicates within the window. Writing the output as Parquet to S3 satisfies the format requirement. This combination directly addresses both the ordering and deduplication needs while leveraging Glue's streaming capabilities.

Why this answer

Glue streaming ETL reading from Kinesis, combined with windowed deduplication on record ID and event timestamp, handles both out-of-order arrival and duplicates. Writing Parquet to S3 meets the format goal. Batch ETL from S3 adds latency and misses cross-batch duplicates, Firehose lacks built-in deduplication and does not use Glue for transformation, and ResolveChoice is not a deduplication tool.

Exam trap

The trap here is assuming that Kinesis Data Firehose or Glue's ResolveChoice transform can deduplicate records by ID, when deduplication requires explicit windowed logic on a key and timestamp.

144
Multi-Selectmedium

A company is using AWS Glue ETL to transform and load data from Amazon S3 to Amazon Redshift. The data engineer notices that the job is taking longer than expected. Which TWO actions can improve the job performance?

Select 2 answers
A.Use Amazon Redshift Spectrum to query data directly.
B.Partition the source data in S3.
C.Increase the number of DPUs for the Glue job.
D.Enable S3 Transfer Acceleration.
E.Use a larger Redshift node type.
AnswersB, C

Partitioning the S3 source data lets AWS Glue read only the relevant partitions rather than scanning the entire dataset, cutting I/O and shuffle volume during the transform stage. This directly addresses the stem's slow job by reducing the bytes read before loading into Amazon Redshift.

Why this answer

Options B and C are correct because partitioning the source data in S3 reduces the amount of data scanned by Glue, improving I/O efficiency, and increasing the number of DPUs adds more parallelism for transformations. Option A is incorrect because Redshift Spectrum is for querying data in S3 directly from Redshift, not for Glue ETL jobs. Option D is incorrect because S3 Transfer Acceleration speeds up uploads to S3 but does not affect Glue job performance during ETL processing.

Option E is incorrect because larger Redshift node types do not impact Glue job execution; they only affect Redshift query performance.

145
MCQhard

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes 500 GB of data. The engineer notices that the job takes several hours and wants to optimize performance. The data is stored in Parquet format and partitioned by date. Which optimization should the engineer implement to improve the job's performance?

A.Convert the Parquet data to CSV to improve read performance.
B.Use predicate pushdown to filter data at the source and partition pruning to read only necessary partitions.
C.Increase the number of DPUs for the Glue job to scale horizontally.
D.Enable job bookmarks to track processed data and avoid reprocessing.
AnswerB

Predicate pushdown and partition pruning allow Glue to read only the relevant partitions and rows from S3, reducing I/O and the amount of data processed. Since the data is partitioned by date, the job can target specific partitions. This significantly speeds up the job and lowers cost. It is a best practice for large datasets in Parquet.

Why this answer

Predicate pushdown and partition pruning are key optimizations for AWS Glue jobs reading partitioned Parquet data. They minimize the amount of data scanned by pushing filters to the source and skipping irrelevant partitions. This reduces I/O and compute time, directly addressing the performance issue without unnecessary cost increases.

Exam trap

The trap here is assuming that simply adding more DPUs will solve performance problems, when actually data layout optimizations like predicate pushdown and partition pruning often yield greater benefits for partitioned Parquet data.

146
MCQeasy

A company needs to ingest data from multiple SaaS applications into Amazon S3. The data sources provide REST APIs. Which AWS service can be used to build a fully managed data ingestion pipeline without writing custom code?

A.Amazon AppFlow
B.Amazon Kinesis Data Streams
C.AWS Lambda with custom code
D.AWS Glue with Python shell
AnswerA

Amazon AppFlow provides managed connectors for SaaS applications and can transfer data into Amazon S3 on a schedule or event trigger, requiring no custom code. Glue or Lambda would demand development effort the stem explicitly excludes.

Why this answer

Amazon AppFlow is a fully managed service designed to transfer data from SaaS applications to AWS services like Amazon S3 without writing custom code. Option B is wrong because Amazon Kinesis Data Streams is primarily for real-time streaming data, not for directly ingesting from SaaS APIs. Option C is wrong because AWS Lambda requires custom code to integrate with SaaS APIs.

Option D is wrong because AWS Glue with Python shell is for ETL transformations and also requires custom scripting.

147
MCQeasy

A data engineer needs to run an AWS Glue ETL job that reads from an Amazon S3 bucket in another AWS account. The bucket owner has granted cross-account access, and the Glue job runs with an IAM role in the engineer's account. The job fails with an access denied error when reading the source objects. Which change is required to allow the Glue job to read the cross-account S3 data?

A.Add an S3 bucket policy in the source account that grants the Glue job's IAM role s3:GetObject and s3:ListBucket on the bucket.
B.Recreate the Glue job in the source account so it uses a role owned by the bucket owner.
C.Attach the AWSGlueServiceRole managed policy to the Glue job's IAM role in the engineer's account.
D.Enable AWS Glue Data Catalog cross-account sharing by creating a resource link to the source account's catalog.
AnswerA

Cross-account S3 access requires both the caller's IAM role to allow the S3 actions and the bucket owner to grant those actions through a bucket policy. Since the role already runs in the engineer's account, the missing piece is the resource-based policy in the source account. Adding s3:GetObject and s3:ListBucket for that role resolves the denial.

Why this answer

S3 authorization evaluates both identity-based policies in the caller's account and resource-based policies on the bucket. Because the Glue job's role lives in a different account than the bucket, the bucket owner must attach a bucket policy granting s3:GetObject and s3:ListBucket to that role. Catalog sharing and service-role policies do not extend to cross-account object reads.

Exam trap

The trap here is assuming a Glue catalog resource link also grants data access, when it only shares metadata and never authorizes S3 object reads.

148
MCQhard

Refer to the exhibit. A data engineer runs the AWS CLI command to describe a Glue job. The job is expected to process new data incrementally using job bookmarks. However, the job reprocesses all data every time it runs. What is the MOST likely reason?

A.The job bookmark option is set to 'job-bookmark-enable' but should be 'job-bookmark-disable'.
B.The job's MaxRetries is set to 0, which disables bookmarks.
C.The ETL script does not use the 'transformation_ctx' parameter in its DynamicFrame transformations.
D.The Glue job's command name is 'glueetl', which does not support job bookmarks.
AnswerC

Job bookmarks rely on the `transformation_ctx` argument to persist state per transformation; without it, Glue cannot track which data each DynamicFrame has already processed, so every run reads the full dataset. Supplying a unique `transformation_ctx` for each source and transformation satisfies the incremental-processing requirement.

Why this answer

AWS Glue job bookmarks rely on the `transformation_ctx` parameter to track state. Without it, Glue cannot identify which data has already been processed, causing the job to reprocess all data on every run. The `transformation_ctx` must be passed to each DynamicFrame transformation (e.g., `apply_mapping`, `filter`, `join`) to enable bookmark-based incremental processing.

Exam trap

The trap here is that candidates often assume bookmarks are controlled only by the job configuration setting (`job-bookmark-enable`) and overlook the critical role of `transformation_ctx` in the ETL script, which is a common oversight in AWS Glue exam questions.

How to eliminate wrong answers

Option A is wrong because `job-bookmark-enable` is the correct setting to enable bookmarks; setting it to `job-bookmark-disable` would disable them, not fix the reprocessing issue. Option B is wrong because `MaxRetries` controls the number of retry attempts on failure and has no effect on job bookmark behavior. Option D is wrong because `glueetl` is the standard command name for ETL jobs and fully supports job bookmarks; the command name does not disable bookmarks.

149
MCQeasy

A company needs to ingest data from an on-premises Oracle database into Amazon S3 on a daily basis. The data volume is about 100 GB per day. Which AWS service is BEST suited for this task?

A.Use AWS DataSync to copy the database files to S3.
B.Use Amazon Kinesis Data Firehose with a database connector.
C.Use AWS Database Migration Service (DMS) to replicate data to S3.
D.Use AWS Glue to extract data from Oracle and write to S3.
AnswerC

AWS DMS performs continuous change data capture from Oracle and writes directly to Amazon S3, handling the 100 GB daily volume without custom extraction code. Its native S3 target endpoint satisfies the daily ingestion requirement, unlike batch-only tools that lack Oracle CDC support.

Why this answer

AWS Database Migration Service (DMS) can continuously replicate data from Oracle to S3, and it supports full load and change data capture (CDC). Option A (AWS DataSync) is for file-based transfers, not database replication. Option B (Amazon Kinesis Data Firehose) is for streaming data, not database pull.

Option D (AWS Glue) is for ETL but does not natively support continuous CDC from Oracle.

150
MCQmedium

A data engineer is configuring an AWS Glue crawler to catalog data stored in Amazon S3. The data is organized as Parquet files under prefixes named by year, month, and day, such as s3://analytics/events/year=2024/month=05/day=17/. Queries in Amazon Athena must use partition pruning to limit scanned data. Which crawler configuration should the engineer choose?

A.Create a crawler for each individual day prefix, such as s3://analytics/events/year=2024/month=05/day=17/, so each day becomes its own table.
B.Create a crawler with the S3 path pointing to s3://analytics/events/ and enable 'Update all new and existing partitions with metadata from the table' so partition metadata stays current.
C.Create a crawler with the S3 path pointing to s3://analytics/events/ and set the 'Schema change policy' to 'Delete tables and columns' so stale partitions are removed.
D.Create a crawler with the S3 path pointing to s3://analytics/events/ and enable the 'Add new columns only' option so new partitions are detected automatically.
AnswerB

The Hive-style year=/month=/day= prefixes are automatically recognized as partitions when the crawler targets the parent prefix. The 'Update all new and existing partitions with metadata from the table' setting ensures that as new date prefixes appear, the crawler adds them and refreshes partition metadata. Athena can then use the partition keys for pruning, scanning only the relevant date ranges instead of the whole bucket.

Why this answer

Hive-style prefixes such as year=/month=/day= are recognized as partition keys when an AWS Glue crawler targets the parent S3 prefix. To keep the catalog current as new date prefixes appear, the crawler should be configured to update all new and existing partitions with metadata from the table. Athena then reads the partition keys from the Data Catalog and prunes to only the relevant prefixes, reducing scanned bytes and query cost.

Exam trap

The trap here is confusing schema-update options such as 'Add new columns only' with partition-update behavior, when partition registration is controlled by the crawler's partition update setting.

← PreviousPage 2 of 6 · 447 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Ingestion Transformation questions.