Courseiva

CCNA Data Ingestion Transformation Questions

75 of 447 questions · Page 1/6 · Data Ingestion Transformation topic · Answers revealed

1
MCQeasy

A data engineer needs to capture change data capture (CDC) events from an Amazon RDS for PostgreSQL database and stream them to Amazon S3 in near real-time. Which AWS service should be used?

A.Amazon S3 Transfer Acceleration
B.Amazon Athena
C.AWS Database Migration Service (AWS DMS)
D.Amazon Kinesis Data Streams
AnswerC

DMS supports ongoing replication (CDC) from databases to S3.

Why this answer

AWS DMS supports continuous replication from PostgreSQL source databases using logical replication slots to capture CDC events in near real-time. It can directly stream these changes to Amazon S3 as a target, making it the correct choice for this use case.

Exam trap

The trap here is that candidates may confuse Kinesis Data Streams as a direct CDC solution, but it requires a separate CDC tool or custom application to capture PostgreSQL changes, whereas AWS DMS natively supports this integration.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Transfer Acceleration is a feature that speeds up uploads to S3 over long distances using edge locations, but it does not capture or stream CDC events from a database. Option B is wrong because Amazon Athena is an interactive query service for analyzing data in S3 using SQL, not a tool for capturing or streaming database changes. Option D is wrong because Amazon Kinesis Data Streams is a real-time streaming service that can ingest data, but it cannot directly capture CDC events from an RDS for PostgreSQL database without additional configuration or a separate CDC connector.

2
MCQhard

A company uses AWS Database Migration Service (DMS) to continuously replicate data from an on-premises Oracle database to Amazon S3 in Parquet format. The replication is used for near-real-time analytics. Recently, the DMS task started failing with an error indicating insufficient memory. The source database is large (2 TB). What should a data engineer do to resolve this issue while minimizing changes to the existing architecture?

A.Change the target format to JSON to reduce memory usage.
B.Split the DMS task into multiple smaller tasks.
C.Use Change Data Capture (CDC) only, without full load.
D.Increase the DMS replication instance size.
AnswerD

DMS memory exhaustion during large-table replication is resolved by scaling the replication instance, adding RAM and CPU. This satisfies the 2 TB source constraint while preserving the existing task, endpoints and Parquet target architecture unchanged.

Why this answer

The error indicates the DMS replication instance is running out of memory during continuous replication of a 2 TB Oracle database to S3 in Parquet format. Increasing the replication instance size (Option D) directly addresses the memory constraint by providing more RAM and processing capacity, which is necessary for handling large volumes of Change Data Capture (CDC) data and Parquet conversion overhead. This solution requires minimal architectural changes, as it only involves modifying the instance class in the DMS task settings.

Exam trap

The trap here is that candidates may think splitting tasks or changing formats reduces memory usage, but the root cause is insufficient instance resources, and AWS DMS tasks require adequate instance sizing for large-scale CDC workloads.

How to eliminate wrong answers

Option A is wrong because changing the target format to JSON would not reduce memory usage; JSON is typically larger than Parquet and would increase memory consumption during serialization, not decrease it. Option B is wrong because splitting the DMS task into multiple smaller tasks would increase complexity and overhead, potentially causing additional memory pressure from multiple connections and task management, and does not directly resolve the insufficient memory error. Option C is wrong because using CDC only without full load ignores the fact that the task is already failing during continuous replication (CDC phase), and the full load may have already completed; disabling full load does not address the memory issue in CDC processing.

3
MCQhard

A data engineer is building an AWS Glue job that reads semi-structured JSON from Amazon S3 and must flatten nested arrays into relational columns before writing to Amazon Redshift. The transformation logic is complex and the engineer wants to unit test it locally without provisioning a cluster. Which Glue capability should the engineer use to develop and test this transformation logic?

A.AWS Glue DataBrew recipe steps executed against a sample dataset.
B.AWS Glue interactive sessions with an AWS Glue Studio notebook.
C.The AWS Glue ETL library (awsglue) run locally with a Python environment and sample datasets.
D.AWS Glue Studio visual job editor with a custom transform node.
AnswerC

The AWS Glue ETL library can be installed and run in a local Python environment, allowing the engineer to execute transformation functions against sample DataFrames without any Glue cluster. This supports true local unit testing of complex flattening logic. It matches the requirement to develop and test without provisioning managed infrastructure, and the same code can later run on Glue with minimal changes.

Why this answer

The AWS Glue ETL library can be installed locally so developers run the same transformation code against sample data without any managed cluster. This enables genuine unit tests of complex nested-array flattening. The visual editor, interactive sessions, and DataBrew all depend on AWS-hosted execution and therefore cannot satisfy the requirement for local testing before deployment.

Exam trap

The trap here is equating AWS Glue Studio notebooks or interactive sessions with local development, when only the installable Glue ETL library actually runs on the developer's machine.

4
MCQmedium

A company has a large volume of CSV files in S3 that need to be transformed into Parquet using AWS Glue. The files are partitioned by date. The engineer wants to minimize costs by processing only new files each day. Which approach should be used?

A.Use S3 partition discovery to automatically read new partitions.
B.Schedule the Glue job to run daily and process all files.
C.Enable job bookmarks in the Glue job.
D.Use S3 Event Notifications to trigger the Glue job on each new file.
AnswerC

Job bookmarks persist state across runs, tracking which S3 objects and partitions were already processed. This satisfies the requirement to process only new files daily, avoiding full re-scans of historical data and cutting Glue DPU costs.

Why this answer

AWS Glue job bookmarks track the state of data processed by a job, so subsequent runs only process new or changed data. When enabled, bookmarks persist the last processed path or partition, allowing the daily job to skip already-transformed files and reduce runtime and cost. This directly satisfies the requirement to process only new files each day.

Exam trap

DEA-C01 often tests the misconception that S3 partition discovery or event notifications provide incremental processing, when only Glue job bookmarks track processed state to avoid reprocessing.

How to eliminate wrong answers

Option A is wrong because S3 partition discovery reads partitions but does not track which files were already processed, so the job would reprocess old data. Option B is wrong because processing all files daily increases cost and runtime, contradicting the goal of minimizing costs. Option D is wrong because S3 Event Notifications trigger the job per file, which can cause excessive job runs and does not inherently track processed state; it also adds complexity and potential duplicate processing.

5
MCQmedium

A data engineer needs to ingest data from an on-premises Oracle database to Amazon S3 daily. The data volume is 500 GB per day, and the network bandwidth is 200 Mbps. The requirement is to minimize the impact on the source database and ensure data integrity. Which combination of AWS services should be used?

A.AWS Database Migration Service (DMS) with S3 as target
B.AWS Glue ETL jobs with JDBC connection
C.Amazon Kinesis Data Firehose with Oracle as source
D.AWS Data Pipeline with SQLActivity
AnswerA

AWS DMS minimizes source impact by using change data capture and supports S3 as a target.

Why this answer

AWS DMS with S3 as target is correct because it supports continuous change data capture (CDC) from Oracle, minimizing impact on the source database by reading redo logs instead of querying tables directly. It can handle 500 GB/day over 200 Mbps (which yields ~2.16 TB/day theoretical max) and ensures data integrity via transactional consistency and validation checksums. DMS also automatically partitions large datasets and can resume from failures, making it ideal for daily bulk loads.

Exam trap

The trap here is that candidates assume AWS Glue or Data Pipeline are suitable for database ingestion, but they lack the CDC and low-impact features of DMS, which is specifically designed for minimal source database load during large-scale migrations or replication.

How to eliminate wrong answers

Option B is wrong because AWS Glue ETL jobs with JDBC connection would pull data via full table scans, placing significant load on the Oracle database and lacking native CDC capabilities, which violates the requirement to minimize source impact. Option C is wrong because Amazon Kinesis Data Firehose cannot use Oracle as a direct source; it ingests from streaming sources like Kinesis Data Streams, not relational databases via JDBC. Option D is wrong because AWS Data Pipeline with SQLActivity uses a polling-based approach that repeatedly queries the source, causing unnecessary overhead, and does not support CDC or optimized large-volume transfers like DMS.

6
Multi-Selectmedium

A data engineer is building an AWS Glue ETL job that reads data from Amazon S3 and must write the output partitioned by year, month, and day for efficient downstream querying in Amazon Athena. The engineer wants the job to create the partition folders and register them in the Glue Data Catalog automatically. (Choose two.)

Select 2 answers
A.Call the Glue catalog's create_partition API for each partition after the job writes files.
B.Set the DynamicFrame write option partitionKeys to the year, month, and day columns.
C.Enable the job's 'Update the Data Catalog' option so new partitions are added during the write.
D.Run an AWS Glue crawler on the output prefix after every job run.
E.Store the output as a single Parquet file per run to preserve partition metadata.
AnswersB, C

The partitionKeys write option tells Glue how to split output files into directory hierarchies based on the specified columns. Setting it to year, month, and day produces the year=/month=/day= folder structure that Athena and other engines expect. This is the primary mechanism for generating partitioned output from a Glue job.

Why this answer

Producing partitioned output and registering it in the Glue Data Catalog from a Glue job is achieved by specifying partitionKeys on the write and enabling the job to update the Data Catalog. Together these create the year/month/day folder structure and add the corresponding catalog partitions during the run. Manual API calls or a separate crawler can work but are not the automatic in-job mechanism the scenario requires.

Exam trap

The trap here is assuming a post-run crawler is the only way to register partitions, when the Glue write itself can update the catalog if configured to do so.

7
Multi-Selecthard

A company is ingesting Apache logs from multiple web servers into AWS. The logs are sent via Amazon CloudWatch Logs to a subscription filter that delivers to a Lambda function. The Lambda function parses the logs and writes to Amazon S3. However, there is a significant backlog. Which THREE actions can reduce the backlog?

Select 3 answers
A.Route the CloudWatch Logs subscription to an Amazon SQS queue first
B.Increase the Lambda function memory allocation
C.Increase the Lambda function reserved concurrency
D.Change the Lambda function runtime from Python to Node.js
E.Increase the Lambda function maximum concurrency (unreserved account concurrency)
AnswersB, C, E

More memory also increases CPU, speeding up processing.

Why this answer

Increasing the Lambda function's memory allocation also increases its CPU allocation, allowing the function to process each log event faster. This reduces the per-invocation processing time, enabling the function to handle more log data per unit time and thus reduce the backlog.

Exam trap

The trap here is that candidates may confuse 'reserved concurrency' with 'maximum concurrency' or think that adding an SQS queue always improves throughput, when in fact it can add latency and does not address the root cause of slow per-invocation processing.

8
MCQmedium

A data engineer is loading data from Amazon S3 into an Amazon Redshift cluster using the COPY command. The S3 bucket contains 500 Parquet files, each about 200 MB, in a single prefix. The COPY job is running slowly and consuming excessive cluster resources. The engineer wants to improve performance without changing the data format or the cluster size. Which action should the engineer take?

A.Use the COPY command with the PARALLEL OFF option to reduce the number of slices used.
B.Split the data into multiple prefixes and run multiple COPY commands in parallel, or use a manifest file to distribute the load across slices.
C.Convert the Parquet files to CSV and load them with the CSV option to enable faster parsing.
D.Add the COMPUPDATE OFF and STATUPDATE OFF options to the COPY command to skip compression and statistics updates.
AnswerB

Redshift COPY parallelizes across slices by reading multiple files concurrently. When many files are in one prefix, the leader node can become a bottleneck and slice distribution may be uneven. Splitting into multiple prefixes or using a manifest file lets COPY distribute files more evenly across slices, improving throughput and reducing resource contention without changing the data format or cluster size.

Why this answer

Redshift COPY achieves high throughput by having multiple slices read files in parallel. When hundreds of files sit in one prefix, the leader node may serialize listing and assignment, and some slices may read more files than others, causing skew and resource contention. Distributing files across prefixes or using a manifest gives COPY finer-grained control over parallel reads, balancing work across slices and improving load speed without altering the format or cluster size.

Exam trap

The trap here is assuming that COPY always parallelizes perfectly regardless of file layout, when in fact a single prefix with many files can create a leader-node and slice-distribution bottleneck.

9
MCQeasy

A data engineer needs to transform JSON data from Amazon S3 into Parquet format using AWS Glue. The data contains nested fields. Which Glue feature should the engineer use to define the schema and handle the nested structure?

A.Use the 'FindMatches' transform to identify duplicates.
B.Use the 'DropFields' transform to remove nested fields.
C.Use the 'Relationalize' transform in a Glue ETL script.
D.Use the 'Spigot' transform to write sample data.
AnswerC

Relationalize flattens nested JSON structures into separate relational tables linked by keys, letting Glue's DynamicFrame handle arrays and structs that a flat schema cannot represent. This satisfies the requirement to define schema and process nested fields before writing Parquet.

Why this answer

The Relationalize transform in AWS Glue is specifically designed to convert nested JSON data into a relational format. It flattens nested structures by creating separate tables for arrays and nested objects, which can then be written to Parquet. This is the appropriate feature to handle nested fields when transforming JSON to Parquet.

Exam trap

DEA-C01 often tests the misconception that other transforms like DropFields or FindMatches can handle nested structures, when Relationalize is the specific tool for that purpose.

How to eliminate wrong answers

Option A is wrong because FindMatches is used for deduplication and record matching, not for schema definition or handling nested structures. Option B is wrong because DropFields simply removes fields, which does not help in defining a schema for nested data. Option D is wrong because Spigot is used for debugging by writing sample data to S3, not for transforming nested structures.

10
MCQeasy

A data engineer needs to run a transformation on a small dataset of 500 MB stored in Amazon S3 and load the result into Amazon Redshift. The transformation logic is simple column renaming and filtering. The engineer wants to minimize operational overhead and avoid managing servers. Which approach is most appropriate?

A.Launch an Amazon EMR cluster, run a Spark job, and terminate the cluster after the job completes.
B.Create an AWS Lambda function to read the S3 object and write directly to Redshift.
C.Use AWS Glue with a Python shell job to perform the transformation and load the data.
D.Provision an Amazon EC2 instance, install Apache Spark, and run the transformation as a scheduled cron job.
AnswerC

AWS Glue Python shell jobs are serverless and designed for small to medium workloads that do not require the distributed processing of Apache Spark. For a 500 MB dataset with simple column renaming and filtering, a Python shell job minimizes operational overhead, starts quickly, and integrates with Redshift for loading. This matches the requirement to avoid managing servers.

Why this answer

AWS Glue Python shell jobs are serverless and optimized for small to medium ETL workloads that do not need Spark's distributed processing. For a 500 MB dataset with simple transformations, they provide the lowest operational overhead while integrating with S3 and Redshift. EMR, EC2, and Lambda each introduce either management burden or technical limits that make them less suitable for this scenario.

Exam trap

The trap here is assuming that any serverless option works equally well, when Lambda's 15-minute timeout and limited resources make it a poor fit for ETL and loading into Redshift.

11
MCQmedium

A data engineer needs to ingest data from an on-premises Apache Kafka cluster into Amazon S3. The data volume is about 10 TB per day. The engineer wants to set up a managed Kafka connector. Which AWS service should they use?

A.AWS Database Migration Service
B.AWS Lambda with Kafka trigger
C.Amazon MSK Connect
D.Amazon Kinesis Data Streams
AnswerC

Amazon MSK Connect is the managed Kafka Connect service, so it runs connectors that stream from the on-premises Kafka cluster into Amazon S3 without you operating Connect workers, meeting the managed-connector and 10 TB per day requirement.

Why this answer

Amazon MSK Connect is a managed Kafka connector service that integrates with Amazon MSK or self-managed Apache Kafka clusters to stream data into Amazon S3 using Kafka Connect. It handles the 10 TB/day volume efficiently with auto-scaling and checkpointing, making it the correct choice for a managed connector setup.

Exam trap

The trap here is that candidates confuse Amazon MSK Connect with Amazon Kinesis Data Streams, thinking both are streaming services, but MSK Connect is specifically a managed Kafka connector service for Kafka-to-S3 ingestion, while Kinesis is a separate streaming platform.

How to eliminate wrong answers

Option A is wrong because AWS Database Migration Service (DMS) is designed for database migrations and continuous replication, not for ingesting data from Kafka into S3. Option B is wrong because AWS Lambda with Kafka trigger is event-driven and not suitable for high-volume, continuous streaming of 10 TB/day due to concurrency limits and lack of managed checkpointing for large-scale Kafka ingestion. Option D is wrong because Amazon Kinesis Data Streams is a separate streaming service, not a managed Kafka connector; it cannot directly connect to an on-premises Kafka cluster as a connector.

12
MCQhard

A company is building a data lake on S3 and needs to ingest data from on-premises Oracle database. The data is 5 TB and changes incrementally. The ingestion must capture changes in near real-time (less than 1 minute latency) and be cost-effective. Which approach should be used?

A.Use AWS Database Migration Service (DMS) with ongoing replication to S3
B.Use Amazon Kinesis Data Firehose with an Oracle JDBC connector
C.Use AWS Glue to perform a full table export daily
D.Use AWS DataSync to sync the Oracle data files to S3
AnswerA

DMS supports CDC and can replicate changes to S3 with low latency.

Why this answer

AWS DMS with ongoing replication captures incremental changes from Oracle using its native change data capture (CDC) mechanism, such as Oracle LogMiner or binary logs, and streams them to S3 in near real-time with latency under 1 minute. This approach is cost-effective because DMS charges only for the compute resources used during replication, and S3 storage is inexpensive, making it ideal for a 5 TB dataset with continuous changes.

Exam trap

The trap here is that candidates often confuse Kinesis Data Firehose's ability to accept data from custom sources with native JDBC support, leading them to choose Option B, but Firehose lacks built-in CDC connectors for relational databases like Oracle.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Firehose does not natively support a JDBC connector for Oracle; it ingests data from sources like Kinesis Data Streams, AWS IoT, or custom HTTP endpoints, and using a JDBC connector would require custom code and add complexity without guaranteeing sub-minute latency for CDC. Option C is wrong because AWS Glue performing a full table export daily cannot meet the near real-time requirement of less than 1 minute latency; it is a batch-oriented service designed for periodic ETL jobs, not continuous change capture. Option D is wrong because AWS DataSync is designed for one-time or scheduled bulk data transfers of files or objects, not for capturing incremental database changes from Oracle; it syncs data at the file level, not the row-level CDC needed for a database.

13
Multi-Selecteasy

A data engineer needs to ingest data from a SaaS application that sends webhooks in JSON format. The data must be stored in S3 for batch analysis. Which AWS services can receive the webhooks and store the data in S3 with minimal custom code? (Choose TWO.)

Select 2 answers
A.AWS Lambda with S3 SDK
B.AWS Glue with a Python shell job
C.Amazon Kinesis Data Streams
D.Amazon API Gateway with S3 integration
E.Amazon API Gateway with Kinesis Data Firehose integration
AnswersD, E

API Gateway can directly write to S3.

Why this answer

Amazon API Gateway with S3 integration (Option D) allows you to create a REST API that directly writes incoming webhook payloads to an S3 bucket without any custom code, using a proxy integration with an AWS service. This minimizes custom code because the integration handles the mapping and storage automatically.

Exam trap

The trap here is that candidates often assume Lambda is the only serverless option for webhook ingestion, overlooking API Gateway's direct S3 integration, or they mistakenly think Kinesis Data Streams can directly receive HTTP requests without a custom producer.

14
MCQmedium

A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 data source, applies a filter transformation, and writes to Amazon S3 in Parquet. The engineer notices that the job is reading all files in the prefix, including files that do not match the expected schema, causing job failures. Which action should the engineer take to ensure only valid files are processed?

A.Enable job bookmarks and set the transformation context to filter invalid records
B.Increase the number of DPUs and enable auto-scaling to handle the invalid files
C.Use a Glue crawler with a custom classifier to catalog only valid files and update the job to read from the catalog table
D.Configure the S3 data source with a path that includes only valid partitions or use a glob pattern to exclude invalid files
AnswerD

Narrowing the S3 source path with a glob pattern or restricting to specific partitions ensures the Glue job only reads files that match the expected schema. This directly prevents invalid files from being processed and avoids job failures. It is the simplest and most targeted fix for the described problem.

Why this answer

The root cause is that the job reads all objects under the prefix, including files that do not match the schema. Restricting the S3 source path with a glob pattern or limiting to valid partitions ensures only conforming files are read. This is a configuration change that directly prevents the failures without adding compute or custom classification logic.

Exam trap

The trap here is reaching for job bookmarks or crawlers to solve a schema mismatch; those features handle incremental processing and cataloging, not selective file exclusion at read time.

15
Multi-Selecthard

A data engineer is designing a data transformation pipeline using AWS Glue. The source data is in Amazon S3 in Parquet format, and the transformed output must be written to another S3 bucket in Parquet format partitioned by year, month, day. The pipeline should handle incremental updates efficiently. Which three features should the engineer use? (Choose THREE.)

Select 3 answers
A.AWS Glue job bookmarks to track processed data
B.Use AWS Glue JobWatch for monitoring job progress
C.Use DynamicFrames instead of Spark DataFrames for schema handling
D.Enable partition pruning in the Glue job
E.Use Spark SQL for transformations
AnswersA, C, D

Enables incremental processing.

Why this answer

AWS Glue job bookmarks track processed data by recording the state of previously processed files and partitions, enabling incremental processing of new or changed data in subsequent runs. This is essential for efficiently handling incremental updates without reprocessing the entire dataset.

Exam trap

The trap here is that candidates may confuse general-purpose tools like Spark SQL or monitoring concepts with the specific AWS Glue features designed for incremental processing and partitioning, leading them to select options that are technically possible but not the three required features.

16
MCQmedium

A company uses AWS Glue crawlers to populate the Data Catalog from data in Amazon S3. The crawler fails to update the schema when new columns are added to the CSV files. What is the most likely cause?

A.The S3 bucket has versioning enabled.
B.The crawler is configured to only crawl new partitions.
C.The IAM role for the crawler lacks permissions to read the new columns.
D.The crawler uses a custom classifier that defines a fixed schema.
AnswerD

A custom classifier overrides the crawler's built-in schema inference, so it applies the fixed column definitions you supplied rather than re-reading the CSV headers. New columns in the S3 files are therefore ignored, which is exactly the failure described. Removing the classifier or updating its schema restores automatic detection.

Why this answer

AWS Glue crawlers use classifiers to infer the schema of data in Amazon S3. If a custom classifier is used and it defines a fixed schema, the crawler will not detect new columns because the classifier's schema is static. The crawler relies on the classifier to parse the data and determine the schema; if the classifier does not include the new columns, they will be ignored.

Therefore, the most likely cause is that the crawler uses a custom classifier with a fixed schema.

Exam trap

DEA-C01 often tests the misconception that IAM permissions or S3 versioning affect schema detection, when the real issue is the use of a custom classifier with a fixed schema that prevents dynamic schema inference.

How to eliminate wrong answers

Option A is wrong because S3 bucket versioning does not affect the crawler's ability to detect schema changes; versioning is about object versions, not schema inference. Option B is wrong because configuring the crawler to only crawl new partitions would affect which partitions are crawled, but if new columns are added to existing partitions, they might still be detected if the classifier is dynamic. However, the question states the crawler fails to update the schema when new columns are added, which is more directly related to the classifier.

Option C is wrong because if the IAM role lacks permissions to read the new columns, the crawler would fail to read the data entirely, not just fail to update the schema; also, permissions are typically at the bucket/object level, not column level.

17
MCQeasy

A data engineer needs to transform JSON data from an S3 bucket using AWS Glue. The JSON contains nested arrays and objects. Which Glue transform is best suited for flattening nested structures?

A.Unnest
B.ResolveChoice
C.Relationalize
D.Map
AnswerC

Relationalize converts nested JSON arrays and objects into separate, related DynamoDB-style tables, unnesting the hierarchy. This flattening produces the tabular structure Glue and Athena require, which ApplyMapping or Filter alone cannot achieve for deeply nested data.

Why this answer

The Relationalize transform is specifically designed to flatten nested JSON structures (arrays and objects) into a set of related tables, making it ideal for this use case. It automatically handles complex nesting by creating separate DataFrames for each nested level and linking them via foreign keys, which is exactly what is needed when ingesting JSON with nested arrays and objects into a relational format.

Exam trap

The trap here is that candidates often confuse the generic Spark SQL function `explode` (or the concept of 'unnesting') with a named AWS Glue transform, leading them to select 'Unnest' even though it does not exist as a Glue transform and would require manual handling of multiple nesting levels.

How to eliminate wrong answers

Option A (Unnest) is wrong because AWS Glue does not have a built-in transform named 'Unnest'; this is a Spark SQL function (e.g., `explode`) but not a named Glue transform, and it would require manual handling of multiple nesting levels. Option B (ResolveChoice) is wrong because it is used to resolve schema ambiguities (e.g., when a column has mixed types like string and int) and does not flatten nested structures. Option D (Map) is wrong because it applies a function to each record in a DynamicFrame for row-wise transformations, but it does not inherently flatten nested arrays or objects—you would need to write custom logic to handle the nesting.

18
Multi-Selecthard

A data engineer is optimizing an AWS Glue ETL job that reads a large dataset from Amazon S3 and writes to Amazon Redshift. The job currently runs slowly and consumes many DPUs. The engineer wants to improve performance and reduce cost. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Increase the number of DPUs to the maximum allowed for the job
B.Partition the source data in Amazon S3 and use predicate pushdown in the Glue job
C.Enable job bookmarks to process only new data on subsequent runs
D.Convert the output to CSV instead of Parquet to reduce write time
E.Disable auto-scaling to keep the job at a fixed capacity
AnswersB, C

Partitioning the S3 data and using predicate pushdown allows Glue to read only the partitions needed for the query, reducing I/O and the amount of data processed. This improves job performance and lowers DPU usage. It is a targeted optimization that addresses the root cause of slow reads from large datasets.

Why this answer

Job bookmarks reduce the data read on subsequent runs by tracking processed files, and partitioning with predicate pushdown reduces the data scanned per run. Together they lower runtime and DPU consumption, improving performance while reducing cost. Increasing DPUs or converting to CSV would increase cost or hurt performance, and disabling auto-scaling removes a cost-saving feature.

Exam trap

The trap here is thinking that adding more DPUs is always the answer to a slow Glue job; the question asks to improve performance and reduce cost, which requires reducing data processed rather than scaling compute.

19
MCQhard

A data pipeline uses AWS Glue to read from an Amazon S3 bucket containing millions of small CSV files (each < 1 MB). The ETL job is slow. Which optimization would most improve performance?

A.Write the ETL script using PySpark instead of Scala
B.Increase the number of Glue workers
C.Use the G.1X worker type for more memory
D.Use S3 file grouping to combine small files
AnswerD

Glue's S3 file grouping (groupFiles) coalesces many small objects into larger partitions per task, cutting per-file listing, opening and request overhead. Millions of sub-1 MB CSVs otherwise dominate runtime, so grouping directly addresses the small-file constraint.

Why this answer

AWS Glue (and Spark generally) performs poorly with millions of tiny files because each file requires a separate S3 GET request, metadata operation, and task scheduling overhead. S3 file grouping (via 'groupFiles' and 'groupSize' job parameters) coalesces multiple small files into a single read per Spark partition, drastically reducing the number of S3 operations and task overhead. This directly addresses the root cause of the slowness.

Exam trap

DEA-C01 often tests the instinct to 'add more workers' for any slow Glue job — candidates miss that small-file overhead is an I/O/metadata problem that horizontal scaling cannot fix.

How to eliminate wrong answers

Option A is wrong because PySpark vs. Scala is a language choice — both compile to the same Spark execution engine, so performance is essentially identical for this workload. Option B is wrong because adding workers increases parallelism but each worker still opens millions of tiny files, so the per-file overhead remains and may even worsen due to more concurrent S3 requests.

Option C is wrong because G.1X provides more memory per worker, but the bottleneck is I/O and metadata overhead, not memory — more memory does not reduce the number of file opens.

20
Multi-Selecthard

A data engineer is implementing a CDC (Change Data Capture) pipeline from a relational database to Amazon S3 using AWS Database Migration Service (DMS). Which TWO configurations are required for continuous replication?

Select 2 answers
A.Define transformation rules in the DMS task.
B.Enable binary logging on the source database.
C.Configure a VPC endpoint for DMS.
D.Enable 'Full load' and 'Ongoing replication' in the task.
E.Pre-create the target table in S3.
AnswersB, D

Binary logging records every row-level change on the source, which DMS reads to capture ongoing inserts, updates and deletes. Without it, the source retains no change history, so continuous replication cannot proceed beyond the initial load.

Why this answer

Option B is correct because AWS DMS continuous replication (CDC) requires the source relational database to expose its change stream — for MySQL/MariaDB that means binary logging (log_bin) enabled with binlog_format=ROW, and for other engines the equivalent (e.g., Oracle supplemental logging, SQL Server MS-CDC). Option D is correct because the DMS task's migration type must be set to 'Full load and ongoing replication' (or 'Ongoing replication' if data is already loaded) so DMS performs the initial load and then continuously applies CDC changes to Amazon S3. Option A is not required — transformation rules are optional and only used to rename, filter, or modify schema/data during migration.

Option C is not required — a VPC endpoint is only needed for private connectivity to services like S3 when using a VPC-attached replication instance, not as a general CDC requirement. Option E is not required — DMS writes to S3 as CSV or Parquet files and does not need a pre-created target table.

Exam trap

DEA-C01 often tests that 'continuous replication' requires both source-side logging and task-level CDC configuration — candidates may select transformation or networking options that sound relevant but don't enable change capture.

21
MCQeasy

A data engineer needs to load data from an Amazon S3 bucket into an Amazon Redshift cluster as part of an ETL pipeline. The source files are already in Parquet format and the engineer wants the fastest load with minimal transformation. Which Redshift load method should the engineer use?

A.Use the COPY command with the PARQUET format option.
B.Use an AWS Glue ETL job to read the Parquet files and write to Redshift using the Redshift connector.
C.Use the Redshift UNLOAD command to move the Parquet files into the cluster.
D.Use an Amazon Kinesis Data Firehose delivery stream with Redshift as the destination.
AnswerA

The Redshift COPY command is the native, high-throughput bulk load utility and it reads Parquet directly from Amazon S3. Using the PARQUET format option lets Redshift parse the columnar file natively, which is faster than loading text formats and requires no intermediate transformation. It is the recommended approach for loading Parquet files into Redshift.

Why this answer

The Redshift COPY command is the native bulk loader and supports Parquet directly via the PARQUET format option, which reads columnar data efficiently from S3. For Parquet files with no required transformation, COPY is faster and simpler than a Glue job, and the other options either move data the wrong direction or add streaming overhead.

Exam trap

The trap here is reaching for a managed ETL service or streaming delivery when the native COPY command already loads Parquet from S3 directly and is the fastest option for untransformed data.

22
MCQmedium

A data engineer needs to transform data in Amazon S3 using SQL statements without managing any infrastructure. The transformations are simple projections and filters, and the engineer wants the results written back to S3 in Parquet. Which AWS service should be used?

A.Amazon Athena with a CREATE TABLE AS SELECT statement
B.AWS Glue ETL with a PySpark script
C.Amazon EMR with Hive on a transient cluster
D.Amazon Redshift Spectrum with an external schema
AnswerA

Athena is serverless and runs standard SQL against data in S3, and CREATE TABLE AS SELECT writes the transformed result back to S3 in a specified format such as Parquet. This matches the requirement for simple SQL transformations with no infrastructure to manage.

Why this answer

Amazon Athena is fully serverless and executes SQL directly against S3 data. A CREATE TABLE AS SELECT statement reads the source, applies projections and filters, and writes the result to S3 in a chosen format like Parquet. This satisfies the SQL and no-infrastructure requirements without cluster or job management.

Exam trap

The trap here is equating SQL-on-S3 with Glue or EMR, when Athena is the serverless option purpose-built for running SQL directly against S3 data.

23
MCQmedium

A company needs to transform JSON data from an Amazon S3 bucket into Parquet format and load it into an Amazon Redshift cluster. The transformation includes joining with a reference table stored in Amazon RDS. Which AWS service is BEST suited for this task?

A.AWS Data Pipeline
B.AWS Glue ETL job
C.Amazon Athena
D.Amazon EMR with Spark
AnswerB

AWS Glue ETL jobs run Apache Spark, so they can read JSON from S3, join it against the RDS reference table via a JDBC connection, and write Parquet into Redshift. This satisfies the transformation and cross-source join requirement in one managed job.

Why this answer

(AWS Glue ETL job) is the best choice because it natively integrates with S3, RDS, and Redshift. Glue can read JSON from S3, connect to RDS via JDBC to join with the reference table, transform the data to Parquet using its built-in converter, and write directly to Redshift. Option A (AWS Data Pipeline) is older and less integrated for this purpose.

Option C (Amazon Athena) can query S3 and convert to Parquet but cannot natively join with RDS without additional services. Option D (Amazon EMR with Spark) is possible but requires more setup and maintenance.

24
MCQmedium

A company uses Amazon Kinesis Data Firehose to ingest application logs into an Amazon S3 bucket. The logs are in JSON format. The data engineering team wants to convert the logs from JSON to Parquet format before landing in S3. What is the most cost-effective way to achieve this?

A.Use Amazon Athena to query the JSON data and write results in Parquet format.
B.Configure the Firehose delivery stream to convert the data to Parquet using a schema from AWS Glue.
C.Use an AWS Lambda function to transform each record to Parquet and send to Firehose.
D.Use an AWS Glue ETL job to run on a schedule and convert JSON to Parquet in S3.
AnswerB

Firehose performs in-flight format conversion to Parquet using an AWS Glue Data Catalog schema, so no separate ETL job or compute cluster is needed. This satisfies the cost-effectiveness constraint by eliminating additional processing infrastructure while landing data directly in S3.

Why this answer

Amazon Kinesis Data Firehose supports record format conversion natively, allowing a delivery stream to convert incoming JSON records to Parquet or ORC using a schema stored in the AWS Glue Data Catalog. This is the most cost-effective approach because the conversion happens within Firehose without provisioning or paying for separate compute such as Lambda or Glue ETL jobs.

Exam trap

The trap is overlooking Firehose's built-in record format conversion and instead choosing Lambda or Glue ETL, which are more expensive and complex; candidates must know that Firehose can convert JSON to Parquet natively using a Glue schema.

How to eliminate wrong answers

Option A is wrong because Athena is a query service, not a transformation pipeline; using it to read JSON and write Parquet would require additional orchestration and cost, and it is not a Firehose-integrated conversion method. Option C is wrong because a Lambda function would need to parse and convert each record to Parquet, which is inefficient, adds latency, and incurs Lambda costs; Firehose already provides this capability. Option D is wrong because a scheduled Glue ETL job would run periodically and incur Glue DPU costs, and it would not convert data before landing in S3 as required; it would process data after it is already in S3.

25
MCQhard

Refer to the exhibit. A data engineer is using a Kinesis Data Stream with 2 shards. The producer uses a partition key that is the user ID (a UUID). The consumer is falling behind. Which change would improve throughput?

A.Switch to Kinesis Data Firehose
B.Increase the number of shards
C.Increase the retention period
D.Change the partition key to a constant value
AnswerB

A stream's throughput ceiling is set by its shard count, and two shards cap parallel consumption. Adding shards increases that ceiling and spreads the UUID-keyed records across more consumers, letting the lagging consumer catch up.

Why this answer

The consumer is falling behind because the total throughput of the stream (1 MB/s or 1,000 records/s per shard for writes, and 2 MB/s per shard for reads) is insufficient for the incoming data volume. Increasing the number of shards scales both the write and read capacity linearly, allowing the consumer to process records faster and catch up. Changing the partition key or retention period does not increase throughput, and switching to Firehose changes the delivery model but does not inherently solve the consumer lag.

Exam trap

The trap here is that candidates may think changing the partition key to a constant value would simplify processing, but it actually destroys parallelism and reduces throughput to a single shard, making the lag worse.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Firehose is a fully managed delivery service that buffers and loads data into destinations like S3 or Redshift; it does not increase the read throughput for a consumer that is falling behind, and it removes the ability for custom consumers to process records in real time. Option C is wrong because increasing the retention period (default 24 hours, max 365 days) only keeps records longer in the stream; it does not increase the ingestion or consumption rate, so the consumer will still lag. Option D is wrong because changing the partition key to a constant value would cause all records to go to a single shard, drastically reducing throughput and making the lag worse, as the other shard would be idle.

26
Multi-Selecthard

A company needs to ingest data from a MySQL database into Amazon S3 using AWS DMS. The data changes frequently and the requirement is to capture changes in near real-time. Which THREE configurations are necessary?

Select 3 answers
A.Create a VPC endpoint for S3.
B.Create an S3 target endpoint in DMS.
C.Enable binary logging (binlog) on the MySQL source database.
D.Create an AWS DMS replication instance.
E.Configure an S3 event notification to trigger DMS.
AnswersB, C, D

DMS requires a target endpoint defining Amazon S3 as the destination, including bucket, folder and IAM role, before any task can write replicated data. Without it, the replication instance has nowhere to deliver the ingested MySQL changes.

Why this answer

Option B is correct because AWS DMS requires an explicitly defined target endpoint that describes the S3 bucket, folder, and IAM role used to write the migrated data. Option C is correct because ongoing change data capture (CDC) from MySQL relies on the source's binary log (binlog), which must be enabled with settings such as log_bin and binlog_format=ROW so DMS can read committed changes in near real-time. Option D is correct because a replication instance is the managed compute resource that runs the DMS migration and CDC tasks between the source and target endpoints.

Option A is not required, since a VPC endpoint for S3 is only an optional networking choice and not a mandatory DMS configuration. Option E is not required, because DMS tasks are started and monitored by DMS itself, not by S3 event notifications.

27
MCQhard

A data pipeline uses Amazon Kinesis Data Firehose to deliver data to an S3 bucket. The delivery stream is configured with a buffer interval of 60 seconds and a buffer size of 5 MB. The data arrives at an average rate of 2 MB per second. What is the expected time interval between S3 writes?

A.Approximately 2.5 seconds
B.Approximately 30 seconds
C.Approximately 60 seconds
D.Approximately 10 seconds
AnswerA

Firehose flushes when either the 5 MB buffer fills or 60 seconds elapses. At 2 MB/s the size threshold is reached first, after roughly 2.5 seconds, so writes occur far more frequently than the configured interval.

Why this answer

Amazon Kinesis Data Firehose writes to S3 when either the buffer interval (60 seconds) or buffer size (5 MB) is reached first. With data arriving at 2 MB/s, the 5 MB buffer fills in 2.5 seconds (5 MB / 2 MB/s), triggering a write before the 60-second interval expires. Thus, the expected time between S3 writes is approximately 2.5 seconds.

Exam trap

The trap here is that candidates assume the buffer interval (60 seconds) is the primary determinant of write frequency, ignoring that the buffer size threshold triggers writes much earlier when data arrival rates are high.

How to eliminate wrong answers

Option B is wrong because 30 seconds would imply a buffer fill rate of ~0.167 MB/s, which does not match the given 2 MB/s arrival rate. Option C is wrong because 60 seconds is the buffer interval, but the buffer size threshold is reached much sooner at 2.5 seconds, making the interval the active trigger only if data arrival is slower. Option D is wrong because 10 seconds would correspond to a buffer size of 20 MB (2 MB/s * 10 s), which is not the configured 5 MB buffer size.

28
MCQhard

A healthcare company is building a data pipeline to ingest electronic health records (EHR) from hospitals. The data is sent as JSON files via SFTP to an on-premises server. The company wants to move this data to AWS using AWS Transfer Family (SFTP) and then process it with AWS Glue. Data sovereignty regulations require that all data remain within the EU (Frankfurt) region. The pipeline must detect when a new file arrives and start the Glue job automatically. The engineer has set up an AWS Transfer Family server in Frankfurt, and files are uploaded to an S3 bucket in the same region. However, the Glue job is not triggering automatically. The engineer needs to implement automated triggering. What should the engineer do?

A.Configure AWS Step Functions to poll the S3 bucket every minute and start the Glue job if new files exist.
B.Configure Amazon CloudWatch Events to trigger the Glue job on a schedule that checks for new files.
C.Use Amazon Simple Queue Service (SQS) to queue file metadata and have a Lambda function poll the queue to start the Glue job.
D.Set up an S3 event notification on the bucket to invoke an AWS Lambda function that starts the Glue job.
AnswerD

S3 event notifications fire on ObjectCreated events, invoking a Lambda function that calls StartJobRun on the Glue job. This provides the event-driven triggering the pipeline lacks, and keeps processing within the Frankfurt region, satisfying the EU data sovereignty constraint.

Why this answer

S3 event notifications can be configured to invoke an AWS Lambda function when objects are created, and the Lambda function can call the Glue StartJobRun API to launch the job. This provides event-driven, near-real-time triggering without polling, and keeps all processing within the Frankfurt region to satisfy data sovereignty requirements.

Exam trap

DEA-C01 often tests the misconception that polling or scheduled checks are needed to trigger Glue jobs, when S3 event notifications with Lambda provide the native, event-driven solution.

How to eliminate wrong answers

Option A is wrong because polling S3 every minute with Step Functions is inefficient, adds latency, and incurs unnecessary cost compared to native event notifications. Option B is wrong because scheduled CloudWatch Events checks for new files rather than reacting to arrivals, introducing delay and wasted executions. Option C is wrong because SQS with Lambda polling adds unnecessary complexity — S3 event notifications can directly invoke Lambda without an intermediary queue for this use case.

29
MCQmedium

A company ingests streaming data into Amazon Kinesis Data Streams. Producers write records using the PutRecords API with explicit partition keys based on customer ID. A data engineer observes that a few shards are consistently at 100 percent write throughput while others are underutilized, causing throttling. Which action should the engineer take to distribute the load more evenly?

A.Change the partition key to include a random suffix or use a higher-cardinality key.
B.Enable enhanced fan-out on the stream to give consumers dedicated throughput.
C.Switch the producers to use the PutRecord API instead of PutRecords.
D.Increase the number of shards in the stream to match the number of partition keys.
AnswerA

Kinesis Data Streams maps records to shards by hashing the partition key. If a small number of customer IDs generate most of the traffic, those keys hash to a few shards and create hot spots. Adding a random suffix or using a higher-cardinality key spreads records across more shards, balancing write throughput and eliminating throttling on the hot shards.

Why this answer

Hot shards occur when partition keys are skewed, causing the hash function to map most traffic to a few shards. Including a random suffix or using a higher-cardinality key distributes records more evenly across all shards. Adding shards or changing the write API does not alter the key-to-shard mapping, and enhanced fan-out only affects consumer reads, not producer write distribution.

Exam trap

The trap here is assuming that adding shards or enabling enhanced fan-out will fix write throttling, when the actual cause is a skewed partition key that keeps sending traffic to the same shards.

30
MCQeasy

A company needs to ingest data from multiple on-premises databases into Amazon S3 for analytics. The databases include Oracle, MySQL, and PostgreSQL. The data must be continuously replicated with minimal latency. Which AWS service should be used?

A.AWS Database Migration Service (AWS DMS)
B.Amazon Kinesis Data Streams
C.AWS Snowball
D.AWS Glue
AnswerA

DMS can continuously replicate from multiple source databases to S3.

Why this answer

AWS DMS supports continuous replication (change data capture, CDC) from Oracle, MySQL, and PostgreSQL to S3 as a target, enabling near-real-time data ingestion with minimal latency. It handles schema conversion and can replicate ongoing changes without interrupting source databases, making it the correct choice for this use case.

Exam trap

The trap here is that candidates may confuse Kinesis Data Streams as a general-purpose ingestion service for databases, but it lacks native CDC connectors for relational databases and is optimized for streaming data from applications, not for replicating transactional changes from databases to S3.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Streams is a real-time streaming service for ingesting high-throughput data from applications or devices, not designed for continuous database replication with CDC from relational databases. Option C is wrong because AWS Snowball is a physical data transfer device for offline, bulk data migration, not suitable for continuous, low-latency replication. Option D is wrong because AWS Glue is a serverless ETL service primarily for batch data transformation and cataloging, not for continuous replication with minimal latency from live databases.

31
MCQmedium

A data engineer has an AWS Glue ETL job that processes JSON files from Amazon S3. The job currently uses the DynamicFrame method to write output to Amazon Redshift. The engineer needs to improve write performance by using a staging Amazon S3 bucket and parallel COPY operations. Which AWS Glue connection option should the engineer configure?

A.Increase the number of Glue DPUs allocated to the job to enable parallel writes.
B.Set the 'useConnectionProperties' flag to true in the Glue job script.
C.Enable the 'redshiftTmpDir' connection property to specify an Amazon S3 temporary directory.
D.Configure the job to write to Amazon Redshift Spectrum using an external schema.
AnswerC

Setting the redshiftTmpDir connection property directs AWS Glue to stage data in Amazon S3 before issuing a COPY command into Amazon Redshift. This enables parallel loading and significantly improves write performance for large datasets. Without this property, Glue may fall back to slower JDBC-based inserts, which are not optimized for bulk transfers.

Why this answer

The redshiftTmpDir connection property tells AWS Glue to stage data in Amazon S3 and then run a COPY command into Redshift, which supports parallel loading and is far faster than JDBC inserts. The other options either reference non-existent properties or confuse read-side services like Redshift Spectrum with write-side optimizations.

Exam trap

The trap here is assuming that adding more DPUs or enabling a generic flag will automatically optimize Redshift writes, when the key is the staging directory property.

32
MCQhard

Refer to the exhibit. A data engineer runs a Glue job manually and receives a ThrottlingException. The engineer checks the job run history and sees a previous failure with the same error. What is the MOST likely cause of the throttling, and which solution is MOST appropriate?

A.Increase the number of DPUs for the job to reduce runtime.
B.Implement retry logic with exponential backoff in the script that calls start-job-run.
C.Use AWS Glue reserved capacity to guarantee API throughput.
D.Delete old job runs to reduce the number of entries in the job run history.
AnswerB

ThrottlingException arises when start-job-run calls exceed the Glue API request rate. Retry logic with exponential backoff spaces repeated calls progressively, allowing the request rate to fall within limits and the job to start successfully without manual intervention.

Why this answer

A ThrottlingException from AWS Glue typically occurs when the StartJobRun API is called too frequently, exceeding the service's request rate limits. Implementing retry logic with exponential backoff in the calling script is the standard AWS-recommended solution to handle transient throttling gracefully. This allows the job to start successfully after a short delay without manual intervention.

Exam trap

DEA-C01 often tests the confusion between performance tuning (DPUs) and API throttling, leading candidates to choose DPU increases when the real issue is request rate limiting.

How to eliminate wrong answers

Option A is wrong because increasing DPUs addresses job performance and runtime, not API throttling on StartJobRun. Option C is wrong because AWS Glue does not offer reserved capacity to guarantee API throughput; Glue is serverless and throttling is managed by AWS service quotas. Option D is wrong because deleting old job runs reduces history storage but does not affect the API request rate that causes throttling.

33
MCQeasy

A data engineer needs to schedule an AWS Glue extract, transform, and load job to run every day at 02:00 UTC and trigger a dependent Amazon Redshift stored procedure only after the Glue job succeeds. The engineer wants a managed orchestration option that avoids provisioning servers. Which approach should the engineer use?

A.Run an Amazon EC2 instance with a cron entry that calls the Glue StartJobRun API and then the Redshift API
B.Configure an Amazon EventBridge schedule rule that invokes the Glue job and rely on the job to call Redshift on completion
C.Use AWS Step Functions Express Workflows with a cron expression to call Glue and then Redshift
D.Create an AWS Glue workflow with a daily trigger that starts the job and a conditional trigger that fires the Redshift step on success
AnswerD

AWS Glue workflows provide serverless orchestration with schedule triggers for the daily 02:00 UTC start and conditional triggers that fire downstream actions only when the job reaches a SUCCEEDED state. This satisfies the dependency requirement without managing any servers, and it keeps orchestration within the Glue service already used for the job.

Why this answer

Glue workflows are serverless and purpose-built for chaining Glue jobs and dependent actions. A schedule trigger handles the daily 02:00 UTC start, and a conditional trigger keyed to the job's success state ensures the Redshift procedure runs only after the job completes successfully, with no servers to manage.

Exam trap

The trap here is assuming any scheduler that can start a Glue job also enforces success-based dependencies, when only a conditional trigger provides that guarantee natively.

34
MCQmedium

A company is using Amazon Kinesis Data Streams with a Lambda consumer to process clickstream data. The data rate is high and the Lambda function is falling behind, resulting in increased processing latency. What is the MOST effective way to improve throughput?

A.Increase the memory allocated to the Lambda function.
B.Increase the Lambda function timeout.
C.Use Kinesis Data Firehose instead of Lambda.
D.Increase the number of shards in the Kinesis stream.
AnswerD

Adding shards raises the stream's parallel capacity, since each shard supports one Lambda invocation per batch and caps ingestion at 1 MB/s or 1,000 records/s. With the Lambda consumer throttled by shard-level concurrency, horizontal scaling of the stream directly relieves the backlog causing the latency.

Why this answer

Increasing the number of shards in the Kinesis stream increases the stream's total read capacity, allowing more concurrent Lambda invocations to process records in parallel. Since each shard supports up to 5 read transactions per second and a maximum of 2 MB/s read throughput, adding shards directly raises the aggregate throughput, enabling the Lambda consumer to keep up with the high data rate.

Exam trap

The DEA-C01 exam often tests the misconception that Lambda performance tuning (memory/timeout) is the primary solution for stream processing backpressure, when in fact the shard count is the fundamental parallelism bottleneck in Kinesis Data Streams with a Lambda consumer.

How to eliminate wrong answers

Option A is wrong because increasing Lambda memory also increases CPU allocation, which can speed up individual function execution, but it does not address the bottleneck of limited shard-level parallelism; the function is falling behind due to insufficient concurrent processing capacity, not per-invocation performance. Option B is wrong because increasing the Lambda timeout only allows the function to run longer before being terminated, but it does not improve throughput; if the function is already timing out, extending the timeout may mask the issue but does not increase the rate at which records are consumed. Option C is wrong because Kinesis Data Firehose is a delivery stream that buffers and loads data to destinations like S3 or Redshift; it does not support real-time per-record processing with custom logic like Lambda, and switching to Firehose would lose the ability to transform or react to each record individually, which is likely required for clickstream processing.

35
Multi-Selecteasy

A company is using AWS Glue to catalog data in Amazon S3. The data is stored in CSV format, but the schema is not consistent across all files. Which TWO actions can the company take to handle schema evolution and ensure the Glue Data Catalog is up to date? (Choose TWO.)

Select 2 answers
A.Configure the Glue crawler to update the table schema on each run.
B.Manually update the Glue Data Catalog tables whenever the schema changes.
C.Disable schema update in the crawler and add partitions manually.
D.Schedule the Glue crawler to run periodically to detect changes.
E.Require all data producers to use a single fixed schema.
AnswersA, D

Configuring the crawler to update the table schema on each run lets it revise column definitions in the Glue Data Catalog when it encounters differing CSV structures, directly handling schema evolution rather than leaving stale definitions in place.

Why this answer

Option A is correct because configuring the Glue crawler to update the table schema on each run allows it to detect new columns or changed data types in the CSV files and automatically revise the Data Catalog table definition, which is essential when schemas vary across files. Option D is correct because scheduling the crawler to run periodically ensures that any schema changes introduced by new or modified files are detected and reflected in the Glue Data Catalog in a timely, automated manner. Together, these two actions provide automated, recurring schema evolution handling.

Option B is not ideal because manual updates are error-prone and do not scale, and the question asks for actions the company can take to handle schema evolution automatically. Option C is incorrect because disabling schema update prevents the crawler from adapting to schema changes, and manual partition addition does not address evolving column structures. Option E is incorrect because enforcing a single fixed schema contradicts the scenario where schemas are already inconsistent and does not help update the catalog for existing variation.

Exam trap

DEA-C01 often tests the features of Glue crawlers for schema evolution, and candidates may overlook the need for both enabling schema update and scheduling the crawler.

36
Multi-Selecthard

A company uses AWS DMS to continuously replicate data from an on-premises SQL Server to Amazon Aurora MySQL. The replication lag is increasing. Which THREE actions can reduce the lag? (Choose three.)

Select 3 answers
A.Use parallel apply on the target endpoint.
B.Filter out unnecessary tables from replication.
C.Enable DMS validation.
D.Enable Multi-AZ for the DMS replication instance.
E.Increase the DMS replication instance size.
AnswersA, B, E

Parallel apply speeds up writes on the target.

37
MCQeasy

A data engineer needs to ingest JSON data from an on-premises relational database into Amazon S3 every hour. Which AWS service should be used to set up a scheduled, incremental data transfer?

A.Amazon S3 Transfer Acceleration with a cron job.
B.AWS Database Migration Service (DMS) with S3 as target.
C.AWS Glue with a JDBC connection and a scheduled crawler.
D.Amazon Kinesis Data Firehose with a database source.
AnswerB

AWS DMS performs ongoing, incremental replication from on-premises relational databases to Amazon S3, tracking changes via CDC rather than full reloads. Scheduled hourly tasks satisfy the stem's requirement for recurring incremental transfer, which SCT or DataSync alone cannot provide from a database.

Why this answer

AWS DMS is purpose-built for migrating databases to AWS targets, including Amazon S3. It supports ongoing replication (change data capture) and scheduled full-load tasks, making it ideal for hourly incremental transfers from an on-premises relational database to S3 without custom scripting.

Exam trap

The trap here is that candidates confuse AWS Glue's ETL capabilities with DMS's managed database migration, assuming Glue's JDBC connections can handle incremental transfers, but Glue lacks built-in change data capture and requires custom logic for scheduled incremental loads.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration only speeds up uploads over long distances via edge locations; it does not provide scheduling, incremental data capture, or database connectivity. Option C is wrong because AWS Glue crawlers are designed for schema discovery and metadata cataloging, not for scheduled incremental data transfer from a database to S3; Glue ETL jobs can do this but require custom code, whereas DMS is the managed service for database migration. Option D is wrong because Kinesis Data Firehose ingests streaming data from producers like Kinesis streams or direct PUT, not from a relational database via JDBC; it lacks built-in change data capture for incremental database loads.

38
Multi-Selecteasy

A company uses Kinesis Data Firehose to deliver streaming data to S3. They need to transform the data by adding a timestamp and removing sensitive fields. Which TWO approaches can achieve this?

Select 2 answers
A.Use Kinesis Data Analytics to transform the stream
B.Use S3 Select to transform data at rest
C.Use AWS Glue ETL to process data after delivery to S3
D.Use Amazon Redshift Spectrum to transform data
E.Configure a Lambda function as a data transformation in Firehose
AnswersC, E

Running AWS Glue ETL over the delivered S3 objects lets you add the timestamp and drop sensitive fields after Firehose lands the data. This satisfies the transformation requirement using a managed, serverless job rather than altering the delivery stream itself.

Why this answer

Option E is correct because Kinesis Data Firehose natively supports Lambda-based data transformation: you attach a Lambda function to the delivery stream, and Firehose invokes it synchronously on each record (or batch) before delivery, allowing the function to add a timestamp and strip sensitive fields. Option C is correct because AWS Glue ETL can process the data after it lands in S3, using Spark-based jobs to enrich records with timestamps and drop sensitive columns, which is a valid post-delivery transformation approach. Option A is not correct because Kinesis Data Analytics is for real-time SQL/Flink analytics on streams, not for modifying records delivered by Firehose.

Option B is not correct because S3 Select only filters and projects data at rest using SQL on individual objects; it cannot add timestamps or rewrite/remove fields. Option D is not correct because Redshift Spectrum queries data in S3 for analytics and does not transform or modify the delivered objects.

39
MCQeasy

A company wants to ingest real-time clickstream data from a website into Amazon S3 with minimal code. The data should be delivered within 60 seconds of generation. Which AWS service should be used?

A.Amazon Kinesis Data Firehose
B.AWS Database Migration Service (DMS)
C.Amazon Kinesis Data Streams
D.Amazon S3 Transfer Acceleration
AnswerA

Amazon Kinesis Data Firehose buffers incoming records and delivers them to Amazon S3 automatically, with a configurable buffer interval as low as 60 seconds, satisfying the near-real-time constraint. It requires no custom consumer code, meeting the minimal-code requirement, unlike Kinesis Data Streams, which needs a separate delivery application.

Why this answer

Amazon Kinesis Data Firehose is a fully managed service that can ingest real-time streaming data and deliver it to Amazon S3 with minimal code. It can buffer data and deliver within 60 seconds, making it ideal for this use case. Firehose requires no custom code for delivery and can scale automatically.

Exam trap

The trap is confusing Kinesis Data Streams with Firehose; candidates may think Data Streams can directly deliver to S3, but it requires custom code, whereas Firehose is the managed solution for minimal code.

How to eliminate wrong answers

Option B is wrong because AWS DMS is designed for database migration, not for ingesting clickstream data into S3. Option C is wrong because Kinesis Data Streams requires custom code to read from the stream and write to S3; it does not natively deliver to S3. Option D is wrong because S3 Transfer Acceleration speeds up uploads to S3 over long distances but does not provide a streaming ingestion pipeline for real-time data.

40
Multi-Selecthard

A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The Flink application reads from a Kinesis Data Streams source, performs aggregations, and writes results to Amazon S3. The application is experiencing high checkpoint failures, and the processing lag is increasing. The data volume is 50 MB/s with an average record size of 1 KB. Which TWO actions would improve checkpoint reliability and reduce lag? (Choose TWO.)

Select 2 answers
A.Decrease the checkpoint interval to complete checkpoints faster.
B.Replace the S3 sink with Kinesis Data Firehose.
C.Decrease the parallelism of the Flink application.
D.Increase the checkpoint interval in the Flink configuration.
E.Increase the number of Kinesis Processing Units (KPUs) for the application.
AnswersD, E

Less frequent checkpoints reduce overhead.

Why this answer

Increasing the checkpoint interval (Option D) reduces the frequency of checkpoint operations, which decreases the overhead on the Flink application and allows it to dedicate more resources to processing data, thereby reducing lag. This is especially effective when checkpoint failures are caused by the system being unable to complete checkpoints within the current interval due to high throughput (50 MB/s).

Exam trap

The trap here is that candidates often think decreasing the checkpoint interval will speed up checkpoints, but in reality, it increases overhead and failure rates, while increasing parallelism (Option C) seems intuitive but actually reduces per-task resources and can worsen backpressure.

41
Multi-Selectmedium

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job must be able to handle a large number of small files efficiently and minimize the number of output files to improve downstream query performance. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Use the AWS Glue groupFiles option to group multiple small input files into a single partition for processing.
B.Enable the job bookmark feature to track previously processed files and avoid reprocessing.
C.Increase the number of DPUs allocated to the job to allow more parallel processing of the small files.
D.Use the AWS Glue Data Catalog to store metadata about the small files and let the job read from the catalog.
E.Set the job's output to use the coalesce transformation to reduce the number of partitions before writing.
AnswersA, E

The groupFiles option in AWS Glue allows the job to combine multiple small input files into a single partition based on a target size. This reduces the overhead of opening and processing many small files, improving read efficiency. It directly addresses the large number of small files problem by consolidating them during the read phase, which is a recommended practice for Glue ETL jobs.

Why this answer

To handle many small input files, AWS Glue provides the groupFiles option, which combines small files into larger partitions for processing. To minimize output files, the coalesce transformation reduces the number of partitions before writing, resulting in fewer, larger files. Together, these actions improve both read efficiency and downstream query performance.

Increasing DPUs, enabling bookmarks, or using the Data Catalog do not directly address these file-count issues.

Exam trap

The trap here is thinking that adding more DPUs will solve the small file problem, when in fact it can increase the number of output files and does not consolidate them.

42
MCQhard

A financial services company processes real-time stock trade data. They use Amazon Kinesis Data Streams with a shard count of 5, each shard receiving about 500 records per second. The consumer application uses the Kinesis Client Library (KCL) with DynamoDB for checkpointing. Lately, some records are being processed multiple times. What is the most likely cause?

A.The consumer application is crashing and restarting, causing re-processing of records.
B.The Kinesis stream's iterator age is exceeding the retention period.
C.The DynamoDB table used for checkpointing is throttling write requests.
D.The record size exceeds the 1 MB API limit, causing retries.
AnswerA

Frequent crashes force the KCL to resume from the last DynamoDB checkpoint, replaying every record consumed after it — at-least-once delivery guarantees duplicates on restart. With five shards at 500 records per second, each restart re-processes a substantial backlog, matching the observed duplicate processing.

Why this answer

The Kinesis Client Library (KCL) uses DynamoDB to track checkpoint progress for each shard. If the consumer application crashes and restarts, the KCL will resume processing from the last committed checkpoint, which may be behind the actual processing point. This causes records that were already processed (but not yet checkpointed) to be re-processed, leading to duplicate processing.

Exam trap

The trap here is that candidates often confuse checkpoint throttling (Option C) with duplicate processing, but throttling would cause checkpoint failures and potential re-processing only if the application cannot recover, whereas the direct cause of duplicates is the gap between processing and checkpointing after a crash.

How to eliminate wrong answers

Option B is wrong because iterator age exceeding the retention period would cause data to expire and become unavailable, not cause duplicate processing. Option C is wrong because DynamoDB throttling on checkpoint writes would cause checkpoint failures and potential re-processing, but the question states checkpointing is occurring and the issue is duplicate processing, not checkpoint failures. Option D is wrong because the 1 MB API limit applies to the total payload per PutRecords request, not per record, and exceeding it would cause write failures or retries, not duplicate processing of already-successful records.

43
Multi-Selectmedium

A company needs to ingest streaming data from thousands of IoT devices. The data must be processed in real-time and stored in Amazon S3. Which TWO services should be used together?

Select 2 answers
A.Amazon Kinesis Data Streams
B.Amazon Kinesis Data Firehose
C.AWS Glue
D.Amazon Simple Queue Service (SQS)
E.AWS Direct Connect
AnswersA, B

Kinesis Data Streams ingests the high-volume IoT telemetry with low latency and durable, replicated storage, buffering thousands of device writes for downstream consumers. It satisfies the real-time processing requirement by feeding analytics or Lambda consumers before data lands in S3.

Why this answer

Amazon Kinesis Data Streams (A) is correct because it provides a highly scalable, real-time streaming ingestion layer that can continuously capture data from thousands of IoT devices with low latency and durable, ordered record storage across shards. Amazon Kinesis Data Firehose (B) is correct because it is the fully managed delivery service that can consume that streaming data and automatically batch, transform, and load it directly into Amazon S3 without writing custom consumer applications. Together they satisfy the requirement for real-time processing plus reliable storage in S3.

AWS Glue (C) is a serverless ETL and data catalog service, not a streaming ingestion or delivery mechanism, so it does not fit this pipeline. Amazon SQS (D) is a message queue for decoupling applications, not a real-time streaming service designed for high-throughput IoT telemetry or direct S3 delivery. AWS Direct Connect (E) is a dedicated network connection from on-premises to AWS, not a data streaming or processing service.

Exam trap

The trap here is that candidates often confuse Kinesis Data Firehose with Kinesis Data Streams, thinking only one is needed, but the question requires both: Data Streams for real-time ingestion from devices and Data Firehose for automated delivery to S3.

44
MCQmedium

A company is ingesting data from multiple sources into S3 using AWS Glue. The data engineer notices that the Glue job is failing with an OutOfMemory error. Which step should be taken to resolve this issue?

A.Reduce the volume of incoming data
B.Configure the job to use a larger memory setting
C.Use a smaller file size for input
D.Increase the number of DPUs allocated to the Glue job
AnswerD

Glue allocates memory per executor across its DPUs; OutOfMemory errors during ingestion indicate insufficient memory for the workload's partitions. Increasing DPU count adds executors and memory capacity, directly resolving the resource constraint causing the job failure.

Why this answer

AWS Glue jobs run on Apache Spark, which distributes data processing across multiple executors. An OutOfMemory error typically indicates that the data being processed exceeds the memory available to the executors. Increasing the number of DPUs (Data Processing Units) allocates more memory and compute resources to the job, allowing it to handle larger datasets without running out of memory.

Exam trap

The trap here is that candidates may think they can directly increase memory settings (Option B) or reduce data volume (Option A), but AWS Glue abstracts memory management through DPUs, and the correct approach is to increase DPU allocation to provide more resources.

How to eliminate wrong answers

Option A is wrong because reducing the volume of incoming data is not a scalable solution and may not be feasible; the job should be able to handle the required data volume. Option B is wrong because AWS Glue does not allow direct configuration of memory settings per executor; memory is tied to DPU allocation, and increasing DPUs is the correct way to increase total memory. Option C is wrong because using a smaller file size for input does not address the root cause of memory exhaustion; Glue can process many small files efficiently, but the issue is the total data volume or skewed partitions, not file size.

45
Multi-Selectmedium

A data engineering team is using AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration must have minimal downtime and needs to capture ongoing changes after the full load. Which THREE resources are required for this task? (Choose three.)

Select 3 answers
A.A DMS source endpoint configured for Oracle.
B.An AWS DMS replication instance.
C.An AWS Snowball Edge device for initial data transfer.
D.An Amazon S3 bucket for staging the data.
E.A DMS target endpoint configured for Amazon RDS PostgreSQL.
AnswersA, B, E

Oracle is the migration source, so a DMS source endpoint must define connection details and credentials for the Oracle database. This endpoint lets the replication instance read the full load and, with change data capture enabled, ongoing redo log changes, satisfying the minimal-downtime requirement.

Why this answer

Option A is correct because AWS DMS requires a source endpoint that defines the Oracle database connection details, including server name, port, SID/service name, and credentials, so the replication instance can read the full load data and ongoing redo/archive log changes. Option B is correct because the AWS DMS replication instance is the managed compute resource that actually runs the migration tasks, performs the full load, and applies change data capture (CDC) from Oracle to the target. Option E is correct because a target endpoint pointing to the Amazon RDS for PostgreSQL instance is mandatory so DMS knows where to write the migrated tables and replicated changes.

Option C is not needed because Snowball Edge is an offline physical transfer device, whereas DMS performs the migration over the network and supports ongoing replication. Option D is not required because DMS does not need an Amazon S3 staging bucket for an Oracle-to-RDS PostgreSQL migration; S3 is only used for certain source/target types such as S3 endpoints or specific native backup/restore workflows.

Exam trap

DEA-C01 often tests whether candidates confuse DMS's core three-component architecture (source endpoint, target endpoint, replication instance) with optional services like S3 staging or Snowball that are only relevant for specific migration patterns.

46
MCQmedium

A company uses AWS Glue ETL to process data from Amazon S3 and write results to Amazon Redshift. The job fails with a memory error when processing large files. Which action should the data engineer take to resolve this issue?

A.Reduce the number of partitions in the Glue job.
B.Increase the number of DPUs allocated to the Glue job.
C.Switch to a smaller instance type in the Glue job configuration.
D.Use S3 Select to filter columns before reading into Glue.
AnswerB

Increasing DPUs adds more executors and memory per worker, directly addressing the out-of-memory failure during large S3 file processing. Glue distributes partitions across these additional workers, so each handles a smaller share. This satisfies the stem's constraint: the job fails with a memory error, which horizontal scaling of compute resolves.

Why this answer

Increasing the number of DPUs (Data Processing Units) allocated to the AWS Glue job provides more memory and compute resources, which directly addresses the out-of-memory error when processing large files. Glue jobs run on Apache Spark, and insufficient DPUs can cause executors to run out of memory during shuffle or aggregation operations on large datasets.

Exam trap

The trap here is that candidates may confuse memory errors with I/O bottlenecks and incorrectly choose S3 Select (Option D) to reduce data volume, when the real issue is insufficient compute memory for Spark transformations.

How to eliminate wrong answers

Option A is wrong because reducing the number of partitions would increase the data size per partition, worsening memory pressure and likely causing the same or a more severe memory error. Option C is wrong because switching to a smaller instance type would reduce available memory per executor, directly contradicting the need to resolve a memory error. Option D is wrong because S3 Select can reduce the amount of data read from S3, but it does not increase the memory available to the Glue job's Spark executors; the memory error occurs during processing, not during data ingestion.

47
MCQhard

A data engineer is setting up an Amazon Kinesis Data Analytics application to process streaming data from a Kinesis data stream named "input-stream". The application uses a reference data source from an S3 bucket. The engineer has attached the IAM policy shown in the exhibit to the application's IAM role. When starting the application, the engineer receives an 'AccessDeniedException' error. Which additional permission is required?

A.kinesis:PutRecord on the input stream
B.s3:GetObject on the S3 bucket containing the reference data
C.kinesis:CreateStream on the input stream
D.kinesis:PutRecords on the input stream
AnswerB

Reading reference data from S3 requires object-level read access, which the attached policy omits. Granting s3:GetObject on the bucket holding the reference data supplies the missing action, resolving the AccessDeniedException thrown when Kinesis Data Analytics loads the reference source at application start.

Why this answer

The Kinesis Data Analytics application needs to read reference data from the S3 bucket, which requires the s3:GetObject permission on the bucket and its objects. The error 'AccessDeniedException' indicates the IAM role lacks this specific permission to retrieve the reference data file. Option B correctly adds the missing s3:GetObject action to allow the application to fetch the reference data from S3.

Exam trap

The trap here is that candidates often confuse the direction of data flow and assume the application needs write permissions (PutRecord/PutRecords) to the input stream, when in fact it only needs read permissions (kinesis:DescribeStream, kinesis:GetShardIterator, kinesis:GetRecords) and the missing permission is for the separate S3 reference data source.

How to eliminate wrong answers

Option A is wrong because kinesis:PutRecord is used to write data to a Kinesis stream, but the application reads from the input stream as a source, not writes to it; the error is not about writing. Option C is wrong because kinesis:CreateStream is an administrative action to create a new stream, which is irrelevant to an existing stream used as input. Option D is wrong because kinesis:PutRecords is for batch writing to a stream, not for reading or for accessing reference data from S3.

48
MCQmedium

A data engineer needs to run an AWS Glue for Apache Spark ETL job that joins a 40 GB Amazon S3 Parquet dataset with a small 8 MB reference lookup table stored as CSV in Amazon S3. The reference table is read on every join and the job's executors are spending a large amount of shuffle time on the join. The reference table changes only once per month. Which approach MOST efficiently reduces shuffle overhead in the Glue job?

A.Partition both datasets by the join key and write them to S3 before the join so Spark can perform a partition-wise merge join.
B.Increase the number of Glue DPUs so that more executors are available to parallelize the shuffle stage of the join.
C.Broadcast the reference table by reading it with the Spark DataFrame API and applying broadcast() before the join.
D.Convert the reference CSV to Parquet and enable Glue job bookmarks so the lookup is only read once per month.
AnswerC

Broadcasting the small reference DataFrame replicates it to each executor, so the large Parquet dataset never needs to be shuffled. This eliminates the expensive shuffle of the 40 GB dataset that the join currently causes. Because the table is only 8 MB and changes monthly, broadcasting is safe and fits comfortably in executor memory, making it the most efficient fix for shuffle overhead in the Glue job.

Why this answer

The small, slowly changing reference table is an ideal broadcast candidate. Replicating it to every executor lets Spark perform a broadcast hash join where the large Parquet side stays in place and is never redistributed. Increasing DPUs, converting formats, or pre-partitioning all leave the shuffle of the large dataset intact, so they fail to address the actual bottleneck.

Exam trap

The trap here is assuming that adding compute resources or changing file formats eliminates shuffle cost, when only broadcasting the small side removes the redistribution of the large dataset.

49
Multi-Selectmedium

A company uses AWS Glue to run ETL jobs daily. The jobs consume data from an Amazon RDS for MySQL database and write results to Amazon S3. The company wants to minimize the impact on the source database during extraction. Which THREE actions should the data engineer take to achieve this? (Choose THREE.)

Select 3 answers
A.Schedule the Glue job to run during off-peak hours.
B.Configure the Glue job to connect to a read replica of the RDS instance.
C.Increase the number of Glue DPUs to process data faster.
D.Disable Glue job bookmarks to force full refresh.
E.Use a JDBC connection with a WHERE clause to extract only incremental data.
AnswersA, B, E

Runs when database load is naturally low.

Why this answer

Scheduling the Glue job to run during off-peak hours minimizes the load on the source RDS for MySQL database by avoiding high-traffic periods, reducing contention for CPU, memory, and I/O resources. This is a straightforward operational practice to reduce impact on production databases during extraction.

Exam trap

The trap here is that candidates often assume increasing DPUs (parallelism) always improves performance without realizing it can amplify the load on the source database, and they may overlook that disabling bookmarks forces full refreshes, which is the opposite of minimizing impact.

50
MCQmedium

A company is using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data must be transformed from JSON to Parquet format before landing in S3. The transformation logic is simple: convert the JSON schema to Parquet. Which approach meets the requirements with the least operational overhead?

A.Use the built-in data format conversion feature of Firehose with an AWS Glue Data Catalog table
B.Use an AWS Lambda function to transform records to Parquet before sending to Firehose
C.Use Amazon Kinesis Data Analytics to convert the stream to Parquet
D.Provision an Amazon EMR cluster to convert the data in micro-batches
AnswerA

Firehose's built-in data format conversion uses an AWS Glue Data Catalog table as the schema reference to transform JSON records into Parquet before delivery. This is serverless and requires no custom code, minimising operational overhead.

Why this answer

Amazon Kinesis Data Firehose provides a built-in data format conversion feature that can automatically convert incoming JSON data to Parquet format using an AWS Glue Data Catalog table as the schema reference. This approach requires no custom code, no additional infrastructure, and no manual transformation logic, making it the simplest solution with the least operational overhead for a straightforward JSON-to-Parquet conversion.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing Lambda or EMR, not realizing that Firehose's built-in format conversion with Glue Data Catalog is the simplest, fully managed option for JSON-to-Parquet conversion without any custom code.

How to eliminate wrong answers

Option B is wrong because using an AWS Lambda function to transform records to Parquet before sending to Firehose introduces unnecessary complexity, additional cost, and operational overhead (e.g., managing Lambda concurrency, packaging Parquet libraries, handling record size limits), whereas Firehose's built-in conversion handles this natively. Option C is wrong because Amazon Kinesis Data Analytics is designed for real-time analytics and stream processing using SQL or Flink, not for simple format conversion; it adds latency, complexity, and cost without benefit for a straightforward schema conversion. Option D is wrong because provisioning an Amazon EMR cluster to convert data in micro-batches is a heavy, over-engineered solution that requires cluster management, scaling, and job orchestration, far exceeding the operational overhead needed for a simple format conversion that Firehose can perform automatically.

51
MCQeasy

Refer to the exhibit. A Lambda function named 'IngestionProcessor' is failing. The engineer checks CloudWatch Logs and sees the log group exists but storedBytes is 0. Why might the logs show no data?

A.The Lambda execution role does not have permission to write logs to CloudWatch
B.The Lambda function is configured with a dead letter queue
C.The Lambda function has not been invoked yet
D.The log group is encrypted with a KMS key and the Lambda function lacks decrypt permission
AnswerA

Without logs:write permission on its execution role, Lambda cannot create log streams or put events, so the log group exists but storedBytes stays 0. The stem's constraint is an empty log group despite invocations, which points to missing CloudWatch Logs write authorisation rather than retention or filtering.

Why this answer

The Lambda execution role must have the `logs:CreateLogStream` and `logs:PutLogEvents` permissions to write logs to CloudWatch Logs. If the role lacks these permissions, the log group will be created (if it doesn't exist) but no log events will be written, resulting in `storedBytes` being 0. This is a common misconfiguration when the IAM policy does not include the necessary CloudWatch Logs actions.

Exam trap

The DEA-C01 exam often tests the distinction between log group creation (which requires `logs:CreateLogGroup`) and log writing (which requires `logs:CreateLogStream` and `logs:PutLogEvents`), leading candidates to confuse the existence of a log group with successful log delivery.

How to eliminate wrong answers

Option B is wrong because a dead letter queue (DLQ) is used to capture failed events for asynchronous invocations, not to prevent logs from being written; it does not affect CloudWatch Logs permissions. Option C is wrong because if the Lambda function had not been invoked, the log group would not exist at all; the presence of the log group with `storedBytes` of 0 indicates the function was invoked but failed to write logs. Option D is wrong because if the log group were encrypted with a KMS key and the Lambda function lacked decrypt permission, the function would fail with an access denied error when trying to write logs, but the log group would still show `storedBytes` as 0; however, the question states the log group exists and `storedBytes` is 0, which is consistent with missing write permissions, not KMS decryption issues (KMS errors would typically produce a different error message in CloudWatch).

52
MCQeasy

A company is using Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data is delivered in 5-minute intervals. The company wants to reduce the delivery frequency to 1 minute to get data faster. Which parameter should be changed in the Firehose delivery stream configuration?

A.Reduce the buffer interval from 300 seconds to 60 seconds.
B.Increase the buffer size to trigger delivery sooner.
C.Enable dynamic partitioning to deliver data more frequently.
D.Enable compression to reduce data size and speed up delivery.
AnswerA

Firehose buffers incoming records and delivers when either the buffer size or buffer interval is reached. Lowering the interval from 300 to 60 seconds forces delivery every minute, satisfying the requirement for faster one-minute delivery.

Why this answer

Kinesis Data Firehose buffers incoming records and delivers them when either the buffer size (e.g., 5 MB) or the buffer interval (e.g., 300 seconds) is reached, whichever comes first. Reducing the buffer interval from 300 seconds to 60 seconds makes Firehose flush every minute, directly achieving the 1-minute delivery frequency the company wants.

Exam trap

The trap is confusing buffer size with buffer interval; candidates often think increasing buffer size speeds delivery, but only the interval controls the time-based flush, and the exam tests whether you know the two thresholds and their defaults.

How to eliminate wrong answers

Option B is wrong because increasing the buffer size makes Firehose wait for more data before delivering, which delays delivery rather than accelerating it. Option C is wrong because dynamic partitioning changes how records are grouped into S3 prefixes based on partition keys; it does not change the buffering interval and therefore does not control delivery frequency. Option D is wrong because compression reduces the size of delivered data and can affect how quickly the buffer size threshold is reached, but it does not set a time-based delivery cadence and is not the parameter for frequency.

53
MCQhard

A data engineer is designing a streaming pipeline using Amazon Kinesis Data Streams with a shard count of 10. The incoming data rate is 1 MB/second. The consuming application uses the Kinesis Client Library (KCL) with a single worker. What is the most likely performance bottleneck?

A.The Lambda function invoked by the stream has a cold start issue
B.The data stream has insufficient write capacity
C.The single KCL worker cannot process all shards in parallel
D.The shard count is too low to handle the data rate
AnswerC

KCL workers should be scaled to match shard count for parallel processing.

Why this answer

The Kinesis Client Library (KCL) uses a 1:1 mapping between shards and record processors by default. With 10 shards and only a single KCL worker, that worker must run all 10 record processors sequentially on a single host, creating a bottleneck. The worker cannot process records from multiple shards in parallel, so the throughput is limited by the single worker's processing capacity, not the stream's write capacity.

Exam trap

The trap here is that candidates often assume the bottleneck is on the write side (insufficient shards or write capacity) because they focus on the incoming data rate, but the question specifically tests the consumer-side limitation of a single KCL worker unable to parallelize across multiple shards.

How to eliminate wrong answers

Option A is wrong because Lambda cold starts are a potential issue only if the consuming application uses Lambda as a consumer, but the question specifies a KCL worker, not a Lambda function. Option B is wrong because the incoming data rate is 1 MB/second, and a single Kinesis shard supports up to 1 MB/second write capacity, so 10 shards provide 10 MB/second—far more than needed. Option D is wrong because the shard count of 10 is more than sufficient to handle the 1 MB/second data rate; the bottleneck is on the consumer side, not the stream's capacity.

54
MCQeasy

A company wants to ingest real-time clickstream data from a website into Amazon S3 with a maximum latency of 60 seconds. The data volume peaks at 500 MB/s. Which service should they use to buffer and deliver the data to S3?

A.Amazon Kinesis Data Firehose
B.Amazon Simple Queue Service (SQS)
C.Amazon Kinesis Data Streams
D.AWS Lambda
AnswerA

Kinesis Data Firehose buffers streaming records and delivers them to Amazon S3 in batches, with configurable buffer intervals that keep latency well under 60 seconds. It scales to handle 500 MB/s peaks without custom consumers, unlike Kinesis Data Streams, which requires separate delivery code.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is designed to ingest streaming data, buffer it, and deliver it to destinations like Amazon S3 with configurable buffer intervals (e.g., 60 seconds) and buffer sizes (e.g., up to 128 MB). It can handle the peak throughput of 500 MB/s by automatically scaling, and it meets the maximum latency requirement of 60 seconds by flushing data to S3 based on time or size thresholds.

Exam trap

The trap here is that candidates often confuse Kinesis Data Streams (a real-time processing stream requiring custom consumers) with Kinesis Data Firehose (a fully managed delivery service), leading them to pick Data Streams for its real-time capabilities, even though Firehose is the correct choice for direct S3 delivery with minimal latency.

How to eliminate wrong answers

Option B (Amazon Simple Queue Service) is wrong because SQS is a message queue for decoupling application components, not a streaming buffer designed for high-throughput data delivery to S3; it lacks native integration to automatically write data to S3 with configurable latency. Option C (Amazon Kinesis Data Streams) is wrong because it is a real-time data streaming service that requires custom consumers (e.g., Lambda or Kinesis Client Library) to read and write data to S3, adding complexity and latency beyond the 60-second requirement; it does not natively buffer and deliver to S3. Option D (AWS Lambda) is wrong because Lambda is a serverless compute service for running code in response to events, not a buffer or delivery mechanism; it cannot handle sustained 500 MB/s ingestion without additional services and would require custom orchestration to meet latency goals.

55
MCQmedium

A data engineer needs to transform data in an AWS Glue job using a custom Python library that is not available by default. The library is packaged as a .whl file stored in Amazon S3. The engineer wants the Glue job to use this library without modifying the job script to install it at runtime. What should the engineer do?

A.Add the library to the Glue Data Catalog as a custom classifier.
B.Use a Glue development endpoint to install the library, then run the job there.
C.Place the .whl file in the Glue job's Python library path parameter (--extra-py-files).
D.Upload the .whl to the Glue job's script bucket and import it directly by filename.
AnswerC

The --extra-py-files job parameter allows specifying additional Python files or wheels stored in S3, which Glue makes available to the job without runtime installation. This directly satisfies the requirement to use a custom library without modifying the script to install it.

Why this answer

Specifying the wheel via the --extra-py-files job parameter makes the custom Python library available to the AWS Glue job without script changes, which is the correct way to include additional dependencies stored in Amazon S3.

Exam trap

The trap here is confusing catalog classifiers or dev endpoints with dependency injection, when only job parameters like --extra-py-files add libraries.

56
MCQhard

A company uses Amazon Kinesis Data Firehose to ingest log data from web servers into Amazon S3. The data is in JSON format and each record is approximately 2 KB. The delivery stream is configured to buffer incoming records for 60 seconds or 5 MB, whichever comes first. The company notices that the data in S3 is delayed by up to 5 minutes during peak hours. Which action would most effectively reduce the delivery latency?

A.Increase the buffer size to 10 MB to allow more records per delivery.
B.Decrease the buffer interval to 15 seconds.
C.Enable compression (GZIP) on the delivery stream.
D.Enable data transformation with AWS Lambda to convert JSON to Parquet.
AnswerB

Shorter buffer interval triggers more frequent deliveries, reducing latency.

Why this answer

The observed delay of up to 5 minutes during peak hours indicates that the buffer size threshold (5 MB) is rarely reached because each record is only ~2 KB, so the delivery stream relies on the buffer interval (60 seconds) to trigger delivery. By decreasing the buffer interval to 15 seconds, Kinesis Data Firehose will push data to S3 more frequently, directly reducing the maximum latency from 60 seconds to 15 seconds per batch, which eliminates the compounding delays caused by queuing during high-throughput periods.

Exam trap

The trap here is that candidates assume increasing buffer size or enabling compression will speed up delivery, but they fail to recognize that with small records, the buffer interval is the bottleneck, and only reducing that interval directly lowers latency.

How to eliminate wrong answers

Option A is wrong because increasing the buffer size to 10 MB would actually increase the time needed to fill the buffer, worsening the latency issue during peak hours when records are small and the buffer interval is the primary trigger. Option C is wrong because enabling GZIP compression reduces storage size and cost but does not affect the delivery frequency or buffer flush timing, so it has no impact on latency. Option D is wrong because converting JSON to Parquet via Lambda adds processing overhead and introduces additional latency from the transformation invocation, which would increase rather than reduce delivery delay.

57
MCQhard

A company uses AWS Glue to process data from multiple S3 buckets. The Glue job runs daily and reads data from a bucket that contains millions of small files (each < 1 MB). The job has been running for hours and is often close to the 8-hour timeout limit. Which optimization would MOST reduce the job's runtime?

A.Pre-process the data to consolidate small files into larger files before the Glue job.
B.Convert the source data from CSV to Parquet format.
C.Increase the number of DPUs allocated to the Glue job.
D.Use a larger Spark shuffle partition size.
AnswerA

Millions of sub-1 MB files force Glue to open each object individually, so per-file overhead dominates runtime. Consolidating them into larger files before the job drastically cuts the number of read operations and task overhead.

Why this answer

Consolidating millions of small files (<1 MB each) into larger files before the Glue job is the most impactful optimization because Spark and Glue incur massive per-file overhead: each file requires a separate S3 GET, task scheduling, and metadata operation. With millions of tiny files, the job spends most of its time on I/O and task overhead rather than actual processing. Merging them into a few hundred MB files dramatically reduces task count and runtime.

Exam trap

The trap is reaching for the 'more resources' answer (more DPUs) — candidates assume scaling compute fixes slow jobs, but the exam tests whether you recognize that small-file I/O overhead is the true bottleneck.

How to eliminate wrong answers

Option B is wrong because converting CSV to Parquet improves compression and scan efficiency but does not solve the small-file problem — millions of tiny Parquet files still cause the same per-file overhead. Option C is wrong because adding DPUs increases parallelism but cannot overcome the serialization bottleneck of millions of file-open operations; it also increases cost without addressing the root cause. Option D is wrong because a larger shuffle partition size affects post-shuffle aggregation, not the initial read of millions of small files, so it does not reduce the dominant cost.

58
MCQmedium

A data engineer is ingesting records from Amazon Kinesis Data Streams into Amazon S3 using AWS Lambda as the consumer. Each stream shard delivers up to 1,000 records per second, and the Lambda function writes each record as an individual small object, causing many tiny S3 files and high PUT costs. The engineer wants fewer, larger objects while keeping near-real-time delivery. What should the engineer do?

A.Replace the Lambda consumer with Amazon Data Firehose, which buffers records and delivers batched objects to S3.
B.Enable enhanced fan-out on the stream so consumers get dedicated throughput per shard.
C.Increase the Lambda function memory so each invocation can write more records per second.
D.Add more shards to the Kinesis data stream to increase parallelism.
AnswerA

Amazon Data Firehose reads from Kinesis Data Streams and buffers incoming records by size and time before delivering consolidated objects to S3. This directly produces fewer, larger files and cuts PUT request costs, while its configurable buffer interval preserves near-real-time delivery. It removes the need for custom batching logic in Lambda.

Why this answer

The inefficiency comes from writing one S3 object per Kinesis record. Amazon Data Firehose natively buffers records by configurable size and time windows and delivers them as consolidated objects, which reduces file count and PUT costs while still meeting near-real-time needs. Tuning Lambda memory, adding shards, or enabling enhanced fan-out changes throughput and parallelism but not the per-record write pattern.

Exam trap

The trap here is tuning Lambda or stream capacity to fix a file-size problem, when the actual fix is introducing a buffering delivery layer that batches records.

59
Multi-Selectmedium

A company uses Amazon Kinesis Data Firehose to ingest data into an S3 bucket. The data is in JSON format and the team wants to convert it to Parquet before storage. Which TWO configurations are required?

Select 2 answers
A.Use Kinesis Data Analytics to transform data to Parquet.
B.Create a Glue table with the schema of the data.
C.Configure a Lambda function to convert data on the fly.
D.Set up an Athena table to read the data.
E.Enable data format conversion in Firehose and set Output format to Parquet.
AnswersB, E

Firehose's Parquet conversion relies on the AWS Glue Data Catalog to resolve the source schema, so a Glue table describing the JSON structure is mandatory. Without it, Firehose cannot map incoming records to Parquet columns, and the conversion configuration fails.

Why this answer

Option B is correct because Firehose data format conversion relies on the AWS Glue Data Catalog: you must create a Glue table (with the appropriate schema and SerDe) that Firehose references so it knows how to interpret the incoming JSON records. Option E is correct because the actual conversion is enabled in the Firehose delivery stream configuration by turning on data format conversion and setting the output format to Parquet (with the Glue table as the schema source). Option A is not required because Kinesis Data Analytics is for SQL/Flink stream processing, not for Firehose's built-in format conversion.

Option C is not required because Firehose performs the JSON-to-Parquet conversion natively via Glue, so a custom Lambda transformation is unnecessary. Option D is not required because Athena is a query service for reading data in S3, not a prerequisite for converting it during ingestion.

Exam trap

DEA-C01 often tests whether candidates know that Firehose Parquet conversion is a native two-part configuration (Glue table + Firehose setting) rather than something requiring Lambda, Glue ETL, or Kinesis Data Analytics.

60
MCQhard

A company uses AWS Glue DataBrew for data preparation. The data source is an S3 bucket with millions of small CSV files (each < 1 MB). The DataBrew project takes a long time to load the sample data. What is the most likely cause and solution?

A.Use Amazon Athena to query the data instead of DataBrew
B.The DataBrew job is under-provisioned; increase the number of DPUs
C.The large number of small files causes S3 LIST overhead; concatenate files into larger files
D.Use AWS Glue ETL instead of DataBrew for this volume
AnswerC

DataBrew must LIST and open each object individually, so millions of sub-1 MB files impose heavy S3 LIST and per-request overhead during sampling. Concatenating them into larger files reduces request count and dramatically speeds up sampling.

Why this answer

DataBrew loads a sample of the data by listing objects in the S3 bucket. With millions of small CSV files, the S3 LIST API call becomes a bottleneck because each list operation has a 1000-object limit per response, requiring multiple paginated requests. Concatenating the small files into larger files reduces the number of objects, dramatically decreasing LIST overhead and speeding up sample loading.

Exam trap

The DEA-C01 exam often tests the misconception that increasing DPUs or switching to a different AWS service will fix performance issues, when the real root cause is S3's small-file overhead and the LIST API's pagination limit.

How to eliminate wrong answers

Option A is wrong because Athena is a query engine, not a data preparation tool; it would still suffer from the same small-file overhead when reading data, and it does not solve the DataBrew sample loading issue. Option B is wrong because DataBrew projects do not use DPUs for sample loading; DPUs are only relevant for running DataBrew jobs (recipes), and the bottleneck here is S3 LIST latency, not compute capacity. Option D is wrong because switching to Glue ETL would not inherently solve the small-file problem; Glue ETL also incurs overhead from listing and processing many small files, and the question specifically asks about DataBrew sample loading, not ETL job performance.

61
MCQeasy

A data engineer needs to run a transformation that processes semi-structured JSON records already stored in Amazon S3 and write the results back to S3 in Parquet format. The team prefers a serverless, Apache Spark-based approach with minimal infrastructure management and wants to use the AWS Glue Data Catalog for metadata. Which approach should the engineer use?

A.Run an AWS Glue crawler to convert the JSON files to Parquet automatically.
B.Use AWS Lambda with pandas to convert each JSON file to Parquet.
C.Launch an Amazon EMR cluster with Spark and submit the job manually.
D.Create an AWS Glue ETL job using the Spark engine and write output with the Glue Parquet writer.
AnswerD

AWS Glue ETL jobs run on a serverless Apache Spark environment, so the engineer avoids managing clusters while gaining Spark's transformation capabilities. Reading JSON from S3, transforming it, and writing Parquet through the Glue Parquet writer produces columnar output optimized for analytics, and the job integrates natively with the Glue Data Catalog for schema and table metadata.

Why this answer

AWS Glue ETL jobs provide a serverless Apache Spark runtime that reads JSON from Amazon S3, applies transformations, and writes Parquet using the Glue Parquet writer. This combination satisfies the serverless and Spark-based preferences while integrating with the AWS Glue Data Catalog for table metadata. EMR adds cluster management overhead, and Lambda or a crawler cannot perform the required rewrite.

Exam trap

The trap here is confusing the role of an AWS Glue crawler, which only catalogs schemas, with an AWS Glue ETL job, which actually transforms and rewrites data.

62
MCQhard

A data engineer runs the describe-stream command and sees the output above. The stream has a retention period of 24 hours. The engineer needs to ensure that consumers can replay data for up to 7 days. Which action is required?

A.Increase the number of shards to allow more data storage.
B.Delete the stream and recreate it with a longer retention period.
C.Use the IncreaseStreamRetentionPeriod API to set retention to 168 hours.
D.Create new consumer applications that read from the stream.
AnswerC

The API can increase retention up to 365 days.

Why this answer

The describe-stream output shows a retention period of 24 hours, but the requirement is to allow consumers to replay data for up to 7 days (168 hours). Amazon Kinesis Data Streams supports modifying the retention period dynamically without recreating the stream, using the IncreaseStreamRetentionPeriod API or the update-shard-count command. Option C correctly uses this API to set retention to 168 hours, which is the maximum supported retention period for Kinesis Data Streams.

Exam trap

The trap here is that candidates often confuse shard count with storage capacity, assuming that more shards allow more data to be stored, when in fact shards only control throughput and retention is a separate, configurable parameter.

How to eliminate wrong answers

Option A is wrong because increasing the number of shards increases the stream's throughput capacity (read/write operations per second), not the data retention period; shards do not affect how long data is stored. Option B is wrong because deleting and recreating the stream is unnecessary and disruptive; Kinesis allows you to modify the retention period on an existing stream without data loss or downtime. Option D is wrong because creating new consumer applications does not change the retention period; consumers can only replay data within the existing retention window, so they would still be limited to 24 hours of replay.

63
MCQeasy

A data engineer needs to ingest streaming data from a social media API into Amazon S3 for batch analytics. The data arrives at a rate of 500 records per second. Which service should be used to capture the stream?

A.Amazon Simple Notification Service (SNS)
B.Amazon Simple Queue Service (SQS)
C.Amazon Kinesis Data Streams
D.Amazon MQ
AnswerC

Kinesis Data Streams is designed for real-time streaming data ingestion.

Why this answer

Amazon Kinesis Data Streams is designed for real-time streaming data ingestion at scale, supporting throughput of up to 1 MB/s or 1,000 records per second per shard. With 500 records per second, Kinesis can reliably capture and store the social media API data for up to 365 days, enabling batch analytics via S3 delivery through Kinesis Firehose or custom consumers.

Exam trap

The trap here is that candidates confuse SQS's message queueing with Kinesis's stream processing, overlooking that SQS lacks ordered, replayable, and high-throughput streaming capabilities required for real-time data ingestion into S3.

How to eliminate wrong answers

Option A is wrong because Amazon SNS is a pub/sub messaging service for push notifications and fan-out, not designed for persistent, ordered streaming data ingestion or high-throughput record capture. Option B is wrong because Amazon SQS is a message queue for decoupling microservices with at-least-once delivery, but it lacks the shard-based parallelism, replay capability, and long-term retention needed for streaming data to S3. Option D is wrong because Amazon MQ is a managed message broker for ActiveMQ or RabbitMQ protocols, optimized for JMS and enterprise messaging, not for high-velocity stream ingestion or direct integration with S3 batch analytics.

64
MCQmedium

A company is streaming IoT data from thousands of devices into Amazon Kinesis Data Streams. The data must be transformed in real time before being stored in Amazon S3. Which service should be used to perform the transformation as the data streams through Kinesis?

A.AWS Glue
B.Amazon Kinesis Data Analytics for Apache Flink
C.Amazon EMR
D.AWS Lambda
AnswerB

Kinesis Data Analytics for Apache Flink runs continuous SQL or Flink applications directly against the stream, transforming records in flight before they land in Amazon S3. This satisfies the real-time transformation constraint, whereas AWS Glue and Lambda-based batch approaches operate after or outside the stream.

Why this answer

Amazon Kinesis Data Analytics for Apache Flink is the correct choice because it is purpose-built for running Apache Flink applications that can perform real-time transformations, filtering, and enrichment on data streaming through Kinesis Data Streams before outputting the results to destinations like Amazon S3. It integrates natively with Kinesis Data Streams as a source and can write transformed data directly to S3 using a Flink sink, making it ideal for this streaming ETL use case.

Exam trap

The trap here is that candidates often choose AWS Lambda because it is a familiar serverless option for event-driven processing, but they overlook its limitations in execution time, payload size, and lack of native state management for complex transformations, which makes Kinesis Data Analytics for Apache Flink the more robust and scalable choice for continuous streaming ETL.

How to eliminate wrong answers

Option A is wrong because AWS Glue is primarily a batch ETL service that processes data in job runs, not a real-time streaming transformation engine; while Glue Streaming exists, it is based on Spark Streaming and requires a separate Glue job with a streaming source, not a native Kinesis Data Streams integration for real-time transformations. Option C is wrong because Amazon EMR is a managed Hadoop/Spark cluster platform that can process streaming data but requires manual cluster management, provisioning, and configuration of Spark Streaming or Flink, adding operational overhead that is unnecessary for a simple transformation before S3 storage. Option D is wrong because AWS Lambda can process Kinesis Data Streams records in near real-time, but it has a maximum execution timeout of 15 minutes and a payload limit of 6 MB per invocation, making it unsuitable for high-throughput, continuous transformations of thousands of devices' data without risk of throttling or data loss.

65
Multi-Selecteasy

Which TWO AWS services can be used as sources for AWS Glue ETL jobs? (Choose two.)

Select 2 answers
A.Amazon Route 53
B.Amazon CloudFront
C.Amazon API Gateway
D.Amazon S3
E.Amazon RDS
AnswersD, E

S3 is a common source for Glue jobs.

Why this answer

Amazon S3 is a fully managed object storage service that serves as a common source for AWS Glue ETL jobs. Glue can read data from S3 using its built-in crawlers and connectors, supporting formats like Parquet, JSON, CSV, and Avro. The Glue Data Catalog can reference S3 locations, and ETL scripts can directly read from S3 buckets via the s3:// protocol.

Exam trap

The DEA-C01 exam often tests the misconception that any AWS service that stores or serves data (like Route 53 for DNS records or CloudFront for cached content) can be a Glue source, but Glue only supports sources that provide a direct data access interface (e.g., object storage, databases, or streaming services like Kinesis).

66
Matchingmedium

Match each AWS service to its primary purpose in data engineering.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Serverless ETL and data catalog

Data warehousing and SQL analytics

Big data processing using Hadoop/Spark

Building and managing data lakes

Real-time streaming data ingestion

Why these pairings

The correct matches are: Amazon S3 → object storage for data lakes (A), AWS Glue → serverless ETL (B), Amazon Athena → interactive SQL queries on S3 (C), and Amazon Redshift → data warehousing (D). Common confusions involve mixing up services like Kinesis (streaming) with Glue (ETL) or Data Pipeline (orchestration) with Kinesis (ingestion).

67
MCQeasy

A data engineer needs to run a one-time transformation on a 500 GB dataset stored in Amazon S3. The transformation is written in Python and uses pandas, which cannot handle the full dataset in memory. The engineer wants a serverless option that can parallelize the work without managing servers. Which AWS service should the engineer use?

A.AWS Lambda with a function that streams the S3 objects and processes them in 15-minute increments.
B.AWS Glue with a Python shell job and the pandas library, splitting the data into chunks.
C.Amazon EMR Serverless with a Spark job that uses Spark DataFrame operations instead of pandas.
D.Amazon Athena with a CREATE TABLE AS SELECT statement that applies the transformation in SQL.
AnswerC

EMR Serverless runs Spark without cluster management and can distribute a 500 GB transformation across many workers. Spark DataFrames provide a distributed alternative to pandas, allowing the engineer to scale out and avoid the single-machine memory limit while remaining serverless.

Why this answer

EMR Serverless provides a serverless Spark environment that can distribute a large transformation across workers, removing the single-node memory limit of pandas. By converting the logic to Spark DataFrames, the engineer can process 500 GB in parallel without provisioning or managing servers, which matches the stated constraints.

Exam trap

The trap here is assuming that any serverless compute (Lambda or Glue Python shell) can handle large datasets, when only a distributed engine like Spark on EMR Serverless provides the needed parallelism and memory.

68
MCQmedium

A data engineering team needs to ingest streaming data from thousands of IoT devices and store it in Amazon S3 for batch processing. The data arrives at a rate of 10 MB/s, with occasional spikes up to 50 MB/s. The data must be processed in near real-time with minimal latency. Which AWS service should be used for ingestion?

A.Amazon DynamoDB Streams
B.Amazon Kinesis Data Streams
C.Amazon SQS
D.Amazon S3
AnswerB

Kinesis Data Streams ingests high-throughput streaming data with sub-second latency and scales elastically to absorb the 50 MB/s spikes, unlike SQS or batch uploads. It satisfies the near-real-time, minimal-latency requirement while buffering records for downstream delivery into Amazon S3.

Why this answer

Amazon Kinesis Data Streams is designed for real-time streaming data ingestion at scale, handling throughput from megabytes to gigabytes per second with low latency. It can absorb the described 10 MB/s baseline and 50 MB/s spikes by sharding, and integrates directly with AWS Lambda or Kinesis Data Firehose to land data into Amazon S3 for batch processing.

Exam trap

The trap here is that candidates confuse Amazon SQS with a streaming service, but SQS is a pull-based queue with no ordering guarantees across multiple consumers, whereas Kinesis Data Streams provides ordered, replayable, and near-real-time data ingestion.

How to eliminate wrong answers

Option A is wrong because DynamoDB Streams captures changes to DynamoDB tables, not arbitrary streaming data from IoT devices, and its throughput is limited by the table's capacity, making it unsuitable for high-volume, low-latency ingestion. Option C is wrong because Amazon SQS is a message queue for decoupling components, not a streaming ingestion service; it does not support real-time processing with sub-second latency for continuous data streams and has a 256 KB message size limit. Option D is wrong because Amazon S3 is an object storage service, not a real-time ingestion endpoint; writing directly to S3 from thousands of devices would cause high latency due to HTTP overhead and lack of streaming semantics, and it cannot handle the required near-real-time processing.

69
MCQhard

A financial services company is building a real-time fraud detection system. Transaction data is ingested via Amazon Kinesis Data Streams and processed by an Amazon Kinesis Data Analytics for Apache Flink application that runs sliding window aggregations. The output is written to an Amazon S3 bucket for downstream analysis. The Flink application is configured with parallelism of 4 and checkpointing every minute. The company has noticed that the application is experiencing high latency and the checkpointing is frequently failing. The CloudWatch metrics show that the Flink application's CPU utilization is near 100% and the checkpoint duration is spiking to over 5 minutes. The data engineer needs to improve performance. Which action should the data engineer take?

A.Increase the number of shards in the source Kinesis stream to improve throughput.
B.Increase the parallelism of the Flink application to distribute the workload across more resources.
C.Increase the heap memory of the Flink application to handle larger state.
D.Decrease the checkpoint interval to 30 seconds to reduce the amount of state being checkpointed.
AnswerB

Raising parallelism spreads the sliding-window aggregation across more task slots, directly relieving the near-100% CPU saturation that is stretching checkpoint duration beyond the one-minute interval. More parallel subtasks also reduce per-subtask state size, so checkpoint barriers complete faster, satisfying the stem's latency and checkpoint-failure constraint.

Why this answer

The Flink application is experiencing high CPU utilization and checkpointing failures, indicating that the current parallelism of 4 is insufficient to handle the workload. Increasing parallelism distributes the workload across more resources, reducing CPU load per task and allowing checkpointing to complete faster. This is a standard scaling approach for Flink applications in Kinesis Data Analytics.

Exam trap

DEA-C01 often tests the confusion between scaling the source (Kinesis shards) and scaling the processing application (Flink parallelism); candidates might think adding shards helps, but the bottleneck is in the Flink application's CPU.

How to eliminate wrong answers

Option A is wrong because increasing Kinesis shards improves source throughput but does not address the Flink application's CPU bottleneck; it could even worsen the problem by providing more data to process. Option C is wrong because increasing heap memory may help with state size but does not address CPU saturation; checkpoint duration is spiking due to CPU, not memory. Option D is wrong because decreasing the checkpoint interval would increase checkpointing frequency, potentially worsening the issue and not reducing state size.

70
Multi-Selecthard

A data engineer is troubleshooting a Kinesis Data Streams consumer application that is falling behind. The stream has 10 shards and is receiving 5 MB/s of data. The consumer uses the Kinesis Client Library (KCL) with a single worker. The worker is processing all 10 shards but is experiencing high latency and checkpointing delays. Which THREE actions should the engineer take to improve consumer performance? (Select THREE.)

Select 3 answers
A.Increase the number of KCL workers to match the number of shards.
B.Enable enhanced fan-out for the consumer.
C.Decrease the checkpoint interval to reduce checkpointing overhead.
D.Increase the KCL maxRecords parameter to process more records per call.
E.Increase the number of shards in the stream.
AnswersA, B, D

Multiple workers can process shards in parallel, reducing per-worker load.

Why this answer

The KCL worker is processing all 10 shards sequentially within a single worker, causing a bottleneck. By increasing the number of KCL workers to match the number of shards, each worker can process one shard in parallel, significantly improving throughput and reducing latency. This is a standard scaling pattern for KCL-based consumers.

Exam trap

The trap here is that candidates may think decreasing the checkpoint interval (Option C) reduces overhead, when in fact it increases the frequency of DynamoDB writes and can degrade performance; the correct approach is to increase the checkpoint interval or use asynchronous checkpointing.

71
MCQhard

A data engineer is troubleshooting a Kinesis Data Streams application that is experiencing high latency. The stream has 2 shards. The application is using a single Kinesis Client Library (KCL) worker to process all shards. Which change will MOST likely reduce latency?

A.Increase the number of shards to 4.
B.Deploy multiple KCL workers to process shards in parallel.
C.Use a larger instance type for the Kinesis stream.
D.Decrease the number of shards to 1.
AnswerB

Deploying multiple KCL workers lets each worker lease a distinct shard, so the two shards are processed concurrently rather than sequentially by one worker. This directly addresses the stem's constraint: a single worker cannot parallelise across shards, so adding workers raises aggregate throughput and reduces processing latency.

Why this answer

The application uses a single KCL worker to process all 2 shards, which processes records sequentially and causes high latency. Deploying multiple KCL workers (ideally one per shard) enables parallel processing of shards, significantly reducing latency. Option A is incorrect because increasing shard count to 4 adds more capacity but does not address the bottleneck of a single worker; the same worker would process all 4 shards sequentially, potentially worsening latency.

Option C is incorrect because Kinesis Data Streams is a managed service; there is no instance type to change for the stream itself. The KCL worker runs on your compute resources, not on the stream. Option D is incorrect because decreasing shards to 1 reduces the level of parallelism, increasing the workload per shard and likely increasing latency further.

72
Multi-Selecteasy

A data engineer needs to transfer 50 TB of data from an on-premises data center to Amazon S3 over a 1 Gbps network. The transfer must be completed within one week. Which TWO AWS services can be used for this task? (Choose TWO.)

Select 2 answers
A.AWS Glue
B.AWS DataSync
C.AWS Snowball
D.Amazon S3 Transfer Acceleration
E.AWS Direct Connect
AnswersB, C

Designed for network-based bulk data transfer.

Why this answer

AWS DataSync is correct because it is designed to efficiently transfer large datasets over the network using a purpose-built agent that parallelizes data transfer and optimizes network utilization. With a 1 Gbps link, DataSync can transfer 50 TB within a week by leveraging its built-in compression, encryption, and incremental transfer capabilities, making it suitable for this time-constrained migration.

Exam trap

The trap here is that candidates assume S3 Transfer Acceleration can accelerate any transfer, but it only optimizes the last-mile upload to S3 and does not address the bottleneck of moving data from on-premises storage to the internet, nor does it provide a mechanism to pull data from on-premises systems.

73
MCQhard

A company is using Amazon Kinesis Data Streams to ingest real-time clickstream data. The data is consumed by a fleet of EC2 instances running a custom application that processes the records and writes to DynamoDB. The application is experiencing high latency and records are being processed slower than they are produced. The stream has 5 shards. Which action would MOST effectively improve processing speed?

A.Use the Kinesis Client Library (KCL) to automatically distribute shards among instances.
B.Increase the EC2 instance size to provide more CPU and memory.
C.Add more EC2 instances consuming from the same stream without changing shard count.
D.Increase the number of shards in the Kinesis stream.
AnswerD

Kinesis shard count sets the stream's total ingest and per-shard consumer throughput ceiling. With five shards saturated, adding shards raises aggregate capacity and lets the EC2 fleet's consumers process records faster, addressing the production-versus-processing imbalance directly.

Why this answer

The bottleneck is the number of shards in the Kinesis stream. Each shard provides a fixed read capacity of 2 MB/s and 5 read transactions per second. With only 5 shards, the total read throughput is limited regardless of how many EC2 instances consume the data.

Increasing the number of shards increases the total read capacity, allowing more records to be consumed in parallel and reducing processing latency.

Exam trap

The trap here is that candidates often think adding more consumers (EC2 instances) will automatically speed up processing, but they fail to recognize that each shard's read throughput is fixed, so without increasing shards, additional consumers cannot consume more data in parallel.

How to eliminate wrong answers

Option A is wrong because the Kinesis Client Library (KCL) manages shard-to-instance assignment and checkpointing, but it does not increase the total throughput of the stream; it only distributes existing shard capacity among consumers. Option B is wrong because increasing EC2 instance size improves compute resources but does not address the fundamental read throughput limit imposed by the number of shards; the application will still be throttled by the shard's 2 MB/s read limit. Option C is wrong because adding more EC2 instances without increasing the number of shards does not increase the total read capacity; each shard can only be consumed by one record processor at a time (within a single KCL application), so additional instances will remain idle or cause contention.

74
MCQeasy

A data engineer is ingesting streaming data from thousands of IoT devices into AWS. The data is JSON-formatted and must be stored in Amazon S3 for long-term analytics. Which service is most appropriate for real-time ingestion and routing to S3?

A.Amazon SQS
B.Amazon Kinesis Data Firehose
C.Amazon Kinesis Data Streams
D.AWS Glue
AnswerB

Kinesis Data Firehose ingests streaming records and delivers them to Amazon S3 with built-in buffering, compression and format conversion, requiring no consumer code. That satisfies real-time ingestion from thousands of IoT devices plus durable S3 storage for later analytics.

Why this answer

Amazon Kinesis Data Firehose is the most appropriate service because it is designed for real-time ingestion of streaming data and can directly deliver data to Amazon S3 without requiring custom code. It automatically handles buffering, compression, and partitioning of JSON data, making it ideal for long-term analytics storage.

Exam trap

The trap here is that candidates often confuse Kinesis Data Streams with Kinesis Data Firehose, assuming both can directly write to S3, but Data Streams requires a downstream consumer to perform the write, making Firehose the correct choice for direct, managed ingestion to S3.

How to eliminate wrong answers

Option A is wrong because Amazon SQS is a message queue service for decoupling application components, not a streaming ingestion service; it lacks built-in data transformation and direct S3 delivery capabilities. Option C is wrong because Amazon Kinesis Data Streams is a real-time data streaming service that requires a separate consumer (e.g., Lambda or Firehose) to write data to S3, adding complexity and latency; it is not a direct ingestion-to-S3 solution. Option D is wrong because AWS Glue is a serverless ETL service for batch data processing and cataloging, not designed for real-time streaming ingestion or direct routing to S3.

75
MCQhard

A healthcare company is ingesting patient data from a legacy system into an Amazon S3 data lake using AWS Glue. The legacy system produces CSV files with inconsistent schemas (columns may appear or disappear in different files). The data engineer needs to create a Glue ETL job that can handle schema evolution and transform the data into a standardized parquet format. The job should also be able to process new files as they arrive. Which approach should the data engineer use?

A.Use AWS Glue crawlers to create a schema in the Data Catalog and then use a standard Spark DataFrame for transformation.
B.Use AWS Glue DynamicFrames to read the CSV files and apply transformations using resolveChoice and applyMapping.
C.Use a Python shell job in Glue to manually parse each file and write to parquet.
D.Use a Glue ETL job with a static schema defined in the script and ignore files that don't match.
AnswerB

DynamicFrames natively accommodate schema evolution: resolveChoice reconciles columns that appear or disappear across files by casting or dropping them, while applyMapping standardises surviving fields before writing Parquet. This satisfies the inconsistent-schema constraint, and Glue job bookmarks let the same job process newly arrived files incrementally.

Why this answer

AWS Glue DynamicFrames are designed to handle schema evolution and inconsistent data. Using resolveChoice to handle columns that appear/disappear and applyMapping to standardize the schema allows the job to process files with varying schemas and output consistent Parquet.

Exam trap

The trap is assuming a static schema or standard Spark DataFrame can handle schema evolution; candidates must recognize DynamicFrames as the Glue-native solution for inconsistent schemas.

How to eliminate wrong answers

Option A is wrong because Glue crawlers create a static schema in the Data Catalog; a standard Spark DataFrame would fail on files with different schemas. Option C is wrong because a Python shell job lacks the distributed processing and built-in schema evolution capabilities of Glue ETL with DynamicFrames. Option D is wrong because ignoring files that don't match a static schema would drop data, which is unacceptable for patient data.

Page 1 of 6 · 447 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Ingestion Transformation questions.