DEA-C01 · domain
Data Ingestion and Transformation
This domain covers moving data into AWS and reshaping it: batch and streaming ingestion, AWS DMS, Kinesis, Glue, and S3-based pipelines. Questions are scenario-based, asking you to pick the right service for source impact, file size, cadence, schema conversion, and least operational overhead.
Focused practice
Practice Data Ingestion and Transformation questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Data Ingestion and Transformation
Be able to match ingestion services to constraints: source impact, volume, file size, frequency, and schema differences. The single most important thing is justifying the choice with the stated constraint, not the most powerful service.
Choosing AWS DMS for on-premises database replication to S3 or RDS with minimal source impact
Selecting AWS Glue, Kinesis Data Firehose, or DataSync for hourly file ingestion into Amazon S3
Using AWS DMS schema conversion and transformation rules to map source tables to a different target schema
Writing IAM policies with CloudWatch Logs ARNs so Lambda can create log streams and put logs
Watch out for
Common Data Ingestion and Transformation exam traps
- ▸Assuming AWS DMS alone transforms schemas; schema conversion or transformation rules are needed when source and target differ.
- ▸Picking heavyweight services like AWS Glue for small hourly files when Firehose or DataSync is cheaper and lower overhead.
- ▸Writing CloudWatch Logs IAM ARNs for the log group only, forgetting the log-stream resource needed to create streams.
Question index
All Data Ingestion and Transformation questions (447)
Click any question to see the full explanation, or start a practice session above.
A data engineer needs to capture change data capture (CDC) events from an Amazon RDS for PostgreSQL database and stream them to Amazon S3 in near real-time. Which AWS service should be used?
Easy2A company uses AWS Database Migration Service (DMS) to continuously replicate data from an on-premises Oracle database to Amazon S3 in Parquet format. The replication is used for near-real-time analytics. Recently, the DMS task started failing with an error indicating insufficient memory. The source database is large (2 TB). What should a data engineer do to resolve this issue while minimizing changes to the existing architecture?
Hard3A data engineer is building an AWS Glue job that reads semi-structured JSON from Amazon S3 and must flatten nested arrays into relational columns before writing to Amazon Redshift. The transformation logic is complex and the engineer wants to unit test it locally without provisioning a cluster. Which Glue capability should the engineer use to develop and test this transformation logic?
Hard4A company has a large volume of CSV files in S3 that need to be transformed into Parquet using AWS Glue. The files are partitioned by date. The engineer wants to minimize costs by processing only new files each day. Which approach should be used?
Medium5A data engineer needs to ingest data from an on-premises Oracle database to Amazon S3 daily. The data volume is 500 GB per day, and the network bandwidth is 200 Mbps. The requirement is to minimize the impact on the source database and ensure data integrity. Which combination of AWS services should be used?
Medium6A data engineer is building an AWS Glue ETL job that reads data from Amazon S3 and must write the output partitioned by year, month, and day for efficient downstream querying in Amazon Athena. The engineer wants the job to create the partition folders and register them in the Glue Data Catalog automatically. (Choose two.)
Medium7A company is ingesting Apache logs from multiple web servers into AWS. The logs are sent via Amazon CloudWatch Logs to a subscription filter that delivers to a Lambda function. The Lambda function parses the logs and writes to Amazon S3. However, there is a significant backlog. Which THREE actions can reduce the backlog?
Hard8A data engineer is loading data from Amazon S3 into an Amazon Redshift cluster using the COPY command. The S3 bucket contains 500 Parquet files, each about 200 MB, in a single prefix. The COPY job is running slowly and consuming excessive cluster resources. The engineer wants to improve performance without changing the data format or the cluster size. Which action should the engineer take?
Medium9A data engineer needs to transform JSON data from Amazon S3 into Parquet format using AWS Glue. The data contains nested fields. Which Glue feature should the engineer use to define the schema and handle the nested structure?
Easy10A data engineer needs to run a transformation on a small dataset of 500 MB stored in Amazon S3 and load the result into Amazon Redshift. The transformation logic is simple column renaming and filtering. The engineer wants to minimize operational overhead and avoid managing servers. Which approach is most appropriate?
Easy11A data engineer needs to ingest data from an on-premises Apache Kafka cluster into Amazon S3. The data volume is about 10 TB per day. The engineer wants to set up a managed Kafka connector. Which AWS service should they use?
Medium12A company is building a data lake on S3 and needs to ingest data from on-premises Oracle database. The data is 5 TB and changes incrementally. The ingestion must capture changes in near real-time (less than 1 minute latency) and be cost-effective. Which approach should be used?
Hard13A data engineer needs to ingest data from a SaaS application that sends webhooks in JSON format. The data must be stored in S3 for batch analysis. Which AWS services can receive the webhooks and store the data in S3 with minimal custom code? (Choose TWO.)
Easy14A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 data source, applies a filter transformation, and writes to Amazon S3 in Parquet. The engineer notices that the job is reading all files in the prefix, including files that do not match the expected schema, causing job failures. Which action should the engineer take to ensure only valid files are processed?
Medium15A data engineer is designing a data transformation pipeline using AWS Glue. The source data is in Amazon S3 in Parquet format, and the transformed output must be written to another S3 bucket in Parquet format partitioned by year, month, day. The pipeline should handle incremental updates efficiently. Which three features should the engineer use? (Choose THREE.)
Hard16A company uses AWS Glue crawlers to populate the Data Catalog from data in Amazon S3. The crawler fails to update the schema when new columns are added to the CSV files. What is the most likely cause?
Medium17A data engineer needs to transform JSON data from an S3 bucket using AWS Glue. The JSON contains nested arrays and objects. Which Glue transform is best suited for flattening nested structures?
Easy18A data engineer is optimizing an AWS Glue ETL job that reads a large dataset from Amazon S3 and writes to Amazon Redshift. The job currently runs slowly and consumes many DPUs. The engineer wants to improve performance and reduce cost. Which two actions should the engineer take? (Choose two.)
Hard19A data pipeline uses AWS Glue to read from an Amazon S3 bucket containing millions of small CSV files (each < 1 MB). The ETL job is slow. Which optimization would most improve performance?
Hard20A data engineer is implementing a CDC (Change Data Capture) pipeline from a relational database to Amazon S3 using AWS Database Migration Service (DMS). Which TWO configurations are required for continuous replication?
Hard21A data engineer needs to load data from an Amazon S3 bucket into an Amazon Redshift cluster as part of an ETL pipeline. The source files are already in Parquet format and the engineer wants the fastest load with minimal transformation. Which Redshift load method should the engineer use?
Easy22A data engineer needs to transform data in Amazon S3 using SQL statements without managing any infrastructure. The transformations are simple projections and filters, and the engineer wants the results written back to S3 in Parquet. Which AWS service should be used?
Medium23A company needs to transform JSON data from an Amazon S3 bucket into Parquet format and load it into an Amazon Redshift cluster. The transformation includes joining with a reference table stored in Amazon RDS. Which AWS service is BEST suited for this task?
Medium24A company uses Amazon Kinesis Data Firehose to ingest application logs into an Amazon S3 bucket. The logs are in JSON format. The data engineering team wants to convert the logs from JSON to Parquet format before landing in S3. What is the most cost-effective way to achieve this?
Medium25Refer to the exhibit. A data engineer is using a Kinesis Data Stream with 2 shards. The producer uses a partition key that is the user ID (a UUID). The consumer is falling behind. Which change would improve throughput?
Hard26A company needs to ingest data from a MySQL database into Amazon S3 using AWS DMS. The data changes frequently and the requirement is to capture changes in near real-time. Which THREE configurations are necessary?
Hard27A data pipeline uses Amazon Kinesis Data Firehose to deliver data to an S3 bucket. The delivery stream is configured with a buffer interval of 60 seconds and a buffer size of 5 MB. The data arrives at an average rate of 2 MB per second. What is the expected time interval between S3 writes?
Hard28A healthcare company is building a data pipeline to ingest electronic health records (EHR) from hospitals. The data is sent as JSON files via SFTP to an on-premises server. The company wants to move this data to AWS using AWS Transfer Family (SFTP) and then process it with AWS Glue. Data sovereignty regulations require that all data remain within the EU (Frankfurt) region. The pipeline must detect when a new file arrives and start the Glue job automatically. The engineer has set up an AWS Transfer Family server in Frankfurt, and files are uploaded to an S3 bucket in the same region. However, the Glue job is not triggering automatically. The engineer needs to implement automated triggering. What should the engineer do?
Hard29A company ingests streaming data into Amazon Kinesis Data Streams. Producers write records using the PutRecords API with explicit partition keys based on customer ID. A data engineer observes that a few shards are consistently at 100 percent write throughput while others are underutilized, causing throttling. Which action should the engineer take to distribute the load more evenly?
Medium30A company needs to ingest data from multiple on-premises databases into Amazon S3 for analytics. The databases include Oracle, MySQL, and PostgreSQL. The data must be continuously replicated with minimal latency. Which AWS service should be used?
Easy31A data engineer has an AWS Glue ETL job that processes JSON files from Amazon S3. The job currently uses the DynamicFrame method to write output to Amazon Redshift. The engineer needs to improve write performance by using a staging Amazon S3 bucket and parallel COPY operations. Which AWS Glue connection option should the engineer configure?
Medium32Refer to the exhibit. A data engineer runs a Glue job manually and receives a ThrottlingException. The engineer checks the job run history and sees a previous failure with the same error. What is the MOST likely cause of the throttling, and which solution is MOST appropriate?
Hard33A data engineer needs to schedule an AWS Glue extract, transform, and load job to run every day at 02:00 UTC and trigger a dependent Amazon Redshift stored procedure only after the Glue job succeeds. The engineer wants a managed orchestration option that avoids provisioning servers. Which approach should the engineer use?
Easy34A company is using Amazon Kinesis Data Streams with a Lambda consumer to process clickstream data. The data rate is high and the Lambda function is falling behind, resulting in increased processing latency. What is the MOST effective way to improve throughput?
Medium35A company is using AWS Glue to catalog data in Amazon S3. The data is stored in CSV format, but the schema is not consistent across all files. Which TWO actions can the company take to handle schema evolution and ensure the Glue Data Catalog is up to date? (Choose TWO.)
Easy36A company uses AWS DMS to continuously replicate data from an on-premises SQL Server to Amazon Aurora MySQL. The replication lag is increasing. Which THREE actions can reduce the lag? (Choose three.)
Hard37A data engineer needs to ingest JSON data from an on-premises relational database into Amazon S3 every hour. Which AWS service should be used to set up a scheduled, incremental data transfer?
Easy38A company uses Kinesis Data Firehose to deliver streaming data to S3. They need to transform the data by adding a timestamp and removing sensitive fields. Which TWO approaches can achieve this?
Easy39A company wants to ingest real-time clickstream data from a website into Amazon S3 with minimal code. The data should be delivered within 60 seconds of generation. Which AWS service should be used?
Easy40A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The Flink application reads from a Kinesis Data Streams source, performs aggregations, and writes results to Amazon S3. The application is experiencing high checkpoint failures, and the processing lag is increasing. The data volume is 50 MB/s with an average record size of 1 KB. Which TWO actions would improve checkpoint reliability and reduce lag? (Choose TWO.)
Hard41A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job must be able to handle a large number of small files efficiently and minimize the number of output files to improve downstream query performance. Which two actions should the engineer take? (Choose two.)
Medium42A financial services company processes real-time stock trade data. They use Amazon Kinesis Data Streams with a shard count of 5, each shard receiving about 500 records per second. The consumer application uses the Kinesis Client Library (KCL) with DynamoDB for checkpointing. Lately, some records are being processed multiple times. What is the most likely cause?
Hard43A company needs to ingest streaming data from thousands of IoT devices. The data must be processed in real-time and stored in Amazon S3. Which TWO services should be used together?
Medium44A company is ingesting data from multiple sources into S3 using AWS Glue. The data engineer notices that the Glue job is failing with an OutOfMemory error. Which step should be taken to resolve this issue?
Medium45A data engineering team is using AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration must have minimal downtime and needs to capture ongoing changes after the full load. Which THREE resources are required for this task? (Choose three.)
Medium46A company uses AWS Glue ETL to process data from Amazon S3 and write results to Amazon Redshift. The job fails with a memory error when processing large files. Which action should the data engineer take to resolve this issue?
Medium47A data engineer is setting up an Amazon Kinesis Data Analytics application to process streaming data from a Kinesis data stream named "input-stream". The application uses a reference data source from an S3 bucket. The engineer has attached the IAM policy shown in the exhibit to the application's IAM role. When starting the application, the engineer receives an 'AccessDeniedException' error. Which additional permission is required?
Hard48A data engineer needs to run an AWS Glue for Apache Spark ETL job that joins a 40 GB Amazon S3 Parquet dataset with a small 8 MB reference lookup table stored as CSV in Amazon S3. The reference table is read on every join and the job's executors are spending a large amount of shuffle time on the join. The reference table changes only once per month. Which approach MOST efficiently reduces shuffle overhead in the Glue job?
Medium49A company uses AWS Glue to run ETL jobs daily. The jobs consume data from an Amazon RDS for MySQL database and write results to Amazon S3. The company wants to minimize the impact on the source database during extraction. Which THREE actions should the data engineer take to achieve this? (Choose THREE.)
Medium50A company is using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data must be transformed from JSON to Parquet format before landing in S3. The transformation logic is simple: convert the JSON schema to Parquet. Which approach meets the requirements with the least operational overhead?
Medium51Refer to the exhibit. A Lambda function named 'IngestionProcessor' is failing. The engineer checks CloudWatch Logs and sees the log group exists but storedBytes is 0. Why might the logs show no data?
Easy52A company is using Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data is delivered in 5-minute intervals. The company wants to reduce the delivery frequency to 1 minute to get data faster. Which parameter should be changed in the Firehose delivery stream configuration?
Easy53A data engineer is designing a streaming pipeline using Amazon Kinesis Data Streams with a shard count of 10. The incoming data rate is 1 MB/second. The consuming application uses the Kinesis Client Library (KCL) with a single worker. What is the most likely performance bottleneck?
Hard54A company wants to ingest real-time clickstream data from a website into Amazon S3 with a maximum latency of 60 seconds. The data volume peaks at 500 MB/s. Which service should they use to buffer and deliver the data to S3?
Easy55A data engineer needs to transform data in an AWS Glue job using a custom Python library that is not available by default. The library is packaged as a .whl file stored in Amazon S3. The engineer wants the Glue job to use this library without modifying the job script to install it at runtime. What should the engineer do?
Medium56A company uses Amazon Kinesis Data Firehose to ingest log data from web servers into Amazon S3. The data is in JSON format and each record is approximately 2 KB. The delivery stream is configured to buffer incoming records for 60 seconds or 5 MB, whichever comes first. The company notices that the data in S3 is delayed by up to 5 minutes during peak hours. Which action would most effectively reduce the delivery latency?
Hard57A company uses AWS Glue to process data from multiple S3 buckets. The Glue job runs daily and reads data from a bucket that contains millions of small files (each < 1 MB). The job has been running for hours and is often close to the 8-hour timeout limit. Which optimization would MOST reduce the job's runtime?
Hard58A data engineer is ingesting records from Amazon Kinesis Data Streams into Amazon S3 using AWS Lambda as the consumer. Each stream shard delivers up to 1,000 records per second, and the Lambda function writes each record as an individual small object, causing many tiny S3 files and high PUT costs. The engineer wants fewer, larger objects while keeping near-real-time delivery. What should the engineer do?
Medium59A company uses Amazon Kinesis Data Firehose to ingest data into an S3 bucket. The data is in JSON format and the team wants to convert it to Parquet before storage. Which TWO configurations are required?
Medium60A company uses AWS Glue DataBrew for data preparation. The data source is an S3 bucket with millions of small CSV files (each < 1 MB). The DataBrew project takes a long time to load the sample data. What is the most likely cause and solution?
Hard61A data engineer needs to run a transformation that processes semi-structured JSON records already stored in Amazon S3 and write the results back to S3 in Parquet format. The team prefers a serverless, Apache Spark-based approach with minimal infrastructure management and wants to use the AWS Glue Data Catalog for metadata. Which approach should the engineer use?
Easy62A data engineer runs the describe-stream command and sees the output above. The stream has a retention period of 24 hours. The engineer needs to ensure that consumers can replay data for up to 7 days. Which action is required?
Hard63A data engineer needs to ingest streaming data from a social media API into Amazon S3 for batch analytics. The data arrives at a rate of 500 records per second. Which service should be used to capture the stream?
Easy64A company is streaming IoT data from thousands of devices into Amazon Kinesis Data Streams. The data must be transformed in real time before being stored in Amazon S3. Which service should be used to perform the transformation as the data streams through Kinesis?
Medium65Which TWO AWS services can be used as sources for AWS Glue ETL jobs? (Choose two.)
Easy66Match each AWS service to its primary purpose in data engineering.
Medium67A data engineer needs to run a one-time transformation on a 500 GB dataset stored in Amazon S3. The transformation is written in Python and uses pandas, which cannot handle the full dataset in memory. The engineer wants a serverless option that can parallelize the work without managing servers. Which AWS service should the engineer use?
Easy68A data engineering team needs to ingest streaming data from thousands of IoT devices and store it in Amazon S3 for batch processing. The data arrives at a rate of 10 MB/s, with occasional spikes up to 50 MB/s. The data must be processed in near real-time with minimal latency. Which AWS service should be used for ingestion?
Medium69A financial services company is building a real-time fraud detection system. Transaction data is ingested via Amazon Kinesis Data Streams and processed by an Amazon Kinesis Data Analytics for Apache Flink application that runs sliding window aggregations. The output is written to an Amazon S3 bucket for downstream analysis. The Flink application is configured with parallelism of 4 and checkpointing every minute. The company has noticed that the application is experiencing high latency and the checkpointing is frequently failing. The CloudWatch metrics show that the Flink application's CPU utilization is near 100% and the checkpoint duration is spiking to over 5 minutes. The data engineer needs to improve performance. Which action should the data engineer take?
Hard70A data engineer is troubleshooting a Kinesis Data Streams consumer application that is falling behind. The stream has 10 shards and is receiving 5 MB/s of data. The consumer uses the Kinesis Client Library (KCL) with a single worker. The worker is processing all 10 shards but is experiencing high latency and checkpointing delays. Which THREE actions should the engineer take to improve consumer performance? (Select THREE.)
Hard71A data engineer is troubleshooting a Kinesis Data Streams application that is experiencing high latency. The stream has 2 shards. The application is using a single Kinesis Client Library (KCL) worker to process all shards. Which change will MOST likely reduce latency?
Hard72A data engineer needs to transfer 50 TB of data from an on-premises data center to Amazon S3 over a 1 Gbps network. The transfer must be completed within one week. Which TWO AWS services can be used for this task? (Choose TWO.)
Easy73A company is using Amazon Kinesis Data Streams to ingest real-time clickstream data. The data is consumed by a fleet of EC2 instances running a custom application that processes the records and writes to DynamoDB. The application is experiencing high latency and records are being processed slower than they are produced. The stream has 5 shards. Which action would MOST effectively improve processing speed?
Hard74A data engineer is ingesting streaming data from thousands of IoT devices into AWS. The data is JSON-formatted and must be stored in Amazon S3 for long-term analytics. Which service is most appropriate for real-time ingestion and routing to S3?
Easy75A healthcare company is ingesting patient data from a legacy system into an Amazon S3 data lake using AWS Glue. The legacy system produces CSV files with inconsistent schemas (columns may appear or disappear in different files). The data engineer needs to create a Glue ETL job that can handle schema evolution and transform the data into a standardized parquet format. The job should also be able to process new files as they arrive. Which approach should the data engineer use?
Hard76An e-commerce company ingests clickstream data from their website into Amazon S3. The data is in JSON format, and each file is about 10 MB. They need to transform the data into a columnar format for analytics and load it into Amazon Redshift nightly. The transformation should be cost-effective and require minimal operational overhead. Which approach meets these requirements?
Medium77A data engineer is tasked with transforming JSON data from an S3 bucket into Parquet format for efficient querying. The transformation should run on a schedule every hour. Which AWS service is best suited for this task?
Easy78A company uses AWS Glue ETL jobs to process data from an S3 data lake. The job reads data in CSV format, transforms it, and writes to Parquet. The job runs daily and takes 2 hours to complete. The data volume is increasing by 20% each month. The engineer wants to reduce the job runtime. Which action is most effective?
Medium79A streaming application sends data to Amazon Kinesis Data Streams. The data must be enriched with reference data from an Amazon DynamoDB table in real-time. Which AWS service can be used to perform this enrichment with minimal latency?
Medium80A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer notices that the job fails with an error indicating that the Redshift table does not exist. The engineer has already created the Redshift cluster and database. What should the engineer do to resolve the error?
Hard81A data engineer needs to ingest data from an external partner's FTP server to Amazon S3. The data arrives once daily as a CSV file. Which AWS service should be used for this ingestion?
Easy82A company is using Amazon Kinesis Data Firehose to ingest data into Amazon S3. The data must be transformed from JSON to Parquet format before delivery. Which feature should be enabled on the Firehose delivery stream?
Easy83A data engineer needs to design a data ingestion pipeline that captures streaming data from mobile app events into Amazon S3 for analytics. The pipeline must support real-time processing of events and allow for schema evolution over time. Which AWS services should the engineer use? (Choose THREE.)
Medium84A data engineer uses AWS Glue DataBrew to clean a dataset stored in Amazon S3. The recipe includes a 'Filter' step that removes rows where the 'status' column equals 'INVALID'. After running the recipe, the engineer notices that the output still contains rows with status 'INVALID'. The recipe was published and the job ran successfully. What is the most likely cause?
Hard85A data engineer is designing a serverless data ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data must be transformed using AWS Lambda before being written to S3. Which two steps are required to enable this transformation? (Select TWO.)
Easy86A data engineer is using AWS Glue to run an ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job processes data in CSV format and the engineer wants to ensure the output is partitioned by year, month, and day based on a timestamp column in the data. The engineer needs to optimize the job for performance and cost. Which approach should the engineer take?
Medium87A data engineer is using AWS Glue Studio to create a visual ETL job that reads from an Amazon S3 bucket containing JSON files, applies a filter transformation, and writes the output to Amazon Redshift. The job must run daily. The engineer notices that the job is taking a long time to complete and wants to improve performance. Which action should the engineer take to optimize the job?
Medium88A company uses Amazon Kinesis Data Firehose to deliver data to an S3 bucket. The data contains personally identifiable information (PII) that must be redacted before storage. Which TWO actions can achieve this requirement? (Choose TWO.)
Hard89A data engineer is using AWS Glue to run an ETL job that reads data from Amazon DynamoDB and writes to Amazon Redshift. The job fails with a 'ThroughputExceededException' error. What is the most likely cause?
Easy90A data engineer is troubleshooting a Kinesis Data Firehose delivery stream that is experiencing high error rates when writing to an S3 bucket. The error logs indicate 'AccessDenied' errors. The S3 bucket policy allows access from the Firehose service, but the errors persist. What is the most likely cause?
Hard91A data engineer is using AWS Glue Studio to build a visual ETL job that reads JSON files from Amazon S3, applies a mapping transform, and writes Parquet to another S3 location. The engineer notices that the job is writing many small files, which hurts downstream query performance. Which action should the engineer take to reduce the number of output files without changing the source data?
Easy92Which TWO AWS services can be used to transform data in transit during ingestion? (Choose 2.)
Easy93A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to improve the performance of the Glue job, which currently takes several hours to complete. The job reads large Parquet files, performs joins, and writes to Redshift. Which two actions should the engineer take to improve performance? (Choose two.)
Hard94A data engineer needs to transform CSV files arriving in an S3 bucket into Parquet format and store them in another S3 bucket. The transformation is simple and on-demand, triggered by data arrival. Which solution is the MOST cost-effective and requires the least operational overhead?
Medium95Order the steps to troubleshoot a failed AWS Glue job that reads from JDBC and writes to S3.
Medium96A data engineering team uses AWS Glue to extract, transform, and load (ETL) data from Amazon RDS for MySQL to Amazon S3. The job runs daily and processes incremental data. The team notices that the job is taking longer than expected. Which TWO actions can improve the job performance? (Choose two.)
Medium97A data engineer is designing a pipeline that ingests JSON logs from an application into Amazon S3. The logs contain a timestamp field. The pipeline must partition the data by date in S3 (e.g., year=2024/month=10/day=01). Which approach minimizes transformation effort?
Medium98A data pipeline ingests streaming data from Kinesis Data Streams into S3 via Kinesis Data Firehose. Occasionally, small files are written to S3, increasing downstream processing costs. What is the most efficient way to reduce the number of small files?
Hard99A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is stored as CSV files and is updated daily. The engineer wants to load only new data each day without duplicating existing records. Which AWS service or feature should the engineer use to automate this process?
Easy100A data engineer must ingest a 4 TB Oracle database into Amazon S3 nightly. The database is on-premises and the network link supports only 200 Mbps. The engineer wants to minimize the total transfer time and avoid impacting production. Which approach should the engineer use?
Medium101Refer to the exhibit. A data engineer deploys this CloudFormation template to create an AWS Glue job. The job fails on the first run with an error: 'AccessDeniedException: User: arn:aws:sts::123456789012:assumed-role/GlueServiceRole/... is not authorized to perform: s3:GetObject on resource: s3://my-bucket/scripts/etl.py'. What is the most likely cause?
Medium102A data engineer is using AWS Glue Studio to build a visual ETL job that joins a large Amazon S3 dataset with a small reference dataset of country codes. The join is currently implemented as a standard join, and the job runs slowly and shuffles large amounts of data. The engineer wants to optimize performance without changing the output. Which change should the engineer make?
Medium103A company uses Kinesis Data Analytics for SQL-based real-time analytics on streaming data. They notice that the application is processing data slower than the incoming rate, causing increased latency. Which action is MOST likely to improve the throughput?
Hard104A company uses AWS Glue to run ETL jobs that process data from Amazon RDS for MySQL and load it into Amazon S3. The job runs daily and processes incremental changes using the JDBC connection. Recently, the job has been failing with a 'Communications link failure' error. The RDS instance is in a private subnet. Which step should the engineer take first to diagnose the issue?
Hard105A company is building a data lake on Amazon S3. Data arrives from multiple sources in JSON, CSV, and Avro formats. The data must be transformed to Parquet and partitioned by date and source. Which TWO services can perform this transformation with minimal custom code? (Choose TWO.)
Medium106A social media company ingests user activity data from multiple sources using Amazon Kinesis Data Firehose. The data is delivered to Amazon S3 in near-real-time. The company wants to transform the data by adding a timestamp and masking email addresses before storing it in S3. The transformation should be applied to all records. What is the most cost-effective way to implement this transformation?
Medium107A data engineer reviews the Glue job configuration. The job fails when processing large datasets. The error message indicates out-of-memory in the executors. Which change to the job configuration will most directly address this issue?
Hard108A data engineer is configuring an AWS Glue crawler to catalog data in an Amazon S3 bucket. The bucket contains CSV files organized in folders by year and month, and new files are added daily. The engineer wants the crawler to detect schema changes automatically and avoid reprocessing unchanged files on subsequent runs. (Choose two.)
Hard109A data engineer is building an AWS Glue ETL job that reads large Parquet datasets from Amazon S3 and must optimize performance and cost. The engineer wants to reduce the number of small files written to the target S3 prefix and improve read efficiency. (Choose two.)
Medium110A data engineering team needs to transform CSV files to Parquet format after they land in an S3 bucket. The transformation should be triggered automatically as soon as a new file arrives. Which AWS service is best suited for this task?
Easy111Which TWO actions can improve the performance of an AWS Glue ETL job that processes large datasets in Amazon S3? (Choose two.)
Medium112A data engineer is configuring an AWS Glue crawler to catalog CSV files stored in Amazon S3. The files are organized in prefixes by year and month, and the engineer wants the crawler to detect new partitions automatically and avoid re-crawling unchanged partitions. (Choose two.)
Medium113A data engineer needs to run an AWS Glue extract, transform, and load (ETL) job that joins an Amazon S3-based Parquet dataset with a slowly changing dimension table in Amazon Redshift. The Redshift cluster is in a private subnet and cannot be reached over the public internet. The engineer wants the Glue job to read from Redshift without exposing credentials in the job script. Which combination of actions should the engineer take to meet these requirements?
Medium114An e-commerce company uses AWS Glue to run ETL jobs that transform clickstream data from Amazon S3. The job reads Parquet files, performs aggregations, and writes the results to Amazon Redshift. The job runs successfully but takes longer than expected. The data volume is increasing. Which design change would MOST improve the job's performance?
Hard115A data engineer is building an AWS Glue job that reads from a large Parquet dataset in Amazon S3 partitioned by year/month/day and writes aggregated results to Amazon Redshift. The job currently reads all partitions and takes several hours. The engineer wants the job to process only partitions from the last seven days and reduce runtime. Which change should the engineer make?
Hard116A data pipeline uses AWS Glue to process large CSV files. The team notices that some jobs fail with out-of-memory errors. Which TWO configuration changes can help mitigate this issue?
Hard117A data engineer is designing a data ingestion pipeline to load millions of small JSON files from an on-premises FTP server into Amazon S3. The pipeline should minimize cost and operational overhead. Which approach is most suitable?
Medium118A data engineer is troubleshooting a daily batch ingestion pipeline that uses AWS Glue to read CSV files from Amazon S3 and write Parquet files to another S3 bucket. The job runs successfully but takes significantly longer than expected. The engineer notices that the input data is highly skewed with many small files. Which is the most effective optimization to reduce job duration?
Hard119A data engineer must transform data in Amazon S3 using Apache Spark. The transformation logic needs to be reused across multiple AWS Glue jobs, and the engineer wants to version-control the code and run it in a serverless environment without managing clusters. Which approach should the engineer take?
Medium120A retail company uses Amazon Kinesis Data Firehose to ingest clickstream data from its website into an Amazon S3 bucket. The data includes fields: user_id, event_type, timestamp, page_url. Recently, the data engineering team noticed that some records have malformed JSON (missing commas, extra brackets) causing delivery failures to S3. The Firehose delivery stream is configured to retry failed records for 300 seconds, after which the records are sent to an S3 bucket for failed records. The team wants to transform the data to correct malformed JSON before delivery to the main S3 bucket. They need a solution that does not require managing servers and can handle high throughput. What should the team do?
Medium121Arrange the steps to create an AWS Glue job that transforms data from Amazon S3 to Amazon Redshift in the correct order.
Medium122A company uses AWS Glue to transform data in S3. The Glue job fails with memory errors. Which THREE actions can help resolve this?
Medium123A company needs to ingest real-time clickstream data from a web application into Amazon Redshift with minimal latency. The data volume is high and requires processing before loading. Which architecture is MOST appropriate?
Hard124A data engineer is configuring an AWS Glue crawler to catalog data stored in an Amazon S3 bucket. The data is partitioned by year, month, and day in a Hive-style structure (for example, s3://bucket/data/year=2023/month=01/day=15/). The engineer wants the crawler to recognize the partitions and add them to the AWS Glue Data Catalog. What should the engineer do?
Medium125Match each AWS database service to its primary use case.
Medium126A company is using AWS Database Migration Service (DMS) to migrate a 2 TB MySQL database to Amazon Aurora MySQL. The migration is taking longer than expected. The source database is in a different AWS region. Which change would MOST likely improve the migration speed?
Hard127A company needs to ingest data from a relational database into Amazon S3 for analytics. The database is an Amazon RDS MySQL instance. Which AWS service should be used for a one-time historical data load?
Easy128A data engineer is using AWS Glue ETL to transform data from an S3 data lake. The job fails with a memory error. Which approach should be used to resolve this issue without major code changes?
Medium129A company uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data delivery is delayed by up to 5 minutes. The engineer wants to reduce the delay to under 1 minute. Which parameter should be adjusted?
Easy130A data pipeline uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The delivery occasionally fails with 'Firehose is throttled'. What should be done to reduce throttling?
Easy131A data engineer is designing an ingestion pipeline where AWS Lambda processes records from an Amazon Kinesis Data Stream. During peak traffic, records are being reprocessed and some are lost. The engineer needs to make the consumer resilient to failures and avoid duplicate processing. (Choose two.)
Hard132A company uses AWS Glue ETL jobs to transform data from Amazon RDS to Amazon S3 daily. The job recently started failing with memory errors. The data volume has grown 3x in the past month. Which change should the data engineer make to resolve the issue?
Medium133A data engineer needs to transform data in Amazon S3 using AWS Glue. The job must handle schema evolution and partition pruning. Which THREE features should be used?
Hard134A data engineer is designing a data ingestion pipeline for a social media analytics platform. The pipeline must ingest tweets in real-time, perform sentiment analysis, and store results in Amazon S3. The sentiment analysis is compute-intensive and must be done as the data arrives. The estimated throughput is 10,000 tweets per second. Which architecture is most suitable?
Hard135A company uses AWS Glue to process streaming data from Amazon Kinesis Data Streams. The data is JSON formatted and includes a timestamp field. The company wants to partition the output in Amazon S3 by date and hour, and ensure exactly-once processing semantics. Which combination of configurations should be used?
Medium136A company is building a data lake on S3. They have a large volume of CSV files (hundreds of GB) in a source bucket. They need to convert them to Parquet, partition by date, and ensure the data is encrypted at rest with SSE-KMS. The pipeline must be triggered automatically when new files arrive. Which THREE steps should be part of the solution? (Choose THREE.)
Hard137A data engineering team needs to ingest streaming data from an application into Amazon S3 for analytics. The data volume is moderate and the team wants the lowest operational overhead. Which AWS service should they use?
Easy138A marketing analytics team needs to ingest customer transaction data from an on-premises PostgreSQL database into Amazon S3 for analysis. The data volume is about 10 GB daily, and the team wants to perform full refresh daily (truncate and load) into S3 as Parquet files. The company has a Direct Connect connection to AWS. The team needs a simple, managed solution that minimizes operational overhead. What should the team use?
Easy139A data engineer is configuring an AWS Glue job bookmark on a job that reads partitioned Parquet data from Amazon S3 and writes to another S3 location. The engineer notices that reprocessing keeps occurring and wants the bookmark to correctly skip already-processed data. Which two actions should the engineer take? (Choose two.)
Medium140A data engineer is designing a data ingestion pipeline for IoT sensor data. The data is generated at a high velocity and must be processed in near real-time. The pipeline must also handle bursty traffic. Which TWO AWS services should be combined to achieve this? (Choose TWO.)
Medium141A data engineer is building an AWS Glue ETL job that must read from an Amazon S3 bucket in the same account and write to an Amazon Redshift cluster in a private VPC. The job must not traverse the public internet and must use least-privilege credentials. (Choose two.)
Hard142A company is using AWS Database Migration Service (DMS) to migrate a 2 TB Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime. The source database is highly active with continuous writes. Which DMS migration type and additional configuration should the engineer use?
Medium143A data engineer is building a pipeline that ingests records from an Amazon Kinesis data stream and writes them to Amazon S3 in Parquet format. The engineer wants to use AWS Glue to perform the transformation and needs the pipeline to handle records that arrive out of order and to deduplicate based on a record ID. Which combination of features should the engineer use?
Medium144A company is using AWS Glue ETL to transform and load data from Amazon S3 to Amazon Redshift. The data engineer notices that the job is taking longer than expected. Which TWO actions can improve the job performance?
Medium145A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes 500 GB of data. The engineer notices that the job takes several hours and wants to optimize performance. The data is stored in Parquet format and partitioned by date. Which optimization should the engineer implement to improve the job's performance?
Hard146A company needs to ingest data from multiple SaaS applications into Amazon S3. The data sources provide REST APIs. Which AWS service can be used to build a fully managed data ingestion pipeline without writing custom code?
Easy147A data engineer needs to run an AWS Glue ETL job that reads from an Amazon S3 bucket in another AWS account. The bucket owner has granted cross-account access, and the Glue job runs with an IAM role in the engineer's account. The job fails with an access denied error when reading the source objects. Which change is required to allow the Glue job to read the cross-account S3 data?
Easy148Refer to the exhibit. A data engineer runs the AWS CLI command to describe a Glue job. The job is expected to process new data incrementally using job bookmarks. However, the job reprocesses all data every time it runs. What is the MOST likely reason?
Hard149A company needs to ingest data from an on-premises Oracle database into Amazon S3 on a daily basis. The data volume is about 100 GB per day. Which AWS service is BEST suited for this task?
Easy150A data engineer is configuring an AWS Glue crawler to catalog data stored in Amazon S3. The data is organized as Parquet files under prefixes named by year, month, and day, such as s3://analytics/events/year=2024/month=05/day=17/. Queries in Amazon Athena must use partition pruning to limit scanned data. Which crawler configuration should the engineer choose?
Medium151A data engineer is using AWS Glue to run a nightly ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3 in Parquet format. The DynamoDB table is large and has a high read capacity. The engineer wants to minimize the impact on the DynamoDB table's performance and reduce the ETL job's runtime. Which approach should the engineer take?
Medium152A data engineer is building an AWS Glue ETL job that reads a large JDBC table from Amazon RDS for PostgreSQL. The job must read the table in parallel to reduce runtime, but the table has no numeric primary key or monotonically increasing column. Which AWS Glue connection property should the engineer configure to enable parallel reads?
Medium153A company uses AWS Glue to process JSON logs from S3. The logs have a nested structure and the schema evolves over time. The data engineer needs to ensure the Glue job can handle schema changes without failing. Which configuration should be used?
Hard154A company runs a SQL Server transactional database on Amazon RDS. They need to capture change data (inserts, updates, deletes) in near real-time and replicate them to an Amazon S3 data lake. Which AWS service is most suitable?
Medium155Which TWO practices improve the performance of AWS Glue ETL jobs? (Choose two.)
Medium156A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream, performs a 1-minute tumbling window aggregation, and writes results to an S3 bucket. Recently, the application started experiencing checkpoint failures and increasing processing delay. Which action should the engineer take FIRST to diagnose the issue?
Hard157A company is ingesting real-time financial transactions into Amazon Kinesis Data Streams. The data is then consumed by a Kinesis Data Analytics for Apache Flink application that calculates running totals. The application is experiencing high latency and checkpoint failures. Which TWO steps should the engineer take to improve performance and reliability? (Select TWO.)
Hard158A data engineer needs to schedule a daily AWS Glue ETL job that transforms data in Amazon S3. The job must run at 2:00 AM UTC every day. What is the simplest way to achieve this?
Easy159A data engineer is building an AWS Glue ETL job that reads JSON files from Amazon S3, flattens nested arrays, and writes Parquet to another S3 bucket. The job runs daily and processes about 2 TB. The engineer notices that job runs are failing intermittently with OutOfMemory errors during the shuffle phase. The job uses 10 G.1X workers. Which change should the engineer make to resolve the memory failures while minimizing cost?
Hard160A data engineer is using AWS Step Functions to orchestrate a daily ingestion workflow. The workflow calls an AWS Glue job, then runs an AWS Lambda function to validate output, and finally starts an Amazon Redshift stored procedure. The engineer needs to ensure that if the Glue job fails, the workflow retries the Glue job up to three times with exponential backoff before failing the entire execution. Which Step Functions feature should the engineer configure?
Hard161A data engineer is building a streaming ingestion pipeline using Amazon Kinesis Data Streams. The producer application writes records with an explicit partition key derived from the device ID, and there are approximately 2,000 active devices. The engineer needs to ensure that records for the same device are processed in order by a downstream consumer. Which configuration should the engineer verify to guarantee per-device ordering?
Medium162A data engineer is using AWS Glue to process a large dataset in Amazon S3. The dataset consists of many small JSON files (average 100 KB each) stored in a single prefix. The Glue job reads these files, performs transformations, and writes the output to Parquet in another S3 location. The job is running slowly and consuming many DPUs. Which action should the data engineer take to improve performance?
Hard163A company wants to ingest streaming data from thousands of IoT devices into Amazon S3 with minimal latency and then transform the data using Spark SQL. Which AWS service should be used for data ingestion?
Medium164A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains a column with inconsistent date formats and another with trailing whitespace. The engineer wants to apply these transformations reproducibly and schedule the recipe to run daily. Which combination of steps should the engineer take?
Hard165A data streaming application uses Kinesis Data Streams with 10 shards. The data producer is throttled frequently. Which action should be taken to resolve this issue?
Hard166A data engineer needs to load data from an Amazon DynamoDB table into an Amazon S3 bucket for analytics. The table is approximately 500 GB and has a high volume of write traffic. The engineer must minimize the impact on the table's read capacity and avoid consuming provisioned throughput. What is the MOST appropriate method to export the data?
Medium167A company uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data is in JSON format and contains a 'timestamp' field with a Unix epoch value. The company wants to partition the S3 objects by year, month, day, and hour based on the timestamp. What is the MOST efficient method to achieve this?
Hard168A company wants to ingest real-time data from a social media API into Amazon S3 for analysis. The API provides data as JSON records. Which AWS service is best suited for this ingestion?
Easy169A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is in CSV format and is updated daily. The engineer wants to use a fully managed service that can handle the load without managing infrastructure. Which AWS service should the engineer use?
Easy170A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains inconsistent date formats in a column named 'transaction_date'. The engineer needs to standardize all dates to ISO 8601 format (YYYY-MM-DD) and then write the cleaned data to a new S3 location. Which transformation should the engineer apply?
Medium171A data engineer is using AWS Glue to run a job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job must be scheduled to run every day at 2:00 AM UTC. Which AWS Glue feature should the engineer use to set up this schedule?
Easy172A company wants to import data from an external FTP server into Amazon S3 on a daily basis. The data volumes are moderate. Which AWS service is MOST suitable for this task?
Easy173A data engineer is building a real-time data pipeline to ingest sensor data from IoT devices. The data is sent to AWS IoT Core, which publishes messages to a Kinesis Data Stream. Each message is about 1 KB in size. The data must be transformed (add a device location field) and then stored in Amazon S3 for long-term analytics. The engineer has set up a Lambda function to transform the records and write to S3. However, the engineer notices that the Lambda function is invoked thousands of times per second, causing high costs and occasional throttling. The Lambda function processes only one record at a time. The engineer wants to reduce the number of Lambda invocations and improve throughput. What should the engineer do?
Medium174A data engineering team is building a data lake on Amazon S3. They need to ingest data from multiple sources: (1) streaming IoT data, (2) daily CSV exports from an on-premises system via SFTP, and (3) change data capture (CDC) from an Amazon Aurora database. Which THREE services should the team use to ingest these data sources?
Hard175Refer to the exhibit. A data engineer has attached this IAM policy to an AWS Glue job role. The Glue job fails when trying to write transformed data to an S3 bucket located in a different AWS account. What is the most likely reason?
Medium176A data engineer needs to run a PySpark transformation on a 2 TB dataset stored in Amazon S3 and write the output back to S3 in Parquet. The team wants to use AWS Glue but does not want to manage clusters or tune Spark configuration manually. They also want to pay only for the time the job runs. Which AWS Glue component should the engineer use?
Easy177A data engineer is designing a data ingestion pipeline using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format and must be converted to Apache Parquet before storage. The engineer wants to minimize costs and operational effort. Which two actions should the engineer take to meet these requirements? (Choose two.)
Medium178A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to ensure that the job handles data quality issues such as duplicate records and missing values before loading. The job must also minimize the amount of data shuffled across the network. Which two actions should the engineer take? (Choose two.)
Medium179A data engineer runs an AWS Glue job that reads from an Amazon Kinesis Data Stream and writes to Amazon S3. The job must process records in order per shard and must checkpoint progress so it can resume after a failure without reprocessing all records. Which Glue configuration BEST supports this?
Hard180A data engineer is configuring an AWS Glue ETL job that reads semi-structured JSON from Amazon S3 and must flatten nested arrays before loading into Amazon Redshift. During testing, the job fails with an AnalysisException stating that a column named 'events' cannot be resolved, even though the AWS Glue Data Catalog table shows the column. The engineer confirms the catalog table was created by a crawler and the S3 data is present. Which action will most directly resolve the schema resolution failure?
Medium181A company is designing a data ingestion pipeline for clickstream data from a website. The data must be ingested in near real-time. Which TWO services can be used together to build this pipeline?
Medium182Refer to the exhibit. An AWS Glue ETL job is failing with an OutOfMemoryError. The job reads from Amazon S3 and performs a GROUP BY on a large dataset. Which change should the data engineer make to resolve this error?
Medium183A data engineer needs to load data from an Amazon DynamoDB table into an Amazon S3 data lake nightly. The table is large and the engineer wants to avoid consuming provisioned read capacity on the live table. Which approach should the engineer use?
Medium184A data engineer is using AWS Glue to read from an Amazon S3 bucket that contains data in Apache Parquet format, partitioned by year/month/day. The Glue job needs to read only the data for the last 7 days. The engineer wants to minimize the amount of data scanned and improve job performance. Which approach should be used to filter the partitions efficiently?
Hard185A company runs an e-commerce platform that generates clickstream data from user interactions on their website. The data is sent as JSON objects via HTTP POST to an API Gateway endpoint, which triggers a Lambda function that writes each record to a Kinesis Data Stream (100 shards). A second Lambda function consumes the stream, transforms the data (enriches with geolocation from a DynamoDB table), and writes to a Kinesis Data Firehose delivery stream that delivers Parquet files to an S3 data lake every 5 minutes. The system has been working for months, but recently the Firehose delivery stream started showing 'DeliveryFailed' errors for a subset of records. The errors point to 'InvalidData' from the Lambda transformation. The engineer reviews the Lambda transformation code and notices that the geolocation lookup occasionally fails because the DynamoDB table has a throttling issue. The engineer needs to handle these failures gracefully so that records that fail enrichment are still delivered to S3 with a null geolocation field, without blocking other records. Which course of action should the engineer take?
Hard186Refer to the exhibit. A data engineer is using a Kinesis Data Stream with one shard. The application writes 2000 records per second, each 1 KB. The put record calls are frequently throttled. What is the most likely cause?
Medium187A company uses AWS Lambda to process messages from an Amazon SQS queue. The messages contain JSON payloads that need to be transformed and written to an Amazon DynamoDB table. Recently, the Lambda function has been timing out and messages are being sent to the dead-letter queue (DLQ). What is the BEST way to troubleshoot and resolve this issue?
Medium188A company uses AWS Glue to run ETL jobs daily. The data engineer wants to reduce costs by optimizing the job configuration. Which two actions will help reduce costs? (Choose TWO.)
Easy189A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an on-premises Oracle database to an Amazon S3 bucket in near real time. The source table has a primary key and the database is configured for ARCHIVELOG mode. The engineer needs change data capture (CDC) to capture INSERT, UPDATE, and DELETE operations. Which AWS DMS task setting should be used to capture ongoing changes?
Medium190A data engineer is designing a data ingestion pipeline for IoT sensor data. The sensors send JSON messages every second. The data must be available in Amazon S3 within 5 minutes and must be transformed (JSON to Parquet) before storage. Which combination of services meets these requirements?
Hard191A data engineer is building a streaming pipeline using Amazon Kinesis Data Streams and AWS Lambda. The Lambda function processes records and writes to Amazon DynamoDB. The engineer notices that the Lambda function is throttled during high traffic. Which action should the engineer take to reduce throttling?
Hard192A company needs to ingest streaming data from multiple sources and store it in Amazon S3. The data volume is up to 5 GB per hour. What is the MOST cost-effective ingestion service?
Easy193A data engineer must load a 250 GB uncompressed CSV dataset from Amazon S3 into Amazon Redshift. The data must be loaded daily, and the engineer wants to minimize load time and avoid saturating the cluster's leader node. Which approach meets these requirements?
Easy194A data pipeline uses AWS Glue to process data from an S3 data lake. The pipeline fails intermittently with a 'ThrottlingException' when writing to a DynamoDB table. What is the MOST likely cause?
Easy195A company stores IoT sensor data in S3 as JSON files. They need to convert the data to Parquet format for efficient querying with Amazon Athena. Which AWS service can perform this transformation with minimal effort?
Easy196A company is using AWS Glue to process streaming data from Amazon Kinesis Data Streams. The job fails intermittently with a 'MemoryError' when the stream has a sudden spike in data volume. Which configuration change would best prevent this error?
Medium197A company uses AWS Glue ETL jobs to transform data in Amazon S3. The data arrives in JSON format but needs to be converted to Parquet for efficient querying. Which AWS Glue feature should be used to infer the schema and generate transformation code?
Easy198A company wants to transform data in Amazon S3 using SQL queries without provisioning servers. The transformations are ad-hoc and run occasionally. Which service should be used?
Easy199An e-commerce company wants to capture clickstream data from its website and store it in Amazon S3 for analytics. The data arrives continuously and the company needs near-real-time processing. Which solution is most appropriate?
Easy200Refer to the exhibit. A data engineer is configuring an IAM policy for a Lambda function that writes transformed data to S3. The function writes to both 'example-bucket/data/' and 'example-bucket/public/'. The policy is intended to enforce server-side encryption with SSE-S3 for all objects written to the 'public/' prefix, while allowing all operations on other prefixes. However, the Lambda function is failing with an AccessDenied error when writing to 'example-bucket/public/'. What is the most likely cause?
Hard201A data engineer is designing a data ingestion pipeline for real-time clickstream data. Which TWO services can be used to ingest the data into Amazon Kinesis Data Streams?
Easy202Refer to the exhibit. A data engineer is troubleshooting a Kinesis Data Streams consumer that is falling behind. The stream has 2 shards and is receiving data at a rate of 2 MB/s. The consumer is an AWS Lambda function with a batch size of 100 records. What should the engineer do to improve consumer throughput?
Medium203A data engineer needs to transform JSON data from an S3 bucket into Parquet format and load it into Amazon Redshift. The transformation must be performed incrementally as new data arrives. Which AWS service is BEST suited for this task?
Easy204A company is ingesting streaming data from multiple sources using Amazon Kinesis Data Streams. The data is then processed by an AWS Lambda function that transforms the records and writes them to an Amazon S3 bucket. The Lambda function is failing intermittently with timeout errors. The average record size is 5 KB, and the shard count is 2. What is the MOST likely cause of the timeout errors?
Hard205Refer to the exhibit. A data engineer runs this AWS Glue Data Catalog DDL statement to create a table. The CSV files in 's3://my-bucket/sales/' use a pipe delimiter (|) instead of a comma. What change is needed to correctly read the data?
Easy206A data engineer needs to transform JSON data into CSV format using AWS Glue. The transformation is simple and must be executed on a schedule. Which Glue component is MOST suitable?
Easy207An e-commerce company is building a near-real-time dashboard to monitor customer clickstream data. The data is ingested via Amazon Kinesis Data Streams, transformed using AWS Lambda, and stored in Amazon S3. The team needs to query the data using Amazon Athena. Which THREE steps should be taken to optimize cost and performance? (Choose three.)
Medium208A data engineer must load data from an Amazon S3 bucket into an Amazon Redshift cluster as part of a nightly batch pipeline. The source files are already in Parquet format and include columns that map directly to the target table. The engineer wants the fastest, most cost-effective load method that avoids staging the data through an external service. Which command should the engineer use?
Easy209A small startup is building a data pipeline to ingest customer orders from a web application into Amazon Redshift for analytics. The orders are written to an Amazon RDS MySQL database. The startup wants to replicate the orders to Redshift in near-real time (within 5 minutes) with minimal operational overhead. The data volume is low, averaging 100 new orders per minute. The startup has a single data engineer who is also responsible for other tasks. What is the simplest solution?
Easy210A company uses Amazon Kinesis Data Streams to ingest IoT sensor data. The data is processed by an AWS Lambda function that transforms the records and writes to an Amazon S3 bucket. Recently, the Lambda function has been failing with 'Rate exceeded' errors for the S3 PUT API calls. The data volume is 10 MB/s with average record size 2 KB. What should be done to resolve this issue?
Hard211A company uses AWS Glue to run ETL jobs that transform data from Amazon S3 (Parquet) into a denormalized format for Amazon Redshift. The Glue job uses the DynamicFrame API. The job is failing with a 'MemoryError' when performing a join operation. The data is skewed on the join key. Which THREE actions can reduce memory usage and improve job stability? (Choose THREE.)
Hard212A data engineer needs to transform JSON data from Amazon S3 into Parquet format using AWS Glue. The source files are in a bucket with thousands of small files. What is the best practice to optimize the Glue job performance?
Easy213A data engineer needs to run a daily AWS Glue ETL job that transforms data in Amazon S3. The job must start at 2:00 AM UTC every day. The engineer wants to minimize operational overhead and ensure the job runs reliably. Which approach should the engineer use?
Easy214A company needs to ingest data from an on-premises database to Amazon S3 with minimal impact on the source database. The data volume is several TB. Which AWS service is best suited for this task?
Easy215A data engineer is designing an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structures and write the output to Amazon Redshift. The engineer needs to ensure the job can handle schema evolution and efficiently process only new data on subsequent runs. (Choose two.)
Medium216A healthcare company processes patient records in near-real-time using Amazon Kinesis Data Streams. Each record contains sensitive personal health information (PHI). The data must be encrypted at rest and in transit. The company also needs to audit access to the data. The data engineer is designing the ingestion pipeline. Which combination of services and configurations meets these requirements?
Hard217A data engineer is configuring an AWS Glue ETL job to read data from an Amazon S3 bucket that contains nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer wants to optimize the job for performance and cost. Which two actions should the engineer take? (Choose two.)
Hard218A data engineer needs to ingest on-premises CSV files into Amazon S3 every hour. The files are less than 1 GB each. Which service is the most cost-effective and requires the least operational overhead?
Easy219A company uses Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format. The delivery stream is configured with a buffer size of 5 MB and a buffer interval of 60 seconds. However, the data engineer notices that S3 objects are being created with sizes much smaller than 5 MB. What is a likely cause?
Medium220Refer to the exhibit. An IAM policy for an AWS Lambda function. The Lambda function is triggered by an S3 event (object created) and needs to read from a Kinesis stream. However, the function fails with access denied when trying to read from Kinesis. What is the most likely cause?
Medium221A data engineer is setting up an Amazon Kinesis Data Firehose delivery stream to load data into Amazon Redshift. The data is coming from an application that produces JSON records. The engineer needs to transform the data to match the Redshift table schema. Which approach is the MOST cost-effective and requires the least operational overhead?
Easy222A company is building a data lake on Amazon S3 and wants to ingest data from multiple AWS services (CloudTrail, VPC Flow Logs, and ALB logs). The data should be stored in a central S3 bucket with a common partitioning scheme. Which service can be used to collect and centralize this data with minimal configuration?
Medium223A data engineer needs to transfer 50 TB of historical data from an on-premises HDFS cluster to Amazon S3. The network bandwidth is limited to 100 Mbps. The transfer must be completed within one week. Which service should be used?
Easy224A company uses AWS Glue to catalog data in Amazon S3. The data arrives in Parquet format, but the crawler fails to update the schema when new columns are added. What is the most likely cause?
Medium225A data engineer runs an AWS Glue job that reads Parquet files from Amazon S3 partitioned by year/month/day and writes to another S3 prefix. The job currently processes all historical partitions on every run, causing long runtimes and high cost. The engineer wants subsequent runs to process only new data. Which configuration should the engineer apply?
Hard226A company ingests application logs into Amazon S3 through Amazon Kinesis Data Firehose. The logs arrive as newline-delimited JSON, and analysts query them with Amazon Athena. Query performance is poor because the JSON files are small and uncompressed. The engineer must improve Athena query performance while keeping the raw JSON available for a downstream legacy system. Which change should the engineer make?
Medium227A company uses AWS Glue to process streaming data from Amazon Kinesis Data Streams. The job fails intermittently with a 'MemoryError'. What is the MOST likely cause?
Medium228A data engineer is using AWS Glue to process a large dataset stored in Amazon S3 in Parquet format. The Glue job performs a join between two tables and writes the result back to S3. The engineer notices that the job is running slowly and consuming excessive DPU hours. The job has 10 workers of type G.1X. Which action should the engineer take to improve performance and reduce cost?
Hard229A data engineer needs to ingest data from an Amazon DynamoDB table into an Amazon S3 data lake. The table is updated frequently and the engineer must capture all item-level changes in near real-time without impacting table performance. The ingested data must be stored in a format that preserves the change type (INSERT, MODIFY, REMOVE). Which solution meets these requirements with the LEAST operational overhead?
Medium230A company uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data is transformed using an AWS Lambda function. Some records fail transformation and are lost because the Lambda function throws an exception. The data engineer needs to capture the failed records for analysis without affecting the pipeline. What should the engineer do?
Hard231A company needs to transfer 20 TB of historical data from an on-premises Hadoop cluster to Amazon S3. The network bandwidth is limited and the transfer must complete within one week. Which service should the company use?
Easy232A company uses AWS Glue to transform data in S3. The transformation job reads Parquet files, filters rows, and writes to another S3 bucket. The job takes longer than expected. Which change would MOST likely reduce the job execution time?
Medium233A company uses AWS Glue to process data from Amazon RDS MySQL into Amazon S3. The Glue job uses a JDBC connection and runs on a schedule. Recently, the job has been failing with a 'Communications link failure' error. The RDS instance is in a private subnet. Which troubleshooting step should the data engineer take FIRST?
Hard234A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream and writes results to an Amazon S3 bucket. Recently, the application has been failing with 'ResourceNotFoundException' for the S3 bucket. What is the MOST likely cause?
Medium235A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job needs to run daily and process only new data since the last run. The data in S3 is partitioned by date in the format `year=YYYY/month=MM/day=DD/`. Which feature should the data engineer use to track processed partitions and avoid reprocessing old data?
Medium236Refer to the exhibit. A data engineer runs this CLI command on an S3 bucket. The data is ingested from multiple sources. Which AWS service would be best to process these files in a single batch transformation?
Easy237A data engineer must ingest data from an Amazon DynamoDB table into an Amazon S3 bucket for analytics. The table receives continuous writes, and the ingestion must capture changes with minimal latency. The engineer wants a fully managed solution that requires no server provisioning. Which approach should the engineer use?
Medium238A data engineering team is designing a batch processing workflow using AWS Glue. The job reads from an S3 bucket, transforms data, and writes to another S3 bucket. The job runs daily and processes new data incrementally. Which THREE features should they use to optimize performance and cost?
Hard239A company is designing a data ingestion pipeline for real-time IoT sensor data. The data volume peaks at 10,000 messages per second. The pipeline must process messages in order per sensor and persist raw data to Amazon S3 for archival. Which TWO services should be used together to meet these requirements? (Choose TWO.)
Easy240A data engineer is designing a pipeline to ingest change data capture (CDC) events from an Amazon RDS for PostgreSQL database into Amazon S3. The CDC events are captured using AWS DMS. The data must be available for querying within 5 minutes of the change. Which approach meets these requirements?
Medium241A company uses AWS Glue DataBrew to clean and normalize data from an Amazon S3 bucket before loading it into Amazon Redshift. The data contains PII such as social security numbers. A compliance policy requires that PII be masked in the DataBrew output. Which DataBrew transformation should the engineer use to replace the last four digits of each SSN with asterisks?
Medium242An IAM policy includes the above resource ARN for CloudWatch Logs. A data engineer needs to allow a Lambda function to create log streams and put logs to the log group 'my-log-group'. However, the Lambda function is failing with access denied. What is the issue?
Easy243A data engineer is using AWS Glue to transform data from an Amazon S3 bucket. The Glue job reads JSON files, applies complex transformations, and writes the output to another S3 bucket in Parquet format. The job runs daily and must complete within a 2-hour window. The engineer notices that the job is taking longer than expected and wants to optimize performance. The source data is partitioned by date, and the job uses a dynamic frame. Which optimization should the engineer implement to improve performance?
Hard244Refer to the exhibit. A data engineer is troubleshooting an AWS Glue ETL job that fails with an access denied error when writing to S3. The IAM role attached to the Glue job has the policy shown. What is the most likely cause of the error?
Medium245A data engineer must transform data in an AWS Glue ETL job. The transform requires calling an external REST API for each record to enrich the data. The Glue job runs on AWS Glue 4.0 with Python. The engineer wants to minimize the number of API calls and improve performance. Which approach should the engineer take?
Medium246A data engineer is designing a data ingestion pipeline to load clickstream data from an Amazon S3 bucket into an Amazon Redshift cluster. The data arrives in 5-minute batches. Which TWO actions should the engineer take to ensure data consistency and avoid duplicates? (Select TWO.)
Medium247A company is ingesting log files from multiple EC2 instances into Amazon S3 using the CloudWatch agent. The logs are delivered to a CloudWatch Logs group, and a subscription filter sends them to a Lambda function for transformation, then to Firehose. The Firehose stream is configured with a buffer interval of 60 seconds and buffer size of 5 MB. The logs are critical and must be available in S3 within 5 minutes. What is the most cost-effective way to reduce the delivery latency?
Medium248A company runs a nightly batch ETL job using AWS Glue to transform data from Amazon RDS for MySQL to Amazon S3. The job reads 100 tables and writes Parquet files partitioned by date. Recently, the job started failing with 'ThrottlingException' from the RDS database. The data volume has increased, and the Glue job is reading large tables without any filtering. The job uses a single Glue job with multiple Spark executors. The engineer needs to reduce the load on the RDS database while maintaining the same processing time. What should the engineer do?
Hard249A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 bucket containing nested JSON, flattens arrays using Relationalize, and writes Parquet to another S3 bucket. The job must run only when new objects land in the source bucket. The engineer wants to avoid unnecessary job runs and minimize cost. Which approach should the engineer use?
Hard250A data engineer is using AWS Glue DataBrew to prepare a dataset stored in Amazon S3. The dataset contains missing values, inconsistent date formats, and duplicate rows. The engineer needs to clean the data and produce a transformed output for downstream analytics. (Choose two.)
Medium251A company runs a data ingestion pipeline that uses AWS Glue to read 500 GB of JSON files from an S3 bucket (s3://raw-data/) every hour. The Glue ETL job transforms the data and writes Parquet files to another S3 bucket (s3://processed-data/). The job is triggered by a time-based CloudWatch Events rule. Recently, the job has started taking over 2 hours to complete, causing delays in downstream processes. The data volume has been consistent, and no changes have been made to the job code or infrastructure. The S3 bucket 's3://raw-data/' receives new files continuously, but the Glue job reads all files in the bucket each run (no incremental processing). The engineer suspects that the job is reprocessing old data. Which action should the engineer take FIRST to reduce the job duration?
Hard252A data engineer is designing a data ingestion pipeline to load data from an on-premises Oracle database to Amazon S3. The pipeline should capture changes in near real-time (within minutes) and minimize impact on the source database. The source table has a 'last_modified' timestamp column. Which service combination would meet these requirements?
Medium253A data engineer is designing a data ingestion pipeline for clickstream data from a mobile app. The data volume varies, with occasional spikes up to 10 MB/s. The pipeline must persist the raw data in Amazon S3 and make it available for near-real-time analytics via Amazon Athena. Which combination of services minimizes cost and operational overhead?
Hard254A data engineer needs to ingest streaming data from thousands of IoT devices and immediately process each record with minimal latency. Which AWS service should be used as the ingestion point?
Easy255A company uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data is transformed using an AWS Lambda function. Recently, the transformation errors have increased due to Lambda timeouts. The data engineer needs to diagnose and resolve the issue without losing data. What should the engineer do?
Hard256A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing sensitive customer records. The job must remove all columns flagged as PII before writing to a target S3 location. The PII columns are not known in advance and vary by file. Which approach should the engineer use to ensure the PII is removed dynamically?
Medium257A company needs to ingest data from a MySQL database into Amazon S3 in near real-time. The database is running on EC2. The data engineer wants to minimize the impact on the source database. Which service should be used?
Easy258A data engineer is designing a streaming pipeline that ingests IoT sensor data from 10,000 devices. Each device sends a 1 KB message every second. The data must be processed in near real-time and stored in S3 for analytics. Which combination of services provides the most cost-effective solution?
Hard259A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes only new data added since the last run. The engineer notices that the job is reprocessing old data, increasing runtime and cost. Which AWS Glue feature should be enabled to ensure only new data is processed?
Medium260A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job must run incrementally and process only new data since the last run. The source data is partitioned by date in S3, and new partitions are added daily. Which AWS Glue feature should the engineer enable to track previously processed data?
Medium261A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains a column with inconsistent date formats (for example, '2023-01-15', '01/15/2023', and '15-Jan-2023'). The engineer needs to standardize all values to ISO 8601 format and then write the cleaned data back to S3. Which approach should the engineer use?
Hard262A data engineer is troubleshooting a Lambda function that reads from a Kinesis Data Stream, processes records, and writes to a Kinesis Data Firehose delivery stream. The Firehose delivery stream is configured to deliver data to an S3 bucket. The Lambda function is failing with an access denied error. The IAM policy attached to the Lambda execution role is shown in the exhibit. Which permission is missing?
Hard263A company wants to ingest streaming data from thousands of IoT devices into AWS for real-time processing. Each device sends JSON payloads of about 2 KB at a rate of 1 message per second. The data must be processed with a durable, ordered stream per device. Which service should the company use as the ingestion layer?
Easy264A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data is then processed by an AWS Lambda function that transforms the records and writes them to an Amazon S3 bucket. Recently, the Lambda function has been timing out and the S3 bucket is not receiving all expected data. The Kinesis stream is not throttling and has sufficient shards. Which step should the company take to resolve this issue?
Medium265A CloudFormation template defines an AWS Glue job. The job fails during execution with the error 'Unable to locate script: s3://scripts-bucket/etl-script.py'. The S3 bucket 'scripts-bucket' exists and the script file is present. What is the most likely cause?
Hard266A company uses AWS Data Pipeline to copy data from DynamoDB to S3 daily. Recently, the pipeline started failing with 'ThrottlingException' errors. The DynamoDB table has on-demand capacity. Which action should be taken to resolve the issue?
Medium267A data engineer is designing a data ingestion pipeline for real-time financial transactions. The pipeline must ensure exactly-once processing semantics and must handle duplicate records that may occur due to retries. Which combination of AWS services can achieve exactly-once processing?
Hard268A company needs to ingest data from multiple SaaS sources (e.g., Salesforce, Marketo) into Amazon S3 for analytics. Which AWS service is designed for this purpose?
Easy269A company is building a data lake on Amazon S3 and needs to ingest data from various on-premises sources. Which TWO AWS services can be used to transfer data securely over the internet?
Easy270A company is building a data lake on Amazon S3. They need to ingest data from multiple sources, including relational databases, streaming data, and log files. Which THREE AWS services can be used to ingest data into the data lake?
Medium271A data engineer must load a 50 GB uncompressed CSV file from Amazon S3 into an Amazon Redshift cluster using the COPY command. The load is taking a long time and the engineer wants to improve performance. Which action should the engineer take?
Easy272A company uses AWS Glue to run ETL jobs that process data from Amazon RDS to Amazon S3. The jobs run nightly and take 3 hours to complete. The data volume is growing by 20% each month. The engineer needs to reduce job runtime and cost. The source RDS is a db.r5.large instance. Which approach would be MOST effective?
Hard273A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data must be transformed and stored in Amazon S3 for batch analytics. The engineer wants to use AWS Lambda for transformation. Which TWO configurations are required? (Choose two.)
Medium274A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The recipe must replace all null values in a specific column with the string 'UNKNOWN' and then convert the column to uppercase. The engineer wants to apply these steps in a repeatable recipe. Which combination of DataBrew transforms should be used?
Medium275A data engineer is designing a data ingestion pipeline to load data from an on-premises Oracle database into Amazon Redshift. The pipeline must capture changes (inserts, updates, deletes) with low latency and minimal impact on the source database. Which combination of AWS services should the engineer use?
Medium276A company uses a Kinesis Data Firehose delivery stream to load data into an S3 bucket. The data is in JSON format and must be converted to Parquet before landing in S3. Which steps are required to achieve this? (Choose THREE.)
Hard277A company uses Amazon S3 to store raw data and needs to transform it into Parquet format for analytics. The transformation job runs daily on a schedule. Which AWS service is BEST suited for this task?
Medium278A data engineer runs an AWS Glue job that reads from a JDBC connection to a PostgreSQL database. The job fails with a 'Connection timed out' error. The Glue job runs in a VPC with the appropriate security group. What is the most likely cause?
Medium279Which THREE factors should be considered when choosing between AWS Glue and Amazon EMR for data transformation? (Choose three.)
Hard280A company uses Amazon Kinesis Data Streams to ingest clickstream data from a website. The data is consumed by an AWS Lambda function that writes to Amazon DynamoDB. The Lambda function is seeing high error rates due to DynamoDB write throttling. Which action should be taken to reduce throttling?
Medium281A company is using AWS Glue to run ETL jobs that transform data from Amazon DynamoDB to Amazon S3. The DynamoDB table has a large number of items (over 10 million) and is heavily used by production applications. The Glue job reads the entire DynamoDB table each time it runs, causing increased read capacity consumption and affecting production performance. The team wants to reduce the impact on the source DynamoDB table while still keeping the S3 data up-to-date. What should the team do?
Easy282A data engineer is troubleshooting a Kinesis Data Firehose delivery stream that ingests JSON log data from web servers. The stream is configured to transform records with an AWS Lambda function and deliver to an Amazon S3 bucket. Recently, the stream has been failing with 'InvalidData' errors. Which action should the engineer take to resolve the issue?
Easy283A data engineer is using AWS Glue to run an ETL job that reads from Amazon S3, performs a join between two large datasets, and writes the result to Amazon Redshift. The job is taking longer than expected, and the engineer suspects data skew. Which technique can help mitigate data skew in the join?
Hard284A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job must handle upserts (inserts and updates) into an existing Redshift table based on a primary key. The engineer needs to ensure that the job efficiently processes only changed records and minimizes data movement. Which two AWS Glue features or techniques should be used to achieve this? (Choose two.)
Hard285A company is designing a data ingestion pipeline for real-time sensor data from thousands of devices. The data must be processed with low latency and stored in Amazon S3. Which TWO services would be appropriate for this use case? (Choose TWO.)
Medium286A data engineer needs to design a data ingestion pipeline that ingests CSV files from an Amazon S3 bucket, transforms the data by adding a timestamp column, and loads it into an Amazon Redshift table. The pipeline should run automatically whenever a new file is uploaded to the S3 bucket. Which AWS service should be used to trigger the transformation?
Medium287A data engineer is using AWS Glue to process a large dataset stored in Amazon S3. The dataset is partitioned by year/month/day and consists of Parquet files. The engineer notices that the Glue job is running slowly and consuming excessive DPU hours. The job performs a join between two large tables and writes the output back to S3. Which optimization technique should the engineer implement to improve performance and reduce cost?
Hard288The exhibit shows an IAM policy attached to a role used by an AWS Glue ETL job. The job reads from an S3 bucket and writes to another S3 bucket. However, the job fails with an access denied error when trying to write to the output bucket. What is the most likely cause?
Hard289A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is in Parquet format and is partitioned by date. The engineer wants to load only the latest partition into Redshift and ensure the load is efficient. Which method should the engineer use?
Medium290A company runs an AWS Glue ETL job that reads data from Amazon S3, transforms it, and writes back to S3 in a different partition structure. The job uses the 'spark.sql.shuffle.partitions' option set to 200. After the job completes, the output has many small files. The data engineer wants to minimize the number of output files while maintaining job performance. Which action should the engineer take?
Hard291A data engineer is troubleshooting an AWS Glue job that reads from Amazon RDS MySQL and writes to Amazon S3. The job runs successfully but takes longer than expected. The engineer wants to optimize performance. Which THREE actions would improve job performance?
Hard292Which TWO AWS services can be used to ingest streaming data from a mobile application into Amazon S3 for near-real-time analytics? (Choose 2.)
Medium293A data engineer is designing a streaming ingestion pipeline using Amazon Kinesis Data Streams. The stream receives records from thousands of IoT devices, and the engineer must ensure that records from the same device are processed in order. The engineer also needs to scale the stream to handle peak loads without manual intervention. Which two actions should the engineer take? (Choose two.)
Hard294A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon S3 in Parquet format. The source data is in JSON format and contains nested structures. The engineer needs to flatten the nested data and write it to Parquet. Which AWS Glue transform should the engineer use to flatten the nested structure?
Medium295A data engineer runs an AWS Glue ETL job that joins a 4 TB Parquet dataset in Amazon S3 with a small 40 MB reference lookup table stored as a single CSV file in S3. The join is taking hours and the job frequently fails with executor out-of-memory errors. The engineer wants to reduce shuffle and memory pressure with the least development effort. Which approach should the engineer take?
Medium296Refer to the exhibit. An S3 event notification is configured to trigger an AWS Lambda function when objects are created in 'my-bucket'. The Lambda function processes the JSON file and writes results to Amazon DynamoDB. The function fails with a timeout error. Which action should the engineer take to resolve the issue?
Easy297A data engineer is designing a pipeline to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real-time with minimal latency, and the engineer wants to use a fully managed service that automatically scales. The data is in JSON format and needs to be converted to Parquet before storage. Which AWS service should be used to achieve this?
Medium298A data engineer needs to ingest JSON files from an S3 bucket into a DynamoDB table. The files are updated hourly and contain new records. Which AWS service should be used to trigger a Lambda function for each new object?
Easy299A data engineer is troubleshooting an AWS Glue ETL job that fails with the error: 'An error occurred while calling o137.pyWriteDynamicFrame. No such file or directory: s3://bucket/output/part-00000.parquet'. The job reads from a JDBC source and writes to S3. What is the most likely cause?
Easy300A data engineer is using AWS Glue Studio to create an ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The source data is in JSON format and contains nested structures. The engineer needs to flatten the nested structures and write the output in Parquet format. The job must be efficient and scalable. Which transformation should the engineer use in Glue Studio to flatten the nested data?
Hard301A data engineer is using AWS Database Migration Service (AWS DMS) to migrate a 4 TB on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must complete in a single maintenance window, and the source database cannot be taken offline for more than 30 minutes. The engineer configures a full load plus change data capture (CDC) task. During testing, the full load phase takes 14 hours. Which configuration change will most effectively reduce the time required for the full load phase?
Medium302A data engineer needs to ingest data from a SaaS application (Salesforce) into Amazon S3 on a daily basis. Which TWO AWS services can be used for this purpose? (Choose TWO.)
Easy303Refer to the exhibit. A data engineer has configured an S3 event notification to send an event to an SQS queue when objects are created in the 'incoming/' prefix. The engineer wants to trigger an AWS Lambda function to process the object. However, the Lambda function is not being invoked. What is the most likely cause?
Hard304A data engineer maintains an AWS Glue ETL job that reads JSON from Amazon S3 and writes Parquet to a second bucket. Downstream consumers report that numeric fields occasionally arrive as strings and timestamps are sometimes null. The engineer must make the job resilient to these schema variations without failing the run. Which approach should the engineer take?
Hard305A data engineer is using AWS Glue Studio to build a job that reads from an Amazon Kinesis Data Stream, performs a 5-minute tumbling window aggregation, and writes results to Amazon S3. The job must run continuously and handle late-arriving records within the window. Which configuration should the engineer use?
Hard306A data engineer needs to run a Python-based transformation on each object as it lands in an Amazon S3 bucket. The objects are small (under 10 MB), arrive sporadically, and must be processed within seconds. Which approach is MOST appropriate?
Easy307A data engineer runs a weekly AWS Glue ETL job that processes data from Amazon DynamoDB to Amazon S3. The job reads the entire table every time, which is slow and expensive. The job needs to process only items that changed since the last run. Which solution should the engineer implement?
Hard308A data engineer needs to transform JSON data into Parquet format using AWS Glue. The input data has nested fields. Which Glue feature should be used to flatten the nested structure?
Medium309A data engineer is using AWS Database Migration Service (AWS DMS) to migrate a large on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must minimize downtime, so the engineer needs to capture ongoing changes while the initial full load runs. Which AWS DMS task configuration should the engineer use?
Medium310A data engineer needs to schedule an AWS Glue ETL job to run every hour and process new data that arrives in an S3 bucket. The job should only process files that have been added since the last run. Which approach should the engineer use to track which files have been processed?
Easy311A company wants to schedule a nightly batch job to copy data from an on-premises PostgreSQL database to Amazon S3. The solution must minimize operational overhead. Which AWS service should be used?
Easy312A data engineer is troubleshooting an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The job runs successfully but 5% of records are missing after the load. The engineer suspects data consistency issues. Which THREE actions could help diagnose and resolve the problem? (Choose THREE.)
Hard313A data engineer is using AWS Glue ETL to transform a large dataset in S3. The job processes 2 TB of data daily and currently runs for 6 hours. The engineer wants to reduce runtime without changing the transformation logic. What is the best approach?
Medium314A data engineer needs to ingest data from an on-premises Oracle database into Amazon S3. The data volume is about 500 GB initially, with daily incremental updates of 10 GB. The pipeline must minimize operational overhead. Which AWS service should be used for the initial and incremental loads?
Easy315Which THREE factors should be considered when choosing between Amazon Kinesis Data Streams and Amazon Kinesis Data Firehose for real-time data ingestion? (Choose three.)
Hard316A company wants to ingest real-time streaming data from thousands of IoT devices into AWS for immediate processing. Which service is designed for ingesting large volumes of streaming data with low latency?
Easy317A company uses AWS Glue to run ETL jobs that process data from Amazon S3 and load into Amazon Redshift. The jobs have recently started failing with 'Out of Memory' errors. The data volume has increased 3x in the past month. Which is the MOST effective solution to resolve this issue without redesigning the job?
Hard318A data engineer needs to load data from an on-premises Oracle database to Amazon S3 daily. The table is 500 GB and grows by 50 MB per day. The load must capture only new and changed rows since the last run. Which solution is MOST cost-effective and requires the least maintenance?
Medium319Refer to the exhibit. A CloudFormation stack outputs the Glue job name and S3 bucket names. The Glue job transforms CSV files from the raw bucket to Parquet in the processed bucket. However, the Glue job is failing with an error that it cannot write to the processed bucket. What is the most likely cause?
Hard320A data engineer is building an AWS Glue ETL job that reads records from an Amazon Kinesis Data Stream and writes them to Amazon S3 in Parquet format. The job must checkpoint its progress so that it can resume without reprocessing data after a failure. Which AWS Glue mechanism should the engineer configure to track the stream position?
Medium321A company ingests JSON data from an S3 bucket into a Glue ETL job. The data contains nested structures and arrays. The team wants to flatten the data into a tabular format for analysis in Athena. Which Glue transformation is appropriate?
Hard322A company runs a daily batch ETL job using AWS Glue. The job processes 500 GB of data from Amazon RDS to Amazon S3. The job currently uses a single DPU and takes 6 hours to complete. The team wants to reduce runtime to under 1 hour without increasing costs significantly. Which approach should they use?
Hard323A company uses AWS DMS to replicate data from an Amazon RDS for MySQL database to Amazon S3. Which TWO configurations are required to enable continuous change data capture (CDC) from MySQL?
Hard324A data engineer is configuring an AWS Glue crawler against an Amazon S3 path that contains CSV files with inconsistent column counts across files. The crawler keeps creating multiple tables for the same data and the engineer wants a single table with a merged schema. Which crawler configuration should the engineer change?
Medium325A company needs to transform JSON data from an S3 bucket into a structured format for Amazon Redshift. The transformation should be done serverlessly. Which service should be used?
Easy326A data engineer maintains an AWS Glue job that incrementally processes new files in Amazon S3 using job bookmarks. After a schema change in the source data added a new column, the engineer updated the Glue Data Catalog table. Subsequent job runs still process only previously seen files and ignore newly arrived objects. The engineer verifies that new files exist in the prefix and that the bookmark state was not reset. Which factor most likely explains why new files are being skipped?
Hard327A company is using Amazon Kinesis Data Streams to ingest real-time clickstream data from a website. The data is consumed by an Amazon Kinesis Data Analytics for Apache Flink application that performs real-time analytics. The Flink application writes its results to an Amazon S3 bucket. The company has noticed that the Flink application is experiencing high checkpoint failure rates, causing delays. The CloudWatch metrics show that the checkpoint size is large and increasing. The data engineer needs to reduce the checkpoint size. Which action should the data engineer take?
Medium328A data engineer runs an AWS Glue Studio job that reads JSON from Amazon S3 and writes to a partitioned Parquet table. The job currently runs for six hours. Profiling shows that a small number of partitions contain millions of rows while most contain a few hundred. Which change will MOST improve runtime?
Hard329A data engineer is building an AWS Glue Studio visual ETL job that reads JSON files from Amazon S3, applies a transformation, and writes to Amazon Redshift. During a test run, the job fails with an error indicating that the dynamic frame could not be written because the target table schema does not match the incoming data. The engineer needs to ensure the job automatically reconciles schema differences such as missing columns and data type mismatches during the write. Which action should the engineer take?
Medium330A data engineer is building an AWS Glue ETL job that reads from an AWS Glue Data Catalog table backed by Amazon S3. The job must process only records added since the last successful run to reduce cost and runtime. The source data is partitioned by year, month, and day. Which approach should the engineer use?
Hard331A data engineer is designing an ingestion pipeline that uses AWS Glue to read from an Amazon RDS for PostgreSQL database. The job must read only rows changed since the previous run and must not scan the entire table each night. The source table has a last_updated timestamp column that is updated on every write. (Choose two.)
Medium332A company is using AWS Glue to run ETL jobs that transform data from S3 to Redshift. The jobs are failing intermittently with out-of-memory errors. Which THREE actions can help resolve this issue? (Choose THREE.)
Medium333A company has a 100 TB dataset stored on-premises in a Hadoop cluster. They want to ingest this data into Amazon S3 for processing with AWS Glue. The company has a limited time window and a slow internet connection. Which strategy is MOST appropriate?
Hard334A company uses AWS Glue to process data from multiple sources. The data is stored in an Amazon S3 data lake. The company needs to transform the data using a custom Python library that is not available in the default Glue environment. What is the MOST efficient way to make this library available to the Glue jobs?
Medium335A data engineering team needs to transform CSV files stored in Amazon S3 into Parquet format using AWS Glue. The files are partitioned by date and are updated hourly. Which AWS Glue feature should be used to automatically detect the schema and partition structure?
Easy336A data pipeline ingests JSON data from an S3 bucket using AWS Glue. The JSON files contain nested structures, and the team wants to flatten them for analysis in Amazon Athena. Which Glue transformation is most appropriate?
Hard337A company needs to ingest streaming data from thousands of IoT devices into Amazon S3 for long-term storage and analytics. The data arrives continuously at a rate of 5 MB per second and must be stored in a compressed format to reduce storage costs. The solution should be highly available and require minimal management. Which AWS service should the company use?
Easy338A data engineer needs to ingest streaming data from thousands of devices sending JSON messages via HTTP POST. The data should be stored in Amazon S3 with minimal latency and also be available for real-time analytics. Which combination of services is MOST appropriate?
Medium339A company uses AWS Glue to transform data stored in S3. The Glue job runs daily and processes data in the range of hundreds of GB. The data engineer wants to optimize the job for cost and performance. Which THREE actions should be taken? (Choose THREE.)
Hard340Which TWO options are valid methods to ingest on-premises relational database data into Amazon S3 for analytics? (Choose 2.)
Medium341The exhibit shows an AWS CLI command and its output. A data engineer wants to copy only objects larger than 10 MB from the S3 bucket to another bucket for processing. Which approach should be used to automate this task?
Medium342A data engineer needs to ingest streaming data from thousands of IoT devices into AWS for real-time processing. The data volume peaks at 5 GB/min. Which AWS service should be used as the ingestion endpoint?
Easy343A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON and writes to a partitioned Parquet table. The engineer wants to reduce job cost and improve read performance. (Choose two.)
Medium344A company is building a data lake on Amazon S3 and needs to ingest data from multiple sources. Which of the following AWS services can be used to ingest and transform data in near real-time? (Select TWO.)
Medium345A data engineer needs to run a transformation on streaming data using SQL-like queries without managing servers, and the output must be written to Amazon S3 in near real time. The source is an Amazon Kinesis Data Stream. Which AWS service is the MOST appropriate to perform the transformation?
Easy346A company wants to ingest streaming data from Apache Kafka into Amazon S3 for long-term storage and analytics. The data is in JSON format and must be delivered to S3 with minimal effort and no custom code. Which AWS service should the data engineer use?
Easy347A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data is then processed by a Kinesis Data Analytics application running SQL queries. The analytics application is falling behind and processing records with increasing latency. The stream has 4 shards, and the average record size is 5 KB. What is the MOST effective way to improve processing latency?
Hard348A data engineer needs to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket for long-term storage. The data is in JSON format and must be delivered within 60 seconds of arrival. The engineer wants a fully managed solution that requires minimal code. Which service should the engineer use?
Easy349A data engineer is using AWS Glue to process a large dataset where a small number of partitions contain disproportionately more rows than others, causing some executors to run much longer than others and the job to take hours. The engineer wants to redistribute the data across partitions before a join operation to improve performance. Which technique should the engineer apply?
Hard350A company is ingesting real-time clickstream data into Amazon S3 using Amazon Kinesis Data Firehose. The data is semi-structured and the company wants to transform the data into Parquet format and partition it by year, month, day, and hour. Which TWO steps should be taken to achieve this? (Choose TWO.)
Medium351A data engineer runs an AWS Glue job that writes Parquet files to Amazon S3. The job frequently fails with an error indicating too many small files are being written, causing slow downstream Athena queries. The engineer wants to reduce the number of output files without changing the transformation logic. Which action should the engineer take?
Hard352Which THREE factors should a data engineer consider when choosing between AWS Glue and Amazon EMR for a data transformation job? (Choose three.)
Hard353A data engineer needs to run a transformation in AWS Glue where each record must be processed independently and the output schema is known ahead of time. The transformation should operate on a DynamicFrame and return a DynamicFrame. Which Glue transform is designed for this row-by-row operation?
Easy354A data engineer is using AWS Glue Studio to create a job that joins data from two Amazon S3 sources: a large fact table and a small dimension table. The job performs a join and then writes the result to Amazon S3 in Parquet format. The engineer notices that the job is running slowly and consuming many DPUs. Which optimization technique should the engineer apply to improve performance?
Hard355A company uses AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration is taking longer than expected. The task status shows 'Full load in progress' with a low 'Table throughput (rows/s)'. Which action would MOST improve throughput?
Hard356A company uses AWS Glue ETL to transform data from Amazon RDS for MySQL to Amazon S3. The Glue job reads from a JDBC connection. The job runs once daily and processes all records, but the data volume is growing. Which change would improve performance and reduce costs?
Medium357A data engineer needs to run a transformation on a large dataset stored in Amazon S3 using AWS Glue Studio. The transformation is a simple column rename and filter that can be expressed visually. The engineer wants to minimize development time and avoid writing PySpark code. Which approach should the engineer use?
Easy358A company needs to ingest real-time clickstream data from a web application into Amazon S3 for analytics. The data must be available within minutes of generation. Which AWS service should be used to capture and deliver this streaming data?
Easy359A data engineer runs an AWS Glue ETL job that reads semi-structured JSON from Amazon S3, flattens nested arrays, and writes Parquet to a partitioned S3 location. Job runs are becoming expensive because Glue reprocesses all historical partitions on every run. The engineer wants subsequent runs to process only newly arrived data. Which approach should the engineer take with the LEAST operational overhead?
Medium360A data engineer is designing a streaming ingestion pipeline using Amazon Kinesis Data Streams. The stream has 10 shards, and the data volume is expected to grow by 50% over the next month. The engineer needs to ensure that the pipeline can scale without manual intervention. Which approach should be used?
Hard361A company wants to migrate on-premises data to Amazon S3 using AWS DataSync. The data is stored on an NFS file server and the total volume is 50 TB. The network bandwidth between the on-premises data center and AWS is 1 Gbps (gigabit per second). What is the primary factor that will determine the total time required for the initial data transfer?
Easy362A data engineer is building an AWS Glue ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3. The DynamoDB table has a large number of items, and the engineer needs to ensure the job reads the data efficiently without consuming too much provisioned throughput. Which method should the engineer use to read from DynamoDB?
Medium363A data engineer needs to ingest log files from multiple EC2 instances into Amazon S3. The logs are written to local disk on each instance. The engineer wants a simple agent-based solution that can collect, compress, and upload logs to S3 with minimal configuration. The solution must support incremental uploads (only new log lines) and handle log rotation. What should the engineer use?
Easy364A data engineer is using AWS Glue DataBrew to profile a dataset stored in Amazon S3. The profile shows that a column named country contains values such as 'US', 'usa', 'United States', and 'U.S.A.' The engineer needs to standardize these values to a single canonical form before loading the data into Amazon Redshift. Which DataBrew transformation should the engineer apply?
Medium365A data engineer needs to ingest streaming data from an IoT fleet into Amazon S3 for near-real-time analytics. The data volume is approximately 5 GB per hour, and each event is less than 1 KB. Which AWS service should be used as the ingestion endpoint?
Easy366A data engineer needs to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real time, and the engineer wants to minimize operational overhead by using a fully managed service that can also transform the data format from JSON to Parquet. Which AWS service should be used?
Easy367A company uses AWS Database Migration Service (DMS) to continuously replicate data from an on-premises Oracle database to Amazon S3. The data is stored as CSV files. The downstream team requires the data to be in Apache Parquet format. Which change should the data engineer make to the DMS task?
Easy368A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 source with many small JSON files and writes to Amazon S3 in Parquet. The job runs slowly and produces many tiny output files. The engineer wants to improve throughput and reduce the number of output files without changing the source data layout. (Choose two.)
Hard369A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift table on a daily schedule. The data is in CSV format and the schema matches. Which service is simplest for this batch ingestion?
Easy370A data engineer needs to transfer 50 TB of historical data from an on-premises HDFS cluster to Amazon S3. The on-premises network has a 1 Gbps link to AWS. The transfer must complete within 5 days. Which solution is MOST cost-effective and meets the requirements?
Medium371A data engineer needs to ingest streaming data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real-time and stored in Parquet format for efficient querying. The engineer wants to minimize custom code. Which solution should the engineer use?
Easy372A data engineer maintains an AWS Glue ETL job that reads JSON files from Amazon S3, applies transformations using a DynamicFrame, and writes Parquet to another S3 bucket. The job currently reads the entire source prefix on every run, but the source is now partitioned by year/month/day and only new partitions need processing. The engineer wants to process only new partitions without changing the job's transformation logic. Which change should the engineer make?
Hard373A data engineer must load data from an Amazon DynamoDB table into an Amazon S3 data lake nightly. The table is approximately 800 GB and the nightly window is tight. The engineer wants a fully managed, serverless option that exports the table to S3 without consuming DynamoDB read capacity or writing custom code. Which solution meets these requirements?
Easy374A data engineer is building an AWS Glue ETL job that reads a large Amazon S3 dataset of nested JSON files and must flatten the nested arrays into separate rows for downstream analytics. The engineer needs the most efficient, code-free way to apply this transformation within the Glue job. Which approach should the engineer use?
Medium375A data engineering team is troubleshooting a slow AWS Glue ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3 in Parquet format. The job processes 50 GB of data. Which action would most effectively improve job performance?
Hard376A company wants to ingest streaming data from IoT devices into Amazon S3 using Amazon Kinesis Data Firehose. The data must be transformed from JSON to Parquet format before landing in S3. What is the SIMPLEST way to achieve this?
Easy377A data engineer is using AWS Glue Studio to build an ETL job that reads semi-structured JSON from Amazon S3. The source files contain nested arrays and inconsistent keys, and the engineer wants the job to automatically infer the schema at runtime without a Data Catalog table. Which transform or configuration should the engineer use to read the data most reliably?
Medium378A data engineer needs to run a one-time AWS Glue ETL job that reads data from an Amazon S3 bucket and writes transformed Parquet files to another S3 bucket. The job does not need a schedule and should be run immediately after creation. Which method should the engineer use to start the job?
Easy379A company needs to ingest data from multiple SaaS applications (e.g., Salesforce, Marketo) into Amazon S3 for analytics. The data sources have different schemas and update frequencies. Which AWS service should be used to build this ingestion pipeline with minimal code?
Medium380A company wants to ingest data from multiple SaaS applications into Amazon S3 using a fully managed service that supports schema discovery and transformation. Which AWS service should they use?
Easy381A data engineer needs to move 50 TB of existing data from an on-premises data center into Amazon S3 as a one-time migration. The data center has a 1 Gbps internet connection that is shared with production traffic, and the migration must complete within two weeks without disrupting production. Which approach should the engineer use?
Easy382A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job uses the write_dynamic_frame.from_jdbc_conf method. The engineer notices that the job is slow and sometimes fails due to connection timeouts. Which action should the engineer take to improve performance and reliability?
Hard383A company uses AWS Glue ETL jobs to transform data stored in Amazon S3. The job reads data in Parquet format, applies transformations, and writes the output back to S3 in Parquet format. The team wants to improve the job's performance and reduce costs. Which action is MOST effective?
Easy384A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes transformed records to an Amazon Redshift cluster. The job must load data in parallel and use an Amazon Redshift IAM role for authentication. Which connection option should the engineer configure in the Glue job to enable parallel loading and IAM-based authentication?
Medium385Refer to the exhibit. A data engineer runs an AWS Glue ETL job that writes output to an S3 bucket. The job fails with the error shown. What is the most likely cause?
Hard386A data engineer is using AWS Glue Studio to build a job that reads from Amazon S3, applies a filter, and writes to Amazon Redshift. The engineer needs to ensure the job can be rerun safely without creating duplicate rows in Redshift if a previous run partially succeeded. Which design choice best meets this requirement?
Hard387A company wants to ingest streaming data from thousands of IoT devices into AWS for real-time analytics. Which AWS service is best suited for this purpose?
Easy388A data engineer needs to ingest data from a relational database (MySQL) into Amazon S3 for analytics. The database is 500 GB and the job must run daily with incremental updates. Which AWS service is BEST suited for this task?
Easy389A data engineer runs an AWS Glue job that reads a large partitioned Parquet dataset from Amazon S3 and writes aggregated results to another S3 prefix. The job runs daily and currently reprocesses the entire dataset each time, which is becoming expensive. The engineer wants subsequent runs to process only data added since the last successful run, based on the job's state. Which feature should the engineer enable?
Hard390A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data must be transformed in real-time using custom Python code before being stored in Amazon S3. Which AWS service should be used to perform this transformation?
Medium391A company needs to ingest data from multiple SaaS applications (Salesforce, Marketo) into Amazon S3 for analytics. The data volume is moderate (~100 GB per day). The pipeline must handle schema changes, deduplicate records, and provide low latency (under 1 hour). Which THREE services should be used? (Choose THREE.)
Hard392A company is migrating a legacy on-premises ETL pipeline to AWS. The pipeline processes daily batch files from an FTP server. The data must be transformed using complex business logic before being loaded into Amazon Redshift. Which THREE AWS services should be used for this migration?
Hard393A data engineer needs to ingest data from multiple SaaS applications (Salesforce, Marketo) into Amazon S3 for a data lake. The data volumes are moderate and the sync needs to be scheduled daily. Which AWS service is most appropriate for this task?
Easy394Match each AWS data migration tool to its primary function.
Medium395A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files and must flatten the nested structures before writing to Amazon Redshift. The job uses the Glue DynamicFrame API. The engineer wants the transformation to run as a single pass without an intermediate shuffle. Which operation should the engineer use?
Hard396A data engineer runs an AWS Glue for Apache Spark job that writes partitioned Parquet output to Amazon S3. The engineer notices thousands of small output files in each partition, which slows downstream Amazon Athena queries. The job reads from a large S3 source and uses default partitioning. Which change should the engineer make to reduce the number of small files?
Medium397A company is streaming IoT sensor data to Amazon Kinesis Data Streams. The data is JSON with a schema that changes occasionally. They want to load the data into Amazon S3 in Parquet format partitioned by date and sensor_id. Which approach is MOST cost-effective and operationally efficient?
Medium398Refer to the exhibit. A data engineer is troubleshooting an AWS Lambda function that reads from an S3 bucket and writes to a Kinesis Data Stream. The Lambda function fails with an AccessDeniedException when calling the kinesis:PutRecords API. Which change is needed to the IAM policy?
Medium399A company wants to use AWS Glue to transform data stored in Amazon S3. The data is partitioned by date and includes both CSV and Parquet files. The transformation should be optimized for cost and performance. Which THREE actions should the data engineer take? (Choose THREE.)
Medium400A media company is building a data pipeline to ingest user activity logs from multiple sources into Amazon S3. The logs are JSON files generated every minute. The company wants to use Amazon Athena to query the logs with minimal latency and cost. The current approach is to use AWS Kinesis Data Firehose to deliver the logs to S3 with a prefix like 'logs/2024/01/01/00/file.json'. However, when running Athena queries, the team notices high query costs because Athena scans all files in the 'logs/' prefix even when querying for a specific date. What should the team do to reduce the amount of data scanned by Athena?
Easy401A company runs a nightly ETL job using AWS Glue. The job reads data from a JDBC connection to an on-premises MySQL database. The job fails with an error indicating that the connection pool is exhausted. What is the most likely cause and solution?
Medium402A company uses AWS Glue ETL jobs to transform data in S3. The job runs successfully but takes longer than expected. The data is in Parquet format and partitioned by date. Which change would most improve performance without increasing cost?
Medium403A data engineer needs to catalog a growing S3 data lake. New CSV files land in s3://analytics/raw/orders/ with a partition structure year=YYYY/month=MM/day=DD/. The engineer must create an AWS Glue Data Catalog table that automatically recognizes these partitions and requires no crawler runs for future dates. Which approach meets these requirements?
Medium404A data engineer is designing an ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver streaming records into an Amazon S3 bucket. The records arrive as JSON, and downstream consumers require Parquet with a stable schema. The engineer must configure the Firehose delivery stream so records are converted to Parquet before landing in S3. (Choose two.)
Medium405A data engineer needs to transform JSON data from Amazon S3 into Parquet using AWS Glue. The JSON is nested and contains arrays. The engineer wants to flatten the nested structure and write the result to S3 partitioned by a 'region' field. Which combination of Glue transforms should the engineer use?
Medium406A data engineer is troubleshooting a Lambda function that reads from the Kinesis stream 'my-data-stream'. The Lambda function is able to read data but occasionally fails with 'KMS.AccessDeniedException'. What is the most likely cause?
Medium407A data engineer must load a 2 GB uncompressed CSV file from Amazon S3 into Amazon Redshift using the COPY command. The cluster is a two-node ra3.xlplus cluster, and the load is running far slower than expected. The engineer wants the fastest reliable improvement without changing the cluster. What should the engineer do?
Easy408A data engineer is using AWS Glue DataBrew to profile a dataset in Amazon S3. The dataset contains a column 'customer_id' that should be unique. The engineer runs a profile job and notices that the 'customer_id' column has a uniqueness metric of 98%. The engineer needs to identify the duplicate values. Which DataBrew feature should the engineer use to display the duplicate values?
Hard409A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer notices that some JSON records have missing fields and inconsistent schemas. Which AWS Glue feature should be used to handle these inconsistencies and ensure the job does not fail?
Hard410A company is ingesting streaming data into Kinesis Data Streams. The consumer application experiences high latency due to a single shard bottleneck. What is the most effective way to reduce latency?
Medium411A company is ingesting streaming data from social media feeds using Amazon Kinesis Data Streams. The data volume peaks at 10,000 records per second, and each record is up to 1 KB. The company needs to archive the raw data in Amazon S3 in near real-time and also make it available for real-time analytics using Amazon Kinesis Data Analytics. What is the MOST efficient architecture to meet these requirements?
Hard412A data pipeline ingests streaming data from thousands of IoT devices into Kinesis Data Streams. The data must be transformed using a simple field mapping before being stored in S3. Which service should be used to perform the transformation with minimal operational overhead?
Easy413A company wants to ingest data from SaaS applications (e.g., Salesforce, Marketo) into Amazon S3 for analytics. The data volume is moderate and updates occur frequently. Which AWS service is BEST suited for this task?
Medium414A data pipeline ingests daily CSV files from an FTP server into an Amazon S3 bucket. The files must be converted to Parquet format and partitioned by date for efficient querying using Amazon Athena. Which AWS service is most suitable for this transformation?
Easy415A data engineer is designing a pipeline to ingest data from an Amazon RDS for PostgreSQL database into Amazon S3 using AWS Database Migration Service (AWS DMS). The source database has a high volume of transactions and the engineer needs to capture ongoing changes with minimal impact on the source. The target S3 bucket must store the data in Parquet format for querying with Amazon Athena. Which two actions should the engineer take to meet these requirements? (Choose two.)
Medium416A data engineer is building a pipeline to ingest data from an on-premises Oracle database into Amazon S3. The pipeline must capture change data (CDC) in near real-time and handle schema changes. Which TWO AWS services should the engineer use?
Hard417A data engineer is using AWS Glue to transform data from Amazon S3. The source data is in CSV format with inconsistent date formats across files (e.g., 'MM/DD/YYYY' and 'YYYY-MM-DD'). The engineer needs to standardize all dates to 'YYYY-MM-DD' format in the output. Which AWS Glue transform should the engineer use to achieve this?
Medium418A company has CSV files in an S3 bucket that need to be converted to Parquet and loaded into a Redshift table daily. The transformation is a simple schema mapping without joins. Which AWS Glue feature is BEST suited for this task?
Easy419A company uses AWS Glue to process data in Amazon S3. The Glue job fails with an error indicating that the partition keys in the catalog do not match the actual S3 partition structure. What is the most likely cause?
Medium420A company is using AWS DMS to replicate data from an on-premises Oracle database to Amazon RDS for MySQL. The replication is working, but the target table has a different schema. Which DMS feature should be used to transform the source schema to match the target?
Hard421A data engineer must load a 500 MB CSV file from Amazon S3 into an existing Amazon Redshift table once per day. The file has a header row and uses a pipe delimiter. The engineer wants the fastest load and the least operational overhead. Which approach should the engineer use?
Easy422A company is using AWS Glue to run ETL jobs that transform data from Amazon S3 to Amazon Redshift. The jobs are failing intermittently with 'Out of Memory' errors. The team wants to resolve this issue without increasing costs significantly. Which TWO actions should the team take?
Medium423A company uses Amazon EMR to process large datasets stored in Amazon S3. The data is in Parquet format and partitioned by date. The EMR cluster uses Spark SQL for transformations. Recently, the job has been slow and some tasks are failing due to 'java.lang.OutOfMemoryError'. The cluster has 10 core nodes of type m5.xlarge. Which configuration change would MOST improve performance and stability?
Hard424A data engineering team uses AWS Glue ETL jobs to process data daily. They notice that job run times are increasing as data volume grows. Which action will most effectively improve performance without changing the code?
Medium425A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket with millions of small JSON files. The job is running slowly and often fails with out-of-memory errors. The engineer needs to improve performance and reliability. What should the engineer do?
Hard426A data engineer is designing an AWS Glue ETL job that reads data from an Amazon S3 bucket containing many small JSON files and writes the output to Amazon S3 in Parquet format. The engineer wants to improve read performance and reduce the number of output files. Which two actions should the engineer take? (Choose two.)
Hard427A company uses AWS Glue to run ETL jobs daily. The data is stored in S3 as Parquet files partitioned by date. Recently, jobs have failed with the error 'No such file or directory' for certain partitions. What is the MOST likely cause?
Easy428A data engineer is designing a data ingestion pipeline for real-time clickstream data from a website. The data must be ingested with low latency (seconds) and made available for multiple consumer applications, including a dashboard that refreshes every minute and a machine learning model that processes data in near-real-time. The engineer needs to choose a streaming ingestion service. Which TWO services meet these requirements? (Select TWO.)
Medium429A company uses Amazon Kinesis Data Streams to ingest real-time logs from thousands of applications. The data must be transformed and enriched with reference data from Amazon S3 before being stored in Amazon S3 in Parquet format. The transformation logic is stateful and requires exactly-once processing. Which AWS service should the data engineer use to perform the transformation?
Medium430A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an Amazon RDS for PostgreSQL database to an Amazon S3 bucket. The source table has a primary key and the engineer needs near-real-time change data capture (CDC) with minimal impact on the source. Which DMS task setting should the engineer configure to meet these requirements?
Hard431A data engineer needs to ingest data from an Amazon S3 bucket into Amazon Redshift for analytics. The data is in CSV format and the Redshift table already exists. Which service can be used to perform this ingestion with minimal configuration?
Easy432A data engineer is building a data pipeline that ingests data from Amazon S3 into Amazon Redshift. The data is in CSV format and includes a timestamp column. The pipeline should load only new data incrementally. Which approach is most efficient?
Medium433A data engineer is configuring an AWS Glue ETL job that reads semi-structured JSON event logs from Amazon S3 and must flatten nested arrays into relational columns before writing to Amazon Redshift. The job must run reliably without writing custom serialization code. Which approach should the data engineer take?
Medium434A company has a nightly batch job that processes 100 GB of data from an Amazon S3 bucket and loads it into an Amazon Redshift table. The job currently runs on an Amazon EMR cluster. Which service would reduce operational overhead while providing similar functionality?
Easy435Which TWO AWS services can be used to ingest streaming data into Amazon S3? (Choose two.)
Easy436A data engineer is designing a real-time streaming pipeline to ingest clickstream data from a website into Amazon S3. The data must be transformed before storage. Which TWO AWS services can be used together to build this pipeline? (Choose TWO.)
Easy437A data engineer is using AWS Glue to read a large dataset from Amazon S3 and write it to Amazon Redshift. The job intermittently fails with 'Communication link failure' errors during the write phase. The dataset is several hundred gigabytes and the Redshift cluster is under heavy query load. Which change is MOST likely to resolve the failures while preserving data integrity?
Hard438A company uses AWS Glue ETL jobs to transform data from Amazon S3 to Amazon Redshift. The job reads JSON files, applies schema mapping, and writes to a Redshift table. Recently, the job started failing with memory errors. The data volume has increased tenfold. Which approach should a data engineer take to resolve this issue with minimal code changes?
Medium439A data engineer is building a data ingestion pipeline using AWS Glue. The source is an Amazon DynamoDB table, and the target is an Amazon S3 data lake in Parquet format. The pipeline must handle large volumes and ensure exactly-once processing. Which THREE features should the engineer use together to achieve this? (Choose THREE.)
Hard440A data engineer uses AWS Glue to process data from S3. The Glue job frequently fails with 'Out of Memory' errors. The job reads several large compressed files. What is the MOST effective way to resolve this issue without changing the code?
Medium441A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data must be transformed in real-time and then stored in Amazon S3. Which AWS service should be used to perform the transformation?
Easy442A company uses AWS DMS to migrate an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. After initial load, ongoing replication is set up. The replication task shows 'Task status: failed with error: The specified LSN is not available in the source database logs.' What is the most likely cause?
Medium443Refer to the exhibit. A data engineer runs the above CLI command to find files smaller than 1000 bytes in a bucket. The command returns an empty array, but the engineer knows there are small files. What is the issue?
Easy444A data engineer is using AWS Glue to process a large dataset stored in Amazon S3. The dataset is partitioned by year, month, and day. The engineer notices that the Glue job is taking a long time and consuming many DPUs. The job reads all partitions, filters the data, and writes the result to another S3 location. The engineer wants to optimize the job to process only the required partitions and reduce cost. Which action should the engineer take?
Hard445A company uses Amazon Kinesis Data Analytics for real-time anomaly detection on clickstream data. The application uses a sliding window of 1 minute. The data engineer notices that the application is producing incorrect results because late-arriving records are not being handled properly. What should the data engineer do to ensure late records are included in the window calculations?
Hard446A data engineer is using AWS Glue to read data from an Amazon Kinesis Data Stream. The Glue job is configured to process the stream in micro-batches. The engineer notices that the job is not processing all records and sometimes skips data. The Kinesis stream has multiple shards, and the Glue job is using the 'kinesis' connection type. What is the most likely cause of the missing records?
Hard447A data engineer is orchestrating a multi-step ingestion workflow where CSV files land in Amazon S3, an AWS Glue job transforms them, and the output is loaded into Amazon Redshift. The engineer wants conditional branching, retry logic, and the ability to pass parameters between steps. Which AWS service should be used to orchestrate this workflow?
HardOther domains
All DEA-C01 exam domains
Frequently asked questions
- What does the Data Ingestion and Transformation domain cover on the DEA-C01 exam?
- Be able to match ingestion services to constraints: source impact, volume, file size, frequency, and schema differences. The single most important thing is justifying the choice with the stated constraint, not the most powerful service.
- How many questions are in this domain?
- This page lists all 447 Data Ingestion and Transformation questions in the DEA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Data Ingestion and Transformation questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.