Courseiva
Data Ingestion and TransformationhardMultiple ChoiceObjective-mapped

DEA-C01 Data Ingestion and Transformation Practice Question

A financial services company ingests stock trade data from multiple exchanges into an Amazon S3 bucket (trade-bucket). Each exchange sends a CSV file every 5 minutes. The data must be transformed into Parquet format and partitioned by exchange and date (trade_date) for efficient querying using Amazon Athena. The pipeline must handle late-arriving data (files up to 2 hours late) and ensure exactly-once processing to avoid duplicates. Currently, a scheduled AWS Glue ETL job runs every hour, reads new CSV files, converts them to Parquet, and writes to an output bucket. However, the team is experiencing data duplication: if the job fails midway, upon retry it reprocesses all files in the input folder, causing duplicates in the output. Additionally, the job takes too long because it scans all files each run. The engineer must redesign the pipeline to eliminate duplicates and improve efficiency. What should the engineer do?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Set up an S3 event notification to invoke an AWS Lambda function that starts a Glue job with a parameter containing the S3 object key of the new file; modify the Glue job to process only that file and use the file key to avoid duplicates.

The best approach. By setting up an S3 event notification to invoke a Lambda function that triggers a Glue job with the new file's S3 key as a parameter, the job processes only that specific file. Using the file key in the job logic ensures idempotency—if the job fails and retries, it reprocesses the same file key, and deduplication can be handled (e.g., by checking if the output partition already contains that file's data or using a job bookmark on the file key). This achieves exactly-once processing and incremental processing (no full scans), improving efficiency. Option A (Glue Workflows) still processes all files each run and doesn't prevent duplicates if a file arrives after the job starts. Option C (moving CSV files to an archive folder) risks race conditions if late-arriving data comes while the job is running, and does not guarantee exactly-once if the job fails mid-way. Option D (EMR with Spark Structured Streaming) is overly complex and expensive for a 5-minute CSV batch ingestion; checkpointing can handle failures but adds significant operational overhead.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Use AWS Glue Workflows to orchestrate the job and add a condition to check for duplicates before writing.

    Why it's wrong here

    Glue Workflows do not provide file-level idempotency; duplicate detection would require additional logic and scans.

  • Set up an S3 event notification to invoke an AWS Lambda function that starts a Glue job with a parameter containing the S3 object key of the new file; modify the Glue job to process only that file and use the file key to avoid duplicates.

    Why this is correct

    This ensures each file is processed exactly once, and the job runs only on new files, improving efficiency.

  • Modify the Glue job to move processed CSV files to an archive folder after successful transformation, and process only unprocessed files.

    Why it's wrong here

    Moving files does not guarantee exactly-once if the job fails after moving some files; also, moving files adds latency and cost.

  • Replace Glue with Amazon EMR and use Spark Structured Streaming with checkpointing to process files incrementally.

    Why it's wrong here

    Using Amazon EMR with Spark Structured Streaming and checkpointing is excellent for incremental processing of new files and ensuring exactly-once semantics for continuous data streams. However, it is not the most robust solution for handling *late-arriving files* that might appear hours after their expected processing window, as a continuous stream might struggle to efficiently re-evaluate past windows without complex watermarking or re-scanning. This approach is tempting because it addresses the incremental processing and exactly-once requirements for real-time or near real-time ingestion, making it suitable for scenarios where data arrives consistently and on time.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,711 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.