Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is troubleshooting an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The job runs successfully but 5% of records are missing after the load. The engineer suspects data consistency issues. Which THREE actions could help diagnose and resolve the problem? (Choose THREE.)

⚠ Common exam trap

Many candidates assume performance tuning (increasing DPUs) or database-level transactions (staging tables) can fix data ingestion gaps, when the actual problem is incomplete or inconsistent file discovery from the source (S3).

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the Redshift COPY command with a manifest file to load data.

Option A is correct because using the Redshift COPY command with a manifest file explicitly lists every S3 object to be loaded, preventing silent omissions of files that can occur with prefix-based loads and making missing records easier to detect. Option C is correct because Glue job bookmarks track which S3 files have already been processed, so enabling them prevents files from being skipped or reprocessed and helps identify gaps in the input data that cause missing records. Option E is correct because reviewing the job's CloudWatch Logs surfaces errors, warnings, and skipped-record messages from Glue and Redshift, which is the primary diagnostic step for understanding why 5% of records are missing. Option B does not belong because increasing DPUs only scales compute capacity and does not address data consistency or missing records. Option D does not belong because a staging table with a transaction ensures atomicity of the load but does not by itself diagnose or resolve records being dropped during extraction or copy.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the Redshift COPY command with a manifest file to load data.

    Why this is correct

    A manifest file explicitly lists every S3 object to load, so COPY fails loudly on missing or unreadable files rather than silently skipping them. This surfaces the 5% discrepancy and prevents partial loads, addressing the suspected consistency issue.

  • ✗

    Increase the number of DPUs for the Glue job.

    Why it's wrong here

    Adding DPUs increases parallel processing capacity and execution speed, but missing records stem from write or commit behaviour, not throughput. It is tempting because scaling DPUs is the standard remedy for slow or memory-constrained Glue jobs, and would be correct if the job were timing out or running long.

  • ✓

    Enable Glue job bookmarks to track processed files.

    Why this is correct

    Glue job bookmarks persist state per transformation context, so reruns skip already-processed S3 objects instead of reprocessing or skipping them inconsistently. This directly addresses the missing-records symptom by ruling out duplicate-suppression or stale bookmark state as the cause of the 5% loss.

  • ✗

    Use a staging table in Redshift with a transaction to commit.

    Why it's wrong here

    A staging table with a transactional commit is already the recommended pattern for atomic Redshift loads, so it addresses the symptom rather than diagnosing the root cause of the 5% loss. It is tempting because staging tables genuinely prevent partial loads during failures, making them the right fix when loads abort midway.

  • ✓

    Review the job's CloudWatch Logs for any error messages.

    Why this is correct

    CloudWatch Logs capture Glue job-level errors, including write failures to Redshift that the job may swallow without failing. Reviewing them surfaces rejected records, connection timeouts, or COPY errors explaining the 5% loss, directly addressing the missing-records symptom before deeper consistency checks.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.