DEA-C01 Data Ingestion and Transformation Practice Question
A company wants to use AWS Glue to transform data stored in Amazon S3. The data is partitioned by date and includes both CSV and Parquet files. The transformation should be optimized for cost and performance. Which THREE actions should the data engineer take? (Choose THREE.)
⚠ Common exam trap
DEA-C01 often tests cost optimization by tempting candidates with 'more DPUs = faster' (Option D) or 'crawl every run' (Option A), when the correct answers are the data-reduction techniques — partition pruning, job bookmarks, and columnar format conversion.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use partition pruning by filtering on the date column in the ETL script.
Option B is correct because partition pruning by filtering on the date column lets AWS Glue read only the relevant S3 partitions instead of scanning the entire dataset, which directly reduces I/O, runtime, and cost. Option C is correct because Glue job bookmarks persist state between runs so the job processes only new or changed data since the last run, avoiding reprocessing of already-transformed partitions and lowering both execution time and DPU consumption. Option E is correct because converting CSV files to Parquet enables columnar storage and compression, so Glue reads less data and benefits from predicate pushdown, improving performance and reducing cost compared with row-based CSV. Option A is not appropriate because running a crawler before every job run adds unnecessary cost and time, and the schema can be supplied directly or updated only when it changes. Option D is not appropriate because maximizing DPUs increases cost rather than optimizing it, and the required performance can be achieved through pruning, bookmarks, and columnar formats.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Run a crawler to update the schema before each job run.
Why it's wrong here
Crawler runs add cost and latency on every job, and schema changes between runs break job bookmarks; partition metadata already exists. Crawlers suit initial schema discovery or infrequent source changes, not repeated per-run refresh in a cost-optimised pipeline.
- ✓
Use partition pruning by filtering on the date column in the ETL script.
Why this is correct
Filtering on the date partition column lets Glue read only relevant S3 prefixes rather than the full dataset. This reduces bytes scanned and DPU hours, directly meeting the cost and performance optimisation constraint in the stem.
- ✓
Use job bookmarks to process only new data.
Why this is correct
Job bookmarks persist state so each Glue run reads only newly arrived S3 objects, skipping previously processed partitions. This directly satisfies the incremental-processing requirement, avoiding full re-scans of historical CSV and Parquet data, which cuts both DPU hours and runtime cost.
- ✗
Increase the number of DPUs to the maximum allowed.
Why it's wrong here
Maximum DPUs raise per-hour compute charges without addressing the real bottleneck, such as file format or partition pruning, so cost rises while performance gains plateau. Scaling DPUs suits genuinely CPU-bound, parallelisable jobs where shuffle and transform stages dominate runtime.
- ✓
Convert all files to Parquet format before processing.
Why this is correct
Converting CSV to Parquet before processing satisfies the cost and performance constraint: Parquet's columnar storage and compression let AWS Glue read only needed columns, sharply reducing bytes scanned and DPU-hours. This lowers both S3 request costs and Glue job runtime compared with row-based CSV scanning.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.