DEA-C01 Data Store Management Practice Question
A data engineer is building a near-real-time ingestion pipeline into Amazon S3. Small JSON files arrive continuously from thousands of devices, and the engineer must optimize the data lake for downstream Amazon Athena queries while minimizing storage cost and query latency. Which TWO actions should the engineer take? (Choose two.)
⚠ Common exam trap
The trap here is focusing on upload speed or storage class alone, when the real issue is the number and format of files that Athena must scan.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use AWS Glue ETL jobs to compact small JSON files into larger Parquet files partitioned by ingestion date.
The pipeline suffers from many small JSON files, which increase Athena query overhead and cost. Compacting files into larger Parquet objects with AWS Glue ETL and using Kinesis Data Firehose to buffer and aggregate records both reduce object count and improve columnar query performance. Transfer Acceleration, Glacier Instant Retrieval, and aborting multipart uploads do not address file size or format for analytics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Store the raw JSON files in S3 Glacier Instant Retrieval to reduce storage cost immediately.
Why it's wrong here
S3 Glacier Instant Retrieval is an archive class with a 90-day minimum storage duration and higher retrieval costs. Using it for data that must be queried by Athena in near-real-time would add latency and expense. It also does not solve the small-file issue, and Athena queries against archived objects can incur retrieval charges.
- ✗
Configure an S3 Lifecycle rule to abort incomplete multipart uploads after 7 days.
Why it's wrong here
Aborting incomplete multipart uploads removes orphaned parts and can reduce storage cost, but it does not affect the size or format of the completed JSON files. It also does not improve Athena query performance. While a good hygiene practice, it is not one of the two actions that directly optimize the data lake for analytics.
- ✓
Use AWS Glue ETL jobs to compact small JSON files into larger Parquet files partitioned by ingestion date.
Why this is correct
Compacting many small JSON files into larger Parquet files reduces the number of S3 objects and takes advantage of columnar storage, which lowers Athena scan costs and improves query latency. Partitioning by ingestion date further enables partition pruning. This directly addresses both the small-file problem and the need for efficient downstream analytics.
- ✓
Use Amazon Kinesis Data Firehose to buffer incoming records and deliver larger aggregated files to S3.
Why this is correct
Kinesis Data Firehose can buffer incoming records based on size or time and deliver them as larger files, which reduces the number of small objects written to S3. It can also convert records to Parquet using AWS Glue schema inference. This directly mitigates the small-file problem and improves downstream Athena efficiency.
- ✗
Enable S3 Transfer Acceleration on the ingestion bucket to speed up uploads from devices.
Why it's wrong here
S3 Transfer Acceleration speeds up uploads over long distances by using AWS edge locations, but it does not address the small-file problem or improve Athena query performance. It also adds cost per GB transferred. The bottleneck here is the number and format of files, not the upload path, so this action does not meet the optimization goals.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.