DEA-C01 Data Store Management Practice Question
A data engineer is designing a data lake on Amazon S3 for a retail company. The company ingests point-of-sale data as small JSON files every few minutes, totaling about 5 GB per day. Analysts query the data with Amazon Athena, and costs are rising due to many small files and full scans. The engineer wants to reduce Athena query costs and improve performance while keeping the data in S3. Which TWO actions should the engineer take? (Choose two.)
⚠ Common exam trap
Test-takers frequently confuse storage-cost optimizations like Intelligent-Tiering with query-cost optimizations, which depend on scanned bytes and file layout.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run an AWS Glue ETL job or AWS Glue compaction to merge small files into larger files of 128 MB or more.
Athena cost and performance depend on bytes scanned and the number of files read. Converting JSON to partitioned Parquet enables column pruning and partition pruning, while compacting small files into larger ones reduces request overhead and metadata processing. These two actions together cut scanned bytes and improve query speed without leaving S3.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the Athena workgroup data usage control limit to allow larger scans.
Why it's wrong here
Raising the data usage control limit permits more bytes to be scanned but does not reduce the amount scanned, so it would increase cost rather than lower it. It is a guardrail setting, not an optimization, and does not address small files or full scans.
- ✓
Run an AWS Glue ETL job or AWS Glue compaction to merge small files into larger files of 128 MB or more.
Why this is correct
Compacting many small files into larger ones reduces the number of S3 GET requests and metadata overhead, which lowers Athena query latency and cost. Glue jobs or Glue compaction can perform this merge while keeping the data in S3, complementing the Parquet and partition changes.
- ✗
Enable Amazon S3 Intelligent-Tiering on the bucket to automatically move data to lower-cost storage classes.
Why it's wrong here
S3 Intelligent-Tiering reduces storage costs by moving objects between access tiers, but Athena query cost is based on bytes scanned, not storage class. It does not reduce the number of files or the columns read, so it does not lower query costs or improve performance for the analysts.
- ✓
Convert the JSON files to Apache Parquet and store them partitioned by date and store ID.
Why this is correct
Converting to Parquet enables columnar reads and compression, so Athena scans only the columns referenced in a query, and partitioning by date and store ID lets partition pruning skip irrelevant prefixes. Together these reduce bytes scanned and query cost, directly addressing the small-file and full-scan problems.
- ✗
Enable S3 Transfer Acceleration on the bucket to speed up Athena queries.
Why it's wrong here
S3 Transfer Acceleration speeds up uploads to S3 over long distances and has no effect on Athena query performance or bytes scanned. It does not address small files or full scans, so it will not reduce query costs for the analysts.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.