DEA-C01 Data Store Management Practice Question
A data engineer manages a large Amazon S3 data lake with millions of small JSON files ingested daily. Amazon Athena queries against this data lake are slow and costly due to high per-query data scanned. The engineer wants to optimize the storage layout to improve query performance and reduce cost, while keeping the data queryable in place. Which solution should the engineer implement?
⚠ Common exam trap
The trap here is assuming that S3 Transfer Acceleration or format changes alone improve Athena query performance, when the key is reducing data scanned through compaction and partitioning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use AWS Glue ETL to compact the small JSON files into larger Parquet files partitioned by commonly filtered columns, then update the AWS Glue Data Catalog.
Compacting small files into larger Parquet files and partitioning by frequently filtered columns directly reduces the amount of data scanned by Athena, improving performance and lowering cost. Parquet's columnar format allows Athena to read only the columns needed, and partitioning enables partition pruning. Updating the Data Catalog ensures queries use the optimized layout. The other options do not address the core issues of small files and inefficient format.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use AWS Glue ETL to compact the small JSON files into larger Parquet files partitioned by commonly filtered columns, then update the AWS Glue Data Catalog.
Why this is correct
Compacting small files into larger Parquet files reduces the number of S3 GET requests and leverages columnar storage, which minimizes data scanned by Athena. Partitioning by frequently filtered columns enables partition pruning, further reducing data scanned. Updating the Data Catalog ensures Athena queries use the new schema and partitions. This directly addresses performance and cost without moving data out of S3.
- ✗
Enable S3 Transfer Acceleration on the bucket to speed up data retrieval by Athena.
Why it's wrong here
S3 Transfer Acceleration speeds up uploads and downloads over long distances by using AWS edge locations, but it does not optimize query performance or reduce data scanned by Athena. It is designed for transfer speed, not for analytical query efficiency. This would not address the small-file problem or the lack of partitioning, and would not lower Athena query costs.
- ✗
Convert the data to CSV format and add more columns to the AWS Glue Data Catalog table.
Why it's wrong here
CSV is row-based and not columnar, so Athena would still scan entire rows, not just the needed columns, leading to high data scanned. Adding more columns to the catalog does not reduce data scanned and could increase it if queries select those columns. This approach does not solve the small-file issue or improve performance meaningfully.
- ✗
Move the data to Amazon Redshift and query it using Redshift Spectrum.
Why it's wrong here
Redshift Spectrum allows querying S3 data, but it still scans the same small JSON files and does not inherently optimize them. Moving to Redshift would incur additional cost and complexity, and the small-file issue persists unless data is transformed. The requirement is to optimize in place, so this is not the best solution.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.