DEA-C01 Data Store Management Practice Question
A data engineer maintains an Amazon S3 data lake with millions of small JSON objects. The engineer needs to improve query performance by reducing the number of objects and compressing them into a columnar format that Amazon Athena can query efficiently. The data must remain partitioned by date. Which solution should the engineer use?
⚠ Common exam trap
The trap here is assuming that simply merging files or changing storage classes solves the small-file problem, when the key is converting to a columnar format and repartitioning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use AWS Glue ETL to read the JSON objects, repartition by date, write them as Parquet with Snappy compression, and register the resulting tables in the AWS Glue Data Catalog.
Converting small JSON files to Parquet with Snappy compression and repartitioning by date reduces storage footprint and the amount of data scanned by Athena. AWS Glue ETL can perform this transformation and update the Data Catalog so Athena queries the new tables. This is the standard pattern for optimizing S3 data lakes for analytics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable S3 Transfer Acceleration on the bucket and run Amazon Athena queries with the existing JSON objects.
Why it's wrong here
S3 Transfer Acceleration speeds up uploads and downloads over long distances, but it does not reduce the number of objects, convert to columnar format, or lower Athena scan costs. Athena still reads the same JSON data, so query performance and cost remain largely unchanged. It does not solve the small-file or format problem.
- ✓
Use AWS Glue ETL to read the JSON objects, repartition by date, write them as Parquet with Snappy compression, and register the resulting tables in the AWS Glue Data Catalog.
Why this is correct
AWS Glue ETL can transform JSON to Parquet, repartition by date, and update the Data Catalog, which Athena uses for schema and partition metadata. Parquet with Snappy reduces storage and scan size, and fewer larger files improve query performance. This directly addresses the small-file problem and columnar requirement while preserving date partitioning.
- ✗
Use Amazon S3 Lifecycle policies to transition the JSON objects to S3 Glacier Instant Retrieval and query them with Athena.
Why it's wrong here
S3 Glacier Instant Retrieval is an archival storage class for rarely accessed data. Moving active query data there increases retrieval costs and latency, and Athena supports it but it is not intended for frequent analytical queries. It also does not convert data to columnar format or reduce object count, so it fails the performance goal.
- ✗
Create an Amazon EMR cluster and run a MapReduce job that merges the JSON files into larger JSON files without changing the format.
Why it's wrong here
Merging small JSON files into larger JSON files reduces object count, but the data remains row-based JSON. Athena must still parse JSON and cannot take advantage of columnar pruning or compression. This approach may slightly improve performance but does not meet the columnar format requirement or optimize scan costs as effectively as Parquet.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.