DEA-C01 Data Store Management Practice Question
A data engineer is managing an Amazon S3 data lake that contains raw JSON data. The engineer needs to optimize the data lake for query performance and cost when using Amazon Athena. The data is currently stored in a single S3 prefix without partitioning, and queries often filter on `event_type` and `event_date`. The engineer wants to implement best practices for Athena. Which TWO actions should the engineer take? (Choose two.)
⚠ Common exam trap
The trap here is considering gzip compression as sufficient, but it does not provide the columnar benefits of Parquet and may not be splittable.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the JSON data to Apache Parquet format.
The two most effective actions are converting JSON to Parquet and partitioning by `event_type` and `event_date`. Parquet's columnar format reduces data scanned, and partitioning enables partition pruning. Together, they minimize query cost and improve performance. Other options either do not affect Athena queries or are less effective. These are core best practices for optimizing Athena on S3 data lakes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable S3 Transfer Acceleration on the bucket.
Why it's wrong here
S3 Transfer Acceleration speeds up uploads and downloads to S3 by using AWS edge locations, but it does not affect Athena query performance or cost. Athena reads data directly from S3 within the same region, so Transfer Acceleration provides no benefit for query optimization. It also adds extra cost. Therefore, it is not a relevant action for improving Athena queries.
- ✓
Convert the JSON data to Apache Parquet format.
Why this is correct
Converting JSON to Parquet reduces the amount of data scanned by Athena because Parquet is columnar and compressed. Athena can read only the columns needed for a query, significantly lowering cost and improving performance. JSON is row-based and not splittable, so queries scan more data. Parquet also supports efficient compression and encoding. This is a fundamental optimization for Athena.
- ✗
Compress the JSON files using gzip.
Why it's wrong here
Compressing JSON files with gzip reduces storage size and the amount of data scanned by Athena, but it is less effective than converting to a columnar format like Parquet. JSON is still row-based, so Athena must read entire rows even if only a few columns are needed. Additionally, gzip-compressed JSON is not splittable, which can limit parallelism. While compression is beneficial, it is not as impactful as Parquet and partitioning. Therefore, it is not one of the two best actions.
- ✓
Partition the data by `event_type` and `event_date` in S3.
Why this is correct
Partitioning the data by commonly filtered columns like `event_type` and `event_date` allows Athena to prune partitions and scan only the relevant data. This reduces the amount of data scanned and lowers query cost. The data must be organized into a directory structure such as `event_type=click/event_date=2023-10-01/`. Partitioning is a best practice for large datasets queried with filters on those columns.
- ✗
Use Amazon S3 Select to filter data before querying with Athena.
Why it's wrong here
Amazon S3 Select allows you to retrieve a subset of data from an S3 object using simple SQL expressions, but it is not integrated with Athena. Athena does not use S3 Select to filter data; it reads objects directly. While S3 Select can reduce data transferred for certain applications, it does not optimize Athena queries. Moreover, S3 Select does not support Parquet. Therefore, this is not a valid optimization for Athena.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.