DEA-C01 Data Store Management Practice Question
A data engineer is setting up an Amazon S3 bucket to store large CSV files that will be queried using Amazon Athena. The engineer wants to minimize query costs and improve performance. The files are currently stored in a single prefix without any partitioning. The most common queries filter data by `year` and `month`. What should the engineer do to optimize the Athena queries?
⚠ Common exam trap
The trap here is thinking that simply defining partition keys in the Glue Data Catalog without reorganizing the data will enable partition pruning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the CSV files to Apache Parquet format and partition the data by `year` and `month`.
Converting data to Apache Parquet and partitioning by commonly filtered columns like `year` and `month` are the most effective ways to reduce data scanned by Athena. Parquet's columnar format allows Athena to read only the columns needed, and partitioning enables partition pruning. Together, they minimize query cost and improve performance. Other options either do not address the core optimizations or introduce unnecessary complexity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Compress the CSV files using gzip and keep them in a single prefix.
Why it's wrong here
Compressing CSV files with gzip reduces storage size and data scanned during queries, but it does not provide the same level of optimization as Parquet. CSV is row-based, so Athena still reads all columns even if only a few are needed. Also, without partitioning, queries filtering by `year` and `month` will scan all data. This approach is less effective than converting to a columnar format and partitioning.
- ✓
Convert the CSV files to Apache Parquet format and partition the data by `year` and `month`.
Why this is correct
Converting to Parquet reduces the amount of data scanned because Parquet is columnar and compressed, and partitioning by `year` and `month` allows Athena to prune partitions based on query filters. This significantly lowers query cost and improves performance. Athena charges based on data scanned, so both optimizations directly reduce cost. This is a best practice for optimizing Athena queries on large datasets.
- ✗
Create an AWS Glue Data Catalog table with partition keys `year` and `month`, and use Amazon Athena to query the CSV files directly.
Why it's wrong here
Creating a table with partition keys without actually organizing the data into partitioned prefixes will not enable partition pruning. The data must be stored in a directory structure like `year=2023/month=10/` for Athena to use partitions. Simply defining partition keys in the catalog does not reorganize the data. Therefore, queries will still scan all data, leading to higher costs and slower performance.
- ✗
Use Amazon Redshift Spectrum to query the CSV files in S3 instead of Athena.
Why it's wrong here
Amazon Redshift Spectrum allows querying S3 data from a Redshift cluster, but it is not a replacement for Athena in this scenario. Spectrum also benefits from columnar formats and partitioning. Using Spectrum would require an active Redshift cluster, adding cost and complexity. The goal is to optimize Athena queries, so this does not address the need. It also does not inherently optimize the CSV files for Athena.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.