DEA-C01 Data Operations and Support Practice Question
A data engineer is using Amazon Athena to query data stored in Amazon S3 in Parquet format. The engineer notices that a specific query is scanning much more data than expected, resulting in high costs and slow performance. The query filters on a column named 'event_date' which is a string in 'YYYY-MM-DD' format. The table is partitioned by 'year', 'month', and 'day' as separate string columns. The engineer wants to reduce the amount of data scanned. Which action should the engineer take?
⚠ Common exam trap
The trap here is assuming that converting a column to date type or enabling caching will reduce data scanned, when the key optimization is to use the existing partition columns in the filter.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the partition columns 'year', 'month', and 'day' in the WHERE clause instead of 'event_date'.
The table is partitioned by year, month, and day, so filtering on those partition columns in the WHERE clause enables Athena's partition pruning. This limits the data scanned to only the relevant partitions, significantly reducing cost and improving performance. Filtering on a non-partition column like event_date does not provide the same benefit.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Compress the Parquet files using Snappy compression.
Why it's wrong here
Compressing Parquet files reduces storage size and can lower the amount of data scanned because Athena reads compressed data, but it does not eliminate scanning of irrelevant partitions. The primary issue is that the query is not leveraging partitions, so compression alone will not solve the problem. It is a good practice but secondary to partition pruning.
- ✗
Enable Amazon Athena query result reuse to cache the results.
Why it's wrong here
Query result reuse caches the results of previously executed queries, which can reduce cost for repeated identical queries. However, it does not reduce the data scanned for a new query with different filters. Since the issue is a specific query scanning too much data, caching would not help unless the exact same query is run again. It is not a solution for the current query's inefficiency.
- ✗
Convert the 'event_date' column to a date type and use it in the WHERE clause.
Why it's wrong here
Converting the column type may improve predicate pushdown if the query uses date functions, but it does not leverage the existing partitions. The partitions are on year, month, and day, so filtering on those columns directly is more effective. Changing the data type alone will not reduce the scanned data if the query does not filter on partition columns. It could also require rewriting data.
- ✓
Use the partition columns 'year', 'month', and 'day' in the WHERE clause instead of 'event_date'.
Why this is correct
Athena uses partition pruning to limit the data scanned based on partition columns in the WHERE clause. Since the table is partitioned by year, month, and day, filtering on these columns allows Athena to scan only the relevant partitions, drastically reducing data scanned. This is the most effective way to optimize the query and reduce costs.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.