DEA-C01 Data Store Management Practice Question
A data engineering team stores clickstream events in an Amazon S3 bucket under the prefix s3://analytics/raw/. New objects arrive continuously, and the team wants Amazon Athena queries to scan only the events for the current day without scanning the entire prefix. The events are written as JSON files partitioned by year/month/day, but queries still scan all partitions because the partition metadata is not registered. Which action should the data engineer take to enable partition pruning in Athena?
⚠ Common exam trap
The trap here is assuming that a columnar file format or faster data transfer alone limits how much data Athena scans, when partition pruning depends on registered partition metadata.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run MSCK REPAIR TABLE on the Athena table to add the year/month/day partitions to the AWS Glue Data Catalog.
Partition pruning requires the query engine to know which partition values exist and where their data lives. Registering the year/month/day folders as partitions in the Data Catalog lets Athena build a plan that reads only the matching day, cutting scanned bytes and cost. Physical optimizations such as acceleration or columnar formats do not supply that partition metadata.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Run MSCK REPAIR TABLE on the Athena table to add the year/month/day partitions to the AWS Glue Data Catalog.
Why this is correct
MSCK REPAIR TABLE scans the S3 prefix, discovers Hive-style partition folders such as year=2024/month=03/day=15, and registers each partition in the Data Catalog so Athena can prune non-matching partitions. This directly addresses the missing partition metadata that causes full-prefix scans in this scenario.
- ✗
Create an Amazon CloudFront distribution in front of the S3 bucket and point Athena at the distribution.
Why it's wrong here
Athena reads data directly from S3 through its own integration and does not query data through a CloudFront distribution. CloudFront caches HTTP content for end users, not SQL table data, so it cannot provide partition metadata or reduce the bytes scanned by an Athena query in this design.
- ✗
Enable S3 Transfer Acceleration on the bucket so Athena can retrieve the daily events faster.
Why it's wrong here
Transfer Acceleration speeds up uploads and downloads to S3 over long distances by using edge locations, but it does not register partition metadata or change how Athena plans a query. Without catalog partitions, Athena still lists and scans every object under the prefix, so query cost and latency remain unchanged.
- ✗
Convert the JSON files to Apache Parquet and rely on columnar storage alone to limit the scan to one day.
Why it's wrong here
Parquet reduces bytes read through column pruning and compression, but it does not restrict a query to a single day. Athena still scans every Parquet file under the prefix unless partitions are registered, so converting formats alone fails to achieve the day-level scan reduction the team requires.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.