DEA-C01 Data Store Management Practice Question
A data engineer manages an Amazon S3 data lake where analytics queries run through Amazon Athena. Monthly partition folders hold Parquet files, and each partition contains tens of thousands of small files averaging 40 KB. Athena queries that scan a single month take much longer than expected and consume far more bytes scanned than the actual data volume. The engineer must improve query performance without changing the table schema or the folder layout. What should the engineer do?
⚠ Common exam trap
The trap here is assuming that a storage-class or network-acceleration change can fix a small-file problem, when the real fix is rewriting the objects into fewer larger files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run an AWS Glue ETL job that reads each partition and rewrites it as fewer, larger Parquet files of roughly 128 MB, then update the partitions in the Data Catalog.
Athena performance degrades when a partition holds a huge number of very small objects because each object requires a separate request and contributes metadata overhead, inflating both runtime and reported bytes scanned. Compacting the files with an AWS Glue job into larger Parquet objects preserves the schema and partition layout while cutting that overhead, which is the standard remedy for this pattern.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable S3 Transfer Acceleration on the bucket so that Athena can retrieve the small objects with lower latency across edge locations.
Why it's wrong here
Transfer Acceleration speeds up uploads and downloads over long geographic distances by routing through edge locations, but Athena reads objects inside the AWS network where that benefit does not apply. The bottleneck here is the sheer number of small file opens, not network distance, so acceleration leaves both query duration and scanned bytes essentially unchanged.
- ✓
Run an AWS Glue ETL job that reads each partition and rewrites it as fewer, larger Parquet files of roughly 128 MB, then update the partitions in the Data Catalog.
Why this is correct
Consolidating many tiny files into fewer large Parquet files drastically reduces the per-file overhead that Athena pays when listing and opening objects, so scans finish faster and less metadata work is repeated. The rewrite keeps the same partition structure and schema, satisfying the constraint that neither the table definition nor the folder layout may change.
- ✗
Attach an S3 Lifecycle policy that transitions the Parquet objects to S3 Glacier Instant Retrieval after 30 days to improve read throughput.
Why it's wrong here
Glacier Instant Retrieval lowers storage cost for cold data and still allows millisecond access, but it does not increase throughput or reduce the number of objects Athena must open. Applying it to actively queried monthly partitions adds retrieval charges without addressing the small-file problem, so query latency and scanned bytes stay high.
- ✗
Increase the number of partitions by splitting each monthly folder into daily folders and repointing the table at the new prefixes.
Why it's wrong here
Finer partitioning can prune more data for selective queries, but it multiplies the number of small files and partitions, worsening the per-file overhead that is causing the slow scans. It also changes the folder layout, which the scenario explicitly forbids, and it does not reduce bytes scanned for queries that already target whole months.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.