DEA-C01 Data Store Management Practice Question
A data engineer is designing a data lake on Amazon S3. Data is ingested from multiple sources in JSON format. The engineer needs to optimize query performance for Amazon Athena while minimizing storage costs. Which storage strategy should the engineer use?
⚠ Common exam trap
Test-takers frequently assume JSON or CSV are acceptable for Athena due to their simplicity, overlooking that columnar formats like Parquet are required for cost-efficient querying in AWS's pay-per-scan model.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert data to Parquet format and partition by date.
Parquet is a columnar storage format that significantly reduces data scan volume in Amazon Athena, which charges per byte scanned. Partitioning by date further limits the data scanned to only relevant partitions, optimizing both query performance and cost. JSON and CSV are row-based formats that require full scans, and Glacier is unsuitable for interactive querying.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Store data as CSV files in a single S3 bucket without prefixes.
Why it's wrong here
CSV lacks the nested structure of the source JSON, and a bucket without prefixes gives Athena no partition pruning, so every query scans all objects. It is tempting as a flat, compact format, but partitioning and columnar formats are required to reduce scanned bytes and cost.
- ✓
Convert data to Parquet format and partition by date.
Why this is correct
Parquet is columnar and compressed, so Athena scans and bills far less data than JSON. Partitioning by date enables partition pruning, restricting each query to relevant prefixes. Together these cut query latency and S3 storage cost.
- ✗
Store data as JSON files in a single prefix without partitioning.
Why it's wrong here
A single unpartitioned prefix forces Athena to scan every JSON object for each query, and uncompressed JSON inflates bytes scanned, raising cost. It is tempting for simplicity of ingestion, but partitioning by source or date plus columnar compression is what actually cuts scanned data and storage.
- ✗
Store compressed JSON files in Amazon S3 Glacier.
Why it's wrong here
Athena cannot query Glacier-stored objects directly, and retrieval adds latency and cost, defeating the performance goal. It is tempting because Glacier minimises storage cost, but it suits archival data rarely accessed, not a data lake queried interactively through Athena.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.