DEA-C01 Columnar Storage Practice Question
A company uses Amazon Athena to query data in S3. Recently, queries have become slow. The data is stored as CSV files in a partitioned table. What is the most effective way to improve query performance?
⚠ Common exam trap
A common trap is to think that simply increasing file size or using a more popular format like JSON will help. However, the key is switching to a columnar format (Parquet or ORC) that minimizes data scanned.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the data to Parquet format and optimize partitioning.
The correct answer. Parquet is a columnar storage format that allows Athena to read only the columns needed for a query, reducing I/O and improving performance. Combined with effective partitioning, it enables partition pruning, which further limits the data scanned. CSV files are row-based and require full scans, even with partitioning. Option A is incorrect because Athena is serverless and users cannot increase nodes; resources are managed automatically. Option C is incorrect: JSON is also row-based and verbose, making it even slower than CSV. Option D is incorrect because larger CSV files still lead to full scans; Parquet's columnar nature is more impactful than file size.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of nodes in the Athena query engine.
Why it's wrong here
Athena is serverless; there are no nodes to scale. Performance depends on bytes scanned, so converting CSV to a columnar format such as Parquet and partitioning the table reduces data read. Adding nodes is tempting because it mirrors cluster scaling in engines like Amazon Redshift or EMR, but Athena offers no such capacity control.
- ✓
Convert the data to Parquet format and optimize partitioning.
Why this is correct
Parquet is columnar, so Athena reads only referenced columns and skips row-by-row parsing of CSV, cutting scanned bytes and cost. Combined with tighter partition pruning, this directly addresses the slow queries against the partitioned S3 table.
- ✗
Convert the data to JSON format.
Why it's wrong here
JSON is row-based and text-encoded, so Athena still scans and parses every record without columnar pruning or compression gains. Converting CSV to Parquet, a columnar format, is what reduces bytes scanned. JSON is tempting as a semi-structured upgrade, but it does not deliver the columnar projection and compression that accelerate Athena queries.
- ✗
Increase the size of the CSV files to reduce the number of files.
Why it's wrong here
Larger CSV files still require Athena to scan and parse every row, and CSV lacks columnar compression and predicate pushdown. Converting to Parquet with partitioning prunes data read. Consolidating files is tempting because it reduces per-file overhead, but it does not reduce the bytes scanned, which drives Athena performance and cost.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on DEA-C01
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A data engineer is troubleshooting a slow-running Amazon Athena query. The query scans a large amount of data. Which TWO actions can improve query performance? (Choose TWO.)
medium- ✓ A.Convert the data to Parquet or ORC format.
- B.Enable encryption at rest.
- C.Increase the Athena query timeout.
- ✓ D.Partition the table on frequently filtered columns.
- E.Use SELECT * to retrieve all columns.
Why A: Option A is correct because converting data to columnar formats like Parquet or ORC lets Athena read only the columns referenced in the query and benefits from compression and predicate pushdown, drastically reducing the bytes scanned and thus improving performance and lowering cost. Option D is correct because partitioning the table on frequently filtered columns (e.g., date or region) enables partition pruning, so Athena skips scanning irrelevant partitions instead of reading the entire dataset. Option B is incorrect because encryption at rest protects stored data but does not reduce the amount of data scanned or speed up query execution. Option C is incorrect because increasing the query timeout only allows a slow query to run longer; it does not make the query faster. Option E is incorrect because using SELECT * retrieves all columns, which increases the data scanned and worsens performance, especially with columnar formats.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.