DEA-C01 Columnar Storage Formats Practice Question
An e-commerce company is building a near-real-time dashboard to monitor customer clickstream data. The data is ingested via Amazon Kinesis Data Streams, transformed using AWS Lambda, and stored in Amazon S3. The team needs to query the data using Amazon Athena. Which THREE steps should be taken to optimize cost and performance? (Choose three.)
⚠ Common exam trap
A common trap is to consider AWS Glue Data Catalog as an optimization step, but it is merely a requirement; the actual optimizations are compression, partitioning, and columnar formats.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the data to Apache Parquet or ORC format.
To optimize cost and performance when querying data with Athena, use columnar formats like Parquet or ORC (C) to reduce data scanned and improve compression. Compress data with gzip or Snappy (D) to reduce storage costs and data transferred during queries. Partition data by date (E) to limit the amount of data scanned per query. Option A (Glue Data Catalog) is a prerequisite, not an optimization step. Option B (JSON) is less efficient than columnar formats for analytical queries.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use AWS Glue Data Catalog to store the table metadata.
Why it's wrong here
Required but not an optimization step; it's foundational.
- ✗
Store the data in JSON format for flexibility.
Why it's wrong here
JSON is not optimized for query performance; columnar formats are better.
- ✓
Convert the data to Apache Parquet or ORC format.
Why this is correct
Columnar formats reduce data scanned and improve compression.
- ✓
Compress the data using gzip or snappy.
Why this is correct
Compression reduces storage and scanning costs.
- ✓
Partition the data by date in S3 (e.g., year/month/day).
Why this is correct
Partitioning allows Athena to scan only relevant partitions, reducing cost.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,711 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.