DEA-C01 Columnar Storage Formats Practice Question
An e-commerce company is building a near-real-time dashboard to monitor customer clickstream data. The data is ingested via Amazon Kinesis Data Streams, transformed using AWS Lambda, and stored in Amazon S3. The team needs to query the data using Amazon Athena. Which THREE steps should be taken to optimize cost and performance? (Choose three.)
⚠ Common exam trap
A common trap is to consider AWS Glue Data Catalog as an optimization step, but it is merely a requirement; the actual optimizations are compression, partitioning, and columnar formats.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the data to Apache Parquet or ORC format.
Option C is correct because converting the clickstream data to columnar formats like Apache Parquet or ORC lets Athena read only the columns referenced in each query and applies columnar compression and predicate pushdown, dramatically reducing bytes scanned and therefore cost and latency. Option D is correct because compressing data with gzip or snappy reduces the amount of data Athena must read from S3, lowering query cost and improving performance; snappy is especially effective with Parquet/ORC since it is splittable and column-oriented. Option E is correct because partitioning the S3 data by date (e.g., year/month/day) enables Athena partition pruning so queries that filter on time only scan the relevant prefixes instead of the entire dataset. Option A is not among the marked answers: while the AWS Glue Data Catalog is the standard metastore Athena uses, defining table metadata is a prerequisite for querying rather than a cost/performance optimization step for the data itself. Option B is not correct because storing data as JSON is row-oriented and non-columnar, which forces Athena to scan and parse entire records, increasing bytes scanned and query cost compared with Parquet or ORC.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use AWS Glue Data Catalog to store the table metadata.
Why it's wrong here
The AWS Glue Data Catalog is the correct metadata store for Athena tables, so this step is required, not a fault. It defines schemas, partitions, and SerDe information that Athena uses to query S3 data, and it integrates with crawlers to keep table definitions current as new partitions arrive.
- ✗
Store the data in JSON format for flexibility.
Why it's wrong here
JSON is row-oriented and unparsed, so Athena reads every column for each query, inflating bytes scanned and cost; columnar formats like Parquet let Athena prune columns. JSON suits flexible, schema-evolving landing data, but for repeated analytical queries over clickstream the correct choice is a columnar, compressed format.
- ✓
Convert the data to Apache Parquet or ORC format.
Why this is correct
Columnar Parquet or ORC lets Athena read only the columns each query references rather than every field, and columnar encoding compresses better. This directly reduces bytes scanned, which is the metric Athena bills on, improving both cost and latency.
- ✓
Compress the data using gzip or snappy.
Why this is correct
Compressing stored objects with gzip or Snappy shrinks the byte volume Athena must read from S3. Since Athena charges per terabyte scanned, smaller compressed files lower query cost, and less data transferred means faster query completion.
- ✓
Partition the data by date in S3 (e.g., year/month/day).
Why this is correct
Partitioning by year/month/day lets Athena prune irrelevant S3 prefixes using partition metadata, so queries scanning a date range skip all other objects. This slashes bytes scanned, cutting cost and latency for the dashboard's time-bounded queries.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.