Courseiva
Data Store Management →mediumMultiple Choice

DEA-C01 Data Store Management Practice Question

A data engineer is designing a data lake on Amazon S3 for a financial analytics workload. The raw data arrives as JSON files from an on-premises system. Analysts need to query the data using Amazon Athena with fast performance and minimal cost for queries that filter on a specific transaction date and customer ID. The engineer wants to convert the data to a columnar format that supports predicate pushdown and compression. Which storage format should the engineer choose?

⚠ Common exam trap

The trap here is assuming that any compressed format will improve Athena performance, when only columnar formats like Parquet or ORC enable column pruning and predicate pushdown.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apache Parquet

Apache Parquet is a columnar storage format that allows Athena to read only the columns referenced in a query and to skip row groups using predicate pushdown. This reduces the amount of data scanned, lowering cost and improving performance for filters on transaction date and customer ID. Parquet also compresses efficiently, further reducing storage and scan volume.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apache Avro

    Why it's wrong here

    Apache Avro is a row-based format optimized for write-heavy workloads and schema evolution, not for columnar analytical queries. Athena would read entire rows, scanning unnecessary columns and increasing data processed. While Avro supports compression, it lacks the columnar predicate pushdown benefits needed for filtering on specific columns. It is better suited for streaming ingestion than for interactive analytics.

  • ✓

    Apache Parquet

    Why this is correct

    Apache Parquet is a columnar format that stores data by column, enabling Athena to read only the columns needed for a query. It supports predicate pushdown, so filters on transaction date and customer ID skip irrelevant data. Parquet also compresses well, reducing data scanned and query cost. This directly meets the requirement for fast, cost-effective queries on filtered columns.

  • ✗

    JSON

    Why it's wrong here

    JSON is a text-based, row-oriented format with high storage overhead and no built-in columnar optimizations. Athena can query JSON, but it must parse every record and cannot skip columns, leading to higher costs and slower performance. JSON also does not support efficient predicate pushdown on nested fields. Keeping raw JSON would not satisfy the performance and cost goals.

  • ✗

    CSV

    Why it's wrong here

    CSV is a row-based text format that lacks compression and columnar pruning. Athena scans the entire file for each query, increasing data processed and cost. CSV also has no schema enforcement or predicate pushdown capabilities. While simple, it is inefficient for analytical queries that filter on specific columns like transaction date and customer ID.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.