Courseiva
Data Store Management →easyMultiple Choice

DEA-C01 Data Store Management Practice Question

A data engineer is setting up an Amazon S3 bucket to store large CSV files that will be queried using Amazon Athena. The engineer wants to minimize query costs and improve performance. The files are currently stored in a single prefix without any partitioning. The most common queries filter data by `year` and `month`. What should the engineer do to optimize the Athena queries?

⚠ Common exam trap

The trap here is thinking that simply defining partition keys in the Glue Data Catalog without reorganizing the data will enable partition pruning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert the CSV files to Apache Parquet format and partition the data by `year` and `month`.

Converting data to Apache Parquet and partitioning by commonly filtered columns like `year` and `month` are the most effective ways to reduce data scanned by Athena. Parquet's columnar format allows Athena to read only the columns needed, and partitioning enables partition pruning. Together, they minimize query cost and improve performance. Other options either do not address the core optimizations or introduce unnecessary complexity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Compress the CSV files using gzip and keep them in a single prefix.

    Why it's wrong here

    Compressing CSV files with gzip reduces storage size and data scanned during queries, but it does not provide the same level of optimization as Parquet. CSV is row-based, so Athena still reads all columns even if only a few are needed. Also, without partitioning, queries filtering by `year` and `month` will scan all data. This approach is less effective than converting to a columnar format and partitioning.

  • ✓

    Convert the CSV files to Apache Parquet format and partition the data by `year` and `month`.

    Why this is correct

    Converting to Parquet reduces the amount of data scanned because Parquet is columnar and compressed, and partitioning by `year` and `month` allows Athena to prune partitions based on query filters. This significantly lowers query cost and improves performance. Athena charges based on data scanned, so both optimizations directly reduce cost. This is a best practice for optimizing Athena queries on large datasets.

  • ✗

    Create an AWS Glue Data Catalog table with partition keys `year` and `month`, and use Amazon Athena to query the CSV files directly.

    Why it's wrong here

    Creating a table with partition keys without actually organizing the data into partitioned prefixes will not enable partition pruning. The data must be stored in a directory structure like `year=2023/month=10/` for Athena to use partitions. Simply defining partition keys in the catalog does not reorganize the data. Therefore, queries will still scan all data, leading to higher costs and slower performance.

  • ✗

    Use Amazon Redshift Spectrum to query the CSV files in S3 instead of Athena.

    Why it's wrong here

    Amazon Redshift Spectrum allows querying S3 data from a Redshift cluster, but it is not a replacement for Athena in this scenario. Spectrum also benefits from columnar formats and partitioning. Using Spectrum would require an active Redshift cluster, adding cost and complexity. The goal is to optimize Athena queries, so this does not address the need. It also does not inherently optimize the CSV files for Athena.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.