Courseiva

DEA-C01 Data Operations and Support Practice Question

A data engineer is using Amazon Athena to query data stored in Amazon S3 in Parquet format. The engineer notices that a specific query is scanning much more data than expected, resulting in high costs and slow performance. The query filters on a column named 'event_date' which is a string in 'YYYY-MM-DD' format. The table is partitioned by 'year', 'month', and 'day' as separate string columns. The engineer wants to reduce the amount of data scanned. Which action should the engineer take?

⚠ Common exam trap

The trap here is assuming that converting a column to date type or enabling caching will reduce data scanned, when the key optimization is to use the existing partition columns in the filter.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the partition columns 'year', 'month', and 'day' in the WHERE clause instead of 'event_date'.

The table is partitioned by year, month, and day, so filtering on those partition columns in the WHERE clause enables Athena's partition pruning. This limits the data scanned to only the relevant partitions, significantly reducing cost and improving performance. Filtering on a non-partition column like event_date does not provide the same benefit.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Compress the Parquet files using Snappy compression.

    Why it's wrong here

    Compressing Parquet files reduces storage size and can lower the amount of data scanned because Athena reads compressed data, but it does not eliminate scanning of irrelevant partitions. The primary issue is that the query is not leveraging partitions, so compression alone will not solve the problem. It is a good practice but secondary to partition pruning.

  • ✗

    Enable Amazon Athena query result reuse to cache the results.

    Why it's wrong here

    Query result reuse caches the results of previously executed queries, which can reduce cost for repeated identical queries. However, it does not reduce the data scanned for a new query with different filters. Since the issue is a specific query scanning too much data, caching would not help unless the exact same query is run again. It is not a solution for the current query's inefficiency.

  • ✗

    Convert the 'event_date' column to a date type and use it in the WHERE clause.

    Why it's wrong here

    Converting the column type may improve predicate pushdown if the query uses date functions, but it does not leverage the existing partitions. The partitions are on year, month, and day, so filtering on those columns directly is more effective. Changing the data type alone will not reduce the scanned data if the query does not filter on partition columns. It could also require rewriting data.

  • ✓

    Use the partition columns 'year', 'month', and 'day' in the WHERE clause instead of 'event_date'.

    Why this is correct

    Athena uses partition pruning to limit the data scanned based on partition columns in the WHERE clause. Since the table is partitioned by year, month, and day, filtering on these columns allows Athena to scan only the relevant partitions, drastically reducing data scanned. This is the most effective way to optimize the query and reduce costs.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.