Courseiva
Data Store Management →mediumMultiple Choice

DEA-C01 Data Store Management Practice Question

A data engineer stores Apache Parquet files in an Amazon S3 data lake partitioned by dt=YYYY-MM-DD. Analysts query the data with Amazon Athena, and monthly reports that scan one month of data are slow and expensive. The engineer confirms that queries filter on the dt column. Which action will MOST effectively reduce the amount of data scanned by these reports?

⚠ Common exam trap

The trap here is assuming that switching file formats or tuning the workgroup fixes slow Athena queries, when the real issue is unregistered partitions that disable pruning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run MSCK REPAIR TABLE on the table, or enable partition projection for the dt partition column.

Athena reduces scan size through partition pruning, which requires the table's partition metadata to reflect the actual S3 prefixes. When files are written as dt=YYYY-MM-DD without catalog registration, filters on dt cannot eliminate prefixes, so monthly reports scan the entire table. Registering partitions with MSCK REPAIR TABLE or defining partition projection restores pruning, cutting both bytes scanned and cost.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert the Parquet files to CSV so that Athena can read them faster.

    Why it's wrong here

    CSV is a row-based, uncompressed text format that prevents Athena from reading only the needed columns. Converting to CSV would increase the bytes scanned and cost for these monthly reports, and it discards the columnar benefits Parquet already provides. The slowness is caused by partition layout, not by the file format, so this change makes the scenario worse rather than better.

  • ✗

    Enable Amazon S3 Transfer Acceleration on the data lake bucket.

    Why it's wrong here

    Transfer Acceleration speeds up uploads and downloads over long geographic distances by routing through edge locations. Athena reads data inside the AWS network, so this feature does not reduce the volume of bytes a query scans and does not affect partition pruning. It would add cost without changing the monthly report's scan size or runtime.

  • ✓

    Run MSCK REPAIR TABLE on the table, or enable partition projection for the dt partition column.

    Why this is correct

    Athena prunes partitions only when the table's partition metadata matches the S3 prefixes. If new dt= prefixes were written without registering partitions in the AWS Glue Data Catalog, every query falls back to scanning the whole table. Running MSCK REPAIR TABLE, or configuring partition projection so Athena derives partitions from a pattern, restores pruning so a one-month filter reads only that month's prefixes.

  • ✗

    Increase the Athena per-query data limit in the workgroup settings.

    Why it's wrong here

    The workgroup data-usage control is a guardrail that cancels queries exceeding a byte threshold; raising it permits larger scans but never makes a query read less. Since the reports are slow and expensive because they scan too much, loosening this limit would keep the same scan volume and cost. It addresses the symptom's cap, not the cause.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.