Courseiva
Data Store Management →hardMultiple Choice

DEA-C01 Data Store Management Practice Question

A media company ingests millions of small JSON files per day into an Amazon S3 bucket. Analysts run Amazon Athena queries over this data and report that each query scans far more data than the files matching their filters, resulting in high cost and slow performance. The files are partitioned by year/month/day in S3. What should a data engineer do to reduce the data scanned per query?

⚠ Common exam trap

The trap here is focusing on ingestion or delivery speedups such as Transfer Acceleration or CloudFront, which do not change how many bytes Athena reads during a query.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert the JSON files to Apache Parquet, store them in the existing partition structure, and define the table in the AWS Glue Data Catalog with the correct partition columns.

Athena's cost model is based on bytes scanned, so the most effective optimization is to store data in a columnar format such as Parquet and rely on partition pruning. Parquet enables column projection and predicate pushdown, while year/month/day partitions let Athena skip entire prefixes. Together they sharply reduce scanned bytes, cutting cost and improving query speed for the analyst workload.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Move the files into an Amazon Redshift cluster using COPY and query them with Redshift Spectrum.

    Why it's wrong here

    This adds a cluster to manage and does not by itself reduce scanned data unless the files are also columnar and partitioned. COPY of millions of tiny JSON files is slow and inefficient. The requirement is to reduce Athena scan volume, and this approach introduces cost and complexity without guaranteeing that outcome.

  • ✗

    Create an Amazon CloudFront distribution in front of the S3 bucket and run Athena against the distribution.

    Why it's wrong here

    CloudFront caches and delivers objects over HTTP; Athena reads directly from S3 using the S3 API and cannot query through a CloudFront distribution. Caching also does nothing to reduce the bytes scanned by columnar or partition pruning. This option misunderstands how Athena accesses data.

  • ✗

    Enable S3 Transfer Acceleration on the bucket and increase the Athena query timeout.

    Why it's wrong here

    Transfer Acceleration speeds up uploads to S3 over long distances and has no effect on how much data Athena scans during a query. Query timeout is not a configurable setting that reduces scanned bytes. This option addresses ingestion performance, not the cost and scan-volume problem described.

  • ✓

    Convert the JSON files to Apache Parquet, store them in the existing partition structure, and define the table in the AWS Glue Data Catalog with the correct partition columns.

    Why this is correct

    Athena charges by data scanned, and columnar Parquet with predicate pushdown lets the engine read only the columns and row groups needed. Combined with partition pruning on year/month/day, queries skip irrelevant prefixes entirely. This directly reduces bytes scanned, lowering cost and improving latency without changing the query interface.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.