Courseiva
Data Store Management →mediumMultiple Choice

DEA-C01 Data Store Management Practice Question

A data engineer manages an Amazon S3 data lake that ingests millions of small JSON files daily from an IoT fleet. Query performance in Amazon Athena has degraded significantly, and each query scans far more data than expected. The engineer wants to reduce per-query cost and improve performance without changing the raw data. Which solution should the engineer implement?

⚠ Common exam trap

The trap here is assuming that a networking or workgroup setting such as S3 Transfer Acceleration or a higher data usage limit can fix slow Athena queries, when the real driver is bytes scanned and file layout.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert the JSON files to Apache Parquet, partition the data, and compact small files with AWS Glue ETL jobs.

Athena performance and cost are driven by how many bytes a query scans, so converting row-oriented JSON to columnar Parquet, partitioning on common filter columns, and compacting small files directly reduce scanned data. These three changes work together: columnar storage reads only referenced columns, partitions prune entire prefixes, and compaction removes per-object overhead. The result is faster queries at lower cost without altering the source data in S3.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Convert the JSON files to Apache Parquet, partition the data, and compact small files with AWS Glue ETL jobs.

    Why this is correct

    Athena charges by bytes scanned, so columnar Parquet with compression and partition pruning dramatically reduces data read. Compacting millions of tiny files into larger row groups removes per-object overhead and improves parallel scan throughput. Partitioning by a common filter column such as device date or region lets Athena skip irrelevant prefixes entirely, cutting both latency and cost.

  • ✗

    Move the data into an Amazon Redshift cluster and query it with Redshift Spectrum.

    Why it's wrong here

    Redshift Spectrum can query S3, but the small-file and row-oriented JSON problems would still apply to external tables, and provisioning a cluster adds cost and operational overhead. The scenario asks for improved Athena performance without changing raw data; migrating engines does not address file sizing or columnar layout. This is a disproportionate architectural change for a file-format issue.

  • ✗

    Increase the Athena workgroup data usage control limit and rerun the queries.

    Why it's wrong here

    Data usage controls in an Athena workgroup cap how many bytes a query or workgroup may scan before it is cancelled. Raising the limit only allows the same inefficient queries to run longer and cost more; it does not improve performance or reduce scanned bytes. This is a governance setting, not an optimization technique, so the underlying small-file and row-format problems remain.

  • ✗

    Enable S3 Transfer Acceleration on the ingestion bucket.

    Why it's wrong here

    S3 Transfer Acceleration speeds up uploads over long geographic distances by routing through edge locations. It does not reduce the number of objects, merge small files, or reduce the bytes scanned by Athena. Query performance and scan cost remain unchanged because Athena still reads every small object and applies no partition pruning. This addresses ingestion latency, not analytical scan efficiency.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.