Courseiva
Data Store Management →mediumMultiple Choice

DEA-C01 Data Store Management Practice Question

A data engineer manages an Amazon S3 data lake with millions of small JSON files. To improve query performance with Amazon Athena, the engineer wants to compact these files into larger Parquet files. The engineer must also minimize ongoing storage costs. Which solution should the engineer implement?

⚠ Common exam trap

The trap here is assuming that S3 Transfer Acceleration or streaming services like Kinesis Data Firehose can optimize existing batch data in S3, when they are designed for different use cases.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use AWS Glue ETL jobs to read the JSON files, transform them to Parquet, and write the output to a new S3 prefix. Then, configure an S3 Lifecycle rule to expire the original JSON objects after a retention period.

Converting JSON to Parquet with AWS Glue ETL compacts files and enables columnar storage, which improves Athena query performance. Writing to a new prefix and using an S3 Lifecycle rule to expire original JSON files after a retention period reduces storage costs while preserving data integrity. This combination addresses both performance and cost requirements.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create an AWS Lambda function that triggers on each S3 PUT event to merge JSON files into larger files.

    Why it's wrong here

    A Lambda function triggered per PUT event would process only new files, not the existing millions of small files. Additionally, merging JSON files into larger JSON files does not provide the columnar storage benefits of Parquet, so Athena performance may not improve significantly. This approach also adds complexity and potential throttling.

  • ✓

    Use AWS Glue ETL jobs to read the JSON files, transform them to Parquet, and write the output to a new S3 prefix. Then, configure an S3 Lifecycle rule to expire the original JSON objects after a retention period.

    Why this is correct

    AWS Glue ETL can efficiently convert JSON to Parquet, reducing file count and improving Athena query performance. Writing to a new prefix preserves the original data until verified. An S3 Lifecycle rule to expire the original JSON objects after a retention period reduces storage costs without immediate data loss. This approach is scalable and aligns with best practices for data lake optimization.

  • ✗

    Use Amazon Kinesis Data Firehose to stream the JSON files into Parquet format in real time.

    Why it's wrong here

    Kinesis Data Firehose can convert incoming streaming data to Parquet, but the data is already stored in S3 as historical JSON files. Firehose is designed for streaming ingestion, not for batch transformation of existing S3 objects. This solution would not process the existing small files and would not reduce storage costs.

  • ✗

    Enable S3 Transfer Acceleration on the bucket to speed up read operations for Athena queries.

    Why it's wrong here

    S3 Transfer Acceleration speeds up uploads and downloads over long distances by using AWS edge locations, but it does not improve Athena query performance on small files. Athena performance is primarily affected by file size, format, and partitioning. This option does not address the root cause of many small files, nor does it reduce storage costs.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.