Courseiva
Data Store Management →mediumMultiple Choice

DEA-C01 Data Store Management Practice Question

A data engineer maintains an Amazon S3 data lake with millions of small JSON objects. The engineer needs to improve query performance by reducing the number of objects and compressing them into a columnar format that Amazon Athena can query efficiently. The data must remain partitioned by date. Which solution should the engineer use?

⚠ Common exam trap

The trap here is assuming that simply merging files or changing storage classes solves the small-file problem, when the key is converting to a columnar format and repartitioning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use AWS Glue ETL to read the JSON objects, repartition by date, write them as Parquet with Snappy compression, and register the resulting tables in the AWS Glue Data Catalog.

Converting small JSON files to Parquet with Snappy compression and repartitioning by date reduces storage footprint and the amount of data scanned by Athena. AWS Glue ETL can perform this transformation and update the Data Catalog so Athena queries the new tables. This is the standard pattern for optimizing S3 data lakes for analytics.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable S3 Transfer Acceleration on the bucket and run Amazon Athena queries with the existing JSON objects.

    Why it's wrong here

    S3 Transfer Acceleration speeds up uploads and downloads over long distances, but it does not reduce the number of objects, convert to columnar format, or lower Athena scan costs. Athena still reads the same JSON data, so query performance and cost remain largely unchanged. It does not solve the small-file or format problem.

  • ✓

    Use AWS Glue ETL to read the JSON objects, repartition by date, write them as Parquet with Snappy compression, and register the resulting tables in the AWS Glue Data Catalog.

    Why this is correct

    AWS Glue ETL can transform JSON to Parquet, repartition by date, and update the Data Catalog, which Athena uses for schema and partition metadata. Parquet with Snappy reduces storage and scan size, and fewer larger files improve query performance. This directly addresses the small-file problem and columnar requirement while preserving date partitioning.

  • ✗

    Use Amazon S3 Lifecycle policies to transition the JSON objects to S3 Glacier Instant Retrieval and query them with Athena.

    Why it's wrong here

    S3 Glacier Instant Retrieval is an archival storage class for rarely accessed data. Moving active query data there increases retrieval costs and latency, and Athena supports it but it is not intended for frequent analytical queries. It also does not convert data to columnar format or reduce object count, so it fails the performance goal.

  • ✗

    Create an Amazon EMR cluster and run a MapReduce job that merges the JSON files into larger JSON files without changing the format.

    Why it's wrong here

    Merging small JSON files into larger JSON files reduces object count, but the data remains row-based JSON. Athena must still parse JSON and cannot take advantage of columnar pruning or compression. This approach may slightly improve performance but does not meet the columnar format requirement or optimize scan costs as effectively as Parquet.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.