Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A media company is building a data pipeline to ingest user activity logs from multiple sources into Amazon S3. The logs are JSON files generated every minute. The company wants to use Amazon Athena to query the logs with minimal latency and cost. The current approach is to use AWS Kinesis Data Firehose to deliver the logs to S3 with a prefix like 'logs/2024/01/01/00/file.json'. However, when running Athena queries, the team notices high query costs because Athena scans all files in the 'logs/' prefix even when querying for a specific date. What should the team do to reduce the amount of data scanned by Athena?

⚠ Common exam trap

DEA-C01 often tests the difference between S3 prefix organization and true Hive-style partitioning — candidates pick 'more granular prefix' thinking it enables pruning, but without Glue Data Catalog partition registration, Athena still scans all files.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a Hive-style partition structure in S3 with keys like 'year=2024/month=01/day=01/hour=00/' and update the Glue Data Catalog accordingly.

Athena reduces data scanned by using partition pruning, which requires a Hive-style partition structure in S3 (e.g., year=2024/month=01/day=01/hour=00/) registered in the Glue Data Catalog. When queries filter on partition columns, Athena only reads the relevant partitions instead of scanning the entire logs/ prefix. This directly addresses the high query cost caused by scanning all files.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create an Athena view that filters by date.

    Why it's wrong here

    An Athena view is a stored query; it does not change how the underlying JSON objects are read, so scanning remains identical. Partition pruning requires the partition columns to be registered in the table metadata. Views suit simplifying repeated query logic, not reducing bytes scanned.

  • ✗

    Increase the number of partitions by using a more granular prefix like 'logs/2024/01/01/00/00/'.

    Why it's wrong here

    Adding an hour-level prefix does not reduce scanning unless the table's partition keys match that S3 layout and queries filter on them. Without matching partition metadata, Athena still lists and reads every object. Granular prefixes suit workloads already using partitioned tables with matching filters.

  • ✗

    Convert the JSON files to Apache Parquet format using AWS Glue ETL jobs.

    Why it's wrong here

    Parquet reduces bytes scanned per column, but the stem's cost stems from reading every date's files, which columnar compression does not fix. Partition projection or partition metadata enables pruning. Parquet conversion is right when queries read few columns from wide tables.

  • ✓

    Create a Hive-style partition structure in S3 with keys like 'year=2024/month=01/day=01/hour=00/' and update the Glue Data Catalog accordingly.

    Why this is correct

    Hive-style partitioning splits data into year/month/day/hour prefixes, letting Athena prune irrelevant partitions via partition projection or the Glue Data Catalog. Queries for one date then scan only that partition rather than every file under logs/, cutting both cost and latency.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.