Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is designing a pipeline that ingests JSON logs from an application into Amazon S3. The logs contain a timestamp field. The pipeline must partition the data by date in S3 (e.g., year=2024/month=10/day=01). Which approach minimizes transformation effort?

⚠ Common exam trap

Test-takers frequently confuse metadata partitioning (e.g., using Glue crawlers or Athena) with physical partitioning in S3, assuming that catalog operations alone reorganize the data, when in fact only ingestion-time partitioning (like Firehose dynamic partitioning) creates the folder structure without extra transformation effort.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Amazon Kinesis Data Firehose with dynamic partitioning

Amazon Kinesis Data Firehose with dynamic partitioning can automatically partition incoming JSON data based on the timestamp field without requiring custom transformation code. It evaluates the timestamp using a JQ expression or inline parsing, then writes records directly to S3 prefixes like year=2024/month=10/day=01. This minimizes transformation effort because the partitioning logic is configured declaratively in the Firehose delivery stream, eliminating the need for Lambda functions or post-ingestion processing.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use Amazon Kinesis Data Firehose with dynamic partitioning

    Why this is correct

    Kinesis Data Firehose dynamic partitioning extracts the timestamp field via a jq expression and writes records into year=/month=/day=/ prefixes automatically, so no downstream ETL job is needed to reorganise objects. This directly satisfies the stem's requirement to minimise transformation effort while delivering date-partitioned JSON into Amazon S3.

  • ✗

    Use AWS Glue crawlers to infer schema and create partitions

    Why it's wrong here

    Glue crawlers infer schema and register partitions in the Data Catalog; they do not physically write objects into year=/month=/day=/ prefixes, so the S3 layout requirement stays unmet. It tempts because crawlers are the standard way to make existing partitioned data queryable, not to produce that partitioning.

  • ✗

    Use AWS Lambda to process each object and copy to the appropriate prefix

    Why it's wrong here

    Lambda per-object copying rewrites every object, adding code and compute to derive date prefixes the writer could emit directly. It tempts when transformation is genuinely unavoidable, such as converting formats or enriching records before landing, but here the timestamp already exists in each JSON log.

  • ✗

    Use Amazon Athena to create partitions on the existing data

    Why it's wrong here

    Athena partitions are metadata registered over existing S3 prefixes; they neither create the year=/month=/day=/ folders nor reorganise objects, so the required layout never materialises. It tempts for querying already-partitioned data cheaply, which is a read-side concern, not ingestion-time partitioning.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.