Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A company is building a data lake on Amazon S3. Data arrives from multiple sources in JSON, CSV, and Avro formats. The data must be transformed to Parquet and partitioned by date and source. Which TWO services can perform this transformation with minimal custom code? (Choose TWO.)

⚠ Common exam trap

Many candidates confuse AWS Lake Formation's data catalog and permission features with actual data transformation capabilities, or they assume Kinesis Data Firehose can transform existing S3 objects when it only processes streaming data in transit.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Amazon EMR with Spark

Amazon EMR with Spark (A) is correct because Spark natively reads JSON, CSV, and Avro, and can write Parquet with partitionBy('date','source') using built-in DataFrame APIs, requiring only a small script rather than a custom transformation engine. AWS Glue ETL jobs (D) are correct because Glue provides managed, serverless Spark with built-in DynamicFrame readers/writers for JSON, CSV, and Avro, plus automatic schema inference and native Parquet output with partition keys, minimizing custom code. AWS Lake Formation (B) is a permission and metadata/catalog governance layer, not a data transformation engine. Amazon Athena CTAS (C) can convert formats and partition results, but it is query-oriented and less suited to general multi-format ETL pipelines. Amazon Kinesis Data Firehose (E) is a streaming delivery service that can convert to Parquet via Glue schema, but it does not perform the required multi-source batch transformation and partitioning logic.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Amazon EMR with Spark

    Why this is correct

    Amazon EMR with Spark provides native readers for JSON, CSV and Avro plus a Parquet writer, and supports partitionBy on date and source columns. Transformations are expressed declaratively in Spark SQL or DataFrames, requiring little bespoke code.

  • ✗

    AWS Lake Formation

    Why it's wrong here

    Lake Formation governs catalog permissions, crawlers and access control; it does not convert JSON, CSV or Avro into Parquet or write date/source partitions. It is tempting because it manages S3 data lakes, and would be correct for securing and granting fine-grained access to already-transformed tables.

  • ✗

    Amazon Athena CTAS queries

    Why it's wrong here

    Athena CTAS converts formats and writes partitioned Parquet, but requires hand-written SQL per source schema, which is custom code the stem seeks to minimise. It is tempting because it transforms S3 data in place, and would be correct for ad-hoc one-off conversions where SQL authoring is acceptable.

  • ✓

    AWS Glue ETL jobs

    Why this is correct

    AWS Glue ETL jobs include built-in transforms and crawlers that read JSON, CSV and Avro, convert to Parquet, and write partitioned output by date and source. The work is defined through the console or generated PySpark, minimising custom code.

  • ✗

    Amazon Kinesis Data Firehose

    Why it's wrong here

    Firehose transforms only streaming records in flight using Lambda or mapping, so it cannot convert existing JSON, CSV and Avro objects already landed in S3. It is tempting because it delivers partitioned Parquet to S3, and would be correct when data arrives continuously as a stream rather than as stored files.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.