Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer needs to run a transformation on a large dataset stored in Amazon S3 using AWS Glue Studio. The transformation is a simple column rename and filter that can be expressed visually. The engineer wants to minimize development time and avoid writing PySpark code. Which approach should the engineer use?

⚠ Common exam trap

The trap here is assuming that any visual data preparation tool, such as DataBrew, is interchangeable with Glue Studio's visual job editor for building ETL jobs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a Glue Studio visual job, add a source node for the S3 bucket, add a Transform node for the rename and filter, and add a target node for the output location.

Glue Studio's visual job editor lets engineers build ETL pipelines by connecting source, transform, and target nodes without writing code. For a simple column rename and filter, this is the fastest path because Glue generates the Spark code and manages execution. The other options either require manual coding, use a different tool not integrated with Glue Studio jobs, or use query-based approaches that do not provide the same visual job experience.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use AWS Glue DataBrew to create a recipe that renames columns and filters rows, then run a DataBrew job.

    Why it's wrong here

    DataBrew is a visual data preparation tool for profiling and cleaning data, and it can rename columns and filter rows. However, DataBrew is oriented toward data preparation and profiling rather than ETL job orchestration within Glue Studio, and it does not integrate with Glue Studio's job graph. For a Glue Studio transformation job, the visual job editor is the native and more direct choice, so DataBrew is not the best answer here.

  • ✗

    Use an AWS Glue crawler to infer the schema, then query the data with Amazon Athena and save the results to a new S3 location.

    Why it's wrong here

    A Glue crawler populates the Data Catalog but does not perform transformation logic such as column renames and filters. Athena can run SQL queries that rename and filter, but saving query results to S3 is not an ETL job and does not provide the visual Glue Studio workflow. This approach also requires SQL authoring and does not meet the goal of a visual, low-code transformation job.

  • ✗

    Develop an AWS Glue ETL script in PySpark using the DynamicFrame API, then run it as a job with the 'glueetl' job type.

    Why it's wrong here

    Writing a PySpark script with the DynamicFrame API is a valid way to perform the transformation, but it requires manual coding and testing, which contradicts the goal of minimizing development time and avoiding PySpark code. The visual editor exists specifically to avoid this effort for straightforward transformations. While this approach is flexible, it is not the best fit for a simple rename and filter expressed visually.

  • ✓

    Create a Glue Studio visual job, add a source node for the S3 bucket, add a Transform node for the rename and filter, and add a target node for the output location.

    Why this is correct

    Glue Studio provides a visual job editor with drag-and-drop nodes for sources, transforms, and targets. A visual job can implement column renames and filters without writing code, and Glue generates the underlying PySpark script automatically. This directly meets the requirement to minimize development time and avoid manual PySpark coding, while still running on the Glue Spark engine for scale.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.