Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data engineer is building an AWS Glue ETL job that reads raw JSON clickstream events from Amazon S3, flattens nested structures, and writes Parquet to a curated S3 prefix for SageMaker training. The job must run daily on only the newly arrived files and must keep the Glue Data Catalog table current so Athena and SageMaker can query it. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)

⚠ Common exam trap

The trap here is treating Glue job bookmarks as a performance tuning knob rather than the stateful mechanism that actually enables incremental reads of new S3 objects.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable job bookmarks on the Glue ETL job.

Job bookmarks give Glue the persistent state needed to process only newly arrived S3 objects, which is exactly the incremental requirement. Keeping the Data Catalog current, either through the job's catalog update option or a crawler over the curated prefix, ensures Athena and SageMaker see the latest schema and partitions. Together these two settings make the daily pipeline both efficient and queryable.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert the input JSON to CSV before processing to improve Glue read performance.

    Why it's wrong here

    Converting input to CSV adds an extra step and loses the nested structure that the job is specifically meant to flatten. Glue reads JSON natively, and the performance gain from CSV is not guaranteed to outweigh the added complexity and potential schema loss. This change does not address incremental processing or catalog updates, so it fails the requirements.

  • ✓

    Enable job bookmarks on the Glue ETL job.

    Why this is correct

    Job bookmarks persist state about which S3 objects and partitions have already been processed, so each daily run reads only newly arrived files. Without bookmarks, every run reprocesses the entire input prefix, which wastes DPU hours and can duplicate curated output. Bookmarks are the standard Glue mechanism for incremental processing and directly satisfy the requirement to handle only new data.

  • ✗

    Set the job's maximum capacity to the smallest possible value to reduce cost.

    Why it's wrong here

    Reducing maximum capacity may lower cost but does not provide incremental file processing or keep the Data Catalog synchronized. In fact, undersizing a job that flattens nested JSON and writes Parquet can cause long runtimes or out-of-memory failures. Capacity tuning is a performance and cost concern, not a mechanism for meeting the incremental-read and catalog-freshness requirements.

  • ✓

    Enable the Glue Data Catalog update option and configure crawlers or partition projection for the output prefix.

    Why this is correct

    Writing Parquet to S3 does not by itself register the new schema and partitions in the Data Catalog. Enabling the catalog update option in the Glue job, or running a crawler over the curated prefix, keeps the table definition and partition metadata current so Athena and SageMaker can query the latest data. This is required for the stated downstream query requirement.

  • ✗

    Use a Development endpoint to run the daily job interactively.

    Why it's wrong here

    Development endpoints are interactive environments for authoring and debugging scripts, not scheduled production execution. Running the daily job there provides no bookmark state, no reliable scheduling, and no catalog update behavior. It also incurs ongoing cost while idle. This option does not satisfy either the incremental processing or the catalog freshness requirement.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.