Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data scientist must join a 50 GB transactional table with a small 5 MB lookup table in AWS Glue before writing Parquet output for SageMaker training. The join currently shuffles the large table across the cluster and the job runs slowly. Which optimization should the data scientist apply?

⚠ Common exam trap

The trap here is responding to a slow Spark join by adding cluster capacity, when the real fix is changing the join strategy so the shuffle disappears entirely.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure a broadcast join by passing the small table to the join as a broadcast hint.

A broadcast join sends the small lookup table to every executor so the large table can be streamed locally without a network shuffle. Because the lookup table is only a few megabytes, it fits comfortably in executor memory. Increasing cluster size, repartitioning the large table, or reshaping nested data all leave the expensive shuffle in place, so none of them addresses the actual bottleneck.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Configure a broadcast join by passing the small table to the join as a broadcast hint.

    Why this is correct

    When one side of a join is small enough to fit in memory, broadcasting it to every executor avoids shuffling the large table across the network. In AWS Glue with Spark, this is expressed through a broadcast hint or by relying on the automatic broadcast threshold, and it directly removes the shuffle that is slowing the job.

  • ✗

    Convert the small lookup table to a Glue DynamicFrame and apply a relationalize transform.

    Why it's wrong here

    Relationalize is intended for flattening nested, semi-structured data such as JSON arrays into relational tables. It has no bearing on join execution strategy and would not reduce the shuffle of the large table. This transform addresses a schema-shaping problem, not the performance issue described.

  • ✗

    Repartition both tables on the join key before the join operation.

    Why it's wrong here

    Repartitioning the large table still requires a full shuffle of its 50 GB of data, which is precisely the expensive operation observed. Co-partitioning helps only when both sides are comparably large and already partitioned consistently. For a tiny lookup table, this approach adds cost without removing the bottleneck.

  • ✗

    Increase the number of DPUs allocated to the Glue job so more executors perform the shuffle.

    Why it's wrong here

    Adding DPUs scales the cluster but does not eliminate the shuffle that moves the 50 GB table between executors. The network and disk cost of the shuffle persists and can even grow with more partitions. More capacity treats the symptom rather than the join strategy causing the slowdown.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.