Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A team is using AWS Glue to process streaming data from Amazon Kinesis. The streaming data contains both structured and semi-structured fields. The team needs to flatten the semi-structured fields into columns for downstream ML training. Which Glue feature is BEST suited?

⚠ Common exam trap

A common mix-up: candidates confuse 'flattening semi-structured data' with simple schema operations like type resolution or column mapping, leading them to choose ResolveChoice or ApplyMapping instead of the specialized Relationalize transform.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Relationalize transform

The Relationalize transform is specifically designed to flatten nested JSON or semi-structured fields into a relational structure, making it ideal for converting complex streaming data from Kinesis into flat columns for ML training. It automatically handles arrays and structs by creating separate tables or columns, which is exactly what the team needs for downstream processing.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Relationalize transform

    Why this is correct

    Relationalize flattens nested semi-structured data into separate relational tables that Glue can join, converting arrays and structs into columns. This directly satisfies the requirement to expose semi-structured Kinesis fields as flat columns for downstream ML training.

  • ✗

    Spigot transform

    Why it's wrong here

    Spigot is a debugging tool that samples and writes records to output for inspection; it performs no flattening of nested structures. It is the right choice when validating or troubleshooting a Glue job's data flow, not when transforming semi-structured fields into columns.

  • ✗

    ResolveChoice transform

    Why it's wrong here

    ResolveChoice resolves type conflicts within a DynamicFrame column (for example, casting between int and string), not nested structures. Flattening semi-structured fields requires Relationalize, which un-nests and pivots them into separate tables. ResolveChoice is tempting when schemas contain mixed types, but it leaves nested data intact.

  • ✗

    ApplyMapping transform

    Why it's wrong here

    ApplyMapping renames, casts and drops existing fields; it cannot explode nested arrays or structs into separate top-level columns. It is the right transform when the schema is already flat and only field names or data types need changing, not for flattening semi-structured data.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.