MLA-C01 Data Preparation for Machine Learning Practice Question
A team is using AWS Glue to process streaming data from Amazon Kinesis. The streaming data contains both structured and semi-structured fields. The team needs to flatten the semi-structured fields into columns for downstream ML training. Which Glue feature is BEST suited?
⚠ Common exam trap
A common mix-up: candidates confuse 'flattening semi-structured data' with simple schema operations like type resolution or column mapping, leading them to choose ResolveChoice or ApplyMapping instead of the specialized Relationalize transform.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Relationalize transform
The Relationalize transform is specifically designed to flatten nested JSON or semi-structured fields into a relational structure, making it ideal for converting complex streaming data from Kinesis into flat columns for ML training. It automatically handles arrays and structs by creating separate tables or columns, which is exactly what the team needs for downstream processing.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Relationalize transform
Why this is correct
Relationalize flattens nested semi-structured data into separate relational tables that Glue can join, converting arrays and structs into columns. This directly satisfies the requirement to expose semi-structured Kinesis fields as flat columns for downstream ML training.
- ✗
Spigot transform
Why it's wrong here
Spigot is a debugging tool that samples and writes records to output for inspection; it performs no flattening of nested structures. It is the right choice when validating or troubleshooting a Glue job's data flow, not when transforming semi-structured fields into columns.
- ✗
ResolveChoice transform
Why it's wrong here
ResolveChoice resolves type conflicts within a DynamicFrame column (for example, casting between int and string), not nested structures. Flattening semi-structured fields requires Relationalize, which un-nests and pivots them into separate tables. ResolveChoice is tempting when schemas contain mixed types, but it leaves nested data intact.
- ✗
ApplyMapping transform
Why it's wrong here
ApplyMapping renames, casts and drops existing fields; it cannot explode nested arrays or structs into separate top-level columns. It is the right transform when the schema is already flat and only field names or data types need changing, not for flattening semi-structured data.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.