MLA-C01 Data Preparation for Machine Learning Practice Question
A company is building a time series forecasting model using SageMaker DeepAR. The raw data is a CSV with columns: timestamp, item_id, and value. What is the correct data format required for DeepAR training?
⚠ Common exam trap
Many candidates assume DeepAR can accept raw CSV data like other SageMaker built-in algorithms (e.g., XGBoost), but DeepAR is a specialized time series algorithm that requires a specific JSON Lines structure with 'start' and 'target' fields, not a simple tabular format.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
JSON Lines files with 'start', 'target', and optional fields per time series
DeepAR requires time series data to be provided in JSON Lines format, where each line represents a single time series with a 'start' timestamp (in ISO 8601 format), a 'target' array of values, and optional fields like 'cat' for categorical features. This structured format allows DeepAR to handle variable-length sequences and missing values natively, which is not possible with simple CSV or wide-format data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
JSON Lines files with 'start', 'target', and optional fields per time series
Why this is correct
DeepAR consumes JSON Lines, where each line holds one time series with a 'start' timestamp and a 'target' array of values. The stem's CSV columns must therefore be pivoted into that per-series structure; CSV is not accepted directly by the algorithm.
- ✗
A wide-format CSV where each column is a different time series
Why it's wrong here
DeepAR requires JSON Lines, where each line holds one time series with start, target and optional cat fields. Wide-format CSV is tempting because classic forecasting tools such as Prophet accept it, but DeepAR's training channel rejects that layout.
- ✗
Parquet files with a schema containing timestamp, item_id, and value
Why it's wrong here
DeepAR's training channel accepts JSON Lines or Parquet, but the schema must nest each series as start, target and cat, not flat timestamp/item_id/value columns. Flat Parquet is tempting because Parquet is supported, yet that column layout fails ingestion.
- ✗
A single CSV file with columns: timestamp, item_id, value
Why it's wrong here
A flat CSV with timestamp, item_id and value is not accepted; DeepAR needs JSON Lines with start, target and cat, or Parquet with that nested schema. The CSV is tempting because it mirrors the raw input, but the training channel rejects it.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.