Courseiva
mediumMultiple Select

MLA-C01 Practice Question: Building a recommender system using implicit…

A company is building a recommender system using implicit feedback (clicks) and explicit feedback (ratings). They plan to use Amazon SageMaker to train a model. The data includes user ID, item ID, timestamp, and rating (if any). Which TWO data preparation steps should the team perform? (Choose TWO.)

⚠ Common exam trap

Test-takers frequently confuse one-hot encoding (which is common in linear models) with the integer indexing required for embedding-based models like matrix factorization, leading them to select Option D instead of Option A.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert user ID and item ID to integer indices for matrix factorization

Matrix factorization algorithms in Amazon SageMaker (e.g., the built-in Factorization Machines algorithm or the Apache Spark-based collaborative filtering) require user and item identifiers to be converted to contiguous integer indices starting from 0. This is necessary for efficient embedding lookup and to avoid memory blowup from sparse categorical features. SageMaker's implementation expects the input data in recordIO-wrapped protobuf format with integer-encoded user and item columns.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Convert user ID and item ID to integer indices for matrix factorization

    Why this is correct

    Matrix factorization algorithms (e.g., in SageMaker's built-in Factorization Machines) require user and item IDs as integers.

  • ✗

    Use target encoding on user ID based on average rating

    Why it's wrong here

    Target encoding can cause data leakage in recommendation contexts and is not typical for matrix factorization.

  • ✗

    Normalize ratings using StandardScaler

    Why it's wrong here

    Normalizing ratings is not standard and may not improve performance; the model can learn the scale naturally.

  • ✗

    One-hot encode user ID and item ID

    Why it's wrong here

    One-hot encoding user and item IDs leads to extremely high-dimensional sparse features; matrix factorization expects integer indices.

  • ✓

    Sort the data by timestamp and use a time-based split for training and validation

    Why this is correct

    Time-based split ensures that the model does not use future interactions to predict past ones, which is critical for time-series recommendation.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.