Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A company runs an online retail business and wants to build a product recommendation system. They have a dataset of customer purchases stored in Amazon S3 as CSV files. The dataset includes columns: 'customer_id', 'product_id', 'purchase_date', 'quantity', 'price', and 'category'. The data science team plans to use Amazon SageMaker to train a factorization machines model. During data exploration, they discover that the 'category' column has 1,200 unique values, and many categories appear only a few times. The 'product_id' column has 50,000 unique values. They want to include both features in the model. The team is concerned about the high cardinality of these features. Which approach should they take to prepare these features for the factorization machines model?

⚠ Common exam trap

It's easy for candidates to default to one-hot encoding (Option A) as the standard categorical encoding technique, not realizing that factorization machines are specifically designed to avoid that explosion by accepting raw integer indices as categorical features.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Encode both columns as integer indices and feed them directly to the factorization machines algorithm as categorical features.

Amazon SageMaker's factorization machines algorithm natively supports categorical features encoded as integer indices (0-based). This avoids the explosion of features from one-hot encoding (which would create 51,200 columns) and leverages the algorithm's ability to learn interactions between high-cardinality features via factorized parameters, making it both memory-efficient and effective for sparse data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apply one-hot encoding to both 'product_id' and 'category' columns.

    Why it's wrong here

    One-hot encoding expands 50,000 products and 1,200 categories into over 51,000 sparse binary columns, producing a matrix too wide for factorization machines to train efficiently. It is tempting because one-hot encoding is standard for low-cardinality categoricals, and would suit a feature with only a handful of distinct values.

  • ✗

    Drop the 'category' column and only use 'product_id' since it has more granularity.

    Why it's wrong here

    Dropping category discards a feature the stem requires in the model, and product_id's 50,000 values still leave high cardinality unresolved. It is tempting because removing sparse columns reduces noise, and it would be correct if category were genuinely irrelevant to recommendations rather than merely high-cardinality.

  • ✓

    Encode both columns as integer indices and feed them directly to the factorization machines algorithm as categorical features.

    Why this is correct

    Factorization machines handle high-cardinality categoricals by learning latent factor vectors per index, so integer-encoding category and product_id directly satisfies the stem's concern. One-hot encoding would explode dimensionality; index encoding lets the algorithm capture interactions between sparse features efficiently.

  • ✗

    Apply principal component analysis (PCA) to reduce the dimensionality of the categorical features.

    Why it's wrong here

    PCA operates on continuous numeric variance and assumes linear relationships, so applying it to categorical identifiers destroys the per-category structure factorization machines need to learn latent interactions. It is tempting because PCA is a standard dimensionality-reduction tool, and would be correct for compressing correlated numeric features such as price and quantity.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.