Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data scientist is building a model to predict…

A data scientist is building a model to predict customer churn based on historical data. The dataset has 10 features and 100,000 records, and the target is binary. Which algorithm is most appropriate for this binary classification problem?

⚠ Common exam trap

AWS often tests the distinction between supervised and unsupervised learning, and the trap here is that candidates may confuse logistic regression with linear regression due to the similar name, or incorrectly choose PCA or K-means because they are familiar with them for feature reduction or segmentation, ignoring that the task is explicitly binary classification.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Logistic regression

Logistic regression is the most appropriate algorithm for this binary classification problem because it directly models the probability of the binary target variable using a logistic (sigmoid) function, making it a natural fit for predicting customer churn (yes/no). It is efficient with 100,000 records and 10 features, providing interpretable coefficients that indicate feature importance, which is crucial for understanding churn drivers.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Principal component analysis

    Why it's wrong here

    PCA is an unsupervised dimensionality-reduction technique; it produces principal components, not class labels, so it cannot predict the binary churn target. It is tempting because it handles 10 features across 100,000 records well, and would be correct as a preprocessing step before classification, or for exploring structure in unlabelled data.

  • ✗

    K-means clustering

    Why it's wrong here

    K-means is unsupervised clustering; it partitions records into k groups by centroid distance and outputs cluster assignments, not predictions of a known binary label. It is tempting because it scales to 100,000 records and segments customers, and would be correct for exploratory customer segmentation without a predefined churn outcome.

  • ✗

    Linear regression

    Why it's wrong here

    Linear regression predicts a continuous numeric value by fitting a straight-line relationship, so its output is unbounded and not a class probability or label. It is tempting because it is simple and handles 10 features, and would be correct for forecasting a continuous target such as predicted revenue or tenure.

  • ✓

    Logistic regression

    Why this is correct

    Logistic regression outputs a probability between 0 and 1 via the sigmoid function, making it suited to the binary churn target. With 100,000 records and only 10 features, it trains quickly, resists overfitting, and yields interpretable coefficients, satisfying the classification constraint.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.