Courseiva
ML Model Development →hardMultiple Choice

MLA-C01 ML Model Development Practice Question

A company uses SageMaker Autopilot to build a binary classification model. The generated leaderboard shows an ensemble model as the best candidate. The team needs a model that can be deployed for real-time inference with latency < 10ms. What should they do?

⚠ Common exam trap

The trap is assuming the 'best' leaderboard model (highest accuracy) is always the right deployment choice — candidates forget that ensemble models trade latency for accuracy, and strict latency SLAs often force selection of a single model.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Select the best single model from the leaderboard (non-ensemble candidate) and deploy it

Ensemble models in SageMaker Autopilot combine multiple base models (e.g., stacking or voting), which increases inference latency because every base model must run and their outputs aggregated. For a strict <10ms real-time latency requirement, the best approach is to select the best-performing single (non-ensemble) model from the leaderboard, which has lower inference overhead. Autopilot's leaderboard explicitly lists both ensemble and individual model candidates, so the team can pick a single model that meets the latency SLA.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use SageMaker Inference Recommender to profile the ensemble model and optimize it

    Why it's wrong here

    Inference Recommender profiles instance types and configurations for an existing model; it cannot restructure an ensemble's sequential scoring into sub-10ms execution. It is tempting because it is the standard SageMaker tool for right-sizing endpoints and would be correct when the model already meets latency and you need instance selection.

  • ✗

    Deploy the ensemble model as a SageMaker endpoint; ensemble models are optimized for low latency

    Why it's wrong here

    Ensemble models aggregate predictions from multiple base learners, so inference latency scales with member count and routinely exceeds 10ms; nothing about ensembles is latency-optimised. It is tempting because Autopilot ranks the ensemble highest on accuracy, and deploying the leaderboard winner directly is the default workflow when latency is unconstrained.

  • ✗

    Retrain the ensemble model with fewer base estimators using a custom container

    Why it's wrong here

    Autopilot's ensemble combines multiple base models, so retraining with fewer estimators requires custom training code outside Autopilot's managed pipeline and still leaves inference latency unguaranteed. It is tempting because reducing estimator count genuinely shrinks ensemble overhead, which would help where retraining autonomy and accuracy trade-offs matter more than a hard latency target.

  • ✓

    Select the best single model from the leaderboard (non-ensemble candidate) and deploy it

    Why this is correct

    Ensemble models combine multiple learners, and their aggregated inference overhead typically exceeds the sub-10ms latency budget. A single non-ensemble candidate from the leaderboard has lower per-request compute, so selecting it satisfies the real-time latency constraint while retaining strong accuracy.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.