Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A trained model needs to be deployed for…

A trained model needs to be deployed for real-time inference with low latency. Which AWS service is best suited for this?

⚠ Common exam trap

AWS often tests the distinction between batch and real-time inference, and the trap here is that candidates confuse SageMaker Batch Transform (which processes data in bulk) with a real-time serving solution, or they overestimate Lambda's ability to handle large model payloads and sustained low-latency requests.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

SageMaker endpoints

SageMaker endpoints are designed for real-time inference by provisioning persistent, auto-scaled HTTPS endpoints that return predictions with millisecond latency. They support automatic scaling, A/B testing, and can be deployed behind a VPC for low-latency access, making them the ideal choice for serving a trained model in production.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    SageMaker Batch Transform

    Why it's wrong here

    Batch Transform processes an entire dataset as an asynchronous job, so it cannot serve individual requests in real time. It is tempting because it suits large-scale offline scoring where latency is irrelevant, but the stem demands synchronous, low-latency inference, which requires a persistent endpoint.

  • ✓

    SageMaker endpoints

    Why this is correct

    SageMaker endpoints host trained models behind a persistent HTTPS endpoint, delivering the low-latency, real-time inference the stem demands. Unlike batch transform, which processes datasets asynchronously, endpoints keep compute provisioned and respond per-request, satisfying the low-latency constraint directly.

  • ✗

    SageMaker Hyperparameter Tuning

    Why it's wrong here

    Hyperparameter Tuning searches for optimal training hyperparameters; it produces a better model, not a deployed real-time endpoint, so it addresses training rather than inference latency. It is correct when the requirement is improving model accuracy before deployment.

  • ✗

    AWS Lambda with model packaged

    Why it's wrong here

    Lambda caps execution duration and lacks GPU-backed instances, so a packaged model cannot meet sustained low-latency inference for larger models. It is tempting for lightweight, event-driven inference on small models, where cold-start latency and payload limits are acceptable.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.