Courseiva
Question 693 of 1,672
ModelinghardMultiple ChoiceObjective-mapped

MLS-C01 Modeling Practice Question

A data scientist is using Amazon SageMaker to deploy a custom model container. The model is a large transformer that requires 16 GB of memory. The scientist wants to minimize inference latency. Which SageMaker hosting option should they choose?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Use a real-time endpoint with an instance that has sufficient memory.

For a large model requiring 16GB memory and minimal inference latency, a real-time endpoint with a suitably sized instance (e.g., ml.p3.2xlarge or ml.g4dn.xlarge) provides dedicated resources and low latency. Option B (asynchronous inference) adds queuing latency and is for non-real-time. Option C (Serverless Inference) has memory limits (up to 6 GB) and may have cold starts, not suitable for a 16GB model. Option D (batch transform) is for offline inference on batches, not real-time.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Use a real-time endpoint with an instance that has sufficient memory.

    Why this is correct

    Real-time endpoints provide low latency and can accommodate large models.

  • Use an asynchronous inference endpoint.

    Why it's wrong here

    Asynchronous adds latency for queuing and processing.

  • Use SageMaker Serverless Inference.

    Why it's wrong here

    Serverless has memory limits up to 6 GB, insufficient for 16 GB model.

  • Use a batch transform job.

    Why it's wrong here

    Batch transform is for offline inference, not real-time.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

Courseiva creates original exam-style practice questions with explanations and wrong-answer analysis. It does not publish real exam questions, exam dumps, or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Last reviewed: Jun 20, 2026

Question Discussion

Share a tip, memory trick, or ask about the reasoning behind this question. Do not post real exam questions, leaked content, braindumps, or copyrighted exam material. Comments are moderated and may be removed without notice.

Loading comments…

Sign in to join the discussion.

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.