Question 693 of 1,672
MLS-C01 Modeling Practice Question
A data scientist is using Amazon SageMaker to deploy a custom model container. The model is a large transformer that requires 16 GB of memory. The scientist wants to minimize inference latency. Which SageMaker hosting option should they choose?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a real-time endpoint with an instance that has sufficient memory.
For a large model requiring 16GB memory and minimal inference latency, a real-time endpoint with a suitably sized instance (e.g., ml.p3.2xlarge or ml.g4dn.xlarge) provides dedicated resources and low latency. Option B (asynchronous inference) adds queuing latency and is for non-real-time. Option C (Serverless Inference) has memory limits (up to 6 GB) and may have cold starts, not suitable for a 16GB model. Option D (batch transform) is for offline inference on batches, not real-time.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a real-time endpoint with an instance that has sufficient memory.
Why this is correct
Real-time endpoints provide low latency and can accommodate large models.
- ✗
Use an asynchronous inference endpoint.
Why it's wrong here
Asynchronous adds latency for queuing and processing.
- ✗
Use SageMaker Serverless Inference.
Why it's wrong here
Serverless has memory limits up to 6 GB, insufficient for 16 GB model.
- ✗
Use a batch transform job.
Why it's wrong here
Batch transform is for offline inference, not real-time.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
About these practice questions
Courseiva creates original exam-style practice questions with explanations and wrong-answer analysis. It does not publish real exam questions, exam dumps, or protected exam content. Learn why practice questions differ from exam dumps →
Last reviewed: Jun 20, 2026
This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.
Question Discussion
Share a tip, memory trick, or ask about the reasoning behind this question. Do not post real exam questions, leaked content, braindumps, or copyrighted exam material. Comments are moderated and may be removed without notice.
Sign in to join the discussion.