MLA-C01 Deployment and Orchestration of ML Workflows Practice Question
A company wants to serve 200 different PyTorch models. Each model is small (under 1 GB) and only a fraction are used at any time. To minimize cost and management overhead, which SageMaker inference option should be used?
⚠ Common exam trap
In AWS SageMaker, the common trap is confusing multi-container endpoints (used for multi-step inference pipelines with different containers) with multi-model endpoints (used for hosting multiple independent models). Candidates often select multi-container endpoints thinking they can serve multiple models, but each container still hosts one model; multi-model endpoints are designed to host many models on a single instance, dynamically loading them as needed.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a multi-model endpoint
A multi-model endpoint (MME) is the correct choice because it allows you to host multiple small PyTorch models (under 1 GB each) on a single endpoint, sharing the underlying compute instance. This minimizes cost by only paying for the active instances, and reduces management overhead since you don't need to create or manage separate endpoints for each model. SageMaker dynamically loads and unloads models from the container's memory based on invocation patterns, which is ideal for a scenario where only a fraction of the 200 models are used at any time.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use batch transform for all models
Why it's wrong here
Batch transform runs offline, job-based inference against stored data, so it cannot serve interactive requests from 200 models. It is tempting because batch transform avoids persistent endpoint costs, but the stem requires on-demand serving with minimal overhead, which multi-model endpoints provide by loading models on invocation.
- ✗
Create a separate real-time endpoint for each model
Why it's wrong here
Separate real-time endpoints per model leave 200 always-on instances billing continuously, even though only a fraction receive traffic. It is tempting because real-time endpoints give low-latency interactive inference, but the stem's cost and management constraints favour a multi-model endpoint hosting many models behind one endpoint.
- ✗
Use a multi-container endpoint
Why it's wrong here
Multi-container endpoints run several containers concurrently on one instance for pipelines or ensembles, not hundreds of selectable models. It is tempting because it consolidates workloads onto shared infrastructure, but the stem needs models loaded on demand; multi-model endpoints store many models in Amazon S3 and load them dynamically.
- ✓
Use a multi-model endpoint
Why this is correct
Multi-model endpoints load models on demand from S3 into shared container memory and unload idle ones, so hosting 200 small PyTorch models costs far less than 200 endpoints. This satisfies the minimise-cost and management-overhead constraint given only a fraction are used concurrently.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on MLA-C01
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company wants to deploy 50 small models (each ~100 MB) for real-time inference. They need to minimize hosting costs while maintaining low latency. Which SageMaker hosting option is most cost-effective?
medium- A.SageMaker Serverless Inference
- B.SageMaker Asynchronous Inference
- ✓ C.SageMaker Multi-Model Endpoint (MME)
- D.SageMaker real-time endpoint with one instance per model
Why C: SageMaker Multi-Model Endpoint (MME) hosts many models on a single endpoint and dynamically loads them into memory/disk on invocation, sharing the underlying instance. For 50 small models (~100 MB each), this dramatically reduces hosting cost versus one endpoint per model while still providing real-time, low-latency inference. Serverless Inference is pay-per-invoke but has cold starts and memory limits that can hurt latency for frequent calls.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.