MLA-C01 Deployment and Orchestration of ML Workflows Practice Question
A company wants to serve a large ensemble of models using NVIDIA Triton Inference Server on SageMaker for high throughput GPU inference. Which SageMaker inference option supports this?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Real-time endpoint with a custom container running Triton
SageMaker supports Triton Inference Server through a custom real-time endpoint container, as Triton is optimized for GPU serving on NVIDIA hardware.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Asynchronous Inference
Why it's wrong here
Asynchronous is for large payloads, not optimized for Triton ensemble serving.
- ✗
Multi-model endpoint
Why it's wrong here
MME does not natively support Triton; it uses SageMaker inference toolkit.
- ✗
Serverless Inference
Why it's wrong here
Serverless does not support custom containers or Triton.
- ✓
Real-time endpoint with a custom container running Triton
Why this is correct
Customers can bring their own Triton container to SageMaker real-time endpoints for optimal GPU inference.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 835 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.