easyMultiple Choice
MLA-C01 Practice Question: A data science team deploys a PyTorch model on…
A data science team deploys a PyTorch model on Amazon SageMaker for real-time inference. The model requires GPU for low latency. Which instance type is MOST cost-effective while meeting the GPU requirement?
⚠ Common exam trap
Many exam-takers assume any GPU instance is equally cost-effective, overlooking that ml.p4d.24xlarge is overprovisioned for typical inference, while CPU-only instances like ml.m5 and ml.c5 are tempting but fail the explicit GPU requirement.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ml.p3.2xlarge
(ml.p3.2xlarge) is correct because it provides a GPU (NVIDIA V100) necessary for low-latency PyTorch inference on SageMaker, while being the most cost-effective among GPU options. The ml.p3.2xlarge offers a single GPU with sufficient compute for many real-time inference workloads, avoiding the higher cost of larger instances like ml.p4d.24xlarge.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
ml.m5.2xlarge
Why it's wrong here
ml.m5.2xlarge is a general-purpose instance without GPU hardware, so inference would fall back to CPU and miss the latency target. It is tempting because it is inexpensive and suits CPU-only workloads, and it would be correct for lightweight models that need no GPU acceleration.
- ✗
ml.p4d.24xlarge
Why it's wrong here
ml.p4d.24xlarge provides eight high-end GPUs, far exceeding what a single low-latency PyTorch endpoint needs, so its cost is unjustified. It is tempting because it is GPU-backed, and it would be correct for large-scale distributed training rather than cost-effective real-time inference.
- ✓
ml.p3.2xlarge
Why this is correct
ml.p3.2xlarge pairs a single NVIDIA V100 GPU with the lowest cost among GPU-backed instances, satisfying the GPU constraint while avoiding the expense of multi-GPU types such as ml.p3.8xlarge. It delivers the required low-latency inference economically.
- ✗
ml.c5.2xlarge
Why it's wrong here
ml.c5.2xlarge is a compute-optimised instance with no GPU, so it cannot satisfy the GPU requirement at all. It is tempting because it is cheap and would be the cost-effective choice for CPU-bound training or inference workloads that need no acceleration.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.