Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data science team deploys a PyTorch model on…

A data science team deploys a PyTorch model on Amazon SageMaker for real-time inference. The model requires GPU for low latency. Which instance type is MOST cost-effective while meeting the GPU requirement?

⚠ Common exam trap

Many exam-takers assume any GPU instance is equally cost-effective, overlooking that ml.p4d.24xlarge is overprovisioned for typical inference, while CPU-only instances like ml.m5 and ml.c5 are tempting but fail the explicit GPU requirement.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ml.p3.2xlarge

(ml.p3.2xlarge) is correct because it provides a GPU (NVIDIA V100) necessary for low-latency PyTorch inference on SageMaker, while being the most cost-effective among GPU options. The ml.p3.2xlarge offers a single GPU with sufficient compute for many real-time inference workloads, avoiding the higher cost of larger instances like ml.p4d.24xlarge.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    ml.m5.2xlarge

    Why it's wrong here

    ml.m5.2xlarge is a general-purpose instance without GPU hardware, so inference would fall back to CPU and miss the latency target. It is tempting because it is inexpensive and suits CPU-only workloads, and it would be correct for lightweight models that need no GPU acceleration.

  • ✗

    ml.p4d.24xlarge

    Why it's wrong here

    ml.p4d.24xlarge provides eight high-end GPUs, far exceeding what a single low-latency PyTorch endpoint needs, so its cost is unjustified. It is tempting because it is GPU-backed, and it would be correct for large-scale distributed training rather than cost-effective real-time inference.

  • ✓

    ml.p3.2xlarge

    Why this is correct

    ml.p3.2xlarge pairs a single NVIDIA V100 GPU with the lowest cost among GPU-backed instances, satisfying the GPU constraint while avoiding the expense of multi-GPU types such as ml.p3.8xlarge. It delivers the required low-latency inference economically.

  • ✗

    ml.c5.2xlarge

    Why it's wrong here

    ml.c5.2xlarge is a compute-optimised instance with no GPU, so it cannot satisfy the GPU requirement at all. It is tempting because it is cheap and would be the cost-effective choice for CPU-bound training or inference workloads that need no acceleration.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.