Cost-Effective Real-Time Inference on GPU
A data science team needs to deploy a trained PyTorch model for real-time inference with sub-100ms latency. The model fits on a single GPU. Which SageMaker inference option is MOST cost-effective while meeting the latency requirement?
Quick Answer
Meeting a sub-100ms latency requirement for real-time inference calls for a persistent, dedicated compute resource that's always ready to respond immediately, which is what a SageMaker real-time endpoint provides, as opposed to options built for asynchronous or batch-style processing where some delay before a response is acceptable. Since the model fits on a single GPU, there's no need to reach for a larger, more expensive multi-GPU instance; the ml.g4dn.xlarge, with its single NVIDIA T4 GPU, provides enough compute to serve the model within the latency requirement while being one of the lower-cost GPU instance types available, which is exactly what being most cost-effective while meeting the latency requirement is asking for. Choosing a bigger or newer GPU instance would likely also meet the latency bar but at a higher cost with no real benefit, since the model doesn't need the extra capacity, while choosing a CPU-based instance would risk missing the sub-100ms target altogether for a GPU-sized model. The pattern worth internalizing is that when a scenario specifies a firm low-latency requirement alongside a model that fits comfortably on a single GPU, the answer is usually the smallest, cheapest GPU-backed real-time endpoint that can still meet that latency bar, not the largest available instance.
⚠ Common exam trap
It's easy for candidates to choose SageMaker Serverless Inference for its cost-saving potential, but they overlook the cold start latency and lack of GPU support, which makes it unsuitable for real-time, sub-100ms inference with PyTorch models.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
SageMaker real-time endpoint on ml.g4dn.xlarge
SageMaker real-time endpoints provide dedicated, persistent instances that can handle synchronous inference with sub-100ms latency. The ml.g4dn.xlarge instance includes a single NVIDIA T4 GPU, which is sufficient for the model size and offers the lowest cost among GPU instances that meet the latency requirement. This option balances performance and cost for real-time, low-latency inference.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
SageMaker Batch Transform
Why it's wrong here
SageMaker Batch Transform processes an entire dataset as an offline job and returns results to storage; it offers no persistent endpoint, so it cannot serve individual real-time requests at all. It is tempting because it is the cheapest inference mode, and would be correct for bulk scoring where responses are not needed immediately.
- ✓
SageMaker real-time endpoint on ml.g4dn.xlarge
Why this is correct
A SageMaker real-time endpoint on ml.g4dn.xlarge provides a persistent, GPU-backed inference host with low single-digit millisecond overhead, meeting sub-100ms latency. Since the model fits one GPU, this single-instance option is more cost-effective than multi-GPU or serverless alternatives.
- ✗
SageMaker Async Inference
Why it's wrong here
Async Inference queues requests and returns results via Amazon S3, so it cannot deliver sub-100ms synchronous responses. It suits long-running or large-payload jobs with relaxed latency. Real-time endpoints, or Serverless Inference for intermittent traffic, satisfy the latency requirement here.
- ✗
SageMaker Serverless Inference
Why it's wrong here
SageMaker Serverless Inference scales to zero between requests, so cold-start initialisation of a PyTorch model on GPU routinely exceeds the sub-100ms latency budget. It is tempting because it charges only per invocation, and would be correct for intermittent, latency-tolerant workloads rather than steady real-time traffic.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on MLA-C01
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A data scientist needs to deploy a single ML model that will serve real-time predictions with low latency (under 10 ms) for a high-traffic web application. The model fits in memory and requires GPU acceleration. Which SageMaker inference option is MOST suitable?
easy- A.Real-time endpoint on ml.m5 instances
- B.Batch Transform
- ✓ C.Real-time endpoint on ml.g4dn instances
- D.Serverless Inference
Why C: Real-time endpoints on GPU instances (ml.g4dn) provide low latency and GPU acceleration, ideal for high-traffic, latency-sensitive workloads.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.