AIF-C01 Applications of Foundation Models Practice Question
A startup is deploying a foundation model on Amazon SageMaker for real-time inference. They notice high latency (over 2 seconds per request). Which action is most likely to reduce latency?
⚠ Common exam trap
AWS often tests the distinction between latency (time per single request) and throughput (requests per second), so candidates mistakenly choose auto-scaling or batch size increases, which improve throughput but not per-request latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Switch to a smaller, distilled version of the model.
Using a smaller, distilled version of the model directly reduces the computational complexity per inference request. Distillation compresses the model by training a smaller student network to mimic a larger teacher model, resulting in fewer parameters and faster forward passes. This is the most direct way to cut latency when the model size is the bottleneck, as it reduces the number of floating-point operations (FLOPs) required per request.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable auto-scaling on the SageMaker endpoint to handle more concurrent requests.
Why it's wrong here
Auto-scaling adds instances for concurrency; it does not shorten the latency of a single request. It is tempting because scaling helps under heavy load, and would be correct if latency rose only when many requests arrived simultaneously, indicating a throughput bottleneck.
- ✓
Switch to a smaller, distilled version of the model.
Why this is correct
A smaller, distilled model reduces the number of parameters and floating-point operations per inference, directly cutting compute time on the SageMaker endpoint. This satisfies the stem's real-time latency constraint, since inference duration scales with model size; distillation preserves much of the original accuracy while delivering sub-second responses.
- ✗
Deploy the model on a CPU-based instance instead of GPU.
Why it's wrong here
CPUs lack the parallel matrix throughput GPUs provide, so inference per request becomes slower, not faster. It is tempting because CPU instances cost less, and would be correct for tiny models or low-traffic batch workloads where GPU cost is unjustified.
- ✗
Increase the batch size parameter in the inference request.
Why it's wrong here
A larger batch size groups more requests per forward pass, increasing the wait before any single response returns. It is tempting because batching raises overall throughput, and would be correct for offline batch transform jobs where per-request latency is irrelevant.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.