AIF-C01 Fundamentals of Generative AI Practice Question
A team is developing a real-time code completion feature using an LLM deployed on Amazon SageMaker. They observe high latency under load. Which optimization technique should they prioritize?
⚠ Common exam trap
The AWS exam often tests the distinction between throughput optimization (batch size, scaling) and latency optimization (quantization, pruning), and the trap here is assuming that scaling out or increasing resources always solves latency issues, when in fact the bottleneck is per-inference computation time.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use model quantization
Model quantization reduces the precision of the model's weights (e.g., from FP32 to INT8), which decreases memory footprint and computational requirements, leading to lower latency per inference. This is the most effective optimization for real-time code completion because it directly reduces the time to generate each token without requiring additional infrastructure changes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase batch size
Why it's wrong here
Increasing batch size groups more requests per forward pass, which raises throughput but adds queuing delay, worsening the per-token latency that real-time completion depends on. It is tempting because batching is a standard GPU-efficiency technique. Latency-sensitive generation instead requires optimisations that shorten each inference, such as quantisation or speculative decoding.
- ✗
Switch to a larger instance type
Why it's wrong here
A larger instance gives more memory and compute but does not change the model's sequential decoding path, so per-token latency under load stays broadly similar. It is tempting because vertical scaling is a familiar fix for resource saturation. Real-time completion instead needs inference-level optimisation, such as quantisation or a distilled model, to cut each request's latency.
- ✗
Increase instance count with Auto Scaling
Why it's wrong here
Auto Scaling adds replicas, which raises aggregate throughput but leaves per-request inference latency unchanged, since each request still traverses the same model. It is tempting because horizontal scaling is the standard remedy for load-related slowdowns in stateless services. Real-time completion needs per-token latency reduced, which model-serving optimisation addresses instead.
- ✓
Use model quantization
Why this is correct
Quantization reduces weight precision, shrinking model size and memory bandwidth demand, which lowers per-token inference latency and increases throughput under concurrent load. This directly addresses the high-latency constraint for real-time code completion without retraining or architectural changes.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.