AIF-C01 Fundamentals of Generative AI Practice Question
Exhibit
A SageMaker notebook cell output: "Model size: 7B parameters\nInference time on ml.g5.2xlarge: 250ms per token\nBatch size: 1\nMemory utilization: 90%"
Refer to the exhibit. A developer is optimizing latency for a generative AI model deployed on SageMaker. Based on the exhibit, which change would most likely reduce per-token latency?
⚠ Common exam trap
Candidates often think that larger instances always reduce latency, when in fact they may increase latency due to higher memory latency and inter-chip communication, while quantization directly addresses the memory bandwidth bottleneck in autoregressive decoding.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reduce model size through quantization
Reducing model size through quantization directly decreases the computational and memory requirements per inference step, which lowers the time to generate each token. This is especially effective on GPU instances where smaller models fit better in GPU memory and reduce memory bandwidth bottlenecks, leading to lower per-token latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a CPU instance
Why it's wrong here
CPU instances lack the matrix-multiply throughput and high-bandwidth memory that transformer decoding depends on, so per-token latency rises sharply. CPU inference is the right pick for small models, sporadic traffic or cost-sensitive endpoints where GPU acceleration is unnecessary.
- ✓
Reduce model size through quantization
Why this is correct
Quantization reduces weight precision, shrinking the model so each forward pass performs fewer and cheaper memory-bound operations per token. This directly lowers per-token latency, satisfying the exhibit's constraint of reducing inference time without retraining or changing the endpoint's instance type.
- ✗
Switch to a larger instance type
Why it's wrong here
A larger instance adds vCPU and memory, which does not raise the GPU's token-generation throughput; per-token latency is bound by accelerator compute and memory bandwidth. Larger instances suit memory-bound models or bigger batch workloads, not the sequential decode path that governs inter-token latency.
- ✗
Increase batch size to 10
Why it's wrong here
Batching ten requests raises aggregate throughput but each token still waits for the whole batch, so individual per-token latency grows. Larger batches suit offline or high-volume throughput scenarios, not the interactive single-stream latency the exhibit targets.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.