AIF-C01 Applications of Foundation Models Practice Question
A company runs a chatbot using a large language model on Amazon Bedrock. They notice high latency during peak hours. Which action would be MOST effective to reduce latency without degrading response quality?
⚠ Common exam trap
AWS often tests the misconception that reducing model size or output length is the primary way to reduce latency, but the real bottleneck in peak-hour scenarios is often infrastructure contention, which Provisioned Throughput resolves without sacrificing quality.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Provisioned Throughput for model inference
Provisioned Throughput on Amazon Bedrock reserves dedicated capacity for model inference, ensuring consistent low latency even during peak hours. This eliminates the variability caused by resource contention in the on-demand tier, directly addressing high latency without altering model size or output quality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of concurrent invocations
Why it's wrong here
Raising concurrent invocations adds parallel capacity but does not reduce the per-request inference time causing peak-hour latency. It is tempting because scaling handles load, and it would be correct for throughput or quota limits rather than single-response speed.
- ✗
Switch to a smaller model
Why it's wrong here
A smaller model reduces compute per token but typically lowers response quality, which the stem forbids. It is tempting because smaller models genuinely cut latency, and it would be correct where cost or speed matters more than answer fidelity.
- ✗
Decrease the maxTokens parameter
Why it's wrong here
Lowering maxTokens truncates generated output, so it cannot cut time-to-first-token, which dominates perceived latency during peak load. It is tempting because capping output length genuinely reduces total generation time for long completions, making it a valid cost or throughput control — but it degrades response quality here, which the scenario explicitly forbids.
- ✓
Use Provisioned Throughput for model inference
Why this is correct
Provisioned Throughput reserves dedicated model units for your Amazon Bedrock model, eliminating the queueing that causes peak-hour latency. Unlike on-demand inference, which shares capacity, it guarantees consistent throughput without altering the model, prompt, or parameters — so response quality stays identical while latency drops.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.