mediumMultiple Choice
AIF-C01 Practice Question: A company uses Amazon Titan Text Express for a…
A company uses Amazon Titan Text Express for a real-time chat application. Users report that responses are too slow. The application uses the InvokeModel API with default settings. Which change is MOST likely to reduce latency?
⚠ Common exam trap
The AWS exam often tests the misconception that all models in a family have identical performance characteristics, but Titan Text Lite and Express are deliberately differentiated by speed versus quality, and candidates may overlook that model selection is the primary latency lever.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Switch from Titan Text Express to Titan Text Lite
Titan Text Lite is a smaller, faster model optimized for low-latency use cases like real-time chat, whereas Titan Text Express prioritizes higher quality and throughput at the cost of speed. Switching to Titan Text Lite directly reduces inference time because it has fewer parameters and lower computational overhead per request, making it the most effective change to reduce latency with default InvokeModel API settings.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use the Converse API instead of InvokeModel
Why it's wrong here
The Converse API unifies request formatting across models but does not change token generation speed, so latency stays the same. It is tempting because it simplifies switching between models, and would be correct when the goal is portable code rather than faster responses.
- ✗
Reduce the temperature parameter to 0
Why it's wrong here
Temperature controls sampling randomness, not token generation count; setting it to 0 makes output deterministic but does not shorten inference time. It is tempting because lower temperature is often assumed to speed generation, and it would be the right choice when consistent, repeatable responses matter more than latency.
- ✓
Switch from Titan Text Express to Titan Text Lite
Why this is correct
Titan Text Lite is the smaller, faster variant in the Titan Text family, trading some capability for lower per-token latency. For a real-time chat workload where speed is the reported problem, this substitution directly reduces response time.
- ✗
Increase the maxTokens parameter to allow longer responses
Why it's wrong here
Raising maxTokens permits longer outputs, so the model may generate more tokens and latency increases rather than falls. It is tempting because token limits sound like a throughput control, and it would be correct when responses are being truncated before completion.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.