Courseiva
mediumMultiple Choice

AIF-C01 Practice Question: A company uses Amazon Titan Text Express for a…

A company uses Amazon Titan Text Express for a real-time chat application. Users report that responses are too slow. The application uses the InvokeModel API with default settings. Which change is MOST likely to reduce latency?

⚠ Common exam trap

The AWS exam often tests the misconception that all models in a family have identical performance characteristics, but Titan Text Lite and Express are deliberately differentiated by speed versus quality, and candidates may overlook that model selection is the primary latency lever.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Switch from Titan Text Express to Titan Text Lite

Titan Text Lite is a smaller, faster model optimized for low-latency use cases like real-time chat, whereas Titan Text Express prioritizes higher quality and throughput at the cost of speed. Switching to Titan Text Lite directly reduces inference time because it has fewer parameters and lower computational overhead per request, making it the most effective change to reduce latency with default InvokeModel API settings.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the Converse API instead of InvokeModel

    Why it's wrong here

    The Converse API unifies request formatting across models but does not change token generation speed, so latency stays the same. It is tempting because it simplifies switching between models, and would be correct when the goal is portable code rather than faster responses.

  • ✗

    Reduce the temperature parameter to 0

    Why it's wrong here

    Temperature controls sampling randomness, not token generation count; setting it to 0 makes output deterministic but does not shorten inference time. It is tempting because lower temperature is often assumed to speed generation, and it would be the right choice when consistent, repeatable responses matter more than latency.

  • ✓

    Switch from Titan Text Express to Titan Text Lite

    Why this is correct

    Titan Text Lite is the smaller, faster variant in the Titan Text family, trading some capability for lower per-token latency. For a real-time chat workload where speed is the reported problem, this substitution directly reduces response time.

  • ✗

    Increase the maxTokens parameter to allow longer responses

    Why it's wrong here

    Raising maxTokens permits longer outputs, so the model may generate more tokens and latency increases rather than falls. It is tempting because token limits sound like a throughput control, and it would be correct when responses are being truncated before completion.

About these practice questions

One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.