Courseiva
mediumMultiple Choice

AIF-C01 Practice Question: Using Amazon Bedrock to build an application that…

A company is using Amazon Bedrock to build an application that requires very low latency responses (under 100ms). They are currently using a large model but need faster inference. Which model selection strategy is MOST appropriate?

⚠ Common exam trap

A common misconception in AWS exams is that adjusting hyperparameters like temperature can affect inference speed, but temperature only influences output diversity and randomness.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Switch to a smaller model that can meet the latency requirement while still providing acceptable quality

Smaller models have fewer parameters, which reduces the computational cost per inference, directly lowering latency. For latency-sensitive applications requiring under 100ms responses, a smaller model can often provide acceptable quality while meeting the strict timing requirement, whereas a larger model would introduce higher inference latency due to increased matrix operations and memory bandwidth demands.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the temperature parameter to speed up generation

    Why it's wrong here

    Temperature governs sampling randomness, not inference speed; it cannot reduce the time a forward pass takes, so the 100ms target stays unmet. Temperature tuning belongs to creativity control, such as diversifying marketing copy, not to latency optimisation.

  • ✗

    Use a larger model with more parameters for better accuracy

    Why it's wrong here

    Adding parameters increases compute per token, raising latency further and moving away from the sub-100ms requirement. Larger models are chosen when accuracy on complex reasoning matters more than response time, not when inference speed is the binding constraint.

  • ✓

    Switch to a smaller model that can meet the latency requirement while still providing acceptable quality

    Why this is correct

    Smaller models have fewer parameters, so each forward pass requires less computation and memory bandwidth, cutting inference latency. This satisfies the sub-100ms constraint while retaining acceptable output quality, unlike prompt tuning or provisioned throughput, which do not reduce per-token compute.

  • ✗

    Use batch processing instead of real-time streaming

    Why it's wrong here

    Batch processing defers inference to asynchronous jobs, so responses cannot arrive within 100ms; it targets throughput over latency. It is tempting because Amazon Bedrock batch inference suits large offline workloads, such as overnight document classification, where turnaround time is measured in hours and per-token cost matters more than immediate replies.

Visual reference

R1 R2 R3 R4 10 100 10 100 OSPF picks R1→R2→R4 (cost 20) over R1→R3→R4 (cost 200)

About these practice questions

One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.