AIF-C01 Fundamentals of Generative AI Practice Question
Which TWO factors are most important when selecting a foundation model in Amazon Bedrock for a text summarization task with strict latency requirements?
⚠ Common exam trap
A common misconception is that model size (parameters) is the primary driver of latency, but in practice, latency depends on inference optimization, model quantization, and hardware, not just parameter count.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Average response latency per request.
Option A is correct because with strict latency requirements the deciding metric is the model's average response latency per request, which directly determines whether the summarization workload can meet its service-level objective. Option D is correct because output quality and token efficiency determine whether the summary is usable and how many tokens must be generated, and fewer generated tokens translate directly into lower end-to-end latency. Option B is not decisive on its own: parameter count correlates loosely with latency but does not guarantee it, since architecture, hardware, and serving optimizations also matter. Option C, the maximum input token limit, matters for handling long documents but does not address the strict latency constraint. Option E, fine-tuning availability, supports domain adaptation and accuracy rather than latency, so it is not one of the two most important factors here.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Average response latency per request.
Why this is correct
Average response latency per request directly measures whether a candidate model can satisfy the stem's strict latency requirement, since Bedrock models vary widely in inference speed. Selecting on this metric ensures the summarisation workload meets its response-time constraint rather than optimising quality alone.
- ✗
Model size in billions of parameters.
Why it's wrong here
Parameter count correlates with capability, not response time; larger models generally add inference latency, which conflicts with the strict requirement. Size matters when maximising reasoning quality on complex tasks, but latency-sensitive summarisation favours smaller, faster models and measured throughput.
- ✗
Maximum input token limit.
Why it's wrong here
A larger input token limit lets the model ingest longer documents, but summarisation latency is driven by output tokens generated and model inference speed, not input capacity. Input limits matter when summarising very long texts; here the constraint is response time.
- ✓
Output quality and token efficiency for summarization tasks.
Why this is correct
Output quality determines summary usefulness, while token efficiency governs how many tokens are generated per summary, directly affecting inference time under the strict latency constraint. Both must be weighed together, since a high-quality but verbose model may breach latency targets.
- ✗
Availability of fine-tuning capability for domain adaptation.
Why it's wrong here
Fine-tuning adapts a model to a domain, yet it does not reduce per-request latency and can even alter output length. It is the right factor when accuracy on specialised vocabulary is the priority, but strict latency demands attention to inference speed and token throughput instead.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.