mediumMultiple Select
AIF-C01 Model Variant Practice Question
A company uses Amazon Bedrock with Anthropic Claude for a question-answering system. They want to reduce costs while maintaining acceptable latency. Which TWO actions would help achieve this? (Choose two.)
⚠ Common exam trap
The trap is that candidates may think increasing context window (D) or using a vector database (E) are cost-effective, but both can increase latency. Vector database retrieval adds overhead, and prompt caching is often overlooked as a cost-saving technique.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a smaller model variant (e.g., Claude Instant instead of Claude)
Option A is correct because choosing a smaller, cheaper model variant such as Claude Instant instead of a larger Claude model directly lowers the per-token inference cost while still providing acceptable latency for question-answering workloads. Option C is correct because enabling prompt caching for frequently reused system prompts avoids re-processing the same input tokens on every request, reducing both token costs and latency for repeated prompt prefixes. Option B is wrong because switching to image generation changes the workload entirely and does not reduce the cost of a text-based question-answering system. Option D is wrong because increasing the context window adds more input tokens, which raises cost and can increase latency rather than reduce them. Option E is wrong because using a vector database to pre-filter documents is a retrieval optimization that can improve relevance, but it does not by itself reduce Bedrock model inference costs or guarantee lower latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a smaller model variant (e.g., Claude Instant instead of Claude)
Why this is correct
Claude Instant processes tokens faster and costs less per token than larger Claude variants, directly lowering both inference spend and response latency. It satisfies the cost-reduction goal provided the smaller model still meets the question-answering quality the system requires.
- ✗
Switch from text generation to image generation
Why it's wrong here
Image generation does not answer text questions and typically costs more per invocation, so it cannot reduce spend while preserving latency. It is tempting because switching modalities sounds like a workload change, but Bedrock pricing differs by model and token, not by task type.
- ✓
Enable prompt caching for frequently used system prompts
Why this is correct
Prompt caching stores the processed system prompt so repeated requests reuse cached context instead of reprocessing those input tokens. Anthropic Claude on Amazon Bedrock bills cached tokens at a reduced rate, cutting cost for frequently reused prompts while latency stays acceptable.
- ✗
Increase the context window to include more documents
Why it's wrong here
Enlarging the context window increases the number of input tokens processed per request, raising cost and latency rather than lowering them. It is tempting because more context can improve answer quality, which is a correctness goal, not the cost-and-latency goal stated.
- ✗
Use a vector database to pre-filter documents before sending to the model
Why it's wrong here
A vector database adds retrieval latency, which may negatively impact latency, so it is not recommended for maintaining acceptable latency.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.