AIF-C01 Fundamentals of Generative AI Practice Question
Which TWO strategies can help reduce inference costs when using Amazon Bedrock? (Select TWO.)
⚠ Common exam trap
Candidates may mistakenly believe that caching responses in Amazon ElastiCache reduces inference costs, but this is not a built-in feature of Bedrock and would require custom implementation, still incurring costs for cache misses. Similarly, adjusting temperature or max tokens does not directly reduce per-token costs. The correct strategies are using provisioned throughput for high-volume predictable workloads and selecting a smaller foundation model variant.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use provisioned throughput for high-volume, predictable workloads
Option C is correct because Provisioned Throughput on Amazon Bedrock lets you purchase dedicated model units at a fixed hourly rate for steady, high-volume, predictable inference traffic, which is typically cheaper than paying on-demand per-token prices once utilization is high. Option E is correct because choosing a smaller foundation model variant (for example, a mini or lite version of a model family) reduces the number of parameters and compute required per inference, directly lowering the per-token or per-request cost while often remaining sufficient for the task. Option A is not correct because temperature controls randomness of sampling, not the number of tokens generated, so raising it does not reduce cost and may even increase output variability. Option B is not correct because increasing max tokens allows longer responses, which increases the number of output tokens billed and therefore raises cost. Option D is not correct because Amazon ElastiCache is a caching service, not a native Bedrock cost-reduction strategy, and caching model responses is not one of the two intended Bedrock cost-optimization approaches in this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a higher temperature setting to generate fewer tokens
Why it's wrong here
Temperature does not affect token count; it affects randomness.
- ✗
Increase the max tokens to allow longer responses
Why it's wrong here
Longer responses cost more because Bedrock charges per token processed and generated.
- ✓
Use provisioned throughput for high-volume, predictable workloads
Why this is correct
Provisioned throughput reserves dedicated model capacity for a committed term, replacing per-token on-demand pricing. For steady, high-volume workloads this committed rate is cheaper than paying on-demand for every request, directly lowering inference cost at predictable scale.
- ✗
Cache frequently used responses in Amazon ElastiCache
Why it's wrong here
Bedrock does not provide built-in caching, and caching is not a direct cost reduction strategy for the Bedrock service itself.
- ✓
Select a smaller foundation model variant
Why this is correct
Smaller foundation model variants use fewer parameters, so each inference consumes less compute and memory, translating into lower per-token charges on Amazon Bedrock. Where the task tolerates reduced capability, this directly reduces inference cost without changing architecture.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.