mediumMultiple Select
AIF-C01 Practice Question: Deploying a generative AI application using…
A company is deploying a generative AI application using Amazon Bedrock and needs to optimize costs for a high-volume, latency-tolerant workload. Which TWO strategies should they implement? (Select TWO.)
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Batch Inference for asynchronous processing
Option A is correct because Amazon Bedrock Batch Inference lets you submit large volumes of prompts as an asynchronous job, which is priced at a discount (typically 50%) compared to on-demand inference and is ideal for latency-tolerant workloads that don't need immediate responses. Option C is correct because selecting a smaller, more efficient foundation model reduces the number of parameters processed per token, directly lowering per-token input/output costs while still meeting the application's quality needs. Option B is not appropriate because fine-tuning a large model adds training costs and does not reduce per-inference pricing for a high-volume workload. Option D is wrong because Provisioned Throughput reserves dedicated capacity at a fixed hourly commitment, which is more expensive for spiky or latency-tolerant batch workloads than on-demand or batch pricing. Option E is incorrect because Bedrock does not offer a built-in response/model caching feature that avoids redundant inference charges, so it is not a valid cost-optimization strategy here.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use Batch Inference for asynchronous processing
Why this is correct
Batch Inference submits large volumes of asynchronous requests at a lower per-token price than on-demand invocation, and results are returned within a target completion window. This satisfies the stem's latency-tolerant constraint while reducing cost for high-volume workloads.
- ✗
Deploy a large model and fine-tune it
Why it's wrong here
Fine-tuning a large model adds training expense and keeps high per-token inference pricing, raising rather than lowering cost for a latency-tolerant workload. It is tempting because fine-tuning improves task accuracy, but cost optimisation here favours smaller models, prompt compression or batch inference instead.
- ✓
Use a smaller, more efficient foundation model
Why this is correct
Smaller foundation models consume fewer tokens and compute per request, lowering inference cost. Because the workload tolerates latency rather than demanding peak responsiveness, this trade-off satisfies the stem's cost-optimisation goal without violating its performance constraint.
- ✗
Enable Provisioned Throughput for guaranteed capacity
Why it's wrong here
Provisioned Throughput reserves fixed hourly capacity, which raises cost when demand is variable and latency tolerance allows slower batch processing. It suits steady, latency-sensitive inference needing guaranteed throughput. The stem's high-volume, latency-tolerant workload instead favours on-demand pricing combined with batch inference.
- ✗
Implement model caching to avoid redundant inferences
Why it's wrong here
Caching stores prior responses, so it only avoids cost when identical prompts recur; a high-volume workload with varied inputs gains little. It is tempting because caching genuinely cuts redundant inference charges, but it does not address per-token cost, where smaller models or provisioned throughput apply.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.