easyMultiple Choice
AIF-C01 Practice Question: Building a generative AI application for code…
A company is building a generative AI application for code generation. They want to minimize costs while maintaining acceptable performance for their workload, which has periodic spikes in demand. Which approach would be MOST cost-effective?
⚠ Common exam trap
AIF-C01 often tests cost optimization strategies; candidates may focus on caching or batch but overlook the fundamental principle of matching model size to task complexity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Right-size model selection: use a smaller model for simple tasks and a larger model only when needed
Right-sizing model selection involves using smaller, cheaper models for simple tasks and reserving larger, more expensive models for complex tasks. This optimizes cost while maintaining acceptable performance, especially for workloads with periodic spikes. It avoids over-provisioning and leverages the most cost-effective model for each request.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Right-size model selection: use a smaller model for simple tasks and a larger model only when needed
Why this is correct
Routing simple requests to a smaller, cheaper model and reserving the larger model for complex tasks cuts inference cost substantially, while elastic scaling absorbs periodic demand spikes. This satisfies the stem's dual constraint of minimising cost without dropping below acceptable performance.
- ✗
Always use the largest available foundation model for all requests
Why it's wrong here
The largest model carries the highest per-token price for every request, so cost rises regardless of demand pattern. It is tempting because large models maximise quality, and this would be the right choice where accuracy on complex reasoning matters far more than spend.
- ✗
Use model caching to store responses for repeated prompts
Why it's wrong here
Caching only helps when prompts repeat; code-generation prompts are largely unique, so hit rates stay low and spikes still need full compute. It is tempting because caching genuinely cuts cost for high-repetition workloads, such as FAQ chatbots or fixed system-prompt calls, where identical requests recur frequently.
- ✗
Use batch inference for all requests
Why it's wrong here
Batch inference trades latency for discount, but code generation needs interactive responses, and periodic spikes still require provisioned capacity. It is tempting because batching genuinely reduces cost for asynchronous, delay-tolerant jobs such as overnight document classification, where results are not needed immediately.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.