AI-102 Implement generative AI solutions Practice Question
Which TWO actions should you take to reduce the cost of using Azure OpenAI for a chatbot that handles high traffic?
⚠ Common exam trap
Test-takers frequently confuse token-related parameters (max_tokens, temperature) with cost-saving mechanisms, when in reality cost reduction for high-traffic chatbots relies on architectural patterns like batching and caching, not on tweaking inference parameters.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the batch API for non-real-time requests.
Option B is correct because the Azure OpenAI Batch API processes asynchronous, non-real-time workloads at a 50% discount compared to standard global pricing, so routing any chatbot requests that don't need immediate responses through batch jobs directly lowers per-token cost. Option C is correct because caching responses to frequently asked questions (for example with Azure Cache for Redis or a semantic cache) means repeated identical or similar prompts are served without calling the model at all, eliminating token charges for that traffic and reducing overall request volume. Option A is not correct because increasing max_tokens raises the maximum output length and therefore can increase, not decrease, token consumption and cost. Option D is not correct because temperature controls randomness in sampling, not the number of tokens generated, so setting it to 0 does not reduce token usage or cost. Option E is not correct because fine-tuning does not shorten the prompt by itself and adds training and hosting costs; prompt length is reduced through techniques like prompt compression or shorter system messages, not fine-tuning alone.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase max_tokens to reduce the number of requests.
Why it's wrong here
Raising max_tokens lengthens each completion, increasing token consumption and cost per call rather than reducing request count. It is appropriate when answers are being truncated. Cost reduction instead comes from capping max_tokens, shortening prompts, and caching repeated responses.
- ✓
Use the batch API for non-real-time requests.
Why this is correct
The batch API processes asynchronous, non-real-time workloads at a 50% discount compared with standard global deployments, directly satisfying the stem's cost-reduction constraint for high-traffic chatbots. Offloading tolerant requests frees real-time capacity, lowering overall spend while preserving interactive latency for genuine user conversations.
- ✓
Implement caching for frequently asked questions.
Why this is correct
Caching stores responses to repeated questions, so identical prompts are served from the cache instead of invoking the model. This cuts token consumption and request volume, directly reducing cost for a high-traffic chatbot where many queries repeat.
- ✗
Lower the temperature to 0 to reduce token usage.
Why it's wrong here
Temperature controls sampling randomness, not token count; setting it to 0 changes output determinism only. It is tempting because lower temperature can produce shorter, more focused completions, but token usage is driven by prompt and max_tokens settings. The scenario needs actions that directly cut billed tokens.
- ✗
Fine-tune the model to reduce prompt length.
Why it's wrong here
Fine-tuning alters model weights to improve task accuracy; it does not shorten prompts or reduce per-token billing. It is tempting because fine-tuning can replace lengthy few-shot examples with a shorter prompt, but that is an indirect benefit, not the mechanism this scenario needs. The question asks for direct cost reduction via token usage.
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.