AI-102 Implement generative AI solutions Practice Question
Your organization is building a chatbot using Azure OpenAI Service. The chatbot must provide citations from a set of internal documents stored in Azure Blob Storage. You need to configure the solution to minimize token usage while ensuring citations are accurate. Which approach should you use?
⚠ Common exam trap
Test-takers frequently confuse fine-tuning with retrieval-augmented generation (RAG), assuming fine-tuning can store factual knowledge for citation, when in reality RAG with a search index is required for accurate, token-efficient document grounding.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Azure OpenAI on your data with Azure Cognitive Search for hybrid retrieval
Azure OpenAI on your data with Azure Cognitive Search for hybrid retrieval combines vector search and keyword search to efficiently find relevant document chunks from Azure Blob Storage, minimizing token usage by only sending the most pertinent content to the model for citation generation. This approach ensures accurate citations without embedding all documents into the prompt or relying on model memory.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Embed all document content into the system prompt
Why it's wrong here
Embedding every document in the system prompt consumes the full corpus on each request, inflating token usage and providing no retrieval mechanism to attribute citations to specific sources. It is tempting for small, fixed instruction sets, and would be correct when the prompt content is short and static.
- ✗
Fine-tune a model on the documents so it can recall them from memory
Why it's wrong here
Fine-tuning adjusts model weights rather than storing documents, so the model cannot reliably reproduce verbatim text or cite source documents, and it cannot be updated without retraining. It is tempting for teaching tone or format, and would be correct for style adaptation, not citation.
- ✗
Use a large context window model (e.g., 32K) and include all documents in the prompt
Why it's wrong here
Loading all documents into a 32K context still bills every token per request and gives no retrieval grounding, so citations cannot be traced to a specific chunk. It is tempting when the corpus fits the window, and would be correct for a small, static document set.
- ✓
Use Azure OpenAI on your data with Azure Cognitive Search for hybrid retrieval
Why this is correct
Azure OpenAI on your data with Cognitive Search hybrid retrieval combines keyword and vector search, returning only the most relevant document chunks as grounding context. That narrows the prompt payload, minimising tokens while preserving accurate citations from the Blob Storage documents.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.