AI-102 Implement generative AI solutions Practice Question
You are deploying a generative AI application using Azure OpenAI Service. The application must generate responses in multiple languages while maintaining high accuracy. You need to minimize token usage. Which approach should you recommend?
⚠ Common exam trap
Azure AI-102 often tests the misconception that translation pipelines or fine-tuning are necessary for multilingual support, when in fact a system message on a base model achieves the goal with lower token usage and complexity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a base model with a system message to output in the desired language
Azure OpenAI Service base models (e.g., GPT-4) natively support multilingual generation via a system message that sets the desired output language. This approach avoids the overhead of translation pipelines or fine-tuning, directly minimizing token usage while maintaining high accuracy through the model's inherent multilingual capabilities.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a base model with a system message to output in the desired language
Why this is correct
A system message steers the base model to respond in the target language, avoiding separate per-language models or lengthy translation prompts. This keeps prompt tokens minimal, satisfying the requirement to minimise token usage while retaining accuracy.
- ✗
Translate all input to English and then translate output back
Why it's wrong here
Translating input to English and back adds two extra translation calls, consuming tokens rather than minimising them, and round-trip translation degrades accuracy across languages. It suits pipelines where downstream models are English-only. Azure OpenAI handles multilingual prompts natively, so direct prompting satisfies both the accuracy and token-reduction requirements.
- ✗
Fine-tune a model for each target language
Why it's wrong here
Fine-tuning per language multiplies training cost and still consumes tokens at inference, so it does not minimise token usage. It is tempting because fine-tuning genuinely improves accuracy for narrow, repetitive tasks with fixed output styles. Here, prompt engineering with concise multilingual instructions and shorter context would reduce tokens directly.
- ✗
Use a separate deployment for each language
Why it's wrong here
Separate deployments per language duplicate model hosting and still consume the same tokens per request, since tokenisation depends on the model, not the deployment. Separate deployments are correct for isolating quota, versioning or traffic between distinct applications.
Go deeper
Related to this question
About these practice questions
This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.