AI-300 Generative AI Optimization Practice Question
You want to evaluate your prompt engineering changes quantitatively. Which method is most reliable for comparing two prompt versions?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Running the prompts against a benchmark evaluation dataset.
A/B testing with a ground-truth dataset allows for objective measurement of performance changes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Monitoring the total request count.
Why it's wrong here
Request count measures volume, not quality.
- ✗
Asking developers for subjective feedback.
Why it's wrong here
Subjective feedback is prone to bias and lacks scientific rigor.
- ✗
Checking if the model is running on GPT-4.
Why it's wrong here
Model version is irrelevant to prompt performance measurement.
- ✓
Running the prompts against a benchmark evaluation dataset.
Why this is correct
Evaluation datasets ensure consistent, objective comparison metrics.
About these practice questions
One of 204 original AI-300 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed August 2026 · checked against the official Microsoft exam blueprint
This AI-300 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-300 exam.