mediumMultiple Choice
Generative AI Leader Practice Question: During a proof-of-concept for a GenAI document…
During a proof-of-concept for a GenAI document summarization tool, the team wants to evaluate whether the summaries are accurate and retain key information before scaling. Which evaluation approach is most appropriate for this stage?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run an A/B test with a small user group and have domain experts manually review a sample of summaries for accuracy
A/B testing with manual review by domain experts provides qualitative and quantitative feedback on accuracy and completeness, which is crucial for a pilot. Automated metrics alone may not capture business relevance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Measure latency and cost as the primary evaluation metrics
Why it's wrong here
Latency and cost measure operational efficiency, not whether summaries preserve key facts or avoid fabrication — the accuracy goal of this proof-of-concept. They become the right focus later, when tuning throughput and budget for a deployed pipeline, but here they cannot detect hallucinated or missing content.
- ✗
Deploy to all users and collect feedback via a survey
Why it's wrong here
A full rollout exposes every user to unvalidated output before accuracy is established, and survey feedback measures satisfaction rather than factual retention. This suits post-deployment monitoring at scale, not a proof-of-concept needing controlled, reference-based accuracy measurement.
- ✓
Run an A/B test with a small user group and have domain experts manually review a sample of summaries for accuracy
Why this is correct
At proof-of-concept stage, accuracy and key-information retention matter more than scale metrics. Expert manual review of a sampled subset directly measures those qualities, whereas a small A/B test captures user preference, not factual fidelity, and is premature before quality is validated.
- ✗
Use ROUGE scores exclusively to compare summaries against human-written ones
Why it's wrong here
ROUGE measures n-gram overlap with reference summaries, so it cannot detect fabricated content or omitted key facts, which is precisely what the proof-of-concept must verify. It is tempting because ROUGE is cheap, reproducible and standard for summarisation benchmarking, and would suit large-scale regression testing once accuracy has already been established.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.