Courseiva
mediumMultiple Choice

Generative AI Leader Practice Question: During a proof-of-concept for a GenAI document…

During a proof-of-concept for a GenAI document summarization tool, the team wants to evaluate whether the summaries are accurate and retain key information before scaling. Which evaluation approach is most appropriate for this stage?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run an A/B test with a small user group and have domain experts manually review a sample of summaries for accuracy

A/B testing with manual review by domain experts provides qualitative and quantitative feedback on accuracy and completeness, which is crucial for a pilot. Automated metrics alone may not capture business relevance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Measure latency and cost as the primary evaluation metrics

    Why it's wrong here

    Latency and cost measure operational efficiency, not whether summaries preserve key facts or avoid fabrication — the accuracy goal of this proof-of-concept. They become the right focus later, when tuning throughput and budget for a deployed pipeline, but here they cannot detect hallucinated or missing content.

  • ✗

    Deploy to all users and collect feedback via a survey

    Why it's wrong here

    A full rollout exposes every user to unvalidated output before accuracy is established, and survey feedback measures satisfaction rather than factual retention. This suits post-deployment monitoring at scale, not a proof-of-concept needing controlled, reference-based accuracy measurement.

  • ✓

    Run an A/B test with a small user group and have domain experts manually review a sample of summaries for accuracy

    Why this is correct

    At proof-of-concept stage, accuracy and key-information retention matter more than scale metrics. Expert manual review of a sampled subset directly measures those qualities, whereas a small A/B test captures user preference, not factual fidelity, and is premature before quality is validated.

  • ✗

    Use ROUGE scores exclusively to compare summaries against human-written ones

    Why it's wrong here

    ROUGE measures n-gram overlap with reference summaries, so it cannot detect fabricated content or omitted key facts, which is precisely what the proof-of-concept must verify. It is tempting because ROUGE is cheap, reproducible and standard for summarisation benchmarking, and would suit large-scale regression testing once accuracy has already been established.

About these practice questions

This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.