easyMultiple Select
Generative AI Leader Practice Question: A project manager wants to measure the success of…
A project manager wants to measure the success of a generative AI feature that summarizes meeting transcripts. Which TWO metrics are MOST appropriate for evaluating quality improvement?
⚠ Common exam trap
Generative AI Leader often tests the confusion between quality metrics and operational metrics (e.g., cost, latency, handle time), so candidates must recognize that only accuracy and user satisfaction directly measure the quality of the generated output.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Accuracy of summaries (e.g., factuality, completeness)
Option B is correct because accuracy metrics such as factuality and completeness directly measure whether the generative AI summaries faithfully and fully capture the meeting transcript content, which is the core quality dimension of a summarization feature. Option E is correct because user satisfaction ratings capture the perceived usefulness and quality of the summaries from the people consuming them, providing a human-centered quality signal that complements automated accuracy checks. Together, these two metrics evaluate both objective correctness and subjective usefulness, which are the most appropriate indicators of quality improvement for a summarization feature. Option A is not appropriate because average handle time measures operational efficiency of follow-up work, not the quality of the generated summaries. Option C is not appropriate because API response time is a latency/performance metric, not a quality metric. Option D is not appropriate because cost per summary is a financial efficiency metric rather than a measure of summary quality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Average handle time for meeting follow-ups
Why it's wrong here
Average handle time for meeting follow-ups measures downstream human effort, not the quality of the generated summary itself. It is tempting because it is an outcome metric linked to productivity, and it would be correct for evaluating operational efficiency or time savings rather than summarisation accuracy, coherence or relevance.
- ✓
Accuracy of summaries (e.g., factuality, completeness)
Why this is correct
Accuracy directly measures whether summaries preserve the transcript's facts and cover its key points, satisfying the stem's quality-improvement focus. Unlike latency or cost metrics, factuality and completeness quantify output correctness, so improvements here demonstrate genuine gains in the summarisation feature's usefulness.
- ✗
API response time
Why it's wrong here
API response time measures latency of the summarisation call, not the quality of the generated summary. It is tempting because it is an easily collected operational metric, and it would be correct for evaluating performance, throughput or user experience rather than summarisation accuracy, coherence or relevance.
- ✗
Cost per summary generated
Why it's wrong here
Cost per summary generated measures the expense of producing each output, not its quality. It is tempting because it is a straightforward financial metric tied directly to the feature, and it would be correct for evaluating cost efficiency, budgeting or return on investment rather than summarisation accuracy, coherence or relevance.
- ✓
User satisfaction rating of the summaries
Why this is correct
User satisfaction ratings capture whether readers judge the summaries as genuinely better, reflecting perceived quality improvement that automated scores can miss. This satisfies the stem's requirement for a quality-improvement metric tied to the summarisation feature's real-world usefulness.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.