Courseiva
easyMultiple Select

Generative AI Leader Practice Question: A project manager wants to measure the success of…

A project manager wants to measure the success of a generative AI feature that summarizes meeting transcripts. Which TWO metrics are MOST appropriate for evaluating quality improvement?

⚠ Common exam trap

Generative AI Leader often tests the confusion between quality metrics and operational metrics (e.g., cost, latency, handle time), so candidates must recognize that only accuracy and user satisfaction directly measure the quality of the generated output.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Accuracy of summaries (e.g., factuality, completeness)

Option B is correct because accuracy metrics such as factuality and completeness directly measure whether the generative AI summaries faithfully and fully capture the meeting transcript content, which is the core quality dimension of a summarization feature. Option E is correct because user satisfaction ratings capture the perceived usefulness and quality of the summaries from the people consuming them, providing a human-centered quality signal that complements automated accuracy checks. Together, these two metrics evaluate both objective correctness and subjective usefulness, which are the most appropriate indicators of quality improvement for a summarization feature. Option A is not appropriate because average handle time measures operational efficiency of follow-up work, not the quality of the generated summaries. Option C is not appropriate because API response time is a latency/performance metric, not a quality metric. Option D is not appropriate because cost per summary is a financial efficiency metric rather than a measure of summary quality.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Average handle time for meeting follow-ups

    Why it's wrong here

    Average handle time for meeting follow-ups measures downstream human effort, not the quality of the generated summary itself. It is tempting because it is an outcome metric linked to productivity, and it would be correct for evaluating operational efficiency or time savings rather than summarisation accuracy, coherence or relevance.

  • ✓

    Accuracy of summaries (e.g., factuality, completeness)

    Why this is correct

    Accuracy directly measures whether summaries preserve the transcript's facts and cover its key points, satisfying the stem's quality-improvement focus. Unlike latency or cost metrics, factuality and completeness quantify output correctness, so improvements here demonstrate genuine gains in the summarisation feature's usefulness.

  • ✗

    API response time

    Why it's wrong here

    API response time measures latency of the summarisation call, not the quality of the generated summary. It is tempting because it is an easily collected operational metric, and it would be correct for evaluating performance, throughput or user experience rather than summarisation accuracy, coherence or relevance.

  • ✗

    Cost per summary generated

    Why it's wrong here

    Cost per summary generated measures the expense of producing each output, not its quality. It is tempting because it is a straightforward financial metric tied directly to the feature, and it would be correct for evaluating cost efficiency, budgeting or return on investment rather than summarisation accuracy, coherence or relevance.

  • ✓

    User satisfaction rating of the summaries

    Why this is correct

    User satisfaction ratings capture whether readers judge the summaries as genuinely better, reflecting perceived quality improvement that automated scores can miss. This satisfies the stem's requirement for a quality-improvement metric tied to the summarisation feature's real-world usefulness.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.