Courseiva
mediumMultiple Choice

Generative AI Leader Practice Question: Piloting a GenAI feature that summarizes customer…

A company is piloting a GenAI feature that summarizes customer support tickets. They want to measure the impact on agent productivity before rolling out to all teams. Which approach BEST evaluates the pilot?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run an A/B test where half the agents use the tool and half do not, then compare average handling time

A/B testing with a control group provides a rigorous comparison of productivity metrics. The other options lack a baseline for comparison.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Survey agents on their perception of productivity after using the tool

    Why it's wrong here

    Perception surveys capture subjective impressions, not the productivity change the pilot must quantify; agents may feel faster without any measurable gain. Tempting because surveys are cheap and reveal usability sentiment, making them suitable for satisfaction or acceptance research rather than impact evaluation.

  • ✓

    Run an A/B test where half the agents use the tool and half do not, then compare average handling time

    Why this is correct

    An A/B test with a control group isolates the tool's effect on average handling time, giving a measurable comparison against agents without it. This satisfies the stem's requirement to evaluate productivity impact before wider rollout, unlike anecdotal or purely qualitative feedback.

  • ✗

    Compare the cost of the API before and after deployment

    Why it's wrong here

    API cost tracks spend, not agent productivity; summarisation costs could fall while agents handle identical ticket volumes. It is tempting because cost is an objective, readily available metric, and this would be the right measure when the goal is budget or return-on-investment tracking.

  • ✗

    Measure the number of summaries generated per day

    Why it's wrong here

    Counting summaries generated per day measures tool usage volume, not productivity impact; agents could generate many summaries without handling more tickets or faster. It is tempting because throughput metrics are simple to collect, and this would suit capacity or adoption tracking rather than an effectiveness pilot.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.