mediumMultiple Choice
Generative AI Leader Practice Question: A startup is developing an AI tool that generates…
A startup is developing an AI tool that generates legal contract summaries. Lawyers will use these summaries for advice. What human oversight mechanism is MOST critical for responsible deployment?
⚠ Common exam trap
This question tests the misconception that automated verification (e.g., a second LLM or confidence scoring) is sufficient for high-stakes AI outputs, when in fact only human expert review provides the necessary accountability and domain-specific validation, as emphasized in Google's responsible AI principles.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Require lawyers to review and approve each AI-generated summary before use
In high-stakes domains like legal contract summarization, the AI output must be validated by a qualified human expert before use. This ensures accountability and mitigates risks of hallucinated clauses or misinterpretations that could lead to legal liability. A confidence score or automated flagging alone is insufficient because even high-confidence outputs can contain subtle errors that only a trained lawyer can catch.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Require lawyers to review and approve each AI-generated summary before use
Why this is correct
Requiring lawyer review before use satisfies the human-in-the-loop constraint for high-stakes legal advice, where errors carry liability. Because generative models can hallucinate clauses or misstate obligations, a qualified legal professional must verify each summary's accuracy and applicability before it informs client advice, ensuring accountability remains with a licensed practitioner.
- ✗
Provide a confidence score for each summary but allow lawyers to bypass it
Why it's wrong here
A confidence score that lawyers may bypass provides no enforced human review, so unreviewed AI output could still inform legal advice. It is tempting because confidence scoring supports transparency, and it would be appropriate where users are trusted to act on warnings voluntarily rather than requiring mandatory oversight.
- ✗
Use a second LLM to verify the first LLM's summaries automatically
Why it's wrong here
A second LLM shares the first model's failure modes and cannot supply the accountable human judgement that legal advice demands. It is tempting because automated verification scales cheaply, and it would be appropriate for internal quality checks on non-advisory content where no professional liability attaches to the output.
- ✗
Automatically flag summaries that contain low-confidence indicators for review
Why it's wrong here
Flagging only low-confidence summaries leaves high-confidence but incorrect output unchecked, and a flag alone does not compel a lawyer to review before advice is given. It is tempting because triage targets scarce reviewer attention, which suits low-stakes drafting where errors are cheap to correct.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.