Generative AI Leader Fundamentals of Generative AI Practice Question
A team is tuning a large language model for a question-answering task. They notice the model gives high confidence scores to answers that are factually incorrect. Which evaluation metric should they primarily use to detect this overconfidence problem?
⚠ Common exam trap
Google Cloud often tests the distinction between intrinsic evaluation metrics (like perplexity) and calibration metrics, leading candidates to mistakenly choose perplexity when the core issue is confidence miscalibration rather than general model uncertainty.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Expected Calibration Error (ECE)
Expected Calibration Error (ECE) directly measures the alignment between a model's predicted confidence and its actual accuracy. In this scenario, high confidence on incorrect answers indicates miscalibration, and ECE quantifies this mismatch by binning predictions by confidence and computing the average absolute difference between accuracy and confidence per bin.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Perplexity
Why it's wrong here
Perplexity measures how fluently a model predicts token sequences, not whether stated confidence matches factual accuracy, so it cannot expose overconfidence. It is tempting because perplexity is a standard intrinsic metric for language modelling and detects distributional drift; it would suit comparing model versions or detecting unfamiliar inputs, not calibration of answer correctness.
- ✓
Expected Calibration Error (ECE)
Why this is correct
Expected Calibration Error directly measures the gap between predicted confidence and actual accuracy across probability bins, so systematically high confidence on wrong answers produces a large ECE value. This satisfies the stem's overconfidence constraint, unlike accuracy or F1, which ignore confidence entirely and cannot expose miscalibration.
- ✗
BLEU score
Why it's wrong here
BLEU compares n-gram overlap between generated and reference text, so a confidently wrong answer sharing wording with a reference scores well and the miscalibration stays hidden. It is tempting because BLEU is a long-standing translation metric; it would be the right choice when scoring surface similarity to reference translations, not for measuring confidence against factual correctness.
- ✗
ROUGE-L
Why it's wrong here
ROUGE-L measures longest-common-subsequence overlap with a reference, which rewards fluent phrasing regardless of factual truth, leaving overconfident errors undetected. It is tempting because ROUGE is standard for summarisation evaluation; it would be correct when judging how much reference content a summary recalls, not for assessing whether confidence scores are calibrated.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.