Courseiva

Generative AI Leader Fundamentals of Generative AI Practice Question

A team is tuning a large language model for a question-answering task. They notice the model gives high confidence scores to answers that are factually incorrect. Which evaluation metric should they primarily use to detect this overconfidence problem?

⚠ Common exam trap

Google Cloud often tests the distinction between intrinsic evaluation metrics (like perplexity) and calibration metrics, leading candidates to mistakenly choose perplexity when the core issue is confidence miscalibration rather than general model uncertainty.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Expected Calibration Error (ECE)

Expected Calibration Error (ECE) directly measures the alignment between a model's predicted confidence and its actual accuracy. In this scenario, high confidence on incorrect answers indicates miscalibration, and ECE quantifies this mismatch by binning predictions by confidence and computing the average absolute difference between accuracy and confidence per bin.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Perplexity

    Why it's wrong here

    Perplexity measures how fluently a model predicts token sequences, not whether stated confidence matches factual accuracy, so it cannot expose overconfidence. It is tempting because perplexity is a standard intrinsic metric for language modelling and detects distributional drift; it would suit comparing model versions or detecting unfamiliar inputs, not calibration of answer correctness.

  • ✓

    Expected Calibration Error (ECE)

    Why this is correct

    Expected Calibration Error directly measures the gap between predicted confidence and actual accuracy across probability bins, so systematically high confidence on wrong answers produces a large ECE value. This satisfies the stem's overconfidence constraint, unlike accuracy or F1, which ignore confidence entirely and cannot expose miscalibration.

  • ✗

    BLEU score

    Why it's wrong here

    BLEU compares n-gram overlap between generated and reference text, so a confidently wrong answer sharing wording with a reference scores well and the miscalibration stays hidden. It is tempting because BLEU is a long-standing translation metric; it would be the right choice when scoring surface similarity to reference translations, not for measuring confidence against factual correctness.

  • ✗

    ROUGE-L

    Why it's wrong here

    ROUGE-L measures longest-common-subsequence overlap with a reference, which rewards fluent phrasing regardless of factual truth, leaving overconfident errors undetected. It is tempting because ROUGE is standard for summarisation evaluation; it would be correct when judging how much reference content a summary recalls, not for assessing whether confidence scores are calibrated.

About these practice questions

This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.