Generative AI Leader Google Cloud's Generative AI Offerings Practice Question
A machine learning engineer is tuning a large language model on Vertex AI for question answering. They want to evaluate the model's performance before deployment. Which THREE metrics should they consider?
⚠ Common exam trap
Many exam-takers confuse operational metrics (like cost or training time) with evaluation metrics that directly measure model output quality, leading them to select options that are irrelevant to performance assessment.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
F1 score
For a question-answering evaluation on Vertex AI, F1 score (B) is correct because it measures token-level overlap between the predicted answer and the ground-truth answer, balancing precision and recall and handling partial matches. Exact match (C) is correct because it reports the percentage of predictions that exactly equal the reference answer, a standard metric in QA benchmarks such as SQuAD. ROUGE-L score (E) is correct because it measures the longest common subsequence between generated and reference text, capturing fluency and recall-oriented overlap useful for generative QA outputs. Cost per training epoch (A) and training time per epoch (D) are operational/training-efficiency measures, not model quality metrics, so they do not evaluate performance before deployment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cost per training epoch
Why it's wrong here
Cost per training epoch is a financial efficiency measure, not a model quality metric, so it cannot assess question-answering accuracy, fluency or groundedness. It is tempting because tracking spend per epoch genuinely matters for budget control during iterative tuning, but evaluation before deployment requires metrics such as exact match, F1 or ROUGE.
- ✓
F1 score
Why this is correct
F1 score is the harmonic mean of precision and recall, capturing the balance between correctly retrieved answers and missed ones. For question answering it exposes whether the tuned model trades precision for coverage, which accuracy alone would hide.
- ✓
Exact match (EM)
Why this is correct
Exact match measures the percentage of predictions matching the reference answer string exactly. It is a strict, unambiguous question-answering metric that reveals whether the tuned model produces precisely correct short answers rather than merely fluent or partially overlapping text.
- ✗
Training time per epoch
Why it's wrong here
Training time per epoch measures tuning cost, not answer quality, so it cannot indicate whether the model answers questions well. It is tempting as an efficiency signal during hyperparameter tuning, but it is correct only when comparing compute budgets, not evaluating deployed question-answering performance.
- ✓
ROUGE-L score
Why this is correct
ROUGE-L measures longest common subsequence overlap between generated and reference answers, capturing fluency and recall without requiring exact n-gram matches. This suits question answering evaluation on Vertex AI, where responses vary in phrasing yet must retain reference content, satisfying the stem's pre-deployment performance assessment need.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.