easyMultiple Choice
Generative AI Leader Practice Question: A data scientist wants to compare the performance…
A data scientist wants to compare the performance of three different foundation models for a text summarization task. They have a labeled dataset of summaries. Which Vertex AI tool should they use to perform this evaluation?
⚠ Common exam trap
Many exam-takers confuse model discovery (Model Garden) or pipeline-specific evaluation (RAG Engine, Agent Builder) with general-purpose generative model evaluation, which lives in Vertex AI Studio.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Vertex AI Studio - Evaluation
Vertex AI Studio includes a built-in Evaluation feature that lets you compare foundation models side-by-side using your own labeled dataset, computing metrics like ROUGE, BLEU, and human-preference scores for tasks such as summarization. It is purpose-built for evaluating generative model outputs against ground-truth references, which matches the data scientist's need exactly.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Vertex AI RAG Engine - Retrieval evaluation
Why it's wrong here
Retrieval evaluation scores how well a RAG system's retrieved context supports grounded answers, not how three foundation models summarise a labelled dataset. It is tempting because it is Vertex AI's evaluation offering, but it would be correct when tuning retrieval quality for a RAG pipeline, not comparing generative summarisation output.
- ✗
Vertex AI Agent Builder - Agent evaluation
Why it's wrong here
Agent evaluation measures an agent's tool-calling, task completion and reasoning trajectories, not summarisation quality against reference summaries. It is tempting because it is Vertex AI's evaluation tooling, but it would be correct when validating a deployed conversational or action-taking agent rather than comparing foundation model summarisation outputs.
- ✗
Model Garden - Model comparison view
Why it's wrong here
Model Garden's comparison view presents model metadata, capabilities and sample prompts side by side; it does not score outputs against a labelled dataset. It is tempting because it is the obvious place to browse and contrast models, and would be correct for shortlisting candidates before committing to a formal evaluation run.
- ✓
Vertex AI Studio - Evaluation
Why this is correct
Vertex AI Studio's evaluation capability scores generative model outputs against a labelled reference dataset using metrics such as ROUGE for summarisation. Running all three foundation models through it produces comparable, quantitative results, satisfying the requirement to compare summarisation performance.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.