mediumMultiple Choice
Generative AI Leader Practice Question: Using Vertex AI to fine-tune a large language…
A company is using Vertex AI to fine-tune a large language model on proprietary customer support transcripts. The transcripts contain personally identifiable information (PII) such as names and email addresses. What is the BEST practice to comply with Google's AI Principles on privacy?
⚠ Common exam trap
The trap here is the plausible-sounding 'consent cures everything' answer — candidates pick C because consent feels like the strongest privacy control, but Google's AI Principles and data-minimization best practice require de-identification regardless of consent.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
De-identify the transcripts by removing or masking PII before fine-tuning
De-identifying the transcripts by removing or masking PII before fine-tuning is the best practice because it enforces data minimization and privacy-by-design, aligning with Google's AI Principles (which explicitly call for avoiding creating or reinforcing unfair bias and protecting privacy). Fine-tuning on raw PII risks the model memorizing and later regurgitating names, emails, or other sensitive data in generated outputs. Masking or tokenizing PII before training removes that exposure while preserving the linguistic patterns needed for the support use case.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
De-identify the transcripts by removing or masking PII before fine-tuning
Why this is correct
Masking or removing names and email addresses before fine-tuning prevents the model from memorising and later regurgitating personal data, directly satisfying the privacy constraint in the stem. De-identification at the data-preparation stage is the control Google's AI Principles expect, since once PII is embedded in weights it cannot be reliably removed.
- ✗
Use a model that has already been trained on similar data to avoid fine-tuning
Why it's wrong here
Reusing a model trained on similar data does not remove the PII from the transcripts being fine-tuned, so the privacy exposure remains. It is tempting because it avoids a training run, but the correct practice is de-identifying or redacting PII before fine-tuning, or using Vertex AI's data governance controls.
- ✗
Obtain consent from all customers whose data appears in the transcripts
Why it's wrong here
Consent alone does not satisfy Google's AI Principles, which require minimising PII use and applying privacy-preserving techniques such as de-identification before training. Consent is relevant for lawful processing under data-protection regimes, but the stem asks specifically about the AI Principles' privacy expectation of data minimisation and purpose limitation.
- ✗
Fine-tune the model directly on the transcripts, as the model will not memorize exact data
Why it's wrong here
Fine-tuning on raw transcripts risks memorisation and regurgitation of names and email addresses, breaching the privacy principle's requirement to limit data exposure. Direct fine-tuning is legitimate when the corpus is already de-identified or synthetic; here the PII must be removed or masked before training, not assumed away.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.