Courseiva
mediumMultiple Select

Generative AI Leader Practice Question: A data scientist is fine-tuning a generative AI…

A data scientist is fine-tuning a generative AI model for customer sentiment analysis. To ensure the fine-tuned model does not inadvertently memorize and reproduce personally identifiable information (PII) from the training data, which THREE practices should they follow? (Select 3)

⚠ Common exam trap

Generative AI Leader often tests privacy-preserving techniques; candidates may confuse model optimization techniques (quantization) or evaluation methods (cross-validation) with privacy practices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply differential privacy during fine-tuning

Option A is correct because applying differential privacy during fine-tuning adds calibrated noise to the training process, mathematically bounding the influence any single training record can have on the model's parameters and thereby limiting memorization of PII. Option C is correct because removing or anonymizing PII in the training data eliminates the sensitive content before it can ever be learned, which is the most direct defense against the model reproducing it. Option D is correct because data minimization restricts training to only the data necessary for the sentiment-analysis task, reducing the volume of PII exposed to the model and shrinking the attack surface for memorization. Option B is not correct because quantization only compresses weights to lower precision to reduce size and inference cost; it does not prevent the model from having memorized PII during training. Option E is not correct because k-fold cross-validation is an evaluation/resampling technique for estimating generalization performance and does not remove, mask, or otherwise protect PII in the training data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Apply differential privacy during fine-tuning

    Why this is correct

    Differential privacy adds calibrated noise during fine-tuning, bounding any single record's influence on learned parameters. This mathematically limits memorisation, so the model cannot reliably reproduce PII from training examples, directly satisfying the stem's requirement to prevent PII reproduction.

  • ✗

    Apply model quantization to reduce model size

    Why it's wrong here

    Quantization compresses weights to lower precision, shrinking memory and speeding inference; it does not remove PII memorised during fine-tuning. It is tempting because quantization is a standard optimisation step, and it would be correct for deployment efficiency, but memorisation is addressed by data sanitisation, differential privacy and output filtering.

  • ✓

    Remove or anonymize all PII from the training data

    Why this is correct

    Removing or anonymising PII before fine-tuning eliminates the sensitive data itself, so there is nothing for the model to memorise or reproduce. This directly satisfies the stem's constraint by addressing the risk at its source rather than relying on post-hoc detection.

  • ✓

    Use only a subset of data that is necessary for the task (data minimization)

    Why this is correct

    Data minimisation restricts fine-tuning to only the records needed for sentiment analysis, shrinking the pool of PII the model could memorise. Fewer sensitive examples means lower reproduction risk, directly satisfying the stem's requirement to prevent PII leakage from training data.

  • ✗

    Use k-fold cross-validation to evaluate the model

    Why it's wrong here

    Cross-validation estimates generalisation performance by partitioning data into training and validation folds; it neither detects nor removes PII, and memorised records persist across every fold. It is tempting because it is the standard technique for tuning hyperparameters and comparing model variants, which is the correct choice when the goal is reliable performance estimation rather than privacy assurance.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.