Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A research lab is fine-tuning a large language model on a small dataset of medical records. They observe that the model overfits, memorizing specific patient details and producing outputs that violate privacy regulations. Which technique should they apply to improve generalization and reduce memorization?
⚠ Common exam trap
Google Cloud often tests the misconception that early stopping or batch size adjustments can prevent memorization, when in fact only techniques like differential privacy directly bound the influence of individual training examples.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply differential privacy (DP-SGD) during fine-tuning
Differential privacy (DP-SGD) is the correct technique because it directly addresses memorization of sensitive patient data by adding calibrated noise to the gradient updates during fine-tuning. This bounds the model's ability to encode any single individual's information, improving generalization and ensuring compliance with privacy regulations like HIPAA.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the batch size to 64
Why it's wrong here
Larger batches change gradient noise and optimisation dynamics but do not prevent a model from memorising rare training examples; with few records, each batch still contains identifiable patient data. It is tempting because larger batches stabilise training and speed convergence, yet the requirement is reducing memorisation, which calls for differential privacy or regularisation.
- ✗
Increase the number of training epochs
Why it's wrong here
More epochs deepen overfitting on a small dataset, reinforcing memorisation of individual patient records rather than improving generalisation. It is tempting because additional training raises accuracy when a model is underfit, but here the training loss is already low and validation performance is degrading, so further passes worsen the privacy violation.
- ✗
Use early stopping based on validation loss
Why it's wrong here
Early stopping halts training when validation loss stops improving, but the model has already memorised patient details by then; it does not alter the training objective that drives memorisation. It is tempting because it genuinely curbs overfitting on standard supervised tasks, yet differential privacy or regularisation is needed to prevent record-level memorisation.
- ✓
Apply differential privacy (DP-SGD) during fine-tuning
Why this is correct
DP-SGD injects calibrated noise into per-sample gradients during fine-tuning, bounding any single record's influence on learned parameters. This directly curbs the memorisation of individual patient details that breaches privacy rules, while the noise also regularises training so the model generalises better to unseen records.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.