Courseiva

Generative AI Leader Google Cloud's Generative AI Offerings Practice Question

A financial services firm needs to generate synthetic data for training models while ensuring that no real customer data leaks. Which technique should they use?

⚠ Common exam trap

Many candidates confuse data masking or redaction (which only hide data in the training set) with techniques that prevent model memorization, overlooking that models can still leak sensitive information through inference even when the input data is obfuscated.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Differential privacy during fine-tuning

Differential privacy during fine-tuning is the correct technique because it adds calibrated noise to the training process, ensuring that the synthetic data generated does not reveal information about any individual real customer record. This approach provides a formal mathematical guarantee of privacy, making it suitable for generating synthetic data that preserves statistical properties while preventing data leakage. In contrast, other methods like redaction, masking, or using a public model do not inherently prevent the model from memorizing and reproducing sensitive information.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Using the Vertex AI PII redaction service

    Why it's wrong here

    PII redaction removes identifiers from existing text but does not create synthetic records, so it cannot supply training data. It is tempting because it directly addresses leakage concerns, and would be correct when scrubbing sensitive fields from real documents before storage or analysis rather than generating synthetic data.

  • ✗

    Using a public foundation model without fine-tuning

    Why it's wrong here

    A public foundation model without fine-tuning generates generic output and cannot produce domain-specific synthetic records for the firm's training needs. It is tempting because no customer data is exposed, and would be correct for general-purpose text generation where bespoke, domain-aligned synthetic data is not required.

  • ✗

    Data masking before training

    Why it's wrong here

    Masking alters existing records, so the model still trains on transformed real data and residual re-identification risk remains; it does not synthesise new records. It is tempting because masking is a standard privacy control, and would be correct when sanitising production datasets for analytics rather than generating synthetic training data.

  • ✓

    Differential privacy during fine-tuning

    Why this is correct

    Differential privacy adds calibrated noise during fine-tuning, bounding any single customer record's influence on the model, so synthetic outputs cannot reveal real individuals. This satisfies the stem's constraint that no actual customer data leaks while still generating usable training data.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.