mediumMultiple Choice
Generative AI Leader Practice Question: Fine-tuning a generative AI model on proprietary…
A company is fine-tuning a generative AI model on proprietary customer data. They are concerned about copyright and IP issues when using the model commercially. What is the BEST practice to mitigate these risks?
⚠ Common exam trap
The trap is assuming that technical measures like watermarking or disclaimers can substitute for proper data licensing; the exam expects candidates to recognize that IP risk is fundamentally a data-rights and documentation problem.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Only use data that is explicitly licensed for commercial use and document its provenance
The most reliable way to mitigate copyright and IP risk when fine-tuning on proprietary data is to ensure every piece of training data is explicitly licensed for commercial use and to maintain clear documentation of its provenance. This creates an auditable chain of custody that supports both legal defensibility and compliance. It addresses the root cause—data rights—rather than trying to patch symptoms after generation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Apply SynthID to all generated content to prove origin
Why it's wrong here
SynthID watermarks AI-generated output for provenance detection; it neither prevents copyrighted training data entering the model nor grants commercial usage rights. It is tempting because watermarking supports transparency, but it would be correct only where the requirement is labelling synthetic media, not licensing clearance.
- ✓
Only use data that is explicitly licensed for commercial use and document its provenance
Why this is correct
Licensed commercial-use data with documented provenance directly satisfies the copyright and IP constraint: each training example carries traceable rights, so commercial output cannot infringe unlicensed material. Provenance records also evidence due diligence if a claim arises.
- ✗
Include a disclaimer on all outputs that the company is not liable for IP infringement
Why it's wrong here
A disclaimer shifts no legal liability and does not prevent infringing output; it addresses neither training-data rights nor commercial licensing. It is tempting because disclaimers are common in AI deployments, but they would be correct only for managing user expectations, not mitigating copyright risk.
- ✗
Use a model trained on publicly available data only
Why it's wrong here
Publicly trained models can still reproduce copyrighted material and carry unclear provenance, so restricting training data does not secure commercial IP rights. It is tempting because public data feels lower-risk, but it would be correct only when the goal is avoiding proprietary-data exposure rather than licensing assurance.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.