Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A team is using Vertex AI Pipelines to deploy a generative AI model for real-time inference. The model sometimes generates harmful content. They want to implement a safety filter that checks the output before returning it to the user, but they need to minimize latency. Which approach best balances safety and performance?
⚠ Common exam trap
Google Cloud often tests the misconception that safety must be integrated into the generative model itself (e.g., via retraining or fine-tuning), when in practice a separate, lightweight post-processing filter is the standard for low-latency production systems.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a secondary lightweight classifier to filter outputs in real-time.
Deploying a secondary lightweight classifier (e.g., a distilled BERT or a small logistic regression model) as a post-processing filter allows real-time inference with minimal latency overhead. This approach decouples safety from the primary generative model, enabling fast rejection of harmful outputs without retraining or blocking the main inference pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a secondary lightweight classifier to filter outputs in real-time.
Why this is correct
A lightweight secondary classifier inspects each generated output and blocks harmful content before it reaches the user. Its low compute overhead adds minimal latency compared with a full model-based filter, satisfying the stem's need to balance safety against real-time inference performance.
- ✗
Retrain the model on every flagged harmful output.
Why it's wrong here
Retraining on each flagged output adds training cycles before any response returns, so latency balloons and safety is not enforced on the current request. It is tempting because retraining reduces future harmful generations, which suits periodic model improvement, not inline filtering of live inference responses.
- ✗
Manually review all outputs before delivery.
Why it's wrong here
Manual review cannot scale to real-time inference volumes and adds human latency, defeating the performance requirement entirely. It is tempting because human reviewers catch nuanced harmful content that automated filters miss, making it valid for low-volume, offline moderation workflows where throughput is not a constraint.
- ✗
Disable safety checks to improve latency.
Why it's wrong here
Removing safety checks eliminates the filtering step entirely, so harmful content reaches users unfiltered, violating the stated safety requirement. It is tempting because latency drops, and it would suit a trusted internal batch workload where output review happens downstream, not real-time user-facing inference.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.