Courseiva

Reduce Toxic Content with Safety Filters

A developer is using Vertex AI Gemini API for a chatbot. The chatbot sometimes outputs harmful content. What is the best first step to mitigate this?

⚠ Common exam trap

Google Cloud often tests the misconception that the first step to mitigate harmful content is to fine-tune the model, when in reality the immediate, low-cost, and recommended first step is to leverage the API's built-in safety filters and settings.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use safety filters and safety settings in the API request

The Vertex AI Gemini API provides built-in safety filters and configurable safety settings (e.g., `safety_settings` parameter with categories like `HARM_CATEGORY_HARASSMENT` and thresholds like `BLOCK_ONLY_HIGH`) that allow developers to block harmful outputs at inference time without retraining. This is the fastest and most direct first step to mitigate harmful content, as it requires no additional infrastructure or model modification.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Fine-tune the model on curated safe data

    Why it's wrong here

    Fine-tuning changes model weights for task or style adaptation; it does not apply the configurable safety filters and thresholds that block harmful output at inference. It is tempting because curated data feels like a durable fix, and it would be right when the goal is domain-specific behaviour, not content safety.

  • ✗

    Add a human-in-the-loop review

    Why it's wrong here

    Human review intercepts outputs after generation, so harmful content still reaches users before anyone checks it. It is tempting because review catches edge cases, and it would be correct where low-volume, high-stakes decisions justify per-response vetting rather than automated filtering at the API layer.

  • ✓

    Use safety filters and safety settings in the API request

    Why this is correct

    Safety filters and safety settings are applied per API request, letting the developer block harmful categories before responses reach users. This is the fastest mitigation, requiring no retraining or architectural change to the Gemini chatbot.

  • ✗

    Switch to a smaller model

    Why it's wrong here

    Model size does not determine safety behaviour; a smaller Gemini variant can still emit harmful content and may degrade quality. It is tempting because smaller models are cheaper and faster, and it would be correct when the requirement is reducing latency or cost, not mitigating unsafe outputs.

About these practice questions

This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.