Courseiva

Generative AI Leader Google Cloud's Generative AI Offerings Practice Question

An organization is using Vertex AI Gemini API for a multimodal chatbot. They notice that the model sometimes provides incorrect information with high confidence. They want to reduce hallucinations without retraining the model. What is the most effective approach?

⚠ Common exam trap

A common misconception is that adjusting generation parameters (like temperature or token limits) can fix factual accuracy issues, when in reality only grounding or retrieval-augmented generation can address hallucinations without retraining.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Provide ground-truth context from a knowledge base using grounding

Grounding with a knowledge base is the most effective approach because it forces the model to base its responses on verified, external data rather than relying solely on its parametric knowledge. By providing ground-truth context via Vertex AI's grounding feature, the model can cross-reference its outputs with authoritative sources, significantly reducing the likelihood of hallucinations. This method directly addresses the root cause of incorrect high-confidence responses without requiring retraining.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Provide ground-truth context from a knowledge base using grounding

    Why this is correct

    Grounding retrieves authoritative content from a knowledge base and injects it as context, so responses are anchored to verified facts rather than parametric memory. This reduces hallucination without retraining, satisfying the stem's constraint of no model fine-tuning.

  • ✗

    Increase the temperature parameter to make the model more creative

    Why it's wrong here

    Raising temperature increases sampling randomness, which amplifies rather than suppresses confident fabrication, so it directly worsens the hallucination problem described. It is tempting because temperature legitimately controls output diversity, and would suit brainstorming or creative-writing tasks where varied, unexpected responses are wanted rather than factual grounding.

  • ✗

    Reduce the maximum output tokens to force concise answers

    Why it's wrong here

    Reducing maximum output tokens constrains response length, not factual grounding; a short answer can still be confidently fabricated, so hallucination rates remain unchanged. It is tempting because token limits control verbosity and cost, and would suit scenarios needing shorter completions or latency reduction — not accuracy improvements.

  • ✗

    Adjust safety settings to filter uncertain responses

    Why it's wrong here

    Safety settings filter based on content categories, not correctness.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.