Courseiva

CCNA Improve Gen Ai Output Questions

75 of 160 questions · Page 2/3 · Improve Gen Ai Output topic · Answers revealed

76
MCQeasy

A support team uses a generative AI assistant to answer customer questions. Agents report that the answers are often too long and include unnecessary background. The team wants the responses to be concise and directly address the question. Which prompt adjustment is most appropriate?

A.Add a system instruction that the model should always provide a detailed explanation for every answer.
B.Lower the model's temperature to zero so answers become deterministic.
C.Tell the model to answer in two sentences or fewer and to skip background unless the customer asks for it.
D.Increase the top-p value so the model considers more possible words.
AnswerC

Adding an explicit length and scope constraint directly addresses the verbosity problem. The model receives a clear instruction about how many sentences to use and what to omit. This is a simple prompt-engineering fix that does not require retraining or infrastructure changes. It also preserves the assistant's ability to answer follow-up questions when background is requested.

Why this answer

The most appropriate adjustment is an explicit prompt constraint that limits answer length and omits background unless requested. Verbosity is a content and instruction-following issue, so a clear directive in the prompt is the right control. Sampling parameters like temperature and top-p do not directly enforce conciseness, and asking for more detail would worsen the problem.

Exam trap

The trap here is confusing randomness controls such as temperature and top-p with instruction-following controls such as explicit length and scope constraints.

77
MCQeasy

A data scientist is using the Gemini API to generate product descriptions for an e-commerce site. The descriptions are often too verbose and include speculative claims that are not in the product specifications. The scientist wants to reduce hallucinations and control the length of the output without retraining the model. What should they do?

A.Increase the max output token count to 2048 and decrease temperature to 0.1.
B.Refine the prompt to be concise and include instructions to stick to facts and limit output to 50 words.
C.Add three few-shot examples of short, factual descriptions.
D.Set temperature to 0.0 and top_k to 1.
AnswerB

Refining the prompt directly constrains generation through the model's instruction-following behaviour, requiring no retraining. Explicit instructions to limit output to 50 words satisfy the length constraint, while directing the model to stick to supplied product specifications reduces speculative claims at inference time. This addresses both the verbosity and hallucination issues within the Gemini API workflow.

Why this answer

Refining the prompt to be concise and include explicit instructions to stick to facts and limit output to 50 words directly addresses both issues without retraining. Prompt engineering is the most effective technique for controlling output length and reducing hallucinations in the Gemini API, as it guides the model's behavior through natural language constraints rather than altering generation parameters. Note that few-shot examples (Option C) are also a form of prompt engineering and could help, but the question asks for the most direct approach: explicit natural-language instructions in the prompt are the simplest and most reliable way to enforce a strict word limit and factual grounding, whereas few-shot examples alone do not guarantee a specific output length.

Exam trap

This exam often tests the misconception that adjusting generation parameters like temperature or top_k is the primary way to control factual accuracy and length, when in fact prompt engineering is the most direct and effective method for these specific requirements without retraining. A related trap is assuming that few-shot examples are always superior to explicit instructions; for strict length control and factual grounding, clear natural-language constraints in the prompt are the most direct solution.

How to eliminate wrong answers

Option A is wrong because increasing the max output token count to 2048 would make the descriptions even more verbose, which is the opposite of what the data scientist wants; decreasing temperature to 0.1 reduces randomness but does not enforce factual adherence or length limits. Option C is wrong because adding three few-shot examples can improve style and structure but does not reliably prevent speculative claims or enforce a strict word count, especially if the examples are not perfectly aligned with the desired constraints. Option D is wrong because setting temperature to 0.0 and top_k to 1 makes the output deterministic and repetitive, which reduces creativity but does not inherently eliminate hallucinations or control verbosity; the model may still generate speculative content based on its training data.

78
MCQmedium

A retail company is deploying a generative AI chatbot on Vertex AI to provide product recommendations. The chatbot uses a base foundation model with no fine-tuning. Users report that the chatbot sometimes gives offensive or insensitive responses. The team must quickly implement safety controls without modifying the model. They also want to reduce irrelevant off-topic answers. Which combination of techniques should they apply?

A.Fine-tune the model on a curated dataset of safe retail conversations.
B.Set temperature to 0.0 and top_p to 0.1.
C.Enable Vertex AI Safety Filters and craft system instructions defining appropriate behavior.
D.Provide 50 few-shot examples of safe interactions.
AnswerC

Safety filters block harmful content at the API layer without retraining the model, while system instructions constrain tone and scope, reducing off-topic replies. Together they deliver fast, model-agnostic guardrails meeting both safety and relevance requirements.

Why this answer

Vertex AI Safety Filters provide out-of-the-box content moderation without modifying the model, and crafting system instructions (system-level prompts) can constrain the chatbot's behavior to stay on-topic and avoid offensive responses. This combination addresses both safety and relevance without requiring fine-tuning or altering model parameters.

Exam trap

Google often tests the distinction between parameter tuning (temperature/top_p) and safety mechanisms—candidates mistakenly think lowering randomness prevents offensive outputs, but safety requires explicit filtering or instruction-based guardrails, not just reduced creativity.

How to eliminate wrong answers

Option A is wrong because fine-tuning requires modifying the model, which contradicts the requirement to not modify the model; it also takes time and resources, not a 'quick' fix. Option B is wrong because setting temperature to 0.0 and top_p to 0.1 reduces randomness and diversity but does not prevent offensive or insensitive responses—it only makes outputs more deterministic, not safer. Option D is wrong because providing 50 few-shot examples of safe interactions can guide the model but does not guarantee safety filtering; it also requires careful curation and may not scale, and the model can still generate off-topic or offensive responses outside the examples.

79
MCQmedium

Refer to the exhibit. A team attempted to start a model tuning job but received the error 'Quota limit exceeded for tuning jobs in region us-central1'. What is the most appropriate action?

A.Request a quota increase for tuning jobs in us-central1
B.Change the region to us-west1 and retry
C.Reduce the size of the training data
D.Use a different base model
AnswerA

The error names a regional quota ceiling, not a configuration fault, so the tuning job cannot start until capacity is granted. Requesting an increase for us-central1 directly lifts that constraint, letting the job run in the required region without redesigning the pipeline or moving data.

Why this answer

The error 'Quota limit exceeded for tuning jobs in region us-central1' indicates that the project has reached its predefined resource quota for model tuning operations in that specific region. The most appropriate action is to request a quota increase from Google Cloud, as this directly resolves the capacity limitation without altering the job's configuration or data. Quotas are per-region limits enforced by the AI Platform to ensure fair resource allocation, and increasing the quota is the standard procedure when legitimate tuning needs exceed the default allowance.

Exam trap

A common pitfall is assuming quota errors can be fixed by changing job parameters (region, data size, model) instead of recognizing that quotas are administrative limits requiring a formal increase request through Google Cloud.

How to eliminate wrong answers

Option B is wrong because changing the region to us-west1 does not address the root cause of the quota limit; the tuning job may still fail if the quota in us-west1 is also insufficient or if the model or data has regional dependencies. Option C is wrong because reducing the size of the training data does not affect the quota limit for tuning jobs; quotas are based on the number of concurrent or total tuning jobs, not on data size. Option D is wrong because using a different base model does not change the quota consumption for tuning jobs; the quota applies to the tuning operation itself, regardless of which base model is selected.

80
Multi-Selectmedium

A team is using a generative AI model to answer customer questions about a complex product. They want to improve the factual accuracy and reduce hallucinations. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Use a grounded prompt that instructs the model to answer only from provided context and to say 'I don't know' if unsure.
B.Ground the model with Retrieval-Augmented Generation (RAG) using an authoritative product knowledge base.
C.Fine-tune the model on a small set of generic customer service dialogues.
D.Reduce the top-k parameter to a very small value.
E.Increase the temperature to make the model more confident.
AnswersA, B

A grounded prompt explicitly constrains the model to rely on the supplied context and to admit uncertainty when the answer is not present. This reduces hallucinations by discouraging the model from inventing information. It works well with RAG and is a simple, effective prompt engineering technique for factual accuracy.

Why this answer

Grounding with RAG supplies authoritative context, and a grounded prompt instructs the model to use only that context and to express uncertainty when appropriate. Together they attack hallucinations at the source: the model no longer relies solely on memorized parameters and is discouraged from fabricating. Other techniques like temperature or top-k affect randomness, not factual grounding, and fine-tuning on generic data does not provide the needed product facts.

Exam trap

The trap here is thinking that sampling parameters or generic fine-tuning can improve factual accuracy, when only grounding techniques like RAG and grounded prompts address hallucination directly.

81
MCQeasy

A marketing team uses a Gemini model through the Vertex AI API to draft campaign copy. The drafts are creative but frequently wander off topic and include unsupported claims. The team wants a low-effort improvement before considering any model customization. Which action should they take first?

A.Rewrite the prompt to include a clear role, the target audience, explicit constraints, and a required output format.
B.Raise the temperature so the model explores more ideas and eventually lands on better copy.
C.Switch the model to a larger version with more parameters to improve instruction following.
D.Create a supervised fine-tuning job with a few hundred labeled campaign examples.
AnswerA

Prompt engineering is the lowest-effort, highest-leverage first step. Specifying a role, audience, constraints, and output format narrows the model's generation space and typically removes off-topic drift and unsupported claims without any training cost. It can be iterated in minutes and does not require new data or infrastructure.

Why this answer

Prompt engineering with an explicit role, audience, constraints, and output format is the fastest and cheapest way to align model output with a brief. Fine-tuning, larger models, and higher temperature all add cost or increase variance without addressing the underlying ambiguity in the instructions.

Exam trap

The trap here is reaching for fine-tuning or a bigger model before exhausting prompt engineering, even though the symptoms point to underspecified instructions.

82
MCQhard

A company is using a fine-tuned LLM for generating financial reports. They need to ensure that the output complies with regulatory standards and does not include speculative content. Which combination of techniques should they implement?

A.Increase the model's safety settings to maximum, use a low top-p value, and limit output tokens.
B.Fine-tune the model on historical compliant reports, use RAG with a regulatory database, and implement a human-in-the-loop review.
C.Use a larger model with more parameters and rely on its inherent knowledge.
D.Use a system instruction to adhere to regulations, set temperature to 0.0, and apply a keyword filter.
AnswerB

Fine-tuning on historical compliant reports instils regulatory tone and structure, while RAG grounds each generation in current regulatory text, preventing speculative output. Human-in-the-loop review provides the final compliance gate. Together these satisfy the stem's dual constraint: regulatory compliance and absence of speculative content.

Why this answer

Fine-tuning on historical compliant reports ensures the model learns from past regulatory requirements, RAG with a regulatory database provides up-to-date compliance information, and human-in-the-loop review adds a final verification layer to catch any non-compliant or speculative content.

Option A (safety settings, low top-p, limit tokens) may reduce harmful content but does not guarantee regulatory compliance. Option C (larger model) alone does not enforce specific regulations. Option D (system instruction, temperature 0.0, keyword filter) is insufficient for complex regulatory standards.

83
MCQeasy

A marketing team is using an LLM on Vertex AI to generate product descriptions. They want to consistently control the creativity and randomness of the output. Which parameter should they adjust?

A.Top-P
B.Temperature
C.Max output tokens
D.Top-K
AnswerB

Temperature controls the randomness of the model's output. Lower values make the output more deterministic and focused, while higher values increase creativity and diversity. In this scenario, adjusting temperature allows the team to consistently control the creativity of the generated product descriptions, aligning with their goal.

Why this answer

Temperature is the primary parameter for controlling the randomness and creativity of generative AI output. Lower temperatures yield more predictable, focused responses, while higher temperatures yield more diverse and creative ones. For a marketing team seeking consistent control over creativity, temperature is the correct choice.

Exam trap

The trap here is confusing sampling parameters like Top-K and Top-P with temperature, which directly governs creativity.

84
MCQhard

The exhibit shows the deployment configuration for a conversational AI model used in a finance application. Users report that responses are creative but often contain factually incorrect financial advice. Which parameter change would most improve factual accuracy?

A.Add grounding sources, such as "EnterpriseSearch" or "Web"
B.Lower temperature to 0.1
C.Increase topP to 1.0
D.Increase maxOutputTokens to 1024
AnswerA

Adding grounding sources constrains generation to retrieved enterprise or web content, so the model cites evidence rather than relying on parametric memory. This directly addresses the stem's factual-accuracy failure, where ungrounded creativity produces incorrect financial advice. Retrieval augmentation, not sampling parameters, is the mechanism that anchors responses in verifiable data.

Why this answer

Adding grounding sources like EnterpriseSearch or Web provides the model with access to authoritative, up-to-date financial data, which directly reduces hallucinations by anchoring responses in verified facts rather than relying solely on the model's parametric knowledge. This is the most effective technique for improving factual accuracy in a domain where correctness is critical.

Exam trap

The Generative AI Leader exam often tests the misconception that adjusting sampling parameters (temperature, topP) can fix factual accuracy issues, when in reality those parameters only control output randomness and diversity, not the truthfulness of the underlying knowledge.

How to eliminate wrong answers

Option B is wrong because lowering temperature to 0.1 makes the model more deterministic and less creative, but it does not introduce new factual information; it only reduces randomness in token selection, which cannot fix incorrect knowledge baked into the model. Option C is wrong because increasing topP to 1.0 includes all possible tokens in the sampling pool, which actually increases the chance of selecting less likely and potentially incorrect tokens, harming factual accuracy. Option D is wrong because increasing maxOutputTokens to 1024 allows longer responses but does not improve the correctness of the content; it may even amplify errors by generating more text based on the same flawed internal knowledge.

85
MCQmedium

After deploying a text-to-image model, the output images often contain distorted objects. The team suspects the prompt is too complex. Which prompt engineering technique should they try first?

A.Increase the guidance scale.
B.Add more descriptive adjectives.
C.Use a negative prompt to exclude distortions.
D.Break the prompt into simpler, separate steps.
AnswerD

Complex prompts overload the model's attention, producing distorted objects. Decomposing the prompt into simpler sequential steps reduces the conditioning burden per generation, letting the model resolve each element correctly before combining them into the final image.

Why this answer

Breaking a complex prompt into simpler, separate steps reduces the cognitive load on the diffusion model, allowing it to focus on generating each element sequentially. This technique, often called 'prompt decomposition' or 'step-by-step prompting,' directly addresses the root cause of distorted objects when the model struggles to attend to multiple conflicting details simultaneously in a single pass.

Exam trap

Candidates often mistakenly think that increasing guidance scale or adding more descriptive details always improves output quality, when in fact these actions can worsen distortions by over-constraining the model's latent space.

How to eliminate wrong answers

Option A is wrong because increasing the guidance scale forces the model to adhere more strictly to the prompt, which can amplify artifacts and distortions rather than reduce them, especially when the prompt is already too complex. Option B is wrong because adding more descriptive adjectives increases prompt complexity, making it harder for the model to disentangle attributes, often leading to more distorted or merged objects. Option C is wrong because using a negative prompt to exclude distortions is a reactive fix that does not address the underlying issue of prompt complexity; it may suppress some artifacts but cannot resolve the model's inability to handle too many simultaneous constraints.

86
MCQeasy

A product team uses Gemini via the Vertex AI API to draft customer emails. The drafts are accurate but often too long and include unnecessary background. The team wants shorter, more direct outputs while keeping the same model. Which approach should they take?

A.Fine-tune the model on a dataset of short emails to permanently change its verbosity.
B.Set the max output tokens parameter to a very low value and leave the prompt unchanged.
C.Revise the prompt to explicitly instruct concise, direct language and specify a target length or format.
D.Increase the top-p value so the model selects from a narrower set of likely tokens.
AnswerC

Prompt instructions are the primary control for output style. Telling the model to be concise, avoid background, and follow a target length or bullet format directly shapes the response. This preserves completeness of key information while meeting the brevity requirement. It works without changing model parameters or retraining.

Why this answer

The most direct and efficient fix for verbose outputs is to change the prompt to request concise, direct language with a specified length or format. Prompt engineering shapes style immediately without retraining or infrastructure changes. Token limits truncate rather than summarize, fine-tuning is overkill for a style tweak, and top-p affects randomness rather than length.

Exam trap

The trap here is reaching for parameter changes like max output tokens or top-p when the real issue is that the prompt never asked for brevity.

87
Multi-Selectmedium

A team wants to reduce hallucinations in a question-answering model. Which THREE techniques should they consider?

Select 3 answers
A.Fine-tune the model on a curated factual dataset
B.Use retrieval-augmented generation (RAG)
C.Apply prompt engineering with specific instructions to cite sources
D.Reduce the number of tokens in output
E.Increase the temperature parameter
AnswersA, B, C

Fine-tuning on factual data improves accuracy.

Why this answer

Fine-tuning on a curated factual dataset directly adjusts the model's weights to prioritize accurate, domain-specific knowledge, reducing the likelihood of generating unsupported or hallucinated content. This technique anchors the model's output in verified data, making it more reliable for question-answering tasks.

Exam trap

Google Cloud often tests the misconception that reducing output length or increasing randomness (temperature) can improve factual accuracy, when in reality these parameters control style and creativity, not truthfulness.

88
MCQmedium

A financial technology company has deployed a custom-tuned PaLM 2 model on Vertex AI to generate personalized investment recommendations for retail clients. The model was fine-tuned on a corpus of historical market data and advisory transcripts. Recently, the compliance team flagged that several recommendations contradicted SEC guidelines, and the model sometimes repeated prohibited statements from outdated training materials. The team has already implemented safety filters (e.g., blocking toxic content) and adjusted the model's system instructions to be more conservative. However, the issues persist. The model's deployment parameters are: temperature=0.4, top_p=0.9, max_output_tokens=500, and no grounding. The company must maintain compliance without significantly increasing latency. What should they do next?

A.Increase temperature to 0.7 to allow more diverse responses, and add a second model to verify outputs
B.Perform an additional fine-tuning round exclusively on the most recent SEC regulatory filings and compliance-approved content
C.Implement a chain-of-thought prompting technique that requires the model to explain its reasoning step by step
D.Configure Vertex AI grounding using a curated data store of real-time SEC regulations and market data
AnswerD

Grounding with a curated Vertex AI data store anchors generation to current SEC regulations, directly addressing the prohibited statements and contradictions sourced from outdated training materials. Unlike fine-tuning, retrieval injects authoritative text at inference time, so compliance updates take effect immediately without retraining. This satisfies the compliance constraint while adding only modest retrieval latency.

Why this answer

Configuring Vertex AI grounding with a curated data store of real-time SEC regulations directly addresses the root cause: the model is generating outputs that contradict current compliance rules. Grounding forces the model to base its responses on authoritative, up-to-date sources, which is more effective than safety filters or system instructions alone, and it avoids the latency increase of a second model or the risk of catastrophic forgetting from additional fine-tuning.

Exam trap

This exam often tests the misconception that fine-tuning or prompt engineering alone can solve compliance issues, when in fact grounding with authoritative data sources is the only reliable method for ensuring outputs adhere to real-time, external regulations without sacrificing latency.

How to eliminate wrong answers

Option A is wrong because increasing temperature to 0.7 would make outputs more random and less deterministic, increasing the likelihood of generating non-compliant statements, and adding a second model for verification would significantly increase latency and cost without fixing the underlying data contamination. Option B is wrong because performing additional fine-tuning on recent SEC filings risks catastrophic forgetting of the original training data and does not guarantee real-time compliance, as fine-tuning is static and cannot adapt to rapidly changing regulations. Option C is wrong because chain-of-thought prompting only improves reasoning transparency but does not constrain the model to use compliant sources; the model could still generate prohibited statements from its outdated training data.

89
MCQmedium

A product team at a retailer is using Vertex AI Studio to build a Gemini-powered assistant that answers questions about their internal return policy. Early tests show the model invents policy details such as a 45-day return window, even though the official policy allows only 30 days. The team wants the assistant to answer strictly from a set of approved policy PDFs stored in a Cloud Storage bucket, and they want to avoid retraining the model. Which technique should they use?

A.Adding a system instruction telling the model to be accurate and to never hallucinate policy details.
B.Supervised fine-tuning of the Gemini model on a labeled dataset of correct policy answers.
C.Lowering the model's temperature to 0 and increasing the top-k value in the generation configuration.
D.Retrieval-augmented generation (RAG) with a Vertex AI Search data store grounded on the approved policy documents.
AnswerD

RAG retrieves relevant passages from the indexed policy PDFs at query time and passes them to Gemini as grounding context, so answers reflect the approved 30-day policy instead of the model's pretrained assumptions. Vertex AI Search provides managed indexing and grounding, and no model retraining is required, which matches the team's constraint.

Why this answer

Grounding with retrieval-augmented generation lets the assistant fetch the relevant passages from the approved policy PDFs and condition its answer on that retrieved text, so the 30-day window is stated correctly and the response can cite the source. It requires no retraining, updates automatically when documents change, and directly addresses hallucinated policy details.

Exam trap

The trap here is assuming that lowering temperature or adding a firm instruction eliminates hallucination, when only supplying authoritative source content through grounding actually fixes a missing factual detail.

90
MCQmedium

A developer deployed a large language model on Vertex AI for real-time chat. Users report slow response times. The model generates sentences one word at a time. Which optimization should be applied to reduce latency?

A.Batch multiple user queries together.
B.Deploy the model with more accelerators.
C.Enable prompt caching to reuse previous queries.
D.Use streaming responses to start output earlier.
AnswerD

Autoregressive generation emits tokens sequentially, so the full response must complete before any output appears. Streaming returns tokens as they are produced, letting the client render the first words immediately and cutting perceived latency without changing the model.

Why this answer

Streaming responses allow the model to send tokens to the client as they are generated, rather than waiting for the full sequence to complete. This reduces perceived latency significantly in real-time chat, as users see the first word appear almost immediately, even though the total generation time remains similar.

Exam trap

The trap here is that candidates often confuse throughput optimization (batching or more accelerators) with latency reduction, failing to recognize that streaming directly minimizes the time users wait for the first visible output in real-time scenarios.

How to eliminate wrong answers

Option A is wrong because batching multiple user queries together increases latency for individual requests, as the system waits to accumulate enough queries before processing, which is counterproductive for real-time chat. Option B is wrong because deploying with more accelerators improves throughput and total generation speed, but does not address the fundamental issue of word-by-word generation latency; the model still outputs one token at a time, and the user must wait for the full response. Option C is wrong because prompt caching reuses previous queries to avoid recomputation, but this optimization targets repeated or similar prompts, not the latency of generating a new response token-by-token.

91
MCQmedium

A company wants to build a customer support chatbot that answers based on internal documentation. They use Vertex AI Search and want to ensure the model only uses retrieved documents. What should they do?

A.Fine-tune the model on the documentation
B.Enable grounding with Vertex AI Search
C.Increase max output tokens
D.Set temperature to 0.0
AnswerB

Grounding with Vertex AI Search constrains generation to retrieved internal documents, satisfying the requirement that answers derive only from that corpus. The model cites source passages rather than relying on parametric memory, which reduces hallucination. This directly meets the stem's constraint that the chatbot use solely retrieved documentation.

Why this answer

Grounding with Vertex AI Search ensures the model's responses are strictly based on the retrieved documents from the internal documentation, preventing hallucination or reliance on pre-trained knowledge. Grounding works by providing the model with a search result context that it must use as the sole source for generating answers, effectively constraining the output to the provided documents.

Exam trap

The exam often tests the distinction between controlling model behavior (temperature, token limits) and controlling the source of information (grounding), leading candidates to mistakenly choose temperature or token adjustments as a solution for hallucination prevention.

How to eliminate wrong answers

Option A is wrong because fine-tuning the model on the documentation would embed the knowledge into the model's parameters, but it does not guarantee that the model will only use that knowledge during inference; the model could still generate responses from its pre-trained weights or hallucinate. Option C is wrong because increasing max output tokens only controls the length of the response, not the source of the information; the model could still generate content not found in the retrieved documents. Option D is wrong because setting temperature to 0.0 makes the model deterministic (greedy decoding) but does not restrict the model to only use retrieved documents; it can still produce answers based on its internal knowledge.

92
MCQhard

A healthcare company is using a generative AI model to draft patient education materials. The model sometimes generates content that includes specific medical advice, which could be harmful if inaccurate. The company wants to ensure that the model's outputs are safe and do not provide medical recommendations. Which technique should they implement?

A.Apply a safety filter that blocks any output containing medical terminology.
B.Reduce the model's temperature to make outputs more deterministic and less likely to include advice.
C.Implement prompt engineering with explicit instructions to avoid providing medical advice.
D.Use reinforcement learning from human feedback (RLHF) to fine-tune the model to avoid giving advice.
AnswerC

Prompt engineering with clear instructions, such as 'Do not provide medical advice; only provide general information,' can effectively steer the model away from generating harmful recommendations. This is a low-cost, immediate solution that can be refined iteratively. It leverages the model's ability to follow instructions when they are explicit and well-crafted.

Why this answer

Prompt engineering with explicit instructions is a direct and efficient way to constrain the model's output. By clearly stating that the model should not provide medical advice, the model is guided to generate only general educational content. This approach is flexible and can be updated as needed without retraining, making it suitable for the healthcare company's requirement to ensure safety.

Exam trap

The trap here is thinking that technical parameter adjustments like temperature will solve content safety issues, when in fact clear instructions in the prompt are often more effective for controlling what the model says.

93
Multi-Selectmedium

Which TWO techniques are effective for reducing bias in generative AI model outputs?

Select 2 answers
A.Increasing model size to learn more patterns
B.Training on diverse and representative datasets
C.Relying solely on post-hoc filters
D.Using adversarial debiasing methods during fine-tuning
E.Limiting the model to only factual prompts
AnswersB, D

Correct: Diverse data helps reduce biased associations.

Why this answer

Training on diverse and representative datasets directly reduces sampling bias and coverage gaps in the training distribution, which are primary sources of stereotypical or skewed outputs. By ensuring the model sees balanced examples across demographics, contexts, and edge cases, it learns more equitable representations and reduces the likelihood of generating biased content.

Exam trap

Google Cloud often tests the misconception that increasing model size or adding post-hoc filters is sufficient to mitigate bias, when in reality these approaches fail to address the root causes of bias in training data and model representations.

94
Multi-Selecthard

A financial analyst uses generative AI to summarize earnings reports. The summaries vary in style. Which THREE methods can improve consistency? (Choose three.)

Select 3 answers
A.Set temperature to 0.2
B.Increase max output tokens
C.Enable citation mode
D.Use few-shot prompting with fixed examples
E.Fine-tune on a curated dataset of desired summaries
AnswersA, D, E

Reduces output randomness.

Why this answer

Setting temperature to 0.2 reduces randomness in token sampling, making the model more deterministic and less likely to produce stylistic variations. Lower temperatures (e.g., 0.1–0.3) narrow the probability distribution, forcing the model to select the most likely next token, which directly improves consistency across multiple summaries.

Exam trap

A common misconception in this Google exam is that increasing max output tokens or enabling citation mode improves consistency. In reality, these features control length and attribution respectively, not stylistic uniformity. The correct methods focus on reducing randomness (low temperature), providing consistent examples (few-shot), or fine-tuning.

95
MCQmedium

A user reports that the model's response to the same prompt varies significantly across different calls. Which parameter change would most likely reduce variability?

A.Decrease topK to 10.
B.Decrease temperature to 0.2.
C.Increase candidateCount to 3.
D.Increase maxOutputTokens to 2000.
AnswerB

Lowering temperature to 0.2 sharpens the softmax probability distribution, so high-probability tokens dominate sampling and near-deterministic output replaces the wide variance seen at higher values. This directly satisfies the stem's requirement to reduce run-to-run variability for an identical prompt, since temperature is the parameter governing sampling randomness.

Why this answer

Temperature controls the randomness of token sampling. Lowering temperature (e.g., to 0.2) makes the model's output more deterministic by reducing the probability of low-likelihood tokens, thus decreasing variability across calls for the same prompt.

Exam trap

Candidates often mistake topK or candidateCount for the primary control of output variability, but in Google's Vertex AI and Gen AI models, temperature is the direct parameter that governs randomness in token selection.

How to eliminate wrong answers

Option A is wrong because decreasing topK to 10 still allows sampling from a limited set of tokens, which can introduce variability if temperature is not also reduced; topK alone does not control randomness as directly as temperature. Option C is wrong because increasing candidateCount to 3 generates multiple independent responses, which increases variability rather than reducing it. Option D is wrong because increasing maxOutputTokens to 2000 only extends the maximum length of the response, not the consistency of the output; it has no effect on token selection randomness.

96
MCQhard

After fine-tuning a model on customer support data, the model starts using profanity. What is the most effective mitigation?

A.Add profanity to training data as negative examples
B.Reduce learning rate and retrain
C.Increase temperature to reduce confidence
D.Enable a safety attribute filter
AnswerD

A safety attribute filter intercepts model outputs and blocks or redacts harmful content such as profanity before it reaches users, directly addressing the fine-tuning side effect. Retraining or prompt engineering may reduce but not reliably prevent toxic generations.

Why this answer

Enabling a safety attribute filter is the most effective mitigation because it acts as a post-processing guardrail that blocks profanity at inference time, regardless of the model's training data. This is a standard practice in production LLM deployments, where safety filters (e.g., using keyword matching or classifier models) intercept and redact harmful outputs before they reach the user, providing immediate and reliable control without requiring retraining.

Exam trap

Google often tests the misconception that modifying training parameters (like learning rate or temperature) can fix output quality issues, when in fact post-processing filters are the standard, immediate solution for content safety in production LLM systems.

How to eliminate wrong answers

Option A is wrong because adding profanity as negative examples in training data can inadvertently reinforce the behavior or cause the model to learn spurious correlations, and it does not guarantee removal of already learned profanity patterns. Option B is wrong because reducing the learning rate and retraining only adjusts the model's weights during fine-tuning, which does not address the root cause of profanity generation and may not eliminate learned toxic patterns without extensive data curation. Option C is wrong because increasing temperature increases randomness in token sampling, which can actually increase the likelihood of generating profanity by making the model less deterministic, not reduce it.

97
MCQmedium

A marketing team uses a Gemini model to generate ad copy. They notice the outputs are repetitive and lack variety across multiple runs for the same prompt. They want more diverse creative options without sacrificing relevance. Which parameter adjustment should they make?

A.Set top-k to 1 so the model always picks the single most likely token.
B.Decrease the temperature to reduce repetition.
C.Reduce the max output tokens to force the model to vary its wording.
D.Increase the temperature to allow more varied token sampling.
AnswerD

Higher temperature flattens the probability distribution, allowing less likely but still relevant tokens to be selected. This produces more diverse creative outputs across runs. The team should balance temperature to avoid incoherence while gaining variety. It directly addresses the repetition issue without changing the prompt or model.

Why this answer

Increasing temperature broadens token sampling, yielding more varied creative outputs while still following the prompt. Lower temperature, top-k of 1, and reduced max output tokens all make outputs more deterministic or shorter, not more diverse. Temperature is the primary control for balancing creativity and coherence in ad copy generation.

Exam trap

The trap here is confusing length controls or greedy decoding with creativity controls, when temperature is the parameter that governs output diversity.

98
MCQhard

A research lab is fine-tuning a large language model on a small dataset of medical records. They observe that the model overfits, memorizing specific patient details and producing outputs that violate privacy regulations. Which technique should they apply to improve generalization and reduce memorization?

A.Increase the batch size to 64
B.Increase the number of training epochs
C.Use early stopping based on validation loss
D.Apply differential privacy (DP-SGD) during fine-tuning
AnswerD

DP-SGD injects calibrated noise into per-sample gradients during fine-tuning, bounding any single record's influence on learned parameters. This directly curbs the memorisation of individual patient details that breaches privacy rules, while the noise also regularises training so the model generalises better to unseen records.

Why this answer

Differential privacy (DP-SGD) is the correct technique because it directly addresses memorization of sensitive patient data by adding calibrated noise to the gradient updates during fine-tuning. This bounds the model's ability to encode any single individual's information, improving generalization and ensuring compliance with privacy regulations like HIPAA.

Exam trap

Google Cloud often tests the misconception that early stopping or batch size adjustments can prevent memorization, when in fact only techniques like differential privacy directly bound the influence of individual training examples.

How to eliminate wrong answers

Option A is wrong because increasing batch size to 64 reduces gradient variance but does not prevent memorization of specific patient details; it may even accelerate overfitting on a small dataset. Option B is wrong because increasing the number of training epochs exacerbates overfitting, causing the model to memorize more training examples and worsen privacy violations. Option C is wrong because early stopping based on validation loss only halts training when validation performance degrades, but it does not impose any privacy guarantee or fundamentally limit memorization of unique patient records.

99
MCQmedium

A team wants to improve the factual accuracy of their chatbot responses regarding internal company policies. What is the most effective approach?

A.Use few-shot prompting with example Q&A pairs
B.Increase the model's maximum tokens
C.Fine-tune the model on policy documents
D.Use RAG with Vertex AI Search indexing the policies
AnswerD

RAG retrieves relevant passages from the indexed policy corpus at query time and supplies them as grounding context, so answers cite actual internal documents rather than relying on parametric memory. Vertex AI Search handles indexing and retrieval, satisfying the requirement for factual accuracy on company-specific policies.

Why this answer

RAG with Vertex AI Search is the most effective approach because it retrieves relevant, up-to-date policy documents from a curated index and injects them into the prompt context at inference time, grounding the chatbot's responses in authoritative sources without modifying the underlying model. This ensures factual accuracy for dynamic or evolving policies, as the model can reference the exact text rather than relying on static training data.

Exam trap

A common misconception in the Google Gen AI Leader exam is that fine-tuning (Option C) is the best approach to improve factual accuracy for dynamic knowledge. In reality, RAG with Vertex AI Search is superior because it retrieves up-to-date policies from a curated index without retraining the model, and provides verifiable source citations.

How to eliminate wrong answers

Option A is wrong because few-shot prompting provides example Q&A pairs but does not guarantee the model will recall or cite the correct policy details, especially for nuanced or updated policies; it relies on the model's parametric memory, which can be incomplete or outdated. Option B is wrong because increasing the maximum tokens only expands the output length, not the factual grounding; it does nothing to improve the accuracy of the content generated. Option C is wrong because fine-tuning on policy documents embeds static knowledge into the model weights, making it difficult to update when policies change and risking catastrophic forgetting of other capabilities; it also does not provide a mechanism to cite specific sources or handle real-time retrieval.

100
MCQeasy

A company notices that their AI chatbot occasionally generates incorrect information. Which technique can best reduce hallucinations without retraining?

A.Use a longer system prompt without examples
B.Use system instructions to constrain the model to only answer from provided context
C.Set top_p to 0.1
D.Increase temperature to 0.9
AnswerB

System instructions steer generation by restricting the model to answer solely from supplied context, so unsupported claims are refused rather than invented. This constrains decoding behaviour at inference time, reducing hallucinations without any retraining, fine-tuning or modification of model weights.

Why this answer

Constraining the model to answer only from provided context directly addresses the root cause of hallucinations—the model generating information not grounded in verified sources. This technique, often implemented via system instructions or retrieval-augmented generation (RAG) pipelines, forces the model to rely on a trusted knowledge base rather than its parametric memory, effectively eliminating unsupported fabrications without requiring retraining.

Exam trap

Google often tests the misconception that adjusting sampling parameters (like top_p or temperature) can fix hallucinations, when in reality these parameters control randomness, not factual grounding, and the correct solution is to constrain the model's output to a trusted context.

How to eliminate wrong answers

Option A is wrong because using a longer system prompt without examples does not prevent hallucinations; it may actually increase the risk by introducing more ambiguous or conflicting instructions, and without explicit grounding constraints, the model can still generate unverified content. Option C is wrong because setting top_p to 0.1 reduces the diversity of token sampling but does not enforce factual accuracy—it merely makes outputs more deterministic, which can still produce confident hallucinations if the model's internal knowledge is flawed. Option D is wrong because increasing temperature to 0.9 increases randomness and creativity in outputs, which exacerbates hallucination risk by making the model more likely to generate improbable or fabricated information.

101
MCQhard

A global software company uses a generative AI model to produce localized release notes. Outputs are accurate but inconsistently formatted: some use tables, some bullet lists, and some paragraphs, which breaks the publishing pipeline. The team wants stable, machine-parseable formatting across all runs. Which technique should they prioritize?

A.Set the model's temperature to a moderate value and rely on repeated runs to average out formatting differences.
B.Instruct the model to think step by step about the release content before formatting the output.
C.Fine-tune the model on a dataset of previously published release notes to teach consistent formatting.
D.Define a strict output schema, such as a JSON or Markdown template, and require the model to conform to it.
AnswerD

A strict output schema removes formatting ambiguity by defining exactly which structure every response must follow. The model fills the template rather than choosing a format, so the publishing pipeline receives consistent, parseable output. This is prompt-level and enforceable with validation, requiring no retraining or sampling changes.

Why this answer

Inconsistent formatting is a specification gap. Declaring a strict output schema, such as a fixed JSON structure or Markdown template, forces every generation into the same shape and makes the output machine-parseable. Temperature averaging, reasoning prompts, and fine-tuning influence variability, reasoning, or style bias but do not guarantee structural conformity.

Exam trap

The trap here is thinking that fine-tuning or repeated sampling stabilizes formatting, when only an explicit output schema reliably constrains the structure of every response.

102
MCQhard

A retail company uses a generative AI model to create personalized product recommendations. The model sometimes generates recommendations that include products the company does not sell. Which technique should be used to prevent the model from generating non-existent products?

A.Increase the model's temperature to allow more creative recommendations.
B.Implement a post-processing filter that checks generated product names against the company's inventory database.
C.Fine-tune the model on a dataset of customer reviews.
D.Use a larger context window to include the entire product catalog in the prompt.
AnswerB

A post-processing filter that validates product names against the actual inventory ensures that only existing products are recommended. This directly prevents the inclusion of non-existent products by catching and removing them before presentation. It is a reliable and straightforward solution for this scenario.

Why this answer

A post-processing filter that cross-references generated product names with the company's inventory database ensures that only real products are recommended. This method is direct, reliable, and does not rely on the model's internal knowledge, effectively preventing hallucinations of non-existent products.

Exam trap

The trap here is assuming that fine-tuning or prompt engineering can completely eliminate hallucinations, when a validation step is often necessary.

103
MCQhard

A streaming platform uses a large generative model for personalized content suggestions. Budget constraints require minimizing inference costs without significantly degrading quality. Which approach is most effective?

A.Deploy the model on higher-end accelerators to save time.
B.Use a distilled version of the model.
C.Implement stronger safety filters to reduce output length.
D.Cache frequent prompts to avoid regeneration.
AnswerB

Distillation trains a smaller model to reproduce the large model's outputs, cutting inference compute and cost while retaining most recommendation quality. This satisfies the budget constraint without the significant quality degradation that cruder reductions would cause.

Why this answer

Distillation trains a smaller 'student' model to mimic a larger 'teacher' model, reducing parameter count and inference latency while retaining most of the recommendation quality. This directly addresses the budget constraint by lowering compute and memory costs per inference, making it the most effective approach among the options.

Exam trap

A common pitfall is assuming that caching frequent prompts or upgrading to higher-end accelerators reduces per-inference costs. Caching only helps with repeated queries, not unique recommendations; hardware upgrades increase fixed costs. Distillation directly reduces model size and inference compute, aligning with cost constraints.

How to eliminate wrong answers

Option A is wrong because deploying on higher-end accelerators increases hardware cost, not reduces it, and while it may save time, the budget constraint demands minimizing inference costs, not just time. Option C is wrong because stronger safety filters do not reduce output length in a meaningful way for cost savings; they add computational overhead for filtering and may degrade user experience by blocking valid suggestions. Option D is wrong because caching frequent prompts only avoids regeneration for identical inputs, but personalized content suggestions are inherently unique per user session, so cache hit rates are low and the approach does not address the core inference cost per unique request.

104
MCQeasy

A developer is using the Gemini API to generate code snippets. They notice the outputs often contain deprecated API calls. Which parameter adjustment or prompt strategy would most effectively encourage the model to use current APIs?

A.Add a system instruction specifying 'Use the most recent API version and avoid deprecated functions.'
B.Set top-p to 0.5 to reduce output diversity
C.Provide one few-shot example of a correct API call
D.Set temperature to 1.5 to increase creativity
AnswerA

A system instruction sets persistent behavioural guidance that conditions every response, so explicitly directing the model toward current API versions and away from deprecated functions steers generation more reliably than per-request wording. This directly targets the deprecated-call pattern observed in the outputs.

Why this answer

Adding a system instruction that explicitly directs the model to 'Use the most recent API version and avoid deprecated functions' directly influences the model's behavior at the prompt level. The Gemini API supports system instructions that act as persistent, high-level guidance, steering the model toward preferred output patterns—in this case, avoiding deprecated API calls. This is the most effective and direct method to enforce current API usage without altering sampling parameters or relying on limited examples.

Exam trap

This question tests the misconception that adjusting sampling parameters (like temperature or top-p) or providing a single example can reliably enforce content constraints, when in fact system instructions are the designed mechanism for persistent behavioral guidance in production-grade APIs.

How to eliminate wrong answers

Option B is wrong because setting top-p to 0.5 reduces the cumulative probability mass of token choices, which narrows output diversity but does not inherently bias the model toward current APIs; it may even suppress rare but correct modern API tokens. Option C is wrong because a single few-shot example provides only one instance of a correct API call, which is insufficient to override the model's training data bias toward deprecated APIs; the model may still default to older patterns. Option D is wrong because increasing temperature to 1.5 amplifies randomness and creativity, which can increase the likelihood of hallucinated or incorrect API calls, including deprecated ones, rather than encouraging adherence to current standards.

105
MCQmedium

A logistics company built a Gemini-powered assistant that answers driver questions about routes and hours-of-service rules. The assistant performs well on common questions but produces fabricated regulatory citations when asked about rare edge cases. The team has a curated set of correct answers for these edge cases and wants the model to adopt that behavior reliably. Which approach best fits?

A.Apply supervised fine-tuning on the curated edge-case examples to teach the desired response pattern.
B.Lower the topP value so the model only considers the most likely tokens and avoids fabricating citations.
C.Add the curated answers to the system instruction and rely on the model to generalize.
D.Increase the model's context window by switching to a long-context variant and pasting the full regulations.
AnswerA

Supervised fine-tuning is designed to teach a model a specific input-to-output behavior using labeled examples. With a curated set of correct edge-case answers, tuning adjusts the model so it reproduces the desired citation style and content, which is more reliable than describing the behavior in a prompt for rare cases.

Why this answer

Supervised fine-tuning uses the curated input-output pairs to adjust the model's weights so it reliably reproduces correct edge-case answers, which is the intended use of labeled examples. Prompt stuffing, sampling controls, and larger context windows do not durably change behavior for rare inputs.

Exam trap

The trap here is believing that narrowing sampling with topP or enlarging the context window will fix factual errors, when those knobs do not teach the model new domain answers.

106
MCQmedium

A financial analyst is using a large language model to generate executive summaries from lengthy earnings call transcripts. The summaries often miss key financial figures and include irrelevant details. Which technique should be used to improve the relevance and accuracy of the summaries?

A.Use few-shot prompting with examples of well-structured summaries.
B.Increase the maximum output token limit.
C.Fine-tune the model on a large corpus of general text.
D.Increase the model's temperature setting.
AnswerA

Few-shot prompting provides the model with examples of desired output, guiding it to focus on relevant information and format. By showing examples that highlight key financial figures and omit irrelevant details, the model can learn to replicate that behavior. This technique is effective for improving relevance and accuracy in summarization tasks.

Why this answer

Few-shot prompting is a powerful technique to guide the model's output by providing examples. In this scenario, showing the model examples of summaries that include key financial figures and exclude irrelevant details helps it learn the desired pattern, thereby improving relevance and accuracy.

Exam trap

The trap here is assuming that increasing token limits or temperature will improve summarization, when actually they can degrade quality.

107
MCQhard

An AI team is building a customer support chatbot for a telecom company using a fine-tuned LLM on Vertex AI. The model performs well on common issues but fails to answer correctly for rare or novel problems, often providing plausible-sounding but incorrect solutions. The team has a large corpus of internal troubleshooting documents. They want to minimize incorrect answers while keeping latency low. Which approach should they take?

A.Switch to a larger base model (e.g., Gemini Ultra) without any retrieval.
B.Implement a retrieval-augmented generation (RAG) pipeline using Vertex AI Search to fetch relevant documents before generating answers.
C.Collect more data on rare issues and continue fine-tuning the model weekly.
D.Use a few-shot prompt with 10 examples of rare problems and solutions.
AnswerB

RAG grounds generation in retrieved internal troubleshooting documents, so rare or novel queries draw on authoritative content rather than parametric guesses. Vertex AI Search supplies relevant passages at inference time, reducing hallucinated answers while keeping latency low.

Why this answer

Implementing a RAG pipeline with Vertex AI Search allows the chatbot to retrieve relevant troubleshooting documents from the internal corpus in real-time, grounding the LLM's responses in authoritative sources. This approach directly addresses the problem of plausible-sounding but incorrect answers for rare/novel issues without requiring retraining, and it keeps latency low by fetching only the most relevant documents before generation.

Exam trap

Google often tests the misconception that fine-tuning or larger models alone can solve knowledge gaps, when in fact retrieval-augmented generation is the standard approach for grounding LLM outputs in up-to-date, domain-specific documents without retraining.

How to eliminate wrong answers

Option A is wrong because switching to a larger base model without retrieval does not solve the core issue of hallucination on rare/novel problems; larger models can still generate plausible-sounding but incorrect answers when they lack specific knowledge, and they often increase latency and cost. Option C is wrong because collecting more data on rare issues and fine-tuning weekly is resource-intensive, may lead to catastrophic forgetting of common issues, and cannot keep pace with the long tail of novel problems that emerge dynamically. Option D is wrong because a few-shot prompt with 10 examples is insufficient to cover the vast space of rare problems, and the model may still hallucinate when the input does not closely match any example, especially without retrieval grounding.

108
Multi-Selecteasy

A company is prompt engineering a model for customer support. They want to reduce hallucination (false information) in responses. Which TWO techniques are most effective? (Choose two.)

Select 2 answers
A.Implement RAG to retrieve relevant documents for context
B.Provide 3 few-shot examples of conversations
C.Reduce max output tokens to 150
D.Add a system instruction: 'Only answer based on the provided context.'
E.Increase temperature to 1.2
AnswersA, D

RAG provides factual grounding, reducing hallucination.

Why this answer

Retrieval-Augmented Generation (RAG) grounds the model's output in external, verifiable documents retrieved from a knowledge base. By providing relevant context at inference time, RAG significantly reduces the likelihood of the model fabricating information, as it can reference and paraphrase from the retrieved sources rather than relying solely on its parametric memory.

Exam trap

A common mistake in this exam is to think that adjusting parameters like temperature or output tokens directly reduces hallucination, when in fact only techniques that constrain the model's knowledge source (like RAG and strict system instructions) are effective.

109
Multi-Selectmedium

A development team is integrating a large language model into a healthcare application. They need to reduce the risk of generating harmful medical advice. Which THREE measures should they implement? (Choose three.)

Select 3 answers
A.Use a safety filter to block outputs containing harmful medical terminology.
B.Implement RAG to retrieve verified medical information from trusted sources.
C.Fine-tune the model on a curated dataset of medical textbooks.
D.Include a disclaimer in the system instruction that the model is not a doctor.
E.Set the temperature to a very high value to ensure diverse outputs.
AnswersA, B, C

Safety filters directly block harmful content at inference time.

Why this answer

Implementing a safety filter that blocks outputs containing harmful medical terminology directly mitigates the risk of generating dangerous advice. This acts as a post-processing guardrail, intercepting model outputs that include terms associated with diagnoses, dosages, or procedures that could lead to patient harm. It is a standard practice in high-stakes domains to layer such filters on top of the generative model.

Exam trap

The Generative AI Leader exam often tests the misconception that disclaimers or system instructions alone are sufficient safety measures, when in fact they do not technically prevent the model from generating harmful content—only post-hoc filtering or architectural controls like RAG and fine-tuning can reduce the risk at the output level.

110
MCQhard

A healthcare organization needs a generative AI model to answer medical questions using proprietary clinical guidelines. They have a large dataset of doctor-patient interactions. Should they fine-tune a pre-trained model or use Retrieval-Augmented Generation (RAG)?

A.Use RAG to reduce inference costs by skipping model updates.
B.Use RAG to retrieve relevant guidelines during inference, avoiding frequent retraining.
C.Use prompt engineering to encode all guidelines into the system prompt.
D.Fine-tune the model on the clinical guidelines and interactions.
AnswerB

RAG retrieves the relevant clinical guideline passages at inference time and supplies them as context, so answers stay grounded in proprietary content without retraining. This satisfies the need to reflect frequently updated guidelines, since the index can be refreshed without touching model weights.

Why this answer

RAG is preferred because it retrieves the most current clinical guidelines from an external knowledge base during inference, avoiding the need for constant retraining when guidelines change. This is especially important in healthcare where regulations update frequently. Fine-tuning a pre-trained model on proprietary interactions might lead to overfitting or outdated knowledge.

Option A is incorrect because RAG does not necessarily reduce inference costs; it adds overhead for retrieval. Option C is incorrect because prompt engineering cannot encode all proprietary guidelines; it is limited by context window size. Option D is incorrect because fine-tuning requires retraining when guidelines change, which is less flexible than RAG.

111
MCQhard

A healthcare organization is using a generative AI model to summarize patient discharge instructions. They need to ensure the summaries are accurate and do not omit critical information. Which technique should they implement to reduce the risk of omissions?

A.Increase the model's token limit to allow longer summaries that can include more details.
B.Fine-tune the model on a large dataset of medical summaries to improve its ability to capture key points.
C.Implement a retrieval-augmented generation (RAG) system that retrieves relevant sections of the discharge instructions and includes them in the prompt.
D.Use chain-of-thought prompting to encourage the model to reason step by step.
AnswerC

RAG ensures the model has access to the actual discharge instructions. By retrieving and including relevant sections, the model is less likely to omit critical information because it is grounded in the source text. This directly addresses the risk of omissions by providing the necessary context.

Why this answer

Retrieval-augmented generation (RAG) is the best choice because it grounds the model in the actual discharge instructions, reducing the chance of omissions. By retrieving and including relevant text in the prompt, the model can generate summaries that are comprehensive and accurate. Other techniques like fine-tuning or chain-of-thought do not directly ensure that all critical information from the source is captured.

Exam trap

The trap here is thinking that a more powerful model or longer output will automatically capture all critical details, when the real issue is ensuring the model references the source document.

112
MCQhard

A healthcare startup fine-tunes a model to generate patient education materials. They want to ensure the model never gives medical advice, only information. They add a safety instruction, but the model sometimes still gives advice. What advanced technique should they apply?

A.Hard-code a list of prohibited phrases in a post-processing script
B.Add a secondary classifier to rewrite any detected advice into general information
C.Use semantic similarity to a 'medical advice' embedding and reject if close
D.Apply RLHF with a reward model that penalizes outputs containing medical advice
AnswerD

RLHF trains a reward model that scores outputs, then optimises the model against it; penalising medical-advice content directly shapes generation away from advice. A single safety instruction only conditions the prompt, so it cannot reliably enforce the constraint across varied inputs.

Why this answer

RLHF (Reinforcement Learning from Human Feedback) directly addresses the model's behavior by training a reward model that penalizes outputs containing medical advice. This aligns the model's generation with the safety instruction at a fundamental level, rather than relying on brittle post-hoc filters or static embeddings that can be easily circumvented by novel phrasings.

Exam trap

Candidates often mistakenly believe that simple post-processing filters or static embedding comparisons are sufficient to enforce safety. However, only advanced alignment techniques like RLHF can truly align the model's generation, as it changes the model's behavior during training rather than applying brittle surface-level checks.

How to eliminate wrong answers

Option A is wrong because hard-coding a list of prohibited phrases is brittle and fails against adversarial or paraphrased advice that doesn't match the exact phrases. Option B is wrong because adding a secondary classifier to rewrite detected advice introduces latency, potential for semantic drift, and cannot handle nuanced contexts where advice is implied rather than explicit. Option C is wrong because semantic similarity to a static 'medical advice' embedding is threshold-dependent and can produce false positives (flagging general information) or false negatives (missing advice phrased differently), and it does not train the model to avoid the behavior.

113
MCQeasy

Which technique allows a model to incorporate real-time data from external APIs?

A.RAG with tool calling
B.Prompt engineering
C.Fine-tuning
D.Model pruning
AnswerA

RAG with tool calling lets the model invoke external APIs at inference time, retrieving live data rather than relying solely on static training weights. This satisfies the stem's real-time external API constraint, which plain retrieval-augmented generation alone cannot meet.

Why this answer

RAG with tool calling is correct because it enables a generative AI model to query external APIs in real-time, retrieve up-to-date information, and incorporate that data into its response. This technique combines retrieval-augmented generation (RAG) with function calling, where the model outputs a structured request (e.g., a JSON object) to invoke an API, receive the result, and then generate a context-aware answer. Unlike static methods, this allows dynamic data integration without retraining.

Exam trap

A common pitfall is thinking that prompt engineering alone can achieve real-time data integration with external APIs. However, in Google Cloud's generative AI context, only RAG with tool calling (function calling) provides the explicit mechanism to execute API calls and incorporate live results into model responses.

How to eliminate wrong answers

Option B is wrong because prompt engineering only modifies the input text to guide model behavior, but it cannot fetch live data from external sources—it relies solely on the model's pre-existing knowledge. Option C is wrong because fine-tuning updates the model's weights on a fixed dataset, which does not enable real-time API access; it only improves performance on static tasks. Option D is wrong because model pruning reduces model size by removing redundant weights, which has no mechanism for external data retrieval or API interaction.

114
MCQmedium

A team is deploying a large language model for legal document summarization. They find the model occasionally omits critical legal clauses. Which improvement technique would be most effective?

A.Design a prompt that explicitly lists required sections
B.Increase the top_p value to 1.0
C.Fine-tune the model on legal summaries
D.Lower the temperature to 0.1
AnswerA

Explicitly listing the required sections in the prompt constrains the model's output structure, directing attention to mandatory clauses that free-form summarisation tends to drop. This prompt-engineering approach improves recall of critical content without retraining or architectural changes.

Why this answer

The most effective technique is to design a prompt that explicitly lists the required sections. This is a form of prompt engineering that directly addresses the omission issue by instructing the model to include all specified clauses, leveraging the model's instruction-following capability. It is immediate, low-cost, and does not require retraining or altering sampling parameters.

By enumerating the sections, the model is guided to cover each one, reducing the chance of missing critical content.

Exam trap

The trap here is confusing sampling parameters (temperature, top_p) with techniques that ensure completeness; candidates might think lowering temperature or increasing top_p would make the model more reliable, but they only affect randomness, not coverage of required content.

How to eliminate wrong answers

Option B is wrong because increasing top_p to 1.0 makes the sampling more diverse (nucleus sampling considers all tokens), which can increase randomness and does not ensure completeness; it may even worsen omissions. Option C is wrong because fine-tuning on legal summaries could improve domain adaptation but is expensive, time-consuming, and does not guarantee that all required sections will be included; it may also overfit to the training summaries' style rather than ensuring clause coverage. Option D is wrong because lowering temperature to 0.1 makes the output more deterministic and focused on high-probability tokens, but it does not enforce inclusion of specific sections; it can actually reduce creativity and may cause the model to stick to common patterns, potentially omitting less frequent but critical clauses.

115
MCQmedium

A developer is building a customer support chatbot using a large language model. The chatbot frequently generates plausible-sounding but incorrect answers to product questions. Which technique should be applied to improve factual accuracy?

A.Provide a few-shot example of correct answers in the prompt.
B.Use a higher temperature setting to encourage more creative responses.
C.Increase the model's context length to include more of the conversation history.
D.Enable Grounding with the company's product knowledge base.
AnswerD

Grounding retrieves authoritative product content and injects it into the prompt context, so the model conditions its response on real documentation rather than parametric memory alone. This directly addresses the hallucination constraint in the stem, replacing plausible fabrication with verifiable, source-anchored answers drawn from the company knowledge base.

Why this answer

Grounding with the company's product knowledge base (D) is the most effective technique to improve factual accuracy. Grounding involves retrieving relevant information from a trusted knowledge source and providing it to the model as context, so the model generates answers based on that data rather than relying solely on its training. This reduces hallucinations and ensures responses are accurate and up-to-date.

Exam trap

Generative AI Leader often tests the difference between prompt engineering techniques and grounding, causing candidates to choose few-shot examples or context length increases when the core issue is lack of factual knowledge.

How to eliminate wrong answers

Option A is wrong because few-shot examples can guide the model's style but do not provide the factual knowledge needed to answer product-specific questions accurately. Option B is wrong because a higher temperature increases randomness and creativity, which would likely worsen factual accuracy. Option C is wrong because increasing context length allows more conversation history but does not inject external factual knowledge; it may even introduce irrelevant information.

116
MCQmedium

A team monitors their generative AI model on Vertex AI. They notice output quality declining. Which metric is most likely the root cause?

A.Input token count per request is increasing.
B.Output token count is decreasing.
C.Prediction latency is stable.
D.Error rate is less than 1%.
AnswerA

Rising input token counts lengthen prompts, pushing the model beyond its effective context and diluting instruction focus, which degrades output quality. This metric directly explains the decline, unlike latency or cost signals that do not affect generation fidelity.

Why this answer

A is correct because an increasing input token count per request can degrade output quality by diluting the model's attention across a longer context window. In transformer-based models like those on Vertex AI, the attention mechanism has a fixed capacity; as input tokens grow, the model may lose focus on critical information, leading to less coherent or relevant outputs. This is a common issue in production systems where users gradually add more context without trimming irrelevant tokens.

Exam trap

A common misconception is that output quality issues are always due to model errors or latency problems, rather than subtle input-side factors like token count inflation that silently degrade attention focus. This question highlights that root cause can be on the input side.

How to eliminate wrong answers

Option B is wrong because a decreasing output token count does not inherently cause quality decline; it may indicate shorter responses, but quality can remain high if the model is well-tuned. Option C is wrong because stable prediction latency suggests consistent infrastructure performance, not a root cause of output quality degradation. Option D is wrong because a low error rate (<1%) indicates the model is responding without failures, but output quality can still suffer from issues like hallucination or incoherence even when error rates are minimal.

117
MCQmedium

A content generation model for e-commerce product descriptions repeats the same phrases across multiple descriptions (e.g., 'high-quality', 'best-in-class'). The team wants more varied and engaging output. Which parameter adjustment is most appropriate?

A.Increase the frequency penalty parameter to 1.0.
B.Decrease the max output tokens to 50.
C.Increase the temperature parameter to 1.5.
D.Set the top-p value to a very small number like 0.1.
AnswerA

Raising frequency penalty to 1.0 penalises tokens proportionally to how often they have already appeared in the generated text, directly discouraging the repeated phrases the stem describes. This satisfies the requirement for varied, engaging product descriptions without altering the model's underlying knowledge or the prompt itself.

Why this answer

Increasing the frequency penalty to 1.0 penalizes tokens that have already appeared in the generated text, directly reducing repetition of phrases like 'high-quality' and 'best-in-class'. This encourages the model to use more diverse vocabulary and sentence structures, leading to varied and engaging product descriptions.

Exam trap

The Generative AI Leader exam often tests the distinction between frequency penalty and temperature, where candidates mistakenly increase temperature to add variety, not realizing that temperature increases randomness and can break coherence, while frequency penalty directly targets repetition without sacrificing quality.

How to eliminate wrong answers

Option B is wrong because decreasing max output tokens to 50 limits the length of each description but does not address the root cause of phrase repetition; the model can still repeat phrases within the shorter output. Option C is wrong because increasing temperature to 1.5 makes the output more random and less coherent, which can lead to nonsensical descriptions rather than controlled variation. Option D is wrong because setting top-p to a very small number like 0.1 restricts the model to only the most likely tokens, which actually increases repetition and reduces diversity, the opposite of the desired outcome.

118
MCQmedium

A product team at a retail company is using a foundation model on Vertex AI to generate short marketing taglines for new products. They find the outputs are often too long and sometimes include extra commentary. They want to constrain the model to produce only a single concise tagline. Which parameter should they adjust?

A.Temperature
B.Top-p
C.Max output tokens
D.Top-k
AnswerC

Max output tokens sets a hard limit on the number of tokens the model can generate in its response. By setting this to a small value that accommodates a single tagline, the team can prevent lengthy outputs and force the model to be concise. It directly addresses the length issue described.

Why this answer

The max output tokens parameter caps the total number of tokens generated, which directly enforces a limit on response length. By setting it appropriately, the team can ensure the model produces only a short tagline and cannot continue with extra commentary. Other parameters like temperature, top-k, and top-p affect randomness and diversity, not length.

Exam trap

The trap here is confusing parameters that control randomness with those that control output length.

119
MCQmedium

A software development team builds an internal code assistant using a generative model. The assistant writes Python functions that often contain security vulnerabilities such as SQL injection or command injection. The team wants to mitigate these vulnerabilities without adding a manual review step for every code snippet, as that would slow development. They have access to a static analysis security scanner API. Which approach best addresses the vulnerabilities while maintaining developer velocity?

A.Increase top-k sampling to generate a wider variety of code tokens.
B.After each generation, automatically run the code through the static analysis scanner, and if vulnerabilities are found, send the output back to the model for revision with the scanner's feedback.
C.Fine-tune the model on a corpus of secure code examples.
D.Add a system prompt: 'Do not generate code with security vulnerabilities.'
AnswerB

Feeding scanner findings back into the model creates an automated generate-scan-revise loop, so vulnerable patterns are corrected before the developer sees the snippet. This satisfies the no-manual-review constraint while preserving velocity, unlike prompt-only hardening, which cannot verify the emitted code.

Why this answer

It creates an automated feedback loop: the static analysis scanner detects vulnerabilities in the generated code, and the model revises the output based on that feedback. This approach directly mitigates security flaws without requiring manual review, preserving developer velocity. It leverages the scanner's precise, rule-based detection to iteratively improve the model's output, which is more reliable than relying on the model's inherent safety.

Exam trap

This exam often tests the misconception that a simple prompt or fine-tuning alone can guarantee safety, when in reality, a closed-loop validation with a dedicated security tool is required for reliable mitigation of injection vulnerabilities.

How to eliminate wrong answers

Option A is wrong because increasing top-k sampling broadens token selection, which can actually introduce more unpredictable and insecure code patterns, not reduce vulnerabilities. Option C is wrong because fine-tuning on secure code examples improves the model's baseline but does not guarantee that every generated snippet will be free of vulnerabilities, especially for novel or context-specific injection attacks. Option D is wrong because a system prompt is a weak, non-enforceable instruction; the model lacks true understanding of security and can easily generate vulnerable code despite the prompt, as it does not perform actual validation.

120
MCQeasy

A developer uses a generative AI model with the system instruction shown. The response is correct but very brief. Which parameter adjustment could encourage more detail without losing accuracy?

A.Add 'Provide a detailed response' to the system instruction.
B.Set temperature to 0 to make output deterministic.
C.Set topK to 1 to focus on most likely tokens.
D.Increase temperature to 1.5 to encourage creativity.
AnswerA

Adding an explicit instruction for detail to the system instruction steers the model's generation toward longer, more thorough responses while the underlying task and data remain unchanged, preserving accuracy. This directly addresses the brevity without altering model parameters.

Why this answer

Modifying the system instruction to explicitly request a detailed response directly influences the model's output behavior without altering its underlying probability distribution. This approach preserves accuracy by keeping temperature, topK, and other sampling parameters at their default values, ensuring the model remains faithful to the training data while simply prompting for more elaboration.

Exam trap

The Google Gen AI Leader exam often tests the misconception that increasing randomness (temperature) or restricting token selection (topK) can improve detail, when in fact these parameters trade off accuracy for diversity or determinism, and the correct approach is to use prompt engineering to guide output length and style.

How to eliminate wrong answers

Option B is wrong because setting temperature to 0 makes the model deterministic by always selecting the highest-probability token, which typically results in shorter, repetitive, and less detailed responses—the opposite of what is needed. Option C is wrong because setting topK to 1 restricts token selection to only the single most likely token at each step, which similarly reduces output diversity and detail, often leading to generic or truncated answers. Option D is wrong because increasing temperature to 1.5 increases randomness in token sampling, which can introduce hallucinations, factual errors, or irrelevant content, thereby sacrificing accuracy for creativity.

121
MCQhard

A generative AI model for chatbot responses sometimes produces toxic language. The team wants to reduce toxicity without significantly affecting the model's helpfulness. Which approach is best?

A.Increase the temperature parameter
B.Reduce the maximum output tokens
C.Fine-tune with a dataset of non-toxic responses and use RLHF
D.Apply a toxicity classifier as a post-processing filter
AnswerC

Supervised fine-tuning on non-toxic responses teaches safer output distributions, and RLHF then optimises the policy against a reward model balancing toxicity reduction with helpfulness. This directly targets the constraint of lowering toxicity without materially degrading response quality.

Why this answer

Fine-tuning with a curated dataset of non-toxic responses directly adjusts the model's weights to reduce the likelihood of generating toxic language, while RLHF (Reinforcement Learning from Human Feedback) further aligns the model with human preferences for helpfulness and safety. This combined approach addresses the root cause of toxicity in the model's behavior without the blunt trade-offs of other methods, preserving the model's utility.

Exam trap

Google Cloud often tests the misconception that post-processing filters (like toxicity classifiers) are sufficient for safety, when in fact they fail to address the model's learned behavior and can degrade helpfulness due to false positives, making fine-tuning with RLHF the superior alignment technique.

How to eliminate wrong answers

Option A is wrong because increasing the temperature parameter increases randomness in token selection, which can actually amplify the probability of generating toxic or nonsensical outputs, not reduce them. Option B is wrong because reducing the maximum output tokens limits response length but does not influence the content or safety of the generated tokens, leaving toxicity unchanged. Option D is wrong because applying a toxicity classifier as a post-processing filter only masks toxic outputs after generation, wasting computational resources and potentially blocking helpful responses that contain false-positive flagged terms, without fixing the underlying model behavior.

122
Multi-Selectmedium

A logistics company uses a generative AI model to draft incident reports from sensor logs. Reviewers find the reports are often incomplete, missing fields such as root cause and corrective action. The team wants to improve output completeness without retraining the model. Which TWO techniques should they use? (Choose two.)

Select 2 answers
A.Shorten the input sensor logs to only the most recent readings.
B.Add few-shot examples of complete incident reports that show every required field filled in.
C.Increase the model's temperature to encourage the model to explore more content.
D.Reduce the maximum output tokens to force the model to be concise.
E.Define a structured output schema listing every required field and instruct the model to populate each one.
AnswersB, E

Complete exemplars demonstrate the expected coverage pattern, teaching the model which sections must appear and how detailed they should be. Combined with a schema, examples reinforce the habit of filling all fields. This is a prompt-level change that improves completeness immediately without weight updates or new training data collection.

Why this answer

Completeness improves when the required output shape is made explicit and demonstrated. A structured schema enumerates every field so omissions become visible and correctable, while few-shot examples of complete reports show the model the expected coverage and detail. Sampling and length controls do not enforce required fields and can reduce or distort content.

Exam trap

The trap here is assuming that more creative sampling or longer/shorter outputs improve report quality, when completeness specifically requires an explicit field schema plus examples of fully populated reports.

123
MCQhard

A financial services firm uses a Gemini model to generate quarterly risk summaries from internal reports. Reviewers note that summaries sometimes contradict the source tables. The team wants the model to reason step by step over the figures before writing the summary. Which technique should they use?

A.Enable a higher top-p value so the model samples more broadly and can reconcile conflicting numbers.
B.Set temperature to zero so the model always produces the same summary and cannot contradict the tables.
C.Use chain-of-thought prompting by instructing the model to compute and list intermediate values before producing the final summary.
D.Lower the maximum output tokens so the model is forced to be concise and avoid contradictory statements.
AnswerC

Chain-of-thought prompting directs the model to work through intermediate calculations and comparisons before drafting the summary, which reduces contradictions because the figures are reconciled first. It makes the reasoning inspectable, so reviewers can spot errors. This technique is well suited to tasks that require multi-step arithmetic and consistency checking.

Why this answer

Chain-of-thought prompting makes the model compute and compare intermediate values before writing, which is exactly what is needed to keep a summary consistent with source tables. Sampling changes and output limits affect randomness or length but not numerical reasoning. When a task requires multi-step reconciliation, prompting for explicit steps is the appropriate technique.

Exam trap

The trap here is treating determinism or brevity as a cure for factual inconsistency, when neither supplies the step-by-step reconciliation the task requires.

124
Multi-Selecthard

Which THREE are best practices for designing prompts for a generative AI model?

Select 3 answers
A.Provide few-shot examples for complex tasks
B.Include specific and clear instructions
C.Break the task into smaller steps
D.Use negative prompts to avoid undesired outputs
E.Always set temperature to 1.0 for creativity
AnswersA, B, C

Correct: Examples guide the model toward desired outputs.

Why this answer

Providing few-shot examples (e.g., 2-5 input-output pairs) helps the model infer the desired pattern, reducing ambiguity for complex tasks like classification or structured extraction. This technique leverages in-context learning, where the model uses the examples as a template without fine-tuning.

Exam trap

A common misconception tested in Google's Gen AI evaluations is that negative prompts can reliably control outputs, but they often fail due to tokenization and probability smoothing, leading to the 'forbidden token' problem where undesired content still appears.

125
Multi-Selectmedium

Which TWO techniques can help improve the factual accuracy of a language model's outputs? (Choose two.)

Select 2 answers
A.Decrease the max output tokens.
B.Increase the temperature parameter.
C.Fine-tune on a domain-specific curated dataset.
D.Implement retrieval-augmented generation (RAG).
E.Use top-k random sampling.
AnswersC, D

Fine-tuning adapts the model to domain facts.

Why this answer

Fine-tuning on a domain-specific curated dataset (C) directly adjusts the model's weights using high-quality, verified examples, teaching it to produce factually correct outputs for that domain. This reduces hallucinations by grounding the model in accurate, relevant data rather than relying solely on its pre-training distribution.

Exam trap

Google Cloud often tests the misconception that adjusting decoding parameters (like temperature, top-k, or max tokens) can improve factual accuracy, when in reality these only control output style, length, or randomness, not the correctness of the underlying information.

126
MCQmedium

A software company is using a large language model to generate code snippets from natural language descriptions. The generated code often has syntax errors and does not follow the company's coding standards. Which approach is most effective to improve the quality of the generated code?

A.Fine-tune the model on a dataset of code that adheres to the company's coding standards.
B.Add a prompt instruction to 'write clean code'.
C.Increase the temperature to generate more diverse code solutions.
D.Use a larger model with more parameters without fine-tuning.
AnswerA

Fine-tuning on a dataset of code that follows the company's standards will teach the model the specific syntax, style, and patterns used. This directly addresses both syntax errors and adherence to coding standards, making it the most effective approach for this scenario.

Why this answer

Fine-tuning on a dataset of code that adheres to the company's coding standards is the most effective way to improve the quality of generated code. It directly teaches the model the desired syntax, style, and patterns, reducing syntax errors and ensuring compliance with standards.

Exam trap

The trap here is relying on prompt engineering alone to enforce coding standards, when fine-tuning provides a more reliable solution.

127
MCQeasy

A data scientist is using a large language model to generate product descriptions. The descriptions are often too verbose. Which parameter adjustment is most appropriate?

A.Decrease the top-k value.
B.Increase the max output tokens.
C.Decrease the temperature.
D.Increase the frequency penalty.
AnswerD

Frequency penalty reduces repetitive phrases, encouraging conciseness.

Why this answer

Increasing the frequency penalty reduces the likelihood of the model repeating the same phrases or ideas, which directly addresses verbosity by discouraging repetitive or overly detailed descriptions. This parameter penalizes tokens that have already appeared in the generated text, promoting more concise and varied output. Other adjustments like temperature or top-k affect randomness and diversity but do not specifically target repetition or length.

Exam trap

Google Cloud often tests the distinction between parameters that control randomness (temperature, top-k) versus those that control repetition (frequency penalty, presence penalty), and the trap here is that candidates confuse 'less verbose' with 'less random' and incorrectly choose temperature or top-k adjustments.

How to eliminate wrong answers

Option A is wrong because decreasing the top-k value restricts the model to a smaller set of high-probability tokens, which can actually make output more predictable and potentially more repetitive, not less verbose. Option B is wrong because increasing the max output tokens allows the model to generate longer text, which would exacerbate verbosity rather than reduce it. Option C is wrong because decreasing the temperature makes the model more deterministic and conservative, often leading to safer but not necessarily shorter or less repetitive text; it does not directly penalize repetition or length.

128
MCQhard

A model generates responses that frequently repeat phrases or words. Which parameter adjustment is most likely to fix this?

A.Increase top_k
B.Increase temperature
C.Increase repetition penalty
D.Increase max output tokens
AnswerC

Repetition penalty directly down-weights tokens already emitted, so previously generated words become less probable at each decoding step. Raising it breaks the loop of recurring phrases, satisfying the stem's requirement to stop repeated words without altering the model itself.

Why this answer

Increasing the repetition penalty directly discourages the model from selecting tokens that have already appeared in the generated sequence, thereby reducing repetitive phrases or words. This parameter works by subtracting a fixed penalty from the logits of previously generated tokens before applying the softmax function, making them less likely to be chosen again.

Exam trap

The trap here is that candidates often confuse repetition penalty with diversity-promoting parameters like temperature or top_k, mistakenly believing that increasing randomness or narrowing token selection will fix repetition, when in fact those adjustments can worsen the problem.

How to eliminate wrong answers

Option A is wrong because increasing top_k limits the sampling pool to the k most likely next tokens, which can actually increase repetition by narrowing the diversity of choices. Option B is wrong because increasing temperature flattens the probability distribution, making all tokens more equally likely, which can lead to more random and potentially more repetitive outputs, not less. Option D is wrong because increasing max output tokens only extends the length of the generated response; it does not address the underlying cause of repetition and may even exacerbate it by allowing more opportunities for the model to loop on repeated phrases.

129
Multi-Selecthard

A financial services firm is using a large language model to generate quarterly investment summaries from raw market data. The summaries occasionally contain fabricated statistics and sometimes omit key risk factors. The team wants to improve factual accuracy and completeness without retraining the model. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Increasing the temperature setting to encourage more creative wording.
B.Fine-tuning the model on a small set of historical summaries to improve style.
C.Retrieval-augmented generation (RAG) using a curated knowledge base of verified financial reports.
D.Implementing a grounding check with a fact-verification API that cross-references generated claims against trusted sources.
E.Using a chain-of-thought prompt that asks the model to reason step by step before writing the summary.
AnswersC, D

RAG grounds the model's output in retrieved, authoritative documents. By fetching relevant passages from verified reports, the model can cite accurate statistics and include risk factors present in those sources. This reduces hallucination and improves completeness because the model conditions on real data rather than relying solely on its parametric memory.

Why this answer

Retrieval-augmented generation supplies the model with verified, up-to-date information, reducing fabrication and improving coverage of key points. A fact-verification API adds a validation layer that catches any remaining inaccuracies. Together, they address both hallucination and omission without retraining the model.

Exam trap

The trap here is assuming that prompt engineering alone, such as chain-of-thought, can eliminate factual errors without external grounding.

130
MCQmedium

A marketing team is using a generative AI model on Vertex AI to create ad copy for a new product launch. The initial outputs are generic and do not reflect the brand's tone. The team wants to quickly improve the outputs without retraining the model. They have a set of example ad copies that exemplify the desired tone. Which technique should they use?

A.Increase the model's temperature setting to encourage more creative outputs.
B.Deploy the model to a new endpoint with higher throughput.
C.Fine-tune the model on the example ad copies.
D.Use few-shot prompting by including the example ad copies in the prompt.
AnswerD

Few-shot prompting involves providing a few examples of the desired output format and style directly in the prompt. This guides the model to generate text that matches the brand's tone without any model retraining, making it a fast and effective solution for this scenario.

Why this answer

Few-shot prompting is ideal when you have example outputs that demonstrate the desired style or format. By including these examples in the prompt, the model can infer the pattern and generate new content that aligns with the brand's tone. This approach requires no training and can be implemented immediately, making it the most efficient solution for the marketing team.

Exam trap

The trap here is assuming that any model improvement requires fine-tuning, when in-context learning via few-shot prompting can often achieve the desired result faster and with less effort.

131
MCQhard

A financial services firm uses a generative AI model to draft client emails. The drafts are accurate but sometimes use an overly casual tone. The firm wants to enforce a consistently formal tone across all drafts. Which approach is most reliable for this requirement?

A.Fine-tune the model on a large corpus of casual emails to broaden its style range.
B.Add a post-processing step that replaces casual words with formal synonyms using a fixed dictionary.
C.Set the temperature to a high value so the model produces more varied language.
D.Provide a system instruction that defines the desired formal tone and include a short example of an approved email.
AnswerD

A system instruction sets persistent behavioral guidance for every request, and a concrete example anchors the model to the desired style. Together they give the model both a rule and a reference, which is more reliable than a one-off user prompt. This approach directly targets tone consistency without retraining. It also scales across many drafts because the instruction applies to all interactions.

Why this answer

System instructions provide persistent behavioral guidance, and an approved example shows the model the exact tone desired. This combination is more reliable than sampling changes, casual fine-tuning, or brittle post-processing. It directly enforces formality across all drafts and scales without retraining, making it the best fit for the firm's consistency requirement.

Exam trap

The trap here is assuming that tone can be enforced by a fixed synonym replacement or by increasing randomness, when tone is best controlled by persistent instructions and examples.

132
MCQmedium

A company uses Vertex AI PaLM for code generation. The code often contains security vulnerabilities. Which improvement should be applied?

A.Set top_k to 1
B.Include a security-focused system instruction
C.Use Codey model instead
D.Increase temperature to 0.8
AnswerB

A security-focused system instruction sets persistent behavioural guardrails, steering the model to avoid insecure patterns such as unsanitised input handling across every generation. This directly targets the vulnerability constraint; prompt-level or post-hoc scanning alone cannot shape generation as reliably.

Why this answer

Including a security-focused system instruction directly guides the model to prioritize secure coding practices, such as input validation and proper error handling, reducing vulnerabilities. This leverages prompt engineering to shape model behavior without altering parameters like temperature or top_k, which control randomness, not security awareness.

Exam trap

Google often tests the misconception that parameter tuning (like temperature or top_k) can fix content quality issues, when in fact prompt engineering—such as system instructions—is the primary tool for guiding model behavior toward specific goals like security.

How to eliminate wrong answers

Option A is wrong because setting top_k to 1 makes the model deterministic (always picks the highest-probability token), which can reduce output diversity but does not address security vulnerabilities—it may even amplify insecure patterns if they are common in training data. Option C is wrong because Codey is a specialized model for code generation, but it does not inherently include security guardrails; the same vulnerabilities can appear if the prompt lacks security context. Option D is wrong because increasing temperature to 0.8 increases randomness and creativity, which can introduce more unpredictable and potentially insecure code, worsening the vulnerability issue.

133
MCQmedium

A company uses a generative model to produce product descriptions. The descriptions are factually inconsistent with the product specs. Which technique would best ensure factual accuracy?

A.Enhance the system prompt with product details
B.Implement retrieval-augmented generation (RAG) with product database
C.Lower the temperature to 0.0
D.Fine-tune the model on product descriptions
AnswerB

RAG retrieves relevant product specifications from the database at inference time and injects them into the prompt, grounding the generated description in authoritative data. This directly addresses factual inconsistency, unlike fine-tuning or prompt wording changes alone.

Why this answer

Retrieval-augmented generation (RAG) is the best technique because it dynamically retrieves relevant, up-to-date product specifications from a trusted database at inference time, grounding the model's output in verified facts. This directly addresses factual inconsistency by ensuring the generated description is based on authoritative source data rather than relying solely on the model's parametric memory.

Exam trap

Google Cloud often tests the misconception that prompt engineering alone (Option A) or deterministic sampling (Option C) can solve factual grounding issues, when in reality they do not provide external knowledge retrieval to correct hallucinations.

How to eliminate wrong answers

Option A is wrong because enhancing the system prompt with product details only provides static context that the model may still hallucinate or misinterpret; it does not enforce retrieval of current or specific factual data. Option C is wrong because lowering the temperature to 0.0 makes the output more deterministic but does not prevent the model from generating factually incorrect content that is confidently wrong. Option D is wrong because fine-tuning on product descriptions can improve style and consistency but does not guarantee factual accuracy for new or updated product specs, and it risks overfitting or memorizing inaccuracies from the training data.

134
MCQeasy

A company uses a text generation model for customer support but notices it occasionally provides outdated information. Which technique should they implement to improve output accuracy?

A.Increase max output tokens
B.Implement retrieval-augmented generation (RAG)
C.Fine-tune the model with more historical support data
D.Increase model temperature to 1.0
AnswerB

RAG retrieves current information, making outputs accurate and up-to-date.

Why this answer

Retrieval-augmented generation (RAG) is the correct technique because it grounds the model's output in real-time, external knowledge sources (e.g., a vector database or document index) rather than relying solely on static training data. This directly addresses the problem of outdated information by allowing the model to retrieve and synthesize current facts at inference time, ensuring accuracy without requiring retraining.

Exam trap

The trap here is that candidates often confuse fine-tuning (which adapts the model's weights to a static dataset) with RAG (which dynamically retrieves external knowledge), leading them to choose fine-tuning as a 'deeper' fix when the core issue is stale information, not model capability.

How to eliminate wrong answers

Option A is wrong because increasing max output tokens only extends the length of the generated response, not its factual accuracy or timeliness; it may even introduce more hallucinated content. Option C is wrong because fine-tuning with more historical support data would reinforce outdated patterns and biases, making the model more likely to repeat stale information rather than adapt to current knowledge. Option D is wrong because increasing model temperature to 1.0 increases randomness and creativity in outputs, which degrades factual precision and reliability, the opposite of what is needed for accurate customer support.

135
MCQmedium

A team uses Vertex AI Generative AI Studio to tune a model via RLHF. After tuning, the model outputs are bland. What likely went wrong?

A.Insufficient training data
B.Too many training steps
C.Low temperature during evaluation
D.Reward model overfits to generic responses
AnswerD

RLHF rewards are learned from human preference data; if that reward model overfits to safe, generic completions, it assigns them inflated scores. Policy optimisation then maximises those scores, collapsing output diversity. The blandness stems from the reward signal, not the base model or sampling temperature.

Why this answer

When the reward model overfits to generic responses, it assigns high rewards to safe, non-committal outputs, causing the RLHF-tuned model to converge toward bland, uninformative text. This happens because the reward model learns to prefer patterns that are statistically common in the training data rather than genuinely high-quality or diverse responses, directly leading to the 'bland' output described.

Exam trap

Google often tests the misconception that bland outputs are caused by inference-time parameters like temperature, rather than by the reward model overfitting during the RLHF training phase.

How to eliminate wrong answers

Option A is wrong because insufficient training data typically causes underfitting or poor generalization, not specifically bland outputs; RLHF can still produce diverse responses if the reward model is well-calibrated. Option B is wrong because too many training steps usually lead to overfitting or reward hacking, where the model exploits the reward model for extreme or repetitive outputs, not blandness. Option C is wrong because low temperature during evaluation reduces randomness and can make outputs more deterministic, but it does not inherently cause blandness; the model would still produce coherent, contextually appropriate responses, just with less creativity.

136
MCQmedium

A chatbot built with Vertex AI PaLM API often provides outdated information about company policies because the training data is months old. Which approach should the team use?

A.Implement grounding by connecting to a knowledge base of current policies.
B.Use prompt engineering to instruct the model to say 'I don't know' if unsure.
C.Increase the context window to include more history.
D.Fine-tune the model on the latest policy documents.
AnswerA

Grounding retrieves current policy documents at inference time and passes them as context, so responses reflect the live knowledge base rather than the model's stale training data. This satisfies the requirement for up-to-date company policy information without retraining the PaLM model.

Why this answer

Grounding connects the PaLM API to a live, authoritative knowledge base (e.g., Cloud Storage, BigQuery, or Vertex AI Search) containing the latest company policies. This allows the model to retrieve and cite current information at inference time without retraining, directly solving the staleness issue. Grounding is the recommended approach in Vertex AI for ensuring factual, up-to-date responses from a foundation model.

Exam trap

In the Google Gen AI Leader exam, a common trap is confusing grounding (dynamic knowledge injection at inference time) with fine-tuning (static model update). Candidates often assume fine-tuning is the best solution for real-time accuracy, but grounding is the correct approach when policies change frequently.

How to eliminate wrong answers

Option B is wrong because prompt engineering to say 'I don't know' does not provide the model with current policy data; it only changes the model's refusal behavior, leaving outdated information uncorrected. Option C is wrong because increasing the context window does not introduce new or updated knowledge; it only allows the model to consider more of the conversation history, which does not address stale training data. Option D is wrong because fine-tuning on the latest policy documents would require significant time, cost, and labeled data, and the model would still be static until the next fine-tuning cycle; grounding provides a dynamic, real-time solution without retraining.

137
Multi-Selecthard

A company is using a generative AI model to create personalized email responses to customer inquiries. The responses sometimes contain factual errors or irrelevant information. The company wants to improve the accuracy and relevance of the responses. Which TWO techniques should they use? (Choose two.)

Select 2 answers
A.Use prompt engineering to provide clear instructions and examples of desired responses.
B.Implement retrieval-augmented generation (RAG) to ground responses in a knowledge base.
C.Fine-tune the model on a dataset of customer inquiries and responses.
D.Increase the model's temperature to make responses more creative.
E.Reduce the model's max output tokens to limit response length.
AnswersA, B

Prompt engineering with explicit instructions and few-shot examples can guide the model to generate more accurate and relevant responses. By specifying the desired format, tone, and content, and providing examples, the model can better align its outputs with the company's requirements, reducing errors and irrelevance.

Why this answer

Retrieval-augmented generation (RAG) grounds the model's responses in a verified knowledge base, ensuring factual accuracy and relevance. Prompt engineering with clear instructions and examples further guides the model to produce desired outputs. Together, these techniques directly target the issues of factual errors and irrelevance without the overhead of fine-tuning.

Exam trap

The trap here is assuming that any model adjustment will improve quality, when in fact techniques like increasing temperature or limiting tokens can worsen accuracy or fail to address the root causes.

138
MCQmedium

What is the primary purpose of a system instruction in the Gemini API?

A.Set the model's temperature and top_p
B.Define the overall behavior and constraints for the model
C.Provide few-shot examples for each query
D.Set the maximum output length
AnswerB

System instructions set persistent behavioural guidance and constraints applied across a Gemini API session, shaping tone, role and boundaries regardless of individual prompts. This satisfies the stem's requirement of defining overall model behaviour rather than specifying a single request's content.

Why this answer

The system instruction in the Gemini API is the primary mechanism to define the overall behavior, persona, constraints, and guardrails for the model across all interactions. Unlike per-query parameters, it sets a persistent context that shapes how the model interprets every user prompt, ensuring consistent adherence to rules such as tone, format, or safety policies.

Exam trap

Google Cloud often tests the distinction between persistent system-level instructions and per-request parameters, so the trap here is confusing the system instruction (which defines the model's role and constraints) with generation controls like temperature, top_p, or max tokens, which only affect the style or length of a single response.

How to eliminate wrong answers

Option A is wrong because temperature and top_p are sampling parameters that control randomness and diversity of output, not the overarching behavioral constraints set by a system instruction. Option C is wrong because few-shot examples are typically provided in the user prompt or as part of a structured conversation, not as the primary purpose of a system instruction, which is for persistent context rather than per-query demonstrations. Option D is wrong because maximum output length is a generation parameter that limits token count, not a behavioral or constraint-setting mechanism like a system instruction.

139
MCQmedium

A global bank uses a Gemini model on Vertex AI to generate personalized investment summaries for clients in multiple regions. Compliance requires that the model never recommend products prohibited in a given region. The team wants a control that enforces these rules regardless of how the prompt is phrased. Which approach should they use?

A.Configure safety filters and content moderation thresholds on the model endpoint.
B.Use a pre-call or post-call validation layer that checks the region against an allowed-product list and blocks violations.
C.Fine-tune the model on examples of compliant and non-compliant recommendations.
D.Add a system instruction listing prohibited products for each region.
AnswerB

A validation layer applies deterministic logic: before or after generation, it checks the client's region against an authoritative allowed-product list and rejects any response that violates it. Because the rule lives in code and data rather than in the prompt, it holds regardless of phrasing and can be updated as regulations change. This gives the auditable enforcement compliance demands.

Why this answer

Absolute compliance rules need deterministic enforcement rather than probabilistic guidance. A validation layer that evaluates the region against an allowed-product list before or after the model call blocks prohibited recommendations no matter how the request is phrased, and it produces an auditable record. System instructions, safety filters, and fine-tuning all influence behavior but cannot guarantee that a disallowed product is never surfaced.

Exam trap

The trap here is assuming that a strongly worded system instruction or safety filter is a compliance guarantee, when prompt-level guidance and safety categories cannot deterministically enforce region-specific product rules.

140
MCQhard

A company is deploying a generative AI model for customer support. They want to reduce hallucinations while maintaining fluency. They have a large dataset of previous support conversations. Which strategy should they prioritize?

A.Increase the beam search width to 10.
B.Implement retrieval-augmented generation (RAG) using the conversation dataset as a knowledge base.
C.Fine-tune the model on the conversation dataset.
D.Set the temperature to 0.1.
AnswerB

Retrieval-augmented generation grounds each response in passages retrieved from the conversation dataset, so answers reflect real support content rather than parametric guesses. This constrains fabrication while the underlying model preserves fluent phrasing, satisfying both the hallucination-reduction and fluency requirements.

Why this answer

Retrieval-augmented generation (RAG) directly addresses hallucinations by grounding the model's responses in factual, retrieved data from the conversation dataset. This approach allows the model to generate fluent, contextually relevant answers while reducing the risk of inventing information, as it retrieves actual support interactions as evidence before generating a response.

Exam trap

Google Cloud often tests the misconception that tuning generation parameters (like temperature or beam search) can fix hallucinations, when in fact only grounding techniques like RAG or knowledge graph integration address the root cause of factual inaccuracy.

How to eliminate wrong answers

Option A is wrong because increasing beam search width to 10 improves output fluency by exploring more candidate sequences but does not reduce hallucinations; it may even amplify incorrect patterns if the model is prone to hallucination. Option C is wrong because fine-tuning on the conversation dataset can improve domain-specific fluency but risks overfitting to noise or biases in the data, and without retrieval, the model may still hallucinate when faced with novel queries. Option D is wrong because setting temperature to 0.1 makes the model more deterministic and less creative, which can reduce variability but does not prevent hallucinations; it may cause the model to repeat common but incorrect patterns from training data.

141
Multi-Selecteasy

Which THREE strategies should be combined to effectively reduce biased outputs in a generative AI model? (Choose three.)

Select 3 answers
A.Implement safety filters targeting hate speech and stereotypes.
B.Conduct human evaluation and feedback loops.
C.Use diverse few-shot examples that represent different demographics.
D.Raise the temperature to increase output variability.
E.Fine-tune the model on a biased dataset to learn patterns.
AnswersA, B, C

Safety filters block explicitly biased content.

Why this answer

Implementing safety filters targeting hate speech and stereotypes directly blocks the generation of biased or harmful content at the output layer. These filters use predefined rule sets or trained classifiers to detect and suppress language that reflects demographic or cultural biases, reducing the risk of the model producing offensive or stereotypical responses.

Exam trap

Google often tests the misconception that increasing randomness (temperature) or training on biased data can somehow reduce bias, when in fact both actions worsen the problem by either amplifying noise or embedding the bias deeper into the model's weights.

142
Multi-Selecthard

A company is using a large language model for automated translation of legal contracts. They find that the translations sometimes alter the meaning of specific clauses. Which TWO approaches would most effectively preserve the original meaning? (Choose two.)

Select 2 answers
A.Provide the full contract context in a single prompt.
B.Set top-p=0.1 to limit the vocabulary to the most likely tokens.
C.Fine-tune the model on a parallel corpus of legal translations.
D.Use a glossary of key legal terms with their translations.
E.Increase the temperature to allow more creative phrasing.
AnswersC, D

Fine-tuning on domain-specific translations improves accuracy.

Why this answer

Fine-tuning on a parallel corpus of legal translations adapts the model's weights to the specific legal domain, improving fidelity for legal terminology and clause structures by learning from ground-truth translations. In addition, providing a glossary of key legal terms with their approved translations constrains the model to use consistent, correct terminology, reducing meaning-altering substitutions. Together, these two approaches most effectively preserve the original meaning.

Prompt engineering alone (Option A) and parameter tweaks such as top-p or higher temperature (Options B and E) do not address the root cause of semantic drift.

Exam trap

A common misconception is that prompt engineering alone (Option A) or parameter tweaks like top-p and temperature (Options B and E) can substitute for domain-specific fine-tuning or explicit glossary control, when in fact these methods do not address the root cause of semantic drift in specialized translations.

143
MCQeasy

A marketing team uses Gemini in Vertex AI to generate campaign taglines. The first drafts are generic and closely mirror the prompt wording. They want more distinctive, varied taglines from the same model without changing the model itself. Which parameter should they adjust?

A.Increase the maximum output tokens to allow longer taglines.
B.Lower the temperature to make the output more deterministic.
C.Raise the temperature to increase randomness and diversity in token selection.
D.Enable a higher top-p value to consider more cumulative probability mass.
AnswerC

Temperature scales the probability distribution over next tokens; raising it flattens that distribution so less likely words become more probable. That yields more varied and less generic phrasing, which fits the goal of distinctive taglines. The model is unchanged, so this is a low-effort tuning knob. Very high values risk incoherence, so moderate increases are appropriate.

Why this answer

Temperature is the direct control over how sharply the model favors likely tokens. Raising it flattens the distribution and encourages less predictable word choices, which produces more distinctive taglines from the same model. Lowering temperature increases repetition, output token limits only change length, and top-p is a related but secondary truncation control rather than the primary diversity knob.

Exam trap

The trap here is confusing controls that change response length or sampling cutoffs with the primary diversity control, and mistakenly lowering temperature when the goal is more varied output.

144
MCQeasy

A marketing company wants to fine-tune a generative AI model to adopt a specific brand voice. Which tuning method is most appropriate?

A.RLHF with general user feedback
B.Grounding with external knowledge base
C.Supervised fine-tuning with labeled examples of the brand voice
D.Prompt engineering with system instructions
AnswerC

Supervised fine-tuning trains the model on labelled input-output pairs, directly shaping its outputs to match the desired brand voice. This satisfies the requirement to adopt a specific style, which prompt engineering alone cannot reliably enforce across all generations.

Why this answer

Supervised fine-tuning (SFT) is the most appropriate method because it directly trains the model on a curated dataset of input-output pairs that exemplify the desired brand voice. By adjusting the model's weights through backpropagation on labeled examples, the model learns to mimic the specific tone, vocabulary, and stylistic patterns of the brand, making it the most precise approach for adopting a fixed voice.

Exam trap

A common misconception is that prompt engineering (Option D) is sufficient for fine-grained style control, when in reality it only provides a weak, non-parametric signal that cannot reliably enforce a consistent brand voice across varied contexts.

How to eliminate wrong answers

Option A is wrong because RLHF with general user feedback optimizes for broad human preferences (e.g., helpfulness, harmlessness) rather than a specific, consistent brand voice; it introduces variance from diverse user ratings that can dilute the target style. Option B is wrong because grounding with an external knowledge base retrieves factual information (e.g., via RAG) to reduce hallucination but does not alter the model's generation style or tone; it cannot teach the model to adopt a brand voice. Option D is wrong because prompt engineering with system instructions provides a static, high-level directive that the model may follow inconsistently, especially for nuanced stylistic constraints; it does not update model weights and fails to embed the brand voice deeply into the model's behavior across diverse prompts.

145
MCQhard

A media company uses a Gemini model to produce short news digests from long articles. Editors complain that digests vary in length and sometimes bury the key fact in the middle. The team wants a repeatable, machine-checkable structure for every digest. Which technique should they use?

A.Provide a response schema with defined fields such as headline, key fact, and supporting details, and request JSON output.
B.Fine-tune the model on a corpus of previously published digests written by editors.
C.Set the temperature to 0 and the topK to 1 to make output fully deterministic.
D.Ask the model to be concise and to put the most important information first in the prompt.
AnswerA

A response schema with controlled generation constrains the model to emit fields in a defined structure, making length and placement of the key fact machine-checkable. The model fills required fields rather than free-forming a paragraph, so editors get consistent digests and downstream systems can validate the JSON programmatically.

Why this answer

A response schema with JSON output turns the digest into a structured object with named fields, so the key fact has a fixed location and length is bounded by the schema. This is directly machine-checkable and repeatable, unlike prose instructions, sampling settings, or style-oriented fine-tuning.

Exam trap

The trap here is treating temperature 0 as a guarantee of structure, when determinism only makes the same input reproducible and says nothing about output shape.

146
MCQhard

A data scientist fine-tunes a generative image captioning model to describe medical images. The model outputs safe but very generic captions (e.g., 'An image of cells'). The goal is to produce more specific, clinically relevant descriptions. Which approach is most effective?

A.Perform incremental fine-tuning on a curated dataset of detailed medical image captions.
B.Use diverse beam search during decoding to generate multiple caption candidates.
C.Adjust top-k sampling to restrict the vocabulary to medical terms only.
D.Increase the temperature to encourage the model to output longer, more varied captions.
AnswerA

Incremental fine-tuning on curated detailed medical captions shifts the model's output distribution towards specific clinical vocabulary while preserving its existing image understanding. Generic captions persist because the original training data lacked that specificity, which targeted examples now supply.

Why this answer

Incremental fine-tuning on a high-quality dataset of specific medical captions directly teaches the model the desired level of detail. Option B is wrong because diverse beam search generates varied outputs but does not inherently improve specificity; it may produce multiple candidates but does not ensure clinical relevance. Option C is wrong because top-k sampling can reduce the vocabulary but does not guarantee medical accuracy or detail.

Option D is wrong because increasing temperature adds randomness, which may produce longer or more varied captions but often introduces irrelevant words rather than specific clinical terms.

147
Multi-Selecthard

A healthcare analytics team uses a Gemini model to generate patient-friendly discharge summaries from clinical notes. Clinicians report that summaries occasionally omit critical follow-up instructions. The team must improve recall of these instructions without retraining the model. Which two techniques should they apply? (Choose two.)

Select 2 answers
A.Reduce the maximum output tokens to force the model to be more concise and include only essentials.
B.Increase the temperature so the model produces more varied summaries and captures more details.
C.Enable deterministic decoding and assume the model will then always include every instruction.
D.Add an explicit instruction to extract all follow-up actions into a mandatory checklist section of the summary.
E.Use few-shot prompting with examples that demonstrate surfacing every follow-up instruction in a dedicated section.
AnswersD, E

A direct, unambiguous instruction to place every follow-up action in a required section gives the model a clear structural target and reduces the chance that an item is dropped during free-form generation. Combined with a defined output schema, it makes omissions visible and easier to catch during review.

Why this answer

Few-shot examples and an explicit mandatory checklist section both shape the model toward systematically extracting follow-up actions, improving recall without retraining. Temperature changes, token limits, and deterministic decoding affect randomness or length, not whether required clinical content is captured in the output.

Exam trap

The trap here is equating deterministic or lower-variance decoding with completeness, when consistency and recall are separate properties of a generative system.

148
MCQeasy

A developer wants to improve the factual accuracy of the model's summaries. Based on the exhibit, what should they do?

A.Enable the support engine.
B.Increase the model's context window.
C.Configure grounding with a knowledge base.
D.Re-train the model with a dataset of facts.
AnswerC

Grounding with a knowledge base retrieves authoritative source content and supplies it to the model at generation time, so summaries cite retrieved facts rather than relying on parametric memory. This directly reduces hallucination and improves factual accuracy of the summaries.

Why this answer

Grounding with a knowledge base is the correct approach because it anchors the model's output to a trusted, external source of facts, directly improving factual accuracy without modifying the model's weights. This technique uses retrieval-augmented generation (RAG) to fetch relevant documents from the knowledge base and inject them into the prompt context, ensuring the summary is based on verified information rather than relying solely on the model's parametric memory.

Exam trap

A common pitfall is mistaking grounding for simply expanding the context window or retraining. Grounding with a knowledge base (RAG) provides direct access to verified facts, which is more effective and efficient than further training or fine-tuning, especially in a production environment.

How to eliminate wrong answers

Option A is wrong because enabling the support engine typically refers to a customer support or troubleshooting tool, not a mechanism for improving factual accuracy in generative AI summaries; it does not provide a knowledge base for grounding. Option B is wrong because increasing the model's context window only allows the model to process more tokens in a single request, but it does not introduce new factual information or correct hallucinations; it may even amplify errors if the additional context is unverified. Option D is wrong because re-training the model with a dataset of facts is a costly, time-consuming process that requires significant computational resources and expertise, and it does not guarantee factual accuracy for unseen or evolving information; grounding with a knowledge base is a more efficient and dynamic solution.

149
MCQmedium

A financial services firm is using a foundation model on Vertex AI to generate investment summaries from quarterly reports. The summaries are accurate but often miss key financial metrics and trends. The team cannot afford to fine-tune the model frequently. Which technique should they use to improve the completeness and relevance of the summaries without modifying the model?

A.Increase temperature to 0.9 to encourage more creative outputs.
B.Provide three few-shot examples in the prompt that highlight the desired metrics.
C.Set stop sequences to [' '] to ensure the model finishes each paragraph.
D.Lower top_p to 0.5 to reduce the sampling pool.
AnswerB

Few-shot prompting supplies in-context examples that steer the foundation model toward including the specific financial metrics and trends the firm needs, without altering weights. This satisfies the constraint of avoiding frequent fine-tuning, since the model remains unmodified and behaviour is shaped purely through prompt content at inference time.

Why this answer

Few-shot prompting provides the model with concrete examples of desired output structure and content, guiding it to include key financial metrics and trends without retraining. This technique leverages in-context learning, where the model generalizes from the examples in the prompt to produce more complete and relevant summaries, while avoiding the cost and latency of fine-tuning.

Exam trap

The trap here is that candidates confuse hyperparameter tuning (temperature, top_p) with prompt engineering, assuming that increasing randomness or restricting token selection will improve output quality, when in fact few-shot examples directly teach the model the desired output structure without modifying the model.

How to eliminate wrong answers

Option A is wrong because increasing temperature to 0.9 encourages randomness and creativity, which would likely make summaries less focused and more prone to missing key metrics, not more complete. Option C is wrong because setting stop sequences to ['

'] only controls when the model stops generating text, but does not influence the content or inclusion of specific financial metrics within the output. Option D is wrong because lowering top_p to 0.5 reduces the sampling pool to only the most likely tokens, which can make outputs more repetitive and less likely to include diverse or specific metrics, not improve completeness.

150
Multi-Selectmedium

A healthcare chatbot must avoid hallucinations. Which TWO techniques should the team implement? (Choose two.)

Select 2 answers
A.Set frequency penalty to 0.0
B.Use chain-of-thought prompting
C.Use higher temperature
D.Increase top_k to 50
E.Enable grounding with a knowledge base
AnswersB, E

Chain-of-thought prompting forces the model to expose intermediate reasoning steps before answering, which surfaces unsupported leaps and improves factual grounding in clinical responses. For a healthcare chatbot where hallucination risk must be minimised, this step-by-step decomposition satisfies the accuracy constraint by making faulty logic detectable rather than hidden inside a single confident output.

Why this answer

Chain-of-thought prompting (B) reduces hallucinations by forcing the model to reason step-by-step, which improves factual accuracy and consistency in complex tasks like medical triage. Enabling grounding with a knowledge base (E) anchors the model's output to verified external data, directly preventing fabrication by restricting responses to retrieved facts.

Exam trap

Google often tests the misconception that increasing randomness (temperature, top_k) or disabling penalties improves output quality, when in fact these parameters increase hallucination risk in safety-critical applications like healthcare.

← PreviousPage 2 of 3 · 160 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Improve Gen Ai Output questions.