Courseiva

CCNA Techniques to Improve Generative AI Model Output Questions

75 of 160 questions · Page 1/3 · Techniques to Improve Generative AI Model Output · Answers revealed

1
MCQhard

A healthcare startup uses a generative model fine-tuned on general medical literature to provide preliminary diagnostic suggestions from patient text. The model frequently misses rare diseases and sometimes suggests common conditions that are unlikely given the symptoms. The startup has a curated dataset of rare disease case reports and wants to improve the model’s sensitivity to rare conditions without sacrificing overall accuracy. They cannot afford to retrain the entire model from scratch. The model is deployed on Vertex AI Prediction with low latency requirement. Which approach should they take?

A.Perform continued fine-tuning on the rare disease dataset using a low learning rate.
B.Add a system prompt instructing the model to consider rare diseases more carefully.
C.Reduce top-p sampling to focus on high-probability tokens, assuming rare diseases have lower probability.
D.Implement a human-in-the-loop system: for outputs with low confidence or suspected rare disease, route to a human expert.
AnswerD

Human-in-the-loop catches edge cases without retraining, preserving accuracy for common conditions.

Why this answer

Implementing a human-in-the-loop process for rare disease flags combines AI with expert review, catching misses while maintaining speed for common cases. Option A is wrong because prompt engineering alone may not teach the model about rare diseases. Option B is wrong because increasing top-p restricts vocabulary but doesn't inject knowledge.

Option C is wrong because fine-tuning again might cause catastrophic forgetting of common conditions.

2
MCQeasy

A retail bank uses a generative AI assistant to answer customer questions about account policies. Compliance requires that every response cite the specific internal policy document section it used. Which approach best enforces this requirement?

A.Fine-tune the model on the bank's policy manuals so it memorizes the sections.
B.Use retrieval-augmented generation to fetch policy passages and instruct the model to cite them.
C.Add a system instruction telling the model to never guess and to be accurate.
D.Lower the model's temperature to zero so responses are deterministic.
AnswerB

Retrieval-augmented generation grounds responses in retrieved internal documents and lets the prompt require a citation to the fetched section. This directly satisfies the compliance rule because the model answers from authoritative policy text rather than parametric memory. It also makes citations verifiable, since each claim maps to a retrievable passage the bank controls.

Why this answer

A citation mandate is a grounding problem, not a sampling or behavior problem. Retrieval-augmented generation supplies authoritative policy passages at inference time and lets the prompt demand a citation for each. Temperature tuning, generic instructions, and fine-tuning either do not provide sources or cannot guarantee verifiable, current section references.

Exam trap

The trap here is believing that a strong instruction to be accurate or a low temperature setting produces verifiable citations, when only grounding the response in retrieved source documents can do that.

3
MCQeasy

A data scientist is using Vertex AI generative AI studio to create a chatbot. The chatbot gives inconsistent answers to similar questions. Which parameter should they adjust to make responses more consistent?

A.Decrease temperature to 0.2
B.Increase top-p to 0.9
C.Increase presence penalty to 0.5
D.Decrease frequency penalty to 0.0
AnswerA

Lowering temperature to 0.2 sharpens the probability distribution over the next token, so the model favours its highest-likelihood continuations rather than sampling broadly. That directly addresses the inconsistent answers described in the stem, since similar prompts then converge on near-identical outputs.

Why this answer

Decreasing the temperature to 0.2 reduces the randomness of the model's token sampling, making the output more deterministic and consistent. Temperature controls the probability distribution over tokens; lower values make the model more likely to choose the highest-probability token, reducing variability in responses to similar questions.

Exam trap

Google Cloud often tests the misconception that increasing top-p or adjusting penalties improves consistency, when in fact temperature is the primary parameter for controlling output determinism.

How to eliminate wrong answers

Option B is wrong because increasing top-p to 0.9 increases the cumulative probability threshold for token sampling, which actually introduces more diversity and randomness, making responses less consistent. Option C is wrong because increasing presence penalty to 0.5 penalizes tokens that have already appeared, encouraging the model to introduce new topics and variability, which reduces consistency. Option D is wrong because decreasing frequency penalty to 0.0 removes the penalty for token repetition, which can lead to repetitive or stuck responses but does not directly control the randomness of token selection; consistency is primarily governed by temperature, not frequency penalty.

4
Multi-Selecthard

A team is fine-tuning a large language model for medical advice. Which TWO techniques are most effective for improving the safety and reliability of the model's outputs?

Select 2 answers
A.Constitutional AI
B.Lowering the temperature to 0.0
C.Increasing training data size
D.Increasing top_p to 1.0
E.Reinforcement learning from human feedback (RLHF)
AnswersA, E

Constitutional AI uses predefined rules to guide model behavior.

Why this answer

Constitutional AI (A) is correct because it embeds a set of ethical principles directly into the model's training process, allowing the model to self-critique and revise its outputs to avoid harmful or unsafe medical advice. This technique proactively enforces safety constraints without requiring extensive human labeling, making it highly effective for high-stakes domains like healthcare.

Exam trap

Google Cloud often tests the misconception that hyperparameter tuning (temperature, top_p) or data scaling alone can solve safety issues, when in fact alignment techniques like Constitutional AI and RLHF are specifically designed for that purpose.

5
MCQmedium

Despite applying safety filters, a generative AI model still produces toxic outputs in some cases. Which additional technique should be applied?

A.Add more examples of toxic content to training
B.Increase the filter threshold
C.Use RLHF with human feedback to reduce toxicity
D.Decrease the model's temperature
AnswerC

RLHF fine-tunes the model using human rankings of outputs, directly penalising toxic responses and steering generation toward safer behaviour. This addresses residual toxicity that static safety filters miss, satisfying the stem's need for an additional mitigation beyond filtering.

Why this answer

RLHF (Reinforcement Learning from Human Feedback) directly addresses toxicity by using human evaluators to rank model outputs, then fine-tuning the model to prefer less toxic responses. This technique teaches the model to avoid harmful patterns that safety filters might miss, as filters are static and can be bypassed by adversarial prompts or nuanced toxicity.

Exam trap

A common misconception is that adjusting static parameters (like temperature or filter thresholds) can solve alignment problems, when in fact dynamic human-in-the-loop methods like RLHF are required for nuanced safety issues.

How to eliminate wrong answers

Option A is wrong because adding more examples of toxic content to training would likely reinforce those patterns, increasing rather than decreasing toxicity. Option B is wrong because increasing the filter threshold would make the filter less sensitive, allowing more toxic content through instead of blocking it. Option D is wrong because decreasing the model's temperature reduces randomness in output but does not specifically target or mitigate toxic content generation.

6
MCQmedium

A financial services firm is using a generative AI model to answer customer queries about account balances. The model sometimes provides outdated information because it relies on its training data. The firm wants to ensure the model always uses the most current account data. Which technique should they use?

A.Fine-tune the model daily with the latest account data.
B.Use retrieval-augmented generation (RAG) to fetch current account data from a database.
C.Increase the model's context window to include more historical data.
D.Apply prompt engineering to instruct the model to use the latest data.
AnswerB

RAG combines a generative model with a retrieval system that fetches relevant, up-to-date information from external sources, such as a database. This ensures the model's responses are based on the most current data without retraining. It is ideal for scenarios requiring real-time or frequently updated information.

Why this answer

Retrieval-augmented generation (RAG) is designed to ground model outputs in external, up-to-date knowledge. By integrating a retriever that queries the firm's database for current account balances, the model can generate responses that reflect the latest information. This approach avoids the cost and latency of frequent fine-tuning and ensures accuracy for time-sensitive queries.

Exam trap

The trap here is assuming that the model can access real-time data through its training or by simply instructing it to do so, when in fact external retrieval is necessary to provide current information.

7
MCQhard

A research team is using a large language model to analyze medical research papers and generate summaries. They need to minimize hallucinations while retaining key details. They have access to a curated database of paper abstracts. Which approach is best?

A.Fine-tune the model on the entire database of papers.
B.Use chain-of-thought prompting to reason step-by-step.
C.Use few-shot prompting with examples of accurate summaries and set temperature=0.0.
D.Implement RAG to retrieve relevant abstracts and incorporate them into the prompt.
AnswerD

RAG grounds generation in the curated abstracts by retrieving relevant passages and inserting them into the prompt, so the model summarises supplied evidence rather than relying on parametric memory. This directly minimises hallucination while retaining key details, satisfying the accuracy constraint.

Why this answer

Retrieval-Augmented Generation (RAG) directly addresses hallucination by grounding the model's output in a curated database of paper abstracts. By retrieving relevant abstracts and injecting them into the prompt, the model generates summaries based on verified facts rather than relying solely on its parametric knowledge, which is the most effective way to minimize hallucinations while retaining key details.

Exam trap

Many candidates mistakenly think that fine-tuning or low temperature alone can solve hallucination, but the trap here is that without external retrieval (RAG), the model has no mechanism to verify facts against a trusted source, so it will still generate plausible-sounding but incorrect details.

How to eliminate wrong answers

Option A is wrong because fine-tuning on the entire database of papers does not prevent hallucinations; it can cause catastrophic forgetting and the model may still fabricate details when asked to summarize unseen or edge-case content. Option B is wrong because chain-of-thought prompting improves reasoning but does not provide external factual grounding, so the model can still hallucinate based on its internal knowledge. Option C is wrong because few-shot prompting with temperature=0.0 reduces randomness but does not supply the model with the actual abstracts to reference; it relies on the model's memory of the examples, which can lead to hallucinated details not present in the source papers.

8
MCQhard

A large e-commerce company deploys a generative AI chatbot on Vertex AI for customer service. The chatbot is powered by a fine-tuned model on the company's historical support tickets. Despite high accuracy on training topics, the chatbot frequently gives irrelevant or off-topic answers when customers ask about new products or promotions. The company maintains a comprehensive product catalog and a knowledge base of current promotions. The chatbot's prompts include a system instruction to 'Answer based on your knowledge' and no other retrieval mechanism. The response time requirement is under 3 seconds. Which course of action should the team take?

A.Implement a RAG pipeline that retrieves relevant product and promotion data from the knowledge base and injects it into the prompt.
B.Increase the temperature to encourage the model to generate more diverse answers.
C.Add additional safety filters to block irrelevant responses.
D.Fine-tune the model again on a larger dataset that includes recent support tickets.
AnswerA

Retrieval-augmented generation fetches current product and promotion content from the knowledge base and injects it into the prompt, grounding answers in up-to-date facts the fine-tuned model never saw. This addresses the off-topic responses while keeping latency within the three-second requirement.

Why this answer

Implementing a RAG (Retrieval-Augmented Generation) pipeline directly addresses the chatbot's inability to answer questions about new products or promotions. By retrieving relevant, up-to-date information from the company's product catalog and knowledge base and injecting it into the prompt, the model gains access to current data beyond its training cutoff. This approach keeps response times under 3 seconds (as retrieval is fast) and avoids the need for costly retraining, while the system instruction 'Answer based on your knowledge' is replaced with grounded context.

Exam trap

Google often tests the misconception that fine-tuning alone can solve knowledge gaps for dynamic or time-sensitive data, when in reality RAG is the appropriate technique for incorporating external, frequently updated information without retraining.

How to eliminate wrong answers

Option B is wrong because increasing the temperature would make the model generate more random and diverse outputs, which would worsen the problem of irrelevant or off-topic answers rather than fix it. Option C is wrong because adding safety filters blocks harmful or inappropriate content but does not solve the core issue of the model lacking knowledge about new products or promotions; it would merely suppress irrelevant responses without providing correct information. Option D is wrong because fine-tuning again on a larger dataset that includes recent support tickets is time-consuming, expensive, and still cannot keep up with rapidly changing promotions or new products; the model would remain static after training, whereas RAG provides dynamic, real-time retrieval.

9
Multi-Selecthard

Which THREE approaches are effective for reducing bias in generative model outputs? (Choose three.)

Select 3 answers
A.Set temperature to a very high value.
B.Use adversarial training.
C.Use a balanced training dataset.
D.Use prompt engineering to specify neutral tone.
E.Fine-tune on a debiased dataset.
AnswersC, D, E

Balanced data reduces representation bias.

Why this answer

A balanced training dataset reduces the risk of the model learning spurious correlations or skewed distributions that lead to biased outputs. By ensuring that all demographic groups, topics, or perspectives are represented proportionally, the model's learned probability distribution is less likely to favor one group over another, directly mitigating representation bias at the data level.

Exam trap

The trap here is that candidates confuse randomness (high temperature) with fairness, or mistake adversarial training (a robustness technique) for a bias mitigation method, when in fact bias reduction requires data-level or fine-tuning interventions like balanced datasets, debiased fine-tuning, or prompt engineering.

10
MCQmedium

A retail bank uses a Gemini model on Vertex AI to answer customer questions about its mortgage products. Testers report that the model sometimes invents interest rates that do not exist in the bank's rate sheet. The team wants the model to ground its answers only in an approved corpus of product documents and cite the passages it used. Which approach should they implement?

A.Lower the temperature parameter to 0 and increase the topK value so the model samples only from the most probable tokens.
B.Configure Vertex AI Search grounding with the bank's document store and enable citations in the generateContent request.
C.Fine-tune the Gemini model on the bank's historical customer support transcripts so it learns the correct rates.
D.Increase the model's output token limit so it has more room to state the correct rate before answering.
AnswerB

Grounding with Vertex AI Search retrieves passages from the indexed rate sheets and product documents, injects them into the prompt, and returns grounding metadata so the response can cite its supporting sources. This directly constrains the model to approved content and satisfies the citation requirement, which sampling parameters alone cannot do.

Why this answer

Grounding with Vertex AI Search connects the model to the bank's authoritative document store, so answers are conditioned on retrieved passages rather than on parametric memory, and the grounding metadata enables citations. Sampling controls only shape randomness, fine-tuning bakes in a snapshot of knowledge without provenance, and token limits only change length.

Exam trap

The trap here is assuming that deterministic sampling settings such as temperature 0 make a model factually accurate, when they only reduce randomness and never supply missing source data.

11
Multi-Selecthard

A team is fine-tuning a model for a legal document summarization task. They need to ensure high accuracy and avoid hallucinations. Which TWO approaches should they combine? (Choose two.)

Select 2 answers
A.Use Retrieval-Augmented Generation to retrieve relevant legal texts
B.Increase temperature to 1.5 during inference
C.Implement early stopping during fine-tuning
D.Incorporate a human-in-the-loop review process
E.Use character-level tokenization to improve spelling
AnswersA, D

RAG grounds the summary in actual documents, reducing hallucination.

Why this answer

Retrieval-Augmented Generation (RAG) is correct because it grounds the model's output in retrieved, authoritative legal texts, directly reducing hallucination by providing factual context during generation. This is critical for legal summarization where accuracy is paramount, as RAG ensures the model references specific statutes or case law rather than relying solely on its parametric memory.

Exam trap

A common misconception is that increasing temperature or using training-time techniques like early stopping can improve inference accuracy, when in fact they either increase randomness or address overfitting, not factual grounding. This trap is frequently tested in Google certification exams.

12
MCQhard

An e-commerce company fine-tunes a model on customer reviews to generate product feedback summaries. They want to ensure the model does not reproduce toxic language from the training data. Besides filtering the training data, which additional technique is most effective at inference time?

A.Set temperature to 0.0 to reduce variance
B.Set top-k to 10 to limit token choices
C.Pass the model output through a toxicity detection model and conditionally regenerate or block
D.Use beam search with a high beam width
AnswerC

Filtering training data alone cannot guarantee toxic-free output, so a post-generation toxicity classifier screens each summary and triggers regeneration or blocking when thresholds are breached. This directly satisfies the inference-time constraint, catching residual toxicity the fine-tuning absorbed. It operates on the model's actual output rather than inputs, making it the most reliable safeguard.

Why this answer

It directly addresses the safety requirement at inference time by introducing a secondary guardrail. A toxicity detection model (e.g., a classifier trained on the Jigsaw Toxic Comment dataset) can score the generated output in real time; if the score exceeds a threshold, the system can either block the response or trigger a regeneration with adjusted parameters. This is the only technique that actively filters for toxic language after generation, rather than merely reducing output variance or exploring alternative sequences.

Exam trap

Google often tests the misconception that controlling randomness (temperature, top-k) or search strategy (beam search) can prevent toxic outputs, when in fact these techniques only affect token probability distributions and do not perform any semantic safety filtering.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0.0 makes the model deterministic (always picks the highest-probability token), which reduces randomness but does not prevent the model from reproducing toxic phrases that were present in the training data. Option B is wrong because top-k sampling limits the token pool to the k most likely tokens, which can still include toxic tokens if they rank highly; it does not perform any semantic or safety filtering. Option D is wrong because beam search with a high beam width explores multiple candidate sequences to find a high-probability output, but it does not incorporate any toxicity detection or safety constraint, so it may still select a toxic sequence if it has high likelihood.

13
MCQeasy

A user provides a long document as context for a question-answering task, but the model outputs irrelevant answers. What is the most likely cause?

A.The document exceeds the model's context window, truncating important details.
B.Safety filters are blocking the relevant response.
C.The model's temperature is too low, making it deterministic.
D.The model is not generating any tokens.
AnswerA

Transformer models process a fixed token budget; text beyond the context window is silently truncated. With a long document, the crucial passages answering the question may fall outside that window, so the model responds from incomplete context, producing irrelevant output.

Why this answer

The most common cause of irrelevant answers when a long document is provided is that the document exceeds the model's fixed context window (e.g., 8K tokens for PaLM 2, 128K for Gemini 1.0 Pro, or up to 1M for Gemini 1.5 Pro). When the input is truncated, critical details needed for accurate retrieval and generation are lost, leading to off-target responses. This is a fundamental limitation of transformer architectures, which cannot attend to tokens beyond their maximum sequence length.

Exam trap

Google often tests the misconception that safety filters or temperature settings are the primary cause of irrelevant outputs, when in fact the context window limit is the most direct and common technical constraint in long-document QA tasks.

How to eliminate wrong answers

Option B is wrong because safety filters block harmful or policy-violating content, not relevant factual answers; they would either suppress the response entirely or return a refusal, not produce irrelevant answers. Option C is wrong because a low temperature (e.g., 0.0) makes the model more deterministic and repetitive, but it does not cause irrelevance—it would still generate answers based on the available context, albeit with less creativity. Option D is wrong because if the model were not generating any tokens, the output would be empty or a failure, not irrelevant answers; the question explicitly states the model outputs irrelevant answers, meaning tokens are being generated.

14
MCQeasy

A social media company uses a generative AI model to moderate user posts. The model occasionally allows offensive content. Which safety technique should be implemented?

A.Use a different tokenizer to avoid offensive words.
B.Configure safety filters on the model endpoint in Vertex AI.
C.Add few-shot examples of safe posts in the prompt.
D.Reduce the temperature to 0.
AnswerB

Configuring safety filters on the Vertex AI model endpoint applies configurable thresholds that block harmful categories before responses reach users, catching offensive content the base model permits. This is a platform-level control applied at inference time, complementing prompt design or fine-tuning.

Why this answer

Configuring safety filters on the model endpoint in Vertex AI directly blocks offensive content at inference time by applying predefined or custom safety thresholds (e.g., toxicity, harassment categories). This is the most reliable technique for real-time moderation, as it prevents harmful outputs regardless of prompt engineering or tokenization changes.

Exam trap

A common misconception tested in the Google Gen AI Leader exam is that prompt engineering (few-shot examples) or parameter tuning (temperature) can substitute for dedicated safety mechanisms, but these do not provide hard guarantees against offensive content.

How to eliminate wrong answers

Option A is wrong because changing the tokenizer does not prevent the model from generating offensive content; tokenizers only split text into tokens and do not understand or filter semantics. Option C is wrong because few-shot examples in the prompt can guide the model but are not a safety mechanism—they can be overridden by the model's training data or adversarial inputs, and they do not enforce hard safety constraints. Option D is wrong because reducing temperature to 0 makes the model deterministic but does not eliminate offensive content; it may even amplify biased or toxic patterns from the training data by always choosing the most likely token.

15
Multi-Selectmedium

A company is deploying a generative AI system that generates customer-facing emails. The system must ensure outputs are not toxic, biased, or harmful. Which TWO techniques are most effective for reducing toxicity in model outputs without significantly affecting performance?

Select 2 answers
A.Increase the maximum output token count to allow more context.
B.Set temperature to a very low value (e.g., 0.1).
C.Fine-tune the model on a dataset of safe emails using reinforcement learning from human feedback (RLHF).
D.Apply a toxicity detection and filtering layer using Vertex AI Safety Filters.
E.Provide 50 few-shot examples of safe emails in every prompt.
AnswersC, D

RLHF fine-tuning adjusts the model's weights toward human-preferred, non-toxic responses, reducing harmful generations at the source. Because the change is learned rather than bolted on, it preserves general capability better than prompt-only or post-hoc filtering approaches.

Why this answer

Option C is correct because fine-tuning with RLHF directly aligns the model's behavior with human preferences for safe, non-toxic content, teaching it to avoid biased or harmful outputs at the model level without degrading overall generation quality. Option D is correct because a toxicity detection and filtering layer such as Vertex AI Safety Filters evaluates generated text against configurable harm categories and blocks or suppresses toxic outputs before they reach customers, adding a reliable guardrail. Option A is not effective because increasing max output tokens only changes length limits and can actually introduce more opportunities for harmful content.

Option B is not effective because a very low temperature reduces randomness but does not remove learned toxic patterns and can hurt creativity and diversity. Option E is not effective because 50 few-shot examples consume significant context, increase cost and latency, and provide only shallow, prompt-level guidance that is easily bypassed.

Exam trap

A common misconception is that simply adjusting generation parameters (like temperature or token count) or adding more examples can effectively control toxicity, when in fact these methods do not address the root cause of harmful patterns in the model's behavior.

16
Multi-Selectmedium

A developer is tuning a text-generation model for creative writing. They want the outputs to be more diverse and less repetitive. Which THREE parameters/changes can help? (Choose three.)

Select 3 answers
A.Increase temperature to 0.9
B.Reduce top-k to 10
C.Increase presence penalty to 0.5
D.Increase top-p to 0.95
E.Reduce frequency penalty to 0.0
AnswersA, C, D

Higher temperature increases randomness and diversity.

Why this answer

Increasing temperature to 0.9 raises the randomness of the probability distribution over the vocabulary, making the model more likely to sample less probable tokens. This directly increases output diversity and reduces repetitiveness by flattening the softmax curve, which is a standard technique for creative generation.

Exam trap

Google Cloud often tests the misconception that reducing top-k or top-p increases diversity, when in fact narrowing the sampling pool (lower top-k or lower top-p) reduces diversity, and the correct approach is to increase these values or increase temperature/penalties.

17
Multi-Selecthard

A media company uses Gemini models on Vertex AI to draft news briefs from long press releases. Editors report that drafts sometimes invent quotes and statistics that do not appear in the source. The team wants to reduce these fabrications while keeping the model's fluent writing. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Raise the top-k sampling value to let the model consider more candidate tokens.
B.Enable grounding with Vertex AI Search over a curated corpus of approved press releases.
C.Provide the press release as context and instruct the model to answer only from the provided text.
D.Shorten the press release by removing paragraphs before sending it to the model.
E.Increase the model's temperature so it paraphrases the release more creatively.
AnswersB, C

Grounding against a curated corpus retrieves verified source passages and ties the draft to them, which directly counters invented quotes and numbers. Because the corpus is controlled, the model has authoritative evidence to draw from, and citations can be surfaced for editorial review. This complements prompt-level constraints and keeps the writing fluent while improving factual fidelity.

Why this answer

Combining an explicit instruction to answer only from the supplied text with retrieval grounding over an approved corpus gives the model both a behavioral constraint and authoritative evidence. Together they reduce invented quotes and statistics while preserving fluent composition. Sampling changes like higher top-k or temperature increase randomness and work against factual fidelity, and deleting source content removes the very evidence the model needs.

Exam trap

The trap here is treating sampling parameters such as top-k or temperature as anti-hallucination controls, when in fact raising them increases randomness and makes fabricated content more likely.

18
Multi-Selecthard

A healthcare company is using a generative AI model to produce patient education materials. They want to ensure the output is accurate, avoids harmful advice, and adheres to medical guidelines. Which TWO techniques should they implement to improve the safety and reliability of the model's output? (Choose two.)

Select 2 answers
A.Reduce the maximum output token limit to shorten responses.
B.Increase the model's temperature to encourage more diverse responses.
C.Implement retrieval-augmented generation (RAG) using a curated medical knowledge base.
D.Fine-tune the model on a dataset of patient education materials without any safety annotations.
E.Use a safety filter or content moderation API to screen generated text.
AnswersC, E

RAG enhances the model's output by retrieving relevant, up-to-date information from a trusted knowledge base. This grounds the generation in accurate medical facts and guidelines, reducing hallucinations and ensuring adherence to current standards. It directly improves accuracy and safety for patient education materials.

Why this answer

Retrieval-augmented generation grounds the model in a trusted medical knowledge base, ensuring accuracy and up-to-date information. Safety filters screen out harmful content, adding a protective layer. Together, these techniques enhance both the reliability and safety of patient education materials.

Exam trap

The trap here is thinking that fine-tuning alone can ensure safety, when in fact it requires curated data and additional safeguards.

19
MCQhard

A data scientist fine-tunes a model on a small proprietary dataset. After fine-tuning, the model repeats training examples verbatim. What is the most effective mitigation?

A.Reduce the temperature during inference to 0.
B.Train for more epochs to improve generalization.
C.Use early stopping based on validation loss.
D.Add regularization like dropout and use a smaller learning rate.
AnswerD

Verbatim repetition indicates overfitting to the small proprietary dataset. Dropout randomly deactivates neurons during training, while a smaller learning rate prevents sharp memorisation of individual examples, directly reducing the model's tendency to reproduce training data.

Why this answer

Overfitting on a small dataset causes the model to memorize training examples rather than generalize. Adding dropout introduces noise that forces the model to learn more robust features, while a smaller learning rate prevents the model from over-optimizing on the limited data. Together, these regularization techniques reduce the model's capacity to memorize verbatim outputs.

Exam trap

Google Gen AI Leader often tests the misconception that reducing temperature or training longer improves generalization, when in fact these actions either increase determinism (and memorization) or worsen overfitting on small datasets.

How to eliminate wrong answers

Option A is wrong because reducing temperature to 0 makes the model deterministic, which can actually increase the likelihood of repeating exact training examples by always choosing the highest-probability token, exacerbating memorization. Option B is wrong because training for more epochs on a small dataset worsens overfitting, causing the model to memorize more training examples verbatim, not improve generalization. Option C is wrong because early stopping based on validation loss can help prevent overfitting, but it does not directly address the model's capacity to memorize; the model may still have high capacity and memorize examples before overfitting occurs, making regularization a more effective mitigation.

20
MCQhard

A healthcare analytics team uses Gemini on Vertex AI to extract structured data from clinical notes. The model occasionally outputs invalid JSON, breaking downstream processing. The team wants to enforce a strict output schema. Which approach should they use?

A.Use controlled generation with a response schema in the Gemini API request to constrain output to the defined JSON structure.
B.Fine-tune the model on a dataset of clinical notes paired with valid JSON outputs.
C.Post-process the model output with a regular expression to fix malformed JSON before parsing.
D.Add 'return JSON' to the prompt and set temperature to 0.2 to encourage valid formatting.
AnswerA

Controlled generation with a response schema constrains the model's decoding to produce valid JSON matching the specified fields and types. This guarantees parseable output and reduces the need for post-processing or retries. It is the recommended approach when downstream systems require strict structure. The schema is provided in the API request alongside the prompt.

Why this answer

Controlled generation with a response schema constrains decoding so the model emits JSON that conforms to the defined structure, guaranteeing parseable output. Prompt hints and low temperature reduce but do not eliminate format errors, fine-tuning is heavier and still not absolute, and regex repair is fragile. Schema-constrained decoding is the correct tool for strict structured output.

Exam trap

The trap here is assuming that telling the model to return JSON or fine-tuning guarantees valid JSON, when only a decoding-time schema constraint enforces it.

21
Multi-Selecteasy

Which TWO methods are most effective for improving factual accuracy in a language model's responses? (Choose two.)

Select 2 answers
A.Use prompt engineering to instruct the model to rely on provided facts.
B.Decrease the temperature to make responses more deterministic.
C.Increase top-k sampling to consider a wider range of tokens.
D.Replace the model with a smaller, more focused model.
E.Implement Retrieval-Augmented Generation (RAG) with a trusted knowledge base.
AnswersA, E

Prompt engineering can explicitly direct the model to verify claims or stick to given knowledge.

Why this answer

Prompt engineering can explicitly instruct the model to base its responses on provided facts, reducing reliance on parametric knowledge that may be outdated or incorrect. By including directives like 'Use only the information in the following text' or 'Answer based solely on the provided context,' the model is guided to prioritize given facts over its internal training data, which improves factual accuracy in the output.

Exam trap

A common misconception is that reducing randomness (temperature) or increasing token diversity (top-k) directly improves factual accuracy, when in fact these parameters affect output style and creativity, not the correctness of the underlying facts.

22
MCQmedium

A company uses a generative AI model to generate product descriptions. They notice variations in style and length across products. How can they enforce consistent formatting?

A.Adjust top-k sampling to include more token candidates.
B.Set a system instruction specifying style and structure.
C.Randomly select few-shot examples from a pool of descriptions.
D.Use a high temperature and vary the prompt slightly.
AnswerB

A system instruction is prepended to every request and constrains the model's output style and structure, so each product description follows the same format regardless of input variation. This directly enforces the consistent formatting the stem requires.

Why this answer

Setting a system instruction explicitly defines the desired style, tone, and structure for the model's output. This is a fundamental technique in prompt engineering, particularly with instruction-tuned models like GPT-4 or Claude, where the system message acts as a persistent directive that overrides the model's default behavior, ensuring consistent formatting across all generated product descriptions.

Exam trap

A common misconception in prompt engineering is that increasing randomness (via temperature or top-k) or varying examples can enforce consistency. In reality, these techniques increase variability, while system instructions provide the deterministic control needed for uniform output formatting. This question tests the ability to identify the best practice for consistent formatting in generative AI models.

How to eliminate wrong answers

Option A is wrong because adjusting top-k sampling to include more token candidates increases randomness and diversity, which would exacerbate style and length variations rather than enforce consistency. Option C is wrong because randomly selecting few-shot examples from a pool of descriptions introduces uncontrolled variability in the examples, leading to inconsistent output formatting instead of a fixed template. Option D is wrong because using a high temperature increases randomness and varying the prompt slightly introduces instability, both of which undermine the goal of consistent formatting.

23
MCQmedium

To improve factuality in generative AI, which is the best approach?

A.Set top_p to 0.1
B.Reduce output length
C.Grounded generation with citations
D.Increase model size
AnswerC

Grounded generation retrieves authoritative source content and conditions the model's output on it, with citations letting users verify each claim. This directly constrains hallucination, satisfying the stem's goal of improving factuality rather than relying on the model's parametric memory alone.

Why this answer

Grounded generation with citations directly addresses factuality by forcing the model to retrieve and cite verifiable sources (e.g., from a knowledge base or document store) before generating an answer. This approach, often implemented via retrieval-augmented generation (RAG), ensures outputs are anchored to external evidence rather than relying solely on the model's parametric memory, which can produce hallucinations. Citations also enable users to verify claims, making this the most effective technique for improving factual accuracy.

Exam trap

Google exams often test the misconception that hyperparameter tuning (like top_p) or scaling model size alone can fix factuality, when in reality these methods do not address the root cause of hallucination—lack of external grounding—and candidates may overlook the importance of retrieval and citation mechanisms, such as Google's Vertex AI Grounding or RAG.

How to eliminate wrong answers

Option A is wrong because setting top_p to 0.1 is a nucleus sampling parameter that reduces output diversity by limiting the cumulative probability mass of token choices, but it does not improve factuality—it only makes outputs more deterministic and potentially repetitive. Option B is wrong because reducing output length limits the amount of text generated but does not prevent the model from fabricating facts within that shorter span; hallucinations can occur in even a single sentence. Option D is wrong because increasing model size generally improves language fluency and knowledge capacity but does not guarantee factuality—larger models can still hallucinate confidently, and scaling alone does not provide a mechanism for grounding outputs in verified sources.

24
MCQmedium

A data science team is fine-tuning a large language model using Vertex AI to generate marketing copy. They notice that the generated text is often repetitive and lacks creativity. Which technique should they apply to improve output diversity?

A.Increase the temperature parameter to 0.9.
B.Decrease the beam search width to 1.
C.Decrease the top-k sampling threshold.
D.Add more examples of repetitive text to the training dataset.
AnswerA

Temperature controls the randomness of token sampling during generation. Raising it to 0.9 flattens the probability distribution, so lower-probability tokens are selected more often, directly countering the repetitive, low-creativity output observed in the marketing copy.

Why this answer

Increasing the temperature parameter to 0.9 raises the randomness of the probability distribution over tokens, allowing less likely tokens to be selected. This directly counteracts repetitive output by encouraging the model to explore more diverse word choices, which is a standard technique for improving creativity in text generation.

Exam trap

Google Cloud often tests the misconception that decreasing sampling thresholds (like top-k or beam width) increases diversity, when in fact they reduce the candidate pool and make output more deterministic.

How to eliminate wrong answers

Option B is wrong because decreasing beam search width to 1 reduces the number of candidate sequences considered, which actually makes output more deterministic and less diverse, worsening repetitiveness. Option C is wrong because decreasing the top-k sampling threshold restricts the model to only the k most likely tokens, which reduces diversity and can increase repetition. Option D is wrong because adding more examples of repetitive text to the training dataset would reinforce the unwanted behavior, making the model more likely to generate repetitive output, not less.

25
MCQmedium

A research team uses a generative AI model to answer questions about internal technical documents. The model sometimes provides outdated information because it relies on its pretrained knowledge instead of the latest documents. The team wants the answers to be based on the current document set. Which technique should they implement?

A.Fine-tune the model on the internal documents to update its knowledge.
B.Increase the model's temperature so it relies less on memorized facts.
C.Retrieval-augmented generation that fetches relevant passages from the current document set and includes them in the prompt.
D.Add a prompt instruction telling the model to always use the most recent information it knows.
AnswerC

Retrieval-augmented generation supplies the model with up-to-date passages at inference time, so answers are grounded in the current documents rather than pretrained memory. This directly solves the outdated-information problem without retraining. It also allows the document set to be updated independently of the model. The model uses the retrieved context to generate accurate, source-based responses.

Why this answer

Retrieval-augmented generation is the right technique because it dynamically fetches relevant passages from the current document set and includes them in the prompt. This grounds the model in up-to-date information without retraining, directly solving the outdated-answer problem. Other options either change randomness, require costly retraining, or rely on an instruction the model cannot fulfill without access to the documents.

Exam trap

The trap here is assuming that a prompt instruction or temperature change can make the model use current documents, when it actually needs the documents to be retrieved and provided.

26
MCQeasy

To ensure that a generative AI model uses the most current information from the web for answering user queries, which Vertex AI feature should be enabled?

A.Grounding with Google Search
B.Safety filters
C.Context caching
D.Model tuning
AnswerA

Grounding with Google Search connects the model to live web results at query time, so answers reflect current information rather than the training cutoff. This satisfies the freshness requirement without retraining or fine-tuning the underlying model.

Why this answer

Grounding with Google Search is the correct feature because it enables the model to retrieve and reference real-time information from the web, ensuring responses are based on the most current data available. This is achieved by integrating Google Search results directly into the model's generation process, allowing it to cite live sources and reduce hallucinations from outdated training data.

Exam trap

Google Cloud often tests the distinction between features that improve output quality through external data retrieval (Grounding) versus those that modify the model's internal behavior (tuning, caching, filtering), leading candidates to confuse safety or optimization features with live data access.

How to eliminate wrong answers

Option B is wrong because safety filters are designed to block harmful or inappropriate content, not to fetch current web information. Option C is wrong because context caching stores frequently accessed context to reduce latency and cost, but it does not provide live web data. Option D is wrong because model tuning adjusts the model's parameters on a specific dataset to improve performance on a task, but it does not enable real-time web retrieval.

27
Multi-Selectmedium

A product team is using a generative AI model to summarize customer feedback from multiple sources. The summaries are sometimes missing key themes or including irrelevant details. The team wants to improve the quality of the summaries without retraining the model. (Choose two.)

Select 2 answers
A.Provide a structured prompt that specifies the sections to include, such as top themes, sentiment, and representative quotes.
B.Reduce the maximum output tokens so the model is forced to include only the most important points.
C.Increase the temperature to encourage the model to explore more diverse themes.
D.Use few-shot examples that show correctly formatted summaries with the desired level of detail.
E.Fine-tune the model on a dataset of customer feedback summaries.
AnswersA, D

A structured prompt tells the model exactly what to cover, reducing omissions of key themes and irrelevant additions. By defining sections like top themes, sentiment, and quotes, the team guides the model's focus. This is a prompt-engineering technique that requires no retraining. It directly addresses the missing-theme and irrelevant-detail problems by making the expected output explicit.

Why this answer

A structured prompt and few-shot examples both guide the model to include the right themes and exclude irrelevant details. The structured prompt defines the required sections, while examples show the desired format and depth. Together they improve summarization quality without retraining, directly addressing the team's concerns about missing themes and extraneous content.

Exam trap

The trap here is thinking that output length limits or higher randomness can fix summarization quality, when the real need is explicit guidance on content and format.

28
MCQeasy

A developer is using Vertex AI Studio to test prompts for a text generation model. They want the model to follow a specific output format (JSON). Which prompt engineering approach is most effective?

A.Set stop sequences to '}'.
B.Include a few-shot example of the exact JSON format in the prompt.
C.Set the system instruction to 'Always output JSON.'
D.Set temperature to 0 to make output deterministic.
AnswerB

Few-shot prompting supplies concrete input-output pairs demonstrating the exact JSON schema, so the model infers field names, nesting and types rather than guessing. This constrains generation to the required structure more reliably than describing the format in prose alone.

Why this answer

Including a few-shot example of the exact JSON format in the prompt provides the model with a concrete pattern to follow, which is the most reliable method for enforcing structured output in generative models. Few-shot prompting leverages in-context learning, where the model uses the provided example to infer the desired schema and formatting rules, reducing ambiguity and improving adherence to the specified JSON structure.

Exam trap

Google Cloud often tests the misconception that system instructions or hyperparameter tuning alone can enforce output format, when in practice, few-shot examples are the most direct and reliable method for guiding model behavior in structured generation tasks.

How to eliminate wrong answers

Option A is wrong because setting stop sequences to '}' would prematurely terminate generation at the first closing brace, which may cut off nested JSON objects or arrays, and does not guarantee the model outputs valid JSON from the start. Option C is wrong because a system instruction like 'Always output JSON' is a high-level directive that models often fail to follow precisely without explicit formatting examples, as they may still produce markdown, extra text, or malformed JSON. Option D is wrong because setting temperature to 0 makes output deterministic but does not enforce a specific output format; the model could still generate non-JSON text or deviate from the required schema, as temperature controls randomness, not structure.

29
MCQhard

A team is using Vertex AI Pipelines to deploy a generative AI model for real-time inference. The model sometimes generates harmful content. They want to implement a safety filter that checks the output before returning it to the user, but they need to minimize latency. Which approach best balances safety and performance?

A.Use a secondary lightweight classifier to filter outputs in real-time.
B.Retrain the model on every flagged harmful output.
C.Manually review all outputs before delivery.
D.Disable safety checks to improve latency.
AnswerA

A lightweight secondary classifier inspects each generated output and blocks harmful content before it reaches the user. Its low compute overhead adds minimal latency compared with a full model-based filter, satisfying the stem's need to balance safety against real-time inference performance.

Why this answer

Deploying a secondary lightweight classifier (e.g., a distilled BERT or a small logistic regression model) as a post-processing filter allows real-time inference with minimal latency overhead. This approach decouples safety from the primary generative model, enabling fast rejection of harmful outputs without retraining or blocking the main inference pipeline.

Exam trap

Google Cloud often tests the misconception that safety must be integrated into the generative model itself (e.g., via retraining or fine-tuning), when in practice a separate, lightweight post-processing filter is the standard for low-latency production systems.

How to eliminate wrong answers

Option B is wrong because retraining the model on every flagged harmful output is computationally expensive, introduces significant latency, and can lead to catastrophic forgetting or overfitting to specific examples, making it impractical for real-time inference. Option C is wrong because manual review of all outputs introduces unacceptable latency and does not scale, violating the requirement to minimize latency. Option D is wrong because disabling safety checks entirely eliminates the safety requirement, which is explicitly needed, and would expose users to harmful content, failing the core objective.

30
Multi-Selectmedium

A data science team is building a document question-answering assistant on Vertex AI. Users report that answers are sometimes fabricated when the retrieved passages do not contain the answer. Which TWO techniques should the team apply to reduce hallucinations? (Choose two.)

Select 2 answers
A.Raise the maximum output token limit so the model has more space to explain its reasoning and avoid mistakes.
B.Instruct the model to answer only from the provided context and to respond with a fixed phrase such as 'not found in the provided documents' when the context is insufficient.
C.Increase the temperature so the model explores alternative answers when the retrieved context seems incomplete.
D.Return citations to the specific retrieved passages used and require the answer to reference them.
E.Fine-tune the model on the full document corpus so all answers are stored in the model weights.
AnswersB, D

An explicit grounding instruction with a defined abstention phrase gives the model a safe path when evidence is missing. It reduces fabrication because the model is told to prefer the supplied passages and to admit when they do not contain the answer. This is a low-cost prompt change that directly targets unsupported claims.

Why this answer

Grounding instructions with an abstention path and mandatory citations both push the model to rely on retrieved passages and to expose when evidence is missing. Higher temperature and larger output limits increase unsupported text, while fine-tuning stores a stale snapshot without providing query-time provenance. Together, the two prompt and output controls reduce fabrication.

Exam trap

The trap here is believing that more randomness or a larger output budget improves reasoning, when both actually increase the chance of unsupported claims.

31
MCQmedium

A developer uses the Gemini API to summarize long articles. The summaries often miss key points from the end of the article. Which technique specifically addresses this length-based loss of information?

A.Increase the max output tokens to 2048
B.Break the article into sections and ask the model to summarize each section, then combine
C.Truncate the article to the first 2000 tokens
D.Use a different model with a larger context window
AnswerB

This structured approach ensures each part is summarized, mitigating attention drop-off.

Why this answer

The Gemini API, like many LLMs, has a limited context window and exhibits a 'lost-in-the-middle' effect where information at the beginning and end of long inputs is retained better, but the middle and far end can be dropped. By breaking the article into sections, summarizing each independently, and then combining those summaries, you ensure that key points from the end are captured in their own focused summary, bypassing the length-based information loss. This technique is a form of 'chunking' and 'recursive summarization' that directly addresses the model's tendency to lose context over long sequences.

Exam trap

Google often tests the misconception that simply increasing token limits or using a larger context window solves all length-related issues, when in fact the underlying attention mechanism and positional biases require explicit chunking strategies to reliably capture information from all parts of a long input.

How to eliminate wrong answers

Option A is wrong because increasing max output tokens only controls the length of the generated summary, not the input context; the model still processes the full article and may drop end content due to its fixed context window. Option C is wrong because truncating the article to the first 2000 tokens removes the end of the article entirely, which is the opposite of what is needed to capture key points from the end. Option D is wrong because using a different model with a larger context window does not guarantee the end content will be retained; even with larger windows, models can still suffer from the 'lost-in-the-middle' phenomenon where information in the middle and end is less reliably recalled.

32
MCQmedium

A financial services firm uses a generative AI model on Vertex AI to answer employees' HR policy questions. The model sometimes invents policy details. The firm wants answers grounded in the official HR handbook and needs to cite the source section. Which solution should they implement?

A.Use retrieval-augmented generation with Vertex AI Search over the HR handbook and instruct the model to cite retrieved sections.
B.Fine-tune the model on the HR handbook text so the knowledge is embedded in the weights.
C.Increase the context window by using a model with a larger token limit and paste the entire handbook into every prompt.
D.Lower the temperature to zero and rely on the model's pretrained knowledge of HR policies.
AnswerA

RAG retrieves the most relevant handbook passages and injects them into the prompt, so answers are based on official text. Instructing the model to cite the retrieved section provides traceability. When the handbook is updated, reindexing the source keeps answers current without retraining. This directly addresses both grounding and citation requirements.

Why this answer

Retrieval-augmented generation over the HR handbook supplies authoritative passages and enables citations, keeping answers grounded and current. Fine-tuning embeds knowledge but does not guarantee citation or easy updates. Low temperature does not add missing knowledge, and stuffing the full handbook into the prompt is inefficient and less precise than retrieval.

Exam trap

The trap here is believing that fine-tuning or a larger context window provides reliable grounding and citations, when retrieval is designed specifically for that purpose.

33
MCQeasy

A company uses a generative AI model to answer customer queries. The model sometimes returns outdated information. Which technique should they apply to ensure responses rely on current data?

A.Fine-tune the model on historical data.
B.Extend the context window to include more tokens.
C.Increase the model's temperature to encourage novelty.
D.Use grounding with a refreshed knowledge base.
AnswerD

Grounding retrieves relevant passages from a refreshed knowledge base at query time and injects them into the prompt, so answers reflect current data rather than stale parametric training. This directly satisfies the requirement that responses rely on up-to-date information.

Why this answer

Grounding with a refreshed knowledge base is the correct technique because it directly connects the generative AI model to an external, up-to-date data source at inference time. This ensures responses are based on current information without retraining the model, addressing the problem of outdated outputs by retrieving fresh data from a vector database or API in real time.

Exam trap

Google often tests the misconception that fine-tuning is the primary method to update model knowledge, when in fact grounding with a refreshed knowledge base is the correct approach for real-time data currency without retraining.

How to eliminate wrong answers

Option A is wrong because fine-tuning on historical data would embed outdated information further into the model's parameters, worsening the problem of stale responses. Option B is wrong because extending the context window only allows the model to process more tokens in a single prompt, but does not introduce new or current data; it merely expands the capacity for existing input. Option C is wrong because increasing the temperature encourages more random or creative outputs, which does not guarantee factual accuracy or recency; it can actually increase hallucinations.

34
MCQeasy

A developer is using the Gemini API to generate creative product taglines. The taglines are often bland and uncreative. The developer wants more variety and novelty in the outputs. Which parameter adjustment would most effectively increase the diversity of the generated taglines?

A.Decrease top_p from 1.0 to 0.5.
B.Set frequency_penalty to 2.0.
C.Increase temperature from 0.2 to 0.9.
D.Decrease temperature from 0.7 to 0.2.
AnswerC

Temperature scales the sampling distribution's randomness; raising it from 0.2 to 0.9 flattens the probability curve, so lower-probability tokens are selected more often. That directly increases tagline variety and novelty rather than reinforcing the bland high-probability output.

Why this answer

Increasing temperature from 0.2 to 0.9 raises the randomness of token sampling, which directly increases the diversity and novelty of generated text. A low temperature (e.g., 0.2) makes the model highly deterministic, always picking the most probable next token, leading to bland outputs. A higher temperature (e.g., 0.9) allows less probable tokens to be selected more often, producing more creative and varied taglines.

Exam trap

The trap here is that candidates often confuse temperature with top_p, incorrectly assuming that lowering top_p increases diversity, when in fact it restricts the token pool and reduces variety.

How to eliminate wrong answers

Option A is wrong because decreasing top_p from 1.0 to 0.5 reduces the nucleus of tokens considered for sampling, which actually decreases diversity by cutting off the long tail of less probable tokens. Option B is wrong because setting frequency_penalty to 2.0 penalizes token repetition too aggressively, which can suppress natural language patterns and may reduce overall output quality without directly increasing novelty. Option D is wrong because decreasing temperature from 0.7 to 0.2 makes the model more deterministic, reducing randomness and thus decreasing diversity, which is the opposite of what the developer wants.

35
MCQhard

A healthcare company is using a generative AI model to draft patient education materials. The model sometimes includes outdated medical advice. The team wants to ensure the content reflects the latest clinical guidelines. They have a database of current guidelines and want to integrate it into the generation process without retraining the model. Which approach should they use?

A.Retrieval-augmented generation (RAG) with the guidelines database as the retrieval source.
B.Increasing the model's temperature to encourage more up-to-date responses.
C.Fine-tuning the model on the latest guidelines.
D.Using a chain-of-thought prompt to ask the model to reason about the latest guidelines.
AnswerA

RAG dynamically retrieves relevant passages from the guidelines database and provides them as context to the model. This ensures the generated content is grounded in the latest clinical guidelines without modifying the model's weights. It is ideal for frequently updated information and avoids the cost of retraining.

Why this answer

Retrieval-augmented generation retrieves relevant, current information from an external database and includes it in the prompt, grounding the model's output in the latest guidelines. This approach avoids retraining and ensures the content is based on verified, up-to-date sources. Other methods like fine-tuning, temperature adjustment, or chain-of-thought do not provide the necessary external knowledge.

Exam trap

The trap here is assuming that prompt engineering or parameter tuning can inject new factual knowledge into a model without external data.

36
Multi-Selecteasy

Which TWO are advantages of using Retrieval-Augmented Generation (RAG) over fine-tuning?

Select 2 answers
A.No need to retrain the base model
B.Requires less data preparation
C.Lower inference latency
D.More secure because model weights are not modified
E.Better suited for rapidly changing knowledge bases
AnswersA, E

Correct: RAG works with the pre-trained model and a retrieval system.

Why this answer

RAG retrieves relevant external knowledge at inference time without modifying the base model's parameters, eliminating the need for retraining (option A). This contrasts with fine-tuning, which requires updating model weights through additional training cycles. Because the knowledge lives in an external, updatable index rather than in the model weights, RAG is also better suited for rapidly changing knowledge bases (option E): documents can be added, updated, or removed without any retraining, whereas fine-tuning would require a new training run each time the knowledge changes.

RAG thus preserves the original model while augmenting its output with up-to-date information.

Exam trap

Google often tests the misconception that RAG is always faster or simpler than fine-tuning, but candidates must remember that retrieval adds latency and requires careful data preprocessing, making options B and C tempting but incorrect.

37
MCQeasy

A developer is using the Gemini API to generate creative marketing copy. They want the output to be more diverse and unexpected. Which parameter should they increase?

A.Temperature.
B.Presence penalty.
C.Top-p.
D.Frequency penalty.
AnswerA

Temperature controls sampling randomness by scaling the probability distribution over candidate tokens before selection. Raising it flattens that distribution, so lower-probability tokens are chosen more often, directly producing the diverse, unexpected marketing copy the developer wants.

Why this answer

Increasing the temperature parameter makes the model's output probabilities more uniform, encouraging it to sample less likely tokens and produce more diverse, unexpected, and creative text. A higher temperature (e.g., >1.0) flattens the probability distribution, so the model is more likely to choose surprising word combinations rather than the most probable ones.

Exam trap

A common pitfall is confusing 'diversity' with 'avoiding repetition.' Increasing temperature or top-p increases randomness and diversity, while presence and frequency penalties reduce repetition. Candidates may incorrectly choose presence or frequency penalties thinking they increase diversity.

How to eliminate wrong answers

Option B (Presence penalty) is wrong because it penalizes tokens that have already appeared in the text, which reduces repetition but does not directly increase diversity or unexpectedness in the same way as temperature; it can still lead to predictable token choices. Option C (Top-p) is wrong because it controls the cumulative probability threshold for token sampling (nucleus sampling), which can limit diversity by cutting off the long tail of low-probability tokens, making outputs less unexpected. Option D (Frequency penalty) is wrong because it reduces the likelihood of tokens based on how often they have appeared, which primarily discourages repetition but does not flatten the probability distribution to encourage surprising token choices.

38
MCQhard

Refer to the exhibit. A Vertex AI endpoint configured with the above deployment is returning HTTP 429 (Too Many Requests) errors during peak traffic. The current CPU utilization reaches 80% consistently. What should the team adjust to resolve this?

A.Increase maxReplicaCount to 10
B.Increase scaleTarget to 0.9
C.Change machineType to n1-highmem-2
D.Increase minReplicaCount to 2
AnswerA

Correct: Higher max allows more replicas to handle traffic spikes.

Why this answer

Increasing maxReplicaCount to 10 allows the Vertex AI endpoint to scale out to more instances during peak traffic, distributing the load and reducing HTTP 429 errors. Since CPU utilization is at 80%, the current maxReplicaCount is insufficient to handle the demand, and raising this limit enables the horizontal pod autoscaler to add replicas up to the new maximum, directly addressing the capacity bottleneck.

Exam trap

The Google Cloud Gen AI Leader exam often tests the distinction between scaling limits (min/maxReplicaCount) and scaling thresholds (scaleTarget), trapping candidates who confuse raising the scaling target with increasing capacity.

How to eliminate wrong answers

Option B is wrong because increasing scaleTarget (the CPU utilization threshold for scaling) to 0.9 (90%) would actually delay scaling, making the 429 errors worse as the endpoint would wait until CPU is even higher before adding replicas. Option C is wrong because changing machineType to n1-highmem-2 (a memory-optimized machine) does not address the CPU bottleneck; the issue is insufficient compute capacity, not memory pressure. Option D is wrong because increasing minReplicaCount to 2 ensures a baseline of 2 replicas but does not raise the upper scaling limit, so during peak traffic the endpoint still cannot scale beyond the current maxReplicaCount, leaving it vulnerable to overload.

39
MCQmedium

A media company uses a generative AI model to draft weekly newsletter articles from bullet points. The drafts are factually correct but read as terse and disjointed. Editors want smoother narrative flow without changing the underlying facts. Which technique should they apply first to improve the output?

A.Fine-tune the model on the company's entire archive of past newsletters.
B.Increase the model's temperature setting so the model chooses less probable words.
C.Reduce the maximum output tokens so the model must compress each article.
D.Provide a few-shot prompt with two or three exemplar newsletter paragraphs demonstrating the desired narrative style.
AnswerD

Few-shot prompting supplies concrete examples that show the model the target structure, tone, and transitions. Because the facts are already correct, the examples teach style rather than content, which is exactly the gap. This is a low-cost, immediate prompt-level change that reliably improves flow without retraining or altering the factual bullet points.

Why this answer

The drafts are factually sound but stylistically weak, so the fastest effective fix is showing the model what good output looks like. Few-shot examples in the prompt communicate structure, tone, and transitions directly. Sampling changes and output-length limits affect randomness or size, not narrative quality, and fine-tuning is unnecessarily heavy for a style-only gap.

Exam trap

The trap here is assuming that any output-quality problem requires model fine-tuning, when prompt-level few-shot examples can fix a purely stylistic gap faster and at far lower cost.

40
MCQmedium

A software company is using a large language model to generate code snippets from natural language descriptions. The generated code often contains syntax errors or uses deprecated functions. The team wants to improve the correctness of the code. Which technique should they use?

A.Increasing the temperature to allow more creative code solutions.
B.Few-shot prompting with examples of correct code and explanations of why deprecated functions are avoided.
C.Using a top-p value of 1.0 to consider all possible tokens.
D.Reducing the max output tokens to force shorter code.
AnswerB

Few-shot prompting provides the model with concrete examples of correct code and reasoning, which helps it mimic the desired patterns and avoid deprecated functions. By including explanations, the model learns the rationale, improving its ability to generate syntactically correct and modern code. This is a direct and effective technique for this scenario.

Why this answer

Few-shot prompting supplies the model with examples of correct code and explanations, which guides it to generate syntactically valid and up-to-date code. This technique leverages in-context learning to improve accuracy without retraining. Other parameter adjustments like temperature, max output tokens, or top-p do not provide the necessary guidance for correctness.

Exam trap

The trap here is thinking that adjusting sampling parameters like temperature can fix systematic issues like syntax errors or deprecated API usage.

41
MCQhard

A financial services firm uses a generative AI model to summarize quarterly earnings calls. The summaries must consistently follow a strict format: an executive summary, key financial metrics, and forward-looking statements. The model sometimes omits sections or changes the order. Which technique should they use to enforce the structure?

A.Use few-shot prompting with examples that demonstrate the exact desired format.
B.Increase the model's temperature to encourage more structured output.
C.Use controlled generation with a response schema (e.g., JSON schema) to constrain the model's output.
D.Reduce the max output tokens to force the model to include all sections.
AnswerC

Controlled generation with a response schema forces the model to produce output that conforms to a defined structure, such as specific fields and their order. This is the most reliable way to enforce sections like executive summary, metrics, and forward-looking statements. It eliminates omission and reordering by constraining decoding to valid schema-compliant tokens.

Why this answer

Controlled generation with a response schema constrains the decoding process so that the model can only produce tokens that satisfy a predefined structure. This guarantees that each required section appears and in the correct order, which is essential for consistent financial summaries. Unlike few-shot examples, schema-based control is deterministic and enforceable at the API level, making it the most reliable choice for strict formatting.

Exam trap

The trap here is assuming that few-shot examples or length limits can guarantee a strict structure, when only schema-based constrained decoding enforces it deterministically.

42
MCQmedium

A developer uses a code generation model to write Python functions. The output frequently contains syntax errors due to incorrect braces and indentation. Which technique should be used to produce syntactically valid code?

A.Increase the temperature to introduce more varied token choices.
B.Apply constrained decoding techniques that enforce a grammar for the target programming language.
C.Fine-tune the model on a large corpus of syntactically correct Python code.
D.Provide a few-shot example of correct Python function in the prompt.
AnswerB

Constrained decoding masks tokens that would violate a Python grammar, so braces and indentation are enforced at generation time rather than corrected afterwards. This directly satisfies the stem's requirement for syntactically valid output, unlike prompt engineering or post-hoc linting, which cannot guarantee validity during sampling.

Why this answer

Constrained decoding (also called grammar-guided generation) enforces the syntax rules of the target language (e.g., Python) during token generation by restricting the model's output to only valid tokens according to a formal grammar (e.g., EBNF or context-free grammar). This directly prevents syntax errors like incorrect braces or indentation, which are structural, not semantic, issues. Techniques such as using a parser-based logit processor or a constrained beam search ensure every generated token sequence is syntactically valid.

Exam trap

A common misconception is that fine-tuning or prompt engineering alone can guarantee syntactic correctness, but only constrained decoding (or grammar-guided generation) provides a hard guarantee against syntax errors by actively restricting the output space to valid tokens per the language's grammar.

How to eliminate wrong answers

Option A is wrong because increasing temperature adds randomness to token selection, which would likely increase syntax errors rather than reduce them, as it encourages less probable (and often malformed) token sequences. Option C is wrong because fine-tuning on syntactically correct code does not guarantee that the model will never produce syntax errors during generation; it only improves the statistical likelihood of correctness, but the model can still deviate from grammar rules, especially with novel or complex prompts. Option D is wrong because few-shot prompting provides examples but does not enforce structural constraints; the model may still produce syntax errors if it fails to generalize the pattern or if the prompt is ambiguous.

43
MCQhard

A real-time customer support chatbot using Gemini is experiencing high latency. The team must maintain response quality while improving speed. Which technique should they implement?

A.Switch to a larger model
B.Increase the batch size
C.Use context caching for frequent queries
D.Decrease the temperature
AnswerC

Context caching stores precomputed attention states for repeated prompt prefixes, so Gemini skips reprocessing them on each call. For a support chatbot with frequent recurring queries, this cuts time-to-first-token while preserving the full model and response quality, directly addressing the latency constraint without downgrading the model.

Why this answer

Context caching reduces latency by storing frequently accessed query responses or intermediate computations, allowing the chatbot to reuse precomputed results instead of reprocessing identical or similar requests through the full model pipeline. This directly addresses high latency in real-time systems while preserving response quality, as the cached outputs are identical to freshly generated ones for the same input.

Exam trap

Google often tests the misconception that latency improvements come from model parameter tuning (like temperature) or throughput adjustments (like batch size), when the real solution for real-time systems is architectural optimization like caching.

How to eliminate wrong answers

Option A is wrong because switching to a larger model increases computational overhead and latency, worsening the problem rather than solving it. Option B is wrong because increasing batch size improves throughput for offline processing but does not reduce per-request latency in a real-time streaming chatbot; it may even increase the time to first token. Option D is wrong because decreasing temperature alters output randomness and creativity but has no meaningful impact on inference speed or latency.

44
Multi-Selectmedium

A media company uses Gemini on Vertex AI to generate short news summaries from long articles. The summaries frequently miss key facts and sometimes include details not in the source. The team wants to improve factual grounding and coverage without retraining the model. (Choose two.)

Select 2 answers
A.Use retrieval-augmented generation (RAG) with Vertex AI Search to supply relevant source passages in the prompt.
B.Enable streaming responses so the model can revise earlier sentences as it generates.
C.Deploy the model on a larger machine type with more vCPUs to increase factual accuracy.
D.Add few-shot examples in the prompt that show correctly grounded summaries with citations to the source.
E.Increase the model temperature to 1.5 so the model explores more of the source content.
AnswersA, D

RAG retrieves authoritative passages from the article or an indexed corpus and injects them into the model context, so the summary is conditioned on real source text. This directly reduces hallucination and improves fact coverage without fine-tuning. It also lets you update the knowledge base independently of the model.

Why this answer

Retrieval-augmented generation supplies authoritative source passages at inference time, and few-shot examples teach the model the desired grounded summarization pattern. Together they improve fact coverage and reduce unsupported claims without retraining. Temperature tuning, larger machines, and streaming do not add source knowledge or enforce faithfulness, so they fail to solve the stated problem.

Exam trap

The trap here is assuming that infrastructure scaling or streaming features improve factual accuracy, when grounding requires better context and prompting rather than more compute.

45
MCQeasy

A company uses a text-to-image model to generate marketing visuals. The outputs often contain distorted human faces. Which technique is most likely to improve face generation?

A.Fine-tune the model on a curated dataset of human faces
B.Increase the output resolution
C.Increase the number of inference steps
D.Reduce the classifier-free guidance scale
AnswerA

Fine-tuning adjusts the model's weights on curated face data, teaching it the specific facial structures and proportions it currently renders poorly. This directly targets the distorted-face failure mode, unlike prompt engineering or higher resolution, which cannot correct learned representation gaps.

Why this answer

Fine-tuning the model on a high-quality dataset of human faces directly addresses the distortion issue by specializing the model for face generation. Option B (increasing output resolution) may improve overall image sharpness but does not specifically correct face distortions. Option C (increasing inference steps) can enhance image coherence but is not targeted at face quality.

Option D (reducing classifier-free guidance scale) decreases prompt adherence, which could actually worsen face generation rather than improve it.

46
MCQhard

A generative AI model for code generation sometimes produces syntactically incorrect code. The team wants to reduce syntax errors without retraining the entire model. Which approach is most effective?

A.Implement constrained decoding with grammar rules
B.Run a syntax checker after generation and regenerate
C.Add a system prompt that instructs the model to produce valid code
D.Increase beam search width
AnswerA

Constrained decoding masks tokens that would violate the target language's grammar at each generation step, so syntactically invalid sequences are never emitted. This directly reduces syntax errors at inference time, satisfying the stem's constraint of no retraining. Unlike fine-tuning, it requires no weight updates, only a grammar specification during decoding.

Why this answer

Constrained decoding with grammar rules directly enforces the syntax of the target programming language during token generation, preventing the model from producing invalid constructs. This approach modifies the decoding process (e.g., using a context-free grammar or a formal syntax specification) to mask or forbid tokens that would lead to a syntax error, without altering the underlying model weights. It is the most effective method because it guarantees syntactically correct output at generation time, rather than relying on post-hoc fixes or probabilistic adjustments.

Exam trap

The trap here is that candidates often choose a post-hoc correction method (Option B) or a prompt-based approach (Option C) because they seem simpler, but they fail to recognize that only a decoding-time constraint can guarantee syntactic validity without retraining, which is the core requirement of the question.

How to eliminate wrong answers

Option B is wrong because running a syntax checker after generation and regenerating is inefficient and does not prevent errors; it relies on trial-and-error, which can be costly and may still produce invalid code if the model repeatedly generates similar errors. Option C is wrong because adding a system prompt is a soft instruction that the model may not reliably follow, especially for complex or edge-case syntax rules, and it does not enforce constraints at the token level. Option D is wrong because increasing beam search width improves the diversity and likelihood of finding high-probability sequences but does not incorporate any syntactic constraints; it may still produce syntactically incorrect code if the highest-scoring beams violate grammar rules.

47
MCQeasy

A company is using a generative AI model to generate product descriptions. They notice the outputs often include factual inaccuracies about product specifications. Which technique would best address this issue without modifying the model's architecture?

A.Implement a Retrieval-Augmented Generation (RAG) pipeline that retrieves product specs from a database
B.Decrease the temperature parameter to 0.1
C.Increase the max output tokens to 1024
D.Use few-shot prompting with 5 examples of correct descriptions
AnswerA

RAG grounds generation in retrieved product specifications from the database, so outputs cite accurate facts rather than relying on parametric memory. This corrects inaccuracies without retraining or altering the model's architecture, satisfying the stem's constraint of no architectural modification.

Why this answer

Retrieval-Augmented Generation (RAG) is the correct technique because it grounds the model's output in factual, up-to-date product specifications retrieved from an external database. This directly addresses factual inaccuracies without modifying the model's architecture, as the model generates text based on retrieved context rather than relying solely on its parametric knowledge.

Exam trap

Google Cloud often tests the misconception that adjusting generation parameters (like temperature or token limits) or providing examples can fix factual accuracy, when in fact only retrieval-augmented methods or fine-tuning on verified data can correct hallucinations without changing the model architecture.

How to eliminate wrong answers

Option B is wrong because decreasing the temperature parameter to 0.1 makes the model more deterministic and reduces randomness, but it does not provide any factual grounding; it can still hallucinate incorrect specifications. Option C is wrong because increasing max output tokens only allows longer generations and does not improve factual accuracy; it may even increase the chance of errors. Option D is wrong because few-shot prompting with examples can guide the style and format but cannot supply specific, dynamic product specs; the model may still invent details not present in the examples.

48
MCQmedium

Refer to the exhibit. The endpoint is experiencing high latency during traffic spikes. The team wants to improve response time by reducing queueing. Which change to the configuration would be most effective?

A.Decrease minReplicaCount to 0
B.Change the model version to '2'
C.Decrease the target value in autoscaling metric to 50
D.Increase maxReplicaCount to 10
AnswerD

Raising maxReplicaCount to 10 lets the autoscaler add more pod instances during traffic spikes, distributing requests across additional workers and cutting queue depth. More concurrent replicas directly reduce queueing latency, which is the stated constraint driving the response-time goal.

Why this answer

Increasing maxReplicaCount to 10 allows the autoscaler to provision more replicas during traffic spikes, distributing the incoming requests across additional endpoints. This directly reduces queueing at each replica because the load is spread over more instances, lowering per-instance latency. The change targets the root cause—insufficient capacity to handle peak load—rather than adjusting thresholds or model versions.

Exam trap

Google Cloud often tests the misconception that lowering the autoscaling target metric (Option C) is the primary fix for high latency, when in fact the maxReplicaCount ceiling is the bottleneck that must be raised to allow sufficient capacity during spikes.

How to eliminate wrong answers

Option A is wrong because decreasing minReplicaCount to 0 would cause the endpoint to scale down to zero replicas during idle periods, leading to cold starts and increased latency when traffic spikes, which worsens queueing. Option B is wrong because changing the model version to '2' does not affect the number of replicas or queueing behavior; it only changes the model artifact, which may have different inference latency but does not address scaling capacity. Option C is wrong because decreasing the target value in the autoscaling metric (e.g., CPU utilization or requests per replica) would cause the autoscaler to add replicas sooner, but without increasing maxReplicaCount, the endpoint may still hit the upper limit and queue requests; the target value adjustment alone does not provide additional capacity during extreme spikes.

49
MCQhard

A law firm uses a generative model to analyze contracts and extract key clauses. The model often outputs irrelevant clauses or misses important ones. They want to improve the relevance of the outputs without retraining the entire model. Which approach is best?

A.Increase the input token limit to provide the entire contract in the prompt.
B.Decrease the temperature to make outputs more deterministic.
C.Implement Retrieval-Augmented Generation (RAG) with a curated legal clause database and a reranker to select the most on-topic passages.
D.Fine-tune the base model on a labeled dataset of contract-clause pairs.
AnswerC

RAG retrieves passages from a curated legal clause database at inference time and a reranker scores them for topical relevance, grounding generation in authoritative clauses. This improves precision and recall of extracted clauses without retraining, satisfying the constraint of no full model retraining.

Why this answer

Retrieval-Augmented Generation (RAG) with a curated legal clause database and a reranker is the best approach because it grounds the model's outputs in a trusted, external knowledge base, ensuring that only the most relevant clauses are retrieved and used for generation. This directly addresses the problem of irrelevant outputs and missed clauses without requiring retraining, as the model can dynamically fetch and rank the most on-topic passages from the curated database.

Exam trap

A common trap in Google's Generative AI exams is assuming that adjusting parameters like temperature or token limits can fix output relevance issues. While these parameters affect randomness and context length, they do not guarantee that the model will generate accurate or relevant content for domain-specific tasks. The correct approach is to use a retrieval system like RAG to ground the model in a curated knowledge base.

How to eliminate wrong answers

Option A is wrong because simply increasing the input token limit to provide the entire contract in the prompt does not improve the model's ability to focus on relevant clauses; it may actually dilute the signal with more noise and exceed context windows, leading to worse performance. Option B is wrong because decreasing the temperature makes outputs more deterministic but does not change the underlying knowledge or relevance of the generated clauses; it only reduces randomness, not the model's tendency to hallucinate or miss key information. Option D is wrong because fine-tuning the base model on a labeled dataset of contract-clause pairs requires retraining, which is explicitly ruled out by the question's constraint of improving outputs 'without retraining the entire model'.

50
MCQeasy

A marketing team is using a text generation model to create ad copy. They notice that the model's output is often bland and lacks creativity. They want to increase the diversity of the generated text while keeping it relevant. Which parameter should they adjust?

A.Top-k
B.Frequency penalty
C.Max output tokens
D.Temperature
AnswerD

Temperature controls the randomness of token selection. A higher temperature makes the model more likely to choose less probable tokens, leading to more diverse and creative outputs. This directly addresses the blandness by encouraging variety, while still being influenced by the prompt for relevance.

Why this answer

Temperature is the primary parameter that controls randomness in generation. Increasing it allows the model to explore less likely token sequences, resulting in more varied and creative outputs. Other parameters like max output tokens, top-k, and frequency penalty address different aspects such as length, token pool restriction, and repetition, respectively.

Exam trap

The trap here is confusing parameters that control repetition or length with those that control randomness and creativity.

51
Multi-Selectmedium

A financial services firm is deploying a generative AI model to answer customer questions about investment products. They want to ensure the model's responses are compliant with regulations and do not provide personalized financial advice. Which TWO techniques can help achieve this? (Choose two.)

Select 2 answers
A.Deploy the model with a lower token limit to keep responses short and less detailed.
B.Use a content filter to block any response that contains words like 'recommend' or 'should invest'.
C.Increase the model's temperature to make responses more varied and less likely to give direct advice.
D.Fine-tune the model on a dataset of compliant customer interactions.
E.Implement a system prompt that explicitly instructs the model to avoid giving personalized advice and to include a disclaimer.
AnswersD, E

Fine-tuning on examples of compliant responses can teach the model the patterns of safe, regulation-abiding language. This helps the model internalize the desired behavior, reducing the need for extensive prompt engineering. It is particularly useful when the compliance rules are complex and require nuanced understanding.

Why this answer

Combining a system prompt with explicit instructions and fine-tuning on compliant interactions provides both immediate guidance and long-term behavioral shaping. The system prompt ensures the model follows rules in real time, while fine-tuning helps the model internalize compliance patterns, making responses more reliably safe. Together, they address the need for regulatory adherence and avoidance of personalized advice.

Exam trap

The trap here is relying on superficial filters or parameter tweaks, which do not reliably enforce nuanced compliance requirements.

52
MCQhard

You are a Generative AI architect at a large financial services firm. The firm has deployed a custom large language model (LLM) fine-tuned on proprietary financial reports to assist analysts in generating quarterly earnings summaries. The model is hosted on Vertex AI using a dedicated endpoint with autoscaling enabled. Recently, the model's output has exhibited two issues: (1) occasional factual inaccuracies about specific financial figures, and (2) a tendency to produce overly verbose and repetitive text in the summaries, sometimes exceeding the desired length of 200 words. The team has already tried adjusting the temperature parameter from 0.7 to 0.2 and increased the top-k sampling from 40 to 50, but the problems persist. The model's training data includes over 10,000 financial reports, and the fine-tuning process used low-rank adaptation (LoRA) with rank 16. The production environment uses a batch size of 1 for inference. You need to recommend a course of action that most directly addresses both the factual accuracy and verbosity issues without requiring a full retraining of the model. Which approach should you take?

A.Increase the LoRA rank to 32 and fine-tune the model for additional epochs on a curated subset of reports that focus on concise and accurate summaries.
B.Implement a retrieval-augmented generation (RAG) pipeline that queries a vector database of verified financial data, and apply constrained decoding with a maximum token limit and a repetition penalty.
C.Switch to a larger pre-trained model (e.g., PaLM 2 or GPT-4) and use the same fine-tuning data with higher rank LoRA to improve capability, then rely on the larger model's inherent accuracy.
D.Experiment with higher temperature (e.g., 0.9) and lower top-k (e.g., 20) to encourage more diverse and concise outputs, and add a post-processing step to truncate summaries to 200 words.
AnswerB

RAG grounds generation in verified figures retrieved from a vector database, correcting factual inaccuracies without retraining, while constrained decoding enforces the 200-word ceiling and penalises repetition, directly resolving verbosity. Temperature and top-k tuning cannot fix hallucinated figures or enforce length limits.

Why this answer

It directly addresses both issues without retraining. A RAG pipeline grounds the model's outputs in verified financial data, eliminating factual inaccuracies. Constrained decoding with a maximum token limit and repetition penalty directly curbs verbosity and repetition, which temperature and top-k adjustments failed to fix.

Exam trap

Google Cloud often tests the misconception that adjusting hyperparameters like temperature or top-k can fix factual accuracy and verbosity, when in reality these issues stem from the model's lack of external knowledge and lack of output constraints, which require architectural changes like RAG and constrained decoding.

How to eliminate wrong answers

Option A is wrong because increasing LoRA rank and fine-tuning on a curated subset still relies on the model's parametric memory, which is prone to hallucination and does not guarantee factual accuracy; it also requires retraining, contradicting the 'no full retraining' constraint. Option C is wrong because switching to a larger model does not inherently solve factual inaccuracies (larger models can still hallucinate) and requires full retraining or significant adaptation, violating the constraint. Option D is wrong because higher temperature (0.9) increases randomness, likely worsening factual inaccuracies, and lower top-k (20) reduces diversity, which may not fix verbosity; post-processing truncation does not address the root cause of repetition or inaccuracy.

53
Multi-Selectmedium

Which TWO techniques effectively reduce bias in generative model outputs? (Choose two.)

Select 2 answers
A.Apply adversarial debiasing during training or fine-tuning.
B.Increase the temperature parameter to introduce more variability.
C.Use a larger model with more parameters.
D.Fine-tune on a dataset with balanced representation across groups.
E.Reduce max output tokens to limit the model's expression.
AnswersA, D

Adversarial debiasing trains a discriminator to detect protected-attribute predictions from the model's representations while the generator learns to defeat it, directly penalising biased internal representations during training or fine-tuning rather than only filtering outputs afterwards.

Why this answer

Adversarial debiasing (A) directly reduces bias by training the model to minimize an adversary's ability to predict protected attributes from the model's outputs, forcing the model to learn representations that are invariant to those attributes. Fine-tuning on a balanced dataset (D) corrects representation bias by ensuring the model sees equal examples across groups, preventing overfitting to majority patterns. Both techniques actively address the root causes of bias in training data or model behavior.

Exam trap

A common misconception is that randomness (temperature) or model size alone can fix bias, when in fact these parameters do not address the systematic skew in training data or model representations.

54
MCQmedium

A team is using a pre-trained language model to summarize legal documents. They find that summaries often miss key dates and parties involved. Which technique would most effectively improve factual accuracy?

A.Fine-tune the model on a dataset of legal summaries with annotated key entities.
B.Use top-p sampling with a low p value.
C.Increase the temperature parameter.
D.Use chain-of-thought prompting.
AnswerA

Fine-tuning on annotated legal summaries teaches the model domain-specific entity patterns — dates and party names — adjusting weights to prioritise those tokens during generation. This directly targets the factual omissions, unlike prompt engineering or retrieval alone.

Why this answer

Fine-tuning on a dataset of legal summaries with annotated key entities directly teaches the model to recognize and reproduce critical factual elements like dates and parties. This supervised learning approach adjusts the model's weights to prioritize entity extraction and accurate generation, which is the most effective method for improving factual accuracy in domain-specific tasks.

Exam trap

Google Cloud often tests the misconception that inference-time parameters (temperature, top-p) or prompting strategies can substitute for targeted training, when in fact only fine-tuning with domain-specific annotated data reliably improves factual accuracy for structured entities.

How to eliminate wrong answers

Option B is wrong because top-p sampling with a low p value restricts the vocabulary to a small set of high-probability tokens, which can reduce creativity but does not address factual accuracy or entity recall—it may even omit rare but important entities. Option C is wrong because increasing the temperature parameter adds randomness to token selection, which typically reduces factual consistency and can lead to hallucinated or missing details. Option D is wrong because chain-of-thought prompting improves reasoning steps for multi-step tasks but does not inherently enforce factual accuracy for specific entities; it relies on the model's existing knowledge, which may still miss key dates and parties without targeted training.

55
MCQmedium

A team is building a generative AI model for customer support. They notice the model often produces overly polite but unhelpful responses. Which technique would best improve response quality without sacrificing helpfulness?

A.Apply reinforcement learning from human feedback (RLHF)
B.Increase the amount of training data
C.Lower the top_k sampling value
D.Increase the temperature parameter
AnswerA

Reinforcement learning from human feedback directly optimises the model against human preference rankings, aligning outputs with helpfulness rather than politeness alone. This satisfies the stem's constraint of improving response quality without sacrificing helpfulness, since reward signals penalise unhelpful replies and reinforce substantive ones during fine-tuning.

Why this answer

RLHF directly addresses the misalignment between the model's training objective (e.g., predicting the next token) and the desired outcome (helpful, not just polite). By using human feedback to train a reward model, the system learns to optimize for response quality and helpfulness, reducing sycophantic or overly polite but uninformative outputs.

Exam trap

Google Cloud often tests the misconception that hyperparameter tuning (temperature, top_k) or more data alone can fix alignment issues, when in fact only RLHF directly optimizes for human-judged helpfulness and quality.

How to eliminate wrong answers

Option B is wrong because simply increasing training data does not correct the model's tendency toward polite but unhelpful responses; it may reinforce existing patterns without addressing alignment. Option C is wrong because lowering top_k sampling reduces diversity by restricting token choices to the top k most likely tokens, which can make responses even more generic and less helpful, not more substantive. Option D is wrong because increasing the temperature parameter increases randomness in token selection, which can lead to less coherent or more erratic responses, not more helpful ones.

56
Multi-Selecteasy

Which TWO techniques are commonly used to control the style and tone of a generative model's output?

Select 2 answers
A.Adjusting the temperature
B.Modifying the top_k value
C.Fine-tuning on a dataset with desired style
D.Prompt engineering with style instructions
E.Changing the top_p value
AnswersC, D

Fine-tuning adapts the model to a specific style.

Why this answer

Fine-tuning on a dataset that embodies the desired style directly adjusts the model's weights, making it consistently produce outputs with that specific tone and style. This is a fundamental technique for customizing generative models, as it teaches the model the exact patterns, vocabulary, and stylistic nuances present in the training data.

Exam trap

Google Cloud often tests the distinction between sampling parameters (temperature, top_k, top_p) that control output randomness and diversity versus training or conditioning techniques (fine-tuning, prompt engineering) that directly influence style and tone, leading candidates to incorrectly select sampling parameters as style-control methods.

57
Multi-Selectmedium

Which TWO techniques are most effective for improving the quality of a generative AI model's output when summarizing complex documents?

Select 2 answers
A.Providing few-shot examples of ideal summaries
B.Using a larger, more capable model (e.g., PaLM 2 instead of PaLM)
C.Increasing max output length significantly
D.Setting top_p to 0.1
E.Adjusting temperature to 0.8
AnswersA, B

Few-shot examples steer the model by conditioning it on concrete input-output pairs, so it infers the desired summary structure, length and tone directly from the prompt. This satisfies the stem's demand for higher output quality on complex documents without retraining, unlike prompt-only instructions that leave format ambiguous.

Why this answer

Option A is correct because few-shot prompting supplies the model with concrete examples of the desired summary format, style, and level of detail, which conditions the model to reproduce that structure and improves fidelity when summarizing complex documents. Option B is correct because a larger, more capable model such as PaLM 2 has greater capacity to comprehend long, intricate documents and generate more accurate, coherent summaries than a smaller predecessor like PaLM. Option C is not appropriate because increasing max output length only allows longer responses; it does not improve summary quality and can even encourage verbosity rather than conciseness.

Option D is not appropriate because setting top_p to 0.1 sharply narrows nucleus sampling, reducing diversity and making the output more rigid without addressing summarization quality. Option E is not appropriate because raising temperature to 0.8 increases randomness, which tends to introduce inaccuracies and inconsistency rather than improve the quality of document summaries.

Exam trap

A common misconception is that increasing output length or adjusting sampling parameters like top_p and temperature universally improves output quality, when in fact these parameters must be tuned carefully for the specific task and can degrade summary quality if misapplied.

58
MCQeasy

A marketing team uses a text generation model to create ad copy. They want the output to be more diverse and creative, exploring unusual angles. Which parameter should they adjust?

A.Top-k
B.Max output tokens
C.Temperature
D.Top-p (nucleus sampling)
AnswerC

Temperature scales the logits before softmax, flattening the distribution at higher values. A higher temperature makes low-probability tokens more likely to be selected, producing more varied and creative text. This directly increases diversity and is the standard parameter for encouraging novel angles in ad copy while still remaining coherent.

Why this answer

Temperature is the parameter that directly controls randomness in token selection. Increasing it flattens the probability distribution, making less likely words and ideas more probable, which yields more diverse and creative ad copy. While top-p and top-k also affect sampling, temperature is the primary and most intuitive knob for encouraging novel angles without changing other constraints.

Exam trap

The trap here is confusing length controls or hard sampling cutoffs with the smooth randomness control that temperature provides for creativity.

59
MCQeasy

Refer to the exhibit. A user wants formal translations from a generative AI model, but the model outputs informal style inconsistently. Which prompt engineering technique would best ensure consistent formal translations?

A.Use context caching
B.Provide a few-shot example with formal and informal pairs
C.Use a longer system prompt with detailed rules
D.Set top_k to 1
AnswerB

Few-shot prompting supplies paired formal and informal examples, letting the model infer the required register from demonstrated patterns rather than vague instructions. This constrains output style consistently, satisfying the requirement for reliably formal translations across varied inputs.

Why this answer

Few-shot prompting provides the model with concrete input/output examples that demonstrate the desired style, which is far more effective at controlling tone and register than abstract instructions. By including formal and informal pairs, the model can infer the transformation pattern and apply it consistently to new translations. This anchors the model's behavior through in-context learning rather than relying on the model to interpret vague stylistic rules.

Exam trap

The trap here is confusing decoding parameters (top_k, temperature) or prompt length with actual style control — candidates assume 'more rules' or 'deterministic sampling' fixes tone, when only concrete demonstrations reliably steer style.

How to eliminate wrong answers

Option A is wrong because context caching only stores and reuses previously processed context to reduce latency and cost; it does not influence the model's stylistic output. Option C is wrong because a longer system prompt with detailed rules still relies on the model's interpretation of abstract instructions, which often produces inconsistent style adherence compared to concrete examples. Option D is wrong because setting top_k to 1 only makes the token sampling deterministic (always picking the highest-probability token); it does not teach the model what 'formal' means and can even amplify a consistently wrong style.

60
MCQmedium

A company uses a text-to-image model to generate marketing visuals. The results often misinterpret the prompt, e.g., 'a red car' generates a blue car. Which technique should they try first to align the output with the prompt?

A.Use a negative prompt to exclude blue
B.Refine the prompt with more adjectives and context, e.g., 'bright red sports car'
C.Upscale the image resolution to 1024x1024
D.Increase the guidance scale to 20
AnswerB

Text-to-image models weight prompt tokens probabilistically, so vague phrasing leaves colour underspecified. Adding explicit descriptors such as 'bright red sports car' increases the token weight for red, steering sampling toward the intended output. This prompt-refinement step is the fastest fix before considering fine-tuning or negative prompts.

Why this answer

Refining the prompt with more adjectives and context directly addresses the root cause of misalignment: insufficient specificity in the text description. Text-to-image models rely on the semantic richness of the prompt to guide the latent diffusion process; adding 'bright red sports car' provides stronger conditioning signals that steer the model's cross-attention layers toward the intended color and object attributes. This is the most efficient first step before adjusting hyperparameters like guidance scale.

Exam trap

The trap here is that candidates often jump to hyperparameter tuning (guidance scale) or post-processing (upscaling) as a first fix, when the most fundamental and cost-effective step is to improve the input prompt's specificity, which directly controls the conditioning signal in the diffusion process.

How to eliminate wrong answers

Option A is wrong because using a negative prompt to exclude 'blue' is a reactive band-aid that does not fix the core issue of the model failing to associate 'red' with the car; it also risks suppressing other unintended features and can degrade image quality by over-constraining the latent space. Option C is wrong because upscaling resolution to 1024x1024 only increases pixel density and does not alter the semantic alignment between the prompt and the generated image; the model's misinterpretation of 'red' would persist at any resolution. Option D is wrong because increasing the guidance scale to 20 excessively amplifies the prompt's influence, often leading to image saturation, artifacts, and mode collapse, while still not correcting the fundamental misassociation of the color attribute.

61
MCQmedium

A media company is using Vertex AI's Imagen model to generate images for marketing campaigns. They have a set of prompts that describe desired scenes, but the generated images often contain artifacts such as distorted faces or unnatural lighting. The team has tried varying the prompt wording but the issues persist. They are using the default parameters (no modifications). They have a budget for additional compute resources and want to improve image quality without switching to a more expensive model. The team has access to a small set of high-quality images in the same style as their target outputs. What should the team do?

A.Increase the guidance scale parameter to make the model follow prompts more closely.
B.Use a more detailed prompt style with negative prompts to avoid artifacts.
C.Fine-tune the Imagen model on the small set of high-quality images to improve output quality.
D.Increase the number of images generated per prompt and manually select the best ones.
AnswerC

Fine-tuning adapts the model to produce images with fewer artifacts and desired style.

Why this answer

Fine-tuning the Imagen model on the small set of high-quality images allows the model to learn the desired style and reduce artifacts like distorted faces and unnatural lighting, improving output quality without switching to a more expensive model. Option A is incorrect because increasing the guidance scale may cause the model to overfit to the prompt and potentially introduce more artifacts rather than fix them, and the team already has issues with prompt adherence. Option B is incorrect because using more detailed prompts with negative prompts might help but the team already tried varying wording without success; the root cause is the model's lack of specific training on the desired quality, which fine-tuning directly addresses.

Option D is incorrect because generating more images per prompt does not improve the per-image quality; it only increases the chance of finding a good one, and the team wants to improve overall image quality.

62
MCQmedium

A marketing team uses a generative AI model on Vertex AI to create ad copy variations. The outputs are often too random and include off-brand phrases. They want to reduce randomness while still allowing some creativity. Which parameter should they adjust?

A.Top-P
B.Top-K
C.Max output tokens
D.Temperature
AnswerD

Temperature controls the randomness of predictions. Lowering it makes the output more deterministic and focused, reducing off-brand phrases while still allowing some variation if set moderately. This directly addresses the need to balance creativity and brand adherence.

Why this answer

Temperature is the primary parameter that scales the randomness of token selection. Lowering temperature makes the model more likely to choose high-probability tokens, resulting in more predictable and brand-safe output. While top-K and top-P also influence diversity, they are secondary to temperature for controlling overall randomness.

Exam trap

The trap here is confusing temperature with top-K or top-P, which also affect diversity but do not directly control the overall randomness scale.

63
MCQhard

A retail analytics team uses a Gemini model to answer questions over a large product catalog stored in BigQuery. Answers are sometimes outdated because the model relies on its training data. They want responses to reflect the latest catalog rows and cite the source table. Which technique should they implement?

A.Increase the model's context window and paste the entire product catalog into every prompt so the model always has all data.
B.Fine-tune the model on a recent export of the product catalog so the weights contain current product data.
C.Enable a higher temperature so the model can combine its training knowledge with newer patterns and produce fresher answers.
D.Use retrieval-augmented generation by querying BigQuery and inserting the retrieved rows into the prompt as grounding context.
AnswerD

Retrieval-augmented generation fetches current rows at query time and passes them to the model as context, so answers reflect the live catalog. Because the retrieved table and rows can be named in the prompt or returned alongside the answer, citations are possible. This solves both freshness and provenance without retraining.

Why this answer

Retrieval-augmented generation grounds the model in data fetched at request time, which keeps answers current and allows the source table to be cited. Fine-tuning freezes a stale snapshot, stuffing the whole catalog is impractical, and temperature changes only affect randomness. When answers must track a live system of record, retrieval is the appropriate pattern.

Exam trap

The trap here is thinking that a larger context window or fine-tuning can substitute for live retrieval, when only query-time retrieval guarantees current data and citations.

64
MCQmedium

A team is building a customer support assistant on Vertex AI using a foundation model. They notice the model occasionally invents policy details that don't exist in their internal documentation. They want the model to ground its answers in the company's knowledge base and provide citations. Which approach should they use?

A.Reduce the max output tokens to limit the length of the model's responses.
B.Increase the temperature parameter to encourage more creative responses.
C.Fine-tune the model on a dataset of general customer service conversations.
D.Use Retrieval-Augmented Generation (RAG) with Vertex AI Search to retrieve relevant documents and include them in the prompt.
AnswerD

RAG retrieves relevant passages from the company's knowledge base and injects them into the model's context, so the model bases its response on actual documentation. Vertex AI Search can also return citations pointing to the source documents. This directly addresses hallucination of policy details by grounding generation in verified internal content.

Why this answer

Retrieval-Augmented Generation retrieves relevant passages from an authoritative knowledge base and supplies them as context, so the model's answer is conditioned on real policy text rather than parametric memory. With Vertex AI Search, the system can also surface citations, letting users verify sources. This combination directly reduces fabricated policy details while preserving natural language fluency.

Exam trap

The trap here is assuming that fine-tuning or temperature adjustment can ground a model in a private, frequently updated knowledge base, when only retrieval at inference time provides that grounding.

65
MCQhard

An enterprise uses a fine-tuned PaLM 2 model for code generation. They want to ensure the generated code passes security audits. Which combination of techniques would be most effective?

A.Integrate a static analysis tool in the pipeline and add a safety filter to reject code containing dangerous functions.
B.Use a few-shot prompt with examples of secure code and set temperature to 1.0.
C.Fine-tune the model on a dataset of insecure code and use top-p=0.9.
D.Increase the model's context window and use a system instruction to 'be secure'.
AnswerA

Static analysis scans generated code for dangerous constructs before merge, while a safety filter rejects outputs containing risky functions at generation time. Together they satisfy the security audit requirement by catching vulnerabilities at two stages rather than relying on prompt engineering alone.

Why this answer

Integrating a static analysis tool (e.g., SonarQube, Checkmarx) into the pipeline provides automated, rule-based scanning for security vulnerabilities like SQL injection or buffer overflows, while a safety filter explicitly blocks generated code containing dangerous functions (e.g., eval(), exec()). This combination creates a defense-in-depth approach that catches both known vulnerability patterns and explicitly prohibited operations, which is essential for passing security audits.

Exam trap

The Generative AI Leader exam often tests the misconception that prompt engineering alone (e.g., system instructions or few-shot examples) is sufficient for security, when in fact deterministic validation and filtering techniques are required to enforce constraints reliably.

How to eliminate wrong answers

Option B is wrong because using a few-shot prompt with secure code examples does not guarantee the model will consistently avoid generating insecure code—temperature=1.0 increases randomness, making the output less deterministic and more likely to deviate from the secure examples. Option C is wrong because fine-tuning on a dataset of insecure code would teach the model to generate vulnerable patterns, which is counterproductive for security; top-p=0.9 does not prevent the model from outputting those learned insecure constructs. Option D is wrong because increasing the context window and using a system instruction to 'be secure' provides no enforcement mechanism—the model can still generate insecure code if the instruction is not followed, and there is no validation step to catch violations.

66
MCQmedium

A marketing team uses a generative AI model to create short social media posts from long product briefs. The posts frequently include made-up statistics and product claims that are not in the brief. The team wants to reduce these unsupported statements without retraining the model. Which technique should they apply first?

A.Reduce the maximum output token limit so the model cannot include extra details.
B.Add grounding instructions that require the model to use only facts from the provided brief and to say when information is missing.
C.Increase the model's temperature setting to encourage more creative output.
D.Fine-tune the model on a dataset of approved social media posts.
AnswerB

Grounding instructions in the prompt explicitly constrain the model to the supplied brief and tell it how to behave when a detail is absent. This directly targets fabricated statistics and claims without retraining. It is a low-cost, immediate prompt-engineering change that improves factual alignment. Because the brief is already available, the model can cite or paraphrase only that content, reducing hallucinations.

Why this answer

Grounding the prompt in the provided brief is the most direct way to stop unsupported statistics and claims. It tells the model exactly which facts are allowed and how to respond when a detail is missing. This approach avoids retraining costs and can be applied immediately, making it the right first step for the marketing team.

Exam trap

The trap here is assuming that creative parameters like temperature control factual accuracy, when they actually affect randomness rather than grounding.

67
Multi-Selecteasy

A team is using a language model for customer feedback analysis. They want to improve the accuracy of sentiment extraction. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Increase the temperature to 0.8 to allow more creative interpretations.
B.Provide few-shot examples of correctly labeled sentiment in the prompt.
C.Use the model's built-in sentiment analysis API instead of prompting.
D.Add a system instruction that asks the model to strictly follow JSON output format.
E.Fine-tune the model on a labeled dataset of customer feedback.
AnswersB, E

Few-shot examples guide the model's output format and accuracy.

Why this answer

Two complementary techniques improve sentiment extraction accuracy. First, providing few-shot examples of correctly labeled sentiment in the prompt guides the model through in-context learning, clarifying the desired output format and classification boundaries without retraining. Second, fine-tuning the model on a labeled dataset of customer feedback adapts the model's weights to the specific domain and label distribution, yielding more accurate and consistent sentiment classification than prompting alone.

Increasing temperature, using a generic built-in sentiment API, or enforcing JSON output format do not directly improve classification accuracy.

Exam trap

A common misconception is that increasing temperature or enforcing output format directly improves accuracy, but these techniques affect creativity or structure, not the correctness of the underlying classification.

68
Multi-Selectmedium

Which TWO techniques are most effective for improving factual accuracy in a generative AI model's responses? (Choose two.)

Select 2 answers
A.Retrieval-Augmented Generation (RAG) with curated datasets.
B.Increasing the model's temperature to 1.5.
C.Grounding with a trusted knowledge base.
D.Using longer system prompts with multiple instructions.
E.Fine-tuning on a large corpus of general text.
AnswersA, C

RAG retrieves relevant passages from a curated corpus at inference time and injects them into the prompt, so responses cite grounded evidence rather than relying solely on parametric memory. The curated datasets constraint directly reduces hallucination and improves factual accuracy.

Why this answer

Option A, Retrieval-Augmented Generation (RAG) with curated datasets, is correct because RAG retrieves relevant, up-to-date passages from an external curated corpus at inference time and injects them into the context, so the model's answer is anchored to verifiable source content rather than relying solely on parametric memory, which directly reduces hallucination and improves factual accuracy. Option C, grounding with a trusted knowledge base, is correct because grounding constrains generation to authoritative, domain-validated sources (e.g., enterprise knowledge bases or vetted document stores), letting the model cite or align with established facts and thereby improving correctness and traceability. Option B is wrong because raising temperature to 1.5 increases sampling randomness and diversity, which makes outputs less deterministic and more prone to fabrication, not more factual.

Option D is wrong because longer system prompts with many instructions can dilute attention, introduce conflicting directives, and do not supply new factual evidence, so they are unreliable for accuracy. Option E is wrong because fine-tuning on a large general-text corpus mainly adapts style and broad language patterns and can even reinforce outdated or incorrect parametric knowledge, whereas accuracy gains require targeted, high-quality, task-specific data or retrieval.

Exam trap

The trap is confusing fine-tuning with factual grounding; candidates assume training on more data improves facts, but fine-tuning on general text teaches patterns, not verified truths — only retrieval and grounding inject authoritative facts at inference time.

69
MCQhard

A healthcare startup has fine-tuned a Vertex AI PaLM 2 model on a dataset of medical records to generate patient summaries. The model produces fluent text but occasionally fabricates diagnoses not present in the input. The team has already tried increasing the training data size by 20% and adjusting the temperature from 0.7 to 0.2, but hallucinations persist. The summaries must be factually accurate for regulatory compliance. What should the team do next?

A.Increase the maximum output tokens to allow the model to generate more detailed summaries.
B.Implement a RAG pipeline using Vertex AI Search to retrieve relevant medical documents before generation.
C.Add more few-shot examples to the prompt for each generation.
D.Switch the base model to Gemini 1.5 Pro without additional changes.
AnswerB

RAG grounds generation in retrieved source documents, so summaries cite actual record content rather than relying on parametric memory that fabricates diagnoses. Fine-tuning and temperature changes cannot guarantee factual accuracy, which regulatory compliance demands here.

Why this answer

Implementing a Retrieval-Augmented Generation (RAG) pipeline with Vertex AI Search grounds the model's output in retrieved, authoritative medical documents. This directly addresses the root cause of hallucination—the model's reliance on its parametric memory—by providing factual context at inference time, which is far more effective for regulatory compliance than adjusting generation parameters or training data size alone.

Exam trap

The trap here is that candidates often assume adjusting model parameters (temperature, tokens) or switching models will fix hallucinations, when in fact the core issue is the lack of external knowledge grounding, which only RAG or similar retrieval-based techniques can reliably address for factual accuracy.

How to eliminate wrong answers

Option A is wrong because increasing maximum output tokens does not improve factual accuracy; it only allows the model to generate longer text, which can actually increase the opportunity for hallucinations. Option C is wrong because adding more few-shot examples to the prompt does not prevent the model from fabricating diagnoses; few-shot learning guides style and format but does not ground the model in external, verifiable facts. Option D is wrong because switching the base model to Gemini 1.5 Pro without additional changes does not solve the hallucination problem; all large language models can fabricate information when relying solely on their training data, and the underlying issue of factual grounding remains unaddressed.

70
MCQeasy

A team notices their text generation model repeats phrases excessively. Which technique would most directly reduce repetition?

A.Use beam search with a beam width of 5
B.Apply a repetition penalty of 1.2
C.Increase top_k to 100
D.Lower temperature to 0.5
AnswerB

A repetition penalty of 1.2 directly down-weights tokens already generated, lowering their probability at each subsequent step. This penalises the excessive phrase recurrence the team observed, satisfying the requirement to reduce repetition without altering the model's weights or retraining.

Why this answer

Applying a repetition penalty directly penalizes the model for generating tokens that have already appeared in the output, reducing the probability of repeated phrases. Unlike sampling or beam search adjustments, this technique explicitly targets the repetition issue by scaling down the logits of previously generated tokens during decoding.

Exam trap

A common misconception is that increasing diversity (via top_k or temperature) directly solves repetition, when in fact these methods can exacerbate the problem by making the model more random or more deterministic without addressing the underlying tendency to repeat high-probability tokens.

How to eliminate wrong answers

Option A is wrong because beam search with a beam width of 5 explores multiple candidate sequences but does not inherently discourage repetition; in fact, beam search often increases repetition by favoring high-probability sequences that may loop. Option C is wrong because increasing top_k to 100 expands the set of candidate tokens from which to sample, which can introduce more diversity but does not directly penalize repeated tokens, so it may not reduce repetition effectively. Option D is wrong because lowering temperature to 0.5 sharpens the probability distribution, making the model more deterministic and often worsening repetition by favoring the same high-probability tokens repeatedly.

71
MCQmedium

A customer support team uses a foundation model via the Gemini API to answer billing questions. Responses are often vague and sometimes omit required steps. The team wants to improve output quality without fine-tuning the model. Which approach should they use?

A.Use prompt engineering with few-shot examples and a clear output format that lists the required steps in order.
B.Lower the top-p value so the model only considers the most probable tokens, which should make the answers more accurate.
C.Increase the maximum output tokens so the model has more room to produce a longer, more complete answer.
D.Raise the temperature setting so the model can explore more possible answers and include more detail in its responses.
AnswerA

Few-shot examples show the model the desired structure and level of detail, while an explicit format forces the required steps to appear. This steers the base model without any training, which matches the constraint of not fine-tuning. It directly addresses both vagueness and missing steps by making the expected answer shape part of the prompt.

Why this answer

Prompt engineering with few-shot examples and an enforced output format is the fastest way to raise quality without changing the model weights. Examples demonstrate the expected depth, and a structured template guarantees that required steps appear. Sampling parameters alter randomness, and token limits alter length, but neither supplies missing procedural knowledge or enforces completeness.

Exam trap

The trap here is assuming that sampling parameters such as temperature or top-p can add missing factual steps, when they only change randomness and cannot supply required content.

72
MCQhard

A healthcare provider uses a Gemini model on Vertex AI to summarize patient intake notes for clinicians. The summaries must consistently follow a fixed structure: chief complaint, history, medications, and plan. Early tests show the model sometimes reorders or omits sections. Which technique most reliably enforces the required structure?

A.Provide few-shot examples in the prompt that show the exact section order and format.
B.Fine-tune the model on a large corpus of unstructured clinical literature.
C.Increase the temperature so the model varies its section ordering.
D.Reduce the maximum output tokens to force the model to be concise.
AnswerA

Few-shot examples demonstrate the precise structure the model should follow, and in-context patterns strongly shape output format for that request. By showing several correctly ordered summaries, the prompt makes the expected sections and sequence explicit, which reliably improves adherence for a fixed template. It requires no model change and can be updated quickly as the template evolves, making it well suited to structured clinical summaries.

Why this answer

Few-shot examples are the most reliable prompt-level technique for enforcing a fixed output structure, because they show the model the exact sections and their order within the same request. They need no retraining and can be revised as the template changes. Temperature changes increase disorder, broad unstructured fine-tuning does not teach the template, and token limits can drop sections rather than order them.

Exam trap

The trap here is reaching for fine-tuning or length limits to fix a formatting problem, when demonstrating the desired structure with in-context examples is the direct and reliable control.

73
MCQhard

A model generates biased output. Which technique is least effective?

A.Use adversarial debiasing
B.Apply safety filters
C.Set frequency penalty to 1.0
D.Fine-tune on diverse data
AnswerC

Frequency penalty only discourages repeated tokens; it cannot correct bias baked into training data or model weights. Bias mitigation requires data curation, fine-tuning or prompt design, so this setting leaves the underlying cause untouched and is therefore least effective.

Why this answer

Setting the frequency penalty to 1.0 is least effective for reducing biased output because frequency penalties reduce repetition of tokens based on their frequency in the generated text, not their association with protected attributes or fairness. This parameter controls lexical diversity, not demographic parity or representational harm, so it does not address the root cause of bias in the model's training data or inference logic.

Exam trap

Google often tests the misconception that any hyperparameter affecting output diversity (like frequency penalty) can mitigate bias, when in fact bias mitigation requires targeted techniques that address representation, fairness, or safety directly.

How to eliminate wrong answers

Option A is wrong because adversarial debiasing directly trains the model to minimize a discriminator's ability to predict protected attributes from the model's representations, actively reducing bias. Option B is wrong because safety filters can block or flag outputs containing harmful stereotypes or slurs, providing a post-hoc mitigation layer against biased content. Option D is wrong because fine-tuning on diverse data reweights the training distribution to include underrepresented groups, reducing statistical bias in the model's predictions.

74
MCQeasy

A company is using a generative AI model to answer customer questions. The answers are often vague and do not use the company's specific terminology. They want to improve the relevance and specificity of the responses. Which technique should they use?

A.Using retrieval-augmented generation (RAG)
B.Increasing the model's temperature
C.Fine-tuning the model on company documents
D.Reducing the max output tokens
AnswerA

RAG combines a generative model with a retrieval system that fetches relevant documents from a knowledge base. By grounding responses in company-specific documents, the model can produce more accurate and specific answers that use the correct terminology, directly addressing the issue.

Why this answer

Retrieval-augmented generation (RAG) enhances a generative model by retrieving relevant information from a knowledge base and including it in the prompt. This grounds the model's responses in factual, company-specific content, improving relevance and ensuring the use of correct terminology. It is efficient because it does not require retraining the model.

Exam trap

The trap here is assuming that fine-tuning is always necessary to inject domain knowledge, when RAG can achieve similar results more quickly and with less data.

75
Multi-Selecthard

Which THREE techniques are commonly used to improve the overall quality and coherence of generative model outputs? (Choose three.)

Select 3 answers
A.Using self-consistency or iterative refinement to choose the best output.
B.In-context learning (few-shot prompting) with relevant examples.
C.Applying output safety filters to remove inappropriate content.
D.Prompt chaining to decompose complex tasks into simpler sub-tasks.
E.Random sampling to increase output diversity.
AnswersA, B, D

Iterative methods improve reliability and coherence by selecting the most consistent response.

Why this answer

Self-consistency (A) improves quality by generating multiple candidate outputs from the same prompt and selecting the most consistent or frequent answer, reducing variance and errors, and iterative refinement lets the model revise its own output via feedback loops. In-context learning (B) with relevant few-shot examples steers the model toward the desired format, style, and reasoning pattern, improving output quality without retraining. Prompt chaining (D) decomposes a complex task into simpler sub-tasks, so each step is handled with a focused prompt and yields more coherent final outputs.

By contrast, output safety filters (C) address appropriateness rather than quality or coherence, and random sampling (E) increases diversity but can reduce coherence, so neither is a quality-improvement technique.

Exam trap

Google often tests the distinction between techniques that improve output quality (e.g., self-consistency, prompt chaining, in-context learning) versus safety or diversity mechanisms, leading candidates to mistakenly select output filters or random sampling as quality-enhancing methods.

Page 1 of 3 · 160 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Techniques to Improve Generative AI Model Output questions.