Courseiva

CCNA Techniques to Improve Generative AI Model Output Questions

10 of 160 questions · Page 3/3 · Techniques to Improve Generative AI Model Output · Answers revealed

151
MCQmedium

A legal firm uses a generative AI to draft contracts. They want the output to follow a specific clause structure. Which technique should they use in the prompt?

A.Include a system instruction that defines the required format.
B.Increase temperature to encourage variance.
C.Use grounding to pull from a database of contracts.
D.Set stop sequences to end generation at certain points.
AnswerA

A system instruction sets persistent behavioural constraints applied before user turns, so the model reliably follows the firm's clause structure across every draft. Embedding the format in the system role enforces consistency that one-off prompt wording cannot guarantee.

Why this answer

A system instruction (or system message) sets the overall behavior and output format for the generative AI model, effectively constraining it to follow a specific clause structure. This is the most direct and reliable technique for enforcing a predefined format in the prompt, as it operates at the model's instruction-following layer.

Exam trap

Candidates may confuse structural/format control (system instructions) with content control (grounding, temperature) or output termination (stop sequences). In Google Gen AI, system instructions are the primary method to enforce output structure, not grounding which pulls external data.

How to eliminate wrong answers

Option B is wrong because increasing temperature encourages more randomness and variance in the output, which is the opposite of what is needed for a consistent, structured clause format. Option C is wrong because grounding (e.g., using Retrieval-Augmented Generation) pulls relevant data from a database but does not enforce a specific output structure; it provides content, not format constraints. Option D is wrong because stop sequences only terminate generation at a specific token or phrase, but they do not guide the model to produce a particular clause structure throughout the entire output.

152
MCQmedium

A marketing agency uses a generative AI model to create slogans for ad campaigns. The model outputs generic slogans like 'Quality you can trust' that lack originality. The agency has a library of past award-winning slogans and wants to generate more creative and brand-specific outputs. They have a requirement that the model must not produce slogans longer than 15 words. Which technique should they prioritize?

A.Use few-shot prompting with 3-5 examples of award-winning slogans in the prompt.
B.Set max tokens to 15 to force shorter, potentially more punchy slogans.
C.Increase the temperature to 1.2 to encourage more creative word combinations.
D.Fine-tune the model on the library of award-winning slogans.
AnswerA

Few-shot prompting supplies the model with concrete award-winning slogans as in-context exemplars, steering its output distribution toward the agency's brand voice and originality rather than generic phrasing. The 15-word limit is enforced separately through explicit instruction or output validation, since examples alone do not guarantee length compliance.

Why this answer

Few-shot prompting (A) is the most direct and efficient technique because it provides the model with concrete examples of the desired output style (award-winning, creative slogans) within the context window, guiding the model's generation without altering its underlying weights. This approach immediately constrains the output to be brand-specific and creative by leveraging in-context learning, while the 15-word limit can be handled via a simple instruction in the prompt, avoiding the need for fine-tuning or risky parameter changes.

Exam trap

A common mistake is to assume that adjusting token limits or temperature can substitute for providing explicit stylistic guidance. Candidates often choose B or C, but few-shot prompting directly addresses the need for creative, brand-specific output with minimal overhead.

How to eliminate wrong answers

Option B is wrong because setting max tokens to 15 does not enforce a word count; tokens are subword units, and 15 tokens could produce a slogan far shorter or longer than 15 words, and it does nothing to improve creativity or brand specificity. Option C is wrong because increasing temperature to 1.2 increases randomness, which can lead to nonsensical or irrelevant slogans, not necessarily more creative or brand-specific ones, and it may violate the 15-word requirement by producing longer outputs. Option D is wrong because fine-tuning on a library of award-winning slogans is resource-intensive, requires significant data and compute, and may cause catastrophic forgetting of general language capabilities; it is overkill when few-shot prompting can achieve the goal with zero training.

153
MCQeasy

A product team uses a translation model to convert English product descriptions into French. The model mixes formal and informal French dialects. Which simple prompt modification likely solves this?

A.Increase the temperature to encourage more consistent output.
B.Add a system prompt specifying 'Use only formal French with no informal expressions.'
C.Fine-tune the model on a corpus of formal French texts.
D.Provide a few-shot example of a formal French translation in the prompt.
AnswerB

A system prompt constrains the model's output style before user input is processed, forcing formal French register and suppressing informal vocabulary. This directly addresses the mixed-dialect problem without retraining, satisfying the requirement for a simple prompt-level modification.

Why this answer

Adding a system prompt that explicitly instructs the model to 'Use only formal French with no informal expressions' directly constrains the output style at inference time without requiring retraining. This leverages the model's instruction-following capability to enforce a specific dialect, which is the simplest and most effective modification for controlling output style in a production translation pipeline.

Exam trap

Google often tests the misconception that fine-tuning or few-shot examples are always necessary for style control, when in fact a system prompt is the simplest and most scalable solution for inference-time behavior modification.

How to eliminate wrong answers

Option A is wrong because increasing temperature adds randomness to token sampling, which would make the output less consistent and potentially increase dialect mixing, not solve it. Option C is wrong because fine-tuning requires a curated dataset and significant compute resources, making it far more complex and time-consuming than a simple prompt change for a style preference. Option D is wrong because a few-shot example can bias the model but does not guarantee consistent enforcement across all outputs, especially if the model's training data contains mixed dialects; a system prompt provides a stronger, persistent constraint.

154
MCQmedium

A team is deploying a text generation model for legal document review. They observe that the model occasionally generates factually incorrect legal citations. Which approach best reduces this issue?

A.Implement retrieval-augmented generation (RAG) with a verified legal database.
B.Lower the temperature to 0.0.
C.Use a larger base model.
D.Increase the max output tokens.
AnswerA

RAG retrieves relevant passages from a verified legal database and supplies them as context, grounding the model's output in authoritative source material. This constrains generation to cited, verifiable content, directly reducing fabricated legal citations that arise from the model's parametric memory alone.

Why this answer

Retrieval-augmented generation (RAG) with a verified legal database grounds the model in factual, up-to-date sources, directly addressing incorrect citations. Option B (lowering temperature) reduces randomness but does not prevent hallucination. Option C (using a larger model) may not guarantee correctness without proper grounding.

Option D (increasing max tokens) has no effect on factual accuracy.

155
MCQeasy

A team is using a generative AI model to create summaries of customer feedback. The summaries are often too long and include unnecessary details. The team wants to make the summaries more concise without losing key information. Which technique should they use?

A.Use prompt engineering to specify a maximum word count for the summaries.
B.Increase the temperature parameter to encourage brevity.
C.Reduce the model's top-k parameter to limit vocabulary diversity.
D.Fine-tune the model on a dataset of short summaries.
AnswerA

Prompt engineering allows the team to include explicit instructions such as 'Summarize in under 50 words.' This directly guides the model to produce concise outputs. It is a simple and effective way to control length without retraining or adjusting technical parameters.

Why this answer

Prompt engineering with a specified word limit is a straightforward way to control the length of generated summaries. By instructing the model to be concise and setting a maximum word count, the team can achieve shorter outputs that still capture essential information. This approach is flexible and can be adjusted easily without model retraining.

Exam trap

The trap here is confusing parameters that control randomness or diversity with those that control length, when in fact length is best managed through explicit instructions in the prompt.

156
MCQeasy

A developer is using Vertex AI PaLM 2 to generate product descriptions. The output is often too verbose and includes irrelevant details. Which technique should the developer apply?

A.Set top_p to 0.1
B.Enable safety filters
C.Use few-shot prompting with examples of concise descriptions
D.Increase temperature to 0.9
AnswerC

Few-shot prompting supplies concrete examples of concise descriptions, steering PaLM 2's output distribution toward brevity and relevance. This conditions the model on the desired style, directly correcting the verbosity and irrelevant detail the developer observes.

Why this answer

The developer needs to constrain the model's output to be concise and relevant. Few-shot prompting provides the model with explicit examples of the desired output format (concise descriptions), guiding it to mimic that style and length. This directly addresses verbosity and irrelevant details without altering the model's fundamental randomness or safety settings.

Exam trap

The trap here is that candidates confuse hyperparameter tuning (top_p, temperature) with prompt engineering techniques, assuming that reducing randomness (top_p) or increasing creativity (temperature) can fix verbosity, when only explicit examples in the prompt can reliably enforce a specific output style.

How to eliminate wrong answers

Option A is wrong because setting top_p to 0.1 reduces the cumulative probability threshold for token sampling, which makes the output less diverse and more deterministic, but it does not teach the model to be concise or omit irrelevant details—it only narrows the pool of possible next tokens. Option B is wrong because safety filters block harmful or sensitive content (e.g., toxicity, violence), not verbose or irrelevant details; they do not control output length or relevance. Option D is wrong because increasing temperature to 0.9 increases randomness and creativity in token selection, which would likely make the output even more verbose and include more irrelevant details, the opposite of what is needed.

157
MCQhard

Refer to the exhibit. The team changed the generation parameters to reduce output variability. However, summaries now often repeat the same phrases. Which parameter change is most likely causing the repetition?

A.Reducing top_p from 0.95 to 0.85
B.Reducing temperature from 0.7 to 0.2
C.Using the same model text-bison@002
D.Reducing top_k from 40 to 10
AnswerB

Low temperature increases determinism and repetition.

Why this answer

Reducing temperature from 0.7 to 0.2 makes the model's token selection much more deterministic, favoring the highest-probability tokens. This low-entropy sampling often causes the model to repeat the same phrases across summaries because it consistently picks the most likely continuation.

Exam trap

Generative AI Leader often tests the confusion between temperature (sharpens distribution, can cause repetition at low values) and top_p/top_k (truncate candidates but preserve randomness), causing candidates to blame the wrong parameter.

How to eliminate wrong answers

Option A is wrong because reducing top_p from 0.95 to 0.85 still leaves a broad nucleus of tokens and is a milder constraint than a temperature drop to 0.2. Option C is wrong because using the same model does not cause repetition — model choice is orthogonal to sampling behavior. Option D is wrong because reducing top_k from 40 to 10 narrows the candidate set but does not force the same token as aggressively as a very low temperature.

158
MCQeasy

A company is using Vertex AI to generate customer support summaries from chat logs. They notice that the summaries sometimes include irrelevant details from the conversation. Which technique should they use to reduce irrelevant details?

A.Use a higher top-k value.
B.Fine-tune the model on a large dataset of general conversations.
C.Add a system instruction to focus on key points.
D.Increase the temperature parameter.
AnswerC

A system instruction sets persistent behavioural guidance applied to every prompt, so the model weights key points over incidental chat content. This directly targets the irrelevant-detail constraint at generation time, unlike post-processing or prompt-by-prompt tweaks.

Why this answer

Adding a system instruction to focus on key points is the most direct and effective technique for reducing irrelevant details in generated summaries. System instructions act as a persistent, high-level directive that guides the model's attention and output structure without altering the underlying model weights. This allows the model to filter out extraneous information from the chat logs by explicitly prioritizing key points, which is a standard practice in prompt engineering for Vertex AI.

Exam trap

A common mistake for Vertex AI summarization is to adjust randomness parameters (top-k or temperature) thinking they will improve focus, but they actually increase variability and can introduce more irrelevant details. The correct technique is to use system instructions to explicitly direct the model to prioritize key points.

How to eliminate wrong answers

Option A is wrong because increasing top-k (e.g., from 40 to 100) actually increases the pool of candidate tokens considered at each step, which can introduce more randomness and irrelevant tokens, making summaries less focused. Option B is wrong because fine-tuning on a large dataset of general conversations would dilute the model's specialization for customer support summaries, potentially worsening the inclusion of irrelevant details rather than reducing them. Option D is wrong because increasing the temperature parameter (e.g., from 0.2 to 0.8) increases the randomness of token selection, which would likely amplify the generation of irrelevant details instead of suppressing them.

159
MCQeasy

A developer is using the Gemini API to build a chatbot. They want the model to always respond in a friendly, professional tone. Which prompt engineering technique should they use?

A.Set system instructions to 'You are a friendly and professional assistant.'
B.Include a few-shot example in every user message.
C.Set the temperature to 0.2.
D.Set max output tokens to 100.
AnswerA

System instructions establish persistent behavioural guidance applied across every turn, so the model consistently adopts the friendly, professional tone. This satisfies the requirement for an always-consistent tone, unlike per-message prompting, which must be repeated and can drift between requests.

Why this answer

Setting system instructions is the most direct and reliable way to define the model's persona and behavioral constraints. In the Gemini API, system instructions act as a persistent, top-level directive that influences every response, ensuring the chatbot consistently adopts a friendly and professional tone without requiring repeated examples or parameter tuning.

Exam trap

Google Cloud often tests the distinction between controlling output style (system instructions) versus controlling output randomness (temperature) or length (max tokens), so the trap here is that candidates may confuse temperature or token limits with persona control, thinking that lowering creativity or capping length will enforce a specific tone.

How to eliminate wrong answers

Option B is wrong because including a few-shot example in every user message is inefficient and not a persistent technique; it would require repeating the example in each turn, increasing token usage and latency, and it does not guarantee consistent tone across all interactions. Option C is wrong because setting the temperature to 0.2 controls randomness and creativity, not tone; a low temperature makes outputs more deterministic but does not enforce a specific persona or style. Option D is wrong because setting max output tokens to 100 limits response length but has no effect on the tone or style of the output; it only truncates the response.

160
MCQeasy

A marketing team uses a text generation model to draft campaign copy. They want the output to consistently follow a specific brand voice and include a call to action at the end of every draft. Which technique should they apply?

A.Provide a system instruction that defines the brand voice and requires a call to action, plus one or two exemplar drafts.
B.Set the temperature to zero so the model always produces the same deterministic wording for every campaign.
C.Increase the top-k value so the model samples from a broader set of words and naturally adopts a more distinctive voice.
D.Shorten the prompt to only the product name so the model has maximum freedom to express the brand personality.
AnswerA

A system instruction sets persistent behavioral rules such as tone and mandatory elements, and exemplars demonstrate the exact style. Together they reliably steer the model toward the brand voice and ensure the call to action appears. This is a prompt-level control that needs no training and can be updated quickly as brand guidelines evolve.

Why this answer

System instructions plus exemplars are the standard way to enforce tone and mandatory content without training. The instruction defines persistent rules, and examples show the desired pattern. Sampling changes affect randomness, and shorter prompts remove guidance, so neither reliably produces consistent brand voice and a guaranteed call to action.

Exam trap

The trap here is equating determinism or broader sampling with stylistic control, when only explicit instructions and examples define voice and required elements.

← PreviousPage 3 of 3 · 160 questions total

Ready to test yourself?

Try a timed practice session using only Techniques to Improve Generative AI Model Output questions.