Courseiva

CCNA Fundamentals of Generative AI Questions

75 of 156 questions · Page 1/3 · Fundamentals of Generative AI · Answers revealed

1
MCQhard

A team is fine-tuning a large language model on custom data using Vertex AI. They find that the training loss decreases but validation loss increases. What is the best course of action?

A.Increase the number of training epochs.
B.Reduce the model size or add dropout regularization.
C.Increase the learning rate.
D.Switch to a smaller batch size.
AnswerB

Adding dropout regularisation directly counteracts the overfitting causing validation loss to diverge from training loss. Reducing model capacity likewise limits memorisation of the custom fine-tuning data. Both address the generalisation gap the stem describes, where training loss falls while validation loss rises, restoring alignment between the two curves.

Why this answer

The increasing validation loss while training loss decreases is a classic sign of overfitting, where the model memorizes the training data but fails to generalize. Reducing model size or adding dropout regularization directly combats overfitting by limiting the model's capacity or introducing noise during training, which forces the model to learn more robust features. This is the best course of action because it addresses the root cause without further exacerbating the problem.

Exam trap

Google Cloud often tests the distinction between underfitting and overfitting, and the trap here is that candidates may confuse increasing validation loss with underfitting and incorrectly choose to increase epochs or learning rate, rather than recognizing the hallmark divergence of overfitting.

How to eliminate wrong answers

Option A is wrong because increasing the number of training epochs would further overfit the model to the training data, worsening the validation loss. Option C is wrong because increasing the learning rate can cause the model to overshoot minima and destabilize training, potentially increasing both training and validation loss, and does not address overfitting. Option D is wrong because switching to a smaller batch size introduces more noise in gradient estimates, which can sometimes help generalization but is not a direct or reliable remedy for overfitting; it may also slow convergence and is not the primary solution for the described loss divergence.

2
Multi-Selecteasy

Which TWO of the following are key differences between generative AI and discriminative AI? (Choose two.)

Select 2 answers
A.Generative models can create new data samples, while discriminative models only assign labels to existing data.
B.Generative models require less training data than discriminative models.
C.Generative models cannot be used for supervised learning tasks like classification.
D.Generative models model the joint probability distribution of inputs and labels, whereas discriminative models model the conditional probability of labels given inputs.
E.Discriminative models always outperform generative models on tasks like image classification.
AnswersA, D

The defining axis is output behaviour: generative models learn the joint distribution to produce novel samples, whereas discriminative models learn decision boundaries mapping inputs to labels. This distinction separates creation of new data from classification of existing data.

Why this answer

Option A is correct because generative AI learns the underlying data distribution so it can produce novel samples (e.g., images, text, audio), whereas discriminative AI learns only a decision boundary or mapping from inputs to labels and therefore just classifies or predicts labels for existing data. Option D is correct because, mathematically, generative models estimate the joint probability P(x, y) (or P(x) for unsupervised generation), allowing them to sample new data, while discriminative models estimate the conditional probability P(y | x) directly to separate classes. Option B is wrong because generative models typically require large amounts of training data, often more than discriminative models, to capture the full data distribution.

Option C is wrong because generative models can be adapted to supervised tasks such as classification (e.g., using class-conditional likelihoods or fine-tuning), so they are not inherently unusable for classification. Option E is wrong because no model class universally outperforms the other; performance depends on the task, data size, and architecture, and discriminative models often excel at classification while generative models excel at synthesis.

Exam trap

Google Cloud often tests the misconception that generative models are only for unsupervised tasks and cannot perform classification, leading candidates to incorrectly select Option C, while also testing the false assumption that discriminative models are universally superior, as in Option E.

3
MCQeasy

Which Google Cloud product provides access to pre-trained foundation models like Gemini?

A.Dataflow
B.Vertex AI Generative AI Studio
C.Cloud Translation
D.Vertex AI Model Registry
AnswerB

Vertex AI Generative AI Studio exposes pre-trained foundation models, including Gemini, through a managed Google Cloud interface for prompting, tuning and deployment. It directly satisfies the requirement for a Google Cloud product offering access to such models.

Why this answer

Vertex AI Generative AI Studio is the correct answer because it is the Google Cloud service specifically designed to provide access to pre-trained foundation models like Gemini, allowing users to test, customize, and deploy them via a managed interface. Unlike other services, Generative AI Studio directly integrates with Gemini's API and offers prompt engineering, tuning, and model evaluation capabilities.

Exam trap

The trap here is that candidates confuse Vertex AI Model Registry (a model management tool) with Generative AI Studio (the actual interface for accessing and experimenting with foundation models), leading them to pick D instead of B.

How to eliminate wrong answers

Option A is wrong because Dataflow is a fully managed stream and batch data processing service based on Apache Beam, not a platform for accessing or interacting with pre-trained foundation models. Option C is wrong because Cloud Translation is a specialized service for language translation using pre-trained models, but it does not provide access to general-purpose foundation models like Gemini or support for multimodal tasks. Option D is wrong because Vertex AI Model Registry is a metadata management service for storing and versioning models, not a tool for directly accessing or experimenting with pre-trained foundation models like Gemini.

4
MCQmedium

A marketing team is using Vertex AI's text generation model to create product descriptions. They want to control the randomness of the output to ensure consistent, focused messaging for a new product line. Which parameter should they adjust to reduce randomness and make the output more deterministic?

A.Max output tokens
B.Temperature
C.Top-k
D.Top-p
AnswerB

Temperature controls the randomness of predictions by scaling the logits before applying softmax. A lower temperature (e.g., 0.2) makes the model more confident and deterministic, producing focused outputs. This directly addresses the need for consistent messaging. Higher temperatures increase diversity but reduce consistency.

Why this answer

Temperature is the parameter that directly controls the randomness of the model's output. Lowering the temperature makes the model more likely to choose high-probability tokens, resulting in more deterministic and consistent text. This is exactly what the marketing team needs for focused product descriptions.

Exam trap

The trap here is confusing top-k or top-p with temperature; while they influence sampling, temperature is the primary control for randomness.

5
MCQmedium

A marketing team is using a generative AI model to create ad copy. They want to control the randomness of the output so that the same prompt produces consistent results for A/B testing. Which parameter should they adjust?

A.Top-p
B.Temperature
C.Top-k
D.Max output tokens
AnswerB

Temperature scales the probability distribution of the next token. Setting it to a low value, such as 0, makes the model choose the most likely token almost deterministically, leading to consistent outputs for the same prompt. This is ideal for A/B testing where reproducibility is needed. Higher temperatures increase randomness.

Why this answer

Temperature directly controls the randomness of token selection. A temperature of 0 (or very close to 0) makes the model deterministically pick the highest-probability token, ensuring the same prompt yields the same output. This is essential for A/B testing where consistent ad copy variants are required.

Other parameters affect diversity but do not guarantee reproducibility.

Exam trap

The trap here is confusing top-k or top-p with determinism; those parameters narrow the sampling pool but still allow random selection within it, unlike temperature set to zero.

6
MCQmedium

A team is tuning a large language model for a question-answering task. They notice the model gives high confidence scores to answers that are factually incorrect. Which evaluation metric should they primarily use to detect this overconfidence problem?

A.Perplexity
B.Expected Calibration Error (ECE)
C.BLEU score
D.ROUGE-L
AnswerB

Expected Calibration Error directly measures the gap between predicted confidence and actual accuracy across probability bins, so systematically high confidence on wrong answers produces a large ECE value. This satisfies the stem's overconfidence constraint, unlike accuracy or F1, which ignore confidence entirely and cannot expose miscalibration.

Why this answer

Expected Calibration Error (ECE) directly measures the alignment between a model's predicted confidence and its actual accuracy. In this scenario, high confidence on incorrect answers indicates miscalibration, and ECE quantifies this mismatch by binning predictions by confidence and computing the average absolute difference between accuracy and confidence per bin.

Exam trap

Google Cloud often tests the distinction between intrinsic evaluation metrics (like perplexity) and calibration metrics, leading candidates to mistakenly choose perplexity when the core issue is confidence miscalibration rather than general model uncertainty.

How to eliminate wrong answers

Option A is wrong because Perplexity measures how well a probability distribution predicts a sample, reflecting model uncertainty over token sequences, but it does not assess calibration of confidence scores against factual correctness. Option C is wrong because BLEU score evaluates n-gram overlap between generated and reference texts for translation quality, not confidence calibration or factual accuracy. Option D is wrong because ROUGE-L measures longest common subsequence recall for summarization tasks, and is unrelated to detecting overconfidence in model predictions.

7
MCQhard

Which of the following is a best practice when using Vertex AI for prompt engineering?

A.Always set temperature to 0
B.Use consistent formatting and delimiters
C.Avoid using examples in the prompt
D.Use very long prompts to include all possible instructions
AnswerB

Consistent formatting and delimiters give the model an unambiguous structure, separating instructions from input data. This reduces parsing ambiguity and variance in outputs, satisfying the prompt engineering best practice of reproducible, predictable responses across repeated runs.

Why this answer

Consistent formatting and delimiters (e.g., using triple backticks, XML tags, or clear section headers) help the model parse instructions and context reliably, reducing ambiguity and improving output quality. This is a core best practice in prompt engineering on Vertex AI because it leverages the model's attention mechanisms to focus on distinct prompt segments, leading to more predictable and accurate responses.

Exam trap

Google Cloud often tests the misconception that 'more is better' in prompts or that deterministic settings like temperature=0 are universally optimal, leading candidates to overlook the importance of structured, concise formatting.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0 always is not a best practice; temperature controls randomness, and while 0 yields deterministic outputs, many tasks benefit from slight variability (e.g., creative generation or diverse suggestions), and Vertex AI supports a range of 0.0 to 1.0. Option C is wrong because including examples (few-shot prompting) is a powerful technique to guide the model's behavior and improve performance, especially for complex or nuanced tasks; avoiding them would reduce effectiveness. Option D is wrong because very long prompts can exceed context windows, dilute key instructions, and increase latency or cost; Vertex AI models have token limits (e.g., 8,192 tokens for Gemini), and concise, well-structured prompts are more efficient.

8
MCQmedium

A marketing team is using a generative AI model to create ad copy. They notice that the outputs sometimes include made-up statistics and false claims about their products. They want to reduce these hallucinations without retraining the model. What should they do?

A.Use grounding with a source of truth such as a product database.
B.Decrease the maximum output tokens.
C.Increase the model's temperature parameter.
D.Fine-tune the model on a small set of correct ad copies.
AnswerA

Grounding allows the model to reference an external, authoritative data source (like a product database) when generating responses. This reduces hallucinations by anchoring outputs in verified facts. It does not require retraining and can be implemented via Vertex AI's grounding features, such as using a corpus or connecting to a database.

Why this answer

Grounding connects the model to verified external data, ensuring responses are based on facts rather than model assumptions. It is a no-retraining approach that directly targets hallucinations by providing context. Other methods like temperature or token limits do not reliably improve factual accuracy.

Exam trap

The trap here is assuming that fine-tuning or parameter tweaks can eliminate hallucinations without external data grounding.

9
MCQeasy

A developer is using Vertex AI PaLM API to generate code snippets. The responses sometimes contain security vulnerabilities. What is the best practice to mitigate this?

A.Implement input validation and output filtering with safety attributes
B.Disable safety filters to allow more output
C.Increase the max output tokens
D.Set safety settings to block all categories
AnswerA

Input validation rejects malicious prompts, while output filtering with safety attributes screens generated code for insecure patterns before delivery. This layered control mitigates vulnerable snippets at both ends, satisfying the requirement to reduce security flaws in Vertex AI PaLM responses.

Why this answer

Input validation and output filtering with safety attributes directly address security vulnerabilities by sanitizing user inputs and filtering model outputs for harmful content. The Vertex AI PaLM API provides safety attribute scores (e.g., toxicity, harassment) that allow developers to programmatically block or flag responses that exceed defined thresholds, reducing the risk of generating insecure code snippets.

Exam trap

The trap here is that candidates may think increasing token limits or disabling filters improves output quality, when in fact the core issue is controlling content safety through validation and filtering, not adjusting generation parameters.

How to eliminate wrong answers

Option B is wrong because disabling safety filters removes all guardrails, allowing the model to generate potentially harmful or insecure code without any mitigation, which increases security risks. Option C is wrong because increasing max output tokens does not affect the content's security; it only allows longer responses, which could include more vulnerabilities. Option D is wrong because setting safety settings to block all categories is overly restrictive and may prevent legitimate code generation, but more importantly, it does not address the root cause of vulnerabilities—input validation and output filtering are needed to catch context-specific issues like insecure code patterns.

10
MCQeasy

A hospital wants to use a generative AI model to draft patient discharge summaries. The model must be able to understand medical terminology and generate coherent, context-aware text. Which type of generative AI model is most suitable?

A.A generative adversarial network (GAN)
B.A variational autoencoder (VAE)
C.A large language model (LLM) like Gemini
D.A rule-based natural language generation (NLG) system
AnswerC

LLMs are trained on vast text corpora, including medical literature, and can understand and generate coherent, context-aware text. They are ideal for drafting discharge summaries because they can handle medical terminology and produce fluent narratives. Their ability to follow instructions and maintain context makes them the best fit for this clinical documentation task.

Why this answer

Large language models like Gemini are pre-trained on extensive text, including medical content, and can generate coherent, context-aware narratives. They are well-suited for drafting discharge summaries because they understand medical terminology and can adapt to different patient cases. Other model types lack the language generation capabilities needed for this task.

Exam trap

The trap here is assuming any generative model can handle text, when only LLMs are designed for complex language understanding and generation.

11
MCQmedium

A marketing analyst at a retail company wants to use a generative AI model to draft product descriptions for 500 new items. The analyst expects the model to understand the context of each product from a short prompt and produce varied, coherent text. Which core capability of generative AI is the analyst primarily relying on?

A.The model's ability to perform sentiment analysis on customer reviews to gauge product reception.
B.The model's ability to generate novel, contextually relevant text by predicting likely sequences of words.
C.The model's ability to retrieve exact matches from a product database and insert them into a template.
D.The model's ability to classify products into predefined categories based on their attributes.
AnswerB

Generative AI models learn statistical patterns from vast text corpora and predict the next token given prior context. This allows them to produce original, context-aware product descriptions from a short prompt. The analyst's task—drafting varied, coherent copy for many products—directly leverages this text-generation capability, which is the hallmark of large language models.

Why this answer

The analyst needs to produce original, varied, and contextually appropriate text for each product. This is achieved through the generative capability of large language models, which predict and generate sequences of words based on the input prompt. Other options describe classification, retrieval, or sentiment analysis—tasks that do not create novel product descriptions.

Exam trap

The trap here is confusing generative AI with traditional NLP tasks like classification or sentiment analysis, which produce labels rather than new text.

12
MCQhard

A gen AI application produces hallucinations (factually incorrect outputs). Which mitigation strategy is LEAST effective?

A.Using prompt templates with constraints
B.Using grounding with a knowledge base
C.Implementing retrieval-augmented generation
D.Increasing model temperature
AnswerD

Raising temperature increases sampling randomness, which makes outputs more varied and creative, so it amplifies rather than reduces hallucination. Grounding, retrieval and lower temperature all constrain the model; temperature is the one lever that moves in the wrong direction.

Why this answer

Increasing model temperature makes the model more random and creative, which directly increases the likelihood of hallucinations. It does not constrain or ground the output in factual data, making it the least effective mitigation strategy among the options.

Exam trap

Google Cloud often tests the misconception that increasing model temperature improves accuracy by making the model 'more confident,' when in reality it increases randomness and hallucination risk.

How to eliminate wrong answers

Option A is wrong because prompt templates with constraints (e.g., specifying 'only answer from the provided context') reduce the model's freedom to generate unverified content, thereby lowering hallucination risk. Option B is wrong because grounding with a knowledge base ties the model's outputs to verified facts, preventing fabrication by restricting the response to a trusted data source. Option C is wrong because retrieval-augmented generation (RAG) explicitly fetches relevant documents from a knowledge base before generation, ensuring the output is based on retrieved evidence rather than parametric memory alone.

13
MCQmedium

A financial services firm wants to deploy a generative AI assistant for employees. The assistant must not reveal any customer data and must comply with internal policies. Which Google Cloud capability should they use to control the assistant's responses based on defined safety and privacy rules?

A.Vertex AI safety filters and configurable thresholds in the generative AI API.
B.Cloud Audit Logs to record all API calls made to the assistant.
C.Vertex AI Model Garden to browse and deploy open models.
D.VPC Service Controls to create a perimeter around the project.
AnswerA

Vertex AI provides safety filters and adjustable thresholds that can block harmful or policy-violating content. By configuring these settings, the firm can reduce the risk of responses that leak sensitive information or violate internal rules. This is the built-in mechanism for applying content safety controls to generative AI responses, making it the appropriate choice for policy enforcement.

Why this answer

To enforce safety and privacy rules on generated content, the firm should use Vertex AI's safety filters and configurable thresholds. These settings allow blocking or adjusting responses that violate defined policies, providing a preventive control at the model output level. This directly supports compliance by reducing the chance of leaking customer data or violating internal rules.

Exam trap

The trap here is confusing infrastructure-level controls like VPC Service Controls or audit logging with content-level guardrails that actually filter model responses.

14
MCQeasy

A marketing team wants to generate product descriptions from a short list of features using a Google Cloud generative AI model. They have no labeled examples and want to avoid any model training. Which approach should they use?

A.Training a new model from scratch using the team's feature list
B.Fine-tuning the model on a dataset of existing product descriptions
C.Zero-shot prompting with a foundation model in Vertex AI
D.Using a traditional rule-based template system to assemble descriptions
AnswerC

Zero-shot prompting sends a task instruction and the feature list directly to a pretrained foundation model, which generates the description without any task-specific examples or training. This matches the team's need for immediate results with no labeled data and no model customization, making it the correct approach here.

Why this answer

Zero-shot prompting lets a pretrained foundation model perform a new task from instructions alone, with no examples or training. Since the team has no labeled data and wants to avoid training, this is the only approach that meets all constraints. Fine-tuning, training from scratch, and rule-based templates either require training or are not generative AI.

Exam trap

The trap here is assuming that any customization, such as fine-tuning, is required to get useful output from a foundation model.

15
MCQeasy

A marketing team wants to use a generative AI model to create ad copy in multiple languages. They are new to Google Cloud and want the fastest way to start experimenting without writing code. Which tool should they use?

A.Cloud Shell Editor to write Python scripts that call the Vertex AI API.
B.Google Cloud CLI with the gcloud ai models command.
C.Vertex AI Generative AI Studio in the Google Cloud console.
D.BigQuery console to run SQL queries against ad performance data.
AnswerC

Generative AI Studio provides a no-code interface in the Google Cloud console for prompting, tuning, and comparing foundation models. The marketing team can experiment with ad copy in multiple languages by typing prompts and adjusting parameters, without writing code. This directly matches their need for a fast, code-free starting point.

Why this answer

Generative AI Studio offers a no-code, console-based environment for prompting and experimenting with foundation models. The marketing team can quickly test multilingual ad copy by entering prompts and adjusting settings, without writing any code. This makes it the fastest and most accessible option for beginners on Google Cloud.

Exam trap

The trap here is equating any Google Cloud console interface with generative AI experimentation, when tools like BigQuery are for analytics, not content generation.

16
MCQmedium

A developer is experimenting with a Gemini model in Vertex AI Studio and notices that for the same prompt, the model produces different answers each time. The developer needs the output to be as consistent and deterministic as possible for a classification task. Which parameter change should they make?

A.Increase the maximum output tokens
B.Enable streaming responses
C.Set the temperature to 2
D.Set the temperature to 0
AnswerD

A temperature of 0 makes the model select the highest-probability token at each step, producing the most deterministic and repeatable output. For classification tasks where consistency matters more than creativity, this is the appropriate setting. It minimizes sampling randomness, so identical prompts tend to yield identical or nearly identical results.

Why this answer

Sampling parameters govern how the model picks the next token. Temperature 0 selects the most probable token deterministically, which is ideal when the same input should map to the same output, as in classification. Other parameters such as token limits and streaming affect output length and delivery, not the randomness of selection, so they cannot deliver the consistency the developer needs.

Exam trap

The trap here is believing that output length controls or streaming settings influence how random or repeatable a model's responses are.

17
MCQmedium

A developer is building a generative AI application using Vertex AI and wants to ensure that the model's responses are safe and aligned with their company's content policies. They need to filter out harmful content such as hate speech and harassment from both user inputs and model outputs. Which Vertex AI feature should they configure?

A.Cloud Armor
B.IAM policies
C.Safety filters
D.VPC Service Controls
AnswerC

Safety filters in Vertex AI allow you to set thresholds for categories like hate speech, harassment, and sexually explicit content. They can be applied to both prompts and responses, automatically blocking or flagging harmful content. This directly meets the requirement to enforce content policies on inputs and outputs.

Why this answer

Safety filters in Vertex AI are specifically designed to detect and block harmful content in both prompts and responses. By configuring these filters, the developer can enforce the company's content policies and reduce the risk of generating or processing unsafe material, ensuring a safer generative AI application.

Exam trap

The trap here is confusing network security controls like VPC Service Controls with content safety filters; only safety filters inspect text for harmful categories.

18
MCQeasy

A graphic design company wants to generate high-quality synthetic images for product mockups. Which Google Cloud generative AI service is most suitable?

A.AutoML Vision
B.Imagen on Vertex AI
C.Codey APIs for code generation
D.Natural Language API
AnswerB

Imagen on Vertex AI generates photorealistic images from text prompts, directly satisfying the demand for high-quality synthetic product mockups. Unlike text-focused models, it specialises in image synthesis with fine detail control, making it the appropriate Google Cloud service for this graphic design scenario.

Why this answer

Imagen on Vertex AI is the correct choice because it is Google Cloud's state-of-the-art text-to-image diffusion model specifically designed to generate high-quality, photorealistic synthetic images from natural language prompts. This directly meets the requirement for creating product mockups, as Imagen can produce custom visuals with fine-grained control over style and composition, and it integrates seamlessly with Vertex AI for deployment and management.

Exam trap

The trap here is that candidates may confuse AutoML Vision's ability to classify or detect objects in images with generative image creation, leading them to select Option A despite it lacking any generative capability.

How to eliminate wrong answers

Option A is wrong because AutoML Vision is a traditional machine learning service for training custom image classification, object detection, or segmentation models on labeled datasets; it does not generate synthetic images from text prompts. Option C is wrong because Codey APIs are specialized for generating code snippets, documentation, and code completions, not for creating visual content like images. Option D is wrong because Natural Language API is designed for analyzing and extracting insights from text (e.g., sentiment, entity recognition), not for generating synthetic images.

19
MCQeasy

A marketing team wants to generate product descriptions from a short bullet list of features. They need a Google Cloud service that provides a web-based console for writing prompts, comparing model responses, and adjusting parameters like temperature without writing code. Which service should they use?

A.Vertex AI Studio
B.Cloud Natural Language API
C.Document AI
D.BigQuery ML
AnswerA

Vertex AI Studio is Google Cloud's console for prompt design, model comparison, and parameter tuning such as temperature and top-p, with no coding required. It directly supports the marketing team's need to draft prompts, view generated product descriptions, and iterate quickly, making it the most appropriate choice for non-developers who want visual experimentation with generative AI models.

Why this answer

Vertex AI Studio is the Google Cloud environment built for prompt design and rapid experimentation with generative models. It provides a visual interface where users can enter prompts, compare outputs across models, and adjust sampling parameters. Because the marketing team wants to generate descriptions without coding, this console directly matches their workflow.

Exam trap

The trap here is assuming any Google Cloud AI service can generate text, when several services such as Natural Language API and Document AI are specialized for analysis rather than open-ended generation.

20
MCQhard

A financial services firm is deploying a generative AI chatbot to answer employee questions about internal policies. The firm must ensure that the chatbot does not reveal sensitive information from other departments. Which technique should they implement?

A.Fine-tune the model on all internal documents to improve accuracy.
B.Set the temperature to a low value to reduce creative responses.
C.Increase the model's context window to include more documents.
D.Use retrieval-augmented generation with document-level access controls.
AnswerD

RAG can retrieve documents from a source that enforces access permissions based on the user's identity. By integrating with an identity-aware search or filtering retrieved documents by department, the chatbot only supplies context the user is authorized to see. This prevents cross-department information leakage while still providing accurate answers.

Why this answer

Retrieval-augmented generation can be combined with access control lists so that the retrieval step only returns documents the requesting user is permitted to view. This ensures the model's context is limited to authorized information, preventing leakage across departments. Fine-tuning or context expansion without access controls would not solve the security requirement.

Exam trap

The trap here is thinking that fine-tuning or larger context windows can enforce security, when in fact access control must be applied at the retrieval layer.

21
MCQmedium

Refer to the exhibit. A developer executed the command to list endpoints. They notice that two models are deployed to the same endpoint. What is the most likely reason for this configuration?

A.It is a canary deployment with traffic splitting
B.The endpoint is misconfigured and will cause conflicts
C.The models are from different frameworks
D.It is a batch prediction endpoint
AnswerA

A canary deployment routes a small slice of live traffic to a new model version while the rest continues to the stable one, so both models must sit behind the same endpoint for traffic splitting to work. This matches the observed configuration without implying an A/B test or blue-green swap.

Why this answer

A is correct because deploying two models to the same endpoint with traffic splitting is a standard canary deployment strategy. In this configuration, a small percentage of inference requests are routed to the new model while the majority go to the stable model, allowing validation of the new model's performance before full rollout. This is commonly supported by Google Cloud's Vertex AI, where you can deploy multiple models to an endpoint and assign traffic percentages to each model variant (e.g., 90% to the stable model and 10% to the canary model).

Exam trap

Google Cloud often tests the misconception that deploying two models to the same endpoint is always an error, when in fact it is a deliberate pattern for canary testing or A/B testing with traffic splitting.

How to eliminate wrong answers

Option B is wrong because deploying two models to the same endpoint with traffic splitting is a deliberate, supported configuration, not a misconfiguration; conflicts are avoided by routing traffic based on defined weights. Option C is wrong because models from different frameworks can be deployed to the same endpoint without issue, as the serving layer handles framework-specific inference containers independently. Option D is wrong because batch prediction endpoints typically use a single model or a single job configuration, not multiple models deployed simultaneously with traffic splitting.

22
MCQmedium

A healthcare organization wants to build a generative AI application that summarizes patient notes. They are concerned about the model generating harmful or inappropriate content. Which Google Cloud service should they use to filter out such content in real time?

A.Vertex AI Pipelines
B.Vertex AI Search
C.Vertex AI Model Garden
D.Vertex AI Safety Filters
AnswerD

Vertex AI Safety Filters are designed to detect and block harmful or inappropriate content in generative AI outputs. They can be configured with thresholds for categories like hate speech, harassment, and sexually explicit content. This service directly addresses the healthcare organization's need to filter out such content in real time, ensuring patient note summaries remain safe and compliant.

Why this answer

Vertex AI Safety Filters are specifically built to detect and filter harmful content in generative AI outputs. For a healthcare application summarizing patient notes, using these filters ensures that inappropriate or harmful content is blocked in real time. Other services like Model Garden, Search, or Pipelines serve different purposes and do not provide content moderation.

Exam trap

The trap here is assuming that any Vertex AI service can filter content, when only Safety Filters provide that specific capability.

23
MCQeasy

A marketing team wants to use a generative AI model to create blog post drafts from short product descriptions. They need the model to produce varied, creative text each time they run it. Which configuration parameter should they adjust to control the randomness of the output?

A.Temperature
B.Top-K
C.Top-P
D.Max output tokens
AnswerA

Temperature controls the randomness of the model's output. A higher temperature (e.g., 0.9) makes the output more diverse and creative, while a lower temperature makes it more deterministic. For generating varied blog post drafts, increasing the temperature is appropriate to encourage creative variation.

Why this answer

Temperature is the parameter that scales the logits before softmax, directly controlling the randomness of the output distribution. Higher values flatten the distribution, making less likely tokens more probable, which increases creativity and variation. For generating diverse blog drafts, the team should increase the temperature setting.

Exam trap

The trap here is confusing temperature with top-K or top-P, which also affect diversity but do not directly control the degree of randomness in the same way.

24
MCQmedium

Refer to the exhibit. A team has deployed a model to an endpoint with the configuration shown. They notice that during peak traffic, the endpoint frequently returns 429 (Too Many Requests) errors. Which action should they take to resolve this issue?

A.Change MACHINE_TYPE to n1-highmem-4
B.Increase MIN_REPLICA_COUNT to 5
C.Decrease MAX_REPLICA_COUNT to 1
D.Disable autoscaling by setting MIN_REPLICA_COUNT equals MAX_REPLICA_COUNT
AnswerB

429 errors indicate the endpoint's replicas cannot absorb peak request volume. Raising MIN_REPLICA_COUNT to 5 keeps more serving replicas warm, increasing concurrent request capacity and distributing traffic so the quota per replica is no longer exceeded.

Why this answer

The 429 (Too Many Requests) errors indicate the endpoint is receiving more concurrent requests than its current replica capacity can handle. Increasing MIN_REPLICA_COUNT raises the baseline number of serving replicas, so the endpoint has more capacity to absorb peak traffic without throttling. This directly addresses the throughput bottleneck rather than changing the instance shape or disabling scaling.

Exam trap

The trap here is confusing vertical scaling (bigger machine type) with horizontal scaling (more replicas); 429 errors are a concurrency/throughput signal, not a memory or CPU-per-instance signal.

How to eliminate wrong answers

Option A is wrong because changing MACHINE_TYPE to n1-highmem-4 alters memory-to-CPU ratio but does not increase the number of serving replicas, so concurrent request capacity is not meaningfully expanded and 429s can persist. Option C is wrong because decreasing MAX_REPLICA_COUNT to 1 caps the endpoint at a single replica, reducing capacity and worsening throttling under peak load. Option D is wrong because setting MIN_REPLICA_COUNT equal to MAX_REPLICA_COUNT disables autoscaling entirely, freezing replica count and preventing the endpoint from scaling out to meet peak demand.

25
MCQhard

Refer to the exhibit. A data scientist is fine-tuning a model. The training loss and accuracy are improving each epoch. However, after training, the model performs poorly on a held-out validation set. What is the most likely issue?

A.Underfitting
B.Inappropriate learning rate
C.Data leakage
D.Overfitting
AnswerD

Overfitting occurs when the model memorises training data, so training loss and accuracy keep improving while generalisation to unseen data degrades. That mismatch between strong training metrics and poor held-out validation performance is precisely the stem's symptom.

Why this answer

The model's training loss and accuracy improve each epoch, but performance on the validation set is poor. This classic symptom indicates overfitting, where the model memorizes the training data (including noise) rather than learning generalizable patterns. In fine-tuning, this often occurs when the model is trained for too many epochs or the dataset is too small relative to model capacity.

Exam trap

Google Cloud often tests the distinction between overfitting and underfitting by presenting improving training metrics alongside poor validation performance, which candidates may misinterpret as a learning rate issue or data leakage if they do not recognize the hallmark divergence pattern.

How to eliminate wrong answers

Option A is wrong because underfitting would show poor performance on both training and validation sets, not improving training metrics. Option B is wrong because an inappropriate learning rate typically causes training instability (e.g., loss divergence or stagnation), not a clear divergence between training and validation performance. Option C is wrong because data leakage would cause both training and validation metrics to be artificially high (since validation data leaks into training), not a gap where training is good and validation is poor.

26
Multi-Selecthard

A company is migrating an on-premises NLP pipeline to Vertex AI. Which three capabilities of Vertex AI align with common MLOps best practices for generative AI? (Choose THREE)

Select 3 answers
A.Automatic model retraining based on performance degradation
B.Local on-premises execution
C.Continuous training with Vertex AI Pipelines
D.Manual data labeling only
E.Model registry for versioning
AnswersA, C, E

Automatic retraining triggered by performance degradation satisfies the stem's MLOps monitoring constraint, closing the loop between evaluation and remediation. Vertex AI pipelines can schedule retraining when drift or quality metrics breach thresholds, keeping the migrated NLP models current without manual intervention.

Why this answer

Option A is correct because Vertex AI Model Monitoring can detect training-serving skew and prediction drift (e.g., via drift/skew thresholds on feature distributions) and trigger automated retraining workflows, which is a core MLOps practice for keeping generative AI models performant over time. Option C is correct because Vertex AI Pipelines (built on Kubeflow Pipelines/Vertex AI Pipelines SDK) enables continuous training by orchestrating reproducible DAGs that ingest new data, retrain, evaluate, and redeploy models as part of a CI/CD/CT pipeline. Option E is correct because Vertex AI Model Registry provides centralized versioning, lineage, and metadata tracking of model artifacts, allowing teams to manage model versions, roll back, and promote models to endpoints—essential for governed MLOps.

Option B is not correct because Vertex AI is a managed cloud service on Google Cloud, not a local on-premises execution environment. Option D is not correct because Vertex AI supports automated and programmatic labeling (e.g., Vertex AI Data Labeling Service with human-in-the-loop and active learning), so restricting to manual labeling only contradicts MLOps best practices.

Exam trap

Google Cloud often tests the misconception that MLOps for generative AI requires on-premises execution or manual-only labeling, but the correct answer emphasizes cloud-native automation and versioning as core best practices.

27
MCQmedium

A team uses PaLM 2 API to generate product descriptions, but the output sometimes contains factual inaccuracies. What is the best approach to improve accuracy?

A.Increase the temperature parameter
B.Reduce the top_k value
C.Use grounding with Google Search
D.Set the max_output_tokens higher
AnswerC

Grounding with Google Search retrieves current, authoritative web evidence and conditions generation on it, so product descriptions reflect verified facts rather than the model's parametric guesses. This directly addresses the factual inaccuracy constraint better than prompt tweaks or temperature changes.

Why this answer

Grounding with Google Search is the correct approach because it allows the PaLM 2 API to retrieve real-time, verifiable information from the web, directly reducing factual inaccuracies in generated product descriptions. Unlike parameter adjustments, grounding provides an external knowledge source that the model can cite, ensuring outputs are based on current and accurate data rather than relying solely on its training data.

Exam trap

Google Cloud often tests the misconception that tuning generation parameters (temperature, top_k, max tokens) can fix factual accuracy issues, when in reality those parameters control randomness and length, not the model's reliance on its training data versus external sources.

How to eliminate wrong answers

Option A is wrong because increasing the temperature parameter makes the model more random and creative, which would likely increase, not decrease, factual inaccuracies. Option B is wrong because reducing the top_k value limits the pool of tokens the model can sample from, which may reduce diversity but does not address the root cause of hallucination or factual errors. Option D is wrong because setting max_output_tokens higher only allows longer responses, which can actually increase the chance of generating more inaccuracies without improving factual correctness.

28
MCQeasy

A developer wants to quickly experiment with different foundation models available in Google Cloud. Which tool should they use?

A.BigQuery ML
B.Cloud Console Compute Engine
C.Gen AI Studio in Vertex AI
D.Vertex AI Model Registry
AnswerC

Gen AI Studio in Vertex AI provides a console playground for prompting and comparing foundation models side by side without writing deployment code, matching the need for quick experimentation. Model Garden is for discovering and deploying models, not rapid prompt-level comparison.

Why this answer

Gen AI Studio in Vertex AI is the correct tool because it provides a unified interface for discovering, testing, and customizing a wide range of foundation models (e.g., PaLM 2, Gemini, Codey, Imagen) directly from Google Cloud. It allows developers to quickly experiment with different models via a web UI or API without provisioning any infrastructure, making it ideal for rapid prototyping and prompt engineering.

Exam trap

Google Cloud exams often test the distinction between tools for model experimentation (Gen AI Studio) versus model management (Model Registry) or data-centric ML (BigQuery ML), leading candidates to choose a wrong option that sounds related but serves a different purpose.

How to eliminate wrong answers

Option A is wrong because BigQuery ML is designed for creating and executing machine learning models using SQL queries on data stored in BigQuery, not for experimenting with pre-built foundation models. Option B is wrong because Cloud Console Compute Engine provides virtual machine instances for custom workloads, requiring manual setup of model environments and lacking the curated model catalog and prompt playground of Gen AI Studio. Option D is wrong because Vertex AI Model Registry is a metadata store for managing and deploying trained models, not a tool for interactive experimentation with foundation models.

29
MCQeasy

A company wants to use generative AI to summarize customer support tickets. Which Google Cloud tool is best suited for this task?

A.Vertex AI Text Generation (Gemini)
B.Dialogflow CX
C.Document AI
D.AutoML Tables
AnswerA

Vertex AI Text Generation with Gemini handles abstractive summarisation, condensing long ticket threads into concise summaries while preserving meaning. It satisfies the scenario's requirement for a managed Google Cloud generative AI service, avoiding custom model training for straightforward text summarisation workloads.

Why this answer

Vertex AI Text Generation (Gemini) is the correct choice because it is a generative AI service specifically designed for natural language understanding and generation tasks, such as summarizing customer support tickets. Gemini models can process long-form text and produce concise, coherent summaries by leveraging transformer-based architectures fine-tuned for instruction-following and text completion. This makes it ideal for extracting key information from support conversations and generating actionable summaries.

Exam trap

The trap here is that candidates may confuse Dialogflow CX (a conversational AI builder) with a generative AI tool, overlooking that Dialogflow is for structured dialogue flows rather than open-ended text generation.

How to eliminate wrong answers

Option B (Dialogflow CX) is wrong because it is a conversational AI platform for building chatbots and virtual agents, not a generative text summarization tool; it focuses on intent recognition and dialogue management rather than free-form text generation. Option C (Document AI) is wrong because it is designed for document processing and extraction of structured data (e.g., OCR, form parsing) from scanned documents, not for generative summarization of unstructured text. Option D (AutoML Tables) is wrong because it is a tabular data modeling service for regression and classification tasks on structured datasets, not a natural language generation tool.

30
MCQhard

A healthcare company is building a generative AI assistant to answer patient questions about medications. The team wants to ensure the model's responses are grounded in approved clinical guidelines and avoid fabricated information. Which approach best addresses this requirement?

A.Implement retrieval-augmented generation (RAG) to fetch relevant passages from a curated clinical guideline database before generating a response.
B.Use a smaller model with fewer parameters to minimize the risk of generating incorrect information.
C.Increase the model's temperature to encourage more creative and varied responses.
D.Fine-tune the model on a large corpus of general medical textbooks without any retrieval mechanism.
AnswerA

RAG combines a retriever that searches a trusted knowledge base with a generator that conditions its output on the retrieved passages. This grounds responses in approved clinical guidelines and reduces hallucinations. For a healthcare assistant, RAG ensures answers are based on current, authoritative content, directly addressing the requirement for accuracy and safety.

Why this answer

Retrieval-augmented generation (RAG) grounds the model's output by retrieving relevant, up-to-date information from a curated database and conditioning the generation on that context. This reduces hallucinations and ensures responses align with approved clinical guidelines. Other options either increase randomness, rely solely on fine-tuning without retrieval, or incorrectly assume smaller models are safer.

Exam trap

The trap here is thinking that fine-tuning alone can eliminate hallucinations, when grounding with retrieval is often necessary for dynamic, factual accuracy.

31
MCQmedium

A data scientist is fine-tuning a large language model using Vertex AI. The training job fails with an out-of-memory error. Which action should they take to resolve this issue?

A.Change the accelerator to TPU
B.Use a larger model
C.Increase the batch size
D.Reduce the batch size
AnswerD

Out-of-memory during fine-tuning stems from the activation and gradient tensors held per training step. Reducing the batch size shrinks those tensors proportionally, lowering peak GPU memory below the accelerator's limit so the Vertex AI job can complete without changing the model architecture.

Why this answer

Reducing the batch size decreases the memory footprint per training step, allowing the model to fit within the available GPU or TPU memory. Out-of-memory errors during fine-tuning on Vertex AI typically occur when the batch size is too large for the allocated accelerator memory, and lowering it directly resolves the issue without changing the model architecture or hardware.

Exam trap

Google often tests the misconception that out-of-memory errors are solved by upgrading hardware (e.g., TPU or larger model) rather than adjusting the batch size, which is the simplest and most direct fix.

How to eliminate wrong answers

Option A is wrong because switching to a TPU does not inherently reduce memory usage; TPUs have their own memory limits and may still run out of memory if the batch size is too large, and the issue is not about accelerator type but memory capacity. Option B is wrong because using a larger model would increase memory consumption, worsening the out-of-memory error rather than resolving it. Option C is wrong because increasing the batch size would require more memory for storing activations and gradients, directly causing or exacerbating the out-of-memory error.

32
Multi-Selectmedium

Which TWO options are best practices for deploying generative AI models on Vertex AI? (Choose two.)

Select 2 answers
A.Disable logging to reduce cost
B.Enable automatic scaling to handle variable traffic
C.Use Vertex AI Model Monitoring to detect drift
D.Manually scale instances based on expected load
E.Serve the model directly without optimization
AnswersB, C

Automatic scaling adjusts the number of serving replicas to match incoming request volume, so generative AI endpoints stay responsive during traffic spikes without over-provisioning. This satisfies the best-practise requirement to handle variable traffic efficiently when deploying models on Vertex AI.

Why this answer

Option B is correct because enabling automatic scaling on a Vertex AI Endpoint lets the deployed model add or remove replicas based on traffic and utilization, so it handles variable load efficiently while controlling cost during idle periods. Option C is correct because Vertex AI Model Monitoring continuously tracks prediction inputs and outputs against a baseline and alerts on training-serving skew and prediction drift, which is essential for maintaining generative AI model quality over time. Option A is wrong because disabling logging removes the audit trail and prediction telemetry needed for debugging, compliance, and monitoring, and the cost savings do not justify losing observability.

Option D is wrong because manually scaling instances is error-prone and cannot react quickly to variable traffic the way autoscaling can. Option E is wrong because serving a model without optimization (for example, without appropriate machine types, batching, or quantization) wastes resources and degrades latency and throughput.

Exam trap

Google Cloud often tests the misconception that manual scaling is more reliable or cost-effective than automatic scaling, but in cloud-native environments, automatic scaling is the standard best practice for variable workloads.

33
MCQmedium

A retail company wants to build a generative AI application that creates personalized product descriptions. They need the model to stay current with weekly inventory changes and brand voice guidelines without retraining the model. Which approach should they use?

A.Deploy the model with a low temperature setting to ensure factual accuracy.
B.Use prompt engineering with few-shot examples that include current inventory and brand voice.
C.Fine-tune a foundation model on the weekly inventory data and brand guidelines.
D.Retrain the foundation model from scratch using the company's entire product catalog.
AnswerB

Prompt engineering with few-shot examples allows the model to generate outputs based on the provided context without changing model weights. By including current inventory details and brand voice examples in the prompt, the model can produce personalized descriptions that reflect the latest data and style guidelines, making it ideal for dynamic content.

Why this answer

Prompt engineering allows the model to generate content based on context provided at inference time, so the application can incorporate the latest inventory and brand guidelines without modifying the model. This approach is flexible, cost-effective, and supports frequent updates, unlike fine-tuning or retraining, which are static and resource-intensive.

Exam trap

The trap here is assuming that fine-tuning is needed to adapt a model to new data, when in fact prompt engineering can handle dynamic context without retraining.

34
MCQmedium

A retail company uses the Vertex AI Gemini API to generate product descriptions. Recently, the model started producing factually incorrect statements about product specifications, such as wrong dimensions and materials. Which strategy should be implemented to improve factual accuracy?

A.Enable model versioning to automatically roll back to a previous version
B.Fine-tune the model on a dataset of product images and descriptions
C.Increase the temperature parameter to 0.9
D.Use grounding with Vertex AI Search to retrieve verified product data
AnswerD

Grounding with Vertex AI Search anchors generation to retrieved, verified product data, so specifications come from an authoritative source rather than the model's parametric memory. This directly reduces the fabricated dimensions and materials the stem describes.

Why this answer

Grounding with Vertex AI Search connects the Gemini API to a verified, structured data source (e.g., product catalog), enabling the model to retrieve and cite factual specifications rather than relying solely on its training data. This directly addresses the hallucination of wrong dimensions and materials by providing a retrieval-augmented generation (RAG) mechanism that overrides the model's parametric knowledge with authoritative, up-to-date information.

Exam trap

Candidates often mistakenly believe that fine-tuning the model or adjusting generation parameters (like temperature) can fix factual inaccuracies. However, these methods do not introduce new, verified data. The correct approach is grounding with Vertex AI Search, which uses retrieval-augmented generation (RAG) to pull authoritative product specifications from a trusted data source, directly addressing hallucination of facts like dimensions and materials.

How to eliminate wrong answers

Option A is wrong because model versioning only reverts to a previous model snapshot; it does not fix the underlying factual inaccuracy if the older version also hallucinates or lacks grounding. Option B is wrong because fine-tuning on product images and descriptions can improve style but does not guarantee factual accuracy for specific specifications; it may even reinforce hallucinations if the training data contains errors. Option C is wrong because increasing the temperature parameter to 0.9 increases randomness and creativity, which would worsen factual accuracy by making the model more likely to generate plausible but incorrect statements.

35
MCQmedium

A company wants to generate images from text descriptions. Which model in Vertex AI Model Garden should they use?

A.Chirp
B.Codey
C.PaLM 2
D.Imagen
AnswerD

Imagen is Google Cloud's text-to-image generation model, converting natural-language prompts directly into photorealistic or artistic images. It satisfies the stem's requirement to generate images from text descriptions, unlike Gemini's multimodal understanding or embedding models. Selecting Imagen from Vertex AI Model Garden is therefore the appropriate choice.

Why this answer

Imagen is the correct choice because it is a text-to-image model in Vertex AI Model Garden specifically designed to generate high-fidelity images from natural language descriptions. Unlike the other options, Imagen uses a diffusion-based architecture to create photorealistic visuals, making it the only option that directly addresses the requirement of generating images from text.

Exam trap

The trap here is that candidates may confuse PaLM 2's multimodal capabilities (which include image understanding but not generation) with Imagen's generative ability, leading them to incorrectly select PaLM 2 for text-to-image tasks.

How to eliminate wrong answers

Option A is wrong because Chirp is a speech-to-text and text-to-speech model in Vertex AI, focused on audio processing, not image generation. Option B is wrong because Codey is a code generation model specialized in writing and completing code, not generating images from text. Option C is wrong because PaLM 2 is a large language model for text generation, reasoning, and chat, but it does not have the capability to produce visual outputs like images.

36
Multi-Selecthard

A team is designing a generative AI application that must avoid generating harmful or biased content. They are considering various techniques to implement safety measures. Which two approaches are recommended for mitigating harmful outputs in generative AI models? (Choose two.)

Select 2 answers
A.Fine-tuning the model on a dataset that contains only positive examples.
B.Using reinforcement learning from human feedback (RLHF) to align model behavior with human values.
C.Reducing the max output tokens to limit the amount of generated text.
D.Increasing the model's temperature to encourage more diverse and creative outputs.
E.Applying content filters that block or flag outputs containing hate speech, violence, or explicit material.
AnswersB, E

RLHF fine-tunes the model using human preferences to reward safe, helpful, and unbiased responses. This aligns the model's outputs with desired ethical standards and reduces harmful content. It is a recommended approach because it directly incorporates human judgment into the training process, making the model more reliable in sensitive applications.

Why this answer

RLHF aligns model behavior with human values by learning from human feedback, while content filters provide a post-generation safety net. Together, they form a robust strategy for reducing harmful outputs. Other options like increasing temperature or limiting output length do not effectively mitigate toxicity and may introduce other issues.

Exam trap

The trap here is thinking that any parameter change, such as increasing temperature or limiting tokens, can serve as a safety measure, when in fact they do not address harmful content.

37
MCQmedium

A developer is deploying a generative AI model on Vertex AI for a production application that requires low latency and high throughput. They need to choose an appropriate endpoint type. Which deployment option should they use?

A.Online prediction with a dedicated endpoint
B.Vertex AI Pipelines
C.Model monitoring
D.Batch prediction
AnswerA

Online prediction with a dedicated endpoint provides synchronous, low-latency responses and can scale to handle high throughput. It is the standard deployment option for real-time applications. Dedicated endpoints allow for consistent performance and can be configured with autoscaling to meet demand.

Why this answer

Online prediction with a dedicated endpoint is the correct choice for low-latency, high-throughput serving. It allows the application to send requests and receive immediate responses, with autoscaling to handle varying loads. Batch prediction, Pipelines, and monitoring serve different purposes and cannot fulfill the real-time inference requirement.

Exam trap

The trap here is confusing model deployment endpoints with other Vertex AI services that manage or monitor models but do not serve predictions.

38
MCQeasy

A marketing team wants to generate product descriptions using generative AI. They need to ensure factual accuracy and avoid hallucinations. Which approach should they use?

A.Use a code generation model to generate structured descriptions.
B.Fine-tune the model on all product descriptions using supervised learning.
C.Implement a retrieval augmented generation (RAG) system that retrieves product facts from a database.
D.Use a large language model with detailed prompt instructions to be accurate.
AnswerC

RAG grounds generation in retrieved product facts, so the model conditions its output on authoritative database content rather than parametric memory alone. This directly satisfies the factual accuracy and hallucination-avoidance constraint, since responses are anchored to verifiable source data instead of plausible-sounding invention.

Why this answer

Retrieval Augmented Generation (RAG) is the correct approach because it grounds the model's output in verifiable, external data sources. By retrieving product facts from a database in real-time, the system ensures that the generated descriptions are based on accurate information, directly mitigating the risk of hallucination. This method combines the generative power of an LLM with a retrieval step that provides factual context, making it ideal for applications where precision is critical.

Exam trap

Google Cloud often tests the misconception that detailed prompting alone (Option D) is sufficient to guarantee factual accuracy, when in reality, without external knowledge retrieval, the model can still generate plausible but incorrect information.

How to eliminate wrong answers

Option A is wrong because code generation models are designed to produce structured code or data formats, not to ensure factual accuracy in natural language product descriptions; they lack a retrieval mechanism to verify facts. Option B is wrong because fine-tuning on all product descriptions using supervised learning can embed training data biases and does not prevent hallucination on unseen or updated product facts; it also requires extensive labeled data and retraining for each change. Option D is wrong because while detailed prompt instructions can guide the model, they do not provide a mechanism to access or verify external facts; the model may still hallucinate based on its parametric knowledge, which can be outdated or incorrect.

39
MCQhard

You are a data scientist at a financial institution. You are using Vertex AI to fine-tune a large language model (LLM) for generating financial reports. You have prepared a dataset of 10,000 examples. During fine-tuning, you notice that the training loss is decreasing steadily, but the validation loss is increasing after 5 epochs. The model's generated reports on the validation set contain many factual errors and nonsensical statements. You suspect overfitting. You have limited compute budget and need to improve generalization. What should you do?

A.Increase the learning rate
B.Increase the number of training epochs to 20
C.Add more training examples from a public dataset
D.Implement early stopping with a patience of 2 epochs
AnswerD

Early stopping halts training once validation loss stops improving, directly countering the divergence observed after 5 epochs. With patience of 2, it preserves the best-generalising checkpoint rather than the final overfitted weights, improving generalisation without extra compute — satisfying the limited-budget constraint.

Why this answer

Early stopping with a patience of 2 epochs is the correct approach because it directly addresses overfitting by halting training when the validation loss fails to improve for a specified number of epochs. This preserves the model's generalization ability without requiring additional compute or data, which aligns with the limited budget constraint. In Vertex AI, early stopping is a built-in hyperparameter tuning strategy that monitors validation metrics and stops the job to prevent further degradation.

Exam trap

The trap here is that candidates often confuse overfitting with underfitting and choose to add more data or increase epochs, failing to recognize that the validation loss increasing while training loss decreases is the classic sign of overfitting, which requires a regularization technique like early stopping.

How to eliminate wrong answers

Option A is wrong because increasing the learning rate would make the optimizer take larger steps, which can cause the loss to diverge or oscillate, worsening the overfitting and factual errors. Option B is wrong because increasing the number of training epochs to 20 would continue training on the same data, likely exacerbating overfitting as the validation loss is already increasing after 5 epochs. Option C is wrong because adding more training examples from a public dataset may introduce domain mismatch or noise, and it does not address the immediate overfitting issue; it also requires additional compute and data curation resources, contradicting the limited budget constraint.

40
MCQmedium

A media company is using a generative AI model to create summaries of news articles. They notice that the summaries sometimes include information not present in the original articles, leading to factual errors. They want to reduce the likelihood of such hallucinations. Which approach should they take?

A.Decrease the top-k parameter to limit the vocabulary used in the summary.
B.Increase the temperature parameter to encourage more creative outputs.
C.Use a smaller model with fewer parameters to limit the amount of information it can generate.
D.Provide the article text as context in the prompt and instruct the model to only use that information.
AnswerD

By including the article text in the prompt and explicitly instructing the model to base its summary only on that content, you ground the model's response in the provided source. This technique, often called grounding or context conditioning, significantly reduces hallucinations because the model is constrained to the given text. It is a best practice for summarization tasks.

Why this answer

Hallucinations occur when a model generates plausible but unsupported information. To mitigate, you can ground the model by providing the source text in the prompt and instructing it to rely solely on that text. This constrains the model's output to the given context, reducing the chance of introducing external or fabricated details.

This is a standard technique in generative AI for summarization.

Exam trap

The trap here is assuming that adjusting model parameters like temperature or top-k can eliminate hallucinations, when grounding with source context is the effective method.

41
MCQmedium

A software team is using a generative AI model to write code snippets. They want to control the model's creativity and ensure it produces consistent, deterministic output for the same prompt. Which parameter should they adjust?

A.Top-p (nucleus sampling)
B.Top-k sampling
C.Maximum output tokens
D.Temperature
AnswerD

Setting temperature to zero makes the model select the most likely next token at each step, producing deterministic output for the same prompt. This is ideal for code generation where consistency is important. Temperature directly controls randomness, making it the correct choice.

Why this answer

Temperature controls the randomness of token selection. Setting it to zero makes the model choose the highest-probability token each time, yielding deterministic output for the same prompt. Top-p, top-k, and max tokens do not provide the same level of determinism.

Exam trap

The trap here is thinking that top-p or top-k alone can produce fully deterministic output, when only temperature zero does.

42
Multi-Selectmedium

What are THREE benefits of using embedding models in a Retrieval Augmented Generation (RAG) system?

Select 3 answers
A.They compress text into dense vectors for efficient retrieval.
B.They allow the model to generate new training data automatically.
C.They enable semantic similarity search beyond keyword matching.
D.They reduce the need for fine-tuning the generator model.
E.They provide deterministic outputs for the same query.
AnswersA, C, D

Embedding models map text to fixed-length dense vectors, so retrieval compares compact numeric representations rather than raw passages. This satisfies the RAG requirement for fast, scalable nearest-neighbour lookup across a large corpus, letting the retriever fetch relevant context efficiently before generation.

Why this answer

Option A is correct because embedding models transform text into dense, fixed-length vector representations that can be indexed (e.g., in a vector database) and compared via distance metrics like cosine similarity, making retrieval fast and scalable. Option C is correct because embeddings capture semantic meaning, so a query can match documents that are conceptually related even when they share no exact keywords, which is the core advantage over lexical/BM25 keyword search. Option D is correct because a RAG pipeline supplies relevant context at inference time through retrieval, so the generator can answer domain-specific questions without being fine-tuned on that domain's data.

Option B is not a benefit of embedding models, since they do not generate training data; that is a separate capability of generative models. Option E is not correct because embedding models produce continuous vectors and retrieval is similarity-based, so outputs are not guaranteed to be deterministic across queries or runs.

Exam trap

Google Cloud often tests the misconception that embedding models are used for generating training data or ensuring deterministic outputs, when in fact their primary role is semantic compression and similarity-based retrieval, while output determinism is controlled by the generator model's parameters, not the embedding model.

43
Multi-Selecteasy

Which TWO statements are true about generative AI models?

Select 2 answers
A.They are typically pre-trained on large datasets.
B.They are deterministic by design.
C.They always produce the same output for the same input.
D.They can generate new content not seen in training.
E.They require no data for training.
AnswersA, D

Pre-training on large datasets is the defining characteristic of generative models: self-supervised learning over billions of tokens or images builds the broad statistical patterns later fine-tuned for specific tasks. This satisfies the stem's requirement for a true statement about how such models are typically built.

Why this answer

Option A is correct because generative AI models such as large language models and diffusion models undergo a pre-training phase on massive, broad datasets (e.g., web text, images) before any fine-tuning or alignment, which is what gives them general-purpose capabilities. Option D is correct because these models learn the underlying probability distribution of the training data and can sample from it to produce novel outputs—new sentences, images, or code—that were not present verbatim in the training set. Option B is incorrect because generative models are inherently stochastic; sampling steps such as temperature, top-k, or top-p introduce randomness, so they are not deterministic by design.

Option C is incorrect for the same reason: with non-zero temperature or random sampling, the same input can yield different outputs across runs. Option E is incorrect because training data is essential—without a corpus to learn from, the model would have no parameters or distribution to generate from.

Exam trap

Google Cloud often tests the misconception that generative AI models are deterministic and always produce the same output for the same input, when in fact they are probabilistic by design, especially at non-zero temperature settings.

44
MCQmedium

A support team is using a generative AI model to answer customer questions from an internal knowledge base. They notice that when the same question is asked twice, the model sometimes gives different answers, and occasionally includes details not found in the knowledge base. They want to reduce variability and keep responses closer to the source content. Which action should they take?

A.Increase the temperature parameter to encourage more creative responses.
B.Decrease the temperature parameter and ground responses with retrieved knowledge base content.
C.Switch to a larger model without changing the prompt or parameters.
D.Increase the maximum output tokens so the model has more room to explain its reasoning.
AnswerB

Lowering temperature reduces randomness so repeated prompts produce more consistent answers, while grounding with retrieved internal content constrains the model to the provided facts. Together they address both symptoms: variable wording and invented details. This is the standard approach for support assistants that must stay faithful to an approved knowledge base.

Why this answer

Reducing temperature makes sampling more deterministic, so identical prompts tend to yield consistent answers. Grounding the model with retrieved knowledge base passages supplies authoritative context and reduces hallucinated details. Used together, these changes directly target both observed issues: inconsistent responses and content not present in the source material.

Exam trap

The trap here is thinking a bigger model or longer output solves hallucination, when the real levers are sampling temperature and providing grounded source content.

45
MCQmedium

A financial services firm is deploying a generative AI model to answer customer questions about investment products. They need to ensure that the model's responses comply with regulatory requirements and do not provide personalized financial advice. Which approach should they take?

A.Use a system prompt that instructs the model to avoid giving personalized advice and to include disclaimers.
B.Fine-tune the model on a dataset of compliant financial conversations.
C.Deploy the model with a high temperature to encourage diverse responses.
D.Limit the model's responses to a predefined set of FAQs.
AnswerA

A system prompt sets the overall behavior and constraints for the model. By instructing the model to avoid personalized advice and include disclaimers, the firm can enforce compliance at the prompt level. This is a flexible and immediate way to guide the model's responses without modifying the model itself.

Why this answer

Using a system prompt is an effective way to set guardrails for the model's behavior. By explicitly instructing the model to avoid personalized financial advice and to include necessary disclaimers, the firm can help ensure compliance with regulations. This approach is flexible and can be updated as regulations change, unlike fine-tuning which requires retraining.

Exam trap

The trap here is assuming that fine-tuning is necessary for compliance, when in fact a well-designed system prompt can enforce behavioral constraints more directly and flexibly.

46
MCQeasy

A marketing team wants to use a generative AI model to create product descriptions from a short list of features. They need the output to be creative but also follow a specific brand voice. Which approach should they take?

A.Use prompt engineering with a foundation model, providing examples of the desired tone and style.
B.Fine-tune a foundation model on a large dataset of existing product descriptions.
C.Deploy the model without any customization and rely on its default output.
D.Train a custom model from scratch using the team's product catalog.
AnswerA

Prompt engineering allows the team to guide the model's output by including instructions and examples in the prompt, achieving the desired creativity and brand voice without training. This is efficient and flexible, enabling quick iteration. Foundation models like Gemini are designed to follow such prompts effectively.

Why this answer

Prompt engineering is the most suitable approach because it allows the team to specify the desired tone, style, and content directly in the prompt. By providing examples and clear instructions, they can steer the foundation model to produce creative yet on-brand descriptions. This method is fast, cost-effective, and requires no model training.

Exam trap

The trap here is assuming that any customization requires fine-tuning or training, overlooking the power of prompt engineering for style adaptation.

47
Multi-Selecteasy

Which THREE components are core to a typical Retrieval Augmented Generation (RAG) system?

Select 3 answers
A.Classifier
B.Vector database
C.Embedding model
D.Rewriter
E.Large language model
AnswersB, C, E

A vector database stores embedded representations of source documents and performs similarity search, retrieving the passages most relevant to the incoming query. This retrieved context is then supplied to the language model, grounding its response in external knowledge and satisfying RAG's requirement for accurate, up-to-date retrieval beyond the model's training data.

Why this answer

In a typical RAG system, the vector database (B) is core because it stores embedded representations of documents and enables efficient similarity search to retrieve relevant context for a query. The embedding model (C) is also essential, as it converts both the source documents and the user query into dense vector representations that allow semantic matching. The large language model (E) is the third core component, since it consumes the retrieved context along with the query to generate a grounded, natural-language answer.

By contrast, a classifier (A) and a rewriter (D) are optional add-ons sometimes used for query routing or query reformulation, but they are not fundamental building blocks of a standard RAG pipeline.

Exam trap

Google Cloud often tests the distinction between core mandatory components (embedding model, vector DB, LLM) and optional auxiliary components (classifier, rewriter, reranker) to see if candidates understand the minimal viable RAG architecture versus extended pipelines.

48
MCQmedium

A media company is using a generative AI model to create short video scripts. They notice that the model sometimes produces content that is factually incorrect or nonsensical. Which term best describes this phenomenon?

A.Hallucination
B.Bias
C.Underfitting
D.Overfitting
AnswerA

Hallucination refers to a generative AI model producing confident but incorrect or nonsensical information. In this scenario, the model generates factually incorrect script content, which is a classic example of hallucination. This occurs because the model predicts likely sequences without grounding in verified facts, leading to plausible-sounding errors.

Why this answer

Hallucination is the term for generative AI producing incorrect or nonsensical information that appears confident. In this scenario, the model's factually incorrect script content is a direct example. Other options like overfitting, bias, or underfitting describe different issues and do not capture the specific problem of generating false but plausible text.

Exam trap

The trap here is conflating hallucination with bias or overfitting, which are distinct issues with different causes and mitigations.

49
MCQeasy

A retail company wants to build an internal assistant that answers employee HR policy questions using their existing policy documents, without exposing sensitive data to external APIs. They need a managed Google Cloud service that supports retrieval-augmented generation (RAG) and integrates with their existing data in Cloud Storage. Which Google Cloud service should they use?

A.Document AI
B.Dialogflow CX
C.Cloud Natural Language API
D.Vertex AI Search and Conversation
AnswerD

Vertex AI Search and Conversation is a managed Google Cloud service designed to build enterprise search and conversational AI applications with RAG. It ingests documents from Cloud Storage, BigQuery, and other sources, and provides grounding to reduce hallucinations, all within Google Cloud's secure environment. This matches the requirement for an internal assistant using existing policy documents without exposing data externally.

Why this answer

Vertex AI Search and Conversation is the only Google Cloud service listed that provides a managed RAG solution for building conversational assistants grounded in enterprise documents. It securely integrates with Cloud Storage and other data sources, enabling the company to deploy an internal HR assistant without exposing sensitive data externally.

Exam trap

The trap here is assuming that any conversational AI service can ground responses in enterprise documents, when only Vertex AI Search and Conversation provides native RAG integration.

50
MCQeasy

A marketing team is using a generative AI model to create ad copy. They notice that the outputs are often too creative and sometimes include exaggerated claims. They want to reduce creativity and make the outputs more predictable and factual. Which parameter should they adjust?

A.Max output tokens
B.Top-p
C.Top-k
D.Temperature
AnswerD

Temperature directly controls the randomness of the model's predictions. Lowering the temperature makes the model more deterministic and less creative, which is exactly what the team needs to reduce exaggerated claims and make outputs more predictable and factual. It is the most straightforward parameter to adjust for this purpose.

Why this answer

Temperature is the parameter that scales the logits before softmax, directly affecting the probability distribution of the next token. A lower temperature makes the distribution sharper, so the model chooses more likely tokens, resulting in more predictable and less creative outputs. This aligns with the team's goal of reducing exaggerated claims and increasing factual consistency.

Exam trap

The trap here is confusing temperature with other sampling parameters like top-p or top-k, which also affect randomness but are not the primary control for creativity.

51
Multi-Selectmedium

A team is designing prompts for a generative AI model to summarize legal documents. They want to improve the quality and relevance of the summaries. Which two prompt engineering best practices should they follow? (Choose two.)

Select 2 answers
A.Keep the prompt as short as possible to avoid confusing the model.
B.Include a few examples of well-written summaries in the prompt.
C.Ask the model to think step by step before providing the summary.
D.Provide clear and specific instructions about the desired summary length and focus.
E.Use the highest possible temperature setting to encourage creativity.
AnswersB, D

Few-shot prompting with examples demonstrates the desired format and style, helping the model generalize. For legal summaries, examples can show how to extract key clauses and maintain neutrality. This technique improves consistency and accuracy without fine-tuning, making it a best practice for specialized tasks.

Why this answer

Clear instructions and few-shot examples are proven prompt engineering techniques that enhance output quality. They reduce ambiguity and guide the model toward the desired format and content. Other options like high temperature or step-by-step reasoning are not appropriate for precise summarization tasks.

Exam trap

The trap here is thinking that creative settings or minimal prompts improve specialized tasks like legal summarization.

52
MCQeasy

A marketing team is using Google Cloud's Vertex AI Studio to generate product descriptions. They want the model to produce output in a very specific JSON format so it can be parsed by their downstream application. Which feature should they use to reliably constrain the model's output to that structure?

A.Adding a system instruction that says 'output JSON'
B.Controlled generation with a response schema
C.Increasing the temperature parameter to 1.0
D.Setting the top-P parameter to 0
AnswerB

Controlled generation lets you supply a response schema (for example, an OpenAPI-style schema) that the model must follow, guaranteeing the output is valid JSON matching the defined fields. This directly addresses the need for a reliably parseable structure without post-processing, and it is a native capability in Vertex AI Studio for Gemini models.

Why this answer

Constrained decoding through a response schema forces the model's tokens to conform to a predefined structure, so the output is guaranteed to be valid JSON with the expected fields. Prompt-level instructions and sampling parameters only influence content probabilistically and cannot guarantee parseable structure. For integration with downstream systems that require strict formatting, the schema-based approach is the only reliable choice.

Exam trap

The trap here is assuming that simply instructing the model to 'return JSON' in a prompt is sufficient to guarantee valid, parseable output.

53
MCQhard

A media company is deploying a generative AI application that must handle bursts of traffic while keeping costs predictable. They want to pay only for the compute resources actually consumed and avoid managing infrastructure. Which Google Cloud option best matches these requirements for serving a foundation model?

A.Host the model on Cloud Run with a GPU-enabled container
B.Use Vertex AI's pay-as-you-go, on-demand model serving endpoints
C.Reserve a fixed number of dedicated prediction nodes in Vertex AI
D.Provision a dedicated GPU cluster on Compute Engine with autoscaling node pools
AnswerB

On-demand serving bills per request or per token and automatically handles capacity, so bursts are absorbed without pre-provisioning and there is no infrastructure to manage. This aligns with both predictable, consumption-based cost and the desire to avoid operational burden. It is the standard serverless path for serving foundation models on Google Cloud.

Why this answer

On-demand serving in Vertex AI charges based on actual consumption and scales automatically, so traffic bursts are handled without provisioning capacity in advance. This satisfies both the cost predictability and the no-infrastructure-management requirements. Dedicated nodes and self-managed clusters bill for reserved capacity regardless of use, and container platforms reintroduce operational responsibilities that the company wants to avoid.

Exam trap

The trap here is equating autoscaling with pay-per-use, when autoscaled clusters still charge for provisioned capacity during idle periods.

54
MCQeasy

An organization wants to ensure their generative AI application does not produce toxic or harmful content. Which Vertex AI feature should they implement?

A.Safety filters and content moderation
B.Explainable AI
C.AutoML
D.Model Monitoring
AnswerA

Safety filters and content moderation apply configurable thresholds that block or flag harmful categories such as hate speech, harassment and dangerous content in both prompts and responses, directly enforcing the requirement that the application not emit toxic output.

Why this answer

Safety filters and content moderation in Vertex AI allow organizations to define and enforce policies that block or flag toxic, harmful, or inappropriate content generated by the model. This feature uses pre-built and customizable classifiers to evaluate prompts and responses against safety attributes (e.g., hate speech, harassment, sexually explicit content) before returning them to the user, directly addressing the requirement to prevent harmful outputs.

Exam trap

Google Cloud often tests the distinction between features that *analyze* model behavior (like Explainable AI or Model Monitoring) versus features that *actively enforce* safety policies (like Safety Filters), leading candidates to confuse monitoring or interpretability tools with content moderation controls.

How to eliminate wrong answers

Option B (Explainable AI) is wrong because it focuses on interpreting model predictions (e.g., feature attributions) rather than blocking toxic content; it provides transparency but no active content filtering. Option C (AutoML) is wrong because it automates model training and deployment for custom ML tasks, not content moderation or safety enforcement. Option D (Model Monitoring) is wrong because it tracks model performance and drift over time (e.g., prediction skew, data drift), not real-time content safety checks on individual outputs.

55
MCQeasy

A marketing team wants to use a generative AI model to create dozens of unique social media captions for a new product launch. They need the outputs to vary in tone and wording each time they run the prompt, while staying on topic. Which model parameter should they adjust to directly control this variability?

A.Top-P
B.Top-K
C.Max output tokens
D.Temperature
AnswerD

Temperature controls the randomness of token selection during generation. A higher temperature increases the likelihood of less probable tokens being chosen, which produces more diverse and creative outputs. For generating varied social media captions from the same prompt, increasing temperature is the correct adjustment to introduce controlled variability while keeping the model on topic.

Why this answer

Temperature is the parameter that scales the logits before sampling, directly controlling how random or deterministic the model's token choices are. Increasing temperature makes the model more likely to pick less probable words, yielding more varied and creative captions. The other parameters influence sampling or output length but do not provide the same direct control over variability.

Exam trap

The trap here is confusing sampling strategies like Top-K or Top-P with the parameter that directly scales randomness, which is temperature.

56
Multi-Selecteasy

Which TWO are benefits of using pre-trained foundation models instead of training from scratch?

Select 2 answers
A.Complete control over model architecture
B.Lower training cost
C.Eliminates the need for prompt engineering
D.Guaranteed absence of bias
E.Faster time to deployment
AnswersB, E

Pre-trained foundation models remove the need for large-scale training runs from scratch, so organisations pay only for adaptation such as fine-tuning or prompting. This directly satisfies the benefit of substantially lower compute and data-labelling expenditure compared with building a model from zero.

Why this answer

Option B (Lower training cost) is correct because pre-trained foundation models let you skip the massive compute, data, and energy expense of pre-training from scratch, requiring only comparatively cheap fine-tuning or inference. Option E (Faster time to deployment) is correct because starting from an already-trained model means you can fine-tune or prompt it and ship a working solution in far less time than a full training pipeline. Option A is wrong because pre-trained models actually constrain architecture choices, since you inherit the provider's design rather than controlling it.

Option C is wrong because prompt engineering is still typically required to steer foundation models effectively. Option D is wrong because pre-trained models can and do carry biases from their training data, so absence of bias is not guaranteed.

Exam trap

Google Cloud often tests the misconception that pre-trained models eliminate the need for any further engineering (like prompt engineering) or that they are completely bias-free, when in fact they still require careful tuning and can perpetuate biases from their training data.

57
MCQmedium

During a RAG pipeline implementation, the retrieval system frequently returns irrelevant documents, causing the generator to produce incorrect answers. Which change is most likely to improve the relevance of retrieved documents?

A.Add a re-ranking step using a cross-encoder model to refine the top results.
B.Increase the number of documents retrieved from the vector store.
C.Use a different embedding model with higher vector dimension.
D.Decrease the chunk size of documents to reduce noise.
AnswerA

A cross-encoder re-ranker scores each query-document pair jointly, capturing semantic interaction that the initial embedding retrieval misses. Re-ordering the top candidates by these scores promotes genuinely relevant documents, directly addressing the irrelevant-retrieval constraint causing incorrect generated answers.

Why this answer

Adding a re-ranking step using a cross-encoder model refines the initial retrieval results by scoring each document against the query more accurately, thus improving relevance. Cross-encoders consider the interaction between query and document, unlike bi-encoders used in vector search, leading to better ranking. The other options are less likely to directly improve relevance.

Exam trap

Generative AI Leader often tests the effectiveness of different RAG optimization techniques; candidates may think increasing retrieved documents or changing embedding models is sufficient, but re-ranking is specifically designed to improve relevance.

How to eliminate wrong answers

Option B is wrong because increasing the number of retrieved documents may include more irrelevant ones, potentially worsening the generator's output. Option C is wrong because using a different embedding model with higher dimension might improve representation but does not guarantee better relevance; it could also increase computational cost. Option D is wrong because decreasing chunk size might reduce context and could increase noise if chunks are too small, but it's not as directly effective as re-ranking.

58
MCQhard

A media company wants to generate personalized video summaries for users based on their viewing history. They plan to use a generative AI model on Vertex AI. Which technique should they use to ensure the summaries are tailored to each user's preferences?

A.Use a high temperature setting to generate diverse summaries that might match user preferences.
B.Apply prompt engineering with a generic template that asks for a summary in a friendly tone.
C.Fine-tune a separate model for each user on their viewing history.
D.Retrieval-augmented generation (RAG) with a vector database of user viewing history and video metadata.
AnswerD

RAG retrieves relevant context from a vector database based on the user's viewing history, then passes it to the model to generate personalized summaries. This ensures each summary reflects the user's preferences by incorporating their past behavior. It is a powerful technique for personalization without fine-tuning per user.

Why this answer

Retrieval-augmented generation (RAG) with a vector database allows the system to fetch relevant user-specific data, such as past viewing history and video metadata, and use it to condition the model's output. This results in summaries tailored to each user's preferences. Fine-tuning per user is unscalable, and generic prompts or temperature adjustments do not provide true personalization.

Exam trap

The trap here is assuming that fine-tuning is necessary for personalization, when RAG can achieve it more efficiently by retrieving user-specific context.

59
MCQmedium

A machine learning engineer is building a text-to-image model using Vertex AI. They want to reduce inference latency. Which strategy is most effective?

A.Use a larger image resolution
B.Use a smaller model variant
C.Enable batch processing
D.Increase the number of inference steps
AnswerB

A smaller model variant reduces the number of parameters and floating-point operations per forward pass, directly cutting compute per generated image. This satisfies the latency constraint because inference time scales with model size, unlike batching or caching, which improve throughput or repeat requests rather than single-pass generation speed.

Why this answer

Using a smaller model variant directly reduces the number of parameters and computational operations required per inference pass, which lowers latency. In text-to-image models like Imagen or Stable Diffusion, the model size is the primary driver of forward-pass time, so a smaller variant (e.g., fewer layers or reduced latent dimensions) yields faster generation.

Exam trap

The trap here is that candidates confuse throughput optimization (batch processing) with latency reduction, or assume that more steps or higher resolution improve quality without considering the latency trade-off.

How to eliminate wrong answers

Option A is wrong because larger image resolution increases the pixel space the model must process, which increases computational load and latency, not reduces it. Option C is wrong because batch processing improves throughput (images per second) but does not reduce per-request latency; it may even increase the time to first token for an individual request. Option D is wrong because increasing the number of inference steps (e.g., diffusion denoising steps) directly increases the sequential computation time, making latency worse.

60
MCQhard

A company is deploying a chatbot that uses a foundation model. They want to minimize latency for user queries. Which action is most effective?

A.Use a larger model with more parameters
B.Disable safety filters
C.Increase the number of tokens
D.Use a smaller distilled model
AnswerD

A smaller distilled model has fewer parameters and lower inference compute per token, directly cutting response latency. Distillation transfers the larger model's behaviour into a compact architecture, satisfying the latency-minimisation constraint while retaining acceptable quality for chatbot queries.

Why this answer

Distilled models are smaller, faster versions of larger foundation models, trained to mimic their behavior while requiring fewer computational resources. This directly reduces inference latency because fewer parameters mean faster forward passes through the network, which is critical for real-time chatbot responses.

Exam trap

Google Cloud often tests the misconception that 'bigger is better' for performance, but in latency-constrained scenarios, model size is inversely related to speed, and candidates may overlook distillation as a standard optimization technique.

How to eliminate wrong answers

Option A is wrong because larger models with more parameters increase computational complexity and memory bandwidth requirements, which actually increases latency rather than reducing it. Option B is wrong because disabling safety filters does not affect model inference speed; safety filters are post-processing steps that add negligible latency compared to the model itself. Option C is wrong because increasing the number of tokens (the output length) forces the model to perform more autoregressive generation steps, which linearly increases latency per additional token.

61
MCQmedium

A developer is using Vertex AI's text generation model to create product descriptions. They want to ensure the output adheres to a specific brand voice and style. Which approach is most effective without retraining the model?

A.Fine-tune the model on a dataset of brand-approved descriptions.
B.Increase the top-k parameter to allow more diverse word choices.
C.Use few-shot prompting with examples of brand-consistent descriptions.
D.Set the temperature to a high value to encourage creativity.
AnswerC

Few-shot prompting provides the model with examples of the desired brand voice, allowing it to mimic the style in new outputs. This is effective without retraining and can be quickly iterated. It leverages the model's in-context learning ability to adapt to specific tones and formats, making it ideal for brand consistency.

Why this answer

Few-shot prompting is a powerful, no-retraining method to steer a model's style by providing examples. It allows the model to infer the desired tone and structure from the prompt. Other options either require retraining or adjust parameters that affect randomness, not style adherence.

Exam trap

The trap here is assuming that parameter tweaks like temperature or top-k can enforce a specific brand voice without examples.

62
Multi-Selectmedium

A company is evaluating Google Cloud's generative AI offerings to build a custom chatbot that can answer questions about their proprietary product manuals. They want to ensure the solution can be grounded in their own data and can be integrated into their existing applications. Which two Google Cloud services should they consider? (Choose two.)

Select 2 answers
A.Cloud Vision API
B.Vertex AI Generative AI Studio
C.Vertex AI Search and Conversation
D.Cloud Translation API
E.Dialogflow CX
AnswersB, C

Vertex AI Generative AI Studio allows developers to experiment with, customize, and deploy generative models. It supports grounding through retrieval augmentation and provides APIs for integration. It can be used to build a custom chatbot grounded in product manuals, especially when combined with a vector store.

Why this answer

Vertex AI Search and Conversation and Vertex AI Generative AI Studio are both Google Cloud services that enable grounding in proprietary data and integration into applications. Vertex AI Search and Conversation provides a turnkey solution for grounded search and chat, while Generative AI Studio offers more customization for building and deploying models with retrieval augmentation.

Exam trap

The trap here is selecting Dialogflow CX as a grounding solution; while it builds chatbots, it lacks native RAG support for proprietary documents.

63
MCQmedium

A media company is using a generative AI model to produce news summaries. They notice that the summaries sometimes include fabricated details not present in the source articles. Which approach should they take to reduce these hallucinations?

A.Increase the model's temperature parameter
B.Fine-tune the model on the company's news articles
C.Ground the model with retrieval-augmented generation (RAG)
D.Use a larger model with more parameters
AnswerC

RAG grounds the model by retrieving relevant passages from a trusted knowledge base and providing them as context, so the model generates summaries based on actual source content rather than relying solely on its internal knowledge. This significantly reduces hallucinations because the model is constrained to use the retrieved facts. It is the recommended approach for factual tasks like news summarization.

Why this answer

Retrieval-augmented generation (RAG) retrieves relevant information from a trusted source and feeds it to the model, ensuring summaries are based on actual content. This reduces hallucinations by anchoring the model's output to provided facts. Other methods like increasing temperature or fine-tuning do not directly address the need for factual grounding.

Exam trap

The trap here is thinking that a larger model or fine-tuning automatically fixes hallucinations, when grounding with retrieved data is the targeted solution.

64
Multi-Selecteasy

A company is deploying a large language model (LLM) for customer support using Vertex AI. Which TWO best practices should they follow to ensure high-quality and cost-effective responses?

Select 2 answers
A.Deploy the model on Spot VMs to reduce infrastructure costs
B.Store prompts in plain text files for easy version control
C.Implement prompt optimization techniques to tailor responses
D.Use Vertex AI Model Monitoring to track input drift and response quality
E.Use a single large model for all query types to maintain consistency
AnswersC, D

Prompt optimisation refines instructions and context sent to the LLM, reducing token consumption and unnecessary retries. On Vertex AI this directly lowers inference cost while improving response relevance, satisfying the cost-effectiveness and quality constraints in the scenario.

Why this answer

Option C is correct because prompt optimization techniques (such as prompt engineering, few-shot examples, and tuning) directly improve the relevance and quality of LLM responses while reducing token usage, which lowers cost per request. Option D is correct because Vertex AI Model Monitoring tracks input drift and response quality metrics, enabling the team to detect degradation in customer support answers and retrain or adjust prompts before quality and cost efficiency suffer. Option A is not appropriate because Spot VMs are preemptible and unsuitable for a production customer support LLM endpoint that requires high availability and low latency.

Option B is not a best practice because storing prompts in plain text files lacks structured versioning, testing, and deployment controls; prompts should be managed with proper prompt management or CI/CD practices. Option E is incorrect because using a single large model for all query types increases cost and latency; routing simple queries to smaller, cheaper models while reserving the large model for complex cases is more cost-effective.

Exam trap

Generative AI Leader often tests the misconception that infrastructure choices like Spot VMs or a single large model are cost-saving best practices, when the real cost levers are prompt optimization and model routing.

65
MCQmedium

An enterprise deploys a large language model (LLM) for internal document summarization. Users complain that summaries sometimes include statements not present in the original document. Which mitigation strategy should the team prioritize to address this hallucination issue?

A.Train a discriminator model to detect hallucinations and perform adversarial training.
B.Implement retrieval-augmented generation (RAG) to ground the model in the original documents and require citations.
C.Apply reinforcement learning from human feedback (RLHF) using a reward model that penalizes hallucinations.
D.Reduce the model's temperature parameter to 0 to make outputs deterministic.
AnswerB

Retrieval-augmented generation grounds each summary in passages retrieved from the source document, so the model conditions its output on actual text rather than parametric memory alone. Requiring citations forces verifiable attribution, directly satisfying the stem's constraint that summaries must not contain statements absent from the original document.

Why this answer

Retrieval-Augmented Generation (RAG) is the most direct and effective mitigation for hallucination in document summarization because it forces the LLM to base its output on retrieved chunks of the original document. By requiring citations, the model must reference specific passages, making it verifiable and reducing the likelihood of fabricating content. This grounds the generation in the source material, addressing the root cause of hallucination—lack of factual grounding—rather than relying on post-hoc correction or output tuning.

Exam trap

Google Cloud often tests the misconception that reducing temperature or applying RLHF alone can solve hallucination, when in fact these methods do not provide the explicit grounding that RAG offers for document-specific tasks.

How to eliminate wrong answers

Option A is wrong because training a discriminator model and performing adversarial training is a complex, resource-intensive approach that does not directly prevent hallucinations at inference time; it only improves robustness against adversarial inputs, not factual grounding. Option C is wrong because RLHF with a reward model that penalizes hallucinations can reduce their frequency over time, but it requires extensive human feedback and fine-tuning, and does not guarantee grounding in specific source documents for each summary. Option D is wrong because reducing the temperature parameter to 0 makes outputs deterministic but does not eliminate hallucinations—it only reduces randomness; the model can still confidently generate false statements that were never in the source document.

66
MCQhard

A financial analyst uses a generative AI model to answer questions about recent market trends. They notice that the model sometimes provides outdated information. They want to ensure the model's responses are based on the most current data. Which approach should they take?

A.Implement retrieval-augmented generation (RAG) to fetch the latest market data from a live source.
B.Increase the model's temperature to encourage more diverse answers.
C.Fine-tune the model on a dataset of historical market reports.
D.Use a larger model with more parameters to improve knowledge retention.
AnswerA

RAG combines a generative model with a retrieval system that queries external, up-to-date sources. By fetching the latest market data at inference time, the model can ground its responses in current information. This directly addresses the need for real-time accuracy without retraining the model.

Why this answer

Retrieval-augmented generation (RAG) allows the model to access and incorporate real-time data from external sources. This ensures responses are based on the most current information, which is critical for market trends. Other options like fine-tuning or increasing temperature do not provide up-to-date data.

Exam trap

The trap here is assuming that a larger model or fine-tuning can provide current information, when they are limited by training data cutoff.

67
MCQmedium

A marketing team wants to generate product descriptions from a short bullet list of features. They need the model to produce creative, varied phrasing rather than a single deterministic output, while keeping the text grammatically correct. Which Vertex AI generative AI parameter should they adjust to introduce randomness into the model's word choices?

A.Maximum output tokens
B.Top-P
C.Temperature
D.Top-K
AnswerC

Temperature controls the randomness of token selection by scaling the model's output logits before sampling. Raising temperature flattens the probability distribution, giving lower-probability words a larger chance of being chosen, which yields more varied and creative phrasing. For a marketing use case that explicitly wants diverse outputs rather than one deterministic answer, temperature is the direct and intended control.

Why this answer

Temperature is the sampling parameter that scales the model's predicted token distribution, making outputs more random and creative when increased. Marketing copy benefits from this controlled variability because it produces distinct phrasings from the same feature list. Other parameters such as Top-K, Top-P, and maximum output tokens influence candidate selection or length but do not by themselves introduce the randomness the team wants.

Exam trap

The trap here is confusing sampling-shaping parameters like Top-K and Top-P with the parameter that actually controls randomness, which is temperature.

68
MCQmedium

An enterprise wants to use a foundation model through Google Cloud but must ensure that their prompts and responses are not used to train the underlying model and that data is encrypted in transit and at rest. They also want to avoid managing infrastructure. Which approach best meets these requirements?

A.Use a third-party public chatbot website to test prompts and then copy the results into internal systems.
B.Fine-tune a model on the enterprise's confidential prompts so the model learns the company's style.
C.Use Vertex AI generative AI models through the Google Cloud console and API under the enterprise's own project and data-governance controls.
D.Download an open-source model and run it on a Compute Engine VM that the team patches and scales manually.
AnswerC

Vertex AI provides managed access to foundation models without infrastructure management, and Google Cloud's data processing terms state that customer prompts and responses are not used to train foundation models. Traffic is encrypted in transit and data at rest is encrypted by default, aligning with all stated requirements within the customer's project boundary.

Why this answer

Vertex AI offers managed access to generative models with enterprise controls inside the customer's Google Cloud project. Google Cloud's terms specify that customer data for foundation models is not used to train those models, and encryption in transit and at rest is provided by default. This combination satisfies the governance and no-infrastructure-management requirements.

Exam trap

The trap here is equating any model access with enterprise-grade governance, when public chatbots and self-managed deployments miss the training-data and infrastructure requirements.

69
MCQhard

A retail company is building a generative AI chatbot to assist customers with product recommendations and order tracking. The chatbot uses Vertex AI with Gemini 1.5 Pro, and the development team has implemented a Retrieval-Augmented Generation (RAG) pipeline using Vertex AI Search for grounding. The pipeline uses a vector store containing product descriptions and order history. During testing, the team observes that the chatbot sometimes provides incorrect order statuses—for example, claiming an order is 'shipped' when it is actually 'pending'. The team suspects the issue is related to how context is retrieved and used. The RAG pipeline currently retrieves the top 5 chunks based on cosine similarity from the vector store, and passes them as context to the model. The team is considering several changes to improve factual accuracy. Which single action would most effectively reduce hallucinations in this scenario?

A.Switch from Vertex AI Search to a different vector database like Pinecone.
B.Reduce the model temperature to 0.0 to make outputs more deterministic.
C.Increase the similarity score threshold for retrieval to 0.85 to filter out less relevant chunks.
D.Increase the top-K retrieval value to 10 to provide more context to the model.
AnswerC

Increasing the similarity threshold to 0.85 filters out low-relevance chunks (e.g., order statuses from other customers), ensuring the model receives only highly relevant context, which directly reduces hallucinated order statuses in this RAG pipeline.

Why this answer

Increasing the similarity score threshold to 0.85 ensures that only highly relevant chunks are passed to the Gemini 1.5 Pro model, directly reducing the risk of the model generating responses based on irrelevant or low-confidence context. In a RAG pipeline using Vertex AI Search, low-similarity chunks can contain order statuses from different customers or products, leading to hallucinations like incorrect order statuses. Filtering out these less relevant chunks improves the factual grounding of the model's output.

Exam trap

Google Cloud often tests the misconception that simply adding more context (higher top-K) or making the model more deterministic (lower temperature) will fix hallucinations, when the real issue is the relevance and quality of the retrieved context in a RAG pipeline.

How to eliminate wrong answers

Option A is wrong because switching to a different vector database like Pinecone does not address the core issue of retrieval relevance; the problem lies in the similarity threshold and chunk selection, not the database technology. Option B is wrong because reducing temperature to 0.0 makes the model more deterministic but does not fix the underlying issue of irrelevant or incorrect context being retrieved; the model will still confidently generate incorrect answers based on poor context. Option D is wrong because increasing top-K to 10 would retrieve more chunks, potentially including even more low-relevance or noisy context, which could worsen hallucinations rather than improve factual accuracy.

70
MCQeasy

Refer to the exhibit. What is the most likely cause of this error?

A.The user does not have the required IAM role
B.The model is too large
C.The network is down
D.The project ID is incorrect
AnswerA

The error stems from an authorisation failure: the caller's identity lacks the IAM role granting the required permission on the resource. Without that binding, the API rejects the request regardless of valid credentials, so granting the missing role resolves it.

Why this answer

The error shown in the exhibit is an HTTP 403 Forbidden response, which indicates that the server understood the request but refuses to authorize it. In Google Cloud, this is most commonly caused by the user's identity lacking the necessary IAM role or permission to call the specific API or access the resource. Even if the project ID is correct and the network is functional, a missing IAM role (e.g., `aiplatform.user` or `roles/aiplatform.user`) will result in this exact error.

Exam trap

Google Cloud often tests the distinction between authentication (who you are) and authorization (what you can do), and the trap here is that candidates confuse a 403 Forbidden with a 404 Not Found or a network error, leading them to pick 'The project ID is incorrect' or 'The network is down' instead of recognizing the IAM permission failure.

How to eliminate wrong answers

Option B is wrong because model size does not cause an HTTP 403 error; a model that is too large would typically result in a 413 Payload Too Large or a resource-exhausted error, not an authorization failure. Option C is wrong because a network outage would produce a connectivity error (e.g., timeout, DNS resolution failure, or HTTP 502/503), not a 403 Forbidden response which requires a successful TCP connection and HTTP request to reach the server. Option D is wrong because an incorrect project ID would cause a 404 Not Found or a 400 Bad Request (e.g., 'Project not found'), not a 403 Forbidden; the 403 specifically indicates the request was received and the project exists, but the caller lacks authorization.

71
MCQeasy

A retail company wants to use a generative AI model to create unique product descriptions for thousands of items. They need the model to produce varied, creative text without requiring them to provide any examples. Which type of model should they use?

A.A convolutional neural network (CNN)
B.A recurrent neural network (RNN)
C.A decision tree classifier
D.A large language model (LLM) such as Gemini
AnswerD

LLMs like Gemini are pre-trained on vast text corpora and can generate varied, creative text from prompts without task-specific examples. They excel at open-ended generation tasks like product descriptions, making them ideal for this scenario. The model's broad training enables it to produce unique outputs for each product, aligning with the requirement for creativity and variety.

Why this answer

Large language models like Gemini are pre-trained on diverse text and can generate creative, varied outputs from prompts alone. They do not require task-specific examples, making them perfect for producing unique product descriptions at scale. Other model types are designed for different tasks and cannot fulfill the creative text generation requirement.

Exam trap

The trap here is assuming that any neural network can generate text, when only generative models like LLMs are designed for open-ended creation.

72
MCQmedium

A developer is using Vertex AI Gemini API for a chatbot. The chatbot sometimes outputs harmful content. What is the best first step to mitigate this?

A.Fine-tune the model on curated safe data
B.Add a human-in-the-loop review
C.Use safety filters and safety settings in the API request
D.Switch to a smaller model
AnswerC

Safety filters and safety settings are applied per API request, letting the developer block harmful categories before responses reach users. This is the fastest mitigation, requiring no retraining or architectural change to the Gemini chatbot.

Why this answer

The Vertex AI Gemini API provides built-in safety filters and configurable safety settings (e.g., `safety_settings` parameter with categories like `HARM_CATEGORY_HARASSMENT` and thresholds like `BLOCK_ONLY_HIGH`) that allow developers to block harmful outputs at inference time without retraining. This is the fastest and most direct first step to mitigate harmful content, as it requires no additional infrastructure or model modification.

Exam trap

Google Cloud often tests the misconception that the first step to mitigate harmful content is to fine-tune the model, when in reality the immediate, low-cost, and recommended first step is to leverage the API's built-in safety filters and settings.

How to eliminate wrong answers

Option A is wrong because fine-tuning on curated safe data is a resource-intensive, secondary step that does not address immediate harmful outputs during inference and may not cover all edge cases of harmful content. Option B is wrong because adding a human-in-the-loop review introduces latency and cost, and is a reactive measure rather than a proactive first step to block harmful content at the API level. Option D is wrong because switching to a smaller model does not inherently reduce harmful outputs; smaller models can still generate harmful content and may have reduced capabilities for safe response generation.

73
MCQhard

An organization wants to use a generative model to automatically generate legal contracts. The model must produce clauses that are not only grammatically correct but also legally enforceable and consistent with current jurisdiction laws. Which combination of techniques best ensures legal compliance?

A.Fine-tune a small model exclusively on legal contracts from a single jurisdiction and use it for generation.
B.Implement retrieval-augmented generation (RAG) with a vector database of all relevant laws.
C.Fine-tune a model on a diverse set of enforceable contracts and incorporate an external compliance verifier that uses rule-based checks.
D.Use a large instruction-tuned model with carefully engineered prompts describing jurisdiction details.
AnswerC

Fine-tuning on enforceable contracts aligns the model’s outputs with jurisdiction-specific clause patterns, while the external rule-based verifier enforces statutory constraints the model cannot guarantee. This satisfies the stem’s requirement for legally enforceable, jurisdiction-consistent clauses by combining learned drafting conventions with deterministic compliance checks, rather than relying on the model’s probabilistic recall of current law.

Why this answer

Fine-tuning on a diverse corpus of enforceable contracts teaches the model patterns of valid legal language across contexts, while an external rule-based compliance verifier checks generated clauses against jurisdiction-specific statutes and regulations. This hybrid approach combines generative fluency with deterministic legal validation, which is necessary because LLMs alone cannot guarantee current legal compliance. The verifier provides an auditable, updatable layer that can be revised as laws change.

Exam trap

Generative AI Leader often tests the belief that RAG or prompt engineering alone guarantees compliance, when deterministic external validation is required for regulated domains.

How to eliminate wrong answers

Option A is wrong because training on a single jurisdiction's contracts limits generalization and still cannot guarantee the model internalizes every current statute or amendment. Option B is wrong because RAG retrieves relevant laws but does not enforce them; the model may still generate non-compliant clauses despite having the correct text in context. Option D is wrong because prompt engineering alone cannot ensure legal enforceability or up-to-date jurisdiction compliance, as the model's parametric knowledge may be stale or incomplete.

74
MCQhard

A data scientist is using Vertex AI to generate product descriptions from a list of features. They notice that the model sometimes omits key features or invents details not present in the input. They want to reduce hallucinations and ensure all provided features are included. Which technique should they apply?

A.Enable the model's grounding feature with a public dataset.
B.Use a few-shot prompting approach with examples that include all features.
C.Set the top-p parameter to 0.1 to narrow the token selection.
D.Increase the temperature to encourage more creative outputs.
AnswerB

Few-shot prompting provides the model with examples of the desired input-output format, showing how to incorporate all features without adding extraneous details. This guides the model to follow the pattern, reducing omissions and hallucinations. It is an effective prompt engineering technique for structured generation tasks.

Why this answer

Few-shot prompting is the most direct way to guide the model to include all features and avoid hallucinations. By showing examples where every feature is incorporated accurately, the model learns the expected pattern. Other parameters like temperature or top-p affect randomness but do not provide the necessary instruction on completeness and fidelity.

Exam trap

The trap here is thinking that lowering randomness parameters like top-p will solve hallucinations, when the real issue is lack of guidance on the task format.

75
MCQmedium

A company fine-tunes a text model on internal HR policies. After deployment, the model sometimes outputs sensitive employee information. What is the most likely cause?

A.The fine-tuning dataset contained personally identifiable information that was not removed.
B.The model was not trained with reinforcement learning from human feedback (RLHF).
C.The model has insufficient parameters to generalize properly.
D.The prompt engineering was too verbose and included misleading instructions.
AnswerA

Fine-tuning embeds the training corpus into the model's weights, so any personally identifiable information left in the HR policy dataset can be reproduced verbatim at inference. The stem's constraint is sensitive employee data appearing in outputs; removing PII before fine-tuning prevents this memorisation, since no filtering layer exists afterwards.

Why this answer

The most likely cause is that the fine-tuning dataset contained personally identifiable information (PII) that was not properly scrubbed. During fine-tuning, the model learns patterns and memorizes specific sequences from the training data. If the dataset includes sensitive employee records, the model can reproduce that information verbatim when prompted, leading to data leakage.

This is a well-known risk in fine-tuning, as models can overfit to rare or unique examples in the training set.

Exam trap

Google Cloud often tests the misconception that RLHF or prompt engineering can fix data leakage issues, but the trap here is that the root cause is always the training data itself—no amount of post-hoc alignment or prompt tweaking can prevent the model from reproducing memorized sensitive content.

How to eliminate wrong answers

Option B is wrong because RLHF is a technique used to align model outputs with human preferences, not to prevent memorization of training data; it does not address the root cause of data leakage from the fine-tuning dataset. Option C is wrong because insufficient parameters would typically cause underfitting or poor generalization, not the exact reproduction of sensitive information; memorization is more likely with larger models that have higher capacity to store training examples. Option D is wrong because verbose or misleading prompt engineering might degrade output quality but cannot cause the model to output specific employee data that was not present in its training or fine-tuning data; the model can only generate information it has learned.

Page 1 of 3 · 156 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Fundamentals of Generative AI questions.