Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 751–825

1008 questions total · 14pages · All types, answers revealed

Page 10

Page 11 of 14

Page 12
751
MCQmedium

An enterprise needs to generate natural-sounding speech from text for a voice assistant. They require low latency and support for custom voice models. Which service should they use?

A.Text-to-Speech API
B.Cloud Translation API
C.Vertex AI Text Generation
D.Speech-to-Text API
AnswerA

The Text-to-Speech API generates natural-sounding speech with low latency and supports custom voice models, matching the voice assistant's requirements exactly. Competing services lack either the custom voice training or the real-time synthesis latency the scenario demands.

Why this answer

The Text-to-Speech API (A) is correct because it is specifically designed to convert text into natural-sounding speech with low latency, and it supports custom voice models through features like Custom Voice and WaveNet voices. This directly meets the enterprise's requirements for a voice assistant that needs real-time, high-quality speech synthesis.

Exam trap

The trap here is confusing the Text-to-Speech API with the Speech-to-Text API, as candidates often mix up the direction of conversion (text-to-audio vs. audio-to-text) under time pressure.

How to eliminate wrong answers

Option B (Cloud Translation API) is wrong because it translates text between languages, not text to speech, and does not generate audio output. Option C (Vertex AI Text Generation) is wrong because it generates text content (e.g., chat responses, summaries) rather than synthesizing speech from text. Option D (Speech-to-Text API) is wrong because it performs the inverse operation—converting audio speech into text—and does not produce speech output.

752
MCQmedium

A marketing team is using a generative AI model on Vertex AI to create ad copy for a new product launch. The initial outputs are generic and do not reflect the brand's tone. The team wants to quickly improve the outputs without retraining the model. They have a set of example ad copies that exemplify the desired tone. Which technique should they use?

A.Increase the model's temperature setting to encourage more creative outputs.
B.Deploy the model to a new endpoint with higher throughput.
C.Fine-tune the model on the example ad copies.
D.Use few-shot prompting by including the example ad copies in the prompt.
AnswerD

Few-shot prompting involves providing a few examples of the desired output format and style directly in the prompt. This guides the model to generate text that matches the brand's tone without any model retraining, making it a fast and effective solution for this scenario.

Why this answer

Few-shot prompting is ideal when you have example outputs that demonstrate the desired style or format. By including these examples in the prompt, the model can infer the pattern and generate new content that aligns with the brand's tone. This approach requires no training and can be implemented immediately, making it the most efficient solution for the marketing team.

Exam trap

The trap here is assuming that any model improvement requires fine-tuning, when in-context learning via few-shot prompting can often achieve the desired result faster and with less effort.

753
MCQhard

A healthcare company is using a generative AI model to summarize patient records. They are concerned about the model generating incorrect medical information. They want to ensure that the summaries are grounded in the provided records and do not include hallucinations. Which technique should they implement?

A.Increase the temperature to allow the model to explore more possibilities.
B.Reduce the max output tokens to limit the summary length.
C.Fine-tune the model on a large corpus of medical textbooks.
D.Use retrieval-augmented generation (RAG) to fetch relevant patient data and include it in the prompt.
AnswerD

RAG combines a retrieval system with the generative model to fetch relevant documents (e.g., patient records) and include them in the prompt. This grounds the model's output in actual data, reducing hallucinations. For summarizing patient records, RAG ensures the model references the provided information rather than relying solely on its parametric knowledge.

Why this answer

Retrieval-augmented generation (RAG) is specifically designed to ground model outputs in external knowledge sources. By retrieving relevant patient records and including them in the prompt, the model can generate summaries that are based on the actual data, significantly reducing the risk of hallucinations. This is critical in healthcare where accuracy is paramount.

Exam trap

The trap here is thinking that fine-tuning on medical data will prevent hallucinations in summaries, when in fact grounding requires the model to reference the specific input records, which RAG provides.

754
MCQeasy

A retail company wants to build an internal tool that generates short product descriptions from bullet points. They have no machine learning engineers on staff and want to avoid managing any infrastructure. Which Google Cloud approach best fits their need?

A.Use BigQuery ML to create a linear regression model that predicts description length.
B.Train a custom image generation model on Vertex AI using their product photos.
C.Deploy an open-source LLM on a Compute Engine VM with a GPU and expose it via a REST API.
D.Use the Gemini API in Vertex AI to generate descriptions from the bullet points with a prompt.
AnswerD

The Gemini API in Vertex AI provides a fully managed, serverless way to call a foundation model for text generation. The team only needs to craft a prompt containing the bullet points and receive generated descriptions, with no infrastructure or ML engineering required. This directly matches the scenario's need for quick, low-effort text generation.

Why this answer

The Gemini API in Vertex AI is a managed service that lets teams generate text without building or operating ML infrastructure. By supplying bullet points in a prompt, the model produces product descriptions on demand. This aligns with the company's lack of ML engineers and desire to avoid infrastructure management, making it the most practical and direct solution.

Exam trap

The trap here is assuming that any Google Cloud AI service can generate text, when many services like BigQuery ML or custom training are designed for predictive or image tasks, not generative text.

755
MCQeasy

You want to use a Google foundation model to generate text summaries of news articles. Which Vertex AI service should you use?

A.Vertex AI Prediction
B.Vertex AI Model Registry
C.Vertex AI Generative AI Studio
D.Vertex AI Feature Store
AnswerC

Vertex AI Generative AI Studio provides prompt design and testing against Google's foundation models, including Gemini, for text summarisation tasks. It is the interface purpose-built for generating summaries from news articles without managing underlying infrastructure.

Why this answer

Vertex AI Generative AI Studio (now part of Vertex AI Agent Builder) provides a no-code/low-code environment to access, test, and tune Google's foundation models, including PaLM 2 and Gemini, specifically for generative tasks like text summarization. It offers built-in prompt templates and safety settings tailored for summarization use cases, making it the correct service for this task.

Exam trap

The trap here is that candidates confuse Vertex AI Prediction (a general model serving service) with the specialized generative AI studio, assuming any model inference task uses Prediction, but Google explicitly separates foundation model access into Generative AI Studio for prompt-based generative workloads.

How to eliminate wrong answers

Option A is wrong because Vertex AI Prediction is designed for deploying and serving custom-trained models or AutoML models for online predictions, not for directly accessing Google's foundation models for generative tasks. Option B is wrong because Vertex AI Model Registry is a metadata store for managing and versioning your own models, not a service for interacting with foundation models or generating summaries. Option D is wrong because Vertex AI Feature Store is a managed repository for storing, serving, and sharing feature data for ML training and online inference, unrelated to text generation or foundation model access.

756
MCQeasy

Which Google Cloud feature enables you to experiment with different prompts and model parameters interactively, and also supports model tuning without writing code?

A.Vertex AI Agent Builder
B.Google Cloud Console
C.Model Garden
D.Vertex AI Studio
AnswerD

Vertex AI Studio provides an interactive console for prompt design, side-by-side comparison and parameter tuning, plus no-code model tuning workflows. This satisfies both requirements in the stem: hands-on prompt experimentation and tuning without writing code.

Why this answer

Vertex AI Studio is the correct answer because it provides an interactive, code-free environment for experimenting with prompts and model parameters, and also supports model tuning through a graphical interface. This aligns directly with the question's requirement for interactive experimentation and no-code tuning. Option A (Vertex AI Agent Builder) is designed for building conversational agents, not for prompt experimentation or tuning.

Option B (Google Cloud Console) is a general management interface but lacks specific interactive experimentation features for generative AI. Option C (Model Garden) allows discovery and selection of models but does not support interactive prompt testing or tuning without code.

Exam trap

The trap here is that candidates may confuse Model Garden's model discovery and selection capabilities with the interactive experimentation and tuning features that are exclusive to Vertex AI Studio.

How to eliminate wrong answers

Option A is wrong because Vertex AI Agent Builder is designed for creating conversational agents and search experiences, not for interactive prompt experimentation or model tuning. Option B is wrong because Google Cloud Console is a general management interface for all GCP services, lacking the specialized, interactive prompt engineering and tuning capabilities of Vertex AI Studio. Option C is wrong because Model Garden is a repository for discovering and accessing pre-trained models, but it does not provide an interactive environment for prompt experimentation or direct model tuning without code.

757
MCQmedium

A company is evaluating ROI for a GenAI-based code review assistant. Which metric set BEST captures both productivity and quality improvements?

A.Number of reviews completed per day and lines of code written
B.Model inference latency and cost per token
C.Developer satisfaction score and reduction in code churn (percentage of code rewritten)
D.Time saved per code review and bug detection rate (percentage of bugs caught before deployment)
AnswerD

Time saved per code review quantifies productivity gains, while bug detection rate before deployment measures quality improvement. Together they capture both dimensions of ROI for a GenAI code review assistant, unlike single-axis metrics such as suggestion count or developer satisfaction.

Why this answer

Time saved per review (productivity) and bug detection rate (quality) directly measure the tool's impact. Code churn and developer satisfaction are secondary. Defect escape rate is important but harder to measure directly for code review.

758
MCQmedium

An enterprise is concerned about the cost of using a large LLM for a high-volume customer support chatbot. They want to reduce token consumption while maintaining response quality. Which strategy would be MOST effective?

A.Use a multimodal model to handle both text and images
B.Always use the largest available model to ensure best quality
C.Increase the max output tokens to capture more detail
D.Implement response caching for common questions and batch similar requests
AnswerD

Caching responses to frequently asked questions avoids re-invoking the LLM for identical prompts, and batching similar requests amortises prompt overhead across calls. Together these cut token consumption substantially at high volume while preserving response quality, meeting the stated cost constraint.

Why this answer

Caching frequent queries reduces costs because the response is served from cache without model inference. Batching requests also saves on per-request overhead. Choosing a smaller model may reduce quality.

759
Multi-Selecthard

A financial services firm is preparing a board presentation on its generative AI program. Directors want assurance that spending is disciplined and that failures are contained. Which TWO practices should the program adopt? (Choose two.)

Select 2 answers
A.Require every generative AI use case to run on the same foundation model to simplify vendor management.
B.Defer all evaluation of model output quality until after the production launch to save time.
C.Fund generative AI initiatives through staged gates, releasing additional budget only when agreed outcome metrics are met.
D.Centralize all generative AI engineering in one team that owns every use case end to end.
E.Define explicit stop conditions and decommission criteria for each initiative before development begins.
AnswersC, E

Staged funding converts a large upfront commitment into a series of smaller, evidence-gated investments. Each gate requires demonstrated progress against the agreed outcome metric, so weak initiatives stop early and capital flows to the strongest candidates. This directly answers the board's demand for disciplined spending while preserving room to scale winners quickly.

Why this answer

Board-level confidence in a generative AI program comes from disciplined capital allocation and bounded downside. Staged funding tied to outcome metrics ensures money follows evidence, while predefined stop and decommission criteria guarantee that underperforming initiatives end quickly and cheaply. Together they make the portfolio's risk explicit and controllable, which is what directors are asking to see.

Exam trap

The trap here is assuming discipline comes from standardization or centralization, when it actually comes from how funding is gated and how failures are bounded.

760
Multi-Selectmedium

An enterprise wants to use Gemini for a customer-facing application. They require the following: data isolation in a VPC, audit logging, and SLA guarantees. Which THREE features of Vertex AI satisfy these requirements?

Select 3 answers
A.Vertex AI SLA (Service Level Agreement)
B.Gemini Nano on-device
C.VPC Service Controls
D.Cloud Audit Logs integration
E.Google AI Studio free tier
AnswersA, C, D

SLA guarantees uptime and performance.

Why this answer

Vertex AI offers a defined Service Level Agreement (SLA) that guarantees uptime and performance metrics for enterprise customers, which is a core requirement for customer-facing applications. The SLA provides contractual assurances, typically covering availability and response times, ensuring the enterprise can meet its own service commitments.

Exam trap

The Generative AI Leader exam often tests the distinction between development tools (like AI Studio free tier) and production-ready enterprise features, and candidates mistakenly assume that any Google AI offering includes SLA and VPC controls by default.

761
MCQmedium

A data scientist notices that a Gemini model generates inconsistent responses to similar prompts. What is the likely cause?

A.Model is not fine-tuned enough
B.The prompt is too short
C.The temperature setting is too low
D.The top_p or temperature parameters are set too high causing randomness
AnswerD

Temperature and top_p control sampling randomness; high values widen the probability distribution over candidate tokens. That directly produces the inconsistent outputs observed for similar prompts, since the model samples differently each call rather than selecting the highest-probability token.

Why this answer

High temperature (e.g., >1.0) or high top_p (e.g., >0.9) increases the randomness of token sampling, causing the model to select less probable tokens. This directly leads to inconsistent responses for similar prompts, as the model's output distribution becomes more uniform and less deterministic.

Exam trap

Google Cloud often tests the misconception that fine-tuning or prompt length is the primary cause of output inconsistency, when in fact the sampling parameters (temperature and top_p) directly control randomness and are the most common culprit.

How to eliminate wrong answers

Option A is wrong because fine-tuning adjusts the model's weights for a specific task, but it does not control the randomness of token generation; even a fully fine-tuned model will produce inconsistent outputs if sampling parameters are set too high. Option B is wrong because prompt length affects context and specificity, not the inherent randomness of the generation process; a short prompt can still yield consistent responses if temperature and top_p are low. Option C is wrong because a low temperature setting (e.g., 0.1) actually reduces randomness, making outputs more deterministic and consistent, not inconsistent.

762
MCQmedium

A financial services firm needs to fine-tune a large language model on proprietary financial data. They require data to never leave their VPC and need full audit logging. Which Gemini access method should they use?

A.Vertex AI
B.Google AI Studio
C.Gemini API directly via Cloud Endpoints
D.Model Garden in Colab
AnswerA

Vertex AI offers VPC Service Controls, audit logging, and data isolation.

Why this answer

Vertex AI is the correct access method because it is the only option that allows fine-tuning of Gemini models within a customer's VPC (Virtual Private Cloud) with full audit logging via Cloud Audit Logs. This ensures proprietary financial data never leaves the secure network boundary, meeting strict compliance and data residency requirements.

Exam trap

The trap here is that candidates may confuse Google AI Studio's free-tier accessibility with enterprise-grade security, overlooking that Vertex AI is the only option with VPC controls and audit logging for fine-tuning proprietary data.

How to eliminate wrong answers

Option B is wrong because Google AI Studio is a web-based prototyping tool that does not support VPC-scoped fine-tuning or enterprise-grade audit logging; data is processed on Google's infrastructure outside the customer's VPC. Option C is wrong because the Gemini API directly via Cloud Endpoints does not provide native VPC controls or fine-tuning capabilities; it is a stateless API call without persistent model customization. Option D is wrong because Model Garden in Colab is a discovery and experimentation environment that lacks VPC isolation, audit logging, and fine-tuning support for production workloads.

763
Multi-Selectmedium

Which THREE of the following are common techniques to reduce harmful biases in generative AI models? (Choose three.)

Select 3 answers
A.Use reinforcement learning from human feedback (RLHF) with a reward model that penalizes biased or unfair outputs.
B.Curate diverse and balanced training datasets that overrepresent underrepresented groups.
C.Decrease the model's temperature parameter to make outputs more deterministic.
D.Apply adversarial training to remove protected attribute information from hidden representations.
E.Conduct a legal review of all generated outputs before release.
AnswersA, B, D

RLHF directly targets bias by training a reward model that assigns lower scores to unfair outputs, then optimising the generative model against it. This satisfies the stem's requirement for a bias-reduction technique, since the penalty signal actively steers the model away from biased generations during fine-tuning rather than merely detecting them afterwards.

Why this answer

Option A is correct because RLHF with a reward model that explicitly penalizes biased or unfair outputs directly optimizes the generative model's policy toward fairness-aligned behavior, using human preference data to shape the reward signal. Option B is correct because curating diverse and balanced training datasets—deliberately overrepresenting underrepresented groups—counteracts sampling bias and spurious correlations learned from skewed data, which is a foundational bias-mitigation technique. Option D is correct because adversarial training that removes protected attribute information from hidden representations (e.g., via adversarial classifiers or fair representation learning) reduces the model's ability to encode and exploit sensitive attributes.

Option C does not belong because lowering the temperature only makes sampling more deterministic; it changes randomness, not the underlying learned biases, and can even amplify the most probable (biased) output. Option E does not belong because a legal review of outputs is a post-hoc compliance check, not a technique that reduces bias within the generative model itself.

Exam trap

Google Cloud often tests the distinction between hyperparameter tuning (like temperature) and actual bias mitigation techniques, so candidates mistakenly think lowering temperature reduces bias when it only affects output randomness.

764
MCQmedium

A company wants to adopt GenAI for internal knowledge management. They plan to start with a small pilot team, gather feedback, and then expand. Which change management approach is MOST aligned with this strategy?

A.Pilot with a small group, identify AI champions, collect feedback, and iteratively expand
B.Conduct mandatory training for all employees before the rollout
C.Deploy the solution to the entire organization at once with a communication campaign
D.Focus solely on the technical deployment without change management activities
AnswerA

A phased pilot with named AI champions, structured feedback capture and iterative expansion matches the stated plan of starting small, learning, then scaling. This incremental, people-led approach reduces adoption risk and builds internal advocacy before wider rollout.

Why this answer

Iterative rollout with a pilot group, AI champions, and measuring adoption is a proven change management pattern. Starting with the entire organization is risky. Mandatory training may cause resistance.

Focusing only on technical deployment ignores the human side.

765
MCQhard

A financial services firm uses a generative AI model to draft client emails. The drafts are accurate but sometimes use an overly casual tone. The firm wants to enforce a consistently formal tone across all drafts. Which approach is most reliable for this requirement?

A.Fine-tune the model on a large corpus of casual emails to broaden its style range.
B.Add a post-processing step that replaces casual words with formal synonyms using a fixed dictionary.
C.Set the temperature to a high value so the model produces more varied language.
D.Provide a system instruction that defines the desired formal tone and include a short example of an approved email.
AnswerD

A system instruction sets persistent behavioral guidance for every request, and a concrete example anchors the model to the desired style. Together they give the model both a rule and a reference, which is more reliable than a one-off user prompt. This approach directly targets tone consistency without retraining. It also scales across many drafts because the instruction applies to all interactions.

Why this answer

System instructions provide persistent behavioral guidance, and an approved example shows the model the exact tone desired. This combination is more reliable than sampling changes, casual fine-tuning, or brittle post-processing. It directly enforces formality across all drafts and scales without retraining, making it the best fit for the firm's consistency requirement.

Exam trap

The trap here is assuming that tone can be enforced by a fixed synonym replacement or by increasing randomness, when tone is best controlled by persistent instructions and examples.

766
MCQmedium

A company uses Vertex AI PaLM for code generation. The code often contains security vulnerabilities. Which improvement should be applied?

A.Set top_k to 1
B.Include a security-focused system instruction
C.Use Codey model instead
D.Increase temperature to 0.8
AnswerB

A security-focused system instruction sets persistent behavioural guardrails, steering the model to avoid insecure patterns such as unsanitised input handling across every generation. This directly targets the vulnerability constraint; prompt-level or post-hoc scanning alone cannot shape generation as reliably.

Why this answer

Including a security-focused system instruction directly guides the model to prioritize secure coding practices, such as input validation and proper error handling, reducing vulnerabilities. This leverages prompt engineering to shape model behavior without altering parameters like temperature or top_k, which control randomness, not security awareness.

Exam trap

Google often tests the misconception that parameter tuning (like temperature or top_k) can fix content quality issues, when in fact prompt engineering—such as system instructions—is the primary tool for guiding model behavior toward specific goals like security.

How to eliminate wrong answers

Option A is wrong because setting top_k to 1 makes the model deterministic (always picks the highest-probability token), which can reduce output diversity but does not address security vulnerabilities—it may even amplify insecure patterns if they are common in training data. Option C is wrong because Codey is a specialized model for code generation, but it does not inherently include security guardrails; the same vulnerabilities can appear if the prompt lacks security context. Option D is wrong because increasing temperature to 0.8 increases randomness and creativity, which can introduce more unpredictable and potentially insecure code, worsening the vulnerability issue.

767
MCQeasy

A startup wants to deploy a custom-tuned large language model for real-time inference on Vertex AI. They need the lowest possible latency for end users. What deployment strategy should they choose?

A.Use Vertex AI Model Garden to deploy the base PaLM 2 model.
B.Wrap the model in a Cloud Function and invoke via HTTP.
C.Deploy the tuned model to a Vertex AI endpoint with GPU acceleration and autoscaling.
D.Use Vertex AI Batch Prediction to process requests in batches.
AnswerC

GPU acceleration provides the compute throughput needed for low-latency token generation, while autoscaling matches capacity to demand without cold starts. Deploying to a Vertex AI endpoint keeps the model resident for real-time inference, directly satisfying the lowest-possible-latency constraint.

Why this answer

Deploying the custom-tuned model to a Vertex AI endpoint with GPU acceleration and autoscaling is the best choice for lowest-latency real-time inference. A dedicated endpoint keeps the model loaded and ready to serve requests, avoiding the per-request startup overhead of serverless wrappers like Cloud Functions. GPU acceleration speeds up the model's forward-pass computation, and autoscaling adds or removes replicas to match traffic so that requests are not queued behind insufficient capacity.

Batch prediction (D) is designed for high-throughput offline jobs and is not suitable for interactive, low-latency use cases.

Exam trap

Candidates often confuse 'lowest possible latency' with 'high throughput' or 'cost efficiency,' leading them to choose batch prediction (D) or serverless options (B) without recognizing that GPU-accelerated endpoints are specifically designed for sub-second inference.

How to eliminate wrong answers

Option A is wrong because using Vertex AI Model Garden to deploy the base PaLM 2 model does not incorporate the custom tuning, so the model would not reflect the startup's specific data or use case, and the base model may not achieve the desired accuracy or latency for the custom task. Option B is wrong because wrapping the model in a Cloud Function introduces additional cold-start latency and HTTP overhead, and Cloud Functions are not optimized for GPU-accelerated inference, leading to higher per-request latency compared to a dedicated endpoint. Option D is wrong because Vertex AI Batch Prediction is designed for asynchronous, high-throughput processing of large datasets, not for real-time inference; it introduces significant latency due to job queuing and batch processing, making it unsuitable for low-latency end-user requests.

768
Multi-Selecteasy

Which TWO safety features are available in Vertex AI Gemini API? (Select TWO.)

Select 2 answers
A.Safety filters for categories like hate speech and harassment
B.Content restrictions based on configurable thresholds
C.Model-level encryption at rest
D.Automatic redaction of personally identifiable information (PII)
E.Integration with Cloud Data Loss Prevention (DLP)
AnswersA, B

Safety filters in the Vertex AI Gemini API let you configure thresholds that block or allow content across harm categories including hate speech, harassment, sexually explicit material and dangerous content. This directly satisfies the stem's requirement for available safety features, as the API exposes these configurable filters on requests and responses.

Why this answer

Option A is correct because the Vertex AI Gemini API applies configurable safety filters that screen both prompts and responses against harm categories such as hate speech, harassment, sexually explicit content, and dangerous content. Option B is correct because these safety filters operate using configurable thresholds (for example BLOCK_LOW_AND_ABOVE, BLOCK_MEDIUM_AND_ABOVE, BLOCK_ONLY_HIGH, BLOCK_NONE) that let developers tune how aggressively content is blocked per category. Together, A and B describe the built-in safety attributes exposed through the API's safetySettings.

Option C is not a Gemini API safety feature but a general Google Cloud storage/platform control (encryption at rest is handled by the underlying infrastructure, not the API's safety settings). Option D is incorrect because the Gemini API does not automatically redact PII as a built-in safety feature. Option E is incorrect because Cloud DLP is a separate Google Cloud service that must be integrated manually and is not a native safety feature of the Gemini API.

Exam trap

A common trap is confusing general Google Cloud security services (like encryption at rest or DLP) with the native safety features of the Vertex AI Gemini API, which only include safety filters and content thresholds.

769
MCQeasy

A retail marketing team wants to generate product descriptions in five languages. The team has no machine learning engineers and wants to avoid managing any infrastructure. Which Google Cloud option should the GenAI Leader recommend first?

A.Provision a Vertex AI training job with custom containers and train a multilingual model from scratch on product data.
B.Call the Gemini API on Vertex AI with prompts requesting the product description in each target language.
C.Create a Google Kubernetes Engine cluster with GPU node pools and self-host an open model behind an inference server.
D.Build a BigQuery ML remote model that calls a public translation endpoint and post-process the output with SQL transformations.
AnswerB

The Gemini API on Vertex AI is a fully managed service, so the team sends prompts and receives multilingual text without provisioning servers, GPUs, or training pipelines. Gemini handles multiple languages natively, letting a small marketing team generate descriptions in five languages through simple API calls or the Google Cloud console, which matches their skill set and no-infrastructure constraint.

Why this answer

A managed foundation model API is the fastest route to multilingual generation for a team without ML engineers, because Google operates the model, scaling, and serving. The other choices require training pipelines, cluster operations, or an analytics-only pattern that does not generate marketing copy from product attributes.

Exam trap

The trap here is equating 'no infrastructure' with 'no cloud service,' when a managed API is precisely the option that removes infrastructure responsibility.

770
MCQmedium

A machine learning engineer wants to convert text into numerical vectors for similarity search. Which Google Cloud service should they use?

A.Vertex AI Embeddings API
B.Natural Language API
C.Vector Search
D.Gemini API
AnswerA

Vertex AI Embeddings API generates dense vector representations of text, directly satisfying the similarity search requirement. Unlike generative models that produce text, embeddings map semantic meaning into numerical space where cosine distance reflects relatedness. This is the purpose-built Google Cloud service for creating embeddings that feed vector databases and nearest-neighbour retrieval.

Why this answer

The Vertex AI Embeddings API is the correct choice because it is specifically designed to convert text (and other data types) into dense numerical vectors (embeddings) that capture semantic meaning. These embeddings are the fundamental input for similarity search, enabling efficient comparison of text based on conceptual closeness rather than exact keyword matching.

Exam trap

The trap here is that candidates confuse the Natural Language API's text analysis capabilities (like entity extraction) with the embedding generation required for similarity search, or they assume Vector Search or Gemini API can generate embeddings directly when they are actually downstream or generative tools.

How to eliminate wrong answers

Option B is wrong because the Natural Language API performs entity extraction, sentiment analysis, and syntax analysis, but it does not generate embeddings for similarity search. Option C is wrong because Vector Search is a service for indexing and querying embeddings at scale, not for generating them from raw text. Option D is wrong because the Gemini API is a multimodal generative model for chat and content generation, not a dedicated embedding service for converting text into vectors.

771
MCQmedium

A company uses a generative model to produce product descriptions. The descriptions are factually inconsistent with the product specs. Which technique would best ensure factual accuracy?

A.Enhance the system prompt with product details
B.Implement retrieval-augmented generation (RAG) with product database
C.Lower the temperature to 0.0
D.Fine-tune the model on product descriptions
AnswerB

RAG retrieves relevant product specifications from the database at inference time and injects them into the prompt, grounding the generated description in authoritative data. This directly addresses factual inconsistency, unlike fine-tuning or prompt wording changes alone.

Why this answer

Retrieval-augmented generation (RAG) is the best technique because it dynamically retrieves relevant, up-to-date product specifications from a trusted database at inference time, grounding the model's output in verified facts. This directly addresses factual inconsistency by ensuring the generated description is based on authoritative source data rather than relying solely on the model's parametric memory.

Exam trap

Google Cloud often tests the misconception that prompt engineering alone (Option A) or deterministic sampling (Option C) can solve factual grounding issues, when in reality they do not provide external knowledge retrieval to correct hallucinations.

How to eliminate wrong answers

Option A is wrong because enhancing the system prompt with product details only provides static context that the model may still hallucinate or misinterpret; it does not enforce retrieval of current or specific factual data. Option C is wrong because lowering the temperature to 0.0 makes the output more deterministic but does not prevent the model from generating factually incorrect content that is confidently wrong. Option D is wrong because fine-tuning on product descriptions can improve style and consistency but does not guarantee factual accuracy for new or updated product specs, and it risks overfitting or memorizing inaccuracies from the training data.

772
MCQeasy

A company wants to generate images from text descriptions using Google Cloud. Which service should they use?

A.Vertex AI Imagen
B.Vertex AI Gemini
C.Cloud Vision API
D.AutoML Vision
AnswerA

Vertex AI Imagen generates images directly from text prompts, satisfying the stem's requirement for text-to-image generation on Google Cloud. Unlike Vertex AI's language or vision models, Imagen is purpose-built for photorealistic image synthesis, offering resolution, editing and watermark controls that match this scenario precisely.

Why this answer

Vertex AI Imagen is Google Cloud's purpose-built service for generating high-fidelity images from text descriptions using diffusion models. It directly addresses the requirement of text-to-image generation, offering capabilities like image editing, upscaling, and style transfer, which are not available in other Vertex AI or Vision services.

Exam trap

The trap here is that candidates may confuse Vertex AI Gemini's multimodal capabilities (understanding images) with generative image creation, or assume that Cloud Vision API or AutoML Vision can be repurposed for generation, when in fact they are strictly analysis or custom training tools.

How to eliminate wrong answers

Option B is wrong because Vertex AI Gemini is a multimodal large language model (LLM) that can process text, images, audio, and video, but it is not optimized or primarily designed for generating images from text; its strength lies in understanding and reasoning across modalities, not in image synthesis. Option C is wrong because Cloud Vision API is a pre-trained model for analyzing and extracting information from images (e.g., object detection, OCR, label detection), not for generating images from text. Option D is wrong because AutoML Vision is a service for training custom image classification or object detection models on labeled datasets, not for generative text-to-image tasks.

773
MCQeasy

A company uses a text generation model for customer support but notices it occasionally provides outdated information. Which technique should they implement to improve output accuracy?

A.Increase max output tokens
B.Implement retrieval-augmented generation (RAG)
C.Fine-tune the model with more historical support data
D.Increase model temperature to 1.0
AnswerB

RAG retrieves current information, making outputs accurate and up-to-date.

Why this answer

Retrieval-augmented generation (RAG) is the correct technique because it grounds the model's output in real-time, external knowledge sources (e.g., a vector database or document index) rather than relying solely on static training data. This directly addresses the problem of outdated information by allowing the model to retrieve and synthesize current facts at inference time, ensuring accuracy without requiring retraining.

Exam trap

The trap here is that candidates often confuse fine-tuning (which adapts the model's weights to a static dataset) with RAG (which dynamically retrieves external knowledge), leading them to choose fine-tuning as a 'deeper' fix when the core issue is stale information, not model capability.

How to eliminate wrong answers

Option A is wrong because increasing max output tokens only extends the length of the generated response, not its factual accuracy or timeliness; it may even introduce more hallucinated content. Option C is wrong because fine-tuning with more historical support data would reinforce outdated patterns and biases, making the model more likely to repeat stale information rather than adapt to current knowledge. Option D is wrong because increasing model temperature to 1.0 increases randomness and creativity in outputs, which degrades factual precision and reliability, the opposite of what is needed for accurate customer support.

774
MCQmedium

A startup wants to integrate a GenAI assistant into Google Workspace (Docs, Gmail, Sheets) to help employees draft emails and create charts. Which Google AI offering is designed for this purpose?

A.Gemini for Workspace
B.Colab Enterprise
C.Vertex AI Agent Builder
D.NotebookLM
AnswerA

Gemini for Workspace embeds generative assistance directly within Docs, Gmail, Sheets and Slides, letting employees draft emails and build charts in-place. It is the Google offering purpose-built for Workspace integration, unlike standalone Vertex AI or the Gemini API.

Why this answer

Gemini for Workspace is Google's AI offering specifically designed to integrate generative AI capabilities into Google Workspace applications like Docs, Gmail, Sheets, and Slides. It provides features such as 'Help me write' in Docs, smart reply and email drafting in Gmail, and formula/chart generation in Sheets. The other options are developer-focused or research-oriented tools not intended for direct Workspace integration.

Exam trap

Generative AI Leader often tests the distinction between end-user AI assistants (Gemini for Workspace) and developer platforms (Vertex AI, Colab Enterprise), tricking candidates into choosing a developer tool when the scenario describes direct productivity features inside Workspace apps.

How to eliminate wrong answers

Option B is wrong because Colab Enterprise is a managed notebook environment for data scientists and ML engineers to build and train models, not an end-user assistant embedded in Workspace apps. Option C is wrong because Vertex AI Agent Builder is a platform for building custom conversational agents and search applications using enterprise data, requiring development effort and not providing out-of-the-box Workspace integration. Option D is wrong because NotebookLM is a research and note-taking assistant that grounds responses in user-uploaded sources; it does not integrate into Gmail, Docs, or Sheets for drafting and chart creation.

775
MCQeasy

A company wants to ensure that its generative AI application complies with the GDPR right to erasure (right to be forgotten) for user data used in model fine-tuning. What is the best approach?

A.Store data with expiration dates and automatically delete after a set period
B.Maintain a mapping of user identities to training data, and upon request, remove the specific data points and retrain the model
C.Use a broad data deletion request on all training data
D.Implement differential privacy during training to prevent memorization
AnswerB

Mapping user identities to their training data points enables precise identification and deletion of the specific records on request, then retraining removes their influence from the model. This directly satisfies GDPR's right to erasure, which requires effective deletion rather than mere suppression.

Why this answer

Only option B fully addresses GDPR compliance by identifying and removing specific user data from the training set, then retraining. The other options do not effectively erase the user's influence from the model.

776
MCQeasy

A small marketing agency wants to let its non-technical staff draft blog posts using generative AI without writing any code or managing infrastructure. The agency already uses Google Workspace. Which Google Cloud offering should they adopt to meet this need most directly?

A.Gemini for Google Workspace
B.Cloud Natural Language API
C.Vertex AI Agent Builder
D.Vertex AI Model Garden
AnswerA

Gemini for Google Workspace embeds generative AI directly into Gmail, Docs, Slides, and other productivity apps the agency already uses. Staff can draft and refine blog content inside Docs with no coding, no model deployment, and no separate AI platform to administer, making it the most direct fit for non-technical users who need writing assistance immediately.

Why this answer

Gemini for Google Workspace brings generative assistance into the productivity applications the agency already relies on, so non-technical staff can draft and edit content without any coding or infrastructure work. The other services either require AI engineering effort or perform analysis rather than content generation.

Exam trap

The trap here is assuming any Google Cloud AI service can satisfy a content-generation need, when several of them only analyze or classify existing text.

777
MCQmedium

A data scientist is using Vertex AI to fine-tune a Gemini model for a specialized legal document summarization task. They have a small set of labeled examples (200 pairs). Which fine-tuning method is MOST cost-effective and likely to perform well?

A.Full fine-tuning of all model parameters
B.Adapter-based fine-tuning (e.g., LoRA)
C.Training a small custom model from scratch
D.Prompt engineering with few-shot examples only
AnswerB

Adapter-based fine-tuning such as LoRA freezes the base weights and trains small injected matrices, so only a fraction of parameters update. With just 200 labelled pairs, this avoids overfitting and full fine-tuning's cost, satisfying the stem's small-dataset, cost-effective constraint.

Why this answer

Adapter-based fine-tuning (like LoRA) updates only a small fraction of parameters, making it efficient with small datasets and low cost, while still adapting the model to the task.

778
MCQeasy

What is the primary purpose of the temperature parameter when configuring a generative AI model?

A.Adjusts the number of highest-probability tokens considered at each step
B.Controls the diversity of the output by scaling the log probabilities before sampling
C.Specifies the minimum probability threshold for token selection
D.Sets the maximum number of tokens in the response
AnswerB

Temperature scales the logits (log probabilities) before the softmax sampling step, so higher values flatten the distribution and yield more diverse tokens, while lower values sharpen it for deterministic output. This directly governs output diversity, the parameter's defined role.

Why this answer

Temperature controls the randomness of token selection. Higher temperature increases creativity/diversity; lower temperature produces more deterministic and focused responses.

779
MCQhard

A company is using generative AI for code generation and wants to evaluate the quality of generated code for security vulnerabilities. Which metric is most appropriate?

A.BLEU score
B.Automatic static analysis
C.Human evaluation
D.Perplexity
AnswerB

Automatic static analysis scans generated code for insecure patterns such as injection flaws, hardcoded credentials and unsafe API calls, giving an objective, repeatable security measure. Functional or similarity metrics cannot detect vulnerabilities, so static analysis directly satisfies the requirement to evaluate security quality.

Why this answer

(Automatic static analysis) is correct because it directly scans code for security vulnerabilities, making it the most appropriate metric for this purpose. Option A (BLEU score) measures text similarity, not security. Option C (Human evaluation) is subjective and less scalable.

Option D (Perplexity) measures language model confidence, not code security.

780
MCQeasy

Which Google Cloud AI service is specifically designed for extracting structured data from scanned documents, such as invoices and receipts?

A.Document AI
B.Natural Language AI
C.Translation AI
D.Vision AI
AnswerA

Document AI applies specialised parsers and OCR models purpose-built for scanned invoices and receipts, returning structured fields rather than raw text. This directly satisfies the stem's constraint of extracting structured data from scanned documents, unlike general-purpose vision or language services that lack invoice-specific schema extraction.

Why this answer

Document AI is the correct answer because it is purpose-built for understanding and extracting structured data from unstructured documents like invoices, receipts, and forms. It uses specialized processors (e.g., the Invoice Parser or Expense Parser) that combine optical character recognition (OCR) with natural language understanding and machine learning models trained on document layouts, enabling it to output structured fields such as vendor name, total amount, and line items.

Exam trap

The trap here is that candidates often confuse Vision AI’s general OCR capability with Document AI’s specialized document understanding, overlooking that Vision AI cannot natively extract structured fields like line items or totals without extensive custom coding.

How to eliminate wrong answers

Option B is wrong because Natural Language AI is designed for analyzing and extracting insights from text (e.g., sentiment, entity recognition, syntax analysis), not for processing scanned document images or extracting structured data from forms. Option C is wrong because Translation AI is a neural machine translation service that converts text between languages, with no capability to parse scanned documents or extract structured fields. Option D is wrong because Vision AI provides general-purpose image analysis (e.g., object detection, OCR for text extraction), but it lacks the specialized document understanding and pre-trained models for extracting structured data from invoices and receipts that Document AI offers.

781
MCQmedium

A media company wants to let its editors query a large archive of internal video transcripts using everyday conversational questions, and the app must return grounded answers that cite the exact source clips. The team has no machine learning engineers and wants the least operational overhead. Which Google Cloud offering should they use?

A.BigQuery ML with a remote Gemini model reference
B.Vertex AI Pipelines with a custom retrieval component
C.Vertex AI Search with a connected transcript data store
D.Gemini via the Gemini API in a stateless prompt loop
AnswerC

Vertex AI Search provides out-of-the-box grounded retrieval over indexed enterprise content, and its data stores support unstructured sources such as transcripts, returning answers with citations to source documents. Because it is a managed offering, no model training or serving infrastructure is required, matching the low-overhead and grounded-citation requirements of the editorial archive scenario.

Why this answer

A managed search-and-grounding service is the right fit when business users need conversational answers over an indexed corpus with citations and no ML engineering effort. Indexing transcripts into a data store and letting the search service handle retrieval, grounding and citation generation satisfies both the accuracy and low-overhead constraints. Custom pipelines or raw model calls push that work back onto the team.

Exam trap

The trap here is assuming that calling a powerful Gemini model directly automatically grounds answers in a private transcript archive, when grounding requires an indexed data store.

782
MCQhard

An enterprise is comparing Google Cloud Vertex AI vs AWS Bedrock vs Azure OpenAI for a generative AI application. Which unique Google differentiator allows the model to reference up-to-date web information and private data with managed retrieval?

A.Vertex AI Agent Builder with search grounding
B.TPU availability
C.Integration with Google Workspace
D.Multimodal understanding
AnswerA

Vertex AI Agent Builder with search grounding lets models retrieve current web results and private enterprise data through managed retrieval, satisfying the up-to-date information constraint without custom pipelines. This grounding capability is Google's distinctive differentiator versus Bedrock and Azure OpenAI.

Why this answer

Vertex AI Agent Builder with search grounding is the unique Google differentiator that enables a generative AI model to reference up-to-date web information and private data through managed retrieval. Search grounding in Vertex AI allows the model to retrieve relevant documents from Google Search or a private corpus (via Vertex AI Search) and incorporate that information into its responses, ensuring factual accuracy and recency. This is a managed service that abstracts away the complexity of building a retrieval-augmented generation (RAG) pipeline, making it a key differentiator for enterprise applications requiring dynamic, context-aware answers.

Exam trap

The trap here is confusing hardware or ecosystem advantages (TPUs, Workspace integration) with a managed retrieval capability; candidates might pick TPU availability thinking it's unique to Google, but the question specifically asks for a differentiator that enables referencing up-to-date web and private data with managed retrieval.

How to eliminate wrong answers

Option B is wrong because TPU availability is a hardware acceleration advantage for training and inference, not a retrieval mechanism for grounding model outputs in external data. Option C is wrong because integration with Google Workspace provides access to productivity tools and data but does not inherently enable managed retrieval of up-to-date web information or private data for grounding; it's an ecosystem integration, not a retrieval service. Option D is wrong because multimodal understanding refers to the model's ability to process multiple input types (text, image, etc.), not to retrieve and reference external information; it's a model capability, not a retrieval-augmented generation feature.

783
MCQhard

A retail company has deployed a generative AI chatbot for customer support. They notice that the model sometimes provides incorrect product information. The team wants to ground the model's responses in their product catalog to improve accuracy. Which Vertex AI feature should they enable?

A.Use Vertex AI RAG Engine
B.Enable Grounding with Google Search
C.Increase the model's temperature setting
D.Fine-tune the model with product catalog updates
AnswerA

Vertex AI RAG Engine retrieves relevant passages from the product catalogue and injects them into the prompt, grounding responses in authoritative source data rather than relying on the model's parametric memory. This directly addresses the accuracy constraint by anchoring outputs to current product information.

Why this answer

Vertex AI RAG Engine enables retrieval-augmented generation by integrating with a private data store (e.g., the company's product catalog). It fetches relevant documents from the catalog at inference time, grounding the model's responses in verified internal data and reducing hallucinations. Option B (Grounding with Google Search) retrieves from public web data, not a private catalog, so it cannot fulfill the requirement for accurate product-specific information.

Options C and D do not address dynamic retrieval from a catalog: temperature affects creativity, not accuracy, and fine-tuning provides only static knowledge updates.

Exam trap

Candidates often confuse Vertex AI Grounding with Google Search (public data) with Vertex AI RAG Engine (custom data store). For private catalogs, RAG Engine is the correct choice, not the public grounding option.

How to eliminate wrong answers

Option A is wrong because Vertex AI RAG Engine (Retrieval-Augmented Generation) is a framework for building custom retrieval pipelines, but it requires additional setup and indexing of the product catalog, whereas Grounding with Google Search provides a simpler, out-of-the-box solution for grounding responses in external data. Option C is wrong because increasing the model's temperature setting would make responses more random and creative, which is the opposite of what is needed to improve accuracy and reduce incorrect product information. Option D is wrong because fine-tuning the model with product catalog updates would require retraining the model on static data, which is inefficient for dynamic catalogs and does not guarantee real-time grounding; it also risks catastrophic forgetting and does not leverage Vertex AI's built-in grounding capabilities.

784
MCQmedium

A team uses Vertex AI Generative AI Studio to tune a model via RLHF. After tuning, the model outputs are bland. What likely went wrong?

A.Insufficient training data
B.Too many training steps
C.Low temperature during evaluation
D.Reward model overfits to generic responses
AnswerD

RLHF rewards are learned from human preference data; if that reward model overfits to safe, generic completions, it assigns them inflated scores. Policy optimisation then maximises those scores, collapsing output diversity. The blandness stems from the reward signal, not the base model or sampling temperature.

Why this answer

When the reward model overfits to generic responses, it assigns high rewards to safe, non-committal outputs, causing the RLHF-tuned model to converge toward bland, uninformative text. This happens because the reward model learns to prefer patterns that are statistically common in the training data rather than genuinely high-quality or diverse responses, directly leading to the 'bland' output described.

Exam trap

Google often tests the misconception that bland outputs are caused by inference-time parameters like temperature, rather than by the reward model overfitting during the RLHF training phase.

How to eliminate wrong answers

Option A is wrong because insufficient training data typically causes underfitting or poor generalization, not specifically bland outputs; RLHF can still produce diverse responses if the reward model is well-calibrated. Option B is wrong because too many training steps usually lead to overfitting or reward hacking, where the model exploits the reward model for extreme or repetitive outputs, not blandness. Option C is wrong because low temperature during evaluation reduces randomness and can make outputs more deterministic, but it does not inherently cause blandness; the model would still produce coherent, contextually appropriate responses, just with less creativity.

785
Multi-Selecteasy

A project manager wants to measure the business impact of a GenAI code review tool. Which THREE metrics should they track to evaluate ROI? (Choose 3)

Select 3 answers
A.Total tokens consumed by the GenAI model
B.Developer satisfaction score
C.Average time saved per code review
D.Number of lines of code generated
E.Defect escape rate (bugs found in production)
AnswersB, C, E

Measures team acceptance and morale, which affects long-term adoption and productivity.

Why this answer

Developer satisfaction score (B) is a critical metric for GenAI code review ROI because it directly measures user adoption and perceived value. If developers find the tool frustrating or inaccurate, they will bypass it, negating any potential time savings or defect reduction. High satisfaction correlates with sustained usage, which is necessary for long-term return on investment.

Exam trap

The Google Gen AI exam often tests the distinction between cost/usage metrics (like tokens consumed) and value/outcome metrics (like time saved or defect reduction), leading candidates to mistakenly select operational metrics instead of business impact metrics.

786
MCQeasy

A company is building a customer support chatbot using Vertex AI Agent Builder. They want the agent to answer questions based on their internal knowledge base. Which feature should they use?

A.Grounding with Google Search
B.Grounding with enterprise data stores
C.Model tuning
D.Prompt engineering
AnswerB

Grounding with enterprise data stores connects the agent to the company's internal knowledge base, letting responses cite retrieved documents rather than rely on parametric memory. This directly satisfies the requirement to answer from proprietary content, reducing hallucination without retraining the underlying model.

Why this answer

Vertex AI Agent Builder supports grounding with enterprise data stores, which allows the agent to retrieve and answer questions based on the company's internal knowledge base (e.g., documents, PDFs, websites) without relying on public web search. This ensures responses are grounded in proprietary, controlled data, making it the correct choice for a customer support chatbot that needs to reference internal policies or product documentation.

Exam trap

The trap here is that candidates may confuse 'grounding with Google Search' (public web) with 'grounding with enterprise data stores' (private data), assuming any grounding feature works for internal knowledge, but only the enterprise data store option provides the necessary data isolation and access control.

How to eliminate wrong answers

Option A is wrong because Grounding with Google Search uses public web data, not the company's internal knowledge base, which could introduce irrelevant or unverified information and violates data privacy requirements. Option C is wrong because model tuning (e.g., fine-tuning a foundation model) adjusts model weights on custom datasets, but it is not designed for real-time retrieval from a specific knowledge base; it also requires significant compute and may not scale for dynamic content. Option D is wrong because prompt engineering involves crafting input prompts to guide model behavior, but it does not provide a mechanism to retrieve and ground answers in a specific enterprise data store; without grounding, the model may hallucinate or rely on its training data.

787
Multi-Selecteasy

A company is adopting generative AI for customer support. Which TWO strategies should they implement to manage risks related to brand reputation?

Select 2 answers
A.Establish a human-in-the-loop escalation process for sensitive interactions.
B.Publish a disclaimer that the AI may make mistakes.
C.Implement automated monitoring for toxic or off-brand language.
D.Deploy the model without any content filters to maximize helpfulness.
E.Disable customer support AI entirely to avoid any risk.
AnswersA, C

Sensitive or ambiguous queries can produce harmful or off-brand replies, so routing them to a human before the response reaches the customer prevents reputational damage. This satisfies the brand-reputation risk constraint by keeping a person accountable for high-stakes interactions.

Why this answer

Option A is correct because a human-in-the-loop escalation process ensures that sensitive or high-stakes customer interactions are reviewed by a person before or during resolution, which directly protects brand reputation by preventing an AI from delivering inappropriate, harmful, or incorrect responses in delicate situations. Option C is correct because automated monitoring for toxic or off-brand language allows the company to detect and remediate reputation-damaging outputs in real time or near real time, maintaining consistent brand voice and catching harmful content that could erode customer trust. Option B is not a sufficient risk-management strategy because a disclaimer only shifts liability and does not prevent reputational harm from bad AI outputs.

Option D is incorrect because deploying a model without content filters increases the risk of toxic, biased, or off-brand responses, which is the opposite of reputation risk management. Option E is incorrect because disabling the AI entirely avoids risk only by eliminating the business benefit, which is not a viable strategy for a company adopting generative AI for customer support.

Exam trap

Google Cloud often tests the distinction between passive risk communication (like disclaimers) and active risk mitigation (like human-in-the-loop or automated monitoring), trapping candidates who think a disclaimer is sufficient to manage brand reputation risk.

788
MCQhard

A financial services firm is deploying a GenAI-powered contract analysis tool. The tool must extract key clauses and flag risky language. Which strategy BEST ensures structured, machine-readable output that downstream systems can parse?

A.Ask the model to write a summary of the contract in natural language
B.Fine-tune the model on a dataset of contracts with clause labels
C.Use a few-shot prompt with examples of JSON output containing the desired fields
D.Rely on the model's pre-trained ability to extract clauses without any formatting instructions
AnswerC

Few-shot examples demonstrating the exact JSON schema constrain the model's output format, yielding machine-readable fields downstream systems can parse reliably. This satisfies the structured-output constraint more directly than free-text prompting, which risks inconsistent or unparseable responses.

Why this answer

Using a few-shot prompt with JSON output examples directly instructs the model to produce structured, machine-readable data. This approach leverages the model's in-context learning ability to follow a specific schema, ensuring downstream systems can parse the extracted clauses without additional transformation. It balances flexibility and precision without requiring costly fine-tuning or relying on unreliable free-form text.

Exam trap

Google often tests the misconception that fine-tuning (Option B) is the only way to achieve structured output, when in fact few-shot prompting with JSON examples can provide a more flexible and cost-effective solution for many business use cases.

How to eliminate wrong answers

Option A is wrong because asking for a natural language summary produces unstructured text that downstream systems cannot reliably parse for key clauses and risk flags, requiring additional NLP processing. Option B is wrong because fine-tuning on labeled contracts improves extraction accuracy but does not guarantee structured output; the model may still return free-form text unless explicitly prompted for a format like JSON. Option D is wrong because relying on the model's pre-trained ability without formatting instructions leads to inconsistent, ad-hoc responses that vary in structure and completeness, making automated parsing impossible.

789
MCQmedium

A chatbot built with Vertex AI PaLM API often provides outdated information about company policies because the training data is months old. Which approach should the team use?

A.Implement grounding by connecting to a knowledge base of current policies.
B.Use prompt engineering to instruct the model to say 'I don't know' if unsure.
C.Increase the context window to include more history.
D.Fine-tune the model on the latest policy documents.
AnswerA

Grounding retrieves current policy documents at inference time and passes them as context, so responses reflect the live knowledge base rather than the model's stale training data. This satisfies the requirement for up-to-date company policy information without retraining the PaLM model.

Why this answer

Grounding connects the PaLM API to a live, authoritative knowledge base (e.g., Cloud Storage, BigQuery, or Vertex AI Search) containing the latest company policies. This allows the model to retrieve and cite current information at inference time without retraining, directly solving the staleness issue. Grounding is the recommended approach in Vertex AI for ensuring factual, up-to-date responses from a foundation model.

Exam trap

In the Google Gen AI Leader exam, a common trap is confusing grounding (dynamic knowledge injection at inference time) with fine-tuning (static model update). Candidates often assume fine-tuning is the best solution for real-time accuracy, but grounding is the correct approach when policies change frequently.

How to eliminate wrong answers

Option B is wrong because prompt engineering to say 'I don't know' does not provide the model with current policy data; it only changes the model's refusal behavior, leaving outdated information uncorrected. Option C is wrong because increasing the context window does not introduce new or updated knowledge; it only allows the model to consider more of the conversation history, which does not address stale training data. Option D is wrong because fine-tuning on the latest policy documents would require significant time, cost, and labeled data, and the model would still be static until the next fine-tuning cycle; grounding provides a dynamic, real-time solution without retraining.

790
MCQmedium

A team is using a generative AI model to create marketing copy. They want the responses to be more focused and less random. Which parameter should they adjust?

A.Decrease temperature
B.Increase top-k
C.Decrease context window
D.Increase temperature
AnswerA

Temperature scales the sampling distribution's randomness; lowering it sharpens probabilities toward high-likelihood tokens, producing more focused, deterministic marketing copy. Raising temperature increases variation, so decreasing it satisfies the stem's requirement for less random, more focused responses.

Why this answer

Decreasing the temperature parameter reduces the randomness of the model's token selection by lowering the probability of sampling less likely tokens. This makes the output more focused and deterministic, which is ideal for marketing copy that needs to stay on-brand and consistent.

Exam trap

Google often tests the misconception that increasing temperature or top-k makes outputs more focused, when in fact both increase randomness and diversity.

How to eliminate wrong answers

Option B is wrong because increasing top-k would actually increase the diversity of token selection by allowing more high-probability tokens to be considered, making responses less focused. Option C is wrong because decreasing the context window limits the amount of input text the model can reference, which can reduce coherence and relevance, not improve focus. Option D is wrong because increasing temperature amplifies randomness, making outputs more creative but less predictable and focused.

791
Multi-Selecthard

A company is deploying a Gemini-based application and needs to ensure low latency for real-time user interactions. They also want to reduce cost. Which THREE strategies should they consider? (Select 3)

Select 3 answers
A.Use Gemini 1.5 Flash instead of Pro
B.Implement response caching for common queries
C.Increase the model's max output tokens to ensure comprehensive answers
D.Use full fine-tuning to make the model faster
E.Keep the context window as short as possible by trimming input
AnswersA, B, E

Gemini 1.5 Flash is a smaller, distilled model offering substantially lower latency and per-token cost than Pro, directly satisfying both the real-time interaction and cost-reduction constraints. It suits high-volume, less complex tasks where Pro's deeper reasoning is unnecessary.

Why this answer

Option A is correct because Gemini 1.5 Flash is a lighter, faster model than Gemini 1.5 Pro, delivering lower latency for real-time interactions and costing less per token, which directly addresses both the latency and cost goals. Option B is correct because caching responses to frequently repeated queries avoids redundant model invocations, cutting both latency (cache hits return instantly) and cost (fewer billed tokens). Option E is correct because transformer inference cost and latency scale with input length, so trimming the context window to only the necessary tokens reduces processing time and token charges.

Option C is not appropriate because increasing max output tokens generates longer responses, raising latency and cost rather than reducing them. Option D is not appropriate because full fine-tuning is expensive and does not inherently make inference faster; latency gains typically come from smaller/distilled models or optimized serving, not full fine-tuning.

Exam trap

A common misconception tested in this exam is that increasing output tokens or fine-tuning improves speed, when in reality these actions increase computational load or add overhead, making them counterproductive for latency and cost goals.

792
Multi-Selecthard

A company is using a generative AI model to create personalized email responses to customer inquiries. The responses sometimes contain factual errors or irrelevant information. The company wants to improve the accuracy and relevance of the responses. Which TWO techniques should they use? (Choose two.)

Select 2 answers
A.Use prompt engineering to provide clear instructions and examples of desired responses.
B.Implement retrieval-augmented generation (RAG) to ground responses in a knowledge base.
C.Fine-tune the model on a dataset of customer inquiries and responses.
D.Increase the model's temperature to make responses more creative.
E.Reduce the model's max output tokens to limit response length.
AnswersA, B

Prompt engineering with explicit instructions and few-shot examples can guide the model to generate more accurate and relevant responses. By specifying the desired format, tone, and content, and providing examples, the model can better align its outputs with the company's requirements, reducing errors and irrelevance.

Why this answer

Retrieval-augmented generation (RAG) grounds the model's responses in a verified knowledge base, ensuring factual accuracy and relevance. Prompt engineering with clear instructions and examples further guides the model to produce desired outputs. Together, these techniques directly target the issues of factual errors and irrelevance without the overhead of fine-tuning.

Exam trap

The trap here is assuming that any model adjustment will improve quality, when in fact techniques like increasing temperature or limiting tokens can worsen accuracy or fail to address the root causes.

793
MCQeasy

Which feature in Vertex AI allows users to browse over 300 foundation models and deploy them with minimal code?

A.Vertex AI Studio
B.Vertex AI Agent Builder
C.Vertex AI Model Garden
D.Vertex AI Pipelines
AnswerC

Vertex AI Model Garden provides a curated catalogue of over 300 foundation models, including Google, open-source and partner models, deployable with minimal code. This directly matches the stem's requirement to browse and deploy foundation models with little configuration effort.

Why this answer

Vertex AI Model Garden, is correct because it provides a centralized repository where users can browse, discover, and deploy over 300 foundation models—including first-party, open-source, and third-party models—with minimal code. It abstracts away infrastructure complexity by offering pre-built deployment templates and one-click integration with Vertex AI endpoints, enabling rapid experimentation and production deployment without deep ML engineering overhead.

Exam trap

The trap here is that candidates confuse Vertex AI Studio's prompt engineering capabilities with Model Garden's model discovery and deployment functionality, leading them to select A instead of C.

How to eliminate wrong answers

Option A is wrong because Vertex AI Studio is a low-code environment for prototyping and tuning generative AI models, but it does not serve as a model repository for browsing and deploying over 300 foundation models; it focuses on prompt design and model customization. Option B is wrong because Vertex AI Agent Builder is designed for creating conversational agents and search experiences using pre-built components, not for browsing and deploying a broad catalog of foundation models. Option D is wrong because Vertex AI Pipelines is an orchestration service for building and managing ML workflows, not a model discovery and deployment interface.

794
MCQmedium

A global logistics firm wants to add a generative AI feature that drafts replies to customer shipment inquiries. The team must prove business value to executives within one quarter, keep engineering effort low, and later swap in a different model without rewriting the application. Which design decision best supports those goals?

A.Embed model-specific SDK calls directly in the application so each provider's latest features are immediately available.
B.Deploy an open-weights model on a single virtual machine and expose it through a custom REST endpoint.
C.Call Gemini through the Vertex AI API and isolate prompt construction and model selection behind an internal abstraction layer.
D.Train a bespoke language model on historical shipment correspondence and serve it from a self-managed GPU cluster.
AnswerC

Vertex AI provides a managed, enterprise-ready endpoint for Gemini, and placing prompt building and model selection behind an internal abstraction lets the team change models or parameters with minimal application changes. This keeps initial engineering effort low while preserving future flexibility, supporting a one-quarter value demonstration.

Why this answer

Using a managed Vertex AI endpoint for Gemini accelerates delivery and removes infrastructure work, while an internal abstraction layer for prompts and model selection decouples the application from any single provider. That combination lets the firm demonstrate value quickly and later change models by adjusting one integration point rather than rewriting business logic.

Exam trap

The trap here is equating rapid delivery with tight coupling to one model's SDK, when the stated requirement to swap models later makes that coupling a liability.

795
MCQmedium

A retail company wants to build an internal assistant that answers employee questions using the company's HR policy documents. The documents change frequently, and the company does not want to retrain a model every time a policy is updated. Which approach best meets these requirements on Google Cloud?

A.Fine-tune a foundation model on the HR documents each time a policy changes
B.Increase the model's context window and paste all HR documents into every prompt
C.Deploy the model with a lower temperature and rely on its pretrained HR knowledge
D.Use retrieval-augmented generation (RAG) with a vector store of the HR documents
AnswerD

RAG retrieves relevant passages from an indexed vector store at query time and supplies them to the model as context, so updated documents only need re-indexing, not retraining. This satisfies the frequent-change requirement and grounds answers in authoritative sources, reducing hallucination. It is the standard Google Cloud pattern using Vertex AI Search or a vector database with embeddings.

Why this answer

Retrieval-augmented generation separates knowledge from the model: documents are embedded and indexed, and relevant chunks are retrieved and passed as context for each query. Updating a policy means re-indexing that document, not retraining, which directly meets the frequent-change requirement. It also grounds responses in source material, improving accuracy and enabling citations, which is essential for an HR policy assistant.

Exam trap

The trap here is conflating fine-tuning with knowledge injection, when fine-tuning is for behavior and style, not for frequently changing factual content.

796
Multi-Selecteasy

A company is using Vertex AI generative models for a high-volume text summarization service. Which two strategies can reduce operational costs?

Select 2 answers
A.Increase the model's max output tokens to 2048.
B.Implement retry logic with exponential backoff.
C.Lower the temperature parameter to 0.
D.Use batch prediction instead of online prediction.
E.Reduce the size of the model (e.g., switch from text-bison@002 to text-bison-light).
AnswersD, E

Batch prediction has lower per-request cost for large jobs compared to online prediction.

Why this answer

Batch prediction reduces costs by processing multiple requests in a single batch job, which avoids the per-request overhead and idle compute time associated with online prediction. This is especially cost-effective for high-volume, non-real-time workloads like text summarization, as you pay only for the compute time used during the batch job rather than for each individual inference.

Exam trap

Google Cloud often tests the misconception that adjusting inference parameters like temperature or output length can reduce costs, when in reality only reducing model size or switching to batch processing directly lowers operational expenses.

797
MCQeasy

A marketing team needs to generate personalized email campaigns for thousands of customers. They want to maintain brand tone consistency and avoid manual writing. Which GenAI approach is BEST suited?

A.Use Vertex AI Studio with prompt design and few-shot examples in the prompt
B.Fine-tune a small model on brand guidelines only
C.Embed a rules-based template engine with no AI
D.Train a custom model from scratch on past campaigns
AnswerA

Vertex AI Studio supports prompt design with few-shot examples, letting the team embed brand-tone exemplars so generated emails stay consistent across thousands of customers without manual writing. This satisfies both the personalisation scale and tone-consistency constraints in the stem.

Why this answer

Vertex AI Studio enables prompt engineering with few-shot examples, allowing the team to generate personalized emails while maintaining brand tone consistency without fine-tuning or custom training. This approach leverages a pre-trained large language model (LLM) with carefully designed prompts that include brand guidelines and a few examples, ensuring output adheres to the desired style and context. It avoids the overhead of fine-tuning or building custom models, making it ideal for rapid deployment and iterative refinement.

Exam trap

Google often tests the misconception that fine-tuning or custom training is always necessary for domain-specific tasks, when in fact prompt engineering with few-shot examples can achieve comparable results with far less effort and cost.

How to eliminate wrong answers

Option B is wrong because fine-tuning a small model on brand guidelines only may lead to catastrophic forgetting or insufficient generalization, as the model might overfit to the narrow dataset and lose the broad language understanding needed for diverse customer personalization. Option C is wrong because a rules-based template engine cannot adapt to the nuanced, context-aware personalization required for thousands of unique customers; it would produce rigid, repetitive content that fails to capture brand tone dynamically. Option D is wrong because training a custom model from scratch on past campaigns is resource-intensive, requires massive labeled datasets, and is unnecessary when pre-trained models with prompt engineering can achieve the same goal more efficiently.

798
MCQmedium

A data scientist is fine-tuning a foundation model for a specialized legal document summarization task. The labeled dataset is only 5,000 examples. Which fine-tuning technique would be MOST efficient to adapt the model without catastrophic forgetting and with minimal computational cost?

A.Low-Rank Adaptation (LoRA)
B.Reinforcement Learning from Human Feedback (RLHF)
C.Full supervised fine-tuning of all model parameters
D.In-context learning with few-shot examples
AnswerA

Low-Rank Adaptation freezes the pretrained weights and injects trainable rank-decomposition matrices into each layer, so only a small fraction of parameters update. With just 5,000 examples, this slashes computational cost and memory while preserving the original weights, directly preventing catastrophic forgetting.

Why this answer

LoRA (Low-Rank Adaptation) is an adapter-based method that trains only a small number of added parameters, making it efficient and less prone to catastrophic forgetting compared to full fine-tuning. Supervised fine-tuning full model is expensive; RLHF is for alignment after fine-tuning; in-context learning requires no training but may not suffice.

799
MCQmedium

A data scientist fine-tunes a large language model on Vertex AI but gets poor results on validation data. What is the most likely cause?

A.Incorrect learning rate
B.Insufficient training data
C.Using wrong model family
D.Overfitting due to too many epochs
AnswerB

Fine-tuning adapts an existing model rather than teaching new knowledge from scratch, so too few examples leave it unable to learn the target pattern. Insufficient training data therefore produces poor validation results, matching the stem's symptom.

Why this answer

Fine-tuning a large language model on Vertex AI with poor validation results is most likely due to insufficient training data. Large language models have billions of parameters and require a substantial amount of high-quality, task-specific data to effectively adapt to a new domain or task; without enough examples, the model cannot learn the desired patterns and will perform poorly on unseen data.

Exam trap

The trap here is that candidates often assume hyperparameter tuning (like learning rate) is the primary cause of poor fine-tuning results, but in generative AI, data quantity and quality are the most common bottlenecks, especially when using pre-trained models on Vertex AI.

How to eliminate wrong answers

Option A is wrong because an incorrect learning rate typically causes training instability (e.g., loss divergence or slow convergence) rather than consistently poor validation results, and Vertex AI's default hyperparameters are often reasonable. Option C is wrong because using the wrong model family (e.g., choosing a text generation model for a classification task) would likely cause immediate, obvious failures or mismatches in output format, not just poor validation performance after fine-tuning. Option D is wrong because overfitting due to too many epochs would manifest as high training accuracy with low validation accuracy, but the question states poor results on validation data without mentioning training performance, and overfitting is less likely with insufficient data (the model would underfit instead).

800
MCQeasy

A project manager wants to track the ROI of a generative AI feature that assists customer support agents. Which metric is MOST directly tied to productivity improvement?

A.Adoption rate of the AI tool
B.Customer satisfaction (CSAT) score
C.Average handle time (AHT) per ticket
D.Cost per API call
AnswerC

Average handle time per ticket measures how long agents spend resolving each contact. Generative AI assistance that drafts responses or surfaces knowledge directly reduces that duration, making AHT the metric most tightly coupled to agent productivity.

Why this answer

Average handle time (AHT) directly measures the time agents spend per interaction, so a reduction indicates productivity gain. CSAT measures satisfaction, not efficiency. Cost per API call is a cost metric.

Adoption rate measures usage.

801
MCQeasy

Which Google AI milestone introduced the Transformer architecture that underpins modern LLMs?

A.AlphaGo
B.Transformer paper
C.AlphaFold
D.BERT
AnswerB

The 2017 paper "Attention Is All You Need" introduced the Transformer, replacing recurrent and convolutional layers with self-attention to process sequences in parallel. This architecture underpins modern LLMs, satisfying the stem's requirement for the Google milestone that originated the Transformer.

Why this answer

The Transformer architecture, which is the foundational technology behind modern large language models (LLMs) like GPT and BERT, was introduced in the 2017 paper 'Attention Is All You Need' by Vaswani et al. This paper proposed the self-attention mechanism and the encoder-decoder structure that replaced recurrent neural networks (RNNs) and long short-term memory (LSTM) networks, enabling parallelized training and superior handling of long-range dependencies in sequence data.

Exam trap

Google often tests the distinction between the original research paper that introduced a concept (the Transformer paper) and later implementations or applications of that concept (like BERT or GPT), causing candidates to confuse the milestone with its derivative products.

How to eliminate wrong answers

Option A is wrong because AlphaGo is a reinforcement learning-based system for playing the board game Go, not a milestone in neural network architecture for language models. Option C is wrong because AlphaFold is a deep learning model for protein structure prediction, not a foundational architecture for LLMs. Option D is wrong because BERT is a pre-trained language model that itself is built on the Transformer architecture, not the original paper that introduced the Transformer.

802
MCQhard

A financial services firm is deploying a generative AI chatbot on Vertex AI. They need to ensure that the model's responses comply with strict regulatory requirements and do not include personally identifiable information (PII). They want to automatically filter out any PII from both user inputs and model outputs. Which Google Cloud service should they integrate?

A.Cloud Data Loss Prevention (DLP) API
B.Security Command Center
C.Cloud Armor
D.Cloud Identity-Aware Proxy (IAP)
AnswerA

The Cloud Data Loss Prevention (DLP) API is designed to discover, classify, and protect sensitive data such as PII. It can inspect text for PII and redact or mask it. By integrating DLP with the chatbot pipeline, the firm can automatically scan and filter user inputs and model outputs, ensuring compliance. This is the appropriate service for PII detection and redaction in text.

Why this answer

To automatically detect and redact PII from text, the Cloud Data Loss Prevention (DLP) API is the correct service. It can be integrated into the chatbot's processing pipeline to scan both incoming user messages and outgoing model responses, identifying and removing sensitive data according to predefined or custom infoTypes. This ensures regulatory compliance and protects user privacy.

Exam trap

The trap here is confusing security services that protect infrastructure with data-level protection services that inspect and redact content.

803
MCQhard

A developer uses Vertex AI to generate code but the output is not syntactically correct. Which parameter should be adjusted?

A.candidate_count
B.max_output_tokens
C.temperature
D.top_k
AnswerC

Temperature controls sampling randomness: higher values increase diversity but raise the chance of malformed syntax, while lower values make token selection more deterministic. Reducing temperature steers Vertex AI toward the most probable, syntactically valid continuations, directly addressing the incorrect code output.

Why this answer

Temperature controls the randomness of token selection during generation. A high temperature increases the likelihood of less probable tokens, which can lead to syntactically incorrect code. Lowering temperature makes the model more deterministic and conservative, favoring higher-probability tokens that are more likely to form valid syntax.

Exam trap

Google Cloud often tests the misconception that increasing candidate_count or max_output_tokens will improve output quality, when in fact these parameters only affect quantity or length, not the underlying token selection logic that determines syntactic correctness.

How to eliminate wrong answers

Option A is wrong because candidate_count controls how many different response candidates are generated, not the syntactic correctness of any single output. Option B is wrong because max_output_tokens limits the length of the generated text, not the quality or validity of the syntax. Option D is wrong because top_k limits the number of highest-probability tokens considered at each step; while it can affect output quality, it does not directly address syntactic correctness as effectively as temperature does.

804
MCQmedium

A company is using Vertex AI Model Garden to discover and test various foundation models. They need a model that can generate code from natural language. Which model should they select?

A.Chirp
B.Codey
C.Med-PaLM
D.Imagen
AnswerB

Codey is Google's foundation model family purpose-built for code, trained on source code and supporting natural-language-to-code generation. It directly satisfies the stem's requirement to generate code from natural language, unlike general-purpose text or multimodal models.

Why this answer

Codey is Google's family of models specifically designed for code generation, including converting natural language descriptions into code. It is built on the PaLM 2 architecture and is optimized for tasks like code completion, code generation, and code chat, making it the correct choice for generating code from natural language.

Exam trap

The trap here is that candidates may confuse Chirp (audio) or Imagen (image) with multimodal models, mistakenly thinking they can handle code generation, when in fact only Codey is purpose-built for code tasks.

How to eliminate wrong answers

Option A is wrong because Chirp is a speech-to-text model designed for audio transcription, not code generation. Option C is wrong because Med-PaLM is a domain-specific model fine-tuned for medical and healthcare applications, not for generating code. Option D is wrong because Imagen is a text-to-image diffusion model for generating images, not code.

805
MCQmedium

A company wants to use a pre-trained language model for customer support summarization. They need to ensure responses are concise and accurate. Which prompt engineering technique is most effective?

A.Zero-shot prompting
B.Few-shot prompting with examples
C.Chain-of-thought prompting
D.Negative prompting
AnswerB

Few-shot prompting supplies labelled input-output pairs, so the model infers the required summary length and factual style directly from the examples. This steers formatting and conciseness without retraining, satisfying the demand for concise, accurate customer support summaries.

Why this answer

Few-shot prompting (B) is most effective because it provides the model with a small set of example input-output pairs (e.g., a customer query and its concise summary), which guides the model to produce outputs that match the desired format, length, and accuracy. This technique is particularly useful for summarization tasks where consistency and adherence to a specific style are critical, as it reduces ambiguity without requiring fine-tuning.

Exam trap

Google Cloud often tests the misconception that zero-shot prompting is sufficient for all tasks, but the trap here is that candidates overlook the need for explicit guidance in format-sensitive tasks like summarization, where few-shot examples provide the necessary constraint for consistency.

How to eliminate wrong answers

Option A (Zero-shot prompting) is wrong because it relies solely on the model's pre-trained knowledge without any examples, which often leads to inconsistent or overly verbose summaries, especially when the task requires a specific format or level of conciseness. Option C (Chain-of-thought prompting) is wrong because it is designed for multi-step reasoning tasks (e.g., arithmetic or logic problems) and is unnecessary for summarization, where the goal is to condense information rather than reason through steps. Option D (Negative prompting) is wrong because it instructs the model on what to avoid (e.g., 'do not include details'), which can be imprecise and may inadvertently suppress relevant information, making it less reliable than providing positive examples of desired outputs.

806
MCQhard

A financial services firm is deploying a generative AI model to answer customer queries about investment products. They want to ensure the model's responses are based on the most current and authoritative internal documents. Which technique should they implement?

A.Using retrieval-augmented generation (RAG) with a vector database
B.Deploying the model with a larger context window
C.Fine-tuning the model on historical customer interactions
D.Increasing the model's temperature parameter
AnswerA

RAG combines a generative model with a retrieval system that fetches relevant documents from a vector database in real time. By indexing the firm's current authoritative documents, the model can ground its responses in the latest information. This ensures accuracy and reduces hallucinations, directly meeting the requirement for responses based on current internal documents.

Why this answer

Retrieval-augmented generation (RAG) dynamically retrieves relevant documents from a vector database and provides them to the model as context, ensuring responses are grounded in the most current authoritative sources. This is ideal for the financial firm's need to answer queries based on up-to-date internal documents. Other options do not provide real-time grounding.

Exam trap

The trap here is thinking that fine-tuning or a larger context window alone can keep responses current, when they lack dynamic retrieval.

807
MCQmedium

A marketing team is using a generative AI model to create ad copy. They want to ensure the output is creative and diverse, but also relevant to the product. They decide to adjust the model's parameters. Which parameter should they increase to make the output more diverse?

A.Top-k
B.Temperature
C.Top-p
D.Max output tokens
AnswerB

Temperature controls the randomness of the model's predictions. A higher temperature increases diversity by making the probability distribution flatter, so less likely words are chosen more often. This leads to more creative and varied outputs. For ad copy, increasing temperature can help generate more diverse ideas while still being relevant if the prompt is clear.

Why this answer

Temperature is the parameter that directly controls the randomness of the model's output. Increasing temperature makes the output more diverse and creative, which is desired for ad copy generation. While top-k and top-p also influence sampling, temperature is the primary knob for adjusting diversity.

Max output tokens only controls length.

Exam trap

The trap here is confusing parameters that control output length or sampling cutoffs with the one that directly scales randomness, which is temperature.

808
MCQmedium

What is the primary purpose of a system instruction in the Gemini API?

A.Set the model's temperature and top_p
B.Define the overall behavior and constraints for the model
C.Provide few-shot examples for each query
D.Set the maximum output length
AnswerB

System instructions set persistent behavioural guidance and constraints applied across a Gemini API session, shaping tone, role and boundaries regardless of individual prompts. This satisfies the stem's requirement of defining overall model behaviour rather than specifying a single request's content.

Why this answer

The system instruction in the Gemini API is the primary mechanism to define the overall behavior, persona, constraints, and guardrails for the model across all interactions. Unlike per-query parameters, it sets a persistent context that shapes how the model interprets every user prompt, ensuring consistent adherence to rules such as tone, format, or safety policies.

Exam trap

Google Cloud often tests the distinction between persistent system-level instructions and per-request parameters, so the trap here is confusing the system instruction (which defines the model's role and constraints) with generation controls like temperature, top_p, or max tokens, which only affect the style or length of a single response.

How to eliminate wrong answers

Option A is wrong because temperature and top_p are sampling parameters that control randomness and diversity of output, not the overarching behavioral constraints set by a system instruction. Option C is wrong because few-shot examples are typically provided in the user prompt or as part of a structured conversation, not as the primary purpose of a system instruction, which is for persistent context rather than per-query demonstrations. Option D is wrong because maximum output length is a generation parameter that limits token count, not a behavioral or constraint-setting mechanism like a system instruction.

809
Multi-Selectmedium

A data scientist is evaluating how to ground a generative AI model to reduce hallucinations when answering questions about a private knowledge base. Which TWO techniques are most suitable?

Select 2 answers
A.Using a larger model like Gemini Ultra
B.Fine‑tuning on the private knowledge base
C.Increasing the temperature to 0.9
D.Retrieval-Augmented Generation (RAG)
E.Prompt engineering to instruct the model to answer based only on the provided context
AnswersD, E

Retrieval-Augmented Generation retrieves relevant passages from the private knowledge base at query time and injects them into the model's context, so answers are conditioned on actual source documents rather than parametric memory alone. This directly satisfies the grounding constraint, reducing hallucination without retraining or fine-tuning the underlying model.

Why this answer

Option D, Retrieval-Augmented Generation (RAG), is correct because it grounds the model by retrieving relevant passages from the private knowledge base at query time and injecting them into the prompt, so answers are based on actual source documents rather than parametric memory, which directly reduces hallucinations. Option E, prompt engineering instructing the model to answer only from the provided context, is correct because it constrains generation to the supplied grounding text and is the standard companion technique used with RAG to keep responses faithful to retrieved content. Option A, using a larger model like Gemini Ultra, does not by itself ground the model in a private knowledge base and can still hallucinate about proprietary data.

Option B, fine-tuning on the private knowledge base, can teach style and some facts but is costly, must be redone as data changes, and does not reliably prevent hallucinations or provide citations. Option C, increasing the temperature to 0.9, raises randomness and would make hallucinations more likely, not less.

Exam trap

In the context of the Google Generative AI Leader exam, fine-tuning is often mistakenly thought to be the primary method for grounding a model on private data, when in fact RAG is preferred for dynamic or large knowledge bases because it avoids retraining and allows real-time updates without modifying model weights.

810
MCQeasy

Which statement best describes the difference between the Gemini Flash and Gemini Pro models on Vertex AI?

A.Gemini Flash is a distilled version of Gemini Pro that requires fine‑tuning before use
B.Gemini Pro is deployed on Google’s TPU v5p chips, while Flash uses TPU v4
C.Gemini Flash is optimized for speed and cost, while Gemini Pro provides higher quality for complex tasks
D.Gemini Flash is only available for image inputs, while Gemini Pro handles text
AnswerC

Gemini Flash targets low-latency, high-volume workloads where throughput and cost efficiency matter most, whereas Gemini Pro delivers stronger reasoning and output quality for demanding tasks. This directly satisfies the stem's requirement to distinguish the two models along the speed/cost versus capability axis on Vertex AI.

Why this answer

Gemini Flash is specifically designed for low-latency, high-throughput, and cost-efficient inference, making it ideal for high-volume, simpler tasks. In contrast, Gemini Pro is a larger, more capable model that delivers superior quality and reasoning for complex, multi-step tasks, though at higher latency and cost. This distinction is fundamental to the Gemini model family on Vertex AI, where Flash serves as the lightweight, fast option and Pro as the premium, high-quality option.

Exam trap

The trap here is that candidates often assume 'Flash' implies a distilled or pruned version of 'Pro' (like a student model), but in reality, Flash is a distinct model trained from scratch with a different architecture optimized for speed, not a compressed version of Pro.

How to eliminate wrong answers

Option A is wrong because Gemini Flash is not a distilled version of Gemini Pro; it is a separate, independently trained model optimized for speed and cost, and it does not require fine-tuning before use—it is available as a pre-trained model for inference. Option B is wrong because both Gemini Flash and Gemini Pro are deployed on Google's TPU v5p chips; the difference in performance is due to model architecture and size, not the underlying TPU generation. Option D is wrong because Gemini Flash handles both text and image inputs (multimodal), just like Gemini Pro; the limitation to image-only inputs is a misconception.

811
MCQeasy

Which of the following best describes the primary benefit of using Grounding with Google Search when building a GenAI chatbot?

A.It reduces the model's latency by caching responses
B.It enables the model to generate images based on text descriptions
C.It provides fine-tuning capabilities for domain-specific data
D.It allows the model to access real-time information from the internet to reduce hallucinations
AnswerD

Grounding with Google Search retrieves live web results and feeds them into the prompt, so answers reflect current facts rather than stale training data. The stem's constraint is real-time information, which grounding supplies to reduce hallucination.

Why this answer

Grounding with Google Search connects the GenAI chatbot to real-time internet data, allowing it to retrieve current facts and events that the model was not trained on. This reduces hallucinations by ensuring responses are based on verified, up-to-date information rather than relying solely on the model's static training data.

Exam trap

The trap here is that candidates often confuse Grounding (a retrieval-based technique for real-time accuracy) with fine-tuning (a training-based technique for domain adaptation), leading them to select Option C incorrectly.

How to eliminate wrong answers

Option A is wrong because Grounding with Google Search does not cache responses to reduce latency; it introduces additional retrieval latency by querying live search results. Option B is wrong because Grounding is a text-based retrieval mechanism and does not enable image generation, which requires a multimodal model or separate image generation service. Option C is wrong because Grounding is a retrieval-augmented generation (RAG) technique, not a fine-tuning method; fine-tuning adjusts model weights on domain-specific data, whereas Grounding retrieves external data at inference time.

812
MCQmedium

A global bank uses a Gemini model on Vertex AI to generate personalized investment summaries for clients in multiple regions. Compliance requires that the model never recommend products prohibited in a given region. The team wants a control that enforces these rules regardless of how the prompt is phrased. Which approach should they use?

A.Configure safety filters and content moderation thresholds on the model endpoint.
B.Use a pre-call or post-call validation layer that checks the region against an allowed-product list and blocks violations.
C.Fine-tune the model on examples of compliant and non-compliant recommendations.
D.Add a system instruction listing prohibited products for each region.
AnswerB

A validation layer applies deterministic logic: before or after generation, it checks the client's region against an authoritative allowed-product list and rejects any response that violates it. Because the rule lives in code and data rather than in the prompt, it holds regardless of phrasing and can be updated as regulations change. This gives the auditable enforcement compliance demands.

Why this answer

Absolute compliance rules need deterministic enforcement rather than probabilistic guidance. A validation layer that evaluates the region against an allowed-product list before or after the model call blocks prohibited recommendations no matter how the request is phrased, and it produces an auditable record. System instructions, safety filters, and fine-tuning all influence behavior but cannot guarantee that a disallowed product is never surfaced.

Exam trap

The trap here is assuming that a strongly worded system instruction or safety filter is a compliance guarantee, when prompt-level guidance and safety categories cannot deterministically enforce region-specific product rules.

813
MCQmedium

A data scientist fine-tunes a foundation model on customer support transcripts. After evaluation, the model's responses are too formal. Which adjustment during fine-tuning is most likely to make responses more conversational?

A.Increase the batch size to stabilize training.
B.Decrease the number of fine-tuning steps to prevent overfitting.
C.Include examples of informal customer interactions in the fine-tuning data.
D.Use a higher learning rate for faster adaptation.
AnswerC

Fine-tuning adapts a model's style to its training distribution, so the formal tone reflects the transcripts used. Adding informal customer interaction examples shifts that distribution toward conversational phrasing, directly satisfying the requirement to make responses less formal without changing architecture or decoding parameters.

Why this answer

The training data directly influences the tone and style of model outputs. Including examples of informal conversations in the fine-tuning dataset teaches the model the desired conversational tone. Other options affect training dynamics but not the style.

814
MCQhard

A global retailer uses a generative AI model to personalize product recommendations. They need to ensure that customer prompts and responses are not logged for model improvement to meet GDPR data minimization principles. Which configuration should they apply?

A.Use a third-party logging service with data deletion policies
B.Enable data logging for six months and then automatically delete
C.Disable prompt/response logging in Vertex AI endpoint settings
D.Anonymize all customer data before logging
AnswerC

Disabling prompt and response logging at the Vertex AI endpoint prevents customer data being stored for model improvement, directly satisfying GDPR data minimisation by ensuring personal data in prompts is not retained beyond serving the request.

Why this answer

To comply with GDPR data minimization, the retailer should disable prompt/response logging. Vertex AI offers settings to control whether prompts and responses are stored for model improvement.

815
MCQeasy

Which Google AI research organization is responsible for AlphaFold, a breakthrough in protein structure prediction?

A.Google Research
B.Google DeepMind
C.Google Brain
D.X Development
AnswerB

Google DeepMind, formed by merging DeepMind with Google Brain, developed AlphaFold, whose deep learning models predict three-dimensional protein structures from amino acid sequences. This directly satisfies the stem's requirement for the Google AI research organisation behind AlphaFold's breakthrough in protein structure prediction, distinguishing it from Google Research and other subsidiaries.

Why this answer

Google DeepMind is the research lab behind AlphaFold, AlphaCode, and other notable achievements.

816
MCQeasy

Which Google Cloud service provides a managed environment for prompt engineering and model evaluation?

A.AI Platform Notebooks
B.Dialogflow CX
C.Vertex AI Generative AI Studio
D.Cloud Composer
AnswerC

Vertex AI Generative AI Studio supplies a managed workspace for prompt design, tuning and model evaluation, directly satisfying the stem's requirement for a managed prompt engineering and evaluation environment. Its integrated prompt gallery and evaluation metrics remove infrastructure overhead, unlike raw Vertex AI endpoints or unmanaged notebook approaches.

Why this answer

Vertex AI Generative AI Studio is the correct answer because it is a managed service within Vertex AI specifically designed for prompt engineering, model tuning, and evaluation of generative AI models. It provides a no-code interface for testing prompts, comparing model outputs, and iterating on prompt design, directly supporting the workflow described in the question.

Exam trap

The trap here is that candidates may confuse Vertex AI Generative AI Studio with AI Platform Notebooks, assuming that any model development environment supports prompt engineering, when in fact Generative AI Studio is the specialized tool for that purpose.

How to eliminate wrong answers

Option A is wrong because AI Platform Notebooks is a managed Jupyter notebook service for custom model development and training, not a dedicated environment for prompt engineering or model evaluation. Option B is wrong because Dialogflow CX is a conversational AI platform for building chatbots and virtual agents, focused on intent classification and dialogue management, not prompt engineering or model evaluation. Option D is wrong because Cloud Composer is a managed Apache Airflow service for workflow orchestration and scheduling, unrelated to prompt engineering or model evaluation.

817
MCQmedium

A data scientist is using the Vertex AI PaLM API for text generation. They notice that the model occasionally generates toxic content. Which parameter should they adjust to reduce the likelihood of toxic outputs?

A.max_output_tokens
B.temperature
C.top_k
D.safety_settings
AnswerD

safety_settings configures per-category harm thresholds, such as harassment, hate speech and dangerous content, that filter model output. Raising the blocking sensitivity reduces toxic generations, directly addressing the observed behaviour without altering temperature or token limits.

Why this answer

Safety settings in the Vertex AI PaLM API allow you to configure thresholds for filtering harmful content categories (e.g., toxicity, harassment, hate speech). By adjusting these settings, you can block or reduce the likelihood of toxic outputs before they are returned, directly addressing the problem without altering the model's creativity or randomness.

Exam trap

The trap here is that candidates often confuse parameters that control output randomness (temperature, top_k) with those that enforce content safety, leading them to incorrectly select temperature or top_k instead of the dedicated safety_settings parameter.

How to eliminate wrong answers

Option A is wrong because max_output_tokens controls the maximum length of the generated text, not the content safety or toxicity. Option B is wrong because temperature adjusts the randomness of token sampling, influencing creativity but not filtering toxic content. Option C is wrong because top_k limits the number of highest-probability tokens considered at each step, affecting diversity but not safety filtering.

818
MCQmedium

A company deploys a fine-tuned text generation model on Vertex AI Endpoints. They want to monitor for data drift and performance degradation over time. Which GCP service should they integrate?

A.Cloud Monitoring
B.Cloud Logging
C.Vertex AI Experiments
D.Vertex AI Model Monitoring
AnswerD

Vertex AI Model Monitoring detects training-serving skew and prediction drift on deployed endpoints, alerting when input distributions or performance shift. Integrating it satisfies the requirement to track data drift and degradation for the fine-tuned text model over time.

Why this answer

Vertex AI Model Monitoring is the correct choice because it is specifically designed to detect data drift (changes in input data distribution) and feature attribution drift in deployed models, including fine-tuned text generation models on Vertex AI Endpoints. It provides automated alerts when model performance degrades due to shifts in production data, enabling proactive retraining or intervention.

Exam trap

The trap here is that candidates confuse general observability tools (Cloud Monitoring, Cloud Logging) with Vertex AI's purpose-built drift detection service, assuming any monitoring tool can handle model-specific data drift analysis.

How to eliminate wrong answers

Option A is wrong because Cloud Monitoring provides infrastructure-level metrics (e.g., CPU, memory, latency) but does not analyze model input data distributions or detect data drift. Option B is wrong because Cloud Logging captures raw log entries for debugging and auditing, not statistical drift detection or performance degradation analysis. Option C is wrong because Vertex AI Experiments tracks training runs and hyperparameters, not post-deployment monitoring of live endpoints.

819
Multi-Selecthard

A healthcare company is evaluating foundation models for a patient triage assistant. They need to ensure the model is appropriate for medical use and that they can control costs. Which two factors should they consider when selecting and using a foundation model in Vertex AI? (Choose two.)

Select 2 answers
A.The model's ability to generate images for patient education materials.
B.The color scheme of the model's documentation website.
C.The token-based pricing and expected request volume to estimate and control costs.
D.The model's training data cutoff and domain relevance to medical knowledge.
E.The number of GPUs available in the company's on-premises data center.
AnswersC, D

Generative AI usage in Vertex AI is typically priced by tokens processed. Understanding token-based pricing and forecasting request volume allows the company to estimate costs and apply controls like quotas or budget alerts. This is essential for managing spend in a patient triage assistant that may handle many interactions, making it a key selection and usage factor.

Why this answer

For a medical triage assistant, the model's training data cutoff and domain relevance are critical to ensure safe, current guidance. Token-based pricing and expected volume determine cost predictability and control. Together, these factors address the dual requirements of clinical appropriateness and financial management when selecting and using a foundation model in Vertex AI.

Exam trap

The trap here is assuming that infrastructure details like on-premises GPUs matter for a managed service, or that unrelated features like image generation are relevant to a text triage assistant.

820
MCQmedium

A company wants to use generative AI to create short product videos from text descriptions. Which Google Cloud service should they consider?

A.Chirp
B.Imagen
C.Gemini
D.Veo
AnswerD

Veo is Google Cloud's generative video model, converting text prompts directly into short video clips. It satisfies the stem's requirement to create product videos from text descriptions, whereas text-only or image-focused services cannot produce motion footage. This makes Veo the appropriate service for the described scenario.

Why this answer

Veo is Google's video generation model that creates videos from text. Imagen creates images, Chirp creates audio, and Gemini is multimodal but not specialized for video generation.

821
MCQhard

An AI company wants to detect whether text was generated by their own model. Which technology developed by Google is specifically designed for this purpose?

A.SynthID
B.Google's confidential computing
C.Differential privacy
D.Federated learning
AnswerA

SynthID embeds imperceptible digital watermarks into AI-generated content, including text, enabling later detection of whether output originated from Google's models. This watermarking mechanism directly satisfies the stem's requirement to identify the company's own generated text.

Why this answer

SynthID is Google DeepMind's watermarking technology that embeds imperceptible digital watermarks into AI-generated content (text, images, audio, video) so it can later be detected as AI-generated. It is specifically designed for identifying content produced by Google's generative models, matching the question's requirement.

Exam trap

The trap is confusing AI safety/privacy technologies — candidates pick differential privacy or federated learning because they sound like 'responsible AI' tools, but only SynthID is a watermarking/detection technology for AI-generated content.

How to eliminate wrong answers

Option B is wrong because confidential computing protects data in use via hardware-based trusted execution environments; it does not detect AI-generated text. Option C is wrong because differential privacy adds statistical noise to datasets to protect individual privacy, not to identify model outputs. Option D is wrong because federated learning trains models across decentralized data without sharing raw data; it has nothing to do with detecting generated content.

822
Multi-Selecthard

A multinational corporation deploys a generative AI chatbot across multiple regions. They need to comply with GDPR and local data residency requirements. Which THREE actions are necessary?

Select 3 answers
A.Obtain explicit consent from every user before collecting any data
B.Anonymize all training data before fine-tuning
C.Store and process data only in approved geographic regions (data residency controls)
D.Implement prompt and response logging with configurable retention policies
E.Encrypt personal data at rest and in transit using customer-managed encryption keys (CMEK)
AnswersC, D, E

Approved geographic regions satisfy data residency by pinning storage and processing to specific locations, preventing cross-border transfers that would breach GDPR or local mandates. Microsoft Entra ID and regional deployments let you enforce these boundaries, so personal data never leaves the jurisdiction the stem requires.

Why this answer

Option C is correct because GDPR and local data residency laws require that personal data be stored and processed only within approved jurisdictions, so implementing geographic data residency controls (e.g., region-pinned storage and compute) is necessary to prevent cross-border transfers. Option D is correct because GDPR accountability and data minimization principles require tracking what personal data the chatbot processes and retaining it only as long as necessary, which is achieved through prompt/response logging with configurable retention and deletion policies. Option E is correct because GDPR Article 32 mandates appropriate technical measures including encryption of personal data at rest and in transit, and using customer-managed encryption keys (CMEK) gives the organization control over key lifecycle and revocation to meet stricter local requirements.

Option A is not universally required because GDPR consent is only one of several lawful bases (e.g., contract, legitimate interest) and is not needed for every data collection. Option B is not necessary because anonymization is not required for all training data; pseudonymization or other safeguards may suffice, and fully anonymized data falls outside GDPR scope but is not a blanket obligation.

823
MCQeasy

A developer is using a generative AI model to summarize long articles. They want to ensure the summaries are concise and do not exceed a certain length. Which parameter should they adjust to control the maximum length of the generated summary?

A.Temperature
B.Max output tokens
C.Top-p
D.Top-k
AnswerB

Max output tokens sets the maximum number of tokens the model can generate in its response. By lowering this value, the developer can ensure summaries stay within a desired length. This parameter directly caps the output size, making it the correct choice for controlling summary length.

Why this answer

The max output tokens parameter directly limits the number of tokens the model can generate. By setting it to a desired value, the developer can prevent summaries from becoming too long. Other parameters like temperature, top-p, and top-k influence randomness and diversity but do not control output length.

Exam trap

The trap here is confusing parameters that control randomness with those that control output length, such as assuming temperature affects length.

824
MCQmedium

A company is piloting a GenAI feature that summarizes customer support tickets. They want to measure the impact on agent productivity before rolling out to all teams. Which approach BEST evaluates the pilot?

A.Survey agents on their perception of productivity after using the tool
B.Run an A/B test where half the agents use the tool and half do not, then compare average handling time
C.Compare the cost of the API before and after deployment
D.Measure the number of summaries generated per day
AnswerB

An A/B test with a control group isolates the tool's effect on average handling time, giving a measurable comparison against agents without it. This satisfies the stem's requirement to evaluate productivity impact before wider rollout, unlike anecdotal or purely qualitative feedback.

Why this answer

A/B testing with a control group provides a rigorous comparison of productivity metrics. The other options lack a baseline for comparison.

825
MCQhard

A company is deploying a generative AI model for customer support. They want to reduce hallucinations while maintaining fluency. They have a large dataset of previous support conversations. Which strategy should they prioritize?

A.Increase the beam search width to 10.
B.Implement retrieval-augmented generation (RAG) using the conversation dataset as a knowledge base.
C.Fine-tune the model on the conversation dataset.
D.Set the temperature to 0.1.
AnswerB

Retrieval-augmented generation grounds each response in passages retrieved from the conversation dataset, so answers reflect real support content rather than parametric guesses. This constrains fabrication while the underlying model preserves fluent phrasing, satisfying both the hallucination-reduction and fluency requirements.

Why this answer

Retrieval-augmented generation (RAG) directly addresses hallucinations by grounding the model's responses in factual, retrieved data from the conversation dataset. This approach allows the model to generate fluent, contextually relevant answers while reducing the risk of inventing information, as it retrieves actual support interactions as evidence before generating a response.

Exam trap

Google Cloud often tests the misconception that tuning generation parameters (like temperature or beam search) can fix hallucinations, when in fact only grounding techniques like RAG or knowledge graph integration address the root cause of factual inaccuracy.

How to eliminate wrong answers

Option A is wrong because increasing beam search width to 10 improves output fluency by exploring more candidate sequences but does not reduce hallucinations; it may even amplify incorrect patterns if the model is prone to hallucination. Option C is wrong because fine-tuning on the conversation dataset can improve domain-specific fluency but risks overfitting to noise or biases in the data, and without retrieval, the model may still hallucinate when faced with novel queries. Option D is wrong because setting temperature to 0.1 makes the model more deterministic and less creative, which can reduce variability but does not prevent hallucinations; it may cause the model to repeat common but incorrect patterns from training data.

Page 10

Page 11 of 14

Page 12